Skip to main content
ClaudeWave
kanishka089 avatar
kanishka089

computer-use-mcp

Ver en GitHub

MCP server that lets Claude operate your REAL desktop — moves the actual mouse/keyboard and reads the real screen, so it works in your own logged-in Chrome and any app.

MCP ServersRegistry oficial0 estrellas0 forksPythonMITActualizado today
ClaudeWave Trust Score
95/100
Verified
Passed
  • Open-source license (MIT)
  • Actively maintained (<30d)
  • Clear description
  • Topics declared
  • Documented (README)
Last scanned: 8/26/2026
Install in Claude Code / Claude Desktop
Method: UVX (Python) · realhands
Claude Code CLI
claude mcp add computer-use-mcp -- uvx realhands
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "computer-use-mcp": {
      "command": "uvx",
      "args": ["realhands"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
Casos de uso

Resumen de MCP Servers

<!-- mcp-name: io.github.kanishka089/realhands -->

# computer-use-mcp ("realhands")

An MCP server that lets **Claude operate your real computer** the way a human does —
moving the **actual mouse**, clicking, typing, and reading the **actual screen**.

Unlike OpenAI Operator, browser-use, or Playwright agents (which spin up a separate,
isolated, logged-out Chrome), this drives the **physical OS cursor and keyboard**. So it
works in **your own Chrome with your own logged-in sessions** — and in every other app —
because it's just a human at the keyboard, as far as any website can tell.

**Status: LIVE and battle-tested.** Registered with Claude Code as the user-scope MCP
**`realhands`** (tool `mcp__realhands__computer`) and ✓Connected since 2026-06-02.
On 2026-06-09 it drove the user's real, logged-in Chrome through a **complete Google
Play Console deployment** (app upload, release notes, submission) end-to-end.

> The server is registered as `realhands` rather than `computer-use` because the name
> "computer-use" is reserved in Claude Code.

## How it works

Claude (Desktop or Code) is the agent loop. You type a task; Claude calls the single
`computer` tool in a see → think → act cycle:

> **See** — `screenshot` returns the real screen (downscaled to ~1280px for grounding accuracy)
> → **Think** — Claude picks the next action + pixel coordinates
> → **Act** — the server moves the real mouse / types on the real keyboard
> → a fresh screenshot comes back automatically after every action, and it repeats.

Two Windows-specific details make clicks land accurately (`src/screen.py`):

- **DPI awareness** — `SetProcessDpiAwareness(2)` is set at import time so screenshot
  pixels == pyautogui cursor coordinates even under display scaling (125% / 150% / …).
- **Stateless coordinate scaling** — screenshots are downscaled (LANCZOS) to at most
  `COMPUTER_USE_MAX_DIM` on the longest side before sending; incoming click coordinates
  are scaled back up to real pixels. The scale factor is a pure function of monitor
  geometry + `MAX_DIM`, so mapping never depends on which screenshot ran last.
  Coordinates are clamped inside the target monitor so a stray click can't fly off-screen.
  Each axis is rounded down to a whole 28px vision patch, so the two axes can scale very
  slightly differently (<2.2%); `to_real()` maps each with its own factor and stays exact.

**Multi-monitor:** every call takes an optional `monitor` index (1 = primary, 2.. =
others, 0 = the whole virtual desktop). `action="monitors"` enumerates the setup.
Origins may be negative for screens left/above the primary — `to_real()` handles the
offset. Use the **same** monitor for a click as for the screenshot you're clicking on.

## Architecture

```
src/realhands/
  server.py   FastMCP server "computer-use"; the single `computer` tool (action enum
              modeled on Anthropic's reference computer_20250124 tool); runs one
              action or a batch of `steps`, then returns status text + at most one
              screenshot
  screen.py   DPI awareness, mss capture, patch-aligned downscale, model-space ->
              real-pixel mapping, unchanged-screen detection
  input.py    pyautogui mouse/keyboard execution; xdotool-style key-name translation
              (Return, Page_Down, ctrl+a, super, ...); clipboard-paste fast path for
              long/Unicode/multiline typing (preserves your existing clipboard);
              activate_window via win32 AttachThreadInput
  safety.py   kill switches + lazy arm / stand-down lifecycle
  config.py   .env-driven configuration (all defaults are sensible; .env is optional)
install.py    one-shot installer: venv, deps, .env, Claude Desktop registration
```

Stack: Python 3.10/3.11 · `mcp` (FastMCP, stdio) · `pyautogui` · `mss` · `pillow` ·
`pynput` · `keyboard` · `pyperclip` · `python-dotenv` — plus `pygetwindow` and `pywin32`
for `activate_window`.

## The `computer` tool

A single tool with an `action` parameter:

| Action | What it does |
|---|---|
| `screenshot` | Capture the screen (always start a task with this) |
| `cursor_position` | Report the real mouse position |
| `monitors` | List detected monitors (for multi-screen setups) |
| `mouse_move` | Glide the cursor to `coordinate` |
| `left_click` / `right_click` / `middle_click` / `double_click` / `triple_click` | Click at `coordinate` (or current position) |
| `left_click_drag` | Drag from `text="x1,y1"` to `coordinate=[x2,y2]` |
| `left_mouse_down` / `left_mouse_up` | Press / release the left button |
| `scroll` | Scroll at `coordinate` (`scroll_direction` + `scroll_amount` notches) |
| `type` | Type `text` (clipboard-paste path for long/Unicode/multiline) |
| `key` | Press a key or chord — `"Return"`, `"ctrl+s"`, `"alt+Tab"` |
| `hold_key` | Hold keys for `duration` seconds |
| `activate_window` | Bring an app to the front by title substring (beats Windows' foreground-lock; far more reliable than clicking the taskbar) |
| `wait` | Sleep `duration` seconds, then screenshot |
| `stop` | Stand down: close the STOP overlay + release the panic hotkey (call as the final action) |

Coordinates are in the pixel space of the most recent screenshot; its size is reported
with every capture. After every non-screenshot action the tool waits ~0.4s for the UI
to settle and returns a fresh screenshot.

### Token economy

Screenshots dominate the cost of driving a desktop, and not just once: every image
stays in the conversation and is re-sent as history on each later turn. Claude bills
vision in 28×28 patches — `tokens = ⌈w/28⌉ × ⌈h/28⌉` — so this server does four things
to keep the bill down.

**Batch steps.** Pass `steps` (a list of action dicts) instead of one call per action
and the whole run shares **one** screenshot at the end:

```jsonc
{"steps": [{"action": "left_click", "coordinate": [420, 300]},
           {"action": "type",  "text": "hello@example.com"},
           {"action": "key",   "text": "Tab"},
           {"action": "type",  "text": "secret"},
           {"action": "key",   "text": "Return"}]}
```

That is 1 125 visual tokens instead of 5 625, and one round trip instead of five — the
larger saving, since each avoided turn also avoids re-sending the entire transcript.
A failing step stops the run, reports which step failed, and still returns the screen.
Add `"screenshot": false` to skip the trailing image too.

**Patch-aligned downscaling.** A dimension that isn't a multiple of 28 pays for a
partial patch row/column carrying almost no pixels. From a 1920×1080 primary the
default 1260×700 is exactly 45×25 patches = 1 125 tokens, versus 1 196 for 1280×720 —
6% off for 1.5% fewer pixels.

**Unchanged-screen suppression.** If under `CHANGE_THRESHOLD` of pixels moved since the
last image sent, the reply is a line of text instead of a screenshot. A real desktop
never produces two byte-identical frames (clock, caret, hover states), so this is a
threshold, not an equality check. After `MAX_SKIPS` suppressions in a row it force-sends
one, so the model can't fly blind if it lost the earlier image to context compaction.

**Right-sized images.** `COMPUTER_USE_MAX_DIM` trades grounding accuracy against cost —
from a 1920×1080 primary: 1792 → 2 304 tokens, 1260 → 1 125 (default), 1036 → 777,
896 → 576. Keep the long edge ≤ 2576 px: an image returned inside a `tool_result` is
*rejected* rather than downscaled when it exceeds the model's limit.

> `COMPUTER_USE_IMAGE_FORMAT` is **not** a token lever — Claude bills by pixel
> dimensions, so a JPEG and a PNG of the same screenshot cost exactly the same. JPEG
> only cuts payload bytes (~977 KB → ~141 KB here), which helps latency at some risk
> to small-text legibility.

## Safety — it controls your REAL machine

This is **fully autonomous**: it does not ask before each action. Three independent
kill switches (`src/safety.py`):

1. **Fail-safe corner** — slam the mouse into the **top-left corner** → pyautogui raises
   `FailSafeException` and the action aborts instantly.
2. **Panic hotkey** — **Ctrl+Alt+Q** (configurable) → hard-kills the server process
   (`os._exit(1)`).
3. **STOP overlay** — an always-on-top window (top-right) showing the current action,
   with a big red **■ STOP AGENT** button that also hard-kills the process.

**Lazy arm / stand-down:** the overlay and the global panic hotkey are armed lazily on
the **first action** of a task, not at server startup — idle sessions show nothing and
grab no hotkeys. They stand down when the agent calls `action="stop"` at the end of a
task, and re-arm automatically on the next action. (The STOP overlay is a single
persistent window that is *hidden* when dormant, never destroyed — recreating it was a
crash hazard.) An optional idle auto-stand-down is available via
`COMPUTER_USE_IDLE_STOP` but is **disabled by default**: an agent's thinking time
between tool calls easily exceeds any short idle window, so a non-zero value would stand
the agent down mid-task.

Pacing also helps you stay in control: every action is followed by a configurable pause
(`COMPUTER_USE_PAUSE`) and the cursor glides rather than teleports
(`COMPUTER_USE_MOVE_DURATION`), so you can watch and interrupt.

**Don't leave it unsupervised on anything that can spend money, send messages, or
delete data.**

## Install

Requires **Python 3.10 or 3.11** (3.13+ untested; avoid the 3.14 beta).

### Recommended: `uvx` (always the latest version)

Nothing to install up front, and **you get every release automatically** — `uvx`
resolves the newest published version each time the server starts, so a restart of your
MCP client is the whole upgrade process. Requires [uv](https://docs.astral.sh/uv/).

```powershell
uvx realhands@latest
```

Drop the `@latest` (`uvx realhands`) if you would rather let uv reuse whatever version
it already has cached.

### From PyPI (pinned)

```powershell
pip install realhands
```

This installs the `realhands` console script and the importable `realhands`
package. Run the server with eit
anthropicautomationclaudecomputer-usedesktop-automationmcpmodel-context-protocolrpa

Lo que la gente pregunta sobre computer-use-mcp

¿Qué es kanishka089/computer-use-mcp?

+

kanishka089/computer-use-mcp es mcp servers para el ecosistema de Claude AI. MCP server that lets Claude operate your REAL desktop — moves the actual mouse/keyboard and reads the real screen, so it works in your own logged-in Chrome and any app. Tiene 0 estrellas en GitHub y su última actualización registrada es del 2026-08-25.

¿Cómo se instala computer-use-mcp?

+

Puedes instalar computer-use-mcp clonando el repositorio (https://github.com/kanishka089/computer-use-mcp) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.

¿Es seguro usar kanishka089/computer-use-mcp?

+

Nuestro agente de seguridad ha analizado kanishka089/computer-use-mcp y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.

¿Quién mantiene kanishka089/computer-use-mcp?

+

kanishka089/computer-use-mcp es mantenido por kanishka089. La última actividad registrada en GitHub es del 2026-08-25, con 0 issues abiertos.

¿Hay alternativas a computer-use-mcp?

+

Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.

Despliega computer-use-mcp en tu cloud

Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.

¿Mantienes este repo? Añade un badge a tu README

Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.

Featured on ClaudeWave: kanishka089/computer-use-mcp
[![Featured on ClaudeWave](https://claudewave.com/api/badge/kanishka089-computer-use-mcp)](https://claudewave.com/repo/kanishka089-computer-use-mcp)
<a href="https://claudewave.com/repo/kanishka089-computer-use-mcp"><img src="https://claudewave.com/api/badge/kanishka089-computer-use-mcp" alt="Featured on ClaudeWave: kanishka089/computer-use-mcp" width="320" height="64" /></a>

Más MCP Servers

Alternativas a computer-use-mcp