MCP server that lets Claude operate your REAL desktop — moves the actual mouse/keyboard and reads the real screen, so it works in your own logged-in Chrome and any app.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add computer-use-mcp -- uvx realhands{
"mcpServers": {
"computer-use-mcp": {
"command": "uvx",
"args": ["realhands"]
}
}
}MCP Servers overview
<!-- mcp-name: io.github.kanishka089/realhands -->
# computer-use-mcp ("realhands")
An MCP server that lets **Claude operate your real computer** the way a human does —
moving the **actual mouse**, clicking, typing, and reading the **actual screen**.
Unlike OpenAI Operator, browser-use, or Playwright agents (which spin up a separate,
isolated, logged-out Chrome), this drives the **physical OS cursor and keyboard**. So it
works in **your own Chrome with your own logged-in sessions** — and in every other app —
because it's just a human at the keyboard, as far as any website can tell.
**Status: LIVE and battle-tested.** Registered with Claude Code as the user-scope MCP
**`realhands`** (tool `mcp__realhands__computer`) and ✓Connected since 2026-06-02.
On 2026-06-09 it drove the user's real, logged-in Chrome through a **complete Google
Play Console deployment** (app upload, release notes, submission) end-to-end.
> The server is registered as `realhands` rather than `computer-use` because the name
> "computer-use" is reserved in Claude Code.
## How it works
Claude (Desktop or Code) is the agent loop. You type a task; Claude calls the single
`computer` tool in a see → think → act cycle:
> **See** — `screenshot` returns the real screen (downscaled to ~1280px for grounding accuracy)
> → **Think** — Claude picks the next action + pixel coordinates
> → **Act** — the server moves the real mouse / types on the real keyboard
> → a fresh screenshot comes back automatically after every action, and it repeats.
Two Windows-specific details make clicks land accurately (`src/screen.py`):
- **DPI awareness** — `SetProcessDpiAwareness(2)` is set at import time so screenshot
pixels == pyautogui cursor coordinates even under display scaling (125% / 150% / …).
- **Stateless coordinate scaling** — screenshots are downscaled (LANCZOS) to at most
`COMPUTER_USE_MAX_DIM` on the longest side before sending; incoming click coordinates
are scaled back up to real pixels. The scale factor is a pure function of monitor
geometry + `MAX_DIM`, so mapping never depends on which screenshot ran last.
Coordinates are clamped inside the target monitor so a stray click can't fly off-screen.
Each axis is rounded down to a whole 28px vision patch, so the two axes can scale very
slightly differently (<2.2%); `to_real()` maps each with its own factor and stays exact.
**Multi-monitor:** every call takes an optional `monitor` index (1 = primary, 2.. =
others, 0 = the whole virtual desktop). `action="monitors"` enumerates the setup.
Origins may be negative for screens left/above the primary — `to_real()` handles the
offset. Use the **same** monitor for a click as for the screenshot you're clicking on.
## Architecture
```
src/realhands/
server.py FastMCP server "computer-use"; the single `computer` tool (action enum
modeled on Anthropic's reference computer_20250124 tool); runs one
action or a batch of `steps`, then returns status text + at most one
screenshot
screen.py DPI awareness, mss capture, patch-aligned downscale, model-space ->
real-pixel mapping, unchanged-screen detection
input.py pyautogui mouse/keyboard execution; xdotool-style key-name translation
(Return, Page_Down, ctrl+a, super, ...); clipboard-paste fast path for
long/Unicode/multiline typing (preserves your existing clipboard);
activate_window via win32 AttachThreadInput
safety.py kill switches + lazy arm / stand-down lifecycle
config.py .env-driven configuration (all defaults are sensible; .env is optional)
install.py one-shot installer: venv, deps, .env, Claude Desktop registration
```
Stack: Python 3.10/3.11 · `mcp` (FastMCP, stdio) · `pyautogui` · `mss` · `pillow` ·
`pynput` · `keyboard` · `pyperclip` · `python-dotenv` — plus `pygetwindow` and `pywin32`
for `activate_window`.
## The `computer` tool
A single tool with an `action` parameter:
| Action | What it does |
|---|---|
| `screenshot` | Capture the screen (always start a task with this) |
| `cursor_position` | Report the real mouse position |
| `monitors` | List detected monitors (for multi-screen setups) |
| `mouse_move` | Glide the cursor to `coordinate` |
| `left_click` / `right_click` / `middle_click` / `double_click` / `triple_click` | Click at `coordinate` (or current position) |
| `left_click_drag` | Drag from `text="x1,y1"` to `coordinate=[x2,y2]` |
| `left_mouse_down` / `left_mouse_up` | Press / release the left button |
| `scroll` | Scroll at `coordinate` (`scroll_direction` + `scroll_amount` notches) |
| `type` | Type `text` (clipboard-paste path for long/Unicode/multiline) |
| `key` | Press a key or chord — `"Return"`, `"ctrl+s"`, `"alt+Tab"` |
| `hold_key` | Hold keys for `duration` seconds |
| `activate_window` | Bring an app to the front by title substring (beats Windows' foreground-lock; far more reliable than clicking the taskbar) |
| `wait` | Sleep `duration` seconds, then screenshot |
| `stop` | Stand down: close the STOP overlay + release the panic hotkey (call as the final action) |
Coordinates are in the pixel space of the most recent screenshot; its size is reported
with every capture. After every non-screenshot action the tool waits ~0.4s for the UI
to settle and returns a fresh screenshot.
### Token economy
Screenshots dominate the cost of driving a desktop, and not just once: every image
stays in the conversation and is re-sent as history on each later turn. Claude bills
vision in 28×28 patches — `tokens = ⌈w/28⌉ × ⌈h/28⌉` — so this server does four things
to keep the bill down.
**Batch steps.** Pass `steps` (a list of action dicts) instead of one call per action
and the whole run shares **one** screenshot at the end:
```jsonc
{"steps": [{"action": "left_click", "coordinate": [420, 300]},
{"action": "type", "text": "hello@example.com"},
{"action": "key", "text": "Tab"},
{"action": "type", "text": "secret"},
{"action": "key", "text": "Return"}]}
```
That is 1 125 visual tokens instead of 5 625, and one round trip instead of five — the
larger saving, since each avoided turn also avoids re-sending the entire transcript.
A failing step stops the run, reports which step failed, and still returns the screen.
Add `"screenshot": false` to skip the trailing image too.
**Patch-aligned downscaling.** A dimension that isn't a multiple of 28 pays for a
partial patch row/column carrying almost no pixels. From a 1920×1080 primary the
default 1260×700 is exactly 45×25 patches = 1 125 tokens, versus 1 196 for 1280×720 —
6% off for 1.5% fewer pixels.
**Unchanged-screen suppression.** If under `CHANGE_THRESHOLD` of pixels moved since the
last image sent, the reply is a line of text instead of a screenshot. A real desktop
never produces two byte-identical frames (clock, caret, hover states), so this is a
threshold, not an equality check. After `MAX_SKIPS` suppressions in a row it force-sends
one, so the model can't fly blind if it lost the earlier image to context compaction.
**Right-sized images.** `COMPUTER_USE_MAX_DIM` trades grounding accuracy against cost —
from a 1920×1080 primary: 1792 → 2 304 tokens, 1260 → 1 125 (default), 1036 → 777,
896 → 576. Keep the long edge ≤ 2576 px: an image returned inside a `tool_result` is
*rejected* rather than downscaled when it exceeds the model's limit.
> `COMPUTER_USE_IMAGE_FORMAT` is **not** a token lever — Claude bills by pixel
> dimensions, so a JPEG and a PNG of the same screenshot cost exactly the same. JPEG
> only cuts payload bytes (~977 KB → ~141 KB here), which helps latency at some risk
> to small-text legibility.
## Safety — it controls your REAL machine
This is **fully autonomous**: it does not ask before each action. Three independent
kill switches (`src/safety.py`):
1. **Fail-safe corner** — slam the mouse into the **top-left corner** → pyautogui raises
`FailSafeException` and the action aborts instantly.
2. **Panic hotkey** — **Ctrl+Alt+Q** (configurable) → hard-kills the server process
(`os._exit(1)`).
3. **STOP overlay** — an always-on-top window (top-right) showing the current action,
with a big red **■ STOP AGENT** button that also hard-kills the process.
**Lazy arm / stand-down:** the overlay and the global panic hotkey are armed lazily on
the **first action** of a task, not at server startup — idle sessions show nothing and
grab no hotkeys. They stand down when the agent calls `action="stop"` at the end of a
task, and re-arm automatically on the next action. (The STOP overlay is a single
persistent window that is *hidden* when dormant, never destroyed — recreating it was a
crash hazard.) An optional idle auto-stand-down is available via
`COMPUTER_USE_IDLE_STOP` but is **disabled by default**: an agent's thinking time
between tool calls easily exceeds any short idle window, so a non-zero value would stand
the agent down mid-task.
Pacing also helps you stay in control: every action is followed by a configurable pause
(`COMPUTER_USE_PAUSE`) and the cursor glides rather than teleports
(`COMPUTER_USE_MOVE_DURATION`), so you can watch and interrupt.
**Don't leave it unsupervised on anything that can spend money, send messages, or
delete data.**
## Install
Requires **Python 3.10 or 3.11** (3.13+ untested; avoid the 3.14 beta).
### Recommended: `uvx` (always the latest version)
Nothing to install up front, and **you get every release automatically** — `uvx`
resolves the newest published version each time the server starts, so a restart of your
MCP client is the whole upgrade process. Requires [uv](https://docs.astral.sh/uv/).
```powershell
uvx realhands@latest
```
Drop the `@latest` (`uvx realhands`) if you would rather let uv reuse whatever version
it already has cached.
### From PyPI (pinned)
```powershell
pip install realhands
```
This installs the `realhands` console script and the importable `realhands`
package. Run the server with eitWhat people ask about computer-use-mcp
What is kanishka089/computer-use-mcp?
+
kanishka089/computer-use-mcp is mcp servers for the Claude AI ecosystem. MCP server that lets Claude operate your REAL desktop — moves the actual mouse/keyboard and reads the real screen, so it works in your own logged-in Chrome and any app. It has 0 GitHub stars and its last recorded update is dated 2026-08-25.
How do I install computer-use-mcp?
+
You can install computer-use-mcp by cloning the repository (https://github.com/kanishka089/computer-use-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is kanishka089/computer-use-mcp safe to use?
+
Our security agent has analyzed kanishka089/computer-use-mcp and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains kanishka089/computer-use-mcp?
+
kanishka089/computer-use-mcp is maintained by kanishka089. The last recorded GitHub activity is dated 2026-08-25, with 0 open issues.
Are there alternatives to computer-use-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy computer-use-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/kanishka089-computer-use-mcp)<a href="https://claudewave.com/repo/kanishka089-computer-use-mcp"><img src="https://claudewave.com/api/badge/kanishka089-computer-use-mcp" alt="Featured on ClaudeWave: kanishka089/computer-use-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
The fastest path to AI-powered full stack observability, even for lean teams.
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!