Windows computer-use MCP server: drive any desktop app via semantic discover-then-act targeting (entities + leases, not pixel coordinates), with per-action perception guards, a native Rust UIA engine, Chrome CDP, and Key Locker credential autofill for ssh/sudo password prompts. Works with Claude, Cursor, and any MCP client.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
- !README contains suspicious pattern: eval\s*\(
claude mcp add desktop-touch-mcp -- npx -y @harusame64/desktop-touch-mcp{
"mcpServers": {
"desktop-touch-mcp": {
"command": "npx",
"args": ["-y", "@harusame64/desktop-touch-mcp"]
}
}
}MCP Servers overview
# desktop-touch-mcp [](https://glama.ai/mcp/servers/Harusame64/desktop-touch-mcp) [日本語](README.ja.md) > **Computer-use MCP server for Windows.** Lets Claude, Cursor, or any MCP client see and operate your Windows 10/11 desktop — screenshots, UI Automation, Chrome CDP, keyboard / mouse, terminal — with **semantic discover-then-act targeting** that avoids pixel-coordinate guessing, and **per-action perception guards** that catch wrong-window typing before it happens. ```bash npx -y @harusame64/desktop-touch-mcp ``` 32 tools, native Rust engine (UIA in 2 ms), zero-config PowerShell fallback, full CJK support, MIT licensed. Add the snippet above to your Claude / Cursor / VS Code Copilot config and Claude can drive Notepad, Excel, Chrome, Windows Terminal, and any other app on your machine. > **Why this over pixel-clicking?** Two ideas run through every tool: **discover-then-act** — `desktop_discover` returns interactive entities with short-lived leases instead of raw coordinates, so `desktop_act` operates on *what* you mean, not *where* it was — and **per-action perception guards** that verify the target window's identity and bounds before input lands, catching wrong-window typing and stale-coordinate clicks before they happen. > > **2.1** reads UI Automation by element count instead of depth: Chrome, Edge and VS Code pages, Explorer and Settings values, and Word's body are read and can be acted on, and what UIA still cannot see falls back to OCR and Set-of-Marks as before. `desktop_act` also types into Word's body, and into Windows Terminal after asking the user. **2.0** refuses an action that cannot be done, with a reason, instead of reporting it done, and an action aimed by `hwnd` reaches that window only. See the [CHANGELOG](CHANGELOG.md). --- ## Features - **🔁 Every act reports what it did** — `desktop_act` answers with what changed: elements that appeared or disappeared, a modal, a focus move, the screen's repaint (`observation`), and with `narrate:"rich"` the values and names that changed. The agent reads the result of its click from the reply instead of taking another screenshot. On UIA-blind targets it can also attach a PNG of just the region that changed (`roiCapture`; on by default for a visible change, `returnCapture:"never"` to suppress, `"always"` to force). - **🛑 An act that cannot be done is refused, not reported as done** — input that would not reach the field, a window blocked by a dialog, a closed window, a window on another virtual desktop: nothing is sent, and the reply names the reason and what to do next. An act aimed by `hwnd` never lands in another window with the same title. - **🌐 Reads deep windows (2.1)** — The UI Automation read goes to depth 64 and 500 elements. Chrome, Edge and VS Code pages, which 2.0 called blind, Explorer and Settings values it left out, and Word's body are read and can be acted on (a Chrome page: 6 elements in 387 ms → 45 in 186 ms). - **🎯 Set-of-Marks (SoM) visual fallback** — Games, RDP sessions and apps with no accessibility tree still return clickable elements: when UIA is blind, `desktop_discover` and `screenshot(detail="text")` switch to a Hybrid Non-CDP pipeline — Rust-powered grayscale + bilinear upscale → Windows OCR → clustering → red bounding boxes with numbered badges (`[1]`, `[2]`…). Two representations come back: a PNG for spatial orientation and an `elements[]` list with `clickAt` coords — no CDP required. - **⌨️ Types where background input does not reach** — Windows Terminal, after asking the user each time (needs an MCP client with elicitation, over stdio), and Word's body, at Word's caret whether it is in front or behind. - **🔐 Key Locker — the terminal autofills your SSH / sudo passwords** — Save a credential once into the locker's own secure dialog (stored encrypted on your machine with Windows DPAPI; never shown to the assistant), then run `ssh` / `sudo` in a console opened by `key_locker(action='launch_console')` — the password is filled in automatically when the hidden prompt appears, with a per-fill confirmation prompt by default. See [Key Locker](docs/guide.md#key-locker-terminal-credential-autofill). - **⚡ Rust native core** — The UIA bridge and image diffing are a Rust native addon (`napi-rs` + `windows-rs`): UIA is called over COM from a dedicated thread instead of spawning PowerShell, and image diffs use SSE2 SIMD. Without the addon, every function falls back to PowerShell transparently. The npm launcher fetches only the GitHub Release matching its version and verifies the Windows runtime zip before extracting it. - **LLM-native design** — Built around how LLMs think, not how humans click. `run_macro` batches multiple operations into a single API call; `diffMode` sends only the windows that changed since the last frame. Minimal tokens, minimal round-trips. - **Reactive Perception Graph** — Register a `lensId` for a window or browser tab, pass it to action tools, and get guard-checked `post.perception` feedback after each action. It reduces repeated `screenshot` / `desktop_state` calls and prevents wrong-window typing or stale-coordinate clicks. - **Full CJK support** — Uses Win32 `GetWindowTextW` for window titles, avoiding nut-js garbling. IME bypass input supported for Japanese/Chinese/Korean environments. - **3-tier token reduction** — `detail="image"` (~443 tok) / `detail="text"` (~100–300 tok) / `diffMode=true` (~160 tok). Send pixels only when you actually need to see them. - **1:1 coordinate mode** — `dotByDot=true` captures at native resolution (WebP). Image pixel = screen coordinate — no scale math needed. With `origin`+`scale` passed to `mouse_click`, the server converts coords for you — eliminating off-by-one / scale bugs. - **Browser capture data reduction** — `grayscale=true` (~50% size), `dotByDotMaxDimension=1280` (auto-scaled with coord preservation), and `windowTitle + region` sub-crops help exclude browser chrome and other irrelevant pixels. Typical reduction for heavy captures: 50–70%. - **Chromium smart fallback** — `detail="text"` on Chrome/Edge/Brave auto-skips UIA (prohibitively slow there) and runs Windows OCR. `hints.chromiumGuard` + `hints.ocrFallbackFired` flag the path taken. - **UIA element extraction** — `detail="text"` returns button names and `clickAt` coords as JSON. Claude can click the right element without ever looking at a screenshot. - **Auto-dock CLI** — `window_dock(action='dock')` snaps any window to a screen corner with always-on-top. Set `DESKTOP_TOUCH_DOCK_TITLE='@parent'` to auto-dock the terminal hosting Claude on MCP startup — the process-tree walker finds the right window regardless of title. - **Emergency stop (Failsafe)** — Park the mouse in the **top-left corner of the primary monitor** (within 10px of 0,0) for 500ms to trigger the emergency stop. --- ## Requirements | | | |---|---| | OS | Windows 10 / 11 (64-bit) | | Node.js | v20+ recommended (tested on v22+) — **to develop or run the test suite, `^22.12 || ^24 || >=26`** — the test runner's own range since #658, which excludes odd majors such as 23 and 25 | | PowerShell | 5.1+ (bundled with Windows) — used only as fallback when the Rust native engine is unavailable | | Claude CLI | `claude` command must be available | > **Note:** nut-js native bindings require the Visual C++ Redistributable. > Download from [Microsoft](https://learn.microsoft.com/en-us/cpp/windows/latest-supported-vc-redist) if not already installed. > **Note (Key Locker):** The credential helper Key Locker uses is an unsigned executable, so on > some machines Windows SmartScreen or antivirus may show an "unknown publisher" warning the > first time it runs. This is expected — the helper ships with desktop-touch-mcp and runs locally > on your machine; you can allow it to proceed. Code signing is planned for a future release. --- ## Installation ```bash npx -y @harusame64/desktop-touch-mcp ``` The npm launcher resolves runtime strictly by npm package version. For package `X.Y.Z`, it fetches only GitHub Release tag `vX.Y.Z`, downloads `desktop-touch-mcp-windows.zip`, verifies its SHA256 digest, and only then expands it under `%USERPROFILE%\.desktop-touch-mcp`. Verified cached releases are reused on later runs. Set `DESKTOP_TOUCH_MCP_HOME` to override the cache root directory. > **On a shared or CI network?** The first run reads the GitHub Releases API to > locate the runtime zip. The anonymous limit is 60 requests/hour per IP, which a > shared public address (CI runners, office NAT) can exhaust before your download > even starts. Set `GITHUB_TOKEN` (or `GH_TOKEN`) in the environment and the > launcher authenticates the request, raising the limit to 5,000 requests/hour. > No token is needed on an ordinary home connection. > **Running the launcher from a source checkout?** A source build's > `bin/launcher.js` carries a placeholder integrity hash (`sha256: "PENDING"`) > instead of a finalized one. Rather than download and run an unverified runtime, > the launcher fails closed — this guard stops an accidentally published or > unfinalized launcher from silently starting unverified code. Published npm > releases always ship a real SHA256, so end users never see this. If you are > intentionally running the launcher from source, set > `DESKTOP_TOUCH_MCP_ALLOW_UNVERIFIED=1` to skip integrity verification > (development only). > **Does your host give up before the launcher finishes?** Some desktop hosts > allow a plugin a fixed budget — 60 seconds is common — to become ready, and a > launcher waiting on an unreachable GitHub can spend all of it. Two environment > variables cover that case. > > `DESKTOP_TOUCH_MCP_FETCH_TIMEOUT_MS` (default `15000`) bounds how long the > launcher waits without hearing from GitHub. It applies to the release lookup > and to the download; for the download it counts silence rather than total > time, so a large runtime
What people ask about desktop-touch-mcp
What is Harusame64/desktop-touch-mcp?
+
Harusame64/desktop-touch-mcp is mcp servers for the Claude AI ecosystem. Windows computer-use MCP server: drive any desktop app via semantic discover-then-act targeting (entities + leases, not pixel coordinates), with per-action perception guards, a native Rust UIA engine, Chrome CDP, and Key Locker credential autofill for ssh/sudo password prompts. Works with Claude, Cursor, and any MCP client. It has 21 GitHub stars and its last recorded update is dated 2026-10-04.
How do I install desktop-touch-mcp?
+
You can install desktop-touch-mcp by cloning the repository (https://github.com/Harusame64/desktop-touch-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is Harusame64/desktop-touch-mcp safe to use?
+
Our security agent has analyzed Harusame64/desktop-touch-mcp and assigned a Trust Score of 85/100 (tier: Trusted). See the full breakdown of passed checks and flags on this page.
Who maintains Harusame64/desktop-touch-mcp?
+
Harusame64/desktop-touch-mcp is maintained by Harusame64. The last recorded GitHub activity is dated 2026-10-04, with 2 open issues.
Are there alternatives to desktop-touch-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy desktop-touch-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/harusame64-desktop-touch-mcp)<a href="https://claudewave.com/repo/harusame64-desktop-touch-mcp"><img src="https://claudewave.com/api/badge/harusame64-desktop-touch-mcp" alt="Featured on ClaudeWave: Harusame64/desktop-touch-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.