Local MCP server that delegates coding tasks to the OpenAI Codex CLI, with explicit model and reasoning-effort selection.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add codex-subagent-mcp -- npx -y codex-subagent-mcp{
"mcpServers": {
"codex-subagent-mcp": {
"command": "npx",
"args": ["-y", "codex-subagent-mcp"]
}
}
}MCP Servers overview
# codex-subagent-mcp [](https://github.com/parisbs/codex-subagent-mcp/actions/workflows/ci.yml) [](https://www.npmjs.com/package/codex-subagent-mcp) [](LICENSE) An MCP server that lets **Claude Code delegate coding tasks to OpenAI's Codex CLI** running on the same machine, with the model and reasoning depth chosen per task. Claude stays the orchestrator. Codex becomes a subagent it can call. An independent project. Not affiliated with, endorsed by, or supported by OpenAI or Anthropic. ## Quick start Requires **Node.js 22+** and the **[Codex CLI](https://developers.openai.com/codex/cli)** installed, on `PATH`, and signed in. ```bash claude mcp add codex-subagent -- npx -y codex-subagent-mcp ``` Ask Claude to run `codex_doctor` if anything is missing. See **[docs/INSTALL.md](docs/INSTALL.md)** for platform setup, Claude Desktop and other installation options, and [Safety](#safety) before enabling writes. ## What you get - You pick the model and the reasoning effort per task, from the catalog your Codex CLI reports live. - Read-only by default, with a sandbox ceiling no tool call can exceed. - Every result states what Codex actually applied (model, effort, sandbox, directory), not only what was asked. - Optional `output_schema` returns JSON that matches your schema, for results Claude can act on directly. - Cancelling, timing out or shutting down stops Codex and every command it started. A real read-only delegation asking for the package name, captured 2026-09-30: ```text codex-subagent-mcp Commands run (1 total): - [exit 0] /bin/zsh -lc "sed -n '1,80p' package.json" model=gpt-5.6-luna | effort=low | sandbox=read-only | working_dir=/Users/you/project | applied=confirmed | duration=8s | tokens=in 35701 (cached 28160, uncached 7541) / out 88 (reasoning 9) | thread_id=01a0f419-9fd2-79e2-8248-91436f04e297 ``` ## Why this exists A single model doing everything has three recurring problems, and delegation solves each one: **Your context window is finite.** Having Claude read forty files to answer one question spends context you need for the actual work. Delegating the investigation returns the answer instead of the forty files. **One model has one set of blind spots.** A second opinion is worth most when it comes from a different model family — different training, different failure modes. Asking the same model twice mostly gets you the same answer twice. **Not every task deserves the same reasoning budget.** Renaming a variable and diagnosing a race condition are not the same job. Here they are separate dials: the *model* sets raw capability, the *reasoning effort* sets how long it deliberates. Cheap work goes to a fast model; a hard problem gets the capable one thinking for as long as it needs. The server runs on your machine and drives the Codex CLI you already have installed; it holds no credentials of its own. Prompts reach OpenAI through Codex, exactly as when you run `codex` yourself. See [How it compares](docs/COMPARISON.md) for a versioned comparison with other Codex MCP servers. ## Using it ### Write a bounded delegation A delegation gets expensive when repeated commands keep adding output to the context carried into later requests. Name the exact question, likely files, stopping condition and evidence the answer must contain; choose higher effort for ambiguity rather than by habit. **[Writing a delegation](docs/DELEGATING.md)** gives the measured cost model, ranked rules and weak-versus-strong examples using the real tool parameters. Background goes in `context` and a persona or extra rules in `system_instructions`, both layered on the built-in quality contract; see the [tool reference](docs/TOOLS.md#codex_delegate). ### Get a second opinion from a different model family The value here is not a second run — it is a different set of blind spots. > Ask Codex to review `src/server.ts` for correctness problems, focusing on error paths. Use a high > reasoning effort and tell it to report each finding with the line and why it matters. A delegated review like this found the `terminate()` defect in this repository's own runner: two code paths could each arm a timer while only one was ever cleared. ### Investigate without spending your context Forty files go into the delegation; one answer comes back. Codex runs its own searches and reads whatever it needs; your conversation receives the conclusion. > Have Codex trace how a reasoning effort travels from the MCP tool call down to the arguments > handed to the Codex CLI, and report just the call chain. ### Run long work in the background while you keep going > Kick off a Codex run in the background that writes unit tests for `src/jobs.ts`, then keep helping > me with the API layer. You get a `job_id` immediately. Ask for the status whenever you want, and read the result when it is done. Up to eight can run at once. ### Buy deep reasoning for one hard problem Raising the reasoning effort for the whole conversation is expensive. Raising it for one delegation is not. > This intermittent test failure has beaten me twice. Ask Codex to work out the root cause at > maximum reasoning effort, give it `test/runner.test.ts` and the CI log, and tell it not to change > anything — I want the diagnosis first. ### Keep the thread going > Ask Codex to expand on its second finding. Follow-ups reuse Codex's context, so they cost a fraction of the original. The server restates the same model, effort and directory on every follow-up because Codex itself does not keep them on resume. ### Let it write, when you mean it > Have Codex apply its first two suggestions. Let it edit files, but keep it inside a git worktree > so my working tree stays clean. `use_worktree` sends the run's edits to `~/.codex/worktrees/`; results list the files it touched and where each landed, up to a thousand distinct files, then report the omitted count. The server does not clean worktrees up: they may hold unapplied work. The experimental feature is enabled only for that invocation; your Codex configuration is unchanged. See [Safety](#safety) for the write boundary. Whatever the sandbox, a delegation that writes reports what it wrote: ```text Files changed (2): - [edit] src/codex/runner.ts - [add] test/runner.test.ts ``` ## Safety This server runs another program on your machine, so it is worth two minutes before you enable writes. ### What protects you Delegations are **read-only by default**. Writing requires an explicit `sandbox: "workspace-write"` or a user-set default. `use_worktree` sends edits to a managed git worktree instead of your checkout; the sandbox and `add_dirs` bound where it can write at all. Unsandboxed runs are unavailable unless you explicitly opt into that ceiling. The confinement is the operating system's own sandbox: Seatbelt on macOS, `bubblewrap` on Linux and WSL2, and a native sandbox on Windows. The table was measured on macOS; [SECURITY.md](SECURITY.md#what-the-sandbox-does-and-does-not-cover) records the measurement details. Linux and Windows have not been measured here, and OpenAI's Windows documentation notes that sandboxed commands can fail to read some directories, so reads may be stricter there: | | `read-only` | `workspace-write` | `danger-full-access` | | --- | --- | --- | --- | | Write inside the working directory | no | yes | yes | | Write outside it (your home) | no | no | yes | | Network access | no | no | yes | | **Read outside the working directory** | **yes** | **yes** | yes | There is no shell in the server's invocation path: the CLI is spawned with an argv array and the prompt is written to its stdin, never interpolated into a command string. Shell metacharacters in a prompt are inert. ### What does not protect you **Reads are not confined.** Codex can read anything your user account can, in every mode — your SSH keys, your cloud credentials. That was measured on macOS, and it is the safe assumption on every platform. Sandboxed network access is blocked so it cannot send them anywhere, but its report comes back to you, and that is a channel. **A prompt is untrusted input, and Codex acts on it.** Content you did not write — an issue body, a web page, a log, a file from someone else's repository — can carry instructions. With `workspace-write` it can direct Codex to modify your repository; even read-only it can direct Codex to read something sensitive and put it in the answer. The sandbox bounds *where* Codex can write. It does not judge *what* it should write, or why it was asked. This is prompt injection, and it is the risk that matters here. **The result is not sanitised.** What comes back is text from a model that just read your files. Treat it as data, not as instructions; review what a delegation did rather than assuming it did what you asked. ### Reducing the risk - Leave the built-in default alone. Read-only handles investigation, review and diagnosis, which is most delegation. - If you never want writes, cap it: `CODEX_SUBAGENT_MAX_SANDBOX=read-only`. No conversation can argue past a ceiling. Register it outside the repository (Claude Code's default `local` scope, `--scope user`, or Claude Desktop's config), not in a project `.mcp.json` that a write-enabled delegation could edit. See [Configuration](docs/TOOLS.md#configuration). - When you enable writes, add `use_worktree` so changes land somewhere you can inspect before they touch your branch. - Do not assemble delegation prompts from untrusted content when you intend to act on the answer. - If this threat matters seriously to you, run Codex under an account or container with no access to your secrets. That solves it at the root instead of bounding it. [SECURITY.md](SECURITY.md) has the full threat model, what a `deny_read` pol
What people ask about codex-subagent-mcp
What is parisbs/codex-subagent-mcp?
+
parisbs/codex-subagent-mcp is mcp servers for the Claude AI ecosystem. Local MCP server that delegates coding tasks to the OpenAI Codex CLI, with explicit model and reasoning-effort selection. It has 0 GitHub stars and its last recorded update is dated 2026-10-01.
How do I install codex-subagent-mcp?
+
You can install codex-subagent-mcp by cloning the repository (https://github.com/parisbs/codex-subagent-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is parisbs/codex-subagent-mcp safe to use?
+
Our security agent has analyzed parisbs/codex-subagent-mcp and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains parisbs/codex-subagent-mcp?
+
parisbs/codex-subagent-mcp is maintained by parisbs. The last recorded GitHub activity is dated 2026-10-01, with 11 open issues.
Are there alternatives to codex-subagent-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy codex-subagent-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/parisbs-codex-subagent-mcp)<a href="https://claudewave.com/repo/parisbs-codex-subagent-mcp"><img src="https://claudewave.com/api/badge/parisbs-codex-subagent-mcp" alt="Featured on ClaudeWave: parisbs/codex-subagent-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.