Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code.
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
- !Licence file present but not machine-readable
git clone https://github.com/JoblessJoe/local-llm-worker{
"mcpServers": {
"local-llm-worker": {
"command": "node",
"args": ["/path/to/local-llm-worker/dist/index.js"]
}
}
}Resumen de MCP Servers
<div align="center">
# local-llm-worker
**Let Claude hand the bulk reading to your local LLM: test logs, big files, web pages. Claude only gets the answer.**
[](https://github.com/JoblessJoe/local-llm-worker/actions/workflows/test.yml)
[](LICENSE)
[](package.json)
[](package.json)
[](#install)
[](#use-it-in-any-mcp-client)
Works with Ollama · llama.cpp · LM Studio · vLLM · LocalAI · any OpenAI-compatible endpoint
</div>
---
Reading a 5,000-line test log or three docs pages costs the same frontier-model tokens as hard
architectural work. **local-llm-worker** is an MCP server and Claude Code plugin that moves
that bulk work onto the GPU (or CPU) you already own:
- **`offload`**: your local model runs the noisy command or reads the big file. Claude gets
the answer.
- **`research`**: your local model searches the web, reads the pages in full and returns a
cited answer. Claude never sees the pages.
- **`delegate`** *(opt-in)*: your local model writes one file in an isolated git worktree
and retries until **a test Claude wrote first** passes. Claude gets a one-line verdict, not
the code.
```text
> offload npm test 2>&1 "Which test fails, and why?"
`fixed coupon never goes below zero` (test/cart.test.js:404): applyCoupon(5, {type:'fixed',
value:10}) returns -5, expected 0.
— qwen3-coder:30b-a3b-q4_K_M · 19,132 input tokens read locally
```
## Why
| | Claude does it | Claude delegates it |
|---|---|---|
| Find one failure in a 50 KB test log | ~19,000 tokens of log in context | a ~130-token answer |
| Answer from three docs pages | the pages, or a summary of them | a cited paragraph + source list, from the full pages |
| Write a helper that tests can pin down | output tokens for the code, then re-reading it | writes the test, reads `PASS` |
| Several independent helpers | sequential edits | parallel delegates, each in its own worktree |
- **Any model, any hardware.** No hardcoded models, no GPU assumptions. CPU-only works, just
slower.
- **Parallel by default.** Several calls in one turn, or from several subagents, run at once.
The concurrency limit is a config value, not a hardcoded lock.
- **Fits the context automatically.** Material is sized to the model's context window per
text (number-heavy logs need more tokens than prose). Anything cut is reported.
- **Zero dependencies.** Plain Node ≥ 18. Installing from git needs no `npm install`.
- **Agent-configurable.** One `configure` call shows the config, the backend and its models,
and changes any setting. It takes effect on the next call, with no restart.
## Install
**Prerequisites:** Node ≥ 18, git, and a running local model server (e.g.
`ollama pull qwen3-coder:30b`).
### Claude Code plugin
```text
/plugin marketplace add JoblessJoe/local-llm-worker
/plugin install local-llm-worker@local-llm-worker
```
Then just ask: *"set up local-llm-worker for my machine"*. Claude finds your backend, picks a
model, and asks which tools it should use on its own (`auto_use`). By default that's `offload`
and `research`; `delegate` is opt-in. Every tool also works whenever you ask for it.
### Make Claude use it every time
A session-start reminder nudges Claude toward the tools you chose. For dependable use, add
this to your `CLAUDE.md`; setup offers to do it for you:
```markdown
## Local LLM (local-llm-worker)
- Run test suites, builds and other noisy commands through `offload` (`command`), and read logs or files over ~300 lines through it, instead of reading the output yourself.
- Use `research` for web lookups instead of WebSearch/WebFetch.
```
With this in place, Claude ran a noisy failing test suite through `offload` every time we
tried. That session cost 40 % less than one that read the output itself.
### Use it in any MCP client
Claude Desktop, Cursor, Windsurf, and others:
```json
{
"mcpServers": {
"local-llm-worker": {
"command": "node",
"args": ["/path/to/local-llm-worker/src/index.js"],
"env": { "LLW_MODEL": "qwen3-coder:30b" }
}
}
}
```
Claude Code without the plugin:
`claude mcp add local-llm-worker -- node /path/to/local-llm-worker/src/index.js`
## How `delegate` works
```mermaid
flowchart LR
A[Claude writes the test<br/>+ spec + invariants] --> B[delegate]
B --> C[fresh git worktree<br/>+ your uncommitted changes]
C --> D[local LLM writes<br/>target file]
D --> E{run test}
E -- fail --> F[parse failures<br/>test source stays hidden]
F --> D
E -- pass --> G[copy file into checkout<br/>unless it changed meanwhile]
G --> H[Claude gets a one-line verdict]
```
- **Isolated.** Each call gets its own `git worktree` with your uncommitted changes mirrored in,
plus symlinked `node_modules` / `.venv`, so parallel calls never collide.
- **Locked scope.** Exactly one target file, jailed to the repo, and never one of the
`test_files` you name.
- **Useful feedback.** Retries get the parsed failing assertions (TAP, pytest, jest, vitest,
go, cargo), not a stack-trace tail.
- **Safe apply.** If you or another delegate touched the target meanwhile, nothing is
overwritten.
- **Honest failure.** After N attempts you get the last failure and the kept worktree. Transport
errors are reported as errors, never as a model FAIL.
- **Self-cleaning.** Each delegate prunes `llw-*` worktrees left behind by a crashed server, and
kept ones older than 24 h. Untracked files over 10 MB are not copied into the worktree (the
verdict says how many were skipped).
The bundled [skill](skills/local-llm-worker/SKILL.md) teaches Claude how to write specs that
pass: one invariant per test, exact API surface in `context`, properties instead of examples, and
never delegating auth or money code.
## Results
On one 24 GB GPU (Tesla P40) with `qwen3-coder:30b-a3b-q4_K_M` on Ollama:
| Task | Read locally | Claude received |
|---|---|---|
| `offload`: find the failing test in a 50 KB test log | ~19,100 tokens | ~130 tokens |
| `offload`: find one ERROR line in a 90 KB log | ~30,400 tokens | the line, quoted |
| `research`: "Latest Node.js LTS and its end of life?" (3 pages) | 7,025 tokens | a cited answer, ~250 tokens |
Several calls run in parallel: three in one turn took as long as the slowest one.
### Which model?
From a reproducible benchmark ([`bench/`](bench/)) of 120 `delegate` runs across six task
types:
| Model | Good for | Speed per call |
|---|---|---|
| `qwen3-coder:30b-a3b` | the best default: parsers, pure functions, Python | 15–55 s |
| `devstral-small-2:24b` | edits to existing files, parsers | 2–9 min |
| `granite4.1:8b` | small, well-specified functions on modest hardware | 25–90 s |
Retries matter: with failure feedback, qwen3-coder's pass rate rose by half from the first
attempt to the third. `offload` and `research` work well with any of these. Every run is logged
(see [Stats](#stats)), so you can measure your own setup.
## Configuration
Agents should use `configure`. Humans can edit JSON. Layers (later wins):
1. built-in defaults
2. user: `~/.config/local-llm-worker/config.json` (respects `$XDG_CONFIG_HOME`)
3. project: `<git root>/.local-llm-worker.json`
4. env: `LLW_<KEY>`, e.g. `LLW_BASE_URL`, `LLW_MODEL`. Integers are digits only, booleans
`true`/`false`/`1`/`0`/`yes`/`no`, `link_dirs` a comma list or JSON array, `headers` a JSON
object. An invalid value is ignored and `configure` lists it under `warnings`.
A config file with invalid JSON is an error that names the file.
| Key | Default | Meaning |
|---|---|---|
| `base_url` | `http://localhost:11434` | Backend root. A trailing `/v1` is stripped. |
| `api` | `auto` | `auto` · `ollama` · `openai`. Auto probes `/api/version`. |
| `api_key` | `""` | Sent as `Authorization: Bearer`. Shown as `(set)`, never printed. Plugin users can set it under the plugin's settings instead, which keeps it in the OS keychain. |
| `headers` | `{}` | Extra HTTP headers for a gateway, e.g. `{"X-Api-Key": "..."}`. Values shown as `(set)`. |
| `model` | `""` | Default model for both tools. |
| `offload_model` | `""` | Override for `offload` (e.g. long-context). |
| `delegate_model` | `""` | Override for `delegate` (e.g. a coder). |
| `research_model` | `""` | Override for `research` (e.g. long-context). |
| `search_url` | `""` | SearXNG instance for `research` (JSON output enabled). |
| `research_sources` | `3` | Pages `research` reads per search (max 10). |
| `page_timeout_ms` | `30000` | Per page download and per search. |
| `allow_private_urls` | `false` | Let `research` read pages on loopback, link-local or private addresses (the `search_url` itself is always allowed). |
| `num_ctx` | `32768` | Context window (Ollama). |
| `max_tokens` | `8192` | Output limit per answer, sent on both APIs (`num_predict` on Ollama). A cut-off answer is reported, not tested. |
| `keep_alive` | `""` | Ollama only: how long the model stays loaded, e.g. `"30m"` or `-1` (forever). `""` keeps the server default (5 min). |
| `concurrency` | `4` | Max in-flight LLM requests per server process. |
| `timeout_ms` | `600000` | Per LLM request, counted from slot acquisition. |
| `test_timeout_ms` | `600000` | Per test / command run. |
| `max_attempts` | `3` | Delegate rounds (cap 10). |
| `link_dirs` | `node_modules, .venv, venv, vendor` | Symlinked into each worktree. |
| `log_path` | `~/.local-llm-worker/runs.jsonl` | Run log; `""` disables. |
| `auto_use` | `offload, research` | Tools Claude uses on its own, via a short session-start reminder. `delegate` is opt-in. Every tool still works when you ask forLo que la gente pregunta sobre local-llm-worker
¿Qué es JoblessJoe/local-llm-worker?
+
JoblessJoe/local-llm-worker es mcp servers para el ecosistema de Claude AI. Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code. Tiene 1 estrellas en GitHub y su última actualización registrada es del 2026-10-04.
¿Cómo se instala local-llm-worker?
+
Puedes instalar local-llm-worker clonando el repositorio (https://github.com/JoblessJoe/local-llm-worker) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar JoblessJoe/local-llm-worker?
+
Nuestro agente de seguridad ha analizado JoblessJoe/local-llm-worker y le ha asignado un Trust Score de 80/100 (tier: Trusted). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene JoblessJoe/local-llm-worker?
+
JoblessJoe/local-llm-worker es mantenido por JoblessJoe. La última actividad registrada en GitHub es del 2026-10-04, con 0 issues abiertos.
¿Hay alternativas a local-llm-worker?
+
Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.
Despliega local-llm-worker en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/joblessjoe-local-llm-worker)<a href="https://claudewave.com/repo/joblessjoe-local-llm-worker"><img src="https://claudewave.com/api/badge/joblessjoe-local-llm-worker" alt="Featured on ClaudeWave: JoblessJoe/local-llm-worker" width="320" height="64" /></a>Más MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.