Skip to main content
ClaudeWave
JoblessJoe avatar
JoblessJoe

local-llm-worker

Ver en GitHub

Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code.

MCP ServersRegistry oficial1 estrellas0 forks● JavaScriptNOASSERTIONActualizado today
ClaudeWave Trust Score
80/100
✓ Trusted
Passed
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Flags
  • !Licence file present but not machine-readable
Last scanned: 10/5/2026
Install in Claude Code / Claude Desktop
Method: Manual
Claude Code CLI
git clone https://github.com/JoblessJoe/local-llm-worker
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "local-llm-worker": {
      "command": "node",
      "args": ["/path/to/local-llm-worker/dist/index.js"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
💡 Clone https://github.com/JoblessJoe/local-llm-worker and follow its README for install instructions.
Casos de uso

Resumen de MCP Servers

<div align="center">

# local-llm-worker

**Let Claude hand the bulk reading to your local LLM: test logs, big files, web pages. Claude only gets the answer.**

[![test](https://github.com/JoblessJoe/local-llm-worker/actions/workflows/test.yml/badge.svg)](https://github.com/JoblessJoe/local-llm-worker/actions/workflows/test.yml)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
[![Node ≥ 18](https://img.shields.io/badge/node-%E2%89%A5%2018-339933.svg)](package.json)
[![Zero dependencies](https://img.shields.io/badge/dependencies-0-brightgreen.svg)](package.json)
[![Claude Code plugin](https://img.shields.io/badge/Claude%20Code-plugin-d97757.svg)](#install)
[![MCP server](https://img.shields.io/badge/MCP-server-6e56cf.svg)](#use-it-in-any-mcp-client)

Works with Ollama · llama.cpp · LM Studio · vLLM · LocalAI · any OpenAI-compatible endpoint

</div>

---

Reading a 5,000-line test log or three docs pages costs the same frontier-model tokens as hard
architectural work. **local-llm-worker** is an MCP server and Claude Code plugin that moves
that bulk work onto the GPU (or CPU) you already own:

- **`offload`**: your local model runs the noisy command or reads the big file. Claude gets
  the answer.
- **`research`**: your local model searches the web, reads the pages in full and returns a
  cited answer. Claude never sees the pages.
- **`delegate`** *(opt-in)*: your local model writes one file in an isolated git worktree
  and retries until **a test Claude wrote first** passes. Claude gets a one-line verdict, not
  the code.

```text
> offload  npm test 2>&1   "Which test fails, and why?"
`fixed coupon never goes below zero` (test/cart.test.js:404): applyCoupon(5, {type:'fixed',
value:10}) returns -5, expected 0.
— qwen3-coder:30b-a3b-q4_K_M · 19,132 input tokens read locally
```

## Why

| | Claude does it | Claude delegates it |
|---|---|---|
| Find one failure in a 50 KB test log | ~19,000 tokens of log in context | a ~130-token answer |
| Answer from three docs pages | the pages, or a summary of them | a cited paragraph + source list, from the full pages |
| Write a helper that tests can pin down | output tokens for the code, then re-reading it | writes the test, reads `PASS` |
| Several independent helpers | sequential edits | parallel delegates, each in its own worktree |

- **Any model, any hardware.** No hardcoded models, no GPU assumptions. CPU-only works, just
  slower.
- **Parallel by default.** Several calls in one turn, or from several subagents, run at once.
  The concurrency limit is a config value, not a hardcoded lock.
- **Fits the context automatically.** Material is sized to the model's context window per
  text (number-heavy logs need more tokens than prose). Anything cut is reported.
- **Zero dependencies.** Plain Node ≥ 18. Installing from git needs no `npm install`.
- **Agent-configurable.** One `configure` call shows the config, the backend and its models,
  and changes any setting. It takes effect on the next call, with no restart.

## Install

**Prerequisites:** Node ≥ 18, git, and a running local model server (e.g.
`ollama pull qwen3-coder:30b`).

### Claude Code plugin

```text
/plugin marketplace add JoblessJoe/local-llm-worker
/plugin install local-llm-worker@local-llm-worker
```

Then just ask: *"set up local-llm-worker for my machine"*. Claude finds your backend, picks a
model, and asks which tools it should use on its own (`auto_use`). By default that's `offload`
and `research`; `delegate` is opt-in. Every tool also works whenever you ask for it.

### Make Claude use it every time

A session-start reminder nudges Claude toward the tools you chose. For dependable use, add
this to your `CLAUDE.md`; setup offers to do it for you:

```markdown
## Local LLM (local-llm-worker)
- Run test suites, builds and other noisy commands through `offload` (`command`), and read logs or files over ~300 lines through it, instead of reading the output yourself.
- Use `research` for web lookups instead of WebSearch/WebFetch.
```

With this in place, Claude ran a noisy failing test suite through `offload` every time we
tried. That session cost 40 % less than one that read the output itself.

### Use it in any MCP client

Claude Desktop, Cursor, Windsurf, and others:

```json
{
  "mcpServers": {
    "local-llm-worker": {
      "command": "node",
      "args": ["/path/to/local-llm-worker/src/index.js"],
      "env": { "LLW_MODEL": "qwen3-coder:30b" }
    }
  }
}
```

Claude Code without the plugin:
`claude mcp add local-llm-worker -- node /path/to/local-llm-worker/src/index.js`

## How `delegate` works

```mermaid
flowchart LR
    A[Claude writes the test<br/>+ spec + invariants] --> B[delegate]
    B --> C[fresh git worktree<br/>+ your uncommitted changes]
    C --> D[local LLM writes<br/>target file]
    D --> E{run test}
    E -- fail --> F[parse failures<br/>test source stays hidden]
    F --> D
    E -- pass --> G[copy file into checkout<br/>unless it changed meanwhile]
    G --> H[Claude gets a one-line verdict]
```

- **Isolated.** Each call gets its own `git worktree` with your uncommitted changes mirrored in,
  plus symlinked `node_modules` / `.venv`, so parallel calls never collide.
- **Locked scope.** Exactly one target file, jailed to the repo, and never one of the
  `test_files` you name.
- **Useful feedback.** Retries get the parsed failing assertions (TAP, pytest, jest, vitest,
  go, cargo), not a stack-trace tail.
- **Safe apply.** If you or another delegate touched the target meanwhile, nothing is
  overwritten.
- **Honest failure.** After N attempts you get the last failure and the kept worktree. Transport
  errors are reported as errors, never as a model FAIL.
- **Self-cleaning.** Each delegate prunes `llw-*` worktrees left behind by a crashed server, and
  kept ones older than 24 h. Untracked files over 10 MB are not copied into the worktree (the
  verdict says how many were skipped).

The bundled [skill](skills/local-llm-worker/SKILL.md) teaches Claude how to write specs that
pass: one invariant per test, exact API surface in `context`, properties instead of examples, and
never delegating auth or money code.

## Results

On one 24 GB GPU (Tesla P40) with `qwen3-coder:30b-a3b-q4_K_M` on Ollama:

| Task | Read locally | Claude received |
|---|---|---|
| `offload`: find the failing test in a 50 KB test log | ~19,100 tokens | ~130 tokens |
| `offload`: find one ERROR line in a 90 KB log | ~30,400 tokens | the line, quoted |
| `research`: "Latest Node.js LTS and its end of life?" (3 pages) | 7,025 tokens | a cited answer, ~250 tokens |

Several calls run in parallel: three in one turn took as long as the slowest one.

### Which model?

From a reproducible benchmark ([`bench/`](bench/)) of 120 `delegate` runs across six task
types:

| Model | Good for | Speed per call |
|---|---|---|
| `qwen3-coder:30b-a3b` | the best default: parsers, pure functions, Python | 15–55 s |
| `devstral-small-2:24b` | edits to existing files, parsers | 2–9 min |
| `granite4.1:8b` | small, well-specified functions on modest hardware | 25–90 s |

Retries matter: with failure feedback, qwen3-coder's pass rate rose by half from the first
attempt to the third. `offload` and `research` work well with any of these. Every run is logged
(see [Stats](#stats)), so you can measure your own setup.

## Configuration

Agents should use `configure`. Humans can edit JSON. Layers (later wins):

1. built-in defaults
2. user: `~/.config/local-llm-worker/config.json` (respects `$XDG_CONFIG_HOME`)
3. project: `<git root>/.local-llm-worker.json`
4. env: `LLW_<KEY>`, e.g. `LLW_BASE_URL`, `LLW_MODEL`. Integers are digits only, booleans
   `true`/`false`/`1`/`0`/`yes`/`no`, `link_dirs` a comma list or JSON array, `headers` a JSON
   object. An invalid value is ignored and `configure` lists it under `warnings`.

A config file with invalid JSON is an error that names the file.

| Key | Default | Meaning |
|---|---|---|
| `base_url` | `http://localhost:11434` | Backend root. A trailing `/v1` is stripped. |
| `api` | `auto` | `auto` · `ollama` · `openai`. Auto probes `/api/version`. |
| `api_key` | `""` | Sent as `Authorization: Bearer`. Shown as `(set)`, never printed. Plugin users can set it under the plugin's settings instead, which keeps it in the OS keychain. |
| `headers` | `{}` | Extra HTTP headers for a gateway, e.g. `{"X-Api-Key": "..."}`. Values shown as `(set)`. |
| `model` | `""` | Default model for both tools. |
| `offload_model` | `""` | Override for `offload` (e.g. long-context). |
| `delegate_model` | `""` | Override for `delegate` (e.g. a coder). |
| `research_model` | `""` | Override for `research` (e.g. long-context). |
| `search_url` | `""` | SearXNG instance for `research` (JSON output enabled). |
| `research_sources` | `3` | Pages `research` reads per search (max 10). |
| `page_timeout_ms` | `30000` | Per page download and per search. |
| `allow_private_urls` | `false` | Let `research` read pages on loopback, link-local or private addresses (the `search_url` itself is always allowed). |
| `num_ctx` | `32768` | Context window (Ollama). |
| `max_tokens` | `8192` | Output limit per answer, sent on both APIs (`num_predict` on Ollama). A cut-off answer is reported, not tested. |
| `keep_alive` | `""` | Ollama only: how long the model stays loaded, e.g. `"30m"` or `-1` (forever). `""` keeps the server default (5 min). |
| `concurrency` | `4` | Max in-flight LLM requests per server process. |
| `timeout_ms` | `600000` | Per LLM request, counted from slot acquisition. |
| `test_timeout_ms` | `600000` | Per test / command run. |
| `max_attempts` | `3` | Delegate rounds (cap 10). |
| `link_dirs` | `node_modules, .venv, venv, vendor` | Symlinked into each worktree. |
| `log_path` | `~/.local-llm-worker/runs.jsonl` | Run log; `""` disables. |
| `auto_use` | `offload, research` | Tools Claude uses on its own, via a short session-start reminder. `delegate` is opt-in. Every tool still works when you ask for
ai-agentsclaude-codeclaude-code-pluginllama-cppllm-toolslm-studiolocal-llmmcpmcp-serverollamatoken-savingvllm

Lo que la gente pregunta sobre local-llm-worker

¿Qué es JoblessJoe/local-llm-worker?

+

JoblessJoe/local-llm-worker es mcp servers para el ecosistema de Claude AI. Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code. Tiene 1 estrellas en GitHub y su última actualización registrada es del 2026-10-04.

¿Cómo se instala local-llm-worker?

+

Puedes instalar local-llm-worker clonando el repositorio (https://github.com/JoblessJoe/local-llm-worker) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.

¿Es seguro usar JoblessJoe/local-llm-worker?

+

Nuestro agente de seguridad ha analizado JoblessJoe/local-llm-worker y le ha asignado un Trust Score de 80/100 (tier: Trusted). Revisa el desglose completo de comprobaciones superadas y flags en esta página.

¿Quién mantiene JoblessJoe/local-llm-worker?

+

JoblessJoe/local-llm-worker es mantenido por JoblessJoe. La última actividad registrada en GitHub es del 2026-10-04, con 0 issues abiertos.

¿Hay alternativas a local-llm-worker?

+

Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.

Despliega local-llm-worker en tu cloud

Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.

¿Mantienes este repo? Añade un badge a tu README

Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.

Featured on ClaudeWave: JoblessJoe/local-llm-worker
[![Featured on ClaudeWave](https://claudewave.com/api/badge/joblessjoe-local-llm-worker)](https://claudewave.com/repo/joblessjoe-local-llm-worker)
<a href="https://claudewave.com/repo/joblessjoe-local-llm-worker"><img src="https://claudewave.com/api/badge/joblessjoe-local-llm-worker" alt="Featured on ClaudeWave: JoblessJoe/local-llm-worker" width="320" height="64" /></a>

Más MCP Servers

Alternativas a local-llm-worker