Skip to main content
ClaudeWave
JoblessJoe avatar
JoblessJoe

local-llm-worker

View on GitHub

Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code.

MCP ServersOfficial Registry1 stars0 forks● JavaScriptNOASSERTIONUpdated today
ClaudeWave Trust Score
80/100
✓ Trusted
Passed
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Flags
  • !Licence file present but not machine-readable
Last scanned: 10/5/2026
Install in Claude Code / Claude Desktop
Method: Manual
Claude Code CLI
git clone https://github.com/JoblessJoe/local-llm-worker
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "local-llm-worker": {
      "command": "node",
      "args": ["/path/to/local-llm-worker/dist/index.js"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
💡 Clone https://github.com/JoblessJoe/local-llm-worker and follow its README for install instructions.
Use cases

MCP Servers overview

<div align="center">

# local-llm-worker

**Let Claude hand the bulk reading to your local LLM: test logs, big files, web pages. Claude only gets the answer.**

[![test](https://github.com/JoblessJoe/local-llm-worker/actions/workflows/test.yml/badge.svg)](https://github.com/JoblessJoe/local-llm-worker/actions/workflows/test.yml)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
[![Node ≥ 18](https://img.shields.io/badge/node-%E2%89%A5%2018-339933.svg)](package.json)
[![Zero dependencies](https://img.shields.io/badge/dependencies-0-brightgreen.svg)](package.json)
[![Claude Code plugin](https://img.shields.io/badge/Claude%20Code-plugin-d97757.svg)](#install)
[![MCP server](https://img.shields.io/badge/MCP-server-6e56cf.svg)](#use-it-in-any-mcp-client)

Works with Ollama · llama.cpp · LM Studio · vLLM · LocalAI · any OpenAI-compatible endpoint

</div>

---

Reading a 5,000-line test log or three docs pages costs the same frontier-model tokens as hard
architectural work. **local-llm-worker** is an MCP server and Claude Code plugin that moves
that bulk work onto the GPU (or CPU) you already own:

- **`offload`**: your local model runs the noisy command or reads the big file. Claude gets
  the answer.
- **`research`**: your local model searches the web, reads the pages in full and returns a
  cited answer. Claude never sees the pages.
- **`delegate`** *(opt-in)*: your local model writes one file in an isolated git worktree
  and retries until **a test Claude wrote first** passes. Claude gets a one-line verdict, not
  the code.

```text
> offload  npm test 2>&1   "Which test fails, and why?"
`fixed coupon never goes below zero` (test/cart.test.js:404): applyCoupon(5, {type:'fixed',
value:10}) returns -5, expected 0.
— qwen3-coder:30b-a3b-q4_K_M · 19,132 input tokens read locally
```

## Why

| | Claude does it | Claude delegates it |
|---|---|---|
| Find one failure in a 50 KB test log | ~19,000 tokens of log in context | a ~130-token answer |
| Answer from three docs pages | the pages, or a summary of them | a cited paragraph + source list, from the full pages |
| Write a helper that tests can pin down | output tokens for the code, then re-reading it | writes the test, reads `PASS` |
| Several independent helpers | sequential edits | parallel delegates, each in its own worktree |

- **Any model, any hardware.** No hardcoded models, no GPU assumptions. CPU-only works, just
  slower.
- **Parallel by default.** Several calls in one turn, or from several subagents, run at once.
  The concurrency limit is a config value, not a hardcoded lock.
- **Fits the context automatically.** Material is sized to the model's context window per
  text (number-heavy logs need more tokens than prose). Anything cut is reported.
- **Zero dependencies.** Plain Node ≥ 18. Installing from git needs no `npm install`.
- **Agent-configurable.** One `configure` call shows the config, the backend and its models,
  and changes any setting. It takes effect on the next call, with no restart.

## Install

**Prerequisites:** Node ≥ 18, git, and a running local model server (e.g.
`ollama pull qwen3-coder:30b`).

### Claude Code plugin

```text
/plugin marketplace add JoblessJoe/local-llm-worker
/plugin install local-llm-worker@local-llm-worker
```

Then just ask: *"set up local-llm-worker for my machine"*. Claude finds your backend, picks a
model, and asks which tools it should use on its own (`auto_use`). By default that's `offload`
and `research`; `delegate` is opt-in. Every tool also works whenever you ask for it.

### Make Claude use it every time

A session-start reminder nudges Claude toward the tools you chose. For dependable use, add
this to your `CLAUDE.md`; setup offers to do it for you:

```markdown
## Local LLM (local-llm-worker)
- Run test suites, builds and other noisy commands through `offload` (`command`), and read logs or files over ~300 lines through it, instead of reading the output yourself.
- Use `research` for web lookups instead of WebSearch/WebFetch.
```

With this in place, Claude ran a noisy failing test suite through `offload` every time we
tried. That session cost 40 % less than one that read the output itself.

### Use it in any MCP client

Claude Desktop, Cursor, Windsurf, and others:

```json
{
  "mcpServers": {
    "local-llm-worker": {
      "command": "node",
      "args": ["/path/to/local-llm-worker/src/index.js"],
      "env": { "LLW_MODEL": "qwen3-coder:30b" }
    }
  }
}
```

Claude Code without the plugin:
`claude mcp add local-llm-worker -- node /path/to/local-llm-worker/src/index.js`

## How `delegate` works

```mermaid
flowchart LR
    A[Claude writes the test<br/>+ spec + invariants] --> B[delegate]
    B --> C[fresh git worktree<br/>+ your uncommitted changes]
    C --> D[local LLM writes<br/>target file]
    D --> E{run test}
    E -- fail --> F[parse failures<br/>test source stays hidden]
    F --> D
    E -- pass --> G[copy file into checkout<br/>unless it changed meanwhile]
    G --> H[Claude gets a one-line verdict]
```

- **Isolated.** Each call gets its own `git worktree` with your uncommitted changes mirrored in,
  plus symlinked `node_modules` / `.venv`, so parallel calls never collide.
- **Locked scope.** Exactly one target file, jailed to the repo, and never one of the
  `test_files` you name.
- **Useful feedback.** Retries get the parsed failing assertions (TAP, pytest, jest, vitest,
  go, cargo), not a stack-trace tail.
- **Safe apply.** If you or another delegate touched the target meanwhile, nothing is
  overwritten.
- **Honest failure.** After N attempts you get the last failure and the kept worktree. Transport
  errors are reported as errors, never as a model FAIL.
- **Self-cleaning.** Each delegate prunes `llw-*` worktrees left behind by a crashed server, and
  kept ones older than 24 h. Untracked files over 10 MB are not copied into the worktree (the
  verdict says how many were skipped).

The bundled [skill](skills/local-llm-worker/SKILL.md) teaches Claude how to write specs that
pass: one invariant per test, exact API surface in `context`, properties instead of examples, and
never delegating auth or money code.

## Results

On one 24 GB GPU (Tesla P40) with `qwen3-coder:30b-a3b-q4_K_M` on Ollama:

| Task | Read locally | Claude received |
|---|---|---|
| `offload`: find the failing test in a 50 KB test log | ~19,100 tokens | ~130 tokens |
| `offload`: find one ERROR line in a 90 KB log | ~30,400 tokens | the line, quoted |
| `research`: "Latest Node.js LTS and its end of life?" (3 pages) | 7,025 tokens | a cited answer, ~250 tokens |

Several calls run in parallel: three in one turn took as long as the slowest one.

### Which model?

From a reproducible benchmark ([`bench/`](bench/)) of 120 `delegate` runs across six task
types:

| Model | Good for | Speed per call |
|---|---|---|
| `qwen3-coder:30b-a3b` | the best default: parsers, pure functions, Python | 15–55 s |
| `devstral-small-2:24b` | edits to existing files, parsers | 2–9 min |
| `granite4.1:8b` | small, well-specified functions on modest hardware | 25–90 s |

Retries matter: with failure feedback, qwen3-coder's pass rate rose by half from the first
attempt to the third. `offload` and `research` work well with any of these. Every run is logged
(see [Stats](#stats)), so you can measure your own setup.

## Configuration

Agents should use `configure`. Humans can edit JSON. Layers (later wins):

1. built-in defaults
2. user: `~/.config/local-llm-worker/config.json` (respects `$XDG_CONFIG_HOME`)
3. project: `<git root>/.local-llm-worker.json`
4. env: `LLW_<KEY>`, e.g. `LLW_BASE_URL`, `LLW_MODEL`. Integers are digits only, booleans
   `true`/`false`/`1`/`0`/`yes`/`no`, `link_dirs` a comma list or JSON array, `headers` a JSON
   object. An invalid value is ignored and `configure` lists it under `warnings`.

A config file with invalid JSON is an error that names the file.

| Key | Default | Meaning |
|---|---|---|
| `base_url` | `http://localhost:11434` | Backend root. A trailing `/v1` is stripped. |
| `api` | `auto` | `auto` · `ollama` · `openai`. Auto probes `/api/version`. |
| `api_key` | `""` | Sent as `Authorization: Bearer`. Shown as `(set)`, never printed. Plugin users can set it under the plugin's settings instead, which keeps it in the OS keychain. |
| `headers` | `{}` | Extra HTTP headers for a gateway, e.g. `{"X-Api-Key": "..."}`. Values shown as `(set)`. |
| `model` | `""` | Default model for both tools. |
| `offload_model` | `""` | Override for `offload` (e.g. long-context). |
| `delegate_model` | `""` | Override for `delegate` (e.g. a coder). |
| `research_model` | `""` | Override for `research` (e.g. long-context). |
| `search_url` | `""` | SearXNG instance for `research` (JSON output enabled). |
| `research_sources` | `3` | Pages `research` reads per search (max 10). |
| `page_timeout_ms` | `30000` | Per page download and per search. |
| `allow_private_urls` | `false` | Let `research` read pages on loopback, link-local or private addresses (the `search_url` itself is always allowed). |
| `num_ctx` | `32768` | Context window (Ollama). |
| `max_tokens` | `8192` | Output limit per answer, sent on both APIs (`num_predict` on Ollama). A cut-off answer is reported, not tested. |
| `keep_alive` | `""` | Ollama only: how long the model stays loaded, e.g. `"30m"` or `-1` (forever). `""` keeps the server default (5 min). |
| `concurrency` | `4` | Max in-flight LLM requests per server process. |
| `timeout_ms` | `600000` | Per LLM request, counted from slot acquisition. |
| `test_timeout_ms` | `600000` | Per test / command run. |
| `max_attempts` | `3` | Delegate rounds (cap 10). |
| `link_dirs` | `node_modules, .venv, venv, vendor` | Symlinked into each worktree. |
| `log_path` | `~/.local-llm-worker/runs.jsonl` | Run log; `""` disables. |
| `auto_use` | `offload, research` | Tools Claude uses on its own, via a short session-start reminder. `delegate` is opt-in. Every tool still works when you ask for
ai-agentsclaude-codeclaude-code-pluginllama-cppllm-toolslm-studiolocal-llmmcpmcp-serverollamatoken-savingvllm

What people ask about local-llm-worker

What is JoblessJoe/local-llm-worker?

+

JoblessJoe/local-llm-worker is mcp servers for the Claude AI ecosystem. Local LLM does Claude's bulk work: reads logs and files, researches the web, writes test-gated code. It has 1 GitHub stars and its last recorded update is dated 2026-10-04.

How do I install local-llm-worker?

+

You can install local-llm-worker by cloning the repository (https://github.com/JoblessJoe/local-llm-worker) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is JoblessJoe/local-llm-worker safe to use?

+

Our security agent has analyzed JoblessJoe/local-llm-worker and assigned a Trust Score of 80/100 (tier: Trusted). See the full breakdown of passed checks and flags on this page.

Who maintains JoblessJoe/local-llm-worker?

+

JoblessJoe/local-llm-worker is maintained by JoblessJoe. The last recorded GitHub activity is dated 2026-10-04, with 0 open issues.

Are there alternatives to local-llm-worker?

+

Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.

Deploy local-llm-worker to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: JoblessJoe/local-llm-worker
[![Featured on ClaudeWave](https://claudewave.com/api/badge/joblessjoe-local-llm-worker)](https://claudewave.com/repo/joblessjoe-local-llm-worker)
<a href="https://claudewave.com/repo/joblessjoe-local-llm-worker"><img src="https://claudewave.com/api/badge/joblessjoe-local-llm-worker" alt="Featured on ClaudeWave: JoblessJoe/local-llm-worker" width="320" height="64" /></a>

More MCP Servers

local-llm-worker alternatives