ollama
Deploy and serve local models with Ollama — pull and run them, then expose the OpenAI-compatible endpoint to apps and agents.
git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness /tmp/ollama && cp -r /tmp/ollama/packages/skills/skills/ollama ~/.claude/skills/ollamaSKILL.md
# Ollama Serving Ollama runs open-weight models locally with automatic GPU detection and an OpenAI-compatible API on `http://localhost:11434`. ## Before you start If the user's message only invokes this skill (e.g. "use ollama skill") without a concrete request, ask the user what they want. Do not run any command until the goal is clear. Ask the user which model to run; if they have no preference, recommend the small default [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) (`ollama pull qwen3.5:0.8b`). The model must fit the machine's RAM/VRAM. Ollama runs everywhere — macOS, Linux and Windows, on CPUs as well as NVIDIA/AMD GPUs — so engine choice follows the user's preference: Ollama is the simple default, while vLLM targets high-throughput GPU serving. Check the current state first: ```bash ollama --version # is Ollama installed? ollama ps # is the service already serving models? ``` If port 11434 is already serving, reuse that instance — never kill an existing Ollama process. ## Suggested workflow 1. Ask the user which model to run; with no preference, recommend [Qwen/Qwen3.5-0.8B](https://huggingface.co/Qwen/Qwen3.5-0.8B) (`qwen3.5:0.8b`). 2. Pick the engine the user prefers: Ollama is the default; vLLM covers high-throughput GPU serving. 3. Install Ollama if missing, then `ollama pull qwen3.5:0.8b`. 4. Verify with `curl http://localhost:11434/v1/models`. 5. Register the endpoint: `penguin config model add ... --client-type openai-chat --base-url http://localhost:11434/v1` — a pulled Ollama model is not visible to Penguin until added. 6. Confirm the new entry with `penguin config model list`. ## Install ```bash curl -fsSL https://ollama.com/install.sh | sh # Linux; macOS/Windows use the desktop app ``` The service then listens on `http://localhost:11434`. ## Pull and run ```bash ollama pull qwen3.5:0.8b # download a model ollama run qwen3.5:0.8b # interactive chat (pulls first if missing) ollama list # downloaded models ollama ps # models loaded in memory ollama stop qwen3.5:0.8b # unload a model ``` ## OpenAI-compatible endpoint The endpoint is `http://localhost:11434/v1`; any non-empty API key is accepted (conventionally `ollama`): ```bash curl http://localhost:11434/v1/models ``` ## Context length The default context window is small, and agent sessions need a large one. Raise it in the server's environment: ```bash OLLAMA_CONTEXT_LENGTH=32768 ollama serve # systemd service: set it via `systemctl edit ollama` ``` Or bake it into a model variant with a Modelfile: ``` FROM qwen3.5:0.8b PARAMETER num_ctx 32768 ``` ```bash ollama create qwen3.5-32k -f Modelfile ``` ## Register with PenguinHarness Model configuration is the penguin CLI's job — `penguin config model add` registers an endpoint and `penguin config model list` shows what has been registered. A pulled Ollama model is not visible to Penguin until you add it: ```bash penguin config model add --provider custom --client-type openai-chat \ --base-url http://localhost:11434/v1 --model-id qwen3.5:0.8b --api-key ollama penguin config model list # the new entry should now be listed ```
Use when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release workflow; writing or auditing changelog entries; writing a blog post or capturing release screenshots; deciding what to do about data already on disk; or auditing prose that reads like a leaked authoring session. Covers the two-repo symlink layout, the CI-parity verification chain, the record-and-ship contract, where blog media is hosted, and the seams that are intentional.
Use when changing the PenguinHarness Web App (`packages/web`) — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a form field, writing user-facing copy, or building a popup. Covers the semantic tone tokens, the icon size/stroke/gap scale, the semantic-versus-formatting rule for explanatory text, the two-dictionary i18n contract, and the portal-panel pattern with its Esc and scroll caveats.
Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
Initialize an Agent's settings from a user requirement by writing AGENTS.md, setting identity metadata, and installing only needed Skills.