outlines
Outlines is a structured text generation library that constrains large language model outputs to valid JSON, XML, regex patterns, or Pydantic models by filtering invalid tokens at the generation level using finite state machines. Use it when you need guaranteed valid structured outputs from local or API-based models while maintaining generation speed and compatibility with Hugging Face Transformers, llama.cpp, or vLLM backends.
git clone --depth 1 https://github.com/NousResearch/hermes-agent /tmp/outlines && cp -r /tmp/outlines/optional-skills/mlops/inference/outlines ~/.claude/skills/outlinesSKILL.md
# Outlines: Structured Text Generation
## When to Use This Skill
Use Outlines when you need to:
- **Guarantee valid JSON/XML/code** structure during generation
- **Use Pydantic models** for type-safe outputs
- **Support local models** (Transformers, llama.cpp, vLLM)
- **Maximize inference speed** with zero-overhead structured generation
- **Generate against JSON schemas** automatically
- **Control token sampling** at the grammar level
**GitHub Stars**: 12,000+ | **From**: dottxt.ai (formerly .txt)
> **API note (Outlines 1.x):** This skill targets the current v1 API.
> The pre-1.0 helpers (`outlines.models.transformers(...)`,
> `outlines.generate.json/choice/regex/...`) have been **removed**. In v1 you
> create a model with `outlines.from_transformers(...)` (or `from_vllm`,
> `from_llamacpp`, `from_openai`) and then **call the model directly** with an
> output type: `model(prompt, output_type)`. JSON/Pydantic outputs are returned
> as a **JSON string** — validate with `YourModel.model_validate_json(result)`.
## Installation
```bash
# Base installation
pip install outlines
# With specific backends
pip install outlines transformers # Hugging Face models
pip install outlines llama-cpp-python # llama.cpp
pip install outlines vllm # vLLM for high-throughput
```
## Quick Start
### Basic Example: Classification
```python
import outlines
from typing import Literal
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
# v1: wrap a Transformers model + tokenizer
model = outlines.from_transformers(
AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
AutoTokenizer.from_pretrained(MODEL_NAME),
)
# Call the model directly with an output type
prompt = "Sentiment of 'This product is amazing!': "
sentiment = model(prompt, Literal["positive", "negative", "neutral"])
print(sentiment) # "positive" (guaranteed one of these)
```
### With Pydantic Models
```python
from pydantic import BaseModel
import outlines
from transformers import AutoModelForCausalLM, AutoTokenizer
class User(BaseModel):
name: str
age: int
email: str
MODEL_NAME = "microsoft/Phi-3-mini-4k-instruct"
model = outlines.from_transformers(
AutoModelForCausalLM.from_pretrained(MODEL_NAME, device_map="auto"),
AutoTokenizer.from_pretrained(MODEL_NAME),
)
# Generate structured output (returns a JSON string)
prompt = "Extract user: John Doe, 30 years old, john@example.com"
result = model(prompt, User, max_new_tokens=200)
user = User.model_validate_json(result) # parse into the Pydantic model
print(user.name) # "John Doe"
print(user.age) # 30
print(user.email) # "john@example.com"
```
## Core Concepts
### 1. Constrained Token Sampling
Outlines constrains token generation at the logit level using a compiled
automaton derived from your output type.
**How it works:**
1. Convert the output type (JSON/Pydantic/regex/`Literal`) to a schema/grammar
2. Compile the grammar into a token-level automaton
3. Filter invalid tokens at each step during generation
4. Fast-forward when only one valid token exists
**Benefits:**
- **Zero overhead**: Filtering happens at token level
- **Speed improvement**: Fast-forward through deterministic paths
- **Guaranteed validity**: Invalid outputs impossible
```python
import outlines
from pydantic import BaseModel
from transformers import AutoModelForCausalLM, AutoTokenizer
class Person(BaseModel):
name: str
age: int
model = outlines.from_transformers(
AutoModelForCausalLM.from_pretrained("microsoft/Phi-3-mini-4k-instruct", device_map="auto"),
AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct"),
)
result = model("Generate person: Alice, 25", Person)
person = Person.model_validate_json(result)
```
### 2. Output Types
In v1 you pass the desired **output type** directly as the second argument.
#### Multiple choice (`Literal`)
```python
from typing import Literal
sentiment = model("Review: This is great!", Literal["positive", "negative", "neutral"])
# Result: one of the three choices
```
#### JSON via Pydantic
```python
from pydantic import BaseModel
class Product(BaseModel):
name: str
price: float
in_stock: bool
result = model("Extract: iPhone 15, $999, available", Product)
product = Product.model_validate_json(result) # valid Product instance
```
#### Regex (pass a regex string)
```python
# Generate text matching a regex pattern
phone = model("Generate phone number:", r"[0-9]{3}-[0-9]{3}-[0-9]{4}")
# Result: "555-123-4567" (guaranteed to match the pattern)
```
#### Numeric types
```python
# Pass the Python type directly
age = model("Person's age:", int) # guaranteed integer
price = model("Product price:", float) # guaranteed float
```
### 3. Model Backends
Outlines supports multiple local and API-based backends via `from_*` factories.
#### Transformers (Hugging Face)
```python
import outlines
from transformers import AutoModelForCausalLM, AutoTokenizer
model = outlines.from_transformers(
AutoModelForCausalLM.from_pretrained("microsoft/Phi-3-mini-4k-instruct", device_map="auto"),
AutoTokenizer.from_pretrained("microsoft/Phi-3-mini-4k-instruct"),
)
result = model(prompt, YourModel)
```
#### llama.cpp
```python
import outlines
from llama_cpp import Llama
llm = Llama("./models/llama-3.1-8b-instruct.Q4_K_M.gguf", n_gpu_layers=35, n_ctx=4096)
model = outlines.from_llamacpp(llm)
result = model(prompt, YourModel)
```
#### vLLM (High Throughput)
```python
import outlines
from vllm import LLM
llm = LLM("meta-llama/Llama-3.1-8B-Instruct", tensor_parallel_size=2)
model = outlines.from_vllm(llm)
result = model(prompt, YourModel)
```
#### OpenAI (server-side constrained JSON)
```python
import outlines
from openai import OpenAI
client = OpenAI()
model = outlines.from_openai(client, "gpt-4o-mini")
# API backends support JSON-schema style structured output
result = model(prompt, YourModel)
```
### 4. Pydantic Integration
OOperate the Antigravity CLI (agy): plugins, auth, sandbox.
Delegate coding tasks to the Blackbox AI multi-model CLI.
Delegate coding to xAI Grok Build CLI (features, PRs).
Configure and troubleshoot Honcho memory for Hermes.
Delegate coding to OpenHands CLI (model-agnostic, LiteLLM).
Read-only EVM client: wallets, tokens, gas across 8 chains.
Hyperliquid market data, account history, trade review.
Query Solana wallets, tokens, txs, and NFTs in USD.