dspy-finetune-bootstrap
This Claude Code skill uses DSPy's BootstrapFinetune optimizer to distill a large teacher language model into a smaller student model with fine-tuned weights, enabling efficient production deployment with reduced inference costs. Use it when you need to compress a working DSPy program into a lightweight model suitable for resource-constrained environments while maintaining performance through knowledge distillation.
git clone --depth 1 https://github.com/OmidZamani/dspy-skills /tmp/dspy-finetune-bootstrap && cp -r /tmp/dspy-finetune-bootstrap/skills/dspy-finetune-bootstrap ~/.claude/skills/dspy-finetune-bootstrapSKILL.md
# DSPy BootstrapFinetune Optimizer
## Goal
Distill a DSPy program into fine-tuned model weights for efficient production deployment.
## When to Use
- You have a working DSPy program with a large model
- Need to reduce inference costs
- Want faster responses (smaller model)
- Deploying to resource-constrained environments
## Inputs
| Input | Type | Description |
|-------|------|-------------|
| `program` | `dspy.Module` | Teacher program to distill |
| `trainset` | `list[dspy.Example]` | Training examples |
| `metric` | `callable` | Validation metric (optional) |
| `train_kwargs` | `dict` | Training hyperparameters |
## Outputs
| Output | Type | Description |
|--------|------|-------------|
| `finetuned_program` | `dspy.Module` | Program with fine-tuned weights |
| `model_path` | `str` | Path to saved model |
## Workflow
### Phase 1: Prepare Teacher Program
```python
import dspy
# Configure with strong teacher model
dspy.configure(lm=dspy.LM("openai/gpt-4o"))
class TeacherQA(dspy.Module):
def __init__(self):
self.cot = dspy.ChainOfThought("question -> answer")
def forward(self, question):
return self.cot(question=question)
```
### Phase 2: Configure Fine-Tuning
Assign the LM directly to predictors before fine-tuning:
```python
import dspy
from dspy.teleprompt import BootstrapFinetune
optimizer = BootstrapFinetune(
metric=lambda gold, pred, trace=None: gold.answer.lower() in pred.answer.lower(),
train_kwargs={
'learning_rate': 5e-5,
'num_train_epochs': 3,
'per_device_train_batch_size': 4,
'warmup_ratio': 0.1
}
)
```
### Phase 3: Fine-tune Student Model
```python
teacher = TeacherQA()
teacher.set_lm(dspy.settings.lm)
finetuned = optimizer.compile(teacher, trainset=trainset)
```
### Phase 4: Deploy
```python
# Save the fine-tuned model (saves state-only by default)
finetuned.save("finetuned_qa_model.json")
# Load and use (must recreate architecture first)
loaded = TeacherQA()
loaded.load("finetuned_qa_model.json")
result = loaded(question="What is machine learning?")
```
## Production Example
```python
import dspy
from dspy.teleprompt import BootstrapFinetune
from dspy.evaluate import Evaluate
import logging
import os
logger = logging.getLogger(__name__)
class ClassificationSignature(dspy.Signature):
"""Classify text into categories."""
text: str = dspy.InputField()
label: str = dspy.OutputField(desc="Category: positive, negative, neutral")
class TextClassifier(dspy.Module):
def __init__(self):
self.classify = dspy.Predict(ClassificationSignature)
def forward(self, text):
return self.classify(text=text)
def classification_metric(gold, pred, trace=None):
"""Exact label match."""
gold_label = gold.label.lower().strip()
pred_label = pred.label.lower().strip() if pred.label else ""
return gold_label == pred_label
def finetune_classifier(trainset, devset, output_dir="./finetuned_model"):
"""Full fine-tuning pipeline."""
# Configure teacher (strong model)
dspy.configure(lm=dspy.LM("openai/gpt-4o"))
teacher = TextClassifier()
teacher.set_lm(dspy.settings.lm)
# Evaluate teacher
evaluator = Evaluate(devset=devset, metric=classification_metric, num_threads=8)
teacher_score = evaluator(teacher)
logger.info(f"Teacher score: {teacher_score:.2%}")
# Fine-tune (train_kwargs passed to constructor)
optimizer = BootstrapFinetune(
metric=classification_metric,
train_kwargs={
'learning_rate': 2e-5,
'num_train_epochs': 3,
'per_device_train_batch_size': 8,
'gradient_accumulation_steps': 2,
'warmup_ratio': 0.1,
'weight_decay': 0.01,
'logging_steps': 10,
'save_strategy': 'epoch',
'output_dir': output_dir
}
)
finetuned = optimizer.compile(
teacher,
trainset=trainset
)
# Evaluate fine-tuned model
student_score = evaluator(finetuned)
logger.info(f"Student score: {student_score:.2%}")
# Save (state-only as JSON)
finetuned.save(os.path.join(output_dir, "final_model.json"))
return {
"teacher_score": teacher_score,
"student_score": student_score,
"model_path": os.path.join(output_dir, "final_model.json")
}
# For RAG fine-tuning
class RAGClassifier(dspy.Module):
"""RAG pipeline that can be fine-tuned."""
def __init__(self, num_passages=3):
self.retrieve = dspy.Retrieve(k=num_passages)
self.classify = dspy.ChainOfThought("context, text -> label")
def forward(self, text):
context = self.retrieve(text).passages
return self.classify(context=context, text=text)
def finetune_rag_classifier(trainset, devset):
"""Fine-tune a RAG-based classifier."""
# Configure retriever and LM
colbert = dspy.ColBERTv2(url='http://20.102.90.50:2017/wiki17_abstracts')
dspy.configure(
lm=dspy.LM("openai/gpt-4o"),
rm=colbert
)
rag = RAGClassifier()
rag.set_lm(dspy.settings.lm)
# Fine-tune (train_kwargs in constructor)
optimizer = BootstrapFinetune(
metric=classification_metric,
train_kwargs={
'learning_rate': 1e-5,
'num_train_epochs': 5
}
)
finetuned = optimizer.compile(
rag,
trainset=trainset
)
return finetuned
```
## Training Arguments Reference
| Argument | Description | Typical Value |
|----------|-------------|---------------|
| `learning_rate` | Learning rate | 1e-5 to 5e-5 |
| `num_train_epochs` | Training epochs | 3-5 |
| `per_device_train_batch_size` | Batch size | 4-16 |
| `gradient_accumulation_steps` | Gradient accumulation | 2-8 |
| `warmup_ratio` | Warmup proportion | 0.1 |
| `weight_decay` | L2 regularization | 0.01 |
| `max_grad_norm` | Gradient clipping | 1.0 |
## Best Practices
1. **Strong teacher** - Use GPUse this skill when you need to QA audit and fix a plugin skill file. Provides a methodology for verifying skill content against official documentation, fixing issues in-place, and producing verification reports.
Use for DSPy adapter selection, JSONAdapter, XMLAdapter, ChatAdapter, native function calling, structured outputs, and multimodal inputs like dspy.Image or dspy.Audio.
Use for composing DSPy modules with Ensemble, MultiChainComparison, ensemble voting, sequential pipelines, and multi-program workflows.
Use for BetterTogether, prompt plus weight optimization, fine-tuning sequences, and strategy chains like p -> w -> p.
Use for BootstrapFewShot, bootstrapped demonstrations, teacher-model demos, and low-data DSPy prompt optimization.
Use for creating custom DSPy modules, extending dspy.Module, reusable components, stateful modules, serialization, and module testing.
Use for debugging DSPy programs, inspect_history, tracing LLM calls, custom callbacks, observability, monitoring, and cost tracking.
Use for DSPy retrieval with dspy.Embedder, dspy.Embeddings, FAISS indexes, semantic search, and local or hosted embedding models.