Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Healthy fork ratio
- ✓Clear description
- ✓Topics declared
- ✓Mature repo (>1y old)
claude mcp add doc-scraper -- docker run -i --rm ghcr.io/sriram-pr/doc-scraper{
"mcpServers": {
"doc-scraper": {
"command": "docker",
"args": ["run", "-i", "--rm", "ghcr.io/sriram-pr/doc-scraper"]
}
}
}MCP Servers overview
# LLM Documentation Scraper (`doc-scraper`)
[](https://golang.org/)
[](https://pkg.go.dev/github.com/Sriram-PR/doc-scraper/v2)
[](https://github.com/Sriram-PR/doc-scraper/blob/main/LICENSE)
[](https://glama.ai/mcp/servers/Sriram-PR/doc-scraper)
> A configurable, concurrent, and resumable web crawler written in Go. Specifically designed to scrape technical documentation websites, extract core content, convert it cleanly to Markdown format suitable for ingestion by Large Language Models (LLMs), and save the results locally.

## Overview
This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a `config.yaml` file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.
### Why Use This Tool?
- **Built for LLM Training & RAG Systems** - Creates clean, consistent Markdown optimized for ingestion
- **Preserves Documentation Structure** - Maintains the original site hierarchy for context preservation
- **Production-Ready Features** - Offers resumable crawls, rate limiting, and graceful error handling
- **High Performance** - Uses Go's concurrency model for efficient parallel processing
## Goal: Preparing Documentation for LLMs
The main objective of this tool is to automate the often tedious process of gathering and cleaning web-based documentation for use with Large Language Models. By converting structured web content into clean Markdown, it aims to provide a dataset that is:
- **Text-Focused:** Prioritizes the textual content extracted via CSS selectors
- **Structured:** Maintains the directory hierarchy of the original documentation site, preserving context
- **Cleaned:** Converts HTML to Markdown, removing web-specific markup and clutter
- **Locally Accessible:** Provides the content as local files for easier processing and pipeline integration
## Key Features
| Feature | Description |
|---------|-------------|
| **Configurable Crawling** | Uses YAML for global and site-specific settings |
| **Scope Control** | Limits crawling by domain, path prefix, and disallowed path patterns (regex) |
| **Content Extraction** | Extracts main content using CSS selectors |
| **HTML-to-Markdown** | Converts extracted HTML to clean GitHub-Flavored Markdown (tables, task lists, strikethrough) |
| **Image Handling** | Opt-in downloading and local rewriting of image links with domain and size filtering (disabled by default; doc-scraper is text-first) |
| **Link Rewriting** | Rewrites internal links to relative paths for local structure |
| **JSONL Output** | Optional one-record-per-page JSONL with a trailing crawl-summary record, for RAG ingestion |
| **Concurrency** | Configurable worker pools and semaphore-based request limits (global and per-host) |
| **Rate Limiting** | Configurable per-host delays with jitter |
| **Robots.txt & Sitemaps** | Respects `robots.txt` and processes discovered sitemaps |
| **State Persistence** | Uses BadgerDB for state; supports resuming crawls via `crawl --resume` |
| **Graceful Shutdown** | Handles `SIGINT`/`SIGTERM` with proper cleanup |
| **HTTP Retries** | Exponential backoff with jitter for transient errors |
| **Observability** | Structured logging (`log/slog`); optional `pprof` endpoint (build with `-tags pprof`) |
| **Modular Code** | Organized into packages for clarity and maintainability |
| **CLI Utilities** | Built-in `config validate` and `config list` commands for configuration management |
| **MCP Server Mode** | Expose as Model Context Protocol server for Claude Code/Cursor integration |
| **Full-Text Search** | Offline BM25 search over crawled docs (SQLite FTS5) via the `search_docs` MCP tool |
| **Auto Content Detection** | Automatic detection of 30+ documentation frameworks (Docusaurus, MkDocs, Sphinx, VitePress, GitBook, and more) with readability fallback |
| **Parallel Site Crawling** | Crawl multiple sites concurrently with shared resource management |
| **Watch Mode** | Scheduled periodic re-crawling with state persistence |
## Getting Started
### Prerequisites
- **Go:** Version 1.26 or later
- **Git:** For cloning the repository
- **Disk Space:** Sufficient for storing crawled content and state database
### Installation
**Option 1: Release Binaries**
Download the archive for your OS and architecture (Linux, macOS, Windows; amd64 and arm64) from the [Releases page](https://github.com/Sriram-PR/doc-scraper/releases/latest). Each release includes a `checksums.txt` to verify the download:
```bash
sha256sum --check --ignore-missing checksums.txt
```
Extract the archive and put `doc-scraper` on your `PATH`. The archive also ships a sample `config.yaml`.
**Option 2: Docker**
```bash
docker run --rm --user "$(id -u):$(id -g)" -v "$PWD":/data ghcr.io/sriram-pr/doc-scraper:latest crawl -site rust_cli_book
```
The image (linux/amd64 and linux/arm64) uses `/data` as its working directory and `doc-scraper` as its entrypoint, so the arguments after the image name are the subcommand. Mount a directory containing `config.yaml` at `/data` and point `output_base_dir` and `state_dir` at relative paths so crawl output lands in the mounted directory. The image runs as a non-root user, so `--user` makes the bind mount writable. Versioned tags (for example `2.9.2`) are published alongside `latest`.
**Option 3: Claude Desktop Extension**
Download `doc-scraper.mcpb` from the [latest release](https://github.com/Sriram-PR/doc-scraper/releases/latest) and open it in Claude Desktop. It prompts for the path to your `config.yaml` and runs the bundled binary as an MCP server (see [MCP Server Mode](#mcp-server-mode)).
**Option 4: Go Install**
Install the latest version directly from GitHub:
```bash
go install github.com/Sriram-PR/doc-scraper/v2/cmd/doc-scraper@latest
```
This installs the `doc-scraper` binary to your `GOPATH/bin` directory (usually `~/go/bin` or `%USERPROFILE%\go\bin`). Make sure this directory is in your `PATH`.
**Option 5: Clone and Build**
1. **Clone the repository:**
```bash
git clone https://github.com/Sriram-PR/doc-scraper.git
cd doc-scraper
```
2. **Install Dependencies:**
```bash
go mod tidy
```
3. **Build the Binary:**
```bash
make build
# or: go build -o doc-scraper ./cmd/doc-scraper
```
This creates an executable named `doc-scraper` in the project root.
### Quick Start
Create a minimal `config.yaml` in the project root:
```yaml
output_base_dir: "./crawled_docs"
state_dir: "./crawler_state"
enable_jsonl_output: true
sites:
rust_cli_book:
start_urls:
- "https://rust-cli.github.io/book/index.html"
allowed_domain: "rust-cli.github.io"
allowed_path_prefix: "/book/"
content_selector: "#content, main"
max_depth: 2 # seed plus one level; set 0 for the whole book
```
Run the crawl:
```bash
./doc-scraper crawl -site rust_cli_book -loglevel info
```
The Markdown, plus `pages.jsonl`, `llms.txt`, and `llms-full.txt`, lands under `./crawled_docs/rust_cli_book/` (output is organized by site key). A small book like this finishes in a few seconds; large sites can take minutes, so start with a low `max_depth` to gauge size before removing the bound.
## Configuration (`config.yaml`)
A `config.yaml` file is **required** to run the crawler. Create this file in the project root or specify its path using the `-config` flag.
### Key Settings for LLM Use
When configuring for LLM documentation processing, pay special attention to these settings:
- `sites.<your_site_key>.content_selector`: Define precisely to capture only relevant text
- `sites.<your_site_key>.allowed_domain` / `allowed_path_prefix`: Define scope accurately
- `skip_images`: Images are **not** downloaded by default (text-first). Set to `false` globally or per-site to download and localize images for offline consumption
- Adjust concurrency/delay settings based on the target site and your resources
### Example Configuration
```yaml
# Global settings (applied if not overridden by site)
default_delay_per_host: 500ms
num_workers: 8
num_image_workers: 8
max_requests: 48
max_requests_per_host: 4
output_base_dir: "./crawled_docs"
state_dir: "./crawler_state"
max_retries: 4
initial_retry_delay: 1s
max_retry_delay: 30s
global_crawl_timeout: 0s
skip_images: true # Default. Set to false to download and localize images
max_image_size_bytes: 10485760 # 10 MiB (applies only when images are downloaded)
enable_jsonl_output: true
jsonl_output_filename: "pages.jsonl"
# HTTP Client Settings
http_client_settings:
timeout: 45s
max_idle_conns_per_host: 6
# Site-specific configurations
sites:
# Key used with -site flag
pytorch_docs:
start_urls:
- "https://pytorch.org/docs/stable/"
allowed_domain: "pytorch.org"
allowed_path_prefix: "/docs/stable/"
content_selector: "article.pytorch-article .body"
max_depth: 0 # 0 for unlimited depth
skip_images: false # Opt in to downloading images for this site
disallowed_path_patterns:
- "/docs/stable/.*/_modules/.*"
- '/docs/stable/.*\.html#.*'
tensorflow_docs:
start_urls:
- "https://www.tensorflow.org/guide"
- "https://www.tensorflow.org/tutorials"
allowed_domain: "www.tensorflow.org"
allowed_path_prefix: "/"
content_selector: ".devsite-article-body"
max_depth: 0
delay_per_host: 1s # Site-specific override
# Disable JSONL output for this site, overriding global
enable_jsonl_output: false
disallowed_path_patterns:
- "/install/.*"
- "/js/.*"
```
### FulWhat people ask about doc-scraper
What is Sriram-PR/doc-scraper?
+
Sriram-PR/doc-scraper is mcp servers for the Claude AI ecosystem. Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data). It has 103 GitHub stars and its last recorded update is dated 2026-10-05.
How do I install doc-scraper?
+
You can install doc-scraper by cloning the repository (https://github.com/Sriram-PR/doc-scraper) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is Sriram-PR/doc-scraper safe to use?
+
Our security agent has analyzed Sriram-PR/doc-scraper and assigned a Trust Score of 100/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains Sriram-PR/doc-scraper?
+
Sriram-PR/doc-scraper is maintained by Sriram-PR. The last recorded GitHub activity is dated 2026-10-05, with 1 open issues.
Are there alternatives to doc-scraper?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy doc-scraper to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/sriram-pr-doc-scraper)<a href="https://claudewave.com/repo/sriram-pr-doc-scraper"><img src="https://claudewave.com/api/badge/sriram-pr-doc-scraper" alt="Featured on ClaudeWave: Sriram-PR/doc-scraper" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.