Skip to main content
ClaudeWave

Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data).

MCP ServersRegistry oficial103 estrellas10 forks● GoApache-2.0Actualizado yesterday
ClaudeWave Trust Score
100/100
✓ Verified
Passed
  • ✓Open-source license (Apache-2.0)
  • ✓Actively maintained (<30d)
  • ✓Healthy fork ratio
  • ✓Clear description
  • ✓Topics declared
  • ✓Mature repo (>1y old)
Last scanned: 10/7/2026
Install in Claude Code / Claude Desktop
Method: Docker · ghcr.io/sriram-pr/doc-scraper
Claude Code CLI
claude mcp add doc-scraper -- docker run -i --rm ghcr.io/sriram-pr/doc-scraper
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "doc-scraper": {
      "command": "docker",
      "args": ["run", "-i", "--rm", "ghcr.io/sriram-pr/doc-scraper"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
Casos de uso

Resumen de MCP Servers

# LLM Documentation Scraper (`doc-scraper`)

[![Go Version](https://img.shields.io/github/go-mod/go-version/Sriram-PR/doc-scraper)](https://golang.org/)
[![Go Reference](https://pkg.go.dev/badge/github.com/Sriram-PR/doc-scraper/v2.svg)](https://pkg.go.dev/github.com/Sriram-PR/doc-scraper/v2)
[![License](https://img.shields.io/github/license/Sriram-PR/doc-scraper)](https://github.com/Sriram-PR/doc-scraper/blob/main/LICENSE)
[![Glama score](https://glama.ai/mcp/servers/Sriram-PR/doc-scraper/badges/score.svg)](https://glama.ai/mcp/servers/Sriram-PR/doc-scraper)

> A configurable, concurrent, and resumable web crawler written in Go. Specifically designed to scrape technical documentation websites, extract core content, convert it cleanly to Markdown format suitable for ingestion by Large Language Models (LLMs), and save the results locally.

![doc-scraper crawling a docs site and answering search queries offline](demo/demo.gif)

## Overview

This project provides a powerful command-line tool to crawl documentation sites based on settings defined in a `config.yaml` file. It navigates the site structure, extracts content from specified HTML sections using CSS selectors, and converts it into clean Markdown files.

### Why Use This Tool?

- **Built for LLM Training & RAG Systems** - Creates clean, consistent Markdown optimized for ingestion
- **Preserves Documentation Structure** - Maintains the original site hierarchy for context preservation
- **Production-Ready Features** - Offers resumable crawls, rate limiting, and graceful error handling
- **High Performance** - Uses Go's concurrency model for efficient parallel processing

## Goal: Preparing Documentation for LLMs

The main objective of this tool is to automate the often tedious process of gathering and cleaning web-based documentation for use with Large Language Models. By converting structured web content into clean Markdown, it aims to provide a dataset that is:

- **Text-Focused:** Prioritizes the textual content extracted via CSS selectors
- **Structured:** Maintains the directory hierarchy of the original documentation site, preserving context
- **Cleaned:** Converts HTML to Markdown, removing web-specific markup and clutter
- **Locally Accessible:** Provides the content as local files for easier processing and pipeline integration

## Key Features

| Feature | Description |
|---------|-------------|
| **Configurable Crawling** | Uses YAML for global and site-specific settings |
| **Scope Control** | Limits crawling by domain, path prefix, and disallowed path patterns (regex) |
| **Content Extraction** | Extracts main content using CSS selectors |
| **HTML-to-Markdown** | Converts extracted HTML to clean GitHub-Flavored Markdown (tables, task lists, strikethrough) |
| **Image Handling** | Opt-in downloading and local rewriting of image links with domain and size filtering (disabled by default; doc-scraper is text-first) |
| **Link Rewriting** | Rewrites internal links to relative paths for local structure |
| **JSONL Output** | Optional one-record-per-page JSONL with a trailing crawl-summary record, for RAG ingestion |
| **Concurrency** | Configurable worker pools and semaphore-based request limits (global and per-host) |
| **Rate Limiting** | Configurable per-host delays with jitter |
| **Robots.txt & Sitemaps** | Respects `robots.txt` and processes discovered sitemaps |
| **State Persistence** | Uses BadgerDB for state; supports resuming crawls via `crawl --resume` |
| **Graceful Shutdown** | Handles `SIGINT`/`SIGTERM` with proper cleanup |
| **HTTP Retries** | Exponential backoff with jitter for transient errors |
| **Observability** | Structured logging (`log/slog`); optional `pprof` endpoint (build with `-tags pprof`) |
| **Modular Code** | Organized into packages for clarity and maintainability |
| **CLI Utilities** | Built-in `config validate` and `config list` commands for configuration management |
| **MCP Server Mode** | Expose as Model Context Protocol server for Claude Code/Cursor integration |
| **Full-Text Search** | Offline BM25 search over crawled docs (SQLite FTS5) via the `search_docs` MCP tool |
| **Auto Content Detection** | Automatic detection of 30+ documentation frameworks (Docusaurus, MkDocs, Sphinx, VitePress, GitBook, and more) with readability fallback |
| **Parallel Site Crawling** | Crawl multiple sites concurrently with shared resource management |
| **Watch Mode** | Scheduled periodic re-crawling with state persistence |

## Getting Started

### Prerequisites

- **Go:** Version 1.26 or later
- **Git:** For cloning the repository
- **Disk Space:** Sufficient for storing crawled content and state database

### Installation

**Option 1: Release Binaries**

Download the archive for your OS and architecture (Linux, macOS, Windows; amd64 and arm64) from the [Releases page](https://github.com/Sriram-PR/doc-scraper/releases/latest). Each release includes a `checksums.txt` to verify the download:

```bash
sha256sum --check --ignore-missing checksums.txt
```

Extract the archive and put `doc-scraper` on your `PATH`. The archive also ships a sample `config.yaml`.

**Option 2: Docker**

```bash
docker run --rm --user "$(id -u):$(id -g)" -v "$PWD":/data ghcr.io/sriram-pr/doc-scraper:latest crawl -site rust_cli_book
```

The image (linux/amd64 and linux/arm64) uses `/data` as its working directory and `doc-scraper` as its entrypoint, so the arguments after the image name are the subcommand. Mount a directory containing `config.yaml` at `/data` and point `output_base_dir` and `state_dir` at relative paths so crawl output lands in the mounted directory. The image runs as a non-root user, so `--user` makes the bind mount writable. Versioned tags (for example `2.9.2`) are published alongside `latest`.

**Option 3: Claude Desktop Extension**

Download `doc-scraper.mcpb` from the [latest release](https://github.com/Sriram-PR/doc-scraper/releases/latest) and open it in Claude Desktop. It prompts for the path to your `config.yaml` and runs the bundled binary as an MCP server (see [MCP Server Mode](#mcp-server-mode)).

**Option 4: Go Install**

Install the latest version directly from GitHub:

```bash
go install github.com/Sriram-PR/doc-scraper/v2/cmd/doc-scraper@latest
```

This installs the `doc-scraper` binary to your `GOPATH/bin` directory (usually `~/go/bin` or `%USERPROFILE%\go\bin`). Make sure this directory is in your `PATH`.

**Option 5: Clone and Build**

1. **Clone the repository:**

   ```bash
   git clone https://github.com/Sriram-PR/doc-scraper.git
   cd doc-scraper
   ```

2. **Install Dependencies:**

   ```bash
   go mod tidy
   ```

3. **Build the Binary:**

   ```bash
   make build
   # or: go build -o doc-scraper ./cmd/doc-scraper
   ```

   This creates an executable named `doc-scraper` in the project root.

### Quick Start

Create a minimal `config.yaml` in the project root:

```yaml
output_base_dir: "./crawled_docs"
state_dir: "./crawler_state"
enable_jsonl_output: true
sites:
  rust_cli_book:
    start_urls:
      - "https://rust-cli.github.io/book/index.html"
    allowed_domain: "rust-cli.github.io"
    allowed_path_prefix: "/book/"
    content_selector: "#content, main"
    max_depth: 2          # seed plus one level; set 0 for the whole book
```

Run the crawl:

```bash
./doc-scraper crawl -site rust_cli_book -loglevel info
```

The Markdown, plus `pages.jsonl`, `llms.txt`, and `llms-full.txt`, lands under `./crawled_docs/rust_cli_book/` (output is organized by site key). A small book like this finishes in a few seconds; large sites can take minutes, so start with a low `max_depth` to gauge size before removing the bound.

## Configuration (`config.yaml`)

A `config.yaml` file is **required** to run the crawler. Create this file in the project root or specify its path using the `-config` flag.

### Key Settings for LLM Use

When configuring for LLM documentation processing, pay special attention to these settings:

- `sites.<your_site_key>.content_selector`: Define precisely to capture only relevant text
- `sites.<your_site_key>.allowed_domain` / `allowed_path_prefix`: Define scope accurately
- `skip_images`: Images are **not** downloaded by default (text-first). Set to `false` globally or per-site to download and localize images for offline consumption
- Adjust concurrency/delay settings based on the target site and your resources

### Example Configuration

```yaml
# Global settings (applied if not overridden by site)
default_delay_per_host: 500ms
num_workers: 8
num_image_workers: 8
max_requests: 48
max_requests_per_host: 4
output_base_dir: "./crawled_docs"
state_dir: "./crawler_state"
max_retries: 4
initial_retry_delay: 1s
max_retry_delay: 30s
global_crawl_timeout: 0s
skip_images: true # Default. Set to false to download and localize images
max_image_size_bytes: 10485760 # 10 MiB (applies only when images are downloaded)
enable_jsonl_output: true
jsonl_output_filename: "pages.jsonl"

# HTTP Client Settings
http_client_settings:
  timeout: 45s
  max_idle_conns_per_host: 6

# Site-specific configurations
sites:
  # Key used with -site flag
  pytorch_docs:
    start_urls:
      - "https://pytorch.org/docs/stable/"
    allowed_domain: "pytorch.org"
    allowed_path_prefix: "/docs/stable/"
    content_selector: "article.pytorch-article .body"
    max_depth: 0 # 0 for unlimited depth
    skip_images: false # Opt in to downloading images for this site
    disallowed_path_patterns:
      - "/docs/stable/.*/_modules/.*"
      - '/docs/stable/.*\.html#.*'

  tensorflow_docs:
    start_urls:
      - "https://www.tensorflow.org/guide"
      - "https://www.tensorflow.org/tutorials"
    allowed_domain: "www.tensorflow.org"
    allowed_path_prefix: "/"
    content_selector: ".devsite-article-body"
    max_depth: 0
    delay_per_host: 1s  # Site-specific override
    # Disable JSONL output for this site, overriding global
    enable_jsonl_output: false
    disallowed_path_patterns:
      - "/install/.*"
      - "/js/.*"
```

### Ful
data-preparationdocumentationgolang-clillmmcpmcp-serverweb-crawlerweb-scraper

Lo que la gente pregunta sobre doc-scraper

¿Qué es Sriram-PR/doc-scraper?

+

Sriram-PR/doc-scraper es mcp servers para el ecosistema de Claude AI. Go web crawler to scrape documentation sites and convert content to clean Markdown for LLM ingestion (RAG, training data). Tiene 103 estrellas en GitHub y su última actualización registrada es del 2026-10-05.

¿Cómo se instala doc-scraper?

+

Puedes instalar doc-scraper clonando el repositorio (https://github.com/Sriram-PR/doc-scraper) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.

¿Es seguro usar Sriram-PR/doc-scraper?

+

Nuestro agente de seguridad ha analizado Sriram-PR/doc-scraper y le ha asignado un Trust Score de 100/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.

¿Quién mantiene Sriram-PR/doc-scraper?

+

Sriram-PR/doc-scraper es mantenido por Sriram-PR. La última actividad registrada en GitHub es del 2026-10-05, con 1 issues abiertos.

¿Hay alternativas a doc-scraper?

+

Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.

Despliega doc-scraper en tu cloud

Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.

¿Mantienes este repo? Añade un badge a tu README

Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.

Featured on ClaudeWave: Sriram-PR/doc-scraper
[![Featured on ClaudeWave](https://claudewave.com/api/badge/sriram-pr-doc-scraper)](https://claudewave.com/repo/sriram-pr-doc-scraper)
<a href="https://claudewave.com/repo/sriram-pr-doc-scraper"><img src="https://claudewave.com/api/badge/sriram-pr-doc-scraper" alt="Featured on ClaudeWave: Sriram-PR/doc-scraper" width="320" height="64" /></a>

Más MCP Servers

Alternativas a doc-scraper