Skip to main content
ClaudeWave

EU-native web scraping for AI agents & developers. official mcp server for crawlbrulee: native scrape and map tools for your agent.

MCP ServersRegistry oficial0 estrellas0 forks● TypeScriptApache-2.0Actualizado today
ClaudeWave Trust Score
95/100
✓ Verified
Passed
  • ✓Open-source license (Apache-2.0)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/6/2026
Install in Claude Code / Claude Desktop
Method: NPX · @crawlbrulee/mcp
Claude Code CLI
claude mcp add crawlbrulee-mcp -- npx -y @crawlbrulee/mcp
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "crawlbrulee-mcp": {
      "command": "npx",
      "args": ["-y", "@crawlbrulee/mcp"],
      "env": {
        "CRAWLBRULEE_API_KEY": "<crawlbrulee_api_key>"
      }
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
Detected environment variables
CRAWLBRULEE_API_KEY
Casos de uso

Resumen de MCP Servers

# 🍮 crawlbrulee mcp

[![npm](https://img.shields.io/npm/v/@crawlbrulee/mcp?style=flat-square&label=npm)](https://www.npmjs.com/package/@crawlbrulee/mcp)
[![license](https://img.shields.io/npm/l/@crawlbrulee/mcp?style=flat-square&label=license)](./LICENSE)

**EU-native web scraping for AI agents & developers.**

plug crawlbrulee into your agent. the official [mcp](https://modelcontextprotocol.io) server for [crawlbrulee](https://crawlbrulee.com) gives mcp-aware agents — Claude Code, Codex, Cursor, Claude Desktop — native tools to scrape pages, map sites, run background jobs, and check usage. one call turns any url into clean markdown, screenshots, metadata and links.

- **everything runs in the EU.** the fetch, the render, the cache and your result never leave EU servers. the proxy exit is the one hop you choose: pick an EU exit and nothing leaves at all. gdpr-aligned, with a data processing agreement.
- **output made for models.** markdown with the page chrome stripped and the links kept, ready for the prompt. full-page screenshots can come back as tiles sized for an image model.
- **the hard parts, handled.** headless Chrome when a page needs it, rotating proxies with country selection, automatic retries, ad and cookie-banner removal, caching, background jobs and signed webhooks.
- **start free.** 750 credits, no credit card.

**get a free api key** → [dashboard.crawlbrulee.com](https://dashboard.crawlbrulee.com)

the server:

- `npx`-runnable — zero install.
- wraps the [`@crawlbrulee/sdk`](https://www.npmjs.com/package/@crawlbrulee/sdk) under the hood; this mcp is just a thin protocol adapter.
- stdio transport for terminal-based agents.
- strict, fully-described tool schemas — agents see what every parameter does without reading docs.

this readme covers the mcp server itself — its tools and how to wire it into a host. for how the api behaves — endpoints, parameters, and error semantics — please see our
[api docs](https://crawlbrulee.com/docs).

---

## install

```bash
# Claude Code
claude mcp add crawlbrulee \
  --env CRAWLBRULEE_API_KEY=cwbl_... \
  -- npx -y @crawlbrulee/mcp

# Cursor — add to ~/.cursor/mcp.json:
{
  "mcpServers": {
    "crawlbrulee": {
      "command": "npx",
      "args": ["-y", "@crawlbrulee/mcp"],
      "env": { "CRAWLBRULEE_API_KEY": "cwbl_..." }
    }
  }
}
```

the same pattern works for Codex, Claude Desktop, and any other host that
accepts a stdio mcp launch command — set `command: npx`, `args: ["-y",
"@crawlbrulee/mcp"]`, and forward `CRAWLBRULEE_API_KEY` via the env block.

## configuration

| env var               | required | description                                                                    |
| --------------------- | -------- | ------------------------------------------------------------------------------ |
| `CRAWLBRULEE_API_KEY` | yes      | api key sent as `Authorization: Bearer …`. get one at https://crawlbrulee.com. |

the mcp reads the env var on first tool invocation — not at startup — so a
typo in your config surfaces as a clear tool-error message rather than the
server failing to come up. see
[authentication](https://crawlbrulee.com/docs/authentication) for how the api consumes keys.

---

## tools

### `scrape`

fetch a single url and return the requested content (markdown, cleaned
html, raw html, links, images, screenshot, page metadata).

**input** — only `url` is required; everything else has sane defaults.

```jsonc
{
  "url": "https://example.com",
  "extract": {
    "markdown": true,
    "links": true,
    "screenshot": { "type": "full_page", "device_mode": "desktop" },
  },
  "require_js": false,
  "proxy": "basic",
  "cleanup": { "ads_and_popups": true, "exclude_selectors": ["nav", "footer"] },
  "cache": { "max_age": 3600 },
  "location": { "locale": "en-US", "country": "US" },
  "zero_data_retention": false,
}
```

`zero_data_retention` (boolean, default `false`) keeps the result out of the shared cache; anything stored to deliver it is kept for 24 hours, then deleted. it adds 1 credit and must be enabled for your organization. `scrape_async` takes it too. see [zero data retention](https://crawlbrulee.com/docs/zero-data-retention).

**output** — full scrape result. page metadata (title, OG tags, etc.) is returned under `metadata`. extracted `images` are returned as absolute urls — query strings are preserved, and relative `src`s are resolved against the page url. screenshots are returned as signed download urls the agent can fetch separately. in rare cases a screenshot can't be captured: when you requested other outputs too, the `screenshot` field is simply left out while the rest is still returned — but a screenshot-only call that can't deliver errors instead (`unsupported_screenshot_output`, HTTP 422, when the content type can't be screenshotted) and isn't billed. the result also carries `page_status_code` and a top-level `response_meta.usage` block:

```jsonc
{
  "url": "https://example.com",
  "requested_url": "https://example.com",
  "page_status_code": 200, // the site's own HTTP status for the final page
  "markdown": "...",
  "metadata": { "title": "Example Domain" },
  "response_meta": {
    "usage": {
      // total_credit_cost = engine_credit_cost × proxy_multiplier + screenshot_slicing_credit_cost + zero_data_retention_credit_cost
      "total_credit_cost": 1,
      "engine_credit_cost": 1, // http 1, browser 3, screenshot 5, cache 0
      "proxy_multiplier": 1, // basic 1, advanced 5
      "screenshot_slicing_credit_cost": 0, // 1 when the screenshot was split into slices, otherwise 0
      "zero_data_retention_credit_cost": 0, // 1 when zero_data_retention added its credit, otherwise 0
      "engine": "http", // "http" | "browser" | "screenshot" | "cache"
      "proxy": "basic", // resolved tier actually used: "basic" | "advanced" (never "auto")
    },
  },
}
```

**a page the site served is a result, not an error.** `page_status_code` is the HTTP status the site answered with for the final page, after redirects. a 404, 410 or 503 page comes back with its content and its status here, so check `page_status_code` before you trust the content: a 404 means the markdown is the site's "not found" page. 2xx and 4xx pages are billed, except 403, 407, 408, 429 and 451; 5xx pages are never billed. when the site can't be reached at all, the tool returns a `target_unreachable` error instead (see [errors](#errors)).

alongside `response_meta.usage`, the result surfaces any non-fatal `warnings` — stable string codes an agent can switch on. an outsized page is truncated rather than refused, and the code names which part was cut:

| code                      | what it means for the payload                                                                         |
| ------------------------- | ----------------------------------------------------------------------------------------------------- |
| `screenshot_truncated`    | the page was taller than the scrolling-capture height cap; the screenshot covers the top of the page. |
| `links_truncated`         | the page had more than 30,000 links; the `links` array is cut at the cap and is incomplete.           |
| `inline_images_truncated` | the page had more than 10,000 inline images; the `images` array is cut at the cap and is incomplete.  |
| `raw_html_truncated`      | the page body exceeded 10,000,000 characters; `raw_html` is cut at a tag boundary, never mid-tag.     |
| `metadata_truncated`      | the page head exceeded 2,000,000 characters; `metadata` can be missing tags that sat past the cut.    |

and if you requested an extract that doesn't apply to the content type (e.g. `markdown` of a pdf), the field name comes back in an `unsupported_fields` list — with the rest of the payload still returned.

every input field, its default, and its constraints are documented under the [scrape endpoint](https://crawlbrulee.com/docs/scrape) — with [extraction](https://crawlbrulee.com/docs/scrape/extraction), [screenshots](https://crawlbrulee.com/docs/scrape/screenshots), [proxies & location](https://crawlbrulee.com/docs/proxies), and [caching](https://crawlbrulee.com/docs/scrape/caching) covering the individual blocks.

### `scrape_async`

submit a scrape job to run **asynchronously** and get back a `job_id` immediately, instead of holding the connection open. use this for long-running scrapes (heavy js rendering, full-page screenshots of long pages); for a quick one-shot fetch prefer the synchronous `scrape` tool. then poll `scrape_status` until the job is `done` and fetch the page with `scrape_result`.

takes the same input as `scrape` plus an optional per-job completion `webhook`:

```jsonc
{
  "url": "https://example.com",
  "extract": { "markdown": true },
  "webhook": {
    // Endpoint that receives one signed `scrape.complete` POST when the job
    // finishes. http/https (HTTPS required in production), max 2048 chars.
    "url": "https://hooks.example.com/cwbl",
    // Opaque correlation object echoed back verbatim in the delivery's
    // `data.metadata`. Serializes to at most 2048 bytes.
    "metadata": { "ref": "order-42" },
  },
}
```

**output** — `{ "job_id": "..." }`.

when a `webhook` is attached, we deliver a single signed `scrape.complete` POST to your endpoint once the job reaches a terminal state, with your `metadata` echoed under `data.metadata` and the job's usage under `data.response_meta.usage` — so you can react to completion (and track cost) without polling. verify the `X-Cwbl-Signature` header with the sdk's `verifyWebhookSignature` (configure the signing secret in the dashboard under account → webhooks).

the job lifecycle is documented under [async scrape](https://crawlbrulee.com/docs/scrape/async); the delivery contract and payload shape under [webhooks](https://crawlbrulee.com/docs/scrape/webhooks), with the signature scheme in [webhook verification](https://crawlbrulee.com/docs/webhook-verification).

### `scrape_status`

look up the current lifecycle status of an as
ai-agentseugdprllmmarkdownmcpmcp-servermodel-context-protocolscreenshotsweb-scraping

Lo que la gente pregunta sobre crawlbrulee-mcp

¿Qué es crawlbrulee/crawlbrulee-mcp?

+

crawlbrulee/crawlbrulee-mcp es mcp servers para el ecosistema de Claude AI. EU-native web scraping for AI agents & developers. official mcp server for crawlbrulee: native scrape and map tools for your agent. Tiene 0 estrellas en GitHub y su última actualización registrada es del 2026-10-05.

¿Cómo se instala crawlbrulee-mcp?

+

Puedes instalar crawlbrulee-mcp clonando el repositorio (https://github.com/crawlbrulee/crawlbrulee-mcp) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.

¿Es seguro usar crawlbrulee/crawlbrulee-mcp?

+

Nuestro agente de seguridad ha analizado crawlbrulee/crawlbrulee-mcp y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.

¿Quién mantiene crawlbrulee/crawlbrulee-mcp?

+

crawlbrulee/crawlbrulee-mcp es mantenido por crawlbrulee. La última actividad registrada en GitHub es del 2026-10-05, con 0 issues abiertos.

¿Hay alternativas a crawlbrulee-mcp?

+

Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.

Despliega crawlbrulee-mcp en tu cloud

Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.

¿Mantienes este repo? Añade un badge a tu README

Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.

Featured on ClaudeWave: crawlbrulee/crawlbrulee-mcp
[![Featured on ClaudeWave](https://claudewave.com/api/badge/crawlbrulee-crawlbrulee-mcp)](https://claudewave.com/repo/crawlbrulee-crawlbrulee-mcp)
<a href="https://claudewave.com/repo/crawlbrulee-crawlbrulee-mcp"><img src="https://claudewave.com/api/badge/crawlbrulee-crawlbrulee-mcp" alt="Featured on ClaudeWave: crawlbrulee/crawlbrulee-mcp" width="320" height="64" /></a>

Más MCP Servers

Alternativas a crawlbrulee-mcp