EU-native web scraping for AI agents & developers. official mcp server for crawlbrulee: native scrape and map tools for your agent.
- ✓Open-source license (Apache-2.0)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add crawlbrulee-mcp -- npx -y @crawlbrulee/mcp{
"mcpServers": {
"crawlbrulee-mcp": {
"command": "npx",
"args": ["-y", "@crawlbrulee/mcp"],
"env": {
"CRAWLBRULEE_API_KEY": "<crawlbrulee_api_key>"
}
}
}
}CRAWLBRULEE_API_KEYMCP Servers overview
# 🍮 crawlbrulee mcp
[](https://www.npmjs.com/package/@crawlbrulee/mcp)
[](./LICENSE)
**EU-native web scraping for AI agents & developers.**
plug crawlbrulee into your agent. the official [mcp](https://modelcontextprotocol.io) server for [crawlbrulee](https://crawlbrulee.com) gives mcp-aware agents — Claude Code, Codex, Cursor, Claude Desktop — native tools to scrape pages, map sites, run background jobs, and check usage. one call turns any url into clean markdown, screenshots, metadata and links.
- **everything runs in the EU.** the fetch, the render, the cache and your result never leave EU servers. the proxy exit is the one hop you choose: pick an EU exit and nothing leaves at all. gdpr-aligned, with a data processing agreement.
- **output made for models.** markdown with the page chrome stripped and the links kept, ready for the prompt. full-page screenshots can come back as tiles sized for an image model.
- **the hard parts, handled.** headless Chrome when a page needs it, rotating proxies with country selection, automatic retries, ad and cookie-banner removal, caching, background jobs and signed webhooks.
- **start free.** 750 credits, no credit card.
**get a free api key** → [dashboard.crawlbrulee.com](https://dashboard.crawlbrulee.com)
the server:
- `npx`-runnable — zero install.
- wraps the [`@crawlbrulee/sdk`](https://www.npmjs.com/package/@crawlbrulee/sdk) under the hood; this mcp is just a thin protocol adapter.
- stdio transport for terminal-based agents.
- strict, fully-described tool schemas — agents see what every parameter does without reading docs.
this readme covers the mcp server itself — its tools and how to wire it into a host. for how the api behaves — endpoints, parameters, and error semantics — please see our
[api docs](https://crawlbrulee.com/docs).
---
## install
```bash
# Claude Code
claude mcp add crawlbrulee \
--env CRAWLBRULEE_API_KEY=cwbl_... \
-- npx -y @crawlbrulee/mcp
# Cursor — add to ~/.cursor/mcp.json:
{
"mcpServers": {
"crawlbrulee": {
"command": "npx",
"args": ["-y", "@crawlbrulee/mcp"],
"env": { "CRAWLBRULEE_API_KEY": "cwbl_..." }
}
}
}
```
the same pattern works for Codex, Claude Desktop, and any other host that
accepts a stdio mcp launch command — set `command: npx`, `args: ["-y",
"@crawlbrulee/mcp"]`, and forward `CRAWLBRULEE_API_KEY` via the env block.
## configuration
| env var | required | description |
| --------------------- | -------- | ------------------------------------------------------------------------------ |
| `CRAWLBRULEE_API_KEY` | yes | api key sent as `Authorization: Bearer …`. get one at https://crawlbrulee.com. |
the mcp reads the env var on first tool invocation — not at startup — so a
typo in your config surfaces as a clear tool-error message rather than the
server failing to come up. see
[authentication](https://crawlbrulee.com/docs/authentication) for how the api consumes keys.
---
## tools
### `scrape`
fetch a single url and return the requested content (markdown, cleaned
html, raw html, links, images, screenshot, page metadata).
**input** — only `url` is required; everything else has sane defaults.
```jsonc
{
"url": "https://example.com",
"extract": {
"markdown": true,
"links": true,
"screenshot": { "type": "full_page", "device_mode": "desktop" },
},
"require_js": false,
"proxy": "basic",
"cleanup": { "ads_and_popups": true, "exclude_selectors": ["nav", "footer"] },
"cache": { "max_age": 3600 },
"location": { "locale": "en-US", "country": "US" },
"zero_data_retention": false,
}
```
`zero_data_retention` (boolean, default `false`) keeps the result out of the shared cache; anything stored to deliver it is kept for 24 hours, then deleted. it adds 1 credit and must be enabled for your organization. `scrape_async` takes it too. see [zero data retention](https://crawlbrulee.com/docs/zero-data-retention).
**output** — full scrape result. page metadata (title, OG tags, etc.) is returned under `metadata`. extracted `images` are returned as absolute urls — query strings are preserved, and relative `src`s are resolved against the page url. screenshots are returned as signed download urls the agent can fetch separately. in rare cases a screenshot can't be captured: when you requested other outputs too, the `screenshot` field is simply left out while the rest is still returned — but a screenshot-only call that can't deliver errors instead (`unsupported_screenshot_output`, HTTP 422, when the content type can't be screenshotted) and isn't billed. the result also carries `page_status_code` and a top-level `response_meta.usage` block:
```jsonc
{
"url": "https://example.com",
"requested_url": "https://example.com",
"page_status_code": 200, // the site's own HTTP status for the final page
"markdown": "...",
"metadata": { "title": "Example Domain" },
"response_meta": {
"usage": {
// total_credit_cost = engine_credit_cost × proxy_multiplier + screenshot_slicing_credit_cost + zero_data_retention_credit_cost
"total_credit_cost": 1,
"engine_credit_cost": 1, // http 1, browser 3, screenshot 5, cache 0
"proxy_multiplier": 1, // basic 1, advanced 5
"screenshot_slicing_credit_cost": 0, // 1 when the screenshot was split into slices, otherwise 0
"zero_data_retention_credit_cost": 0, // 1 when zero_data_retention added its credit, otherwise 0
"engine": "http", // "http" | "browser" | "screenshot" | "cache"
"proxy": "basic", // resolved tier actually used: "basic" | "advanced" (never "auto")
},
},
}
```
**a page the site served is a result, not an error.** `page_status_code` is the HTTP status the site answered with for the final page, after redirects. a 404, 410 or 503 page comes back with its content and its status here, so check `page_status_code` before you trust the content: a 404 means the markdown is the site's "not found" page. 2xx and 4xx pages are billed, except 403, 407, 408, 429 and 451; 5xx pages are never billed. when the site can't be reached at all, the tool returns a `target_unreachable` error instead (see [errors](#errors)).
alongside `response_meta.usage`, the result surfaces any non-fatal `warnings` — stable string codes an agent can switch on. an outsized page is truncated rather than refused, and the code names which part was cut:
| code | what it means for the payload |
| ------------------------- | ----------------------------------------------------------------------------------------------------- |
| `screenshot_truncated` | the page was taller than the scrolling-capture height cap; the screenshot covers the top of the page. |
| `links_truncated` | the page had more than 30,000 links; the `links` array is cut at the cap and is incomplete. |
| `inline_images_truncated` | the page had more than 10,000 inline images; the `images` array is cut at the cap and is incomplete. |
| `raw_html_truncated` | the page body exceeded 10,000,000 characters; `raw_html` is cut at a tag boundary, never mid-tag. |
| `metadata_truncated` | the page head exceeded 2,000,000 characters; `metadata` can be missing tags that sat past the cut. |
and if you requested an extract that doesn't apply to the content type (e.g. `markdown` of a pdf), the field name comes back in an `unsupported_fields` list — with the rest of the payload still returned.
every input field, its default, and its constraints are documented under the [scrape endpoint](https://crawlbrulee.com/docs/scrape) — with [extraction](https://crawlbrulee.com/docs/scrape/extraction), [screenshots](https://crawlbrulee.com/docs/scrape/screenshots), [proxies & location](https://crawlbrulee.com/docs/proxies), and [caching](https://crawlbrulee.com/docs/scrape/caching) covering the individual blocks.
### `scrape_async`
submit a scrape job to run **asynchronously** and get back a `job_id` immediately, instead of holding the connection open. use this for long-running scrapes (heavy js rendering, full-page screenshots of long pages); for a quick one-shot fetch prefer the synchronous `scrape` tool. then poll `scrape_status` until the job is `done` and fetch the page with `scrape_result`.
takes the same input as `scrape` plus an optional per-job completion `webhook`:
```jsonc
{
"url": "https://example.com",
"extract": { "markdown": true },
"webhook": {
// Endpoint that receives one signed `scrape.complete` POST when the job
// finishes. http/https (HTTPS required in production), max 2048 chars.
"url": "https://hooks.example.com/cwbl",
// Opaque correlation object echoed back verbatim in the delivery's
// `data.metadata`. Serializes to at most 2048 bytes.
"metadata": { "ref": "order-42" },
},
}
```
**output** — `{ "job_id": "..." }`.
when a `webhook` is attached, we deliver a single signed `scrape.complete` POST to your endpoint once the job reaches a terminal state, with your `metadata` echoed under `data.metadata` and the job's usage under `data.response_meta.usage` — so you can react to completion (and track cost) without polling. verify the `X-Cwbl-Signature` header with the sdk's `verifyWebhookSignature` (configure the signing secret in the dashboard under account → webhooks).
the job lifecycle is documented under [async scrape](https://crawlbrulee.com/docs/scrape/async); the delivery contract and payload shape under [webhooks](https://crawlbrulee.com/docs/scrape/webhooks), with the signature scheme in [webhook verification](https://crawlbrulee.com/docs/webhook-verification).
### `scrape_status`
look up the current lifecycle status of an asWhat people ask about crawlbrulee-mcp
What is crawlbrulee/crawlbrulee-mcp?
+
crawlbrulee/crawlbrulee-mcp is mcp servers for the Claude AI ecosystem. EU-native web scraping for AI agents & developers. official mcp server for crawlbrulee: native scrape and map tools for your agent. It has 0 GitHub stars and its last recorded update is dated 2026-10-05.
How do I install crawlbrulee-mcp?
+
You can install crawlbrulee-mcp by cloning the repository (https://github.com/crawlbrulee/crawlbrulee-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is crawlbrulee/crawlbrulee-mcp safe to use?
+
Our security agent has analyzed crawlbrulee/crawlbrulee-mcp and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains crawlbrulee/crawlbrulee-mcp?
+
crawlbrulee/crawlbrulee-mcp is maintained by crawlbrulee. The last recorded GitHub activity is dated 2026-10-05, with 0 open issues.
Are there alternatives to crawlbrulee-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy crawlbrulee-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/crawlbrulee-crawlbrulee-mcp)<a href="https://claudewave.com/repo/crawlbrulee-crawlbrulee-mcp"><img src="https://claudewave.com/api/badge/crawlbrulee-crawlbrulee-mcp" alt="Featured on ClaudeWave: crawlbrulee/crawlbrulee-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.