Skip to main content
ClaudeWave

Check whether a page is scrapeable before the scraper is written, and say when it cannot tell

MCP ServersOfficial Registry1 stars0 forks● RustMITUpdated today
ClaudeWave Trust Score
95/100
✓ Verified
Passed
  • ✓Open-source license (MIT)
  • ✓Actively maintained (<30d)
  • ✓Clear description
  • ✓Topics declared
  • ✓Documented (README)
Last scanned: 10/1/2026
Install in Claude Code / Claude Desktop
Method: NPX · scrape-le-mcp
Claude Code CLI
claude mcp add scrape-le -- npx -y scrape-le-mcp
claude_desktop_config.json (Claude Desktop)
{
  "mcpServers": {
    "scrape-le": {
      "command": "npx",
      "args": ["-y", "scrape-le-mcp"]
    }
  }
}
1. Run the command above in your terminal (Claude Code), or paste the JSON config into claude_desktop_config.json (Claude Desktop).
2. Replace any <placeholder> values with your API keys or paths.
3. Restart Claude. The MCP server and its tools appear automatically.
Use cases

MCP Servers overview

<p align="center">
  <img src="src/assets/images/icon.png" alt="Scrape-LE Logo" width="96" height="96"/>
</p>
<h1 align="center">Scrape-LE: Zero Hassle Scrapeability Checks</h1>
<p align="center">
  <b>Load a URL in headless Chromium and see what will block your scraper — before you write it</b><br/>
  <i>Anti-bot vendors, rate limits, robots.txt rules, login walls, console errors, screenshots</i>
</p>

<p align="center">
  <a href="https://marketplace.visualstudio.com/items?itemName=nolindnaidoo.scrape-le">
    <img src="https://img.shields.io/badge/Install%20from-VS%20Code-blue?style=for-the-badge&logo=visualstudiocode" alt="Install from VS Code Marketplace" />
  </a>
  <a href="https://open-vsx.org/extension/OffensiveEdge/scrape-le">
    <img src="https://img.shields.io/open-vsx/dt/OffensiveEdge/scrape-le?style=for-the-badge&label=Open%20VSX&color=blue" alt="Open VSX downloads" />
  </a>
  <a href="https://www.npmjs.com/package/scrape-le-mcp">
    <img src="https://img.shields.io/npm/v/scrape-le-mcp?style=for-the-badge&label=MCP%20server&color=blue&logo=npm" alt="scrape-le-mcp on npm" />
  </a>
  <a href="https://crates.io/crates/scrape-le">
    <img src="https://img.shields.io/crates/v/scrape-le?style=for-the-badge&label=Rust%20CLI&color=blue&logo=rust" alt="scrape-le on crates.io" />
  </a>
  <a href="https://letools.dev/tools/scrape-le">
    <img src="https://img.shields.io/badge/LE%20Tools-letools.dev-blue?style=for-the-badge" alt="LE Tools" />
  </a>
</p>

---

<p align="center">
  <img src="src/assets/images/demo.gif" alt="Scrapeability Check Demo" style="max-width: 100%; height: auto;" />
</p>

> **Useful?** A star or rating is how other developers find it —
> [★ GitHub](https://github.com/nolindnaidoo/scrape-le) ·
> [★ Open VSX](https://open-vsx.org/extension/OffensiveEdge/scrape-le/reviews) ·
> [★ Marketplace](https://marketplace.visualstudio.com/items?itemName=nolindnaidoo.scrape-le&ssr=false#review-details)

## What it does

Run `Scrape-LE: Check URL Scrapeability` (`Ctrl+Alt+S` / `Cmd+Alt+S`), enter a URL, and the page loads in a real headless Chromium. The report lands in the output channel: HTTP status, page title, load time, console errors, a full-page screenshot, and four detections. Works in VS Code and VS Code–based editors like Cursor and VSCodium (installable from Open VSX).

One-time setup: run `Scrape-LE: Setup Browser` to install Chromium (~130MB, into Playwright's browser cache).

## Install

| Where | What you get | Install |
|---|---|---|
| **VS Code** | The same check, in your editor, on a keystroke | [Marketplace](https://marketplace.visualstudio.com/items?itemName=nolindnaidoo.scrape-le) |
| **Cursor, VSCodium, Windsurf** | The same extension | [Open VSX](https://open-vsx.org/extension/OffensiveEdge/scrape-le) |
| **A terminal or a CI step** | The same run over a whole tree, with exit codes | `cargo install scrape-le` · [crates.io](https://crates.io/crates/scrape-le) |
| **Any MCP agent, via Node** | `analyze_robots_txt` over stdio | `npx scrape-le-mcp` · [npm](https://www.npmjs.com/package/scrape-le-mcp) |
| **Zed** | The MCP server as a context server | [add it by hand](https://zed.dev/docs/ai/mcp) *(no listing yet)* |

## Use it from an AI agent

The same engine runs as an [MCP](https://modelcontextprotocol.io) server, so an agent can call it directly instead of you running a command.

| Editor | How |
|---|---|
| **VS Code** 1.101+ | Nothing to install — the extension registers `analyze_robots_txt` with agent mode |
| **Zed** | No listing yet — [add the MCP server by hand](https://zed.dev/docs/ai/mcp) |
| **Claude Code** | `claude mcp add scrape-le -- npx -y scrape-le-mcp` |
| **Cursor, Windsurf, anything else** | point it at `npx scrape-le-mcp` |

```
analyze_robots_txt(content, path, agent?, maxResults?)
```

Given robots.txt contents and a path, reports whether the rules permit crawling it: the group naming `agent` when one does, otherwise the generic (`User-agent: *`) rules, with `agent` in the answer saying which. Plus the crawl delay, disallowed patterns and any sitemaps.

The server takes content and returns data — it reads no files and makes no network requests of its own. Published as [`scrape-le-mcp`](https://www.npmjs.com/package/scrape-le-mcp) on npm and as `io.github.nolindnaidoo/scrape-le` in the [MCP registry](https://registry.modelcontextprotocol.io).

<details>
<summary><b>Configuring it by hand</b> — any host with an MCP config file</summary>

Most hosts read a JSON config. Add one entry:

```json
{
  "mcpServers": {
    "scrape-le": {
      "command": "npx",
      "args": ["-y", "scrape-le-mcp"]
    }
  }
}
```

`-y` skips the install prompt on first run. Pin a version if you would rather not track releases — `scrape-le-mcp@2.3.0`.

Prefer not to go through `npx` on every launch? Install it once and point at the binary instead:

```bash
npm install -g scrape-le-mcp
```

```json
{
  "mcpServers": {
    "scrape-le": { "command": "scrape-le-mcp" }
  }
}
```

It speaks MCP over stdio and needs no environment variables, no API key and no configuration of its own. To check it before wiring it into anything:

```bash
echo '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | npx -y scrape-le-mcp
```

That prints the tool list and exits — if you see `analyze_robots_txt`, the server works.

</details>

## The CLI

The same check runs from a terminal or an agent loop: a Rust CLI in [`crate/`](crate/) of this repository, sharing one signature corpus with the extension — [`crate/signatures/`](crate/signatures/) and [`crate/fixtures/`](crate/fixtures/) — so CI fails if the two ever disagree about a URL.

```bash
scrape-le https://example.com/search   # JSON on stdout, summary on stderr
scrape-le --input urls.txt             # a batch, streamed as it completes
scrape-le mcp                          # the same check over MCP on stdio
```

The exit code is the answer: **0 clear · 1 a real no · 2 the question was malformed.** ## Detections

| Detection | How it works |
|---|---|
| Anti-bot vendors | Response headers, script sources, DOM elements, and window globals fingerprint Cloudflare (incl. Turnstile challenges), reCAPTCHA, hCaptcha, DataDome, and PerimeterX |
| Rate limiting | `X-RateLimit-*` / `RateLimit-*` / `Retry-After` response headers, plus HTTP 429 |
| robots.txt | Fetches `<origin>/robots.txt` and evaluates it against your URL with RFC 9309 semantics — the group naming `scrape-le.robotsTxt.agent` when set and named, otherwise `User-agent: *`, and the report says which answered; grouped agents, `Allow`/`Disallow` longest-match, `*` wildcards, `$` anchors, crawl-delay, sitemaps |
| Authentication | HTTP 401/403, login forms (password + username fields), auth keywords in page text, auth path segments in the final URL |

Honest limitations: signatures are best-effort fingerprints of public integration patterns — a detected widget means the page *can* challenge you, not that it will, and a clean result is not proof a site allows scraping. Agent-specific robots.txt groups are ignored (only the `*` rules are reported). Pages get up to 5 seconds to go network-idle after load, so content rendered later than that can be missed by the page-level detections.

## Commands

| Command | Description |
|---|---|
| `Scrape-LE: Check URL Scrapeability` (`Ctrl+Alt+S` / `Cmd+Alt+S`) | Prompt for a URL and run the full check |
| `Scrape-LE: Check Selected URL` | Run the check on the URL in the current selection (also in the right-click menu) |
| `Scrape-LE: Setup Browser` | Install or verify the Chromium browser |
| `Scrape-LE: Open Settings` | Open Scrape-LE settings |
| `Scrape-LE: Help & Troubleshooting` | Built-in documentation |

## Settings

| Setting | Default | Description |
|---|---|---|
| `scrape-le.browser.timeout` | `30000` | Page-load timeout in ms (5000–120000) |
| `scrape-le.browser.viewport.width` | `1280` | Viewport width |
| `scrape-le.browser.viewport.height` | `720` | Viewport height |
| `scrape-le.browser.userAgent` | `""` | Custom User-Agent (empty = Chromium default) |
| `scrape-le.retry.userAgents` | `false` | On a blocked or failed check, retry under common User-Agents and report which worked |
| `scrape-le.screenshot.enabled` | `true` | Save a full-page screenshot per check |
| `scrape-le.screenshot.path` | `.vscode/scrape-le` | Screenshot directory (workspace-relative or absolute) |
| `scrape-le.screenshot.format` | `png` | `png` or `jpeg` |
| `scrape-le.screenshot.quality` | `90` | JPEG quality 0–100 (ignored for png) |
| `scrape-le.checkConsoleErrors` | `true` | Capture console and page errors while loading |
| `scrape-le.detections.antiBot` | `true` | Anti-bot vendor detection |
| `scrape-le.detections.rateLimit` | `true` | Rate-limit detection |
| `scrape-le.detections.robotsTxt` | `true` | robots.txt fetch + evaluation |
| `scrape-le.detections.authentication` | `true` | Authentication-wall detection |
| `scrape-le.notificationsLevel` | `important` | `all` = every notification, `important` = warnings + errors, `silent` = errors only |
| `scrape-le.statusBar.enabled` | `true` | Show the status bar item |

## Languages

Twelve languages besides English:

German · Spanish · French · Indonesian · Italian · Japanese · Korean ·
Portuguese (Brazil) · Russian · Ukrainian · Vietnamese · Chinese (Simplified)

Both halves are covered — the manifest (command titles, setting names and
descriptions) and everything shown while the extension runs (notifications,
the status bar, quick-picks and prompts). The extension follows VS Code's
display language, so it matches whatever the editor is already set to; no
setting of its own.

## Privacy & security

- **Network access is the feature, and it is scoped.** A check talks to exactly two things: the URL you enter (loaded in headless Chromium, which fetches that page's own resources like any browser) and that origin's `/robots.txt`. Nothing is sent anywhere else — no telemetry, no analytics.
- **Screenshots stay local**, wr
anti-botcaptchaclicloudflaredeveloper-toolsheadlessletoolsmcpplaywrightrate-limitreachabilityrobots-txtrustscraperstatic-analysistypescriptvscodevscode-extensionweb-scraping

What people ask about scrape-le

What is nolindnaidoo/scrape-le?

+

nolindnaidoo/scrape-le is mcp servers for the Claude AI ecosystem. Check whether a page is scrapeable before the scraper is written, and say when it cannot tell It has 1 GitHub stars and its last recorded update is dated 2026-09-30.

How do I install scrape-le?

+

You can install scrape-le by cloning the repository (https://github.com/nolindnaidoo/scrape-le) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.

Is nolindnaidoo/scrape-le safe to use?

+

Our security agent has analyzed nolindnaidoo/scrape-le and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.

Who maintains nolindnaidoo/scrape-le?

+

nolindnaidoo/scrape-le is maintained by nolindnaidoo. The last recorded GitHub activity is dated 2026-09-30, with 0 open issues.

Are there alternatives to scrape-le?

+

Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.

Deploy scrape-le to your cloud

Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.

Maintain this repo? Add a badge to your README

Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.

Featured on ClaudeWave: nolindnaidoo/scrape-le
[![Featured on ClaudeWave](https://claudewave.com/api/badge/nolindnaidoo-scrape-le)](https://claudewave.com/repo/nolindnaidoo-scrape-le)
<a href="https://claudewave.com/repo/nolindnaidoo-scrape-le"><img src="https://claudewave.com/api/badge/nolindnaidoo-scrape-le" alt="Featured on ClaudeWave: nolindnaidoo/scrape-le" width="320" height="64" /></a>

More MCP Servers

scrape-le alternatives