Fast MCP server that turns any web page into clean Markdown, structured metadata and classified links for LLMs. SSRF-hardened, paginated, typed.
claude mcp add web-fetcher -- npx -y mcp-server-web-fetcher{
"mcpServers": {
"web-fetcher": {
"command": "npx",
"args": ["-y", "mcp-server-web-fetcher"]
}
}
}MCP Servers overview
# mcp-server-web-fetcher
[](https://github.com/vojtisprime11/mcp-server-web-fetcher/actions/workflows/ci.yml)
[](https://www.npmjs.com/package/mcp-server-web-fetcher)
[](LICENSE)
[](tsconfig.json)
[](https://nodejs.org)
[](https://modelcontextprotocol.io)
[](https://registry.modelcontextprotocol.io/)
A fast, dependency-light [Model Context Protocol](https://modelcontextprotocol.io) server that turns
any web page into something a language model can actually read: clean Markdown, structured metadata,
and a classified list of links.
Web pages are 90% chrome. Scripts, cookie banners, navigation, sidebars and tracking pixels burn
context and derail reasoning. This server strips all of that on the way in, so the model sees the
article and nothing else.
```
┌────────────┐ stdio/JSON-RPC ┌──────────────────────┐ HTTPS ┌──────────┐
│ MCP client │ ─────────────────► │ mcp-server-web- │ ────────► │ web page │
│ (Claude, │ ◄───────────────── │ fetcher │ ◄──────── │ │
│ Kiro, …) │ Markdown + JSON │ fetch → clean → md │ └──────────┘
└────────────┘ └──────────────────────┘
```
## Features
- **Clean Markdown output.** Readability-style main-content detection, GFM tables, language-tagged
code fences, absolute links. No leftover HTML, even for gnarly nested tables.
- **Context-window aware.** Long pages are paginated with `startIndex` / `nextStartIndex` instead of
being silently cut in half.
- **Structured metadata.** Title, description, canonical, `lang`, author, publish/modify dates,
Open Graph, Twitter cards, JSON-LD, `hreflang` alternates, RSS/Atom feeds, h1–h6 outline and raw
HTTP headers.
- **Link intelligence.** Absolute URLs, anchor text, `rel`, `nofollow`, internal vs external
classification, scope filters, de-duplication.
- **SSRF hardening.** Loopback, private, link-local and cloud-metadata addresses are blocked on
_every_ redirect hop, not just the first request. Credentials in URLs are stripped.
- **Predictable failures.** Every error carries a stable machine-readable `code`
(`TIMEOUT`, `HTTP_ERROR`, `BLOCKED_HOST`, `RESPONSE_TOO_LARGE`, …), a retryability flag and a
recovery hint, so the model can self-correct instead of guessing.
- **Typed end to end.** Strict Zod input schemas plus published `outputSchema`, so clients get
validated `structuredContent`, not prose they have to re-parse.
- **Well behaved.** Byte caps, timeouts, redirect limits, charset sniffing (including legacy
encodings), a small TTL cache, transient-failure retries and optional `robots.txt` enforcement.
- **115 tests, no network required.** Vitest unit tests plus in-memory MCP protocol tests.
## Quickstart
Requires Node.js 20.18.1 or newer (inherited from cheerio, which needs undici 7).
```bash
# run it without installing
npx -y mcp-server-web-fetcher
# or install globally
npm install -g mcp-server-web-fetcher
mcp-server-web-fetcher
```
The server speaks MCP over stdio, so on its own it just waits for a client. Point a client at it:
### Claude Desktop
`~/Library/Application Support/Claude/claude_desktop_config.json` on macOS,
`%APPDATA%\Claude\claude_desktop_config.json` on Windows:
```json
{
"mcpServers": {
"web-fetcher": {
"command": "npx",
"args": ["-y", "mcp-server-web-fetcher"]
}
}
}
```
With configuration, using a global install:
```json
{
"mcpServers": {
"web-fetcher": {
"command": "mcp-server-web-fetcher",
"env": {
"WEB_FETCHER_TIMEOUT_MS": "20000",
"WEB_FETCHER_RESPECT_ROBOTS": "true"
}
}
}
}
```
Restart Claude Desktop, then ask it to summarise a URL.
### Claude Code
```bash
claude mcp add web-fetcher -- npx -y mcp-server-web-fetcher
```
### Kiro, Cursor, Windsurf and other MCP clients
Same shape, in `.kiro/settings/mcp.json` / the client's MCP config file:
```json
{
"mcpServers": {
"web-fetcher": {
"command": "npx",
"args": ["-y", "mcp-server-web-fetcher"],
"disabled": false,
"autoApprove": ["fetch_page_markdown", "extract_metadata", "extract_links"]
}
}
}
```
### MCP Registry
The server is listed in the [official MCP Registry](https://registry.modelcontextprotocol.io/)
as `io.github.vojtisprime11/web-fetcher`, so clients that read the registry can discover it
directly:
```bash
curl "https://registry.modelcontextprotocol.io/v0.1/servers?search=io.github.vojtisprime11/web-fetcher"
```
### From source
```bash
git clone https://github.com/vojtisprime11/mcp-server-web-fetcher.git
cd mcp-server-web-fetcher
npm install
npm run build
node dist/index.js # or: npm run inspect
```
```json
{
"mcpServers": {
"web-fetcher": {
"command": "node",
"args": ["/absolute/path/to/mcp-server-web-fetcher/dist/index.js"]
}
}
}
```
## Tools
All three tools are read-only, idempotent and open-world (annotated as such in the protocol), and
all take an absolute `http(s)` URL.
### `fetch_page_markdown`
Downloads a page and returns clean, LLM-friendly Markdown.
| Parameter | Type | Default | Description |
| ----------------- | ------------------- | ------------- | -------------------------------------------------------------- |
| `url` | string | required | Absolute http(s) URL. |
| `maxLength` | integer 500–1000000 | `25000` | Markdown characters to return per call. |
| `startIndex` | integer ≥ 0 | `0` | Character offset; use `nextStartIndex` from the previous call. |
| `mainContentOnly` | boolean | `true` | Drop nav/header/footer/sidebar, keep the densest block. |
| `includeLinks` | boolean | `true` | Keep Markdown links (`false` inlines the text only). |
| `includeImages` | boolean | `false` | Keep images as ``. |
| `includeMetadata` | boolean | `true` | Attach a short metadata summary. |
| `timeoutMs` | integer 1000–120000 | env / `15000` | Per-request timeout. |
Request:
```json
{
"name": "fetch_page_markdown",
"arguments": {
"url": "https://example.com/blog/caching",
"maxLength": 8000,
"mainContentOnly": true,
"includeImages": false
}
}
```
`structuredContent`:
```json
{
"url": "https://example.com/blog/caching",
"requestedUrl": "https://example.com/blog/caching",
"status": 200,
"contentType": "text/html; charset=utf-8",
"title": "How caching works",
"markdown": "# How caching works\n\nCaching is the art of **not** doing work twice...",
"markdownLength": 7984,
"totalLength": 21874,
"startIndex": 0,
"endIndex": 7984,
"nextStartIndex": 7984,
"truncated": true,
"wordCount": 1203,
"bytesDownloaded": 148213,
"elapsedMs": 412,
"fromCache": false,
"redirects": [],
"metadata": {
"title": "How caching works",
"description": "A deep dive into HTTP caching.",
"canonical": "https://example.com/blog/caching",
"language": "en",
"author": "Ada Lovelace",
"publishedTime": "2026-01-15T09:00:00Z"
}
}
```
To read the rest, call again with `"startIndex": 7984`. Keep going while `nextStartIndex` is not
`null`.
### `extract_metadata`
Everything a model needs to classify a page, without spending context on its body.
| Parameter | Type | Default | Description |
| -------------------- | ------------------- | ------------- | ----------------------------------------------------------- |
| `url` | string | required | Absolute http(s) URL. |
| `includeJsonLd` | boolean | `true` | Include parsed JSON-LD blocks (malformed ones are skipped). |
| `includeHeadings` | boolean | `true` | Include the h1–h6 outline. |
| `includeHttpHeaders` | boolean | `true` | Include response headers, lower-cased. |
| `timeoutMs` | integer 1000–120000 | env / `15000` | Per-request timeout. |
```json
{
"name": "extract_metadata",
"arguments": { "url": "https://example.com/blog/caching", "includeJsonLd": true }
}
```
`structuredContent` (abridged):
```json
{
"url": "https://example.com/blog/caching",
"status": 200,
"charset": "utf-8",
"title": "How caching works",
"description": "A deep dive into HTTP caching.",
"canonical": "https://example.com/blog/caching",
"language": "en",
"author": "Ada Lovelace",
"publishedTime": "2026-01-15T09:00:00Z",
"modifiedTime": null,
"robots": "index, follow",
"favicon": "https://example.com/favicon.ico",
"openGraph": {
"title": "How caching works",
"type": "article",
"image": "https://example.com/img/cover.png"
},
"twitter": { "card": "summary_large_image", "site": "@example" },
"alternates": [{ "hreflang": "de", "href": "https://example.com/de/blog/caching" }],
"feeds": [
{ "title": "Feed", "href": "https://example.com/feed.xml", "type": "application/rss+xml" }
],
"jsonLd": [{ "@type": "Article", "headline": "How caching works" }],
"hWhat people ask about mcp-server-web-fetcher
What is vojtisprime11/mcp-server-web-fetcher?
+
vojtisprime11/mcp-server-web-fetcher is mcp servers for the Claude AI ecosystem. Fast MCP server that turns any web page into clean Markdown, structured metadata and classified links for LLMs. SSRF-hardened, paginated, typed. It has 0 GitHub stars and was last updated today.
How do I install mcp-server-web-fetcher?
+
You can install mcp-server-web-fetcher by cloning the repository (https://github.com/vojtisprime11/mcp-server-web-fetcher) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is vojtisprime11/mcp-server-web-fetcher safe to use?
+
vojtisprime11/mcp-server-web-fetcher has not been audited yet by our security agent. Review the original repository on GitHub before using it in production.
Who maintains vojtisprime11/mcp-server-web-fetcher?
+
vojtisprime11/mcp-server-web-fetcher is maintained by vojtisprime11. The last recorded GitHub activity is from today, with 0 open issues.
Are there alternatives to mcp-server-web-fetcher?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy mcp-server-web-fetcher to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/vojtisprime11-mcp-server-web-fetcher)<a href="https://claudewave.com/repo/vojtisprime11-mcp-server-web-fetcher"><img src="https://claudewave.com/api/badge/vojtisprime11-mcp-server-web-fetcher" alt="Featured on ClaudeWave: vojtisprime11/mcp-server-web-fetcher" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
The fastest path to AI-powered full stack observability, even for lean teams.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!