Arquivo.pt MCP — full-text search over the Portuguese web archive.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
claude mcp add mcp-arquivo-pt -- npx -y @pipeworx/mcp-arquivo-pt{
"mcpServers": {
"mcp-arquivo-pt": {
"command": "npx",
"args": ["-y", "@pipeworx/mcp-arquivo-pt"]
}
}
}MCP Servers overview
# @pipeworx/arquivo-pt
Full-text search of archived web pages in Arquivo.pt, the Portuguese web
archive run by FCCN (captures since 1996). Where the Internet Archive's
Wayback Machine (pack `wayback`) needs the URL, Arquivo.pt indexes the text
of what it captured, so a caller can find archived pages by phrase, then read
a capture's extracted text or list every preserved version of a URL.
Part of [Pipeworx](https://pipeworx.io) — an MCP gateway connecting AI agents to 1743+ live data sources. This is an independent, unofficial integration — not affiliated with, endorsed by, or published by the upstream provider.
## Tools
- `arquivo_search_pages(query, from?, to?, site?, type?, limit?, offset?, per_site?)`
— full-text hits for terms or a quoted phrase: title, original URL, capture
date, decoded snippet, archived replay URL, extracted-text and screenshot
links, plus `estimated_total` and `next_offset`. Zero hits answer
`found:false` with `reason:"no_match"` and a hint.
- `arquivo_url_history(url, from?, to?, limit?, offset?)` — every preserved
capture of a domain, host or full URL, newest first, with capture timestamp,
crawl HTTP status, MIME type, size, digest and replay links. Zero captures
answer `found:false` with `reason:"no_captures"`.
- `arquivo_page_text(url, timestamp, max_chars?)` — the plain text Arquivo.pt
extracted from one capture (`original_url` + `captured_at_ts` of a hit).
A capture that does not exist answers `found:false, reason:"capture_not_found"`.
Every response carries `source` (the exact upstream URL) and `data_as_of`.
A transport failure, a non-JSON body or a changed response shape throws a loud
error naming Arquivo.pt and the HTTP status — never an empty list.
## Auth
Keyless.
## Coverage, honestly
- Centred on the Portuguese web (`.pt` sites and Portuguese-language pages),
but international pages linked from them are captured too. Probed 2026-10-08:
"climate change" ≈ 52-53M estimated hits (academia.edu, eea.europa.eu,
ipcc.ch), "quantum computing" ≈ 0.8-1.1M (techtarget.com, microsoft.com,
en.wikipedia.org), "Federal Reserve interest rates" ≈ 1.8-1.9M
(federalreserve.gov, vox.com). The estimate varies by index node. English
queries work; the top hits skew to pages Portuguese sites cite.
- **The full-text index lags the crawl by years.** Probed 2026-10-08, a
search restricted to `from=2022` returns 0 hits on most calls, one node
answered with 2023 captures, and `arquivo_url_history` lists captures from
mid-2026. Use the URL history for anything recent.
- **Arquivo.pt load-balances across index nodes that disagree.** From a live
Cloudflare Worker on 2026-10-08 the same URL returned, minutes apart: real
hits in 30-37 s; real hits in 1-6 s with `FAWP*` collection ids and a
`hostKey` field; `response_items: []` with `estimated_nr_results` 3.5M;
and `response_items: [{}, {}, {}]` (429 bytes, every row empty) — the last
one for every unquoted multi-word query for a stretch, while a laptop got
real hits for the identical URL and a quoted phrase answered correctly from
the Worker. The pack retries once with `maxItems` nudged and then throws an
error naming the fault; rows with no URL/timestamp are never returned as
hits (`dropped_malformed_rows` counts them). Some nodes also ignore
`maxItems` (9 rows for 5); hits are capped at `limit`.
- The pack sends `to` only when you pass one. The upstream default (the
previous calendar year) already covers the whole index, and forcing
`to=<now>` made the same query return zero items from a live Worker while
`estimated_nr_results` stayed at 3.8M (probed twice, 2026-10-08).
- Hits are deduplicated to 2 per site by default (`per_site`); raise it to see
more captures of one site.
## Data sources
- <https://arquivo.pt/textsearch?q=…> — full-text search (`q`, `from`, `to`,
`siteSearch`, `type`, `maxItems` ≤ 500, `offset`, `dedupValue`, `dedupField`).
- <https://arquivo.pt/textsearch?versionHistory=…> — URL version history.
- <https://arquivo.pt/textextracted?m=…> — extracted text
of one capture. The `m` value is the original URL followed by `/` and the
14-digit timestamp, so a URL ending in `/` yields `//` before the stamp —
that is correct, do not "fix" it.
- API reference: <https://github.com/arquivo/pwa-technologies/wiki/Arquivo.pt-API>.
A URL passed as `q` is an HTTP 400 upstream; the pack refuses it first and
points at `arquivo_url_history`.
- Snippets come back HTML-escaped with Latin-1 named entities
(`Inteligência`); the pack decodes them.
## Quick Start
Add to your MCP client (Claude Desktop, Cursor, Windsurf, etc.):
```json
{
"mcpServers": {
"arquivo-pt": {
"url": "https://gateway.pipeworx.io/arquivo-pt/mcp"
}
}
}
```
### What this endpoint actually serves
`tools/list` at `https://gateway.pipeworx.io/arquivo-pt/mcp` returns the tools in the table
above **plus the shared Pipeworx meta-tools** — `ask_pipeworx`,
`discover_tools`, `search_within`, `remember`/`recall` and the rest of the
gateway-wide set. So the tool count you see is larger than this table: a
single-pack endpoint currently lists roughly 30 shared tools alongside the
pack's own. The connection's `initialize` response states its exact scope, and
is the authoritative answer for a given day.
This is deliberate, not multiplexing by accident. The meta-tools are what let a
scoped connection answer a question this pack does not cover — via
`ask_pipeworx`, which routes across the whole catalog — without you adding a
second MCP server. There is currently no way to mount a pack endpoint without
them; if the extra schemas cost you more context than the routing is worth,
connect to the full gateway once rather than to several pack endpoints.
Or connect to the full Pipeworx gateway to get every pack's tools listed
directly, instead of just this one's:
```json
{
"mcpServers": {
"pipeworx": {
"url": "https://gateway.pipeworx.io/mcp"
}
}
}
```
Both URLs reach the same gateway and the same 1743+ data sources. The
only difference is which pack's tools are listed **directly**; `ask_pipeworx`
reaches all of them from either one.
## No MCP client? Call it over HTTP
```bash
curl -X POST https://gateway.pipeworx.io/v1/tools/arquivo_search_pages \
-H 'Content-Type: application/json' \
-d '{"query":"\"inteligência artificial\"","limit":5}'
```
No account needed for the first calls. Inspect any tool: `GET https://gateway.pipeworx.io/v1/tools/arquivo_search_pages`. Find one: `POST https://gateway.pipeworx.io/v1/tools/search_packs` with `{"query":"..."}`.
## Standalone (no gateway account)
This package also runs as a local stdio MCP server — no Pipeworx account, no
gateway round-trip:
```json
{
"mcpServers": {
"arquivo-pt": {
"command": "npx",
"args": ["-y", "@pipeworx/mcp-arquivo-pt"]
}
}
}
```
Or run it directly to confirm it starts:
```bash
npx -y @pipeworx/mcp-arquivo-pt
```
It speaks MCP over stdin/stdout and answers `initialize`/`tools/list`/`tools/call`
for **only** this pack's tools — none of the shared meta-tools the gateway
connection above adds. Same source, same tools, no ask_pipeworx routing.
## Using with ask_pipeworx
Instead of calling tools directly, you can ask questions in plain English —
this works on the pack endpoint above as well as on the full gateway:
```
ask_pipeworx({ question: "your question about Arquivo Pt data" })
```
The gateway picks the right tool and fills the arguments automatically.
## More
- [Docs and guides](https://pipeworx.io/docs)
- [pipeworx.io](https://pipeworx.io)
## License
MIT
What people ask about mcp-arquivo-pt
What is pipeworx-io/mcp-arquivo-pt?
+
pipeworx-io/mcp-arquivo-pt is mcp servers for the Claude AI ecosystem. Arquivo.pt MCP — full-text search over the Portuguese web archive. It has 0 GitHub stars and its last recorded update is dated 2026-10-08.
How do I install mcp-arquivo-pt?
+
You can install mcp-arquivo-pt by cloning the repository (https://github.com/pipeworx-io/mcp-arquivo-pt) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is pipeworx-io/mcp-arquivo-pt safe to use?
+
Our security agent has analyzed pipeworx-io/mcp-arquivo-pt and assigned a Trust Score of 95/100 (tier: Verified). See the full breakdown of passed checks and flags on this page.
Who maintains pipeworx-io/mcp-arquivo-pt?
+
pipeworx-io/mcp-arquivo-pt is maintained by pipeworx-io. The last recorded GitHub activity is dated 2026-10-08, with 0 open issues.
Are there alternatives to mcp-arquivo-pt?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy mcp-arquivo-pt to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/pipeworx-io-mcp-arquivo-pt)<a href="https://claudewave.com/repo/pipeworx-io-mcp-arquivo-pt"><img src="https://claudewave.com/api/badge/pipeworx-io-mcp-arquivo-pt" alt="Featured on ClaudeWave: pipeworx-io/mcp-arquivo-pt" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.