The missing layer between AI and the web. Open-source escalating web unlocker + read/search/transcribe/grab across web, GitHub, YouTube, Reddit, Twitter, LinkedIn, and RSS for AI agents.
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Topics declared
- ✓Documented (README)
git clone https://github.com/capad-xyz/searchts && cp searchts/*.md ~/.claude/agents/Resumen de Subagents
# searchts <!-- mcp-name: io.github.capad-xyz/searchts --> **The missing layer between AI and the web.** A Python CLI and library that lets an AI agent read and search the internet, fronted by a fully open-source "unlocker" that gets through common bot-walls with no paid proxy and no API key. [](https://github.com/capad-xyz/searchts/actions/workflows/pytest.yml) [](https://pypi.org/project/searchts/) [](https://pypi.org/project/searchts/) [](https://pepy.tech/projects/searchts) [](https://github.com/capad-xyz/searchts/blob/main/LICENSE) <p align="center"> <img src="https://raw.githubusercontent.com/capad-xyz/searchts/main/demo/demo1.gif" alt="A Claude agent's fetch hits a 403 bot wall, so it routes through searchts, reads the page, and answers the question" width="860"> <br> <a href="https://github.com/capad-xyz/searchts/releases/download/v0.7.0/searchts-demo-v5.mp4">▶ Watch the full 1-minute demo</a> </p> ## Why searchts? - Reads pages behind common bot walls - Reads complete ChatGPT / Claude / Gemini / Grok / Poe / DeepSeek / Perplexity / Copilot shared conversations - Works with Claude, Codex, and MCP agents - Extracts clean Markdown, ready to feed a model - Says when a page has more than it returned (a next page, a feed, folded text) and rebuilds search results and feeds the extractor mangles - Searches the web without API keys - Downloads a page's assets (images, fonts, palette) - Transcribes videos, subtitles-first ## Why it's free AI agents constantly need to read web pages, but the naive way they fetch is trivially blocked by modern anti-bot systems (Cloudflare, PerimeterX, DataDome). Paid unlocker services solve this, but the thing they really charge for is a large pool of clean residential IP addresses. `searchts` runs on your own machine, from your own connection, at personal volume, so it sidesteps that cost and gets through most of those walls for free. ## The unlocker `searchts` reads any URL through an escalating ladder and stops at the first tier that returns real content: 1. **curl_cffi**: a fetch that impersonates a real Chrome's TLS/JA3 and HTTP2 fingerprint. Beats user-agent and fingerprint filters. Fast, local, private. 2. **Jina Reader**: a JavaScript-rendering relay (`r.jina.ai`), for pages that only fill in content after running JS. **Default on** — the target URL is sent to Jina's servers on this rung. Opt out with `SEARCHTS_NO_JINA=1` or config `jina: false` (local curl + stealth only). 3. **stealth browser**: an undetected headless Chromium (patchright), launched lazily only when the cheaper tiers fail, for live JS / Cloudflare managed challenges. If no tier comes back with real content, an optional human-in-the-loop step opens a real browser so you can clear the page once and continue. That covers interactive CAPTCHAs and soft walls alike: a login page served as HTTP 200 is not a challenge, but it is still a page only a human gets past. Block detection is phrase-based (not vendor-name based), so legitimate pages that merely embed a bot-sensor script are not falsely rejected. Content is extracted to clean Markdown with `trafilatura`. **Walls (F12 playbook, not a bypass):** fail loud on login/challenge/thin. Do not cut a release that claims Reddit/LinkedIn now read (**N7**). Order: stealth already retries `page.content` after a navigation race (**P3.11**) → next is a persistent Chromium profile so clearance can survive across reads (**F1**, not shipped) → then `--human` / device session for extras only (**F7**, never silent, never inside `read_url`). Never paid residential as default (**N1**). Never a keyed commercial unlocker as default (**N3**). ## AI-chat share links Share links from AI chat apps are a special kind of hard: the conversation never appears in the page HTML as extractable text, so generic readers (and most AI agents' built-in fetch) return an empty shell or a fragment cut off mid-chat. `searchts read` recognizes these URLs and decodes each provider's own data channel instead, returning the **complete conversation** as role-labeled Markdown — keyless, no login: | Provider | Share URL | How it's read | |----------|-----------|---------------| | ChatGPT | `chatgpt.com/share/…`, `chatgpt.com/s/…` | turbo-stream payload embedded in the page | | Claude | `claude.ai/share/…` | keyless snapshot API (behind Cloudflare) | | Gemini | `gemini.google.com/share/…` | keyless batchexecute RPC | | Grok | `grok.com/share/…`, `x.com/i/grok/share/…` | keyless share-links API | | Poe | `poe.com/s/…` | `__NEXT_DATA__` payload embedded in the page | | DeepSeek | `chat.deepseek.com/share/…` | stealth render, scrolled to the end | | Perplexity | `perplexity.ai/search/…`, `perplexity.ai/page/…` | stealth render, scrolled to the end | | Copilot | `copilot.microsoft.com/shares/…`, `…/shares/pages/…` | stealth render, scrolled to the end | The first five need no browser. The last three are JavaScript shells with nothing in the initial HTML, so those reuse the stealth tier: wait for the conversation to render, auto-scroll until the page height stops changing (list virtualization will otherwise truncate a long chat), then expand the collapsed sections before reading. The benchmark currently covers the five that read without a browser and passes all five; the three that need one are not in it yet. ChatGPT issues two shapes: `/share/<uuid>` for a whole conversation, and the newer `/s/<prefix>_<id>` short links for a single shared turn (`t_` thread, `m_` message, `dr_` deep research, `cd_` Codex). Both are read. Each provider is a drop-in plugin module (`searchts/share_extractors/`); if a provider changes its format, extraction falls back to the normal unlocker ladder instead of failing. ## Install Keep it (global isolated CLI, MCP extra included): ```bash pipx install "searchts[mcp]" ``` Try it without installing (one-shot, copy-paste): ```bash uvx --from "searchts[mcp]" searchts <verb> ``` venv / packaging only (not the recommended path for the CLI): ```bash python -m venv .venv source .venv/bin/activate # Windows: .venv\Scripts\activate pip install "searchts[mcp]" ``` Stealth browser (installs into the same env as the running CLI). With uvx or uv tool, put `browser` in the spec itself, for example `uvx --from "searchts[mcp,browser]"`, in every command you use, the MCP one included. Only Chromium is shared between environments: ```bash searchts install --browser ``` ## Quickstart ```bash searchts read https://en.wikipedia.org/wiki/Ada_Lovelace # fetch a page as clean Markdown searchts search "open source vector db" # multi-provider web search (keyless by default) searchts transcribe https://youtu.be/... # transcript of a YouTube/TikTok/Instagram/Reddit video searchts grab https://example.com # download a page's assets + extract palette/fonts searchts get https://example.com/logo.png # download one asset (image/PDF/font/file) searchts doctor # see what is configured and working ``` `read` flags: `--json`, `--backend <tier>`, `--human` (hand off a CAPTCHA or login wall to a real browser), `--scrub` (redact injection). `search` flags: `-n <count>`, `--json`, `--provider <name>`. Content goes to stdout (pipeable); status to stderr. `grab` flags: `--out <dir>`, `--kinds <images,icons,css,fonts,svg>`, `--read` (also save page.md), `--max <n>`, `--json`. ## Use it from your AI agent Add searchts to your agent in one line - as an MCP server, or as a Claude Code slash command: <p align="center"> <img src="https://raw.githubusercontent.com/capad-xyz/searchts/main/demo/demo2.gif" alt="Installing searchts as an MCP server with claude mcp add, or as a Claude Code slash command with searchts skill install" width="820"> </p> Two ways, both one command: ```bash # 1) MCP: always-on read_url + web_search + fetch_asset + grab_site + get_status + transcribe # Try / no install / Claude cannot see PATH: claude mcp add searchts -- uvx --from "searchts[mcp]" searchts mcp serve # Keep (after pipx install "searchts[mcp]"): # claude mcp add searchts -- searchts mcp serve # Desktop / Cursor JSON: `searchts mcp install` (or uvx the same serve command) # First read: Wikipedia — example.com is thinner than _MIN_CHARS and looks like a failed install. # 2) Slash command: type /searchts <url-or-query> in Claude Code searchts skill install # writes ~/.claude/commands/searchts.md ``` See the [MCP server reference](https://github.com/capad-xyz/searchts/blob/main/docs/mcp.md) for all six tools (`read_url`, `web_search`, `fetch_asset`, `grab_site`, `get_status`, `transcribe`), their inputs and outputs, and when to use each. ## Features - **Escalating open-source unlocker**: curl_cffi, then Jina Reader, then a stealth browser. - **Multi-provider search with rank fusion**: DuckDuckGo (keyless default), plus SearXNG, Exa, Brave, and Tavily when configured; results merged with reciprocal rank fusion and de-duplicated. - **Video transcription**: yt-dlp audio plus Whisper for YouTube, TikTok, Instagram, and Reddit videos. - **Asset + design grabber**: `searchts grab <url>` downloads a page's images/icons/css/fonts and extracts a color palette plus the fonts in use; `searchts get <url>` pulls a single asset. Both go through the same escalating unlock ladder, so they work on fingerprint-gated CDNs, not just open ones. - **Prompt-injection scrubbing**: strips invisible/bidi characters, flags injection indicators, optional redaction, so untrusted page content is safer to feed a model. - **Per-domain backend memory**: remembers which tier worked per domain and tries it first (`SEARCHTS_NO_MEMORY=1` to disable). - **Jina opt-out**: the J
Lo que la gente pregunta sobre searchts
¿Qué es capad-xyz/searchts?
+
capad-xyz/searchts es subagents para el ecosistema de Claude AI. The missing layer between AI and the web. Open-source escalating web unlocker + read/search/transcribe/grab across web, GitHub, YouTube, Reddit, Twitter, LinkedIn, and RSS for AI agents. Tiene 2 estrellas en GitHub y su última actualización registrada es del 2026-10-02.
¿Cómo se instala searchts?
+
Puedes instalar searchts clonando el repositorio (https://github.com/capad-xyz/searchts) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar capad-xyz/searchts?
+
Nuestro agente de seguridad ha analizado capad-xyz/searchts y le ha asignado un Trust Score de 95/100 (tier: Verified). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene capad-xyz/searchts?
+
capad-xyz/searchts es mantenido por capad-xyz. La última actividad registrada en GitHub es del 2026-10-02, con 8 issues abiertos.
¿Hay alternativas a searchts?
+
Sí. En ClaudeWave puedes explorar subagents similares en /categories/agents, ordenados por popularidad o actividad reciente.
Despliega searchts en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/capad-xyz-searchts)<a href="https://claudewave.com/repo/capad-xyz-searchts"><img src="https://claudewave.com/api/badge/capad-xyz-searchts" alt="Featured on ClaudeWave: capad-xyz/searchts" width="320" height="64" /></a>Más Subagents
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
The agent that grows with you
Java 面试 & 后端通用面试指南,覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
The agent engineering platform.