MCP server for Opensolr — hybrid BM25+kNN search, server-side embeddings, and RAG answers as agent tools
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Clear description
- ✓Documented (README)
- !README contains suspicious pattern: eval\s*\(
claude mcp add opensolr-mcp -- uvx opensolr-mcp{
"mcpServers": {
"opensolr-mcp": {
"command": "uvx",
"args": ["opensolr-mcp"],
"env": {
"OPENSOLR_API_KEY": "<opensolr_api_key>"
}
}
}
}OPENSOLR_API_KEYResumen de MCP Servers
# opensolr-mcp
mcp-name: com.opensolr/opensolr-mcp
MCP (Model Context Protocol) server for [Opensolr](https://opensolr.com) —
gives any AI agent **managed Apache Solr search** as tools: hybrid
(BM25 + kNN) retrieval, server-side GPU embeddings, document indexing (local
files and folders too, PDFs page by page, scans through OCR), and grounded RAG
answers.
**See it live (real news index, hybrid + AI answer):** https://search.opensolr.com/news__dense?q=how+am+I+supposed+to+save+money%3F
No embedding model to configure. No vector database to run. One API key.
## Tools
| Tool | What it does |
|---|---|
| `opensolr_search` | Hybrid (keyword + semantic) or pure semantic search, with Solr filters and a date range; a PDF comes back as its best page, with a link that opens it at that page |
| `opensolr_search_by_image` | Search with a **photo** — Opensolr reads its visual labels, OCR text and any barcode/QR, then searches with those words (no image vector stored) |
| `opensolr_read_document` | Read one result in full — a PDF page by page, any other document as one text |
| `opensolr_ai_answer` | Grounded RAG answer: top hybrid hits become the LLM context — same pipeline as the hosted search UI |
| `opensolr_index_files` | Index a local file or a whole folder: PDFs one document per page, Word / RTF / OpenDocument / text / HTML, scans and pictures through OCR |
| `opensolr_extract_text` | Read local files without indexing them: text, OCR of their pictures, a PDF page by page |
| `opensolr_describe_image` | What a picture shows (a sentence), the text printed on it, its labels and barcodes |
| `opensolr_add_documents` | Index plain text + metadata (embedded server-side) |
| `opensolr_delete_documents` | Remove documents by id |
| `opensolr_list_indexes` / `opensolr_index_info` | Inspect the account's indexes |
| `opensolr_create_index` | Provision a vector index in a region of `opensolr_vector_regions` (an environment id, or a short form like `us` matched against that live list) |
| `opensolr_vector_regions` | Live list of vector-enabled regions |
| `opensolr_check_schema` | Whether an index can do what each tool needs (Data Ingestion, vector and hybrid search, PDF pages, keyword search), and exactly why not |
| `opensolr_apply_schema` | Install the Opensolr vector schema on an index of a vector region (replaces its configuration; `force` when it holds documents) |
## Setup
Get a free Opensolr account (free forever, no card) at
[opensolr.com/register](https://opensolr.com/register) and copy your API key
from **Account**.
### Try it without an account
There is a public demo account. Point the package at it and everything in this
README works immediately, with no signup:
```bash
export OPENSOLR_EMAIL=mcp@opensolr.com
export OPENSOLR_API_KEY=420b8b23e7b12dc8ab838932145a5065
```
`mcp_demo_d1__dense` is already loaded with 300 news articles, so search, filtering
and grounded answers work the moment you connect. You also get the write path:
create your own index on the account, ingest into it and query it. Deletion is not
available on this shared key: no index here can be deleted or reconfigured by hand.
Whatever you create is removed automatically after 3 days.
Know what you are working with:
- **Anything you create there is deleted after 3 days.** Automatically, without warning
or export. That includes indexes you created and every document in them.
- **The account is shared with everyone reading this.** Your index is visible to them
and they can add documents to it, as you can to theirs. Nobody can delete or
reconfigure an index here (that is switched off for this key), but never put anything
real, private or client-owned in it.
- **The limits are per index, and deliberately small.** 200 MB of bandwidth and 50 MB
of disk per index. Bandwidth is the one you will hit first: it covers a demo, a
tutorial and a proof of concept, and it will not carry an application.
When you want an index that is private, yours and still there next week, get your own
key — [free forever, no card](https://opensolr.com/register) — and change the two
variables above. Nothing else in your code changes.
### Claude Desktop / Claude Code
```json
{
"mcpServers": {
"opensolr": {
"command": "uvx",
"args": ["opensolr-mcp"],
"env": {
"OPENSOLR_EMAIL": "you@example.com",
"OPENSOLR_API_KEY": "YOUR_OPENSOLR_API_KEY"
}
}
}
}
```
### Cursor / Windsurf / any MCP client
Same shape — stdio transport, command `uvx opensolr-mcp` (or
`pipx run opensolr-mcp`), with `OPENSOLR_EMAIL` and `OPENSOLR_API_KEY` in env.
## Example agent session
> **You:** Index our FAQ answers, then find everything about refunds.
>
> The agent calls `opensolr_add_documents(index="faq__dense", texts=[...])`,
> then `opensolr_search(index="faq__dense", query="refund policy", hybrid=True)`
> — BM25 catches the exact word "refund", kNN catches "giving customers their
> money back", and the scores fuse per document.
## Search with a photo
`opensolr_search_by_image` lets the agent search with a **picture** instead of a
text query. Opensolr reads the image three ways — visual labels (what it
depicts), OCR text (words printed on it), and any barcode / QR code — turns that
into words, and runs the normal search. No image vector is stored.
```
opensolr_search_by_image(index="catalog__dense", image_path="/tmp/shelf.jpg",
using="auto", k=5)
# using: "auto" (engine's choice) | "meaning" (visual labels) |
# "text" (OCR only) | "code" (exact barcode/QR) | "all" (everything)
# -> { "read": {text, mode, labels, codes}, "results": [...] }
```
`search_mode`, `mode`, `alpha`, `fresh_bias`, `filter_query`, `date_from`, `date_to`
and `group_pages` behave exactly as in `opensolr_search`.
## Local files and folders: PDFs page by page, scans through OCR
`opensolr_index_files` indexes a file or a whole folder from the machine the server
runs on. Each file is read on Opensolr's servers — its text, the text in its
pictures (OCR, more than 100 languages), a short description of the pictures that
have no text — and a **PDF is indexed one document per page**, so a search lands on
the page that answers it instead of a 200-page file that mentions it somewhere.
Word, RTF, OpenDocument, text, HTML and picture files (scans, photos of documents)
are one document each.
```
opensolr_index_files(index="contracts__dense", path="~/Documents/contracts",
pattern="*.pdf", metadata={"team": "legal"})
# -> {"documents": [{path, source, uri, id}], "pending": [], "failed": [], "jobs": [...]}
```
- The bytes travel once: a file the account already sent (same content) is not
uploaded again, and sending the same files again replaces their documents.
- Each document's `uri` is `https://ingest.opensolr.com/<index>/<source>`, where
`source` is the file's path from the folder you sent (also in `meta_source`); its
date is the file's modification time, so date ranges work on local files too.
- Big files are read in the background: whatever is still being read after
`wait_seconds` comes back under `pending`; call the tool again with the same path
to index those (nothing is sent twice).
- Pictures read by OCR count 0.5 of an AI request each; a file read before is free.
`opensolr_extract_text` does the same reading without indexing anything — the text of
each file, a PDF as its pages (`page`, `text`, `is_ocr`, `lang`) — for an agent that
wants to read a contract, a scanned invoice or a folder of reports itself.
`opensolr_describe_image` answers what a picture shows and what is written on it:
`caption` (a sentence), `text` (OCR), `labels` and barcode / QR `codes`.
## PDF pages in search results
On an index where PDFs are stored page by page (files indexed above, PDFs sent to
Data Ingestion with `rtf`, PDFs found by the Web Crawler), `opensolr_search` returns
each PDF **once**, as its best page:
```
{"title": "Master services agreement", "page": 14,
"page_url": "https://example.com/msa.pdf#page=14", # opens the PDF at that page
"document_id": "9f2c...", # for opensolr_read_document
"more_pages": [{"page": 3, "page_url": "...#page=3"}, ...], # its other matching pages
"matching_pages": 4, "uri": "https://example.com/msa.pdf", ...}
```
`group_pages=false` returns the pages as separate results. `opensolr_read_document`
reads any result in full: a PDF as its pages in order (`page_from` / `page_to`), any
other document as one text.
## Searching a period
`date_from` and `date_to` (`YYYY-MM-DD`, both included, either one can be left out)
keep only the documents dated in that period (`creation_date`). They combine with
every other option, `fresh_bias` included.
## Search operators
The query string understands the operators people already expect from a search box. They work
in every mode — keyword, hybrid and pure vector.
| Operator | Meaning | Example |
| --- | --- | --- |
| `"word1 word2"` | Phrase — those words together, in that order | `"machine learning"` |
| `+word` | Required — every result must contain it | `+laptop 15 inch gaming` |
| `+"word1 word2"` | Required phrase | `+"13 inch"` |
| `-word` | Excluded — drop any document containing it | `laptop -refurbished` |
| `-"word1 word2"` | Excluded phrase | `-"open box"` |
They compose: `+laptop +"13 inch" -refurbished` returns only 13-inch laptops and never a
refurbished one.
A prefixed term (`+` or `-`, word or phrase) becomes a **filter**, applied to the whole result
set. That matters as soon as a vector is involved: the semantic side of a hybrid search has no
concept of negation, so left as query text `-refurbished` would actually pull refurbished
listings *towards* the top rather than removing them. As a filter it binds every document,
whichever side of the search found it, and the exclusion is absolute.
An unprefixed phrase (`"machine learning"` with no `+` in front) is a keyword-side relevance
signal raLo que la gente pregunta sobre opensolr-mcp
¿Qué es phpcip/opensolr-mcp?
+
phpcip/opensolr-mcp es mcp servers para el ecosistema de Claude AI. MCP server for Opensolr — hybrid BM25+kNN search, server-side embeddings, and RAG answers as agent tools Tiene 0 estrellas en GitHub y su última actualización registrada es del 2026-10-09.
¿Cómo se instala opensolr-mcp?
+
Puedes instalar opensolr-mcp clonando el repositorio (https://github.com/phpcip/opensolr-mcp) o siguiendo las instrucciones del README en GitHub. ClaudeWave también te ofrece bloques de instalación rápida en esta misma página.
¿Es seguro usar phpcip/opensolr-mcp?
+
Nuestro agente de seguridad ha analizado phpcip/opensolr-mcp y le ha asignado un Trust Score de 77/100 (tier: Trusted). Revisa el desglose completo de comprobaciones superadas y flags en esta página.
¿Quién mantiene phpcip/opensolr-mcp?
+
phpcip/opensolr-mcp es mantenido por phpcip. La última actividad registrada en GitHub es del 2026-10-09, con 0 issues abiertos.
¿Hay alternativas a opensolr-mcp?
+
Sí. En ClaudeWave puedes explorar mcp servers similares en /categories/mcp, ordenados por popularidad o actividad reciente.
Despliega opensolr-mcp en tu cloud
Lleva este repo a producción en minutos. Cada plataforma genera su propio entorno con variables de entorno editables.
¿Mantienes este repo? Añade un badge a tu README
Pega el badge en tu README de GitHub para mostrar que está auditado por ClaudeWave. Cada badge enlaza de vuelta a esta página y muestra el Trust Score actual.
[](https://claudewave.com/repo/phpcip-opensolr-mcp)<a href="https://claudewave.com/repo/phpcip-opensolr-mcp"><img src="https://claudewave.com/api/badge/phpcip-opensolr-mcp" alt="Featured on ClaudeWave: phpcip/opensolr-mcp" width="320" height="64" /></a>Más MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.