MCP server for RCSB PDB APIs
- ✓Open-source license (MIT)
- ✓Actively maintained (<30d)
- ✓Documented (README)
claude mcp add rcsb-mcp -- uvx rcsb-mcp{
"mcpServers": {
"rcsb-mcp": {
"command": "uvx",
"args": ["rcsb-mcp"]
}
}
}MCP Servers overview
<!-- mcp-name: io.github.rcsb/rcsb-mcp --> # rcsb-mcp An [MCP](https://modelcontextprotocol.io) server for **interrogating Protein Data Bank structures** — discover, inspect, and cross-reference — from LLM clients (Claude Desktop, MCP Inspector, Cursor, etc.). It spans three RCSB APIs: - **Discover** — find structures with the [Search API](https://search.rcsb.org) (keyword, attribute, sequence, chemistry, 3D shape, motif). - **Inspect** — fetch entry / entity / assembly / ligand details and annotations from the [Data API](https://data.rcsb.org/graphql). - **Relate** — map sequences and positional features across PDB, UniProt, and NCBI with the [Sequence Coordinates API](https://sequence-coordinates.rcsb.org/graphql). ## Tools ### Search (search.rcsb.org) Searching is **two steps**: build a query with an `rcsb_query_*` tool, then execute it with `rcsb_search_request`. The builders are pure — they return a query document (readable JSON plus a digest) and touch no network, so **nothing is searched until you call `rcsb_search_request`**. `rcsb_query_composer` joins two or more documents with AND/OR, which is also how a single search mixes services (e.g. sequence similarity AND an organism filter). **Build** | Tool | What it does | |------|--------------| | `rcsb_query_fulltext` | Free-text keyword query (e.g. `"CRISPR Cas9"`). | | `rcsb_query_attribute` | Structured query on one or more indexed attributes (resolution, organism, release date, ...) joined by a single `logical_operator`. Each condition supports `exists`, `negation`, `case_sensitive`; `chemical_attributes=True` selects the chemical-component catalog. | | `rcsb_query_sequence` | MMseqs2 sequence-similarity query (BLAST-like), with identity / e-value cutoffs. | | `rcsb_query_chemical` | Chemical query by SMILES/InChI descriptor (whole-molecule or substructure) or molecular formula. | | `rcsb_query_structure` | 3D shape-similarity query against a reference PDB assembly or chain. | | `rcsb_query_seqmotif` | Short **sequence**-motif query (PROSITE pattern, regex, or simple wildcards). | | `rcsb_query_strucmotif` | 3D **structural**-motif query: a geometric arrangement of specific residues (e.g. a catalytic triad). | | `rcsb_query_composer` | Join 2+ query documents with AND/OR — nested boolean logic, and the only way to combine different services in one search. | **Run** | Tool | What it does | |------|--------------| | `rcsb_search_request` | Execute a query document and return matching **identifiers only**. Carries every output option: `return_type`, `limit`/`offset`, `all_hits`, `facets`, `sort_by`/`sort_direction`, `group_by`/`group_by_ranking`, `include_computed_models`. | **Discover** | Tool | What it does | |------|--------------| | `rcsb_list_pdb_search_attributes` | Discover searchable attribute paths, types, and operators. `schema="structure"` (default, ~683) or `schema="chemical"` (~61: `chem_comp.*`, `drugbank_info.*`, ...). Records also carry `enum` (closed value sets) and `nested_group` (attributes stored in nested objects — see below). | | `rcsb_find_go_terms` | Resolve a free-text molecular function / biological process / cellular component to Gene Ontology ids (via EBI QuickGO), annotated with PDB entry counts — then search by `rcsb_polymer_entity_annotation.annotation_lineage.id`. | | `rcsb_find_interpro_domains` | Resolve a free-text protein domain / family / fold to InterPro and Pfam ids (via EBI Search), annotated with PDB entry counts — then search by `rcsb_polymer_entity_annotation.annotation_id`. | | `rcsb_find_enzyme_classes` | Resolve a free-text enzyme / reaction to Enzyme Commission (EC) numbers (via EBI Search/IntEnz), annotated with PDB entry counts — then search by `rcsb_polymer_entity.rcsb_ec_lineage.id` (hierarchical). | | `rcsb_find_disease_terms` | Resolve a free-text disease / condition to MONDO ids (via EBI OLS), annotated with PDB entry counts — then search by `rcsb_uniprot_annotation.annotation_lineage.id` (hierarchical, UniProt-based). | | `rcsb_find_organisms` | Resolve a free-text organism / common name / clade to NCBI Taxonomy ids (via UniProt taxonomy), annotated with PDB entry counts — then search by `rcsb_entity_source_organism.taxonomy_lineage.id` (hierarchical: a clade id matches every organism beneath it). | Both attribute catalogs are generated from the live metadata schemas by [`scripts/generate_search_attributes.py`](scripts/generate_search_attributes.py). To search chemical-component attributes, find the path with `rcsb_list_pdb_search_attributes(schema="chemical")`, pass `chemical_attributes=True` to `rcsb_query_attribute`, and usually set `return_type="mol_definition"`. **Nested attributes.** An object can hold many annotations, many binding affinities, many citations. For attributes carrying a `nested_group`, the query shape selects the semantics: conditions built in **one** `rcsb_query_attribute` call, with nothing else in it, must hold on the **same** record; conditions in separate calls are matched independently against any record. Both are valid and mean different things — `type=Kd` with `value<1` describes one measurement, while an InterPro id and a GO type are necessarily two different annotations. **Counting and faceting** are output options on `rcsb_search_request`, not separate tools: every response includes `total_count` (the full match count — for "how many ..." run the search with `limit=1` and read it), and passing `facets` returns a breakdown (terms/histogram/date_histogram/range/cardinality) instead of hits. A terms facet also tells you what your hits **share**, which is how a handful of results becomes a re-searchable value. **Grouping.** `group_by` returns one representative per cluster — `seqid_30` … `seqid_95` for sequence-identity clusters, or `uniprot` to collapse by accession — with `group_by_ranking` choosing the representative. Requires `return_type="polymer_entity"`; the response reports `group_count` alongside `total_count`. **Sorting.** `sort_by` (an attribute path) + `sort_direction` (`asc`/`desc`) replaces the default score ordering (for similarity searches this overrides the similarity-ranked order). Only attributes indexed for sorting work — those exposing `exact_match` (strings) or `equals` (numbers/dates) in `rcsb_list_pdb_search_attributes`; sorting is not available for `return_type="mol_definition"`. **Paging.** `rcsb_search_request` accepts `limit` (1–100, default 10) and `offset` (default 0), and each response reports `total_count`, `has_more`, and `next_offset` — call again with the same query document and `offset=next_offset`. For an explicit "ALL ..." request, `all_hits=True` returns the complete set in one call (refused above 10,000 hits, and it cannot be combined with `offset`). ### Data (data.rcsb.org/graphql) There is one tool per Data API GraphQL root field. Each takes a **list of IDs** (singular lookups = a one-element list) plus an optional `fields` argument to override the curated default selection with your own GraphQL sub-selection. Unknown IDs are reported under `not_found`. Discover the paths to put in `fields` with `rcsb_describe_data_object` — browse a level, drill into a nested object with `into=`, or search the schema by keyword with `query=` + `max_depth=`. Every path it returns is verified against the live schema, so don't guess field names. | Tool | Object | Example ID | |------|--------|----------------------------------| | `rcsb_get_entries` | PDB entries | `"4HHB"` | | `rcsb_get_polymer_entities` | Polymer entities (protein/NA) | `"4HHB_1"` | | `rcsb_get_nonpolymer_entities` | Ligand/cofactor entities | `"4HHB_3"` | | `rcsb_get_branched_entities` | Carbohydrate entities | `"5FMB_2"` | | `rcsb_get_polymer_entity_instances` | Polymer chains | `"4HHB.A"` | | `rcsb_get_nonpolymer_entity_instances` | Bound-ligand instances | `"4HHB.E"` | | `rcsb_get_branched_entity_instances` | Glycan chains | `"5FMB.C"` | | `rcsb_get_assemblies` | Biological assemblies | `"4HHB-1"` | | `rcsb_get_interfaces` | Assembly interfaces | `"1BMV-1.1"` | | `rcsb_get_chem_comps` | Chemical components / ligands | `"HEM"`, `"ATP"` | | `rcsb_get_entry_groups` | Entry groups | `"G_1002266"` | | `rcsb_get_polymer_entity_groups` | Polymer entity groups (seq. clusters) | `"85_70"` | | `rcsb_get_nonpolymer_entity_groups` | Non-polymer entity groups | `"ATP"` | | `rcsb_get_uniprot` | UniProt record (single) | `"P69905"` | | `rcsb_get_pubmed` | PubMed record (single, integer) | `6726807` | | `rcsb_get_group_provenance` | Grouping provenance (single) | `"provenance_sequence_identity"` | | `rcsb_describe_data_object` | Introspect an object's live GraphQL schema to build a `fields=` selection: browse a level, drill into a nested object with `into=`, or search by keyword with `query=` + `max_depth=` (flat, incl. nested + cross-object paths). Returns verified dotted paths. The Data API analogue of `rcsb_list_pdb_search_attributes`. | — | The Search API only returns identifiers, so a search is the first step: batch the returned ids into the matching `rcsb_get_*` tool to fetch titles, organisms, and other metadata (these tools query the GraphQL endpoint, batching every requested ID into one request). All 16 typed tools are generated from a single registry in [`queries.py`](src/rcsb_mcp/queries.py) (`DATA_OBJECTS`), so adding a field or endpoint is a one-line change. ### Sequence Coordinates (sequence-coordinates.rcsb.org/graphql) Maps alignments and positional annotations between sequence reference systems (`UNIPROT`, `NCBI_PROTEIN`, `NCBI_GENOME`, `PDB_ENTITY`, `PDB_INSTANCE`). Each t
What people ask about rcsb-mcp
What is rcsb/rcsb-mcp?
+
rcsb/rcsb-mcp is mcp servers for the Claude AI ecosystem. MCP server for RCSB PDB APIs It has 4 GitHub stars and its last recorded update is dated 2026-10-05.
How do I install rcsb-mcp?
+
You can install rcsb-mcp by cloning the repository (https://github.com/rcsb/rcsb-mcp) or following the README instructions on GitHub. ClaudeWave also provides quick install blocks on this page.
Is rcsb/rcsb-mcp safe to use?
+
Our security agent has analyzed rcsb/rcsb-mcp and assigned a Trust Score of 82/100 (tier: Trusted). See the full breakdown of passed checks and flags on this page.
Who maintains rcsb/rcsb-mcp?
+
rcsb/rcsb-mcp is maintained by rcsb. The last recorded GitHub activity is dated 2026-10-05, with 5 open issues.
Are there alternatives to rcsb-mcp?
+
Yes. On ClaudeWave you can browse similar mcp servers at /categories/mcp, sorted by popularity or recent activity.
Deploy rcsb-mcp to your cloud
Ship this repo to production in minutes. Each platform spins up its own environment with editable env vars.
Maintain this repo? Add a badge to your README
Drop the badge into your GitHub README to show it's tracked on ClaudeWave. Each badge links back to this page and reflects the live Trust Score.
[](https://claudewave.com/repo/rcsb-rcsb-mcp)<a href="https://claudewave.com/repo/rcsb-rcsb-mcp"><img src="https://claudewave.com/api/badge/rcsb-rcsb-mcp" alt="Featured on ClaudeWave: rcsb/rcsb-mcp" width="320" height="64" /></a>More MCP Servers
Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
An open-source AI agent that brings the power of Gemini directly into your terminal.
Real-time global intelligence dashboard. AI-powered news aggregation, geopolitical monitoring, and infrastructure tracking in a unified situational awareness interface
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
The fastest path to AI-powered full stack observability, even for lean teams.