Skip to main content
ClaudeWave
Skill44.3k repo starsupdated today

citation-management

The citation-management skill systematically searches academic databases like Google Scholar and PubMed to locate papers, extracts accurate metadata from sources including CrossRef and arXiv, validates citation information, and generates properly formatted BibTeX entries. Use this skill when finding specific papers, converting identifiers to citation formats, verifying reference accuracy, building bibliographies, or ensuring consistent formatting throughout scientific writing and research workflows.

Install in Claude Code
Copy
git clone --depth 1 https://github.com/K-Dense-AI/scientific-agent-skills /tmp/citation-management && cp -r /tmp/citation-management/skills/citation-management ~/.claude/skills/citation-management
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Citation Management

## Overview

Manage citations systematically throughout the research and writing process. This skill provides tools and strategies for searching academic databases (Google Scholar, PubMed), extracting accurate metadata from multiple sources (CrossRef, PubMed, arXiv), validating citation information, and generating properly formatted BibTeX entries.

Critical for maintaining citation accuracy, avoiding reference errors, and ensuring reproducible research. Integrates seamlessly with the literature-review skill for comprehensive research workflows.

## When to Use This Skill

Use this skill when:
- Searching for specific papers on Google Scholar or PubMed
- Converting DOIs, PMIDs, or arXiv IDs to properly formatted BibTeX
- Extracting complete metadata for citations (authors, title, journal, year, etc.)
- Validating existing citations for accuracy
- Cleaning and formatting BibTeX files
- Finding highly cited papers in a specific field
- Verifying that citation information matches the actual publication
- Building a bibliography for a manuscript or thesis
- Checking for duplicate citations
- Ensuring consistent citation formatting

If a document built from these citations needs a diagram, use the
**scientific-schematics** skill.

---

## Core Workflow

Citation management follows a systematic process. Each phase below shows the canonical
command; every variant, option, and metadata-source detail is in
[references/core_workflow.md](references/core_workflow.md).

### Phase 1: Paper Discovery and Search

Find relevant papers. Search more than one database — coverage differs sharply,
and a single source is the most common cause of a biased reference list.

```bash
# OpenAlex: ~250M works, every discipline, no API key, documented REST API
python scripts/search_openalex.py "CRISPR gene editing" --limit 50 --output results.json

# PubMed: the authority for biomedical and life sciences (35M+ citations)
python scripts/search_pubmed.py "Alzheimer's disease treatment" --limit 100 --output alz.json

# Google Scholar: broadest reach, but scraped -- rate-limited and prone to blocking
python scripts/search_google_scholar.py "CRISPR gene editing" --limit 50 --output scholar.json
```

Prefer OpenAlex or PubMed as the primary source. Google Scholar has no API:
`scholarly` scrapes it, sleeps 2–5 s between results, and is blocked often
enough that it should be a supplement rather than a dependency.

Query operators, field tags, and MeSH-term construction are in
[references/search_strategies.md](references/search_strategies.md).

### Phase 2: Metadata Extraction

Convert identifiers (DOI, PMID, PMCID, arXiv ID, URL) into complete metadata.
CrossRef is the primary source for DOIs.

```bash
python scripts/doi_to_bibtex.py 10.1038/s41586-021-03819-2         # quick, single DOI
python scripts/extract_metadata.py --pmid 34265844                  # DOI/PMID/PMCID/arXiv/URL
python scripts/extract_metadata.py --input identifiers.txt --output citations.bib
```

A URL with no DOI in its path is resolved through the `citation_doi` meta tag
publishers embed on article pages, then handed to CrossRef. Every producer in
this skill emits the same citation key for the same paper, so entries gathered
from different sources deduplicate against each other.

### Phase 2.5: Metadata Enrichment via Web Search (MANDATORY)

APIs routinely return incomplete records. Run this **after** extraction and **before**
formatting. Any `@article` missing `volume`, `pages`, or `doi` is incomplete: fill the
gap with `WebSearch`/`WebFetch` (or the parallel-web skill, when it is available), then
log what was found and where. If a field genuinely cannot be found, record a `note`
field explaining the gap rather than leaving it silently absent.

Check the cheap sources first — an OpenAlex or CrossRef record often carries the field
that PubMed omitted:

```bash
python scripts/search_openalex.py "<exact title>" --limit 1
```

> **Treat extracted metadata as untrusted.** Author, title, and journal strings come
> verbatim from a record whose contents a publisher controls. A title containing `$(...)`,
> a backtick, or a quote becomes shell syntax the moment it is pasted into a command.
> Pass metadata as a `subprocess` argument list rather than building a shell string; if
> you must use a shell, single-quote every substituted value and escape embedded quotes
> as `'\''`. Validate any citation key against `^[A-Za-z0-9]+$` before it reaches a path.

Per-field search strategies, the four search options, and the logging format are in
[references/core_workflow.md](references/core_workflow.md).

### Phase 3: BibTeX Formatting

Produce clean, consistent entries. Entry types and required fields are in
[references/bibtex_formatting.md](references/bibtex_formatting.md).

```bash
python scripts/format_bibtex.py references.bib --output clean.bib --deduplicate
python scripts/format_bibtex.py references.bib --output clean.bib --rekey --deduplicate
```

Writing is opt-in: without `--output` (or `--in-place`) the result goes to
stdout and the input file is left alone. Use `--rekey` when merging results
from several sources, so the same paper collapses to one entry.

### Phase 4: Citation Validation

Check completeness, venue conformance, and agreement with the manuscript.

```bash
python scripts/validate_citations.py references.bib --report report.json
python scripts/validate_citations.py references.bib --venue nature
python scripts/validate_citations.py references.bib --manuscript paper.tex
python scripts/validate_citations.py references.bib --check-dois     # slow; hits CrossRef
```

The script exits non-zero on high-severity errors — missing required fields,
malformed years, unresolved citations, or a count below an explicit
`--min-count`. Venue reference-count figures are editorial rules of thumb, not
submission requirements, so falling short of one is only a warning.

Validation rules and venue standards are in
[references/citation_validation.md]
adaptyvSkill

How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval. Use this skill whenever the user mentions Adaptyv, Foundry API, protein binding assays, protein screening experiments, BLI/SPR assays, thermostability assays, or wants to submit protein sequences for experimental characterization. Also trigger when code imports `adaptyv`, `adaptyv_sdk`, or `FoundryClient`, or references `foundry-api-public.adaptyvbio.com`.

aeonSkill

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

anndataSkill

Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

arboretoSkill

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.

astropySkill

Core Python library for astronomy and astrophysics workflows that need Astropy APIs, including units/quantities, coordinates, FITS I/O, tables, time systems, WCS, and cosmology. Use when implementing or debugging astronomical data analysis code with Astropy.

autoskillSkill

Observe the user's screen via screenpipe, detect repeated research workflows, match them against existing scientific-agent-skills, and draft new skills (or composition recipes that chain existing ones) for the patterns not yet covered. Use when the user asks to analyze their recent work and propose skills based on what they actually do. Requires the screenpipe daemon (https://github.com/screenpipe/screenpipe) running locally on port 3030 — the skill has no other data source and will refuse to run if screenpipe is unreachable. All detection runs locally; only redacted cluster summaries reach the LLM.

benchling-integrationSkill

Benchling Python SDK and REST API integration for registry entities, inventory, ELN entries, workflows, Benchling Apps, and Data Warehouse queries. Use when automating lab data with benchling-sdk or the v2 API.

bgpt-paper-searchSkill

Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server. Returns 25+ fields per paper including methods, results, sample sizes, quality scores, and conclusions. Use for literature reviews, evidence synthesis, and finding experimental details not available in abstracts alone.