Skip to main content
ClaudeWave
Skill281 repo starsupdated 4d ago

model-evaluation

>

Install in Claude Code
Copy
git clone --depth 1 https://github.com/Aperivue/medsci-skills /tmp/model-evaluation && cp -r /tmp/model-evaluation/skills/model-evaluation ~/.claude/skills/model-evaluation
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Model-Evaluation Skill

## Purpose

This skill makes a medical-imaging model's **held-out evaluation task-correct and honest**: the right
metric for the task and the prevalence, with uncertainty, calibration, and subgroup performance. It
emits a **per-case metric table** that the publication statistics build on, and gates the metric choice
against Metrics Reloaded (Maier-Hein & Reinke et al., *Nat Methods* 2024) and CLAIM 2024.

It sits between `/model-validation` (which audits the split / design) and `/analyze-stats` (which owns
the comparative inference). It computes the imaging-specific per-case metrics (surface distances, FROC,
ECE of a softmax head); `/analyze-stats` owns DeLong / NRI / IDI / decision curves / MRMC. Like
`/analyze-stats`, it **generates and executes** code on your predictions — numbers are never hand-typed.

## When to use
- You have held-out predictions + ground truth and need task-correct metrics with CIs, calibration, and
  subgroup slices, plus a per-case table for the manuscript statistics.

## When NOT to use
- Auditing the validation design / leakage → `/model-validation`.
- DeLong / NRI / IDI / decision curves / MRMC reader study → `/analyze-stats`.
- Building / training the model → `/model-scaffold`; LLM / MLLM → `/mllm-eval`.
- Figure rendering → `/make-figures`.

## Workflow

### Phase 1 — Fix the analysis unit and the task
State the task (segmentation / classification / detection / interactive / generative) and the **analysis unit** the metric must
respect (per-patient vs per-lesion vs per-image). A per-lesion metric must not be reported as
per-patient.

### Phase 2 — Compute task-correct metrics
Generate evaluation code that computes, on the held-out predictions:
- **segmentation**: Dice/IoU **and** a boundary metric (HD95 / NSD), **per structure** not only a global
  mean, with bootstrap 95% CIs.
- **classification**: **AUROC and AUPRC** with bootstrap CIs, sensitivity/specificity, and PPV/NPV **at
  the deployment prevalence** (not a balanced set).
- **detection**: **FROC / mAP** with the **IoU match criterion stated**.
- **interactive / promptable segmentation** (SAM2 / MedSAM2 / nnInteractive): the segmentation
  metrics above **plus** the interaction axis — Dice-vs-interactions / number-of-clicks (NoC) to a
  target threshold, initial-vs-converged (or peak) Dice, and per-case interaction/inference time
  (see the metric guide; the study design is in `/design-study` + `/model-validation`).
- **generative / synthesis** (image generation or modification): full-reference similarity
  (MSE/RMSE/PSNR/SSIM) or no-reference quality (SNR/CNR, standardized visual scores), **plus a
  downstream-task evaluation** — image quality is not clinical utility (Park et al., *Radiol Med* 2024).
  For **multiclass** classification, state the aggregation scheme (one-vs-rest / macro / micro /
  pairwise / Obuchowski); **time-to-event** discrimination (Harrell's C, time-dependent ROC) is handed
  to `/analyze-stats`.
Add **calibration** (reliability diagram / ECE) and **subgroup** slices (the Model Card Factors).
See `${CLAUDE_SKILL_DIR}/references/metric_guide.md`. Emit a **per-case CSV** for `/analyze-stats`.

### Phase 3 — Gate the metric choice (deterministic)
```bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_metric_reporting.py \
  --report results.md --task segmentation|classification|detection|interactive|generative --strict
```
`PIXEL_ACCURACY_SEG` / `NO_BOUNDARY_METRIC` / `ACCURACY_ONLY` / `DETECTION_METRIC_MISSING` must be zero.

### Phase 4 — Hand off
The per-case table → `/analyze-stats` (DeLong / NRI / IDI / decision curves, publication tables);
figures → `/make-figures`; the numbers + subgroup performance → `/model-card`; Methods/Results →
`/write-paper`; compliance → `/check-reporting`.

## Anti-Hallucination

- **Never fabricate a metric value.** Every number comes from executed code on the supplied predictions;
  if predictions or ground truth are missing, say so and stop — do not invent a result.
- **Never report pixel/voxel accuracy for segmentation or bare accuracy under imbalance** — the gate
  flags these; report Dice + a boundary metric, or AUROC + AUPRC with CIs.
- **Never report a per-lesion metric as if it were per-patient** — respect the analysis unit.
- If a metric definition or its CI method is uncertain, flag `[VERIFY]` and ask.

## Deterministic gate
`scripts/check_metric_reporting.py` — flags a task-metric mismatch / missing uncertainty (stdlib,
network-free). Reproducible challenge: `bash ${CLAUDE_SKILL_DIR}/scripts/metric_reporting_challenge/verify.sh`.

## Reference Files

Load on demand (keep SKILL.md short):
- `${CLAUDE_SKILL_DIR}/references/metric_guide.md` — operational checklist: the task-correct metric
  per task (segmentation Dice + HD95/NSD per structure; classification AUROC + AUPRC + sens/spec at
  deployment prevalence; detection FROC/mAP with a stated IoU), plus calibration, subgroup slices,
  run-variance, and the per-case CSV hand-off.
- `${CLAUDE_SKILL_DIR}/references/metric_selection_grounding.md` — the standards grounding behind
  those choices: the Metrics Reloaded task-fingerprint principle, why each metric pairing is
  required, calibration vs discrimination, disaggregated reporting, and the CLAIM 2024
  reporting-fit map (`/check-reporting` owns the item audit).

## Boundaries

```
model-validation (design) -> model-evaluation (this skill: per-case task-correct metrics + CIs)
  -> analyze-stats (DeLong / NRI / IDI / decision curves, publication tables) -> make-figures
  -> model-card (numbers + subgroup) -> write-paper + check-reporting
```
skillsSkill
academic-aioSkill

Medical AI paper optimization for AI search engines (Perplexity, ChatGPT web, Elicit, Consensus, SciSpace) and RAG-based literature tools. Applies when drafting or reviewing titles, abstracts, structured summary boxes (Key Points / Research in Context / Plain-Language Summary), manuscripts for high-impact medical AI journals (Lancet Digital Health, Radiology, Radiology-AI, npj Digital Medicine, Nature Medicine), preprints (medRxiv/arXiv), GitHub README + CITATION.cff + Zenodo archives, and Hugging Face model/dataset cards. Integrates TRIPOD+AI, CLAIM 2024, STARD-AI, TRIPOD-LLM, DECIDE-AI reporting requirements with generative engine optimization (GEO) principles. Produces a visible pass/fail checklist.

add-journalSkill

>

analyze-statsSkill

Statistical analysis for medical research papers. Generates reproducible Python/R code with publication-ready tables and figures. Supports diagnostic accuracy, inter-rater agreement, meta-analysis, survival analysis, survey data, group comparisons, regression, propensity score, and repeated measures.

author-strategySkill

PubMed author profile analysis. Author name → PubMed fetch → study-type classification → visualization → strategy report → optional trajectory-archetype classification.

batch-cohortSkill

Generate N analysis scripts from a single methodology template × multiple exposure/outcome combinations. The "80-person team" pattern — same validated method, swap variables only. Produces batch R/Python code + summary matrix.

calc-sample-sizeSkill

>

check-reportingSkill

Check manuscript compliance with medical research reporting guidelines. Supports 49 guidelines including STROBE, STROBE-MR, RECORD, REMARK (prognostic tumor-marker studies), TARGET (target trial emulation), GATHER (burden-of-disease / health-estimate modeling), CONSORT, CONSORT-AI, STARD, STARD-AI, TRIPOD, TRIPOD+AI, TRIPOD-LLM, PGS-RS, ARRIVE, PRISMA, PRISMA 2020 for Abstracts, PRISMA-DTA, PRISMA-P, PRISMA-ScR (scoping reviews), CARE, SPIRIT, SPIRIT-AI, CLAIM, DECIDE-AI, MI-CLEAR-LLM, SQUIRE 2.0, CLEAR, MOOSE, GRRAS, SWiM, AMSTAR 2, CHEERS 2022, CROSS (survey studies), SRQR and COREQ (qualitative research), and risk of bias tools (QUADAS-3, QUADAS-2, QUADAS-C, RoB 2, ROBINS-I, ROBINS-E, ROBIS, ROB-ME, PROBAST, PROBAST+AI, NOS, COSMIN, RoB NMA). Generates item-by-item assessment with PRESENT/MISSING/PARTIAL status.