Skip to main content
ClaudeWave
Skill1.6k repo starsupdated 4d ago

verify

Verify apps/desktop frontend changes visually without launching the Tauri app or a live model

Install in Claude Code
Copy
git clone --depth 1 https://github.com/ai4s-research/open-science /tmp/verify && cp -r /tmp/verify/apps/desktop/.claude/skills/verify ~/.claude/skills/verify
Then start a new Claude Code session; the skill loads automatically.

SKILL.md

# Verifying desktop frontend changes

Full-app runs need an OpenCode session + live model turn. For pure frontend
component changes, mount the component in a throwaway vite page instead:

1. Create `apps/desktop/verify-<x>.html` (plain html, `<div id="root">`,
   `<script type="module" src="/src/verify-<x>.tsx">`) and
   `apps/desktop/src/verify-<x>.tsx` (ReactDOM.createRoot, import
   `./index.css` for Tailwind, import the component via `@/`). Vite serves
   any .html under `apps/desktop/` automatically.
2. `cd apps/desktop && npx vite --port 5199 --strictPort` (background).
   The npx wrapper may double-spawn and report "port in use" failure while
   the first instance is fine — check `lsof -iTCP:5199` before retrying.
3. Drive with the browser-control skill (Chrome HTTP API, port 9528):
   `createWindow` → `waitForSelector` → `readDom` → `screenshotTab`
   (screenshots land in `~/Downloads/`).
   Gotchas: `evalScript` is blocked by CSP on the extension — use
   `readDom`/`scrollTab` instead; `readDom` requires an explicit
   `"attributes": []` or it errors with "Value is unserializable".
4. Clean up: `closeTab`, kill the vite pid, delete both harness files,
   confirm `git status` is clean.
my-skillSkill

A test skill that says hello. Use when you want to test skill loading or verify that the skill system is working.

domain-checkSkill

Use whenever you write or run scientific analysis code (physics, earth/geo, biology, chemistry, or social science) in this workspace — before executing it and again after generating results. Runs a deterministic domain-correctness gate that catches code which runs but is scientifically wrong (unit/dimension mismatch, Euclidean distance on lat/lon without a CRS, 0-based/1-based coordinate and strand errors, impossible SMILES valence, uncorrected multiple comparisons, averaging a categorical code). Surfaces structured findings; never claims the code is correct.

large-fileSkill

Use BEFORE reading any data file that could be large (CSV/TSV, Parquet, HDF5, FITS, NetCDF, NDJSON, genomics FASTQ/FASTA/VCF/BAM, GRIB, ROOT, or big text/simulation logs like VASP OUTCAR). Returns a compact memory pointer — header/schema/shape/sample/key numbers — by introspection and sampling in bounded memory, so you never load a file bigger than the context window into the model. Reference data via the pointer; read specific ranges deterministically.

modal-runSkill

Use when the user asks to run heavy or GPU work on Modal (the cloud compute platform) — writing a Modal function in the workspace, running it with the user's own `modal` CLI + token, and bringing results back. Data-to-compute for jobs too big for the laptop, without a Slurm cluster.

publication-figuresSkill

Use whenever you generate or review a chart, plot, table, or paper figure in this workspace, including work delegated by paper-writing, literature-survey, and experiment skills. Applies the Open Science publication style, enforces readable final-size layout for figures and tables, and rejects generic diagram-tool output as a publication figure. Interactive Plotly/HTML may be used for exploration, but paper delivery requires a static publication-ready export.

remote-computeSkill

Use when the user asks to run, submit, monitor, or cancel a job on a remote machine over SSH — their own GPU/CPU server, a workstation, or a Slurm cluster ("the cluster", a login node, "my 3090 box", "the compute server"). Picks a saved machine, runs the work directly over SSH (or via Slurm when present), tracks it, and fetches results back into the workspace.

stats-integritySkill

Use whenever you run statistical analysis for the social sciences (regression, hypothesis tests, econometrics) or read Stata (.dta) / SPSS (.sav) data in this workspace. Enforces an execute-don't-interpret boundary (surface estimates, don't volunteer causal claims), checks the analysis against a preregistration plan for HARKing, verifies reproducible seeds, and reproduces .dta/.sav estimates via R. Flags integrity risks; never certifies the analysis is sound.

traceability-reviewSkill

Use when the user asks to review, verify, or audit a report, manuscript, or analysis in the workspace for traceability — resolving citations, flagging numbers with no source, and checking figures against the code that generated them. Emits a structured review block the app renders as reviewer findings. Verifies traceability, never "correctness".