verify
Verify apps/desktop frontend changes visually without launching the Tauri app or a live model
git clone --depth 1 https://github.com/ai4s-research/open-science /tmp/verify && cp -r /tmp/verify/apps/desktop/.claude/skills/verify ~/.claude/skills/verifySKILL.md
# Verifying desktop frontend changes Full-app runs need an OpenCode session + live model turn. For pure frontend component changes, mount the component in a throwaway vite page instead: 1. Create `apps/desktop/verify-<x>.html` (plain html, `<div id="root">`, `<script type="module" src="/src/verify-<x>.tsx">`) and `apps/desktop/src/verify-<x>.tsx` (ReactDOM.createRoot, import `./index.css` for Tailwind, import the component via `@/`). Vite serves any .html under `apps/desktop/` automatically. 2. `cd apps/desktop && npx vite --port 5199 --strictPort` (background). The npx wrapper may double-spawn and report "port in use" failure while the first instance is fine — check `lsof -iTCP:5199` before retrying. 3. Drive with the browser-control skill (Chrome HTTP API, port 9528): `createWindow` → `waitForSelector` → `readDom` → `screenshotTab` (screenshots land in `~/Downloads/`). Gotchas: `evalScript` is blocked by CSP on the extension — use `readDom`/`scrollTab` instead; `readDom` requires an explicit `"attributes": []` or it errors with "Value is unserializable". 4. Clean up: `closeTab`, kill the vite pid, delete both harness files, confirm `git status` is clean.
A test skill that says hello. Use when you want to test skill loading or verify that the skill system is working.
Use whenever you write or run scientific analysis code (physics, earth/geo, biology, chemistry, or social science) in this workspace — before executing it and again after generating results. Runs a deterministic domain-correctness gate that catches code which runs but is scientifically wrong (unit/dimension mismatch, Euclidean distance on lat/lon without a CRS, 0-based/1-based coordinate and strand errors, impossible SMILES valence, uncorrected multiple comparisons, averaging a categorical code). Surfaces structured findings; never claims the code is correct.
Use BEFORE reading any data file that could be large (CSV/TSV, Parquet, HDF5, FITS, NetCDF, NDJSON, genomics FASTQ/FASTA/VCF/BAM, GRIB, ROOT, or big text/simulation logs like VASP OUTCAR). Returns a compact memory pointer — header/schema/shape/sample/key numbers — by introspection and sampling in bounded memory, so you never load a file bigger than the context window into the model. Reference data via the pointer; read specific ranges deterministically.
Use when the user asks to run heavy or GPU work on Modal (the cloud compute platform) — writing a Modal function in the workspace, running it with the user's own `modal` CLI + token, and bringing results back. Data-to-compute for jobs too big for the laptop, without a Slurm cluster.
Use whenever you generate or review a chart, plot, table, or paper figure in this workspace, including work delegated by paper-writing, literature-survey, and experiment skills. Applies the Open Science publication style, enforces readable final-size layout for figures and tables, and rejects generic diagram-tool output as a publication figure. Interactive Plotly/HTML may be used for exploration, but paper delivery requires a static publication-ready export.
Use when the user asks to run, submit, monitor, or cancel a job on a remote machine over SSH — their own GPU/CPU server, a workstation, or a Slurm cluster ("the cluster", a login node, "my 3090 box", "the compute server"). Picks a saved machine, runs the work directly over SSH (or via Slurm when present), tracks it, and fetches results back into the workspace.
Use whenever you run statistical analysis for the social sciences (regression, hypothesis tests, econometrics) or read Stata (.dta) / SPSS (.sav) data in this workspace. Enforces an execute-don't-interpret boundary (surface estimates, don't volunteer causal claims), checks the analysis against a preregistration plan for HARKing, verifies reproducible seeds, and reproduces .dta/.sav estimates via R. Flags integrity risks; never certifies the analysis is sound.
Use when the user asks to review, verify, or audit a report, manuscript, or analysis in the workspace for traceability — resolving citations, flagging numbers with no source, and checking figures against the code that generated them. Emits a structured review block the app renders as reviewer findings. Verifies traceability, never "correctness".