git clone --depth 1 https://github.com/UnicomAI/wanwu /tmp/scgpt && cp -r /tmp/scgpt/configs/microservice/bff-service/configs/agent-skills/claude-science/scgpt ~/.claude/skills/scgptSKILL.md
# scGPT — Single-Cell Foundation Model
## Prerequisites
| Requirement | Minimum | Recommended |
| ----------- | ------- | ----------- |
| Python | 3.10+ | 3.11 |
| CUDA | 12.1+ | 12.4+ |
| GPU VRAM | 16 GB | 24 GB+ |
## How to run
### Loading the vocabulary and checkpoint
scGPT checkpoints are **raw directories** (`args.json`, `best_model.pt`,
`vocab.json`) — not Hugging Face hub repos. Point at the directory, not an HF
repo id.
```python
from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv)) # 60697 for the released human checkpoint
```
### Embedding an AnnData
```python
import anndata as ad
from scgpt.tasks import embed_data
adata = ad.read_h5ad("dataset.h5ad") # var must contain a gene-name column
emb = embed_data(
adata,
model_dir="/path/to/scgpt-human",
gene_col="feature_name",
use_fast_transformer=False, # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]
```
## Output format
`embed_data` returns an `AnnData` whose `.obsm["X_scGPT"]` is the per-cell
embedding (`n_cells × emb_dim`, 512 by default). Downstream: feed to
`scanpy.pp.neighbors` / `scanpy.tl.umap`.
## Remote compute
Needs ≥24 GB VRAM and the released human checkpoint (~200 MB:
`args.json`, `best_model.pt`, `vocab.json`). Read
`compute_details({provider, mode:'read'})` for an environment with `scgpt`
and a pre-cached checkpoint directory, then:
```python
c = host.compute.create(provider)
job = c.submit_job(
intent="scGPT embed 50k cells — 1×GPU, ~5 min",
inputs=[
{"src": "dataset.h5ad", "dst_filename": "dataset.h5ad"},
{"src": "embed.py", "dst_filename": "embed.py"},
],
command="python3 embed.py",
environment=..., # env name from compute_details
outputs=["embedded.h5ad"],
timeout_seconds=1800,
)
print(job.job_id) # cell ends here — kernel never blocks on compute
```
Then call the `wait_for_notification` brain-tool. When the
`compute_done` notification arrives, act on its payload:
```python
save_artifacts(payload["featured_files"]) # paths under hpc/<job_id>/
```
For the full result dict (`output_files`, `remote_workdir`, …), re-enter the
kernel and bind the compute handle separately — `.close()` lives on the
handle, not on the job object:
```python
h = host.compute.create(provider)
res = h.attach_job(job_id).result()
h.close()
```
See the `remote-compute-ssh` / `remote-compute-modal` skill for the
orchestration details.
In `embed.py`, pass `model_dir=` the checkpoint path from `compute_details`.
If `flash-attn` is unavailable in that environment, set
`use_fast_transformer=False`.
## Gotchas
- **`use_fast_transformer` default is `True`** but resolves to a FlashAttention
path that may not import in every env. Pass `use_fast_transformer=False`
unless you've confirmed `flash_attn` loads cleanly.
- The package historically depended on `torchtext.vocab.Vocab`; in
environments without torchtext a pure-Python shim provides `Vocab` —
functionally identical for `GeneVocab`, but if you hit
`AttributeError: 'Vocab' object has no attribute …`, you're on a stale shim.
- Gene names must match the vocab; unmatched genes are dropped. Set
`gene_col` to the column in `adata.var` that holds symbols.
## Troubleshooting
| Symptom | Fix |
| ------------------------------------------------- | ------------------------------------------------ |
| `flash_attn is not installed` warning at import | Harmless; pass `use_fast_transformer=False` |
| `'Vocab' object has no attribute 'vocab'` | Env has an old torchtext shim — update the env |
| Nearly all genes dropped | Wrong `gene_col`; check `adata.var.columns` |
| "scgpt not in manifest" / env-detection misses scGPT | The baked env manifest lists the distribution as `scGPT` (and `flash_attn`), pip's canonical casing — normalize manifest keys before lookup: `name.lower().replace('-', '_')` |
---
**Next**: cluster/annotate the embedding with the scanpy library
(`sc.pp.neighbors` → `sc.tl.leiden` / `sc.tl.umap`), or compare to an
scvi-tools latent space on the same data.万悟平台 SSE 子会话递归嵌套与三明治序列渲染架构指南。涵盖 parentId 领养、order 绝对排序、动静 Chunk 分层及 Vue 2 响应式引用协议。
Creating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems. Create original algorithmic art rather than copying existing artists' work to avoid copyright violations.
Applies Anthropic's official brand colors and typography to any sort of artifact that may benefit from having Anthropic's look-and-feel. Use it when brand colors or style guidelines, visual formatting, or company design standards apply.
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
Build apps with the Claude API or Anthropic SDK. TRIGGER when: code imports `anthropic`/`@anthropic-ai/sdk`/`claude_agent_sdk`, or user asks to use Claude API, Anthropic SDKs, or Agent SDK. DO NOT TRIGGER when: code imports `openai`/other AI SDK, general programming, or ML/data-science tasks.
Guide users through a structured workflow for co-authoring documentation. Use when user wants to write documentation, proposals, technical specs, decision docs, or similar structured content. This workflow helps users efficiently transfer context, refine content through iteration, and verify the doc works for readers. Trigger when user mentions writing docs, creating proposals, drafting specs, or similar documentation tasks.
Use this skill whenever the user wants to create, read, edit, or manipulate Word documents (.docx files). Triggers include: any mention of 'Word doc', 'word document', '.docx', or requests to produce professional documents with formatting like tables of contents, headings, page numbers, or letterheads. Also use when extracting or reorganizing content from .docx files, inserting or replacing images in documents, performing find-and-replace in Word files, working with tracked changes or comments, or converting content into a polished Word document. If the user asks for a 'report', 'memo', 'letter', 'template', or similar deliverable as a Word or .docx file, use this skill. Do NOT use for PDFs, spreadsheets, Google Docs, or general coding tasks unrelated to document generation.
Create distinctive, production-grade frontend interfaces with high design quality. Use this skill when the user asks to build web components, pages, artifacts, posters, or applications (examples include websites, landing pages, dashboards, React components, HTML/CSS layouts, or when styling/beautifying any web UI). Generates creative, polished code and UI design that avoids generic AI aesthetics.