Skip to content
FunCoding

Search

Search docs, Skills and MCP

scgpt

Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For probabilistic single-cell models (scVI etc.), use the scvi-tools library.

科研5.5kresources/skills/scgpt/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/aipoch/open-science/scgpt/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

scGPT — Single-Cell Foundation Model

Prerequisites

RequirementMinimumRecommended
Python3.10+3.11
CUDA12.1+12.4+
GPU VRAM16 GB24 GB+

How to run

Loading the vocabulary and checkpoint

scGPT checkpoints are raw directories (args.json, best_model.pt, vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF repo id.

from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv))   # 60697 for the released human checkpoint

Embedding an AnnData

import anndata as ad
from scgpt.tasks import embed_data

adata = ad.read_h5ad("dataset.h5ad")        # var must contain a gene-name column
emb = embed_data(
    adata,
    model_dir="/path/to/scgpt-human",
    gene_col="feature_name",
    use_fast_transformer=False,             # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]

Output format

embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell embedding (n_cells × emb_dim, 512 by default). Downstream: feed to scanpy.pp.neighbors / scanpy.tl.umap.

Remote compute

Needs ≥24 GB VRAM and the released human checkpoint (~200 MB: args.json, best_model.pt, vocab.json). Read compute_details({provider, mode:'read'}) for an environment with scgpt and a pre-cached checkpoint directory, then:

c = host.compute.create(provider)
job = c.submitJob(
    intent="scGPT embed 50k cells — 1×GPU, ~5 min",
    inputs=[
        {"src": "dataset.h5ad", "dstFilename": "dataset.h5ad"},
        {"src": "embed.py", "dstFilename": "embed.py"},
    ],
    command="python3 embed.py",
    environment=...,   # env name from compute_details
    outputs=["embedded.h5ad"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking host.compute.create(provider).attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for the orchestration details.

See the remote-compute-ssh / remote-compute-modal skill for the orchestration details.

In embed.py, pass model_dir= the checkpoint path from compute_details. If flash-attn is unavailable in that environment, set use_fast_transformer=False.

Gotchas

  • use_fast_transformer default is True but resolves to a FlashAttention path that may not import in every env. Pass use_fast_transformer=False unless you've confirmed flash_attn loads cleanly.
  • The package historically depended on torchtext.vocab.Vocab; in environments without torchtext a pure-Python shim provides Vocab — functionally identical for GeneVocab, but if you hit AttributeError: 'Vocab' object has no attribute …, you're on a stale shim.
  • Gene names must match the vocab; unmatched genes are dropped. Set gene_col to the column in adata.var that holds symbols.

Troubleshooting

SymptomFix
flash_attn is not installed warning at importHarmless; pass use_fast_transformer=False
'Vocab' object has no attribute 'vocab'Env has an old torchtext shim — update the env
Nearly all genes droppedWrong gene_col; check adata.var.columns
"scgpt not in manifest" / env-detection misses scGPTThe baked env manifest lists the distribution as scGPT (and flash_attn), pip's canonical casing — normalize manifest keys before lookup: name.lower().replace('-', '_')

Next: cluster/annotate the embedding with the scanpy library (sc.pp.neighbors → sc.tl.leiden / sc.tl.umap), or compare to an scvi-tools latent space on the same data.

Similar Skills

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

Science

cellprofiler
K-Dense-AI/scientific-agent-skills48k

cellprofiler

Runs reproducible CellProfiler microscopy pipelines for nuclear segmentation, cell counts, and per-object fluorescence measurements. Supports image/channel manifests, headless batch execution, segmentation overlays, and measurement QC for 2D fluorescence assays.

Science

autoskill
K-Dense-AI/scientific-agent-skills48k

autoskill

Analyzes user-requested Screenpipe history windows to detect repeated research workflows, match existing scientific skills, and stage new skill drafts or composition recipes for review. Requires a reachable Screenpipe HTTP API, normally on localhost:3030. Detection and embedding inference run locally; the selected LLM receives redacted app/title cluster summaries and matched skill descriptions. Use only when the user explicitly asks to analyze their recent work and propose skills.

Science

arbor
K-Dense-AI/scientific-agent-skills48k

arbor

Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experiment research runs. Includes a standard-library state manager and guidance for the RUC-NLPIR Arbor CLI.

Science

cantera
K-Dense-AI/scientific-agent-skills48k

cantera

Runs Cantera homogeneous chemical reactors and evaluates ignition delay with mechanism provenance, conservation checks, and numerical refinement. Use for combustion kinetics, closed adiabatic ideal-gas constant-volume or constant-pressure ignition, temperature histories, or mechanism-specific ignition-delay comparisons.

Science

bgpt-paper-search
K-Dense-AI/scientific-agent-skills48k

bgpt-paper-search

Searches BGPT scientific papers by topic or DOI and retrieves claim-level evidence extracted from full text, including experiments, reported statistics, scope, limitations, and provenance. Use for literature reviews, evidence synthesis, and finding experimental details beyond abstracts.

Science