跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

scvi-tools

Probabilistic single-cell RNA-seq with scvi-tools — scVI for a batch-corrected latent space, scANVI for semi-supervised label transfer, and Bayesian differential expression. Reach for this skill to integrate scRNA-seq batches, embed cells for clustering, transfer annotations from a reference onto a query, or score differentially expressed genes per cluster. For spatial deconvolution / mapping use the cell2location, DestVI, or Tangram methods instead.

科研5.5kresources/skills/scvi-tools/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/aipoch/open-science/scvi-tools/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

scvi-tools — scVI / scANVI

scvi-tools (Gayoso et al. 2022, github.com/scverse/scvi-tools, BSD-3-Clause) wraps a family of deep generative models for single-cell omics. The scRNA-seq core is scVI (unsupervised batch-corrected latent embedding) and scANVI (scVI + a classifier head for semi-supervised cell-type label transfer). Both expect raw integer UMI counts and emit a low-dimensional X_scVI / X_scANVI that drops into the scanpy neighbors → leiden → umap pipeline.

Setup (any agent, no API key)

This is a pure skill — kernel.py is deterministic Python and you (the base model) do all the reasoning. There is no host runtime and no LLM API. The helpers are prepare_scvi_counts for input validation and h5ad_safe_obs for obs/var frames that anndata.write_h5ad() can serialize. Load them once per session in a Python cell:

exec(open("scvi-tools/kernel.py", encoding="utf-8").read())   # path to this skill's kernel.py

Nothing auto-loads it outside Claude Science. Then call the helpers directly. If a helper raises NameError, you haven't exec'd kernel.py.

Dependencies: pip install scvi-tools scanpy anndata. Training needs a CUDA-capable GPU — see Remote compute to fall out to a rented GPU when you don't have one locally.

How to run

scVI — batch-corrected latent space

prepare_scvi_counts(adata) checks and preserves existing counts. If absent, it checks .X before copying it to counts. Values must be finite, nonnegative, and integer-valued, with at least one positive count. Invalid existing counts raise instead of being replaced by .X; sparse matrices stay sparse.

These numerical checks cannot prove raw-count provenance. Check the dataset documentation or preprocessing history; do not round or exponentiate transformed values to make them pass. If the history is unclear, resolve it before training. The helper does not use .raw.X, which may be normalized or have a different gene axis. Use an in-memory AnnData object; materialize backed data or copy views only within the available memory budget.

import scanpy as sc
import scvi

adata = sc.read_h5ad("dataset.h5ad")
# Verify count provenance from the input documentation or preprocessing history.
# Preserve counts if present; otherwise validate .X before copying it to counts.
counts_record = prepare_scvi_counts(adata)
print(counts_record)       # numerical checks only; does not certify provenance
adata.X = adata.layers["counts"].copy()        # derive plotting/HVG data from verified counts
sc.pp.normalize_total(adata); sc.pp.log1p(adata) # optional, for HVG / plotting only
sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key="batch", subset=True)

scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="batch")
model = scvi.model.SCVI(adata, n_latent=30)
model.train(max_epochs=200, early_stopping=True, accelerator="gpu", devices=1)

adata.obsm["X_scVI"] = model.get_latent_representation()
adata.layers["scvi_normalized"] = model.get_normalized_expression(library_size=1e4)

scANVI — label transfer from a partially-annotated reference

lvae = scvi.model.SCANVI.from_scvi_model(
    model, labels_key="cell_type", unlabeled_category="Unknown",
)
lvae.train(max_epochs=20, n_samples_per_label=100, accelerator="gpu", devices=1)

adata.obsm["X_scANVI"] = lvae.get_latent_representation()
adata.obs["pred_cell_type"] = lvae.predict()

accelerator="gpu", devices=1 is the PyTorch-Lightning spelling; the legacy use_gpu= kwarg was removed in scvi-tools 1.x and now raises TypeError.

Differential expression

de = model.differential_expression(
    groupby="leiden", group1="3",   # group2=None → vs. all other cells
    mode="change", delta=0.25,
)
top = de.sort_values("proba_de", ascending=False).head(50)

For one-vs-rest leave group2 out — "rest" is scanpy's rank_genes_groups convention, not scvi-tools'; here group2 is a literal category name and "rest" would match zero cells.

scvi-tools ≥1.4 defaults to mode="vanilla", whose result columns are exactly:

['proba_m1', 'proba_m2', 'bayes_factor', 'scale1', 'scale2', 'raw_mean1',
 'raw_mean2', 'non_zeros_proportion1', 'non_zeros_proportion2',
 'raw_normalized_mean1', 'raw_normalized_mean2', 'comparison', 'group1',
 'group2']

— no lfc_*, no proba_de, no is_de_fdr_*. Pass mode="change" to get lfc_mean / lfc_median / proba_de / is_de_fdr_0.05. Sort on proba_de (or on bayes_factor if you deliberately stayed in vanilla mode).

Output format

KeyWhat
adata.obsm["X_scVI"]n_cells × n_latent batch-corrected embedding
adata.obsm["X_scANVI"]label-aware embedding (better separates known classes)
adata.obs["pred_cell_type"]scANVI predicted label per cell
adata.layers["scvi_normalized"]decoded expression, library-size normalized
DE dataframeper-gene lfc_* / proba_de (with mode="change")

Remote compute (rent a GPU)

An A100-class GPU is recommended for >50k cells. Training is a plain Python script (pipeline.py) that reads counts, trains scVI/scANVI, and writes the output .h5ad — run it on whatever GPU you have (a local/cluster CUDA box, or a serverless GPU host such as Modal). There is no Claude-Science compute broker here; drive the GPU host directly.

Modal (serverless GPU) — wrap pipeline.py in a Modal app and run it with the Modal CLI (modal run pipeline.py), which blocks until the job finishes, so you read the result synchronously (no notification tool needed):

# pipeline.py — run with:  modal run pipeline.py
import modal
image = (modal.Image.debian_slim()
         .pip_install("scvi-tools==1.4.2", "scanpy==1.11.5", "anndata==0.11.4"))
app = modal.App("scvi-run", image=image)
vol = modal.Volume.from_name("scvi-data", create_if_missing=True)  # holds dataset.h5ad / out.h5ad

@app.function(gpu="A100", timeout=3600, volumes={"/data": vol})
def train():
    import scanpy as sc, scvi   # noqa
    adata = sc.read_h5ad("/data/dataset.h5ad")
    # ... setup_anndata / scVI / scANVI / DE — see the recipe above ...
    adata.obs = h5ad_safe_obs(adata.obs)   # paste the helper into THIS script (below)
    adata.write_h5ad("/data/out.h5ad")
    vol.commit()

@app.local_entrypoint()
def main():
    train.remote()   # blocks until done; then read /data/out.h5ad from the volume

Helpers loaded via exec in your local session (see Setup) are not defined in a remote pipeline.py. Include prepare_scvi_counts from kernel.py before the remote preprocessing/training recipe, and h5ad_safe_obs before .write_h5ad(). Copy the helper definitions into that script or ship and load kernel.py there; keep the documented count-source selection and validation.

For a local/cluster GPU, just run pipeline.py directly where CUDA is visible — no wrapper needed. (For a fuller Modal workflow see the remote-compute-modal skill.)

Gotchas

GotchaWhat happens / fix
differential_expression() defaults to mode="vanilla" (scvi-tools ≥1.4)KeyError: 'lfc_mean' / 'proba_de' when sorting — pass mode="change" to get lfc_*/proba_de/is_de_fdr_*; in vanilla mode sort on bayes_factor.
adata.obs index/columns are string[pyarrow] (ArrowStringArray).write_h5ad() dies with IORegistryError: No method registered for writing <class 'pandas.arrays.ArrowStringArray'> (anndata #2377). Coerce before writing: adata.obs = h5ad_safe_obs(adata.obs) (kernel helper — load it via exec locally, see Setup; inline the coercion in a remote pipeline.py). .astype(str) alone is not enough — on a pyarrow-backed Index/Series it returns another Arrow-backed array; round-trip through np.asarray(..., dtype=object). anndata.settings.allow_write_nullable_strings = True does not cover Arrow-backed strings.
use_gpu= kwargRemoved in 1.x → TypeError: train() got an unexpected keyword argument 'use_gpu'. Use accelerator="gpu", devices=1.
Log-normalized data fed to setup_anndataSilent garbage — scVI's NB likelihood needs raw integer counts. Verify count provenance, call prepare_scvi_counts, and pass layer="counts". Never overwrite counts from an unverified .X.

Troubleshooting

SymptomFix
KeyError: 'lfc_mean' (or 'proba_de', 'is_de_fdr_0.05') on DE resultAdd mode="change" to differential_expression(); the default vanilla mode has no LFC columns.
IORegistryError: No method registered for writing <class 'pandas.arrays.ArrowStringArray'> on .write_h5ad()adata.obs = h5ad_safe_obs(adata.obs) (and adata.var if needed) before writing. The allow_write_nullable_strings flag does not help here.
TypeError: ... unexpected keyword argument 'use_gpu'Replace with accelerator="gpu", devices=1.
ValueError: ... non-negative integers / NB loss explodeslayer="counts" points at log/float data — restore raw counts.
MisconfigurationException: No supported gpu backend foundNo CUDA visible — drop accelerator/devices to fall back to CPU, or dispatch via Remote compute.
UnicodeEncodeError: 'ascii' codec can't encode character ... writing a summary / printingContainer has no LANG so Python defaults to ASCII. Open files with encoding="utf-8" and/or sys.stdout.reconfigure(encoding="utf-8") at script top, or set PYTHONIOENCODING=utf-8 in the image.

Next: cluster on X_scVI with scanpy (sc.pp.neighbors(use_rep="X_scVI") → sc.tl.leiden → sc.tl.umap); for spatial deconvolution train cell2location / DestVI / Tangram on the scRNA-seq reference.

相似的 Skill

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

科研

13c-metabolic-flux
K-Dense-AI/scientific-agent-skills48k

13c-metabolic-flux

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional isotopomers, parallel tracer experiments, and determining whether labeling data constrain a pathway flux. Distinguishes measured-label inference from COBRA flux balance analysis and flags experiments requiring nonstationary MFA.

科研

datamol
K-Dense-AI/scientific-agent-skills48k

datamol

Pythonic wrapper around RDKit with simplified interface and sensible defaults. Preferred for standard drug discovery including SMILES parsing, standardization, descriptors, fingerprints, clustering, 3D conformers, parallel processing. Returns native rdkit.Chem.Mol objects. For advanced control or custom parameters, use rdkit directly.

科研

biopython
K-Dense-AI/scientific-agent-skills48k

biopython

Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Supports batch processing, custom molecular-biology pipelines, BLAST automation, structure analysis, and motif analysis.

科研

bulk-rnaseq
K-Dense-AI/scientific-agent-skills48k

bulk-rnaseq

Prepares bulk RNA-seq FASTQ, Salmon, STAR or featureCounts output for gene-level differential expression. Covers nf-core/rnaseq and standalone quantification, biological replication, strandedness, reference provenance, validated count assembly and a PyDESeq2 handoff. Use for FASTQ-to-counts analysis, nf-core/rnaseq configuration, STAR/Salmon quantification, or building a counts matrix for DESeq2. For single-cell data use scanpy; for statistical fitting alone use pydeseq2.

科研

alphagenome
K-Dense-AI/scientific-agent-skills48k

alphagenome

Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), scores variants or scans windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and builds Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.

科研