Skip to content
FunCoding

Search

Search docs, Skills and MCP

evo2

Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model. Use this skill when: (1) Computing per-nucleotide or per-sequence likelihoods for variant effect scoring, (2) Embedding genomic windows for downstream classification, (3) Generating DNA conditioned on a prefix, (4) Scoring regulatory or coding regions across species.

科研5.5kresources/skills/evo2/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/aipoch/open-science/evo2/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Evo 2 — DNA Language Model

Prerequisites

RequirementMinimumRecommended
Python3.113.12 (<3.13)
CUDA12.1+12.4+
GPU VRAM24 GB (7B bf16)80 GB (40B)
RAM32 GB128 GB

How to run

Installation

pip install evo2
# Weights pulled from Hugging Face on first model load.

Loading and scoring

from evo2 import Evo2

model = Evo2("evo2_7b")        # or "evo2_40b" — see model table
seqs = ["ATCG" * 50, "GGGCTTAA" * 25]
ll = model.score_sequences(seqs)   # → list[float], mean per-token log-likelihood
print(ll)

Generation

out = model.generate(
    prompt_seqs=["ATGAAAGCT"],
    n_tokens=256,
    temperature=0.7,
)
print(out.sequences[0])

Models

NameParamsContextVRAM (bf16)Notes
evo2_7b7 B1 M nt~22 GBDefault; fits on a single 24 GB+ GPU
evo2_40b40 B1 M nt~78 GBH100 80 GB or multi-GPU
evo2_1b_base1 B8 K nt~6 GBFP8 path requires sm_89+ (H100)

Output format

score_sequences returns a list[float] (or np.ndarray) of mean log-likelihoods, one per input sequence. More negative ⇒ less likely under the model. For variant effect, compute Δll = ll_alt - ll_ref over a fixed window.

generate returns a GenerationOutput with .sequences (list[str]), .logits (list[Tensor]), and .logprobs_mean (list[float]) — always populated, no flag required.

Decision tree

Need a DNA model?
│
├─ Per-base/per-sequence likelihood, generation → Evo 2 ✓
├─ Predict experimental tracks (expression, accessibility) → borzoi
└─ Protein, not DNA → fair-esm2 / esmfold2

Remote compute

7B/40B inference is GPU-bound (≥24 GB / 80 GB VRAM). Read compute_details({provider, mode:'read'}) for an environment with evo2 + flash-attn and a pre-cached HF weight mount, then submit:

c = host.compute.create(provider)
job = c.submitJob(
    intent="Evo2-7B score 200bp variant window — 1×GPU, ~2 min",
    inputs=[{"src": "score_evo2.py", "dstFilename": "score_evo2.py"}],
    command="python3 score_evo2.py",   # env selection is host-specific — see compute_details for your provider
    outputs=["scores.json"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking c.attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for details.

Inside score_evo2.py, point HF_HOME at the provider's weight-cache mount (path is in compute_details) and set HF_HUB_OFFLINE=1 so the loader doesn't try to write refs/ into a read-only mount. Weight footprint: ~15 GB (7B), ~80 GB (40B).

Typical performance

Task7B on H100Notes
Model load (cached)~5-7 minFirst call hydrates weights
score_sequences, 200×200bp~10-20 sAfter load
generate, 1×512 nt~15 s

Troubleshooting

SymptomCauseFix
Transformer Engine not installedNo FP8 — falls back to bf16Informational only on non-H100; ignore
OOM on load40B on <80 GB GPUUse evo2_7b or shard with device_map
HF tries to write refs/mainHF_HOME points at RO mountSet HF_HUB_OFFLINE=1
dtype mismatch in score_sequencesPassing tensors not stringsPass list[str]; the API tokenises for you

Next: pair with borzoi to predict track-level effects of the same variants.

Similar Skills

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

Science

adaptyv
K-Dense-AI/scientific-agent-skills48k

adaptyv

Uses the Adaptyv Bio Foundry API and Python SDK to design protein characterization experiments, estimate costs, submit sequences, monitor laboratory progress, and retrieve results. Applies to Adaptyv Foundry, its target catalog, binding screening and affinity assays, thermostability, expression, fluorescence, epitope binning, and enzyme activity workflows, including code using adaptyv or FoundryClient.

Science

anndata
K-Dense-AI/scientific-agent-skills48k

anndata

Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

Science

deepchem
K-Dense-AI/scientific-agent-skills48k

deepchem

Builds molecular property prediction and MoleculeNet workflows with DeepChem, including SMILES featurization, scaffold or grouped holdouts, masked labels, graph models and explicit pretrained encoder transfer. Used for ADMET, toxicity, solubility and chemistry ML when DeepChem data/model contracts and scientific validation are needed.

Science

arboreto
K-Dense-AI/scientific-agent-skills48k

arboreto

Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3. Use for transcription factor-target association ranking, compatible Dask execution, sparse expression inputs, and network stability checks.

Science

cirq
K-Dense-AI/scientific-agent-skills48k

cirq

Google quantum computing framework. Use when targeting Google Quantum AI hardware, designing noise-aware circuits, or running quantum characterization experiments. Best for Google hardware, noise modeling, and low-level circuit design. For IBM hardware use qiskit; for quantum ML with autodiff use pennylane; for physics simulations use qutip.

Science