跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

fair-esm2

Embed proteins with Meta AI's ESM-2 (`fair-esm` package). Use this skill when: (1) Extracting per-residue or per-sequence embeddings for downstream ML, (2) Masked-LM likelihood / mutation effect scoring, (3) Contact prediction from a sequence.

科研5.5kresources/skills/fair-esm2/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/aipoch/open-science/fair-esm2/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

fair-esm2 — ESM-2 (Meta AI)

ESM-2 code and weights are MIT (Meta AI, github.com/facebookresearch/esm).

Package disambiguation. pip install fair-esm gives you import esm with esm.pretrained.* (ESM-1/2). Biohub's github.com/Biohub/esm fork (MIT) gives you from esm.models.esmfold2 import ESMFold2InputBuilder — see the esmfold2 skill. Both share the esm namespace but are different libraries. This skill covers fair-esm (the Meta package).

Prerequisites

RequirementMinimumRecommended
Python3.8+3.11
CUDA11.7+12.x
GPU VRAM8 GB (8M), 16 GB (650M)24 GB+ (650M / 3B)

How to run

Embeddings

import torch, esm

model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
model = model.eval().cuda()
bc = alphabet.get_batch_converter()

_, _, toks = bc([("ubq", "MQIFVKTLTGKTITLEVEPSDTIENVK")])
with torch.no_grad():
    out = model(toks.cuda(), repr_layers=[33])
emb = out["representations"][33]      # (1, L+2, 1280) — includes BOS/EOS
seq_emb = emb[0, 1:-1].mean(0)        # per-sequence mean

Masked-LM scoring

with torch.no_grad():
    out = model(toks.cuda(), repr_layers=[33])
logits = out["logits"][0, 1:-1]       # (L, |vocab|)
# WT marginal log-likelihood; for mutation scoring, mask the position and
# compare logit[mut] − logit[wt].

Contact prediction

with torch.no_grad():
    out = model(toks.cuda(), repr_layers=[33], return_contacts=True)
contacts = out["contacts"][0]         # (L, L)

Models

NameLayersDimParamsUse
esm2_t6_8M_UR50D63208 MFast smoke / tiny embeddings
esm2_t33_650M_UR50D331280650 MDefault embedding model
esm2_t36_3B_UR50D3625603 BBest embeddings, 24 GB+

Output format

out["representations"][layer] is (B, L_max+2, D), where L_max is the longest tokenized residue sequence in the batch. ESM-2 adds BOS/EOS and pads shorter sequences after EOS. The single-sequence 1:-1 slice above is valid without padding; in a mixed-length batch it includes EOS and may include padding for shorter sequences.

For a batch, count non-padding tokens separately for each sequence, then remove BOS/EOS before pooling. Here toks and out must come from the same batch:

token_lengths = (toks != alphabet.padding_idx).sum(1).tolist()  # includes BOS/EOS
emb = out["representations"][33]
residue_embs = [emb[i, 1 : token_length - 1] for i, token_length in enumerate(token_lengths)]
seq_embs = torch.stack([residues.mean(0) for residues in residue_embs])  # (B, D)

Use nonempty protein sequences. Keep the batch order when associating embeddings with sequence IDs. out["contacts"] (when return_contacts=True) has shape (B, L_max, L_max); for sequence i, retain only out["contacts"][i, : token_lengths[i] - 2, : token_lengths[i] - 2].

Remote compute

Needs ≥16 GB VRAM (650M model) and either pre-cached .pt checkpoints or egress to dl.fbaipublicfiles.com. Read compute_details({provider, mode:'read'}) for an environment with fair-esm and a torch-hub weight cache, then:

c = host.compute.create(provider)
job = c.submitJob(
    intent="ESM-2 650M embeddings for 200 sequences — 1×GPU, ~2 min",
    inputs=[
        {"src": "seqs.fasta", "dstFilename": "seqs.fasta"},
        {"src": "embed_esm2.py", "dstFilename": "embed_esm2.py"},
    ],
    command="python3 embed_esm2.py",
    environment=...,   # env name from compute_details
    outputs=["embeddings.pt"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking c.attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for details.

Inside embed_esm2.py, set TORCH_HOME to the provider's torch-hub cache mount (path is in compute_details) so esm.pretrained.* resolves locally.

Troubleshooting

SymptomCauseFix
ModuleNotFoundError: No module named 'esm.models'You want Biohub's esm fork, not fair-esmSee esmfold2 skill; this skill uses esm.pretrained.*
Slow first callDownloading weights via torch.hubSet TORCH_HOME to a cached location

Next: feed embeddings to a classifier. For structure prediction, use esmfold2.

相似的 Skill

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

科研

adaptyv
K-Dense-AI/scientific-agent-skills48k

adaptyv

Uses the Adaptyv Bio Foundry API and Python SDK to design protein characterization experiments, estimate costs, submit sequences, monitor laboratory progress, and retrieve results. Applies to Adaptyv Foundry, its target catalog, binding screening and affinity assays, thermostability, expression, fluorescence, epitope binning, and enzyme activity workflows, including code using adaptyv or FoundryClient.

科研

anndata
K-Dense-AI/scientific-agent-skills48k

anndata

Handles annotated matrices in single-cell analysis, .h5ad and Zarr files, and integration with the scverse ecosystem. This is the data format skill—for analysis workflows use scanpy; for probabilistic models use scvi-tools; for population-scale queries use cellxgene-census.

科研

deepchem
K-Dense-AI/scientific-agent-skills48k

deepchem

Builds molecular property prediction and MoleculeNet workflows with DeepChem, including SMILES featurization, scaffold or grouped holdouts, masked labels, graph models and explicit pretrained encoder transfer. Used for ADMET, toxicity, solubility and chemistry ML when DeepChem data/model contracts and scientific validation are needed.

科研

arboreto
K-Dense-AI/scientific-agent-skills48k

arboreto

Infers candidate gene regulatory networks from bulk or single-cell expression data using AertsLab Arboreto GRNBoost2 and GENIE3. Use for transcription factor-target association ranking, compatible Dask execution, sparse expression inputs, and network stability checks.

科研

cirq
K-Dense-AI/scientific-agent-skills48k

cirq

Google quantum computing framework. Use when targeting Google Quantum AI hardware, designing noise-aware circuits, or running quantum characterization experiments. Best for Google hardware, noise modeling, and low-level circuit design. For IBM hardware use qiskit; for quantum ML with autodiff use pennylane; for physics simulations use qutip.

科研