Skip to content
FunCoding

Search

Search docs, Skills and MCP

fair-esm2

Embed proteins with Meta AI's ESM-2 (`fair-esm` package). Use this skill when: (1) Extracting per-residue or per-sequence embeddings for downstream ML, (2) Masked-LM likelihood / mutation effect scoring, (3) Contact prediction from a sequence.

科研5.5kresources/skills/fair-esm2/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/aipoch/open-science/fair-esm2/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

fair-esm2 — ESM-2 (Meta AI)

ESM-2 code and weights are MIT (Meta AI, github.com/facebookresearch/esm).

Package disambiguation. pip install fair-esm gives you import esm with esm.pretrained.* (ESM-1/2). Biohub's github.com/Biohub/esm fork (MIT) gives you from esm.models.esmfold2 import ESMFold2InputBuilder — see the esmfold2 skill. Both share the esm namespace but are different libraries. This skill covers fair-esm (the Meta package).

Prerequisites

RequirementMinimumRecommended
Python3.8+3.11
CUDA11.7+12.x
GPU VRAM8 GB (8M), 16 GB (650M)24 GB+ (650M / 3B)

How to run

Embeddings

import torch, esm

model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
model = model.eval().cuda()
bc = alphabet.get_batch_converter()

_, _, toks = bc([("ubq", "MQIFVKTLTGKTITLEVEPSDTIENVK")])
with torch.no_grad():
    out = model(toks.cuda(), repr_layers=[33])
emb = out["representations"][33]      # (1, L+2, 1280) — includes BOS/EOS
seq_emb = emb[0, 1:-1].mean(0)        # per-sequence mean

Masked-LM scoring

with torch.no_grad():
    out = model(toks.cuda(), repr_layers=[33])
logits = out["logits"][0, 1:-1]       # (L, |vocab|)
# WT marginal log-likelihood; for mutation scoring, mask the position and
# compare logit[mut] − logit[wt].

Contact prediction

with torch.no_grad():
    out = model(toks.cuda(), repr_layers=[33], return_contacts=True)
contacts = out["contacts"][0]         # (L, L)

Models

NameLayersDimParamsUse
esm2_t6_8M_UR50D63208 MFast smoke / tiny embeddings
esm2_t33_650M_UR50D331280650 MDefault embedding model
esm2_t36_3B_UR50D3625603 BBest embeddings, 24 GB+

Output format

out["representations"][layer] is (B, L_max+2, D), where L_max is the longest tokenized residue sequence in the batch. ESM-2 adds BOS/EOS and pads shorter sequences after EOS. The single-sequence 1:-1 slice above is valid without padding; in a mixed-length batch it includes EOS and may include padding for shorter sequences.

For a batch, count non-padding tokens separately for each sequence, then remove BOS/EOS before pooling. Here toks and out must come from the same batch:

token_lengths = (toks != alphabet.padding_idx).sum(1).tolist()  # includes BOS/EOS
emb = out["representations"][33]
residue_embs = [emb[i, 1 : token_length - 1] for i, token_length in enumerate(token_lengths)]
seq_embs = torch.stack([residues.mean(0) for residues in residue_embs])  # (B, D)

Use nonempty protein sequences. Keep the batch order when associating embeddings with sequence IDs. out["contacts"] (when return_contacts=True) has shape (B, L_max, L_max); for sequence i, retain only out["contacts"][i, : token_lengths[i] - 2, : token_lengths[i] - 2].

Remote compute

Needs ≥16 GB VRAM (650M model) and either pre-cached .pt checkpoints or egress to dl.fbaipublicfiles.com. Read compute_details({provider, mode:'read'}) for an environment with fair-esm and a torch-hub weight cache, then:

c = host.compute.create(provider)
job = c.submitJob(
    intent="ESM-2 650M embeddings for 200 sequences — 1×GPU, ~2 min",
    inputs=[
        {"src": "seqs.fasta", "dstFilename": "seqs.fasta"},
        {"src": "embed_esm2.py", "dstFilename": "embed_esm2.py"},
    ],
    command="python3 embed_esm2.py",
    environment=...,   # env name from compute_details
    outputs=["embeddings.pt"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking c.attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for details.

Inside embed_esm2.py, set TORCH_HOME to the provider's torch-hub cache mount (path is in compute_details) so esm.pretrained.* resolves locally.

Troubleshooting

SymptomCauseFix
ModuleNotFoundError: No module named 'esm.models'You want Biohub's esm fork, not fair-esmSee esmfold2 skill; this skill uses esm.pretrained.*
Slow first callDownloading weights via torch.hubSet TORCH_HOME to a cached location

Next: feed embeddings to a classifier. For structure prediction, use esmfold2.

Similar Skills

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

Science

cellprofiler
K-Dense-AI/scientific-agent-skills48k

cellprofiler

Runs reproducible CellProfiler microscopy pipelines for nuclear segmentation, cell counts, and per-object fluorescence measurements. Supports image/channel manifests, headless batch execution, segmentation overlays, and measurement QC for 2D fluorescence assays.

Science

autoskill
K-Dense-AI/scientific-agent-skills48k

autoskill

Analyzes user-requested Screenpipe history windows to detect repeated research workflows, match existing scientific skills, and stage new skill drafts or composition recipes for review. Requires a reachable Screenpipe HTTP API, normally on localhost:3030. Detection and embedding inference run locally; the selected LLM receives redacted app/title cluster summaries and matched skill descriptions. Use only when the user explicitly asks to analyze their recent work and propose skills.

Science

arbor
K-Dense-AI/scientific-agent-skills48k

arbor

Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experiment research runs. Includes a standard-library state manager and guidance for the RUC-NLPIR Arbor CLI.

Science

cantera
K-Dense-AI/scientific-agent-skills48k

cantera

Runs Cantera homogeneous chemical reactors and evaluates ignition delay with mechanism provenance, conservation checks, and numerical refinement. Use for combustion kinetics, closed adiabatic ideal-gas constant-volume or constant-pressure ignition, temperature histories, or mechanism-specific ignition-delay comparisons.

Science

bgpt-paper-search
K-Dense-AI/scientific-agent-skills48k

bgpt-paper-search

Searches BGPT scientific papers by topic or DOI and retrieves claim-level evidence extracted from full text, including experiments, reported statistics, scope, limitations, and provenance. Use for literature reviews, evidence synthesis, and finding experimental details beyond abstracts.

Science