跳到正文
FunCoding

搜索

搜索文档、Skill 和 MCP

solublempnn

Inverse-fold a backbone with SolubleMPNN — ProteinMPNN retrained on a soluble-PDB subset (Dauparas et al. 2022) — for sequences biased toward cytosolic expression and reduced aggregation. Reach for this skill when designs from vanilla ProteinMPNN are aggregating or going to inclusion bodies, when redesigning a membrane-adjacent fold for soluble expression, or when an E. coli expression screen is the next step.

科研5.5kresources/skills/solublempnn/SKILL.md

安装

把这段话发给 Claude Code、Codex 或 Cursor。智能体会先检查安全性,你确认后才安装。

读取 https://funcoding.ai/skills/aipoch/open-science/solublempnn/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

SolubleMPNN

SolubleMPNN is not a separate package — it is the ProteinMPNN architecture retrained on a soluble-PDB subset, which shifts the output distribution away from the surface hydrophobics that the full-PDB model happily places (because many of them are buried at crystallographic or membrane interfaces in the training set). Reach for it when the goal is soluble yield in a heterologous host; stick with proteinmpnn when native-like recovery matters more, since the soluble prior trades a few points of recovery for the surface bias. Code and weights are MIT (github.com/dauparas/ProteinMPNN, soluble_model_weights; also exposed via github.com/dauparas/LigandMPNN). The model is small enough to run on CPU — for a handful of sequences on one backbone that is seconds and usually faster than dispatching; a GPU helps for batched campaigns. Either way the repo is cloned in-job (no PyPI dist; checkpoints bundled).

Running it

pip install torch numpy   # if not already present
git clone --depth 1 https://github.com/dauparas/ProteinMPNN.git proteinmpnn
cd proteinmpnn
python protein_mpnn_run.py \
  --pdb_path backbone.pdb --pdb_path_chains "A" \
  --out_folder out --num_seq_per_target 16 \
  --sampling_temp "0.1" --use_soluble_model

The runner uses repo-relative imports, so the cd line is load-bearing — invoking the script by absolute path from elsewhere fails with ModuleNotFoundError. If you want threaded designed-sequence PDBs as well, the LigandMPNN runner accepts --model_type soluble_mpnn (see ligandmpnn for that path; it needs ProDy in addition to torch). The flag surface is otherwise identical to proteinmpnn (or ligandmpnn for the second form), including the string-typed temperature and the fixed-position JSONL keyed by PDB stem — see proteinmpnn for the parsing quirks. The repo ships soluble weights at v_48_010 and v_48_020 only; asking for --model_name v_48_002 --use_soluble_model errors on a missing checkpoint, so leave --model_name at its default.

Output is out/seqs/<stem>.fa with score= and seq_recovery= in each header. Expect recovery against a native structure to drop a few points relative to vanilla — that is the prior working, not a bug.

Hydrophobic surface patches still recur where the fold needs them

Soluble weights shift the distribution; they do not enforce a hydrophobicity ceiling. If a particular surface patch keeps coming back hydrophobic, that patch is likely structurally load-bearing and the network is paying the solubility cost to keep the fold. Layering --omit_AAs "CW" or a per-position bias on top is fine, but check that the resulting designs still fold (via boltz or esmfold2) before assuming the constraint was free.

"Crystallisable" training set ≠ "soluble in your host" — keep an orthogonal filter

The training set is "structures that were soluble enough to crystallise," which correlates with but is not the same as "expresses solubly in E. coli at 37 °C." For campaigns where expression yield is the bottleneck, rank the soluble-MPNN output by an orthogonal sequence-based predictor before committing wet-lab slots; treat the MPNN bias as widening the funnel, not replacing the filter.


Next: fold the designs with boltz or esmfold2 to confirm the backbone is still recovered, then carry survivors into the expression screen.

相似的 Skill

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

科研

cellprofiler
K-Dense-AI/scientific-agent-skills48k

cellprofiler

Runs reproducible CellProfiler microscopy pipelines for nuclear segmentation, cell counts, and per-object fluorescence measurements. Supports image/channel manifests, headless batch execution, segmentation overlays, and measurement QC for 2D fluorescence assays.

科研

autoskill
K-Dense-AI/scientific-agent-skills48k

autoskill

Analyzes user-requested Screenpipe history windows to detect repeated research workflows, match existing scientific skills, and stage new skill drafts or composition recipes for review. Requires a reachable Screenpipe HTTP API, normally on localhost:3030. Detection and embedding inference run locally; the selected LLM receives redacted app/title cluster summaries and matched skill descriptions. Use only when the user explicitly asks to analyze their recent work and propose skills.

科研

arbor
K-Dense-AI/scientific-agent-skills48k

arbor

Applies Arbor Hypothesis Tree Refinement to research artifacts with repeatable evaluators, including model training, agent harnesses, data synthesis and benchmark optimization. Uses persistent hypotheses, isolated experiments, evidence propagation and held-out candidate comparison for multi-experiment research runs. Includes a standard-library state manager and guidance for the RUC-NLPIR Arbor CLI.

科研

cantera
K-Dense-AI/scientific-agent-skills48k

cantera

Runs Cantera homogeneous chemical reactors and evaluates ignition delay with mechanism provenance, conservation checks, and numerical refinement. Use for combustion kinetics, closed adiabatic ideal-gas constant-volume or constant-pressure ignition, temperature histories, or mechanism-specific ignition-delay comparisons.

科研

bgpt-paper-search
K-Dense-AI/scientific-agent-skills48k

bgpt-paper-search

Searches BGPT scientific papers by topic or DOI and retrieves claim-level evidence extracted from full text, including experiments, reported statistics, scope, limitations, and provenance. Use for literature reviews, evidence synthesis, and finding experimental details beyond abstracts.

科研