Skip to content
FunCoding

Search

Search docs, Skills and MCP

oma-scholar

Search academic literature and generate, validate, or compare Knows paper sidecars. Use for claim/evidence analysis and literature synthesis.

科研1.3kskills/oma-scholar/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent checks the Skill for safety first and installs it only after you confirm.

读取 https://funcoding.ai/skills/first-fluke/oh-my-agent/oma-scholar/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Scholar - Research Paper Sidecar Companion

Scheduling

Goal

Search, fetch, generate, validate, analyze, review, and compare scholarly paper sidecars using the Knows .knows.yaml spec for token-efficient research workflows.

Intent signature

  • User asks for academic literature search, sidecar generation, sidecar validation, paper claims/evidence summary, structural paper comparison, or peer review as sidecar.
  • User references Knows, .knows.yaml, knows.academy, OpenAlex, claims, evidence, relations, or paper sidecars.

When to use

  • Reading research papers token-efficiently via Knows sidecars (~700 tokens for claims-only vs ~10K for full PDF)
  • Generating .knows.yaml sidecars from your own paper drafts, LaTeX, or research notes
  • Validating sidecar structure (rule-based) before sharing
  • Producing peer reviews as sidecars
  • Querying or summarizing existing sidecars
  • Structurally comparing two papers (claims, methods, evidence)
  • Searching/fetching sidecars from knows.academy (2026 papers only; current counts via /api/proxy/jobs/stats)

When NOT to use

  • General web search or non-academic content -> use oma-search
  • Translating papers -> use oma-translation
  • PDF parsing only (no sidecar) -> use oma-pdf
  • Submitting sidecars back to knows.academy -> out of scope (host LLM only consumes/produces locally)
  • Full peer-review workflow with editor system -> out of scope

Expected inputs

  • Paper, abstract, draft, LaTeX, research notes, sidecar file, DOI, OpenAlex ID, Knows record ID, or search query
  • Desired mode: generate, validate, review, analyze, compare, or remote fetch
  • Optional strictness, section filter, or CI behavior

Expected outputs

  • .knows.yaml sidecar, review sidecar, lint report, search/fetch result, natural-language analysis, or structural comparison
  • Sidecars conforming to v0.9.0 / paper@1 profile
  • Validation status and warnings before sharing generated sidecars

Dependencies

  • oma scholar CLI subcommands
  • knows.academy public API; OpenAlex and Semantic Scholar fallbacks
  • resources/sidecar-spec.md, API endpoints, OpenAlex setup, upstream cache, checklist, and execution protocol

Control-flow features

  • Branches by mode, source availability, Knows/OpenAlex coverage, strict vs lenient validation, and fetched section
  • Reads/writes YAML sidecars and may call public APIs
  • Avoids fabrication when source evidence is missing

Structural Flow

Entry

  1. Identify mode and source artifact/query.
  2. Resolve paper identity through Knows or OpenAlex when needed.
  3. Load sidecar spec and mode-specific protocol.

Scenes

  1. PREPARE: Select mode and gather source or remote identifiers.
  2. ACQUIRE: Fetch paper metadata, sidecar sections, or local source text.
  3. REASON: Extract claims, evidence, relations, provenance, or comparison structure.
  4. ACT: Generate, lint, review, analyze, compare, or fetch sidecar data.
  5. VERIFY: Check the supported local contract, IDs, relations, and provenance; report lint limitations. Full canonical conformance requires the exact schema validator.
  6. FINALIZE: Return sidecar, report, summary, or comparison with caveats.

Transitions

  • If knows.academy lacks the paper, fall back to OpenAlex metadata/abstract.
  • If generating a sidecar, run lint before sharing.
  • If consuming third-party sidecars with dangling references, use lenient mode when appropriate.
  • If source evidence is absent, omit fields instead of guessing.

Failure and recovery

  • If remote API times out, retry or use OpenAlex fallback.
  • If YAML fails parsing, fix indentation and scalar types.
  • If relation density or orphan statements warn, review coverage and type-appropriate links. Preserve source-anchored questions/definitions; never add unsupported relations to satisfy a quota.

Exit

  • Success: requested sidecar operation completes with validation status.
  • Partial success: missing metadata, fallback source, or validation warnings are explicit.

Logical Operations

Actions

ActionSSL primitiveEvidence
Select modeSELECTGenerate/Validate/Review/Analyze/Compare/Remote
Read paper or sidecarREADSource files or YAML
Request remote dataREQUESTKnows/OpenAlex APIs
Infer claims/evidence/relationsINFERSidecar generation/analysis
Write sidecarWRITE.knows.yaml outputs
Validate sidecarVALIDATEoma scholar lint
Report resultNOTIFYSummary or lint report

Tools and instruments

  • oma scholar search|resolve|get|lint
  • Knows public API, OpenAlex fallback, sidecar spec, checklist

Canonical command path

oma scholar search "<query>"
oma scholar resolve "<title-or-doi>"
oma scholar get "<record-id-or-doi>"
oma scholar lint "<paper.knows.yaml>"

Resource scope

ScopeResource target
LOCAL_FSPaper drafts, sidecar YAML, review sidecars
NETWORKknows.academy and OpenAlex APIs
PROCESSoma scholar CLI and lint
USER_DATAUser-provided paper content and research notes

Preconditions

  • Mode and source are identifiable.
  • Spec rules are available for generation or validation.

Effects and side effects

  • May create local sidecar or review sidecar files.
  • May query public scholarly APIs.
  • Does not submit sidecars back to knows.academy.

Guardrails

  1. Local generation targets v0.9.0 production compatibility / paper@1 (review sidecars use review@1). This is narrower than complete canonical schema conformance; see resources/sidecar-spec.md.
  2. Host LLM generates sidecars: never shell out to anthropic SDK or external LLM CLI; this skill runs inside an agent
  3. Anti-fabrication: if DOI/venue/year is not visible in source, omit the key entirely; never write doi: TODO or guess
  4. Top-level metadata: title, authors, venue, year live at the top level (no metadata wrapper)
  5. Field names are exact: statement_type, evidence_type, predicate, artifact_type (not type/claim)
  6. Provenance has SINGLE actor: provenance.actor is one object, NOT a provenance.actors array
  7. Confidence is an object: {claim_strength: ..., extraction_fidelity: ...}, both from high|medium|low
  8. Coverage is an object: coverage.statements (4-value enum) + coverage.evidence (3-value enum)
  9. Closed enums: actor tool|person|org (never ai/llm/model); artifact role subject|supporting|cited; predicates in present tense
  10. Numbers unquoted: value: 22, never value: '22'
  11. Relation coverage: review the local ratio/orphan warnings; add only source-supported, type-appropriate relations. Claims with evidence use supported_by; definitions and open questions can remain source-anchored. A ratio of 1.5 is diagnostic, not a graph-padding quota.
  12. ID format: descriptive kebab-case with prefix: stmt:privacy-budget-tradeoff, ev:cifar10-accuracy-table, art:paper
  13. Validate before sharing: run oma scholar lint after Generate and Review; cross-record refs (record_id#local_id) in review sidecars are recognized and not flagged as dangling
  14. Remote API has no auth: https://knows.academy/api/proxy/* is public; do not invent auth headers
  15. Partial fetch param is section (singular): fixed enum statements|evidence|relations|artifacts|citation
  16. Fallback API keys are optional: OPENALEX_API_KEY (metadata enrichment) and S2_API_KEY (Semantic Scholar dedicated rate limit) both improve throughput but the cascade degrades gracefully without them
  17. Sidecar content stays English: schema fields, IDs, statement text follow upstream convention; user-facing responses follow oma-config.yaml language
  18. Contract scope: local generation rules and lint track production compatibility, not the complete canonical v0.9 schema. Preserve canonical imported shapes and report local-lint incompatibilities. Do not claim full schema compliance or add $schema without validation against that exact schema; see resources/upstream-spec-cache.md.

Modes

ModeTriggerOutput
Generate"create sidecar from this paper / abstract / draft", "generate .knows.yaml"{paper}.knows.yaml (host LLM emits, then oma scholar lint validates)
Validate"lint this sidecar", "validate .knows.yaml"Pass/fail report with file:line issues
Review"peer review this paper as sidecar"{paper}.review.knows.yaml
Analyze"summarize this sidecar", "what claims does it make?"Natural-language answer
Compare"compare paper A and paper B structurally"Diff table (claims/methods/evidence)
Remote"find papers on X", "fetch sidecar :id", "get claims only for :id"Search results / sidecar payload

Provider Fallback (knows.academy → OpenAlex → Semantic Scholar)

knows.academy currently indexes only 2026 papers (mostly arXiv). For older or non-2026 papers (Transformer 2017, BERT 2018, classics, journals), the skill automatically falls back to OpenAlex for metadata and abstract, then to Semantic Scholar (AI TL;DR, citation + influential-citation counts) when OpenAlex also has nothing. See resources/fallback-providers.md.

Use the oma scholar CLI subcommands:

# Hybrid search: knows first, OpenAlex fallback
oma scholar search "vision language action"

# Cross-source resolve: figures out which source has the right paper
oma scholar resolve "Attention Is All You Need"

# Get by id (knows record_id, OpenAlex W-id, DOI, arXiv:<id>, CorpusId:<n>)
oma scholar get "10.48550/arXiv.1706.03762"

When OpenAlex returns the answer (knows.academy lacks the paper), use the returned abstract as input to Mode 1 Generate to produce a local sidecar.

Quick Reference

Search (knows + auto OpenAlex fallback)
oma scholar search "diffusion super resolution"
oma scholar search --year-min 2024 "vision language action"
Find one specific paper
oma scholar resolve "Attention Is All You Need"
# returns top hit from each source + recommendation
Fetch a sidecar or work
# knows.academy full sidecar
oma scholar get "knows:generated/reconvla/1.0.0"

# Partial fetch (claims only, ~700 tokens, 93% reduction vs PDF)
oma scholar get --section statements "knows:generated/reconvla/1.0.0"

# By DOI or OpenAlex W-id (works regardless of knows.academy availability)
oma scholar get "10.48550/arXiv.1706.03762"

# Semantic Scholar: AI TL;DR + citation/influential-citation counts
oma scholar get "arXiv:1706.03762"

When knows.academy is unreachable, get knows:... automatically falls back to OpenAlex by extracting the slug from the record_id. The result is marked with fallback: "openalex" and contains metadata + abstract, useful for running Mode 1 Generate locally.

Validate
# Strict mode for own Generate output (default)
oma scholar lint paper.knows.yaml

# Lenient mode for third-party / fetched sidecars
oma scholar lint --lenient remote.knows.yaml

# Treat warnings as failures (CI mode)
oma scholar lint --fail-on-warning paper.knows.yaml

About 47% of knows.academy-served sidecars contain at least one dangling cross-reference (typo in subject_ref/object_ref, measured across 15 production samples). Use --lenient when consuming third-party records so these surface as warnings rather than blocking errors.

Raw API (when CLI is unavailable)
curl -s "https://knows.academy/api/proxy/search?q=..."
curl -s "https://knows.academy/api/proxy/sidecars/<encoded-id>"
curl -s "https://knows.academy/api/proxy/partial?record_id=<id>&section=statements"
curl -s "https://knows.academy/api/proxy/jobs/stats"   # platform health

Configuration

Project-specific settings: config/scholar-config.yaml. One key is user-tunable and lives elsewhere — read scholar.base_url from .agents/oma-config.yaml first and fall back to api.base_url in the skill config. Everything else there (endpoint paths, timeouts, id prefixes, lint rules) is protocol shape, not preference.

Troubleshooting

IssueSolution
[ERROR] *.value: numeric value '22' is quotedRemove quotes: value: '22' -> value: 22
[ERROR] provenance.actor.type: 'ai' is not allowedChange to tool, person, or org
[ERROR] *.type: use \statement_type` instead of `type``Rename type -> statement_type (or evidence_type/predicate/artifact_type)
[ERROR] provenance.actors: v0.9 spec uses singular \actor``Replace actors: [{...}] array with actor: {...} object
[ERROR] *.object_ref: reference 'X' does not match any defined idFix the subject_ref/object_ref to point to a real id, OR use --lenient if consuming third-party data (cross-record record_id#local_id refs are recognized and never flagged)
[WARN] relations: avg relations/statement is N.NN (target ≥ 1.5)Review coverage; add only supported, type-appropriate links, never pad the graph
[WARN] statements: only N statements; most papers warrant ≥ 8Expected when generating from abstract only; full-paper Generate should hit 15+
[WARN] *.predicate: past-tense '...' is suspiciousSwitch to present tense (evaluated_on -> evaluates_on)
Remote API returns empty resultsTry broader query; check /api/proxy/jobs/stats; CLI auto-falls-back to OpenAlex
knows.academy search failed: fetch failed (stderr)Platform timeout; fallback to OpenAlex is automatic; retry later for sidecars
OpenAlex 403/429Set OPENALEX_API_KEY (see resources/setup-openalex.md)
semanticscholar error: HTTP 429 (stderr)Anonymous pool throttled; CLI retries once then skips the source. Set S2_API_KEY (free form at semanticscholar.org/product/api) for a dedicated limit
YAML won't parseCheck indentation; numbers/booleans must be unquoted; strings with : need quotes

References

  • Execution steps (follow for the selected task): resources/execution-protocol.md
  • Sidecar spec rules: resources/sidecar-spec.md
  • API endpoints: resources/api-endpoints.md
  • OpenAlex setup: resources/setup-openalex.md
  • Upstream spec snapshot: resources/upstream-spec-cache.md
  • Post-generation checklist: resources/checklist.md
  • CLI subcommands: oma scholar search|resolve|get|lint (implementation under cli/commands/scholar/)
  • Context loading: ../_shared/core/context-loading.md
  • Quality principles: ../_shared/core/quality-principles.md
  • i18n rules: ../../rules/i18n-guide.md

Similar Skills

lead-research-assistant
ComposioHQ/awesome-claude-skills77k

lead-research-assistant

Identifies high-quality leads for your product or service by analyzing your business, searching for target companies, and providing actionable contact strategies. Perfect for sales, business development, and marketing professionals.

Science

13c-metabolic-flux
K-Dense-AI/scientific-agent-skills48k

13c-metabolic-flux

Estimates intracellular metabolic fluxes from steady-state carbon-13 isotope-tracing measurements using validated atom maps, mfapy isotope simulation, constrained multistart fitting, and flux-profile diagnostics. Use for 13C-MFA, carbon tracing, mass isotopomer distributions (MDVs/MIDs), positional isotopomers, parallel tracer experiments, and determining whether labeling data constrain a pathway flux. Distinguishes measured-label inference from COBRA flux balance analysis and flags experiments requiring nonstationary MFA.

Science

datamol
K-Dense-AI/scientific-agent-skills48k

datamol

Pythonic wrapper around RDKit with simplified interface and sensible defaults. Preferred for standard drug discovery including SMILES parsing, standardization, descriptors, fingerprints, clustering, 3D conformers, parallel processing. Returns native rdkit.Chem.Mol objects. For advanced control or custom parameters, use rdkit directly.

Science

biopython
K-Dense-AI/scientific-agent-skills48k

biopython

Provides Biopython workflows for sequence manipulation, file parsing (FASTA/GenBank/PDB), phylogenetics, and programmatic NCBI/PubMed access (Bio.Entrez). Supports batch processing, custom molecular-biology pipelines, BLAST automation, structure analysis, and motif analysis.

Science

bulk-rnaseq
K-Dense-AI/scientific-agent-skills48k

bulk-rnaseq

Prepares bulk RNA-seq FASTQ, Salmon, STAR or featureCounts output for gene-level differential expression. Covers nf-core/rnaseq and standalone quantification, biological replication, strandedness, reference provenance, validated count assembly and a PyDESeq2 handoff. Use for FASTQ-to-counts analysis, nf-core/rnaseq configuration, STAR/Salmon quantification, or building a counts matrix for DESeq2. For single-cell data use scanpy; for statistical fitting alone use pydeseq2.

Science

alphagenome
K-Dense-AI/scientific-agent-skills48k

alphagenome

Looks up precomputed AlphaGenome Atlas effects for any GRCh38 single-nucleotide variant (AVI score with Phred and 18 SHAP feature attributions, plus raw and quantile scores for RNA-seq, DNase, ATAC, ChIP-TF, ChIP-histone, CAGE, PRO-cap, splicing, polyadenylation and contact-map tracks), scores variants or scans windows on demand with the AlphaGenome model for human and mouse (variant scoring, in silico mutagenesis, REF-versus-ALT track prediction), and builds Atlas website deep links. Use when the user mentions AlphaGenome, AlphaGenome Atlas, AVI or AlphaGenome Variant Impact, DeepMind variant effect prediction, or wants to prioritise or mechanistically interpret non-coding, regulatory, splicing, enhancer, promoter, or chromatin-accessibility effects of SNVs from a VCF, credible set, or region. Research use only; not a clinical tool.

Science