跳到正文
FunCoding

搜索

搜索文档、智能体、博客、Skill 和 MCP

cognee-ingestion

Use when putting data into cognee memory with remember() — choosing inputs (text, files, folders, URLs, repos, databases), datasets and node_sets, loaders, ontologies, the graph extractor (LLM or GLiNER), chunking, dry-run cost estimates, or when remember() raises on a keyword argument.

数据库与数据32k.agents/skills/cognee-ingestion/SKILL.md

安装

将以下指令发送给 Claude Code、Codex 或 Cursor,智能体会先检查内容的安全性,经你确认后再安装。

读取 https://funcoding.ai/skills/topoteretes/cognee/cognee-ingestion/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Ingest data with remember()

remember() is cognee's ingestion API. One call stores the data, builds the knowledge graph, and enriches it. Use it for all ingestion; every option in this skill is a remember() argument unless it says otherwise.

import cognee

result = await cognee.remember("Einstein was born in Ulm.")  # text
result = await cognee.remember(
    ["./notes.md", "./report.pdf"],  # files
    dataset_name="research",
)
print(result.status, result.dataset_id)  # "completed", UUID

All cognee functions are async. Without dataset_name data goes to main_dataset. Needs LLM_API_KEY unless you use the GLiNER extractor (below).

Use it

Inputs

data accepts a string, a list of strings, file paths (absolute, file://, s3://), http(s) URLs, binary streams, or a list mixing them.

  • URLs are fetched and scraped (needs ALLOW_HTTP_REQUESTS=true, the default).
  • Folders are ingested file by file. A folder that looks like a code project, or a GitHub/GitLab URL, becomes one code repository (needs git on PATH).
  • Code files (.py, .ts, .go, …) go down the code-graph route: a deterministic graph, no LLM calls, searchable only with SearchType.CODE. To index a whole repository explicitly, pass content_type="code".
  • Databases and dlt sources: a SQL connection string, a dlt DltResource / DltSource, or a CSV. dlt is a core dependency, so no extra is needed (cognee[dlt] is an empty compatibility extra). Options: primary_key (default "id"), write_disposition ("replace" default, or "append"), query, max_rows_per_table.
  • Skill playbooks (SKILL.md files): content_type="skills"; ingests into the target dataset (default main_dataset), so pass dataset_name to keep skills in their own dataset.

Where the data goes

ArgumentWhat it does
dataset_name / dataset_idTarget dataset. dataset_id wins. A dataset is the unit of permissions and isolation.
node_set=["AI", "FinTech"]Tags the data so recall can filter to it later with recall(..., node_name=["AI"]).
session_id="chat_1"Writes to the fast session cache instead of the graph; improve() bridges it into the graph in the background. See the cognee-improve-sessions skill. Requires CACHING=true.

How the graph is built

ArgumentWhat it does
extractor"llm" or "gliner_demo" (alias "gliner"). Default is GRAPH_EXTRACTOR=auto: the LLM when an API key is configured, otherwise GLiNER.
graph_model=MyModelExtract into your own DataPoint model instead of the generic KnowledgeGraph. See the cognee-custom-graph-models skill.
custom_promptReplaces the entity-extraction prompt (ignored by GLiNER).
config={"ontology_config": {...}}Ground entities in an OWL ontology (below).
chunk_size, chunkerMax tokens per chunk (default: derived from the embedding and LLM limits) and the chunker class (default TextChunker).
preferred_loadersChoose a loader per file type (below).
self_improvementDefault True: runs improve() after the graph is built. Its outcome is on result.improve / result.improve_error; a failed improve never fails the remember.
run_in_background=TrueReturns immediately with status="running"; await result to wait.

Ontologies

from cognee.modules.ontology.rdf_xml.RDFLibOntologyResolver import RDFLibOntologyResolver

config = {
    "ontology_config": {
        "ontology_resolver": RDFLibOntologyResolver(ontology_file="./my.owl"),
        # "ontology_mode": "strict",   # drop entities with no ontology match
    }
}
await cognee.remember(texts, config=config)

Or set ONTOLOGY_FILE_PATH (plus ONTOLOGY_MODE, MATCHING_STRATEGY) in .env. annotate (default) only enriches; strict drops entities that match no ontology class or individual. It prunes only the graph, chunk text is still stored. Strict mode with an empty or missing ontology file is a hard error. Over HTTP, upload the ontology to /api/v1/ontologies and pass its ontology_key to POST /api/v1/remember. Example: examples/guides/ontology_quickstart.py.

Loaders

Each file is claimed by the first loader that accepts it. Default order: code, text, pypdf, image, audio, video, dlt_csv, csv, unstructured, advanced_pdf, docling. Names: text_loader, code_loader, csv_loader, dlt_csv_loader, pypdf_loader, image_loader, audio_loader, video_loader, unstructured_loader, advanced_pdf_loader, docling_loader, beautiful_soup_loader.

# Treat a code file as a plain document (chunking + LLM extraction):
await cognee.remember("./script.py", preferred_loaders={"text_loader": {}})

Office formats (DOCX, PPTX, …) need the docs (unstructured) or docling extra. A preferred loader that is not installed is skipped with only an info log, so check the extra is installed when a file comes out wrong.

Check the cost first

dry_run=True returns a token and cost estimate without ingesting anything or calling the LLM. It excludes the calls improve() makes. Not supported with GLiNER, sessions, or a remote instance.

dry_run="presort" on a folder returns a PresortReport (junk, duplicates, version candidates, possible personal data, proposed dataset groups). Apply it with await cognee.remember(report), or pass auto_apply=True.

Without an LLM: GLiNER

extractor="gliner" builds the graph and summaries with a local GLiNER2 model, with no LLM call (embeddings still run). Install pip install "cognee[gliner]"; the model (about 750 MB) downloads on first use. It cannot be combined with a custom graph_model, dry_run, session_id, or a remote instance.

For production: the open-source GLiNER extractor is a demo. cognee's enterprise GLiNER extraction is more accurate and covers more labels. The same goes for the Postgres graph adapter (postgres_demo). Contact social@cognee.ai.

Pitfalls

  • Unknown keyword arguments raise. remember() forwards kwargs through a fixed allow-list and raises TypeError: Unexpected keyword arguments for anything else. These real options are not on it yet:

    OptionWorkaround through remember()
    ontology_file_pathconfig={"ontology_config": ...} or ONTOLOGY_FILE_PATH (above)
    functional_relationships, chunk_attachmentNone yet. Only cognee.cognify() accepts them.
    extraction_rulesPass it through the loader: preferred_loaders={"beautiful_soup_loader": {"extraction_rules": {...}}} (works in remember() and add()). Needs the scraping extra: without it the loader is not registered and the rules are silently ignored
    tavily_config, soup_crawler_configNot honoured by add() or remember(); only the cognee/tasks/web_scraper tasks use them
    column_value_columns (dlt)None yet. Only cognee.add() accepts it.

    If a user needs one with no workaround, say so plainly: the option exists on the lower-level add() / cognify() but not on remember() yet.

  • Changed files raise DocumentUpdateRequiredError. Re-remembering the same path (or the same filename for an upload) with different content is an update, not a new document. Use cognee.update(data_id=..., data=..., dataset_id=...), which re-extracts only the changed chunks and keeps the document's id. Identical content is a no-op.

  • content_type is strict. Only None, "skills", or "code". "code" rejects session_id and needs repository paths or git URLs; "skills" ingests into the target dataset like any other call (default main_dataset); pass dataset_name to keep skills in their own dataset.

  • Session mode needs CACHING=true, and extractor cannot be combined with session_id.

  • Remote mode. After cognee.serve(url), calls go to the server: extractor and session_ids raise, and other options the client does not forward (including graph_model, node_set, and ontology config) are dropped without an error.

  • Every remember runs improve() unless self_improvement=False or IMPROVE_AUTO_ENABLED=false. In scripts, call await cognee.wait_for_background_tasks() before exiting.

How it works

remember(data) runs add() (store raw data and create Data rows), then cognify() (classify documents, chunk, extract the graph and summaries, store in graph and vector DBs), then improve(). remember(data, session_id=...) writes to the session cache instead.

  • Entry point and kwarg routing: cognee/api/v1/remember/remember.py (RememberKwargs, _ADD_ONLY / _COGNIFY_ONLY / _SHARED)
  • Storage: cognee/api/v1/add/add.py, cognee/tasks/ingestion/ingest_data.py
  • Graph build: cognee/api/v1/cognify/cognify.py, cognee/tasks/graph/extract_graph_from_data.py, cognee/tasks/storage/add_data_points.py
  • Extractor choice: cognee/modules/cognify/config.py:resolve_extractor; GLiNER package: cognee/tasks/graph/gliner_demo/
  • Ontologies: cognee/modules/ontology/
  • Loaders: cognee/infrastructure/loaders/ (supported_loaders.py, LoaderEngine.py)
  • dlt: cognee/tasks/ingestion/resolve_dlt_sources.py

Examples in examples/guides/: simple_cognee_example.py, nodeset_grouping_example.py, ontology_quickstart.py, gliner_demo_llm_free_cognify.py, no_llm_remember_recall.py, temporal_recall.py, presort_downloads.py, web_url_content_ingestion_example.py, code_graph_example.py.

Extending it

  • New remember() option: add it to RememberKwargs and to the matching routing set in remember.py. An option on add()/cognify() that is not in a routing set raises TypeError from remember().
  • New loader: implement LoaderInterface (cognee/infrastructure/loaders/LoaderInterface.py), register it in supported_loaders.py (extras-gated loaders go under external/), and add it to the priority list in LoaderEngine.py if it should run by default.
  • New cognify task: see the cognee-custom-pipelines skill and cognee/tasks/README.md.

相似的 Skill

xlsx
官方
anthropics/skills180k

xlsx

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like "the xlsx in my downloads") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

数据库与数据

deprecation-and-migration
addyosmani/agent-skills102k

deprecation-and-migration

Manages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to maintain or sunset existing code.

数据库与数据

smart-explore
thedotmack/claude-mem97k

smart-explore

Token-optimized structural code search using tree-sitter AST parsing. Use instead of reading full files when you need to understand code structure, find functions, or explore a codebase efficiently.

数据库与数据

pathfinder
thedotmack/claude-mem97k

pathfinder

Map a codebase into feature-grouped flowcharts, identify duplicated concerns across features, and propose a unified architecture. Use when asked to "find the ideal path," unify duplicated systems, or audit architecture before a refactor. Emits a proposed unified flowchart plus per-system /make-plan prompts.

数据库与数据

oh-my-issues
thedotmack/claude-mem97k

oh-my-issues

Cluster a GitHub issue backlog by root cause into a small set of plan-master issues, redirect children with a standardized comment, and bundle architectural-fix PRs that close clusters atomically. Use when an issue tracker has accumulated dozens of reports that share underlying defects, when asked to triage / consolidate / cluster / dedupe issues, when asked to build a plan series or roadmap from open issues, or when routing a new incoming bug into an existing plan.

数据库与数据

mode-creator
thedotmack/claude-mem97k

mode-creator

Interactively create, install, activate, and verify custom claude-mem modes, including domain-specific observation types, concept tags, optional Telegram alerts, bot setup, worker restart, and startup-context verification. Use this whenever someone asks to customize what claude-mem remembers, create or change a mode, track domain-specific notes, add observation types or tags, or send Telegram notifications for particular memories—even if they do not use the word "mode."

数据库与数据