Skip to content
FunCoding

Search

Search docs, agents, posts, Skills and MCP servers

cognee-performance

Use when cognee is slow, expensive, or hitting rate limits — speeding up or throttling ingestion (remember/cognify batching, chunk size, LLM and embedding rate limits, per-stage models), cutting read latency, estimating cost before ingesting, running work in the background, or planning how to scale a deployment.

数据库与数据32k.agents/skills/cognee-performance/SKILL.md

Install

Send this to Claude Code, Codex or Cursor. The agent reviews the Skill for safety first and installs it after you confirm.

读取 https://funcoding.ai/skills/topoteretes/cognee/cognee-performance/install.md ,按里面的步骤帮我安装这个 Skill。

SKILL.md

Tune cognee's performance

The LLM is the bottleneck. Building the graph costs two LLM calls per chunk (graph extraction and summarization), and almost every other tuning question comes down to how many of those calls run, how fast, and on which model. Database and embedding work is small next to that.

Use it

1. Estimate before you ingest

estimate = await cognee.remember("./docs", dry_run=True)
print(estimate)  # per-stage token counts and approximate cost; no LLM calls

The estimate covers graph extraction and summarization only, not improve(), embeddings, or contradiction detection. Not available with GLiNER or a remote instance.

2. Ingestion knobs

All are remember() arguments (and cognify() ones):

KnobDefaultWhat it controls
data_per_batch20How many documents are processed at the same time (a concurrency limit, not a batch). Lower it to reduce load; raise it for many small files.
chunks_per_batch2000 (env CHUNKS_PER_BATCH)How many chunks each extraction/storage call receives. Lower it to spread load and fail smaller. The DLT pipeline defaults to 100.
chunk_sizemin(embedding max tokens, LLM max tokens / 2); about 8191 tokens with default OpenAI modelsMax tokens per chunk. Fewer, bigger chunks mean fewer LLM calls; smaller chunks mean finer-grained graphs.
run_in_backgroundFalseReturn immediately (status="running"); await result later.
self_improvementTrueimprove() after the graph is built. False (or IMPROVE_AUTO_ENABLED=false) skips that extra work.

Worst-case concurrent LLM calls are about data_per_batch × min(chunks per document, chunks_per_batch) × 2, and there is no separate concurrency cap. The rate limiter below is the brake.

Several datasets in one cognify(datasets=[...]) call run one after another, not in parallel.

3. LLM throttling and models

Env varDefaultNotes
LLM_RATE_LIMIT_ENABLEDfalseTurn on a requests-per-interval limit
LLM_RATE_LIMIT_REQUESTS / LLM_RATE_LIMIT_INTERVAL60 / 60 s10 requests for local providers (Ollama, llama.cpp, LM Studio)
AUTO_RATE_LIMITtrueOn a 429/503/529 or timeout, switches the limiter on for 15 minutes (extended while errors continue)
LLM_EXTRACTION_MODEL / _PROVIDER / _ENDPOINT / _API_KEY / _API_VERSIONthe main modelUse a cheaper or faster model just for graph extraction
LLM_SUMMARIZATION_*, LLM_QUERY_*the main modelSame, for summaries and for answering queries
LLM_MAX_COMPLETION_TOKENS16384Also caps the default chunk size (half of it)
LLM_ARGS—Extra litellm arguments merged into every call, e.g. a request timeout

Structured-output calls retry with exponential backoff for at least two attempts and four minutes before failing; auth, not-found, and quota errors are not retried.

The rate limiter is built from the main model's settings. A stage routed to a local model does not get the local default of 10 requests; set LLM_RATE_LIMIT_REQUESTS yourself.

4. Embeddings

Env varDefault
EMBEDDING_BATCH_SIZE36 texts per request
EMBEDDING_MAX_CONCURRENT_DATA_POINTS150 (so 150 / 36 = 4 concurrent requests)
EMBEDDING_RATE_LIMIT_ENABLED / _REQUESTS / _INTERVALoff / 60 / 60 s

Embedding calls retry transient errors, rate-limit 429s included, with exponential jitter for up to 128 s (budget-exhaustion errors are terminal, not retried), but AUTO_RATE_LIMIT does not apply to them; enable EMBEDDING_RATE_LIMIT_ENABLED yourself for sustained load. Without any credentials cognee embeds locally with fastembed (BAAI/bge-small-en-v1.5), which runs on the CPU and blocks while it works.

5. Skip the LLM for extraction

remember(data, extractor="gliner") builds the graph and summaries with a local GLiNER model: no LLM calls, embeddings still run. The runtime installs on first use (GLINER_AUTO_INSTALL=false to disable; then pip install "cognee[gliner]"). See the cognee-ingestion skill for its limits.

6. Read latency

  • AUTO_FEEDBACK=false is the biggest win for chat-style use: by default every answered turn (the default session is used when no session_id is passed) makes one extra LLM call to analyze feedback. Keep CACHING=true so session memory still works.
  • Pick a type without an LLM when you need passages, not an answer: CHUNKS, CHUNKS_LEXICAL, SUMMARIES, CODE, SKILLS. Completion types make one LLM call; GRAPH_COMPLETION_COT, _DECOMPOSITION, _CONTEXT_EXTENSION, GRAPH_SUMMARY_COMPLETION and TEMPORAL make several; NATURAL_LANGUAGE makes one and retries (up to 3 attempts) only on an empty or failed query; FEELING_LUCKY adds one to pick the type.
  • Narrow the search: pass datasets=[...] (one search runs per dataset) and a smaller top_k (per dataset, default 15; HYBRID caps each lane at 10, so only values below 10 shrink its context).
  • SESSION_SEARCH_MODE=concurrent (default) overlaps the feedback analysis with the answer; sequential runs them back to back.

7. Several datasets at once

With access control on (the default), each dataset has its own databases. DATASET_QUEUE_MAX_CONCURRENT (default 6, from DATABASE_MAX_LRU_CACHE_SIZE) caps how many datasets are processed at once in one process, and SUBPROCESS_IDLE_TTL_SECONDS (600) keeps idle database workers warm. DATASET_QUEUE_ENABLED (default true) enforces that cap, releases subprocess engines when a dataset's last scope exits (they stay warm for SUBPROCESS_IDLE_TTL_SECONDS and close at once only when it is 0) and pins in-use engines against eviction. Setting it false removes the cap rather than disabling parallelism, and risks file-lock leaks and engine eviction under parallel load, so keep it on.

8. Scaling a deployment

Open-source cognee scales vertically, in one process:

  • The API server runs one worker (gunicorn -w 1), and the improve and session locks, caches and semaphores are in-process only. More workers or replicas against the same data are not coordinated (the Helm chart README says to validate before scaling past one replica).
  • The default embedded databases (SQLite, LanceDB, Ladybug) live on local disk. For production, use Postgres + PGVector and a graph-native backend such as Neo4j.
  • distributed/deploy/ holds one-click deploy templates (Modal, Fly, Railway, Render, Daytona), not a distributed runner.

For production scale: horizontal scaling (distributed ingestion across workers), cognee's enterprise GLiNER extraction, and the production Postgres graph adapter are part of cognee's proprietary offering. Contact social@cognee.ai.

Pitfalls

  • CHUNK_SIZE, CHUNK_OVERLAP, CHUNK_STRATEGY (in .env.template) and cognee.config.set_chunk_size() etc. have no effect on ingestion. Pass chunk_size= to remember() instead.
  • LLM_RATE_LIMIT_TOKENS and EMBEDDING_RATE_LIMIT_TOKENS do nothing; only request-count limits are enforced.
  • FALLBACK_MODEL is used only when the main model rejects content on policy grounds, not when it is slow, down, or rate-limited.
  • data_per_batch=None crashes. Omit the argument to get the default.
  • A background cognify() is not tracked by cognee.wait_for_background_tasks() (a background remember() or improve() is). Await its pipeline status yourself.
  • Big default chunks with a small local embedder. The default chunk size follows the configured embedding limit (8191), even for fastembed's 512-token model. Pass a smaller chunk_size for local embedding models.
  • Benchmark with CACHING=true: with it off you measure cognee without its memory layer.

How it works

remember() → add() → cognify(): every document of the dataset is scheduled at once, data_per_batch of them run concurrently, and each runs the task chain (classify, chunk, extract graph + summarize, store). Tasks receive their input in batches of chunks_per_batch. extract_graph_and_summarize runs extraction and summarization together, each gathering one LLM call per chunk in the batch. Every LLM call passes through one process-wide rate limiter.

  • Cognify task list and defaults: cognee/api/v1/cognify/cognify.py
  • Concurrency (data_per_batch semaphore): cognee/modules/pipelines/operations/run_tasks.py
  • Task batching: cognee/modules/pipelines/tasks/task.py, run_tasks_base.py
  • LLM calls per chunk: cognee/tasks/graph/extract_graph_and_summarize.py
  • Chunk-size default: cognee/infrastructure/llm/utils.py:get_max_chunk_tokens
  • LLM config, rate limits, stage routing: cognee/infrastructure/llm/config.py, cognee/shared/rate_limiting.py, cognee/infrastructure/llm/overload_policy.py, cognee/infrastructure/llm/retry_config.py
  • Embedding config: cognee/infrastructure/databases/vector/embeddings/config.py
  • Dataset queue: cognee/infrastructure/databases/dataset_queue/queue.py
  • Cost estimate: cognee/modules/cognify/estimator.py

Extending it

  • Benchmarks: cognee/tests/performance/ (batch_add_cognify_test.py, locust_performance_analysis.py, and statistics_percentile/ for p50–p99 runs with mocked or replayed LLM calls). Measure before and after a change.
  • Retrieval quality and cost: cognee/eval_framework/ (python -m cognee.eval_framework; Modal runs cost money).
  • A new pipeline task that calls an LLM should keep needs_llm=True and go through LLMGateway, so it shares the rate limiter and retry policy.

Similar Skills

xlsx
Official
anthropics/skills180k

xlsx

Use this skill any time a spreadsheet file is the primary input or output. This means any task where the user wants to: open, read, edit, or fix an existing .xlsx, .xlsm, .xltx, .csv, or .tsv file (e.g., adding columns, computing formulas, formatting, charting, cleaning messy data); create a new spreadsheet from scratch or from other data sources; or convert between tabular file formats. Trigger especially when the user references a spreadsheet file by name or path — even casually (like "the xlsx in my downloads") — and wants something done to it or produced from it. Also trigger for cleaning or restructuring messy tabular data files (malformed rows, misplaced headers, junk data) into proper spreadsheets. The deliverable must be a spreadsheet file. Do NOT trigger when the primary deliverable is a Word document, HTML report, standalone Python script, database pipeline, or Google Sheets API integration, even if tabular data is involved.

Databases & data

deprecation-and-migration
addyosmani/agent-skills102k

deprecation-and-migration

Manages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to maintain or sunset existing code.

Databases & data

smart-explore
thedotmack/claude-mem97k

smart-explore

Token-optimized structural code search using tree-sitter AST parsing. Use instead of reading full files when you need to understand code structure, find functions, or explore a codebase efficiently.

Databases & data

pathfinder
thedotmack/claude-mem97k

pathfinder

Map a codebase into feature-grouped flowcharts, identify duplicated concerns across features, and propose a unified architecture. Use when asked to "find the ideal path," unify duplicated systems, or audit architecture before a refactor. Emits a proposed unified flowchart plus per-system /make-plan prompts.

Databases & data

oh-my-issues
thedotmack/claude-mem97k

oh-my-issues

Cluster a GitHub issue backlog by root cause into a small set of plan-master issues, redirect children with a standardized comment, and bundle architectural-fix PRs that close clusters atomically. Use when an issue tracker has accumulated dozens of reports that share underlying defects, when asked to triage / consolidate / cluster / dedupe issues, when asked to build a plan series or roadmap from open issues, or when routing a new incoming bug into an existing plan.

Databases & data

mode-creator
thedotmack/claude-mem97k

mode-creator

Interactively create, install, activate, and verify custom claude-mem modes, including domain-specific observation types, concept tags, optional Telegram alerts, bot setup, worker restart, and startup-context verification. Use this whenever someone asks to customize what claude-mem remembers, create or change a mode, track domain-specific notes, add observation types or tags, or send Telegram notifications for particular memories—even if they do not use the word "mode."

Databases & data