跳到正文
FunCoding

搜索

搜索文档、文章、Skill 和 MCP

Query modes, prompts, and models

How much conversation the sub-agent sees, how eager it is about returning memory, and how its model is resolved.

Query modes

config.queryMode controls how much conversation the blocking sub-agent sees. Pick the smallest mode that still answers follow-ups well; grow timeoutMs as context size grows, from message to recent to full.

message

Only the latest user message is sent.

Latest user message only

Use when you want the fastest behavior, the strongest bias toward stable preference recall, and follow-up turns do not need conversational context. Start around 3000-5000 ms for config.timeoutMs.

recent

The latest user message plus a small recent conversational tail.

Recent conversation tail:
user: ...
assistant: ...
user: ...

Latest user message:
...

Use for a balance of speed and conversational grounding, when follow-up questions often depend on the last few turns. Start around 15000 ms.

full

The full conversation is sent to the blocking sub-agent.

Full conversation context:
user: ...
assistant: ...
user: ...
...

Use when recall quality matters more than latency, or important setup is far back in the thread. Start around 15000 ms or higher depending on thread size.

Prompt styles

config.promptStyle controls how eager or strict the sub-agent is about returning memory:

StyleBehavior
balancedGeneral-purpose default for recent mode
strictLeast eager; minimal bleed from nearby context
contextualMost continuity-friendly; conversation history matters more
recall-heavySurfaces memory on softer but still plausible matches
precision-heavyAggressively prefers NONE unless the match is obvious
preference-onlyOptimized for favorites, habits, routines, taste, recurring personal facts

Default mapping when config.promptStyle is unset:

message -> strict
recent -> balanced
full -> contextual

An explicit config.promptStyle always overrides the mapping.

Model fallback policy

If config.model is unset, active memory resolves a model in this order:

explicit plugin model (config.model)
-> current session model
-> agent primary model
-> optional configured fallback model (config.modelFallback)
modelFallback: "google/gemini-3-flash"

If nothing in that chain resolves, active memory skips recall for the turn. config.modelFallbackPolicy is a compatibility field kept for older configs, deprecated in v2026.4.12; it no longer changes runtime behavior — modelFallback is strictly the last resort in the chain above, not a runtime failover that swaps in another model when the resolved one errors.

Speed recommendations

Leaving config.model unset (inherit the session model) is the safest default: it follows your existing provider, auth, and model preferences. For lower latency, use a dedicated fast model instead — recall quality matters, but latency matters more here than on the main answer path, and the tool surface is narrow (only memory recall tools).

Good fast-model options:

  • cerebras/gpt-oss-120b, a dedicated low-latency recall model
  • google/gemini-3-flash, a low-latency fallback without changing your primary chat model
  • your normal session model, by leaving config.model unset

Cerebras setup

{
  models: {
    providers: {
      cerebras: {
        baseUrl: "https://api.cerebras.ai/v1",
        apiKey: "${CEREBRAS_API_KEY}",
        api: "openai-completions",
        models: [{ id: "gpt-oss-120b", name: "GPT OSS 120B (Cerebras)" }],
      },
    },
  },
  plugins: {
    entries: {
      "active-memory": {
        enabled: true,
        config: { model: "cerebras/gpt-oss-120b" },
      },
    },
  },
}

Confirm the Cerebras API key has chat/completions access for the chosen model — /v1/models visibility alone does not guarantee it.