POST /query retrieves Knowledge (documents and app sources), Memories (user preferences, conversation history, inferred content), or both. The response contains matching passages (chunks) and their source details, which your application can put in a model prompt. Search combines meaning, keywords, and relationships between the content. It does not generate the answer.
For the full request and response schema, see Query: API Reference.
1. Use-case recipes
Pick the row that matches your goal and use the parameters as a starting point:2. Parameter reference
When to senddatabase and collection: Multi-tenant. The older names tenant_id and sub_tenant_id still work as deprecated aliases.
Retrieval
Scope
Graph
Shaping results
Access control
3. Tuning heuristics
Most of the time the defaults are right. When they aren’t, here’s where to start:mode: Pick"fast"or"thinking"explicitly when you need predictable retrieval behavior instead of automatic routing.alpha: Start at0.8. Lower toward0.3to0.5when the query contains literal tokens (error codes, SKUs, product names). Raise toward0.9for conceptual questions.max_results: Start at10. Drop to5for tight context windows; raise to20if you rerank downstream.additional_context: Use it when the query alone is ambiguous. Keep it short and factual.graph_context: Keep it on (the default) when answers depend on entity relationships (multi-hop questions, “how does X relate to Y”). Pair withmode: "thinking", because in"fast"mode the graph slice is shallow.query_apps: Leave it on for app data (Slack, Gmail, Jira, and so on), and pair it withmode: "thinking"so relations and threads expand.
4. Minimal working example
A personalized-answer flow takes a single call:POST /query with type: "all" returns merged knowledge and per-user memory in one ranked result set. Both stores are read from the same scope, so if shared knowledge lives in its own collection, list it next to the user’s in collections.
Setup
One call: Knowledge and Memories together
Merge into the LLM prompt
The response is a singleRetrievalResult containing chunks[], sources[], and, when applicable, graph_context and additional_context. Chunks from Knowledge and Memories are already interleaved and ranked by relevance, so no manual merging is required. Pass the result through the helper in How to Use API Results to turn it into a context string for your prompt:
collections. A list gives every scope equal normalized weight; an object applies relative ranking weights with at most one decimal place before the final merged ranking. When max_results is set, it caps the final merged response across all selected collections:
/query twice in parallel with type: "knowledge" and type: "memory", then merge client-side.
Production checklist
- Set per-call timeouts: Generous for
thinking(3 to 5 s), tight forfast(500 ms or less). Formode: "auto", size the timeout for thethinkingcase, since it can resolve to either pipeline and defaults towardthinkingwhen the routing signal is inconclusive. - Pass
additional_contextwith known session state (page, feature, role). It sharpens retrieval without extra calls.
5. Common mistakes
6. Advanced patterns
Hybrid + text in two parallel calls. When a query mixes a literal token (error code, SKU, function name) with natural-language intent, runquery_by: "hybrid" and query_by: "text" in parallel, dedupe by chunk ID, and treat text hits as a “must include” floor.
Recall-then-rerank. Ask for more chunks than you actually need (max_results: 20) and apply your own reranker (recency windows, compliance filters, business rules) before picking the final top-k for the prompt.
Cache the prompt context: Include the database, collection scopes, caller ACL, and all query parameters in the cache key. Expire cached results when content or permissions change.
Related
- Knowledge: shared document context
- Memories: user-scoped dynamic context
- Metadata: designing filterable fields
- Context Graphs: how graph traversal enriches retrieval
- How to Use API Results: turning
RetrievalResultinto an LLM prompt - Query: API Reference: full parameter and response schema
