Skip to main content

1. What it is

Knowledge is the source material your agent searches to answer questions: product docs, policy PDFs, Slack threads, wiki pages, CSVs, and emails. One uploaded file or app record is a source. HydraDB splits its text into searchable passages called chunks. Use Memories for preferences and history you want to carry between conversations. The content category does not set permissions: knowledge can be private, and either category can be stored in a collection you choose.
One endpoint, two stores. Knowledge and memories are stored separately, but a single POST /query call can hit both via type: "all". Both stores are read from the same collection scope, and results are merged and re-ranked together. Pick type: "knowledge" for document questions, type: "memory" for per-user context, or type: "all" for personalized answers grounded in both. See Query for the full picture.

2. What it does

When you call POST /context/ingest with type=knowledge, HydraDB walks each source through a pipeline:
  1. Parses the input: PDF, DOCX, Markdown, CSV, plain text, or app-source JSON.
  2. Splits the text into passages that can be retrieved independently.
  3. Finds entities (named people, teams, services, or topics) and their relationships (see Context Graphs).
  4. Builds search indexes for matching by meaning and by keywords.
  5. Indexes the result so it becomes queryable via POST /query.
The pipeline runs asynchronously: ingest returns 202 Accepted with ids, and the content becomes queryable once indexing_status reaches graph_creation. Section 5 shows how to poll.

3. When to use it

Reach for Knowledge ingestion when the content is:
  • Reference material: product docs, policy PDFs, runbooks, wikis, or tickets. Choose collections and document permissions separately.
  • Static or infrequently updated: versioned by replacing the source rather than appending.
  • Document-structured rather than conversational: files, threads, pages, articles.
Reach for Memories instead when the content is:
  • User preferences or behavioral signals tied to one person.
  • Per-user conversation history that should shape future answers for that user.
  • Anything that should personalize a single user’s experience.

4. Two ingestion paths

POST /context/ingest with type=knowledge accepts two content formats, and you can send either or both in one multipart/form-data request. Pick the path that matches what you have:
  • Path A when you have binary files HydraDB should parse for you.
  • Path B when you have pre-extracted text from an app or connector and want to skip parsing.

Path A: Files

Binary files HydraDB should parse and chunk: PDF, DOCX, Markdown, CSV, TXT.

Path B: App sources

Pre-parsed records from connected apps. You supply app-native fields (kind, provider, external_id, and fields), metadata, and optional attachments/comments. Use this when your application already extracts content from Slack messages, Notion pages, Gmail messages, tickets, or CRM records and you want HydraDB to preserve app structure instead of treating the item as generic text.
Response (both paths):
The /context and /query endpoints return this envelope shape. Access the payload via response.data (e.g. response.data.results), not at the top level.

5. Verifying processing

Ingestion is async, so you need to confirm each source actually finished before it’ll show up in query. Poll GET /context/status with the ids you got back from ingestion. The snippets below loop until every item reaches a terminal state (completed or errored):
Each status in the data payload looks like this:
Status values move forward through queued, processing, graph_creation, and completed (or land in errored on failure). Poll every few seconds; most documents complete in 1 to 5 minutes. The full table is on Ingestion Status.
graph_creation is already queryable. Items in this state are retrievable via POST /query. Wait for completed only when you specifically need full graph traversal (graph_context: true).

6. Key parameters

The full schema lives on POST /context/ingest. Below are the fields you’ll touch most often, grouped by where they live in the request.

File metadata (document_metadata array items)

Send one item per file in documents, in the same order. The item count must match the file count. Any other key, including title, type, url, and timestamp, returns 400. App sources accept title, url, and timestamp.

App source model (app_knowledge array items)

Request-level fields


7. Forceful relations

Forceful relations let you declare links between Knowledge sources at ingestion time. When a query matches a source that has forceful relations attached, HydraDB pulls the linked sources into the additional_context field of the query response alongside the primary chunks. This is separate from the entity and relationship context graph HydraDB builds automatically. Graph extraction is derived from content; forceful relations are declared by you, which is useful when you know two documents belong together (a contract and its addendum, a runbook and its troubleshooting guide) but the text doesn’t make the connection explicit. Forceful relations link whole sources; to declare the full entity/relation graph within a single document (replacing extraction for it), use Bring Your Own Graph. For file uploads, set relations.ids on the matching document_metadata item. For app sources, prefer the app-native relations[] shape so each link can carry a predicate and provider-aware target:
At query time, linked sources are returned in the additional_context field of the query response. query_forceful_relations defaults to true, so you only need to set it explicitly if you want to disable forceful-relation expansion.
Knowledge items only. Forceful relations work within the same store. Pointing a Knowledge item at a Memory ID will silently return nothing at query time.
thinking mode only. Forceful-relation context is fetched only when mode: "thinking" is set on the query request (or when mode: "auto" routes to thinking). In fast mode, query_forceful_relations is ignored even if set to true.

8. Metadata on knowledge

Metadata is how you bridge structured filters and semantic search. Attach metadata at ingestion to enable deterministic narrowing at query time:
Then at query time:
Declare the metadata keys you filter on in the database metadata schema. In a database with a schema, ingest rejects undeclared keys, so filtering on one returns nothing. additional_metadata is stored alongside the source for display and bookkeeping; to filter on it, nest the keys under additional_metadata in metadata_filters (e.g. {"metadata_filters": {"additional_metadata": {"author": "Alice"}}}). document_metadata is accepted only as a legacy alias for that nested filter namespace. See Metadata for schema design, filter examples, and the metadata vs additional_metadata decision rule. See Create Database for the schema field configuration (enable_dense_embedding, enable_sparse_embedding).

9. Common mistakes