> ## Documentation Index
> Fetch the complete documentation index at: https://cortex-e852fafe-t3code-rewrite-docs-declutter.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Knowledge

> What Knowledge is, how it differs from Memories, and how to ingest it correctly.

export const Field = ({name, type, required, recommended}) => {
  const label = required ? 'required' : recommended ? 'recommended' : null;
  const typeLabel = typeof type === 'string' ? type : null;
  const ariaParts = [name, typeLabel && `${typeLabel}`, label].filter(Boolean);
  return <span aria-label={ariaParts.join(', ')} className={label ? 'field-wrap has-field-tip' : 'field-wrap'} style={{
    position: 'relative',
    cursor: label ? 'default' : undefined
  }} tabIndex={label ? 0 : undefined}>
      <span className="field-name-row">
        <code>{name}</code>
        {required && <span className="field-req"> *</span>}
        {recommended && <span className="field-rec"> ●</span>}
      </span>
      {type && <span className="field-type">{type}</span>}
      {label && <span className="field-tip" role="tooltip">
          {label}
        </span>}
    </span>;
};

## 1. What it is

Knowledge is the source material your agent searches to answer questions: product docs, policy PDFs, Slack threads, wiki pages, CSVs, and emails. One uploaded file or app record is a **source**. HydraDB splits its text into searchable passages called **chunks**.

Use [Memories](/essentials/v2/memories) for preferences and history you want to carry between conversations. The content category does not set permissions: knowledge can be private, and either category can be stored in a collection you choose.

| | Knowledge | Memories |
| - | - | - |
| **Content** | Documents, files, app-generated content | User preferences, conversations, inferred traits |
| **Scope** | Searches the selected collection; use [ACLs](/essentials/v2/access-control) to restrict documents | Usually stored in a user's collection; your application chooses it |
| **Mutability** | Versioned: replaced or deleted explicitly | Updated when your application sends new information |
| **Query via** | [`POST /query`](/api-reference/v2/endpoint/query) with `type: "knowledge"` | [`POST /query`](/api-reference/v2/endpoint/query) with `type: "memory"` |
| **`vectorstore_status`** | `vectorstore_status.knowledge` | `vectorstore_status.memories` |

<Info>
  **One endpoint, two stores.** Knowledge and memories are stored separately, but a single [`POST /query`](/api-reference/v2/endpoint/query) call can hit both via `type: "all"`. Both stores are read from the same collection scope, and results are merged and re-ranked together. Pick `type: "knowledge"` for document questions, `type: "memory"` for per-user context, or `type: "all"` for personalized answers grounded in both. See [Query](/essentials/v2/query) for the full picture.
</Info>

***

## 2. What it does

When you call [`POST /context/ingest`](/api-reference/v2/endpoint/ingest-context) with `type=knowledge`, HydraDB walks each source through a pipeline:

1. **Parses** the input: PDF, DOCX, Markdown, CSV, plain text, or app-source JSON.
2. **Splits** the text into passages that can be retrieved independently.
3. **Finds entities** (named people, teams, services, or topics) **and their relationships** (see [Context Graphs](/essentials/v2/context-graphs)).
4. **Builds search indexes** for matching by meaning and by keywords.
5. **Indexes** the result so it becomes queryable via [`POST /query`](/api-reference/v2/endpoint/query).

The pipeline runs **asynchronously**: ingest returns `202 Accepted` with `id`s, and the content becomes queryable once `indexing_status` reaches `graph_creation`. [Section 5](#5-verifying-processing) shows how to poll.

***

## 3. When to use it

Reach for Knowledge ingestion when the content is:

* **Reference material**: product docs, policy PDFs, runbooks, wikis, or tickets. Choose collections and document permissions separately.
* **Static or infrequently updated**: versioned by replacing the source rather than appending.
* **Document-structured** rather than conversational: files, threads, pages, articles.

Reach for [Memories](/essentials/v2/memories) instead when the content is:

* User preferences or behavioral signals tied to one person.
* Per-user conversation history that should shape future answers for that user.
* Anything that should personalize a single user's experience.

***

## 4. Two ingestion paths

[`POST /context/ingest`](/api-reference/v2/endpoint/ingest-context) with `type=knowledge` accepts two content formats, and you can send either or both in one `multipart/form-data` request. Pick the path that matches what you have:

* **Path A** when you have **binary files** HydraDB should parse for you.
* **Path B** when you have **pre-extracted text** from an app or connector and want to skip parsing.

### Path A: Files

Binary files HydraDB should parse and chunk: PDF, DOCX, Markdown, CSV, TXT.

<CodeGroup>
  ```python Python SDK theme={"dark"}
  import json

  with open("/path/to/policy.pdf", "rb") as f1, open("/path/to/runbook.pdf", "rb") as f2:
      result = client.context.ingest(
          type="knowledge",
          database="acme_corp",
          documents=[
              ("policy.pdf", f1, "application/pdf"),
              ("runbook.pdf", f2, "application/pdf"),
          ],
          document_metadata=json.dumps([
              {"id": "policy_v2", "metadata": {"document_type": "policy", "status": "approved"}},
              {"id": "runbook_deploy", "metadata": {"document_type": "runbook"}},
          ]),
      )
  ```

  ```typescript TypeScript SDK theme={"dark"}
  const result = await client.context.ingest({
    type: "knowledge",
    database: "acme_corp",
    documents: [
      { path: "/path/to/policy.pdf", filename: "policy.pdf", contentType: "application/pdf" },
      { path: "/path/to/runbook.pdf", filename: "runbook.pdf", contentType: "application/pdf" },
    ],
    documentMetadata: JSON.stringify([
      { id: "policy_v2", metadata: { document_type: "policy", status: "approved" } },
      { id: "runbook_deploy", metadata: { document_type: "runbook" } },
    ]),
  });
  ```

  ```bash cURL theme={"dark"}
  curl -X POST 'https://api.hydradb.com/context/ingest' \
    -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
    -H "API-Version: 2" \
    -F "type=knowledge" \
    -F "database=acme_corp" \
    -F "documents=@/path/to/policy.pdf" \
    -F "documents=@/path/to/runbook.pdf" \
    -F 'document_metadata=[
      { "id": "policy_v2", "metadata": { "document_type": "policy", "status": "approved" } },
      { "id": "runbook_deploy", "metadata": { "document_type": "runbook" } }
    ]'
  ```
</CodeGroup>

### Path B: App sources

Pre-parsed records from connected apps. You supply app-native fields (`kind`, `provider`, `external_id`, and `fields`), [metadata](/essentials/v2/metadata), and optional attachments/comments. Use this when your application already extracts content from Slack messages, Notion pages, Gmail messages, tickets, or CRM records and you want HydraDB to preserve app structure instead of treating the item as generic text.

<CodeGroup>
  ```python Python SDK theme={"dark"}
  import json

  result = client.context.ingest(
      type="knowledge",
      database="acme_corp",
      app_knowledge=json.dumps([
          {
              "database": "acme_corp",
              "collection": "default",
              "id": "slack_C01_1716213600_000100",
              "title": "Pricing discussion - Slack #product",
              "kind": "message",
              "provider": "slack",
              "external_id": "1716213600.000100",
              "fields": {
                  "kind": "message",
                  "body": "We agreed on three tiers: Starter at $29, Pro at $79, Enterprise at $199.",
                  "author": "alice",
                  "thread_id": "1716213600.000100",
                  "created_at": "2026-05-20T10:00:00Z",
              },
              "metadata": {"channel": "product", "status": "finalized"},
              "additional_metadata": {"slack_ts": "1716213600.000100"},
          }
      ]),
  )
  ```

  ```typescript TypeScript SDK theme={"dark"}
  const result = await client.context.ingest({
    type: "knowledge",
    database: "acme_corp",
    appKnowledge: JSON.stringify([
      {
        database: "acme_corp",
        collection: "default",
        id: "slack_C01_1716213600_000100",
        title: "Pricing discussion - Slack #product",
        kind: "message",
        provider: "slack",
        external_id: "1716213600.000100",
        fields: {
          kind: "message",
          body: "We agreed on three tiers: Starter at $29, Pro at $79, Enterprise at $199.",
          author: "alice",
          thread_id: "1716213600.000100",
          created_at: "2026-05-20T10:00:00Z",
        },
        metadata: { channel: "product", status: "finalized" },
        additional_metadata: { slack_ts: "1716213600.000100" },
      },
    ]),
  });
  ```

  ```bash cURL theme={"dark"}
  curl -X POST 'https://api.hydradb.com/context/ingest' \
    -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
    -H "API-Version: 2" \
    -F "type=knowledge" \
    -F "database=acme_corp" \
    -F 'app_knowledge=[
      {
        "database": "acme_corp",
        "collection": "default",
        "id": "slack_C01_1716213600_000100",
        "title": "Pricing discussion - Slack #product",
        "kind": "message",
        "provider": "slack",
        "external_id": "1716213600.000100",
        "fields": {
          "kind": "message",
          "body": "We agreed on three tiers: Starter at $29, Pro at $79, Enterprise at $199.",
          "author": "alice",
          "thread_id": "1716213600.000100",
          "created_at": "2026-05-20T10:00:00Z"
        },
        "metadata": { "channel": "product", "status": "finalized" },
        "additional_metadata": { "slack_ts": "1716213600.000100" }
      }
    ]'
  ```
</CodeGroup>

**Response (both paths):**

```json theme={"dark"}
{
  "success": true,
  "data": {
    "success": true,
    "message": "Knowledge uploaded successfully",
    "results": [
      { "id": "policy_v2", "filename": "policy.pdf", "status": "queued", "error": null }
    ],
    "success_count": 1,
    "failed_count": 0
  },
  "error": null,
  "meta": {
    "request_id": "9d13aef4-02f4-4e73-8c62-4c2601d04f9d",
    "latency_ms": 12.3
  }
}
```

The `/context` and `/query` endpoints return this envelope shape. Access the payload via `response.data` (e.g. `response.data.results`), not at the top level.

***

## 5. Verifying processing

Ingestion is async, so you need to confirm each source actually finished before it'll show up in [query](/essentials/v2/query). Poll [`GET /context/status`](/api-reference/v2/endpoint/source-status) with the `id`s you got back from ingestion. The snippets below loop until every item reaches a terminal state (`completed` or `errored`):

<CodeGroup>
  ```python Python SDK theme={"dark"}
  import time

  ids = ["policy_v2", "runbook_deploy"]
  terminal = {"completed", "errored"}

  while True:
      status = client.context.status(
          database="acme_corp",
          collection="default",
          ids=ids,
      )
      if all(s.indexing_status in terminal for s in status.data.statuses):
          break
      time.sleep(5)
  ```

  ```typescript TypeScript SDK theme={"dark"}
  const ids = ["policy_v2", "runbook_deploy"];
  const terminal = new Set(["completed", "errored"]);

  let status;
  while (true) {
    status = await client.context.status({
      database: "acme_corp",
      collection: "default",
      ids: ids,
    });

    const allDone = status.data.statuses.every((s) => terminal.has(s.indexingStatus));
    if (allDone) break;

    await new Promise((r) => setTimeout(r, 5_000));
  }
  ```

  ```bash cURL theme={"dark"}
  while true; do
    response=$(curl -s -G 'https://api.hydradb.com/context/status' \
      -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
      -H "API-Version: 2" \
      --data-urlencode "database=acme_corp" \
      --data-urlencode "collection=default" \
      --data-urlencode "ids=policy_v2" \
      --data-urlencode "ids=runbook_deploy")

    statuses=$(echo "$response" | jq -r '.data.statuses[].indexing_status')
    if echo "$statuses" | grep -qv '^\(completed\|errored\)$'; then
      sleep 5
      continue
    fi
    echo "$response" | jq
    break
  done
  ```
</CodeGroup>

Each status in the `data` payload looks like this:

```json theme={"dark"}
{
  "statuses": [
    {
      "id": "policy_v2",
      "indexing_status": "completed",
      "success": true
    }
  ]
}
```

Status values move forward through `queued`, `processing`, `graph_creation`, and `completed` (or land in `errored` on failure). Poll every few seconds; most documents complete in 1 to 5 minutes. The full table is on [Ingestion Status](/api-reference/v2/endpoint/source-status).

<Info>
  **`graph_creation` is already queryable.** Items in this state are retrievable via [`POST /query`](/api-reference/v2/endpoint/query). Wait for `completed` only when you specifically need full graph traversal (`graph_context: true`).
</Info>

***

## 6. Key parameters

The full schema lives on [`POST /context/ingest`](/api-reference/v2/endpoint/ingest-context). Below are the fields you'll touch most often, grouped by where they live in the request.

### File metadata (`document_metadata` array items)

Send one item per file in `documents`, in the same order. The item count must match the file count.

| Field | Description |
| - | - |
| <Field name="id" type="string" /> | Optional stable identifier. Used as upsert key when `upsert: true`. |
| <Field name="metadata" type="object" /> | Filterable fields. Keys must match the database metadata schema. Up to 16 KiB as compact JSON. |
| <Field name="additional_metadata" type="object" /> | Free-form fields for display and bookkeeping. Filterable only when nested under `additional_metadata` in `metadata_filters`. Up to 1 KiB as compact JSON. |
| <Field name="relations" type="object" /> | Forcefully connect this source to others with `{ "ids": [...] }`. |

Any other key, including `title`, `type`, `url`, and `timestamp`, returns `400`. App sources accept `title`, `url`, and `timestamp`.

### App source model (`app_knowledge` array items)

| Field | Description |
| - | - |
| <Field name="id" type="string" recommended /> | Stable ID. Used as the upsert/delete/status key. |
| <Field name="database" type="string" recommended /> | Database that owns the item. Must match form-level `database` (formerly `tenant_id`) when present. |
| <Field name="collection" type="string" recommended /> | Logical partition. Must match form-level `collection` when present. |
| <Field name="title" type="string" recommended /> | Short title, subject, ticket title, or page name. |
| <Field name="kind" type="email, message, ticket, knowledge_base, comment, meeting, custom" recommended /> | App object kind. Selects the app parser for `fields`. |
| <Field name="provider" type="string" recommended /> | App namespace such as `slack`, `gmail`, `jira`, `notion`, `salesforce`, or `linear`. |
| <Field name="external_id" type="string" recommended /> | Stable provider ID for this exact item. Enables exact lookup and relation resolution. |
| <Field name="fields" type="object" required /> | App-native content and structure. Put primary text here (`body`, `description`, `title`, or `data` depending on `kind`). |
| <Field name="url" type="string" /> | Canonical URL or reference link. |
| <Field name="timestamp" type="ISO-8601 datetime" /> | Creation, sent, or last-updated time. |
| <Field name="metadata" type="object" /> | Filterable fields matching the database schema. |
| <Field name="additional_metadata" type="object" /> | Free-form fields. Filterable only when nested under `additional_metadata` in `metadata_filters`. |
| <Field name="attachments" type="array" /> | Optional related attachments. Include extracted text in `attachments[].content.text` or `.markdown`. |
| <Field name="comments" type="array" /> | Optional embedded snapshot comments. For live incremental comments, prefer a first-class `kind: "comment"` source. |
| <Field name="relations" type="array" /> | App-native relation array: `[{ "predicate": "linked_to", "target": { "external_id": "AUTH-123", "provider": "jira" } }]`. |

### Request-level fields

| Field | Default | Description |
| - | - | - |
| <Field name="database" type="string" required /> | None | Target database. |
| <Field name="collection" type="string" /> | default collection | Scope this knowledge to a collection. If omitted, knowledge goes to the database's default collection, which a query naming another `collection` does not read. |
| <Field name="upsert" type="boolean" /> | `true` | Replace existing items with the same `id`. |

***

## 7. Forceful relations

Forceful relations let you **declare** links between Knowledge sources at ingestion time. When a query matches a source that has forceful relations attached, HydraDB pulls the linked sources into the `additional_context` field of the query response alongside the primary chunks.

This is separate from the entity and relationship [context graph](/essentials/v2/context-graphs) HydraDB builds automatically. Graph extraction is *derived* from content; forceful relations are *declared* by you, which is useful when you know two documents belong together (a contract and its addendum, a runbook and its troubleshooting guide) but the text doesn't make the connection explicit. Forceful relations link whole **sources**; to declare the full entity/relation graph *within* a single document (replacing extraction for it), use [Bring Your Own Graph](/essentials/v2/bring-your-own-graph).

For file uploads, set `relations.ids` on the matching `document_metadata` item. For app sources, prefer the app-native `relations[]` shape so each link can carry a predicate and provider-aware target:

```json theme={"dark"}
{
  "relations": [
    {
      "predicate": "linked_to",
      "target": {
        "id": "auth-troubleshooting"
      }
    },
    {
      "predicate": "related_to",
      "target": {
        "external_id": "AUTH-FAQ",
        "provider": "notion"
      }
    }
  ]
}
```

<CodeGroup>
  ```python Python SDK theme={"dark"}
  client.context.ingest(
      type="knowledge",
      database="acme_corp",
      app_knowledge=json.dumps([
          {
              "database": "acme_corp",
              "collection": "default",
              "id": "auth-guide",
              "title": "Authentication Guide",
              "kind": "knowledge_base",
              "provider": "notion",
              "external_id": "AUTH-GUIDE",
              "fields": {"kind": "knowledge_base", "title": "Authentication Guide", "body": "..."},
              "relations": [
                  {"predicate": "linked_to", "target": {"id": "auth-troubleshooting"}},
                  {"predicate": "related_to", "target": {"external_id": "AUTH-FAQ", "provider": "notion"}},
              ],
          }
      ]),
  )
  ```

  ```typescript TypeScript SDK theme={"dark"}
  await client.context.ingest({
    type: "knowledge",
    database: "acme_corp",
    appKnowledge: JSON.stringify([
      {
        id: "auth-guide",
        title: "Authentication Guide",
        kind: "knowledge_base",
        provider: "notion",
        external_id: "AUTH-GUIDE",
        fields: { kind: "knowledge_base", title: "Authentication Guide", body: "..." },
        relations: [
          { predicate: "linked_to", target: { id: "auth-troubleshooting" } },
          { predicate: "related_to", target: { external_id: "AUTH-FAQ", provider: "notion" } },
        ],
      },
    ]),
  });
  ```

  ```bash cURL theme={"dark"}
  curl -X POST 'https://api.hydradb.com/context/ingest' \
    -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
    -H "API-Version: 2" \
    -F "type=knowledge" \
    -F "database=acme_corp" \
    -F 'app_knowledge=[
      {
        "id": "auth-guide",
        "title": "Authentication Guide",
        "kind": "knowledge_base",
        "provider": "notion",
        "external_id": "AUTH-GUIDE",
        "fields": { "kind": "knowledge_base", "title": "Authentication Guide", "body": "..." },
        "relations": [
          { "predicate": "linked_to", "target": { "id": "auth-troubleshooting" } },
          { "predicate": "related_to", "target": { "external_id": "AUTH-FAQ", "provider": "notion" } }
        ]
      }
    ]'
  ```
</CodeGroup>

At query time, linked sources are returned in the `additional_context` field of the [query response](/api-reference/v2/endpoint/query). `query_forceful_relations` defaults to `true`, so you only need to set it explicitly if you want to disable forceful-relation expansion.

<Warning>
  **Knowledge items only.** Forceful relations work within the same store. Pointing a Knowledge item at a Memory ID will silently return nothing at query time.
</Warning>

<Warning>
  **`thinking` mode only.** Forceful-relation context is fetched only when `mode: "thinking"` is set on the query request (or when `mode: "auto"` routes to thinking). In `fast` mode, `query_forceful_relations` is ignored even if set to `true`.
</Warning>

***

## 8. Metadata on knowledge

[Metadata](/essentials/v2/metadata) is how you bridge structured filters and semantic search. Attach `metadata` at ingestion to enable deterministic narrowing at query time:

```json theme={"dark"}
{
  "metadata": { "document_type": "runbook", "environment": "production" }
}
```

Then at query time:

```json theme={"dark"}
{
  "query": "How do I rotate API keys?",
  "metadata_filters": { "document_type": "runbook", "environment": "production" }
}
```

Declare the `metadata` keys you filter on in the database metadata schema. In a database with a schema, ingest rejects undeclared keys, so filtering on one returns nothing. `additional_metadata` is stored alongside the source for display and bookkeeping; to filter on it, nest the keys under `additional_metadata` in `metadata_filters` (e.g. `{"metadata_filters": {"additional_metadata": {"author": "Alice"}}}`). `document_metadata` is accepted only as a legacy alias for that nested filter namespace.

See [Metadata](/essentials/v2/metadata) for schema design, filter examples, and the `metadata` vs `additional_metadata` decision rule. See [Create Database](/api-reference/v2/endpoint/create-tenant) for the schema field configuration (`enable_dense_embedding`, `enable_sparse_embedding`).

***

## 9. Common mistakes

| Mistake | What goes wrong | Fix |
| - | - | - |
| Querying Knowledge with `type: "memory"` | Nothing returned: wrong store | Use `type: "knowledge"` (or `"all"`) on [`POST /query`](/api-reference/v2/endpoint/query). |
| Using `additional_metadata` for hot filterable fields | Filters work only with nested syntax and are less efficient than schema-backed metadata filters | Put hot filter keys in `metadata` and declare them in the [database schema](/api-reference/v2/endpoint/create-tenant), or filter under `metadata_filters: {"additional_metadata": {...}}` for occasional/free-form fields. See [Metadata](/essentials/v2/metadata). |
| Not polling [`/context/status`](/api-reference/v2/endpoint/source-status) before querying | Chunks aren't indexed yet, so query results are empty | Wait for `indexing_status: completed` (or `graph_creation` for partial availability). |
| Sending `document_metadata` as a JSON object instead of a JSON string | `multipart/form-data` validation fails | Stringify `document_metadata` before sending in the `-F` field. |
| Omitting `database` | `400` error (`database is required`) | Required on every call. |

***

## Related

* [Memories](/essentials/v2/memories): the per-user counterpart to Knowledge
* [Query](/essentials/v2/query): how [`POST /query`](/api-reference/v2/endpoint/query) retrieves Knowledge
* [Metadata](/essentials/v2/metadata): designing filterable schemas
* [Context Graphs](/essentials/v2/context-graphs): how the graph layer enriches retrieval
* [Ingest Content: API Reference](/api-reference/v2/endpoint/ingest-context)
* [Ingestion Status: API Reference](/api-reference/v2/endpoint/source-status)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.