> ## Documentation Index
> Fetch the complete documentation index at: https://cortex-e852fafe-t3code-rewrite-docs-declutter.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Metadata and Filters

> How HydraDB uses metadata to scope query results with schema-backed fields, free-form additional metadata, source metadata edits, and exact metadata filters.

Metadata is structured data attached to [Knowledge](/essentials/v2/knowledge) and [Memories](/essentials/v2/memories). Use it to require a matching field before results are ranked, such as `department=legal`, `region=us`, `status=published`, or `author=alice`.

HydraDB has two metadata layers:

| Layer | Sent at ingest as | Query filter shape | Best for |
| - | - | - | - |
| Database metadata | `metadata` | Top-level keys in `metadata_filters` | Fields you filter on often, such as department or region. Declare their names and types in `database_metadata_schema`. |
| Additional metadata | `additional_metadata` | Nested under `metadata_filters.additional_metadata` | Free-form per-source fields for display, citations, debugging, external IDs, or occasional filters. |

If you will filter on a field in most queries, it belongs in `metadata` and the [database schema](/api-reference/v2/endpoint/create-tenant). If it's ad-hoc or unique to one document, use `additional_metadata` instead.

```json theme={"dark"}
{
  "metadata_filters": {
    "department": "legal",
    "additional_metadata": {
      "author": "alice"
    }
  }
}
```

***

## 1. Choose the right metadata layer

| I want to… | Put it in… | Filter with… | Notes |
| - | - | - | - |
| Scope most queries by a field like department, region, plan, customer, or status | `metadata` | `metadata_filters: { "department": "legal" }` | Declare the field in `database_metadata_schema`. |
| Store source-specific fields like author, Slack timestamp, external ID, document version | `additional_metadata` | `metadata_filters: { "additional_metadata": { "author": "alice" } }` | No schema required. Better for occasional filters and display metadata. |
| Combine hard scoping with semantic search | Both | `metadata_filters` plus your natural-language `query` | Filters narrow candidates; ranking still uses the query. |
| Search semantically over a metadata text field | `metadata` with `enable_dense_embedding: true` | Put the desired concept in `query` | Only supported for `VARCHAR` fields. Do not put fuzzy concepts in `metadata_filters`. |
| Search by keyword over a metadata text field | `metadata` with `enable_sparse_embedding: true` | Use normal `/query` text/BM25 behavior | Only supported for `VARCHAR` fields. |
| Partition by user, workspace, or team | `collection` | Send `collection` on every request | Use metadata filters inside that partition, not as a replacement for it. |

When to send `database` and `collection`: [Multi-tenant](/essentials/v2/multi-tenant#2-when-to-use-each). The older names `tenant_id` and `sub_tenant_id` still work as deprecated aliases.

***

## 2. How metadata filters run

HydraDB applies `metadata_filters` in three steps:

1. **Prefilter:** `metadata` and `additional_metadata` filters are resolved to matching source IDs in MongoDB.
2. **Scoped retrieval:** vector and BM25 retrieval search only those source IDs.
3. **Recheck:** the retrieved passages are checked again against the requested metadata. This also protects graph expansion and fallback/retry paths from leaking excluded sources.

That means a valid filter with no matching sources returns an empty result set; HydraDB does not silently widen it into an unfiltered search.

### Filter semantics

| Behavior | Contract |
| - | - |
| Multiple keys | AND logic. Every provided key/value must match. |
| Database + additional metadata | AND logic across both layers. |
| Operators | Each `metadata` (top-level key) takes `{"equals": value}`, `{"contains": value}`, or `{"contains_any": [values]}`. Not available inside `additional_metadata`; use the bare forms there. `contains` and `contains_any` need a `VARCHAR` field; `equals` works on any type. See [Filter operators](#filter-operators). |
| Values | `equals` is exact equality against the **whole** stored value. |
| Lists | **ANY, not ALL.** A list matches a source holding **any one** of the listed values (OR/IN), so `tags: ["alpha", "beta"]` matches a source tagged `"alpha"` and a source tagged `"beta"`. Write it as `{"contains_any": [...]}` in `metadata` (top-level key), or as a bare list in `additional_metadata`. Adding values widens the result set; there is no ALL/AND operator inside a single key. Lists are supported on `VARCHAR` fields only; a list passed for a declared field of another type is rejected with `400`. |
| Null values | Rejected with `400` (`metadata_filters values cannot be null`). |
| Range / fuzzy | Not supported in user `metadata_filters`. Put fuzzy concepts in `query`, use semantic metadata fields, or post-process client-side. |
| OR | Within one key, use `contains_any`. Across different keys, run multiple queries and union client-side. |
| Graph context | Respects metadata filters; graph paths from excluded sources are removed. |
| Size | Each list holds at most **500** values, and the whole `metadata_filters` object is capped at **64 KiB**. See [Filter size limits](#filter-size-limits). |

<Warning>
  `metadata_filters` are hard constraints, not semantic hints. A filter like `{ "mood": "happy" }` requires an exact stored value; it does not expand to related values like "joyful" or "cheerful". To search metadata text semantically, declare a `VARCHAR` field with `enable_dense_embedding` and include the concept in the main `query`.
</Warning>

### Filter operators

Each `metadata` (top-level) key in `metadata_filters` takes an operator object naming the comparison you want.

| Operator | Value | Matches |
| - | - | - |
| `equals` | a single value | sources whose field is **exactly** that value |
| `contains` | a single value | sources whose field **holds** that value, at any position in a multi-value field |
| `contains_any` | an array | sources holding **any one** of the listed values (OR/IN) |

```json theme={"dark"}
{
  "metadata_filters": {
    "department": { "equals": "legal" },
    "attendee_emails": { "contains": "b@company.com" },
    "tags": { "contains_any": ["alpha", "beta"] }
  }
}
```

Given four sources with these `emails` values:

| Source | `emails` |
| - | - |
| A | `a@company.com` |
| B | `a@company.com,b@company.com` |
| C | `c@company.com` |
| D | `x@company.com,a@company.com,z@company.com` |

| Filter | Returns |
| - | - |
| `{"equals": "a@company.com"}` | A |
| `{"contains": "a@company.com"}` | A, B, D |
| `{"contains_any": ["a@company.com", "c@company.com"]}` | A, B, C, D |

`contains` is position-independent: it matches whether the value is the only one, the first, or in the middle.

Operators apply to `metadata` (top-level keys) only. Inside `additional_metadata`, use a bare scalar for an exact match or a bare array to match any listed value.

<Warning>
  An operator used inside `additional_metadata` is **not** rejected. It is read as an exact-match filter against a stored object, so on a normal field it matches nothing and the request returns `200` with an empty result rather than an error.
</Warning>

A known operator given the wrong operand type, or several operators in one object, is rejected with `400` rather than silently returning no results. A **misspelled** operator is not: `{"contian": "x"}` is indistinguishable from a filter for a stored object with that key, so it is left alone and matches nothing.

<Warning>
  **`contains`, `contains_any` and `equals` are reserved key names:** An object built only from them is read as an operator, so it can no longer be used to exact-match a stored object:

  | Filter value | Read as |
  | - | - |
  | `{"contains": "x"}` | the `contains` operator |
  | `{"contains": "a", "equals": "b"}` | rejected with `400` (every key is an operator name) |
  | `{"contains": "a", "other": 1}` | an exact-match object filter, unchanged |

  This only affects a `JSON`-typed field storing an object whose keys are all drawn from those three words. If you need to match such an object, rename the nested key or the field.
</Warning>

<Note>
  The bare forms still work and are unchanged, but are **deprecated** in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as `equals`, a bare array as `contains_any`, and a bare single-element array as `contains`, so `{"emails": "a@x"}` and `{"emails": ["a@x"]}` differ by one character and return different results.
</Note>

### Storing multiple values in one field

A declared schema field holds a single value. To store several values in one field, declare it as `VARCHAR` and join the values with commas:

```json theme={"dark"}
{
  "metadata": {
    "attendee_emails": "a@company.com,b@company.com"
  }
}
```

Then filter for one member with `contains`:

```json theme={"dark"}
{
  "metadata_filters": {
    "attendee_emails": { "contains": "b@company.com" }
  }
}
```

Points to know:

* `data_type: "array"` is **not** accepted on a declared field. Declaring one is rejected with `400`.
* Sending an array **value** for a declared `VARCHAR` field is also rejected: `metadata field "attendee_emails" must be of type string, got array`. Join the values yourself.
* The comma is the separator, so a value that itself contains a comma will not match as expected. Use a field per value, or a different value format, if your values can contain commas.
* Size the field for the whole joined string. `max_length` defaults to `1024` and its maximum is `65535`, and it **cannot be raised after the field is created**, so declare it large enough up front.
* `equals` compares the entire joined string, so it is rarely what you want on a multi-value field. Use `contains`.

***

## 3. Plan the metadata schema

Before your first ingest:

* **Plan scoping fields before first ingest:** [`PATCH /databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema) can add fields later, but existing fields cannot be renamed, retyped, reflagged, or deleted. Undeclared scope keys are not ignored at query time: they match no sources, because a database with a schema rejects undeclared keys on ingest and edit. The exception is `connector_id` and `provider`, which [connectors](/essentials/v2/connectors#metadata-on-synced-objects) write on every synced object. If you'll scope on it more than once, declare it.
* **Pick `metadata` for hot paths, `additional_metadata` for cold ones:** Top-level `metadata` filters are prefiltered; `additional_metadata` filters that cannot be prefiltered need a post-retrieval pass with over-fetch.
* **Keep keys stable:** Renaming a metadata key requires re-ingesting affected sources.
* **Ingested metadata is editable in place:** Once a source is indexed, [`PATCH /context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata) merges new `metadata` and `additional_metadata` values into it without re-ingesting. See [section 7](#7-update-metadata-without-re-ingesting).
* **Don't substitute `metadata_filters` for `collection`:** `metadata_filters` scopes results *inside* a partition. For partitioning by user, team, or workspace, use [`collection`](/essentials/v2/multi-tenant).

***

## 4. Minimal working example

Two phases: **set up metadata** (declare the schema, then attach values at ingest), then **scope at query**.

### Step 1: Create metadata

The schema lives at the database level; values land on each source at ingest time. Both happen before any query.

#### Step 1a: Declare the schema at database creation

<CodeGroup>
  ```bash cURL theme={"dark"}
  curl -X POST 'https://api.hydradb.com/databases' \
    -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
    -H "API-Version: 2" \
    -H "Content-Type: application/json" \
    -d '{
      "database": "acme_corp",
      "database_metadata_schema": [
        {
          "name": "department",
          "data_type": "VARCHAR"
        },
        {
          "name": "priority",
          "data_type": "INT64"
        },
        {
          "name": "summary_label",
          "data_type": "VARCHAR",
          "enable_dense_embedding": true,
          "enable_sparse_embedding": true
        }
      ]
    }'
  ```

  ```python Python SDK theme={"dark"}
  client.databases.create(
      database="acme_corp",
      database_metadata_schema=[
          {"name": "department", "data_type": "VARCHAR"},
          {"name": "priority", "data_type": "INT64"},
          {
              "name": "summary_label",
              "data_type": "VARCHAR",
              "enable_dense_embedding": True,
              "enable_sparse_embedding": True,
          },
      ],
  )
  ```

  ```typescript TypeScript SDK theme={"dark"}
  await client.databases.create({
    database: "acme_corp",
    databaseMetadataSchema: [
      { name: "department", dataType: "VARCHAR" },
      { name: "priority", dataType: "INT64" },
      {
        name: "summary_label",
        dataType: "VARCHAR",
        enableDenseEmbedding: true,
        enableSparseEmbedding: true,
      },
    ],
  });
  ```
</CodeGroup>

### Schema field options

| Field | Type / values | Purpose |
| - | - | - |
| `name` | string | Metadata key. Must start with a letter, contain only letters/numbers/underscores, and not use reserved system names such as `source_id`, `chunk_id`, or `metadata`. |
| `data_type` | `VARCHAR`, `BOOL`, `INT8`, `INT16`, `INT32`, `INT64`, `FLOAT`, `DOUBLE`, `JSON` or friendly aliases `string`, `boolean`, `integer`, `float`, `object` | Defaults to `VARCHAR`. `ARRAY` is **not** supported and is rejected with `400`; for multi-value fields see [Storing multiple values in one field](#storing-multiple-values-in-one-field). |
| `max_length` | integer | Max length for `VARCHAR`. Default `1024`, maximum `65535`. This sizes one declared schema field, and is **not** the limit on how much metadata you can send per request (for that see [Size limits](#size-limits)). |
| `enable_dense_embedding` | boolean | Adds a dense semantic-search lane for a `VARCHAR` metadata field. |
| `enable_sparse_embedding` | boolean | Adds a sparse/BM25 keyword-search lane for a `VARCHAR` metadata field. |
| `searchable` | boolean | Backward-compatible input accepted by the API, but do not rely on it to replace explicit `enable_dense_embedding` / `enable_sparse_embedding`. Set those flags directly. |

Limits and guardrails:

* Up to 32 custom metadata fields.
* Up to **6 embedding-enabled** fields per database. `enable_dense_embedding` and `enable_sparse_embedding` each count as one, so a field with both set counts as two. Exceeding the limit fails database creation with `400` before anything is provisioned.
* Field names are unique case-insensitively.
* Dense/sparse embedding flags are only valid on `VARCHAR` fields.
* Runtime `metadata` values must match the declared type when the database has a schema.
* Unknown `metadata` keys are rejected on ingest and edit when the database has a non-empty schema.
* Each metadata layer has a byte budget per request. See [Size limits](#size-limits).

### Size limits

Every request that attaches metadata is checked against two caps, on ingest and on
metadata edit alike:

| Layer | Send as | Cap |
| - | - | - |
| Database metadata | `metadata` at ingest, `database_metadata` on edit | **16 KiB** (16,384 bytes) |
| Document metadata | `additional_metadata` | **1 KiB** (1,024 bytes) |

Older spellings are still accepted, but **not uniformly**. Which one works
depends on the endpoint:

| Alias | On `/context/ingest` | On `PATCH /context/{id}/metadata` |
| - | - | - |
| `tenant_metadata` | accepted | accepted (`database_metadata` wins if both are sent) |
| `document_metadata` | accepted | **rejected with `400`**: send `additional_metadata` |

Accepted aliases have the same cap as the canonical field.

The cap applies to the **whole map**, not to any one value, and it is measured on
the map's compact JSON encoding in UTF-8 bytes. Three consequences worth planning
around:

* **Keys and punctuation count:** Quotes, colons, commas and braces all count toward
  the byte limit.
* **Bytes, not characters:** Accented Latin characters cost 2 bytes, most CJK
  characters 3, and emoji 4.
* **Budget in bytes from the start:** A 950-character summary sounds comfortably
  under a 1 KiB cap, but with two small sibling keys it serializes to 1,015
  bytes (65 bytes of that is structure alone). Push the summary to 1,000
  characters and the request is rejected at 1,065 bytes.

```json Document metadata: 1,015 bytes, just inside the 1 KiB cap theme={"dark"}
{"title":"Q3 Board Deck","author":"ada@example.com","summary":"<950 characters>"}
```

Exceeding either cap fails the whole request with `400` before anything is
ingested. The message names the offending field and reports both numbers, so you
can see exactly how far over you are:

```json theme={"dark"}
{
  "success": false,
  "data": null,
  "error": {
    "code": "INVALID_INPUT",
    "message": "additional_metadata is too large (1065 bytes when serialized; the maximum is 1024). Reduce the number or size of metadata fields."
  }
}
```

On [`PATCH /context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata)
the same message is prefixed with `invalid metadata edit:`.

<Tip>
  If a document needs more than 1 KiB of descriptive metadata, put the long text
  in the document body where it gets chunked and embedded, and keep
  `additional_metadata` for the short values you actually filter on.
</Tip>

### Filter size limits

The caps above bound the metadata you **store**. `metadata_filters` on
[`/query`](/api-reference/v2/endpoint/query) has its own, separate pair. These
bound what you **send at query time** and are unrelated to how much metadata a
source carries:

| Limit | Cap |
| - | - |
| Values in any one list | **500** |
| Whole `metadata_filters` object | **64 KiB** (65,536 bytes) |

Measured the same way (compact JSON, UTF-8 bytes, field names and punctuation
counted), and the object total includes the nested `additional_metadata` dict.

The object total is measured **after** operator objects are reduced to their
values, so `{"contains": "x"}` counts as `["x"]` and the operator keyword itself
costs nothing. Both spellings of the same filter cost the same, because the cap
bounds the expression sent to the vector store, which the spelling does not
change.

Both exist because every value in a list is expanded into the filter expression
sent to the vector store. The per-list cap catches one runaway list; the object
cap catches many individually-legal lists adding up. Twenty lists of 500 values
are each within the element cap but total roughly 127 KiB, so the object cap is
what rejects them.

Over either limit returns `400` before the query runs, naming the offending key
or the actual byte count:

```
metadata_filters.customer_id (got 743) must contain at most 500 values

metadata_filters is too large (130251 bytes when serialized; the maximum is 65536).
Reduce the number or size of filter values.
```

<Tip>
  Needing far more than 500 values in one filter usually means the constraint
  belongs in the data rather than the query: add a metadata field that groups
  those values (a segment, tier, or cohort key) and filter on that instead.
</Tip>

### Add schema fields later

You can add database metadata fields after database creation with [`PATCH /databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema). This is additive only:

* add new fields: yes
* delete fields: no
* change type/flags of existing fields: no

<Warning>
  Fields added this way cannot turn on dense or sparse search lanes: the endpoint rejects `enable_dense_embedding` and `enable_sparse_embedding`. Declare semantic or keyword metadata fields when you create the database.
</Warning>

***

## 5. Attach metadata at ingest

For knowledge ingestion, send `metadata` and `additional_metadata` on each `document_metadata[]` item or `app_knowledge[]` item.

<CodeGroup>
  ```bash cURL theme={"dark"}
  curl -X POST 'https://api.hydradb.com/context/ingest' \
    -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
    -H "API-Version: 2" \
    -F "type=knowledge" \
    -F "database=acme_corp" \
    -F 'app_knowledge=[
      {
        "database": "acme_corp",
        "collection": "default",
        "id": "auth-controls-001",
        "kind": "knowledge_base",
        "fields": {
          "title": "Authentication controls",
          "body": "Authentication controls are reviewed quarterly for SOC2."
        },
        "metadata": {
          "department": "security",
          "priority": 7,
          "summary_label": "quarterly access control review"
        },
        "additional_metadata": {
          "author": "alice",
          "doc_version": 3
        }
      }
    ]'
  ```

  ```python Python SDK theme={"dark"}
  import json

  client.context.ingest(
      type="knowledge",
      database="acme_corp",
      app_knowledge=json.dumps([
          {
              "database": "acme_corp",
              "collection": "default",
              "id": "auth-controls-001",
              "kind": "knowledge_base",
              "fields": {
                  "title": "Authentication controls",
                  "body": "Authentication controls are reviewed quarterly for SOC2.",
              },
              "metadata": {
                  "department": "security",
                  "priority": 7,
                  "summary_label": "quarterly access control review",
              },
              "additional_metadata": {"author": "alice", "doc_version": 3},
          }
      ]),
  )
  ```

  ```typescript TypeScript SDK theme={"dark"}
  await client.context.ingest({
    type: "knowledge",
    database: "acme_corp",
    appKnowledge: JSON.stringify([
      {
        database: "acme_corp",
        collection: "default",
        id: "auth-controls-001",
        kind: "knowledge_base",
        fields: {
          title: "Authentication controls",
          body: "Authentication controls are reviewed quarterly for SOC2.",
        },
        metadata: {
          department: "security",
          priority: 7,
          summary_label: "quarterly access control review",
        },
        additional_metadata: { author: "alice", doc_version: 3 },
      },
    ]),
  });
  ```
</CodeGroup>

<Warning>
  For `type=memory`, the `memories` multipart field is already JSON-stringified, and each memory item's `metadata` and `additional_metadata` are plain objects inside it. Do not stringify them a second time.
</Warning>

***

## 6. Query with metadata filters

Mix database-level (top-level) and document-level (nested) scopes in the same `metadata_filters` object:

<CodeGroup>
  ```bash cURL theme={"dark"}
  curl -X POST 'https://api.hydradb.com/query' \
    -H "Authorization: Bearer $HYDRA_DB_API_KEY" \
    -H "API-Version: 2" \
    -H "Content-Type: application/json" \
    -d '{
      "database": "acme_corp",
      "collection": "default",
      "query": "How do access control reviews work?",
      "type": "knowledge",
      "query_by": "hybrid",
      "metadata_filters": {
        "department": "security",
        "priority": 7,
        "additional_metadata": {
          "author": "alice"
        }
      }
    }'
  ```

  ```python Python SDK theme={"dark"}
  result = client.query(
      database="acme_corp",
      collection="default",
      query="How do access control reviews work?",
      type="knowledge",
      query_by="hybrid",
      metadata_filters={
          "department": "security",
          "priority": 7,
          "additional_metadata": {"author": "alice"},
      },
  )
  ```

  ```typescript TypeScript SDK theme={"dark"}
  const result = await client.query({
    database: "acme_corp",
    collection: "default",
    query: "How do access control reviews work?",
    type: "knowledge",
    queryBy: "hybrid",
    metadataFilters: {
      department: "security",
      priority: 7,
      additional_metadata: { author: "alice" },
    },
  });
  ```
</CodeGroup>

Use the legacy alias only when maintaining older clients:

```json theme={"dark"}
{
  "metadata_filters": {
    "document_metadata": { "author": "alice" }
  }
}
```

If both aliases are present and both are objects, `additional_metadata` wins on conflicts.

***

## 7. Update metadata without re-ingesting

Use [`PATCH /context/{id}/metadata`](/api-reference/v2/endpoint/update-source-metadata) when you know the source ID and need to update metadata in place.

```json theme={"dark"}
{
  "database": "acme_corp",
  "collection": "default",
  "database_metadata": {
    "department": "legal"
  },
  "additional_metadata": {
    "author": "grace"
  }
}
```

Behavior:

* The source must already exist.
* `collection` is required.
* At least one of `database_metadata` (deprecated alias `tenant_metadata`), `additional_metadata`, or `acl` is required.
* The update is a merge/upsert: sent keys are inserted or overwritten; omitted keys are preserved.
* `document_metadata` is not accepted on this endpoint; use `additional_metadata`.
* The same endpoint accepts `acl` to change who may retrieve the source. Unlike metadata, `acl` **replaces** rather than merges, and an `acl`-only body is a valid edit. See [Access Control](/essentials/v2/access-control).
* Updates to database metadata fields without an embedding flag are MongoDB-only and take effect for filters and listing.
* If an edited database metadata field has `enable_dense_embedding` or `enable_sparse_embedding`, HydraDB synchronously syncs the relevant vector store lane and reports `vector_sync_required` / `vector_synced` in the response (the `milvus_sync_required` / `milvus_synced` aliases are still emitted, deprecated).

For full document/content replacement, re-ingest with `upsert: true` and the same source `id`. Upsert replaces the source payload and metadata supplied by ingestion.

***

## 8. Listing with metadata filters

Use [`POST /context/list`](/api-reference/v2/endpoint/list-documents) when you want to browse or page sources rather than run semantic retrieval:

```json theme={"dark"}
{
  "database": "acme_corp",
  "collection": "default",
  "type": "knowledge",
  "filters": {
    "metadata": { "department": "legal" },
    "additional_metadata": { "author": "alice" },
    "source_fields": { "type": "slack" }
  }
}
```

`/context/list` also accepts legacy aliases `tenant_metadata` for `metadata` and `document_metadata` for `additional_metadata`.

***

## 9. Common mistakes

| Symptom | Cause | Fix |
| - | - | - |
| `metadata_filters` returns no results | Key isn't declared in `database_metadata_schema`, so no source holds it (undeclared keys are not ignored) | Add the field with [`PATCH /databases/{database}/metadata-schema`](/api-reference/v2/endpoint/update-metadata-schema) and ingest or edit sources with a value, or move the field to `additional_metadata` and nest the scope under `additional_metadata`. |
| Scope on an `additional_metadata` key returns nothing | Scope passed as top-level key, so it is matched against `metadata` | Nest it: `metadata_filters: { additional_metadata: { author: "alice" } }`. |
| Schema field changes don't take effect | Existing schema fields cannot be renamed, retyped, reflagged, or deleted; you can only add fields | Create a new database with the corrected schema. There's no in-place migration for existing fields. |
| Query returns 0 results after adding a scope | Over-scoping: combined constraints exclude everything | Start with the minimum hard constraints; add scopes one at a time and re-check counts. |
| Range / fuzzy scope doesn't work | `metadata_filters` supports only `equals`, `contains`, and `contains_any` | Move that constraint into the `query` text, use a different `query_by`, or apply it in your own post-scoping. |
| `additional_metadata` scope feels slow | It can be: when it cannot be prefiltered, it over-fetches about 3x to compensate for post-retrieval matching | Move the hot field into `metadata` and declare it in the schema. |
| Metadata edit returns 400 | Unknown key, wrong type, over-size payload (16 KiB database metadata / 1 KiB document metadata), too-deep nesting, reserved key, or missing `collection` | Check schema and [size limits](#size-limits); send only `database_metadata` / `additional_metadata` / `acl`. |
| Dense/sparse metadata edit rejects `null` | Null would leave stale semantic metadata vectors | Set a non-null value or re-ingest with the desired metadata. |
| Need a semantic or keyword metadata field after creation | The schema update endpoint rejects dense and sparse flags | Declare those fields at database creation, or re-ingest into a new database with the final schema. |

***

## 10. Advanced patterns

**Stacked scopes with collection partitioning:** Use `collection` for the partition (per-user, per-workspace), and use `metadata_filters` to scope *inside* that partition. They're complementary, not interchangeable. See [Multi-Tenant](/essentials/v2/multi-tenant).

**Published vs draft:** Add a `status` field to your schema; tag every source with `metadata.status = "draft" | "published"`; pass `metadata_filters: { status: "published" }` on user-facing queries. Keeps work-in-progress out of customer answers automatically.

**Multi-language corpora:** Add a `language` field; route each query to the right language by passing `metadata_filters: { language: detect_language(query) }`.

**Review schema changes:** Treat `database_metadata_schema` as a data contract. Existing field definitions are immutable, so changing them requires re-ingestion.

***

## Related

* [Knowledge](/essentials/v2/knowledge): what attaches to knowledge sources
* [Memories](/essentials/v2/memories): metadata fields on memory items
* [Multi-Tenant Support](/essentials/v2/multi-tenant): partitioning vs scoping
* [Query](/essentials/v2/query): how `metadata_filters` interact with ranking
* [Create Database (API Reference)](/api-reference/v2/endpoint/create-tenant): full `database_metadata_schema` reference
* [Query (API Reference)](/api-reference/v2/endpoint/query): full `metadata_filters` reference
* [List Context](/api-reference/v2/endpoint/list-documents): source browsing filters
* [Access Control](/essentials/v2/access-control): restricting who may retrieve a document, which is not a metadata filter
* [Update Source Metadata](/api-reference/v2/endpoint/update-source-metadata): point metadata edits
* [Update Metadata Schema](/api-reference/v2/endpoint/update-metadata-schema): additive database schema changes


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.