Skip to main content
Metadata is structured data attached to Knowledge and Memories. Use it to require a matching field before results are ranked, such as department=legal, region=us, status=published, or author=alice. HydraDB has two metadata layers: If you will filter on a field in most queries, it belongs in metadata and the database schema. If it’s ad-hoc or unique to one document, use additional_metadata instead.

1. Choose the right metadata layer

When to send database and collection: Multi-tenant. The older names tenant_id and sub_tenant_id still work as deprecated aliases.

2. How metadata filters run

HydraDB applies metadata_filters in three steps:
  1. Prefilter: metadata and additional_metadata filters are resolved to matching source IDs in MongoDB.
  2. Scoped retrieval: vector and BM25 retrieval search only those source IDs.
  3. Recheck: the retrieved passages are checked again against the requested metadata. This also protects graph expansion and fallback/retry paths from leaking excluded sources.
That means a valid filter with no matching sources returns an empty result set; HydraDB does not silently widen it into an unfiltered search.

Filter semantics

metadata_filters are hard constraints, not semantic hints. A filter like { "mood": "happy" } requires an exact stored value; it does not expand to related values like “joyful” or “cheerful”. To search metadata text semantically, declare a VARCHAR field with enable_dense_embedding and include the concept in the main query.

Filter operators

Each metadata (top-level) key in metadata_filters takes an operator object naming the comparison you want.
Given four sources with these emails values: contains is position-independent: it matches whether the value is the only one, the first, or in the middle. Operators apply to metadata (top-level keys) only. Inside additional_metadata, use a bare scalar for an exact match or a bare array to match any listed value.
An operator used inside additional_metadata is not rejected. It is read as an exact-match filter against a stored object, so on a normal field it matches nothing and the request returns 200 with an empty result rather than an error.
A known operator given the wrong operand type, or several operators in one object, is rejected with 400 rather than silently returning no results. A misspelled operator is not: {"contian": "x"} is indistinguishable from a filter for a stored object with that key, so it is left alone and matches nothing.
contains, contains_any and equals are reserved key names: An object built only from them is read as an operator, so it can no longer be used to exact-match a stored object:This only affects a JSON-typed field storing an object whose keys are all drawn from those three words. If you need to match such an object, rename the nested key or the field.
The bare forms still work and are unchanged, but are deprecated in favour of the operators, because the comparison they perform is inferred from the JSON shape rather than stated. A bare scalar behaves as equals, a bare array as contains_any, and a bare single-element array as contains, so {"emails": "a@x"} and {"emails": ["a@x"]} differ by one character and return different results.

Storing multiple values in one field

A declared schema field holds a single value. To store several values in one field, declare it as VARCHAR and join the values with commas:
Then filter for one member with contains:
Points to know:
  • data_type: "array" is not accepted on a declared field. Declaring one is rejected with 400.
  • Sending an array value for a declared VARCHAR field is also rejected: metadata field "attendee_emails" must be of type string, got array. Join the values yourself.
  • The comma is the separator, so a value that itself contains a comma will not match as expected. Use a field per value, or a different value format, if your values can contain commas.
  • Size the field for the whole joined string. max_length defaults to 1024 and its maximum is 65535, and it cannot be raised after the field is created, so declare it large enough up front.
  • equals compares the entire joined string, so it is rarely what you want on a multi-value field. Use contains.

3. Plan the metadata schema

Before your first ingest:
  • Plan scoping fields before first ingest: PATCH /databases/{database}/metadata-schema can add fields later, but existing fields cannot be renamed, retyped, reflagged, or deleted. Undeclared scope keys are not ignored at query time: they match no sources, because a database with a schema rejects undeclared keys on ingest and edit. The exception is connector_id and provider, which connectors write on every synced object. If you’ll scope on it more than once, declare it.
  • Pick metadata for hot paths, additional_metadata for cold ones: Top-level metadata filters are prefiltered; additional_metadata filters that cannot be prefiltered need a post-retrieval pass with over-fetch.
  • Keep keys stable: Renaming a metadata key requires re-ingesting affected sources.
  • Ingested metadata is editable in place: Once a source is indexed, PATCH /context/{id}/metadata merges new metadata and additional_metadata values into it without re-ingesting. See section 7.
  • Don’t substitute metadata_filters for collection: metadata_filters scopes results inside a partition. For partitioning by user, team, or workspace, use collection.

4. Minimal working example

Two phases: set up metadata (declare the schema, then attach values at ingest), then scope at query.

Step 1: Create metadata

The schema lives at the database level; values land on each source at ingest time. Both happen before any query.

Step 1a: Declare the schema at database creation

Schema field options

Limits and guardrails:
  • Up to 32 custom metadata fields.
  • Up to 6 embedding-enabled fields per database. enable_dense_embedding and enable_sparse_embedding each count as one, so a field with both set counts as two. Exceeding the limit fails database creation with 400 before anything is provisioned.
  • Field names are unique case-insensitively.
  • Dense/sparse embedding flags are only valid on VARCHAR fields.
  • Runtime metadata values must match the declared type when the database has a schema.
  • Unknown metadata keys are rejected on ingest and edit when the database has a non-empty schema.
  • Each metadata layer has a byte budget per request. See Size limits.

Size limits

Every request that attaches metadata is checked against two caps, on ingest and on metadata edit alike: Older spellings are still accepted, but not uniformly. Which one works depends on the endpoint: Accepted aliases have the same cap as the canonical field. The cap applies to the whole map, not to any one value, and it is measured on the map’s compact JSON encoding in UTF-8 bytes. Three consequences worth planning around:
  • Keys and punctuation count: Quotes, colons, commas and braces all count toward the byte limit.
  • Bytes, not characters: Accented Latin characters cost 2 bytes, most CJK characters 3, and emoji 4.
  • Budget in bytes from the start: A 950-character summary sounds comfortably under a 1 KiB cap, but with two small sibling keys it serializes to 1,015 bytes (65 bytes of that is structure alone). Push the summary to 1,000 characters and the request is rejected at 1,065 bytes.
Document metadata: 1,015 bytes, just inside the 1 KiB cap
Exceeding either cap fails the whole request with 400 before anything is ingested. The message names the offending field and reports both numbers, so you can see exactly how far over you are:
On PATCH /context/{id}/metadata the same message is prefixed with invalid metadata edit:.
If a document needs more than 1 KiB of descriptive metadata, put the long text in the document body where it gets chunked and embedded, and keep additional_metadata for the short values you actually filter on.

Filter size limits

The caps above bound the metadata you store. metadata_filters on /query has its own, separate pair. These bound what you send at query time and are unrelated to how much metadata a source carries: Measured the same way (compact JSON, UTF-8 bytes, field names and punctuation counted), and the object total includes the nested additional_metadata dict. The object total is measured after operator objects are reduced to their values, so {"contains": "x"} counts as ["x"] and the operator keyword itself costs nothing. Both spellings of the same filter cost the same, because the cap bounds the expression sent to the vector store, which the spelling does not change. Both exist because every value in a list is expanded into the filter expression sent to the vector store. The per-list cap catches one runaway list; the object cap catches many individually-legal lists adding up. Twenty lists of 500 values are each within the element cap but total roughly 127 KiB, so the object cap is what rejects them. Over either limit returns 400 before the query runs, naming the offending key or the actual byte count:
Needing far more than 500 values in one filter usually means the constraint belongs in the data rather than the query: add a metadata field that groups those values (a segment, tier, or cohort key) and filter on that instead.

Add schema fields later

You can add database metadata fields after database creation with PATCH /databases/{database}/metadata-schema. This is additive only:
  • add new fields: yes
  • delete fields: no
  • change type/flags of existing fields: no
Fields added this way cannot turn on dense or sparse search lanes: the endpoint rejects enable_dense_embedding and enable_sparse_embedding. Declare semantic or keyword metadata fields when you create the database.

5. Attach metadata at ingest

For knowledge ingestion, send metadata and additional_metadata on each document_metadata[] item or app_knowledge[] item.
For type=memory, the memories multipart field is already JSON-stringified, and each memory item’s metadata and additional_metadata are plain objects inside it. Do not stringify them a second time.

6. Query with metadata filters

Mix database-level (top-level) and document-level (nested) scopes in the same metadata_filters object:
Use the legacy alias only when maintaining older clients:
If both aliases are present and both are objects, additional_metadata wins on conflicts.

7. Update metadata without re-ingesting

Use PATCH /context/{id}/metadata when you know the source ID and need to update metadata in place.
Behavior:
  • The source must already exist.
  • collection is required.
  • At least one of database_metadata (deprecated alias tenant_metadata), additional_metadata, or acl is required.
  • The update is a merge/upsert: sent keys are inserted or overwritten; omitted keys are preserved.
  • document_metadata is not accepted on this endpoint; use additional_metadata.
  • The same endpoint accepts acl to change who may retrieve the source. Unlike metadata, acl replaces rather than merges, and an acl-only body is a valid edit. See Access Control.
  • Updates to database metadata fields without an embedding flag are MongoDB-only and take effect for filters and listing.
  • If an edited database metadata field has enable_dense_embedding or enable_sparse_embedding, HydraDB synchronously syncs the relevant vector store lane and reports vector_sync_required / vector_synced in the response (the milvus_sync_required / milvus_synced aliases are still emitted, deprecated).
For full document/content replacement, re-ingest with upsert: true and the same source id. Upsert replaces the source payload and metadata supplied by ingestion.

8. Listing with metadata filters

Use POST /context/list when you want to browse or page sources rather than run semantic retrieval:
/context/list also accepts legacy aliases tenant_metadata for metadata and document_metadata for additional_metadata.

9. Common mistakes


10. Advanced patterns

Stacked scopes with collection partitioning: Use collection for the partition (per-user, per-workspace), and use metadata_filters to scope inside that partition. They’re complementary, not interchangeable. See Multi-Tenant. Published vs draft: Add a status field to your schema; tag every source with metadata.status = "draft" | "published"; pass metadata_filters: { status: "published" } on user-facing queries. Keeps work-in-progress out of customer answers automatically. Multi-language corpora: Add a language field; route each query to the right language by passing metadata_filters: { language: detect_language(query) }. Review schema changes: Treat database_metadata_schema as a data contract. Existing field definitions are immutable, so changing them requires re-ingestion.