Skip to main content
TopK provides a data frame-like syntax for querying documents. It features built-in semantic search, text search, vector search, and metadata filtering capabilities. With TopK’s declarative query builder, you can easily select fields, chain filters, and apply vector/text search in a composable manner.

Query structure

In TopK, a query consists of multiple stages:
1

Select stage

Select static or computed fields that will be returned in the query results. These fields can be used in stages such as Filter or TopK.
2

Filter stage

Filter the documents that will be returned in the query results. Filters can be applied to static fields, computed fields such as vector_distance() or semantic_similarity(), or custom properties computed inside select().
3

Sort stage

Order results by an expression (ascending or descending).
4

Limit stage

Return at most k results.

Count stage

Return the total number of documents matching the query.
All queries must have either Sort + Limit or Count collection stage.
You can stack multiple select and filter stages in a single query.
A typical query in TopK looks as follows:

Select

The select() function is used to initialize the select stage of a query. It accepts a key-value pair of field names and field expressions:

Select expressions

Use a field() function to select fields from a document. In the select stage, you can also rename existing fields or define computed fields using function expressions.

Function expressions

Function expressions are used to define computed fields that will be included in your query results. TopK currently supports four main function expressions:
  • vector_distance(field, vector): Computes distance between vectors for vector search. This function is available for all dense and sparse vector types.
  • bm25_score(): Calculates relevance scores using the BM25 algorithm for keyword search
  • semantic_similarity(field, query): Measures semantic similarity between the provided text query and the field’s embedding
  • multi_vector_distance(field, matrix): Computes MaxSim distance for multi-vector (matrix) fields. Requires a multi_vector_index() on the field. See multi-vector search.

Vector distance

The vector_distance() function is used to compute the vector score between a query vector and a vector field in a collection. There are multiple ways to represent a query vector:
  • Dense vectors:
    • [0.1, 0.2, 0.3, ...] - Array of numbers resolved as a dense float32 vector
    • f32_vector([...]) - Helper function returning a dense float32 vector
    • f16_vector([...]) - Helper function returning a dense float16 vector
    • f8_vector([...]) - Helper function returning a dense float8 vector
    • u8_vector([...]) - Helper function returning a dense u8 vector
    • i8_vector([...]) - Helper function returning a dense i8 vector
    • binary_vector([...]) - Helper function returning a binary vector
  • Sparse vectors:
    • { 0: 0.1, 1: 0.2, 2: 0.3, ... } - Mapping from index → value resolved as a sparse float32 vector
    • f32_sparse_vector({ ... }) - Helper function returning a sparse float32 vector
    • u8_sparse_vector({ ... }) - Helper function returning a sparse u8 vector
Optionally, uses can provide skip_refine=True to bypass the internal distance refinement step. This will improve performance for queries with large top_k at the cost of lower accuracy.
We don’t recommend using skip_refine=True unless you’re using a large top_k.
To use the vector_distance() function, you must have a vector index defined on the field you’re computing the vector distance against:

BM25 Score

The BM25 score is a relevance score that can be used to score documents based on their text content. To use the fn.bm25_score() in your query, you must include a match predicate in your filter stage.
To use the fn.bm25_score() function, you must have a keyword index defined in your collection schema.

Semantic similarity

The semantic_similarity() function is used to compute the similarity between a text query and a text field in a collection. To use the semantic_similarity() function, you must have a semantic index defined on the field you’re computing the similarity on.

Multi-vector distance

The multi_vector_distance() function computes the MaxSim score between a query matrix and a matrix field in a collection. Use it for multi-vector (late-interaction) retrieval when documents are stored as N x D embedding matrices. To use multi_vector_distance(), you must have a multi_vector_index() defined on the field. The query matrix can be a list of lists (defaults to f32), a numpy array (type inferred from dtype), or a matrix() instance. The optional candidates parameter limits the number of candidate vectors considered during search.
See multi-vector search for schema setup and ingestion details.

Advanced select expressions

TopK doesn’t only let you select static fields from your documents or computed fields using function expressions. You can also use TopK powerful expression language to select fields by chaining arbitrary logical expressions:

Filtering

You can filter documents by metadata, keywords, custom properties computed inside select() (e.g. vector similarity or BM25 score) and more. Filter expressions support all

Metadata filtering

The match() function is the backbone of keyword search in TopK. It allows you to search for documents that contain specific keywords or phrases. You can configure the match() function to:
  • Match on multiple terms
  • Match only on specific fields
  • Use weights to prioritize certain terms
The match() function accepts the following parameters:
string
required
String token to match. Can also contain multiple terms separated by a delimiter which is any non-alphanumeric character.
string
Field to match on. If not provided, the function will match on all fields.
number
Weight to use for matching. If not provided, the function will use the default weight(1.0).
boolean
Use all parameter when a text must contain all terms(separated by a delimeter)
  • when all is false (default) it’s an equivalent of OR operator
  • when all is true it’s an equivalent of AND operator
Searching for a term like "catcher" in your documents is as simple as using the match() function in the filter stage of your query:

Match multiple terms

The match() function can be configured to match all terms when using a delimiter. A term delimiter is any non-alphanumeric character. To ensure that all terms are matched, use the all parameter:

Give weight to specific terms

You can give weight to specific terms by using the weight parameter:

Boost ranking with optional terms

The should() function adds an optional BM25 scoring term without filtering documents from the result set. Documents containing the term receive a higher BM25 score, while documents that do not contain it remain eligible for the results. Use should() together with match() when some terms are required and others should only influence ranking. When used on its own, should() matches the entire collection and ranks documents according to how well they match the term.
This returns only documents matching hobbit or rings, while boosting documents that also match lord. The should() function accepts the following parameters:
string
required
String token to score against.
string
Keyword-indexed field used for scoring. Searches all eligible fields when omitted.
number
Multiplier applied to the term’s BM25 contribution. Defaults to 1.0.

Combine keyword search and metadata filtering

You can combine metadata filtering and keyword search in a single query by stacking multiple filter stages. In the example below, we’re searching for documents that contain the keyword "catcher" and were published in 1997, or between 1920 and 1980.

Operators

When writing queries, you can use the following operators for:
  • field selection
  • filtering
  • topk collection

Logical operators

Logical operators combine multiple expressions by applying boolean logic and conditions.

and

The and operator can be used to combine multiple logical expressions.

or

The or operator can be used to combine multiple logical expressions.

not

The not helper can be used to negate a logical expression. It takes an expression as an argument and inverts its logic.

all

The all() helper evaluates to true if each expression in the array is true. It’s equivalent to applying the logical AND operator across all expressions.
This is equivalent to:

any

The any() helper evaluates to true if at least one expression in the array is true. It’s equivalent to applying the logical OR operator across all expressions.
This is equivalent to:

choose

The choose operator evaluates a condition and returns the first argument if the condition is true, else the second argument.

boost

The boost operator multiplies the scoring expression by the provided boost value if the condition is true. Otherwise, the scoring expression is unchanged (multiplied by 1).

coalesce

The coalesce operator replaces null values with a provided value.

Comparison operators

Comparison operators provide various logical, numerical and string functions that evaluate to true or false.

eq

The eq operator can be used to match documents that have a field with a specific value.

ne

The ne operator can be used to match documents that have a field with a value that is not equal to a specific value.

is_null

The is_null operator can be used to match documents that have a field with a value that is null.

is_not_null

The is_not_null operator can be used to match documents that have a field with a value that is not null.

gt

The gt operator can be used to match documents that have a field with a value greater than a specific value. For strings, it uses lexicographic order.

gte

The gte operator can be used to match documents that have a field with a value greater than or equal to a specific value. For strings, it uses lexicographic order.

lt

The lt operator can be used to match documents that have a field with a value less than a specific value. For strings, it uses lexicographic order.

lte

The lte operator can be used to match documents that have a field with a value less than or equal to a specific value. For strings, it uses lexicographic order.

starts_with

The starts_with operator can be used on string fields to match documents that start with a given prefix. This is especially useful in multi-tenant applications where document IDs can be structured as {tenant_id}/{document_id} and starts_with can then be used to scope the query to a specific tenant. Also supports list-of-string fields for prefix filtering on array elements (e.g. field("tags").starts_with("fiction")).

contains

The contains operator can be used on both text fields and list fields to match documents that include a specific value. For text fields, it matches documents that include a specific substring (case-sensitive). For list fields, it matches documents where the field of type list contains the specified value.
  • Text fields: Matches documents that include a specific substring. It is case-sensitive and avoids the text processing pipeline (tokenization and stemming) used by the match() function. This makes it particularly useful when you need exact substring matching or want to provide your own pre-processed tokens. Unlike match(), the contains operator can be used without requiring a keyword index.
  • List fields: Matches documents where the list field contains the specified value. The value can be a literal or a field reference. You can also use a list of strings with a keyword index if you want to provide your own tokens instead of using the text processing pipeline.
The contains operator works exactly the same as the in operator, but with reversed operands: x CONTAINS y is equivalent to y IN x. Both operators are provided for convenience and to make queries more readable.

in

The in (or in_ in Python) operator checks if a field value is present in a list of values, string literal or another field. It can be used in several ways:
  • Field in list: Checks if a field value is present in a list of literal values.
  • Field in string: Checks if a string field is a substring of another string. Unlike the match(), this avoids the text processing pipeline (tokenization and stemming) and performs exact substring matching.
  • Field in field: Checks if a field value is present in another field.
The in operator works exactly the same as the contains operator, but with reversed operands: y IN x is equivalent to x CONTAINS y. Both operators are provided for convenience and to make queries more readable.

match_all

The match_all operator returns true if all terms in the query are present in the field with a keyword index.
When using a match_all operator against a text field, it must be used in conjunction with a keyword index defined in your collection schema.

match_any

The match_any operator returns true if any term in the query is present in the field with a keyword index.
When using a match_any operator against a text field, it must be used in conjunction with a keyword index defined in your collection schema.

regexp_match

The regexp_match operator returns true if the field value matches the regular expression. Internally, this uses Rust’s regex crate to evaluate the regular expression.

Mathematical operators

Mathematical operators perform computations on numbers.

add

The add operator can be used to add two numbers.

sub

The sub operator can be used to subtract two numbers.

mul

The mul operator can be used to multiply two numbers.

div

The div operator can be used to divide two numbers.

abs

The abs operator returns the absolute value of a number, which is useful for calculating distances or differences.

min

The min operator returns the smaller of two values, commonly used for clamping or setting upper bounds. It can work with both scalar values and other fields or expressions. For strings, it uses lexicographic order.

max

The max operator returns the larger of two values, commonly used for clamping or setting lower bounds. It can work with both scalar values and other fields or expressions. For strings, it uses lexicographic order.

ln

The ln operator calculates the natural logarithm, useful for logarithmic scaling and dampening large values.

exp

The exp operator calculates the exponential function (e^x), useful for exponential scaling and boosting.

sqrt

The sqrt operator calculates the square root, useful for dampening values and creating non-linear transformations.

square

The square operator multiplies a number by itself (x²), useful for amplifying differences and creating quadratic transformations.

Collection

All queries must have a collection stage. Currently, we support topk(), count() and group_by() collectors.

topk

Use the topk() function to return the top k results. The topk() function accepts the following parameters:
LogicalExpression
required
The logical expression to sort the results by.
number
required
The number of results to return.
boolean
required
Whether to sort the results in ascending order.
To get the top 10 results with the highest title_similarity, you can use the following query:
The topk() stage is equivalent to applying sort(expr, asc) followed by limit(k). It is a convenience shorthand for the common pattern of ordering results by a scoring expression and returning the top k.

limit

Use the .limit(k) stage to return at most k results. Can be used with or without topk().

offset

Use the .offset(n) stage to skip the first n results.

sort

Use the .sort(expr, asc) stage to sort results by an expression. Use asc=True for ascending order, asc=False for descending order.
To sort by more than one key, pass a list of sort expressions instead of a single expression. They are applied in order, each breaking ties left by the previous ones:

count

Use the count() function to get the total number of documents matching the query. If there are no filters then count() will return the total number of documents in the collection.

group_by

Use the group_by(keys, aggs) stage to group documents by one or more key expressions and compute aggregations for each group. The query returns one row per group, containing the group keys and the aggregated values.
The following aggregate functions are supported, exposed via the agg module in the SDKs (topk_sdk.query.agg in Python, agg from topk-js/query in JavaScript) and as standard COUNT/SUM/MIN/MAX/AVG functions in SQL: Multiple aggregations can be computed in a single group_by():
The group_by() stage can be combined with preceding filter() and select() stages — group keys and aggregations can reference computed columns projected by an earlier select():
When writing queries, remember that they all require the topk, count or group_by function at the end.

Query options

The query() method accepts a consistency parameter to control read consistency: indexed, balanced (default), or strong. See the consistency concept for details.

LSN-based Consistency

TopK supports LSN (Log Sequence Number) based consistency for ensuring read-after-write consistency. When you perform a write operation (like upsert), you receive an LSN as a string that represents the sequence number of that write in the system’s log. You can use this LSN in subsequent queries to ensure that the query only returns results that are at least as recent as that write operation.

How it works

  1. Write operation: When you call lsn = client.collection().upsert(), you receive an LSN
  2. Query with LSN: Pass that LSN to client.collection().query(..., lsn=lsn)
  3. Consistency guarantee: If the write is not yet available in the read path, the query will be rejected and the client will automatically retry
This approach ensures that your queries always see the results of your recent writes, providing strong consistency guarantees when needed.
Using LSN-based consistency may increase query latency as the system needs to verify that the specified LSN has been processed before returning results.