Skip to main content

Text search expressions

similarity, isimilarity, and fulltext return numeric values using the constraint expression grammar. Compare these values with a threshold to select matches, combine conditions with all, any, and not, and use sort expressions to rank results. Return any number of scores through results.expressions.

Filtering does not choose an ordering or combine scores automatically. Specify an explicit sort expression for ranked retrieval. Ordinary arithmetic functions such as add and multiply can combine scores with application-defined weights. Output names are aliases for returned values; they are not stored properties and cannot be referenced by another expression.

Fuzzy matching​

{"isimilarity": [text, query]} compares two strings after Unicode full case folding. {"similarity": [text, query]} compares the original strings and is case-sensitive. Both functions take exactly two arguments and return normalized optimal string alignment (OSA) similarity: 1 - distance / max(codepoint_length(text), codepoint_length(query)). Insertion, deletion, substitution, and adjacent transposition each cost one. OSA is restricted Damerau–Levenshtein: a substring cannot be edited more than once. For example, CA versus ABC has distance 3. Whole values are compared, including spaces and punctuation.

For query cat, cat scores 1, cats scores 0.75, cut, bat, and act score approximately 0.666667, czz scores approximately 0.333333, and dog scores 0. aaa versus aaaa scores 0.75. Two empty strings score 1; exactly one empty string scores 0. Case-insensitive distance and lengths are measured after folding. Neither function normalizes Unicode forms or removes accents. Arguments can reference properties or use other value expressions, in either order. JSON scalar strings are unwrapped. Missing or null arguments return null; non-string arguments and invalid UTF-8 raise errors.

Filtering and ranking​

This query returns up to ten names with a case-insensitive similarity score of at least 0.2, ordered from highest to lowest score:

[{
"FindEntity": {
"with_class": "Person",
"constraints": [{"isimilarity": ["$name", "alen"]}, ">=", 0.2],
"sort": {
"expression": {"isimilarity": ["$name", "alen"]},
"order": "descending"
},
"limit": 10,
"results": {
"count": true,
"list": ["name"],
"expressions": {"nameScore": {"isimilarity": ["$name", "alen"]}}
}
}
}]

Optional similarity index​

Similarity functions work without an index. To accelerate positive-score filters on a stored string property, create a similarity index with a matching case mode. similarity uses case_sensitive (the default) or dual_case; isimilarity uses case_insensitive or dual_case.

[{
"CreateIndex": {
"index_type": "entity",
"class": "Person",
"property_key": "name",
"kind": "similarity",
"params": {"text": "dual_case"}
}
}]

Trigram indexes accelerate substring searches. Use a similarity index for fuzzy matching. Computed properties can be used in similarity expressions but cannot be indexed.

Full-text matching​

{"fulltext": ["$body", "image search"]} scores a direct property reference against a literal query string. The optional third argument is the literal "all" (default), requiring every distinct query term, or "any", requiring at least one. A present native string that does not match has score 0. A missing property returns null. Present values of other types raise an error, including JSON-wrapped strings, JSON null, and arrays. Computed properties are not supported.

Each evaluated entity or connection uses a ready fulltext index on its own class and the specified property. with_class is optional. When it is specified, a missing required index is reported even if that class has no candidates. Without with_class, every class reached by the function must have the required index, even if the particular record lacks that property. Ordinary constraint short-circuiting can avoid evaluating the function for unrelated records.

Find articles containing both words and return their relevance scores:

[{
"CreateIndex": {
"index_type": "entity",
"class": "Article",
"property_key": "body",
"kind": "fulltext"
}
}, {
"FindEntity": {
"with_class": "Article",
"constraints": [{"fulltext": ["$body", "image search"]}, ">", 0],
"sort": {
"expression": {"fulltext": ["$body", "image search"]},
"order": "descending"
},
"results": {
"list": ["title"],
"expressions": {"score": {"fulltext": ["$body", "image search"]}}
}
}
}]

Create the index once, then use FindEntity for subsequent searches.

Text analysis and scores​

The unicode_v1 analyzer validates UTF-8, applies Unicode NFKC case folding, and splits words using Unicode word boundaries with the root locale. It treats compatibility forms, case differences, and equivalent Unicode spellings consistently while preserving accents. Matching uses complete terms, not arbitrary substrings. The analyzer does not stem words, remove stop words, or expand synonyms. Repeated query terms are considered once. Queries with no retained terms score 0 for string values.

Scores use BM25 with k1 = 1.2 and b = 0.75. Term frequency and document length come from the selected property. Corpus statistics include all values with at least one indexed term in that class/property index, independently of constraint or relationship filtering. Matching scores are positive and have no fixed upper bound; they can change when other indexed values change. Scores from different properties or classes use different corpora; any combination is an explicit application choice.

Terms exceeding 256 bytes after normalization are omitted from documents and queries. An all query requires only the remaining terms. Strings with no remaining terms do not contribute to corpus counts or average document length.

Combining full-text and fuzzy matching​

This example creates a full-text index and finds published articles whose body contains both query words or whose title meets a similarity threshold. It ranks matches using the body score plus twice the title score and returns both scores:

[{
"CreateIndex": {
"index_type": "entity",
"class": "Article",
"property_key": "body",
"kind": "fulltext"
}
}, {
"FindEntity": {
"with_class": "Article",
"constraints": ["all",
["$published", "==", true],
["any",
[{"fulltext": ["$body", "image search", "all"]}, ">", 0],
[{"isimilarity": ["$title", "image search"]}, ">=", 0.5]
],
["not", ["$status", "==", "archived"]]
],
"sort": {
"expression": {"add": [
{"coalesce": [{"fulltext": ["$body", "image search"]}, 0]},
{"multiply": [2,
{"coalesce": [{"isimilarity": ["$title", "image search"]}, 0]}
]}
]},
"order": "descending"
},
"limit": 20,
"results": {
"list": ["title"],
"expressions": {
"bodyScore": {"fulltext": ["$body", "image search"]},
"titleScore": {"isimilarity": ["$title", "image search"]}
}
}
}
}]

Use coalesce when a missing score should contribute zero to arithmetic. Index creation, property changes, and deletions participate in the transaction. Functions evaluated later in the transaction observe those changes. Removing a required index makes subsequent full-text evaluation fail until it is recreated.

Projection and pagination​

results.expressions maps output names to value expressions and places their values beside the ordinary properties in each returned entity or connection. A missing expression value appears explicitly as JSON null. Names must follow ordinary property naming rules, be distinct ignoring case, and not overlap results.list, including wildcard selections. Expression projection cannot accompany all_properties or results.group. Expression sorting is also unsupported for grouped aggregates.

Constraints select the candidate set before sorting, offset, and limit. results.count counts matches before pagination. Sorting breaks complete ties by ascending object ID. With group_by_source and per_group, the specified sort and pagination apply independently to each source group.

Resource limits​

Text inputs must be valid UTF-8. Size limits count UTF-8 bytes; KiB and MiB use powers of 1024.

InputLimit
Each similarity or isimilarity operand16 MiB, both before and after case folding.
Full-text query64 KiB and at most 1024 distinct indexed terms.
Text analyzed for full-text matching16 MiB, both before and after normalization.
Individual full-text termTerms longer than 256 bytes after normalization are omitted.

Text search input limits

Searches are also subject to memory and computation limits. A query that exceeds these limits fails with an error; it does not return truncated matches. Long strings and broad searches can exceed these limits even with a small limit, which controls the number of returned records. coalesce substitutes for null values and does not catch errors. The constraint resource limits also apply.