Text search expressions
similarity, isimilarity, and fulltext return numeric values using
the constraint expression grammar.
Compare these values with a threshold to select matches, combine
conditions with all, any, and not, and use sort
expressions to rank results. Return any number of scores through
results.expressions.
Filtering does not choose an ordering or combine scores automatically.
Specify an explicit sort expression for ranked retrieval. Ordinary
arithmetic functions such as add and multiply can combine scores
with application-defined weights. Output names are aliases for returned
values; they are not stored properties and cannot be referenced by
another expression.
Fuzzy matching
{"isimilarity": [text, query]} compares two strings after Unicode full
case folding. {"similarity": [text, query]} compares the original
strings and is case-sensitive. Both functions take exactly two arguments
and return normalized optimal string alignment (OSA) similarity:
1 - distance / max(codepoint_length(text), codepoint_length(query)).
Insertion, deletion, substitution, and adjacent transposition each cost
one. OSA is restricted Damerau–Levenshtein: a substring cannot be edited
more than once. For example, CA versus ABC has distance 3. Whole
values are compared, including spaces and punctuation.
For query cat, cat scores 1, cats scores 0.75, cut, bat, and
act score approximately 0.666667, czz scores approximately 0.333333,
and dog scores 0. aaa versus aaaa scores 0.75. Two empty strings
score 1; exactly one empty string scores 0. Case-insensitive distance
and lengths are measured after folding. Neither function normalizes
Unicode forms or removes accents. Arguments can reference properties or
use other value expressions, in either order. JSON scalar strings are
unwrapped. Missing or null arguments return null; non-string arguments
and invalid UTF-8 raise errors.
Filtering and ranking
This query returns up to ten names with a case-insensitive similarity score of at least 0.2, ordered from highest to lowest score:
[{
"FindEntity": {
"with_class": "Person",
"constraints": [{"isimilarity": ["$name", "alen"]}, ">=", 0.2],
"sort": {
"expression": {"isimilarity": ["$name", "alen"]},
"order": "descending"
},
"limit": 10,
"results": {
"count": true,
"list": ["name"],
"expressions": {"nameScore": {"isimilarity": ["$name", "alen"]}}
}
}
}]
Optional similarity index
Similarity functions work without an index. To accelerate positive-score
filters on a stored string property, create a similarity index with a
matching case mode. similarity uses case_sensitive (the default) or
dual_case; isimilarity uses case_insensitive or dual_case.
[{
"CreateIndex": {
"index_type": "entity",
"class": "Person",
"property_key": "name",
"kind": "similarity",
"params": {"text": "dual_case"}
}
}]
Trigram indexes accelerate substring searches. Use a similarity index for fuzzy matching. Computed properties can be used in similarity expressions but cannot be indexed.
Full-text matching
{"fulltext": ["$body", "image search"]} scores a direct property
reference against a literal query string. The optional third argument is
the literal "all" (default), requiring every distinct query term, or
"any", requiring at least one. A present native string that does not
match has score 0. A missing property returns null. Present values of
other types raise an error, including JSON-wrapped strings, JSON null,
and arrays. Computed properties are not supported.
Each evaluated entity or connection uses a ready fulltext index on its
own class and the specified property. with_class is optional. When
it is specified, a missing required index is reported even if that class
has no candidates. Without with_class, every class reached by the
function must have the required index, even if the particular record
lacks that property. Ordinary constraint short-circuiting can avoid
evaluating the function for unrelated records.
Find articles containing both words and return their relevance scores:
[{
"CreateIndex": {
"index_type": "entity",
"class": "Article",
"property_key": "body",
"kind": "fulltext"
}
}, {
"FindEntity": {
"with_class": "Article",
"constraints": [{"fulltext": ["$body", "image search"]}, ">", 0],
"sort": {
"expression": {"fulltext": ["$body", "image search"]},
"order": "descending"
},
"results": {
"list": ["title"],
"expressions": {"score": {"fulltext": ["$body", "image search"]}}
}
}
}]
Create the index once, then use FindEntity for subsequent searches.
Text analysis and scores
The unicode_v1 analyzer validates UTF-8, applies Unicode NFKC case
folding, and splits words using Unicode word boundaries with the root
locale. It treats compatibility forms, case differences, and equivalent
Unicode spellings consistently while preserving accents. Matching uses
complete terms, not arbitrary substrings. The analyzer does not stem
words, remove stop words, or expand synonyms. Repeated query terms are
considered once. Queries with no retained terms score 0 for string
values.
Scores use BM25 with k1 = 1.2 and b = 0.75. Term frequency and
document length come from the selected property. Corpus statistics
include all values with at least one indexed term in that class/property
index, independently of constraint or relationship filtering. Matching
scores are positive and have no fixed upper bound; they can change when
other indexed values change. Scores from different properties or classes
use different corpora; any combination is an explicit application
choice.
Terms exceeding 256 bytes after normalization are omitted from documents
and queries. An all query requires only the remaining terms. Strings
with no remaining terms do not contribute to corpus counts or average
document length.
Combining full-text and fuzzy matching
This example creates a full-text index and finds published articles whose body contains both query words or whose title meets a similarity threshold. It ranks matches using the body score plus twice the title score and returns both scores:
[{
"CreateIndex": {
"index_type": "entity",
"class": "Article",
"property_key": "body",
"kind": "fulltext"
}
}, {
"FindEntity": {
"with_class": "Article",
"constraints": ["all",
["$published", "==", true],
["any",
[{"fulltext": ["$body", "image search", "all"]}, ">", 0],
[{"isimilarity": ["$title", "image search"]}, ">=", 0.5]
],
["not", ["$status", "==", "archived"]]
],
"sort": {
"expression": {"add": [
{"coalesce": [{"fulltext": ["$body", "image search"]}, 0]},
{"multiply": [2,
{"coalesce": [{"isimilarity": ["$title", "image search"]}, 0]}
]}
]},
"order": "descending"
},
"limit": 20,
"results": {
"list": ["title"],
"expressions": {
"bodyScore": {"fulltext": ["$body", "image search"]},
"titleScore": {"isimilarity": ["$title", "image search"]}
}
}
}
}]
Use coalesce when a missing score should contribute zero to
arithmetic. Index creation, property changes, and deletions participate
in the transaction. Functions evaluated later in the transaction observe
those changes. Removing a required index makes subsequent full-text
evaluation fail until it is recreated.
Projection and pagination
results.expressions maps output names to value expressions and places their values beside the ordinary properties in each returned entity or connection. A missing expression value appears explicitly as JSON null. Names must follow ordinary property naming rules, be distinct ignoring case, and not overlap results.list, including wildcard selections. Expression projection cannot accompany all_properties or results.group. Expression sorting is also unsupported for grouped aggregates.
Constraints select the candidate set before sorting, offset, and limit. results.count counts matches before pagination. Sorting breaks complete ties by ascending object ID. With group_by_source and per_group, the specified sort and pagination apply independently to each source group.
Resource limits
Text inputs must be valid UTF-8. Size limits count UTF-8 bytes; KiB and MiB use powers of 1024.
| Input | Limit |
|---|---|
Each similarity or isimilarity operand | 16 MiB, both before and after case folding. |
| Full-text query | 64 KiB and at most 1024 distinct indexed terms. |
| Text analyzed for full-text matching | 16 MiB, both before and after normalization. |
| Individual full-text term | Terms longer than 256 bytes after normalization are omitted. |
Text search input limits
Searches are also subject to memory and computation limits. A query that
exceeds these limits fails with an error; it does not return truncated
matches. Long strings and broad searches can exceed these limits even
with a small limit, which controls the number of returned records.
coalesce substitutes for null values and does not catch errors. The
constraint resource limits also apply.