Full-Text Search in Your Database β BM25, No Separate Engine
Adding search usually means adding infrastructure: an external engine, an
indexing pipeline, and a sync job that drifts. In Aito v2, any Text column
is full-text searchable in place β analyzed, BM25-ranked, and composable with
ordinary database conditions in the same JSON query. The same index also
powers prediction from text, so search and inference share one copy of the
data.
The $search operator
A Lucene-subset query language: bare terms (AND), "quoted phrases", OR,
NOT / -term, and ( β¦ ) groups.
{
"from": "products",
"where": { "name": { "$search": "(milk OR bread) AND NOT chocolate" } },
"select": ["name"],
"limit": 20
}
{ "offset": 0, "total": 12, "hits": [
{ "name": "VAASAN Ruispalat 660g 12 pcs fullcorn rye bread" },
{ "name": "Fazer Puikula fullcorn rye bread 9 pcs/500g" },
{ "name": "Vaasan Ruispalat thin sliced rye bread 6pcs/195g" } ] }
It composes with plain conditions β full-text match and structured filter
in one where:
{
"from": "products",
"where": {
"name": { "$search": "milk OR juice" },
"category": "104"
},
"select": ["name", "category"]
}
Relevance-ranked search with highlighting
The search query type ranks by BM25 β the same scorer Lucene and
Elasticsearch use β and $highlight marks the matched tokens for your UI:
{
"from": "products",
"search": { "field": "name", "text": "milk" },
"select": ["name", "$score", "$highlight"],
"limit": 3
}
$why on a search result breaks the score down per token; $matches reports
the matched token positions.
From search to understanding
Because the analyzed text is a first-class column, the same tokens are prediction evidence β this is where an embedded search engine beats a bolted-on one:
{
"from": "products",
"where": { "name": { "$match": "semi-skimmed milk" } },
"predict": "category",
"select": ["$value", "$p", "$why"],
"limit": 1
}
One system answers "which rows match" (search), "which category is this" (classification), and "why" (explanation) β over the same live data, with no sync pipeline.
Why Aito for this
- Zero-infrastructure: per-language analyzers configured in the schema
(
"type": "Text", "analyzer": "English"); the index is built and merged with the data. - Exact, reproducible results β real BM25 scores, deterministic ties.
- Composability:
$searchis awhereproposition like any other β it works underpredict, in population restrictions, and alongside links.
Related: Query Reference Β· Vector Search for semantic (embedding-based) retrieval Β· Sandbox