Inference (v2)
This guide explains the inference operations on the v2 (rep2) engine: predict,
recommend, relate, ranking by context, and explainable results with
$why. The statistical model is the same family as
v1 inference, exposed through the simpler Query2 syntax.
The idea behind Aito is that the same from/where you'd use to filter data is
also the evidence for a prediction. You don't train a model per question β you
ask the database.
Predict
Point predict at a field and Aito ranks its values by probability given your
where evidence:
{
"from": "products",
"where": { "name": { "$match": "rye bread" } },
"predict": "category",
"select": ["$p", "$value"]
}
Each hit has a $value (a candidate category) and $p (its probability). Results
are ordered most-probable first.
Candidates, evidence, and $f
where is evidence, not a filter on the answer. A predict ranks the
target's whole value set, and the evidence decides each value's probability.
Which values are candidates depends on the target:
| target column | candidates |
|---|---|
a plain column (PaymentMethod, category) | every value that occurs in the column |
a link (GLCode β glCodes) | every row of the linked table β including a row no other row points to yet |
So a predict with evidence still returns values that never occurred with that
evidence. Select $f to see which ones did: $f is the number of rows that
match the where and carry the value.
The one exception: a where on the predicted field itself, or on a path under
it. {"GLCode": {"$or": [...]}} or {"GLCode.Department": "IT & Infrastructure"} while
predicting GLCode is not evidence β it is a statement about the answer, and it
restricts the candidate set to the values it admits. That is what
scoping below is built on.
{
"from": "invoices",
"where": { "Processor": "Emily Davis" },
"predict": "GLCode",
"select": ["$value", "$p", "$f"],
"limit": 4
}
{ "offset": 0, "total": 10, "hits": [
{ "$f": 20, "$p": 0.847, "$value": "F001" },
{ "$f": 0, "$p": 0.042, "$value": "R001" },
{ "$f": 0, "$p": 0.032, "$value": "E001" },
{ "$f": 0, "$p": 0.027, "$value": "E002" } ] }
Emily Davis's 20 invoices are all coded F001, so F001 has
$f: 20. The other accounts have $f: 0 β never seen with
this evidence β and still get a small probability rather than zero:
- Why not zero? Every estimate is smoothed: an account's probability given
the evidence is pulled toward how common that account is overall, and the
pull weakens as the evidence gets more rows of its own. An account that is
frequent elsewhere (
R001) keeps more of the remaining probability than a rare one. Nothing is ever certain on finite data, and a$pof exactly 0 would claim that it is. - Why is that useful? A GL account added to
glCodestoday is a candidate immediately, at a small$p, and rises as invoices are coded to it β there is no cold-start gap to engineer around.
To return only the values seen with the evidence, add having:
{
"from": "invoices",
"where": { "Processor": "Emily Davis" },
"predict": "GLCode",
"select": ["$value", "$p", "$f"],
"having": { "$f": { "$gte": 1 } }
}
{ "offset": 0, "total": 1, "hits": [
{ "$f": 20, "$p": 0.847, "$value": "F001" } ] }
having filters the ranked list after scoring: the remaining $p values are
not renormalised, and $f is the only field it accepts.
Scoping candidates per customer
If one Aito database serves several customers (tenants), decide how each
predict's candidates are scoped. A where term that is evidence will not do
it, and neither will a
nested from when the
target is a link: the restriction scopes base rates and counts, but a link
target's candidates stay every row of the linked table.
What works today:
-
A
whereon the target's own path β the linked table carries the scoping column, and the query names it under the predicted field:{ "from": "invoices", "where": { "Description": "monthly cloud hosting", "Processor.Department": "Finance" }, "predict": "Processor", "select": ["$value", "$p"] }Processor.Departmentsits under the predicted field, so it scopes the candidate set rather than acting as evidence: on the demo this ranks the two employees in Finance, where the same predict without it ranks all ten. The rest of thewhereis ordinary evidence. In a multi-tenant schema the scoping column is the tenant β{"GLCode.tenant": "acme"}while predictingGLCoderanks only that tenant's accounts, including ones it has never used yet.Equality,
$in, ranges and other operators, and$and/$or/$notall scope the same way, as long as every leaf of a combinator sits under the link path β a clause satisfiable through a term outside it ({"$or": [{"GLCode.tenant": "acme"}, {"TotalAmount": 10}]}) constrains nothing, because a row can satisfy it without the linked condition holding.This is the option to reach for first: it is one query, and it needs no per-tenant schema. It scopes the candidates only: the probabilities are still learned from every tenant's rows, so another tenant's bookings move this tenant's ranking. To keep them out as well, put the tenant into a nested
frombeside the link filter β see Multi-tenant isolation. Neither is access control, so both belong beside a server-side check that the caller may see that tenant, not instead of one.Available since: v2.8.3 for the operator and combinator forms, and v2.9.0 for all of them again β v2.8.4 regressed link-target scoping by an attribute of the target and ignored it.
-
A per-tenant linked table β the target links to a table that holds only that tenant's values (for example one
glCodes_<tenant>table per tenant, or one database per tenant). The candidates are scoped by construction, and a new value is a candidate as soon as its row exists. This is the design to choose when tenants must never see each other's values. The statistics still come from the rows the query sees, so on a shared invoice table pair it with a nestedfromas well. -
Two steps over one shared table, when the tenant is not a column on the linked table and only the co-occurrence history distinguishes the values.
- Ask for the tenant's own values: a predict whose
whereholds only the tenant, with"having": { "$f": { "$gte": 1 } }. - Run the real predict β the invoice's evidence in
where(and the tenant as a nestedfrom, so the base rates are the tenant's) β and keep only the hits whose$valueis in the set from step 1.
Do not put the invoice's evidence into step 1:
$fcounts rows that match the wholewhere, so with real evidence beside the tenant it is 0 for most values andhavingreturns nothing. - Ask for the tenant's own values: a predict whose
having shapes what a query returns; it is not an access control.
Across linked tables
Evidence can come from a linked table using a dotted path, and basedOn projects
linked-table fields into the prediction context as extra evidence:
{
"from": "invoices",
"where": { "Description": "monthly cloud hosting", "Processor.Department": "IT & Infrastructure" },
"basedOn": ["Department"],
"predict": "GLCode",
"select": ["$p", "$value"]
}
Cross-table predict / recommend / relate run end-to-end on v2, and cross-table priors (the dominant accuracy contributor) are on by default.
Evidence through the target's attributes. basedOn names columns of the
linked table, and those attributes then carry evidence of their own: predicting
GLCode with basedOn: ["Department"] lets a department the invoice's evidence
implies lift every account belonging to it, including accounts this evidence has
never co-occurred with. Left out, the default is deliberately narrow β only a
linked column that a where term also names is projected β so the fields a
prediction is allowed to draw on stay the ones the caller declared.
Scoping the candidate set. Apart from the target's own path
(above), where supplies evidence, not an output
domain β a predict ranks the field's global value set, with values never seen
under your where in the tail at $f: 0. For a link target the value set
is the linked table (every target row, matching the v1 engine), so an entity
that exists but was never referenced in training is still a ranked candidate β
scored by its cross-table priors and name evidence rather than co-occurrence.
That is what lets a brand-new employee whose name appears in an invoice's text
be predicted as its handler. When the tail must not show them at
all (the multi-tenant case: a tenant-scoped predict must not list other
tenants' values as low-probability alternatives), add the candidate filter:
{
"from": "invoices",
"where": { "ReceiverName": "Acme Grocery Store" },
"predict": "Processor",
"having": { "$f": { "$gte": 1 } }
}
having keeps only candidates that actually co-occur with the where
($f β₯ 1), applied after ranking and before offset/limit, with no
renormalisation. On the demo that is the difference between ranking all ten
employees and ranking the four who have actually processed this receiver's
invoices. Only $f with one comparison operator is accepted β anything
else is a loud 400.
Cost follows the candidate count, not just the row count. A predict scores
every candidate against the evidence, so a link target over a large table is the
expensive shape β not because the table has many rows, but because each row is
a value to rank. Scoping with a where on the target's own path is therefore a
latency lever as much as a correctness one: it cuts the pool before scoring,
where an unscoped prediction pays for every row of the linked table on every
request. On a large linked table that difference dominates the response time. having does not help here β it filters after ranking, so
everything is still scored. Two rules of thumb follow: scope by construction
where you can, and treat βhow many candidates does this query rank?β as the
first question when a prediction is slower than you expect.
Available since: v2.8.0 for having, and for the
linked-table candidate universe on the v2 engine (before it, v2 ranked only
link values seen in training). You can also
predict the link itself (e.g. predict processor and let the engine draw on
the linked employee's attributes), and match into array/multi-value columns
with $has (membership), so a tag list or a basket of products is first-class
evidence.
Recommend
recommend is predict paired with a goal β rank candidates by how well they
achieve a desired outcome, not just by similarity:
{
"from": "impressions",
"where": { "context.user": "larry" },
"recommend": "product",
"goal": { "purchase": true },
"limit": 10
}
This ranks products by P(purchase | context.user=larry, product), i.e. what Larry is most likely to actually buy β not merely what's popular.
Rank by context
You can rank candidates by the probability of a condition evaluated against the
context (from) table, using $context inside orderBy:
{
"from": "impressions",
"get": "product",
"orderBy": { "$p": { "$context": { "purchase": true } } },
"select": ["$value", "$p"]
}
"Rank products by P(purchase | product)". It reuses the recommend machinery β the context condition becomes the goal β so there's no separate scoring path.
Relate
relate returns statistical relations instead of rows β which features move a
target, with lift and information-gain style statistics:
{
"from": "impressions",
"where": { "purchase": true },
"relate": ["product.category"]
}
Use it to discover what's associated with an outcome.
Each hit reports the relation from several angles: lift and info (the
smoothed strength), fs/ps (exact counts and within-population rates), and
relation (the raw 2Γ2 contingency the numbers derive from).
Date columns relate as ranges. A Date field's values are grouped into
calendar buckets sized by the field's span β days under ~3 months, months up
to ~4 years, years beyond β so a two-year order history relates as monthly
ranges instead of one relation per distinct day. Each bucket's condition is
the re-queryable range form, e.g.
{"$and":[{"order_date":{"$gte":"2024-03-01"}},{"order_date":{"$lt":"2024-04-01"}}]},
which you can paste straight back into a where.
Available since: v2.8.0 for the relation block and
the date-range grouping.
Explainable results β $why
Predictions you can't explain are hard to trust. Add $why to select and every
ranked candidate comes back with the factor tree that produced its score. For a
collection where color predicts label (redβa, blueβb):
{
"from": "obs",
"where": { "color": "red" },
"predict": "label",
"select": ["$p", "$value", "$why"]
}
returns, for the top candidate:
{
"$p": 0.8,
"$value": "a",
"$why": {
"type": "product",
"factors": [
{ "type": "baseP", "value": 0.5, "proposition": { "label": { "$has": "a" } } },
{ "type": "relatedPropositionLift", "proposition": { "color": "red" }, "value": 1.6 }
]
}
}
Read it back: the base rate for label=a is 0.5, the evidence color=red lifts
it by 1.6Γ, and 0.5 Γ 1.6 β the predicted 0.8. The explanation isn't a
post-hoc approximation β it's the actual arithmetic the engine did.
The example above is trimmed to the two factors that carry the story. A real tree
also contains structural normalizer nodes and a calibration node, so
the exact $p is a little below the 0.5 Γ 1.6 you'd get from just those two β
every node the engine multiplied is shown, which is what makes the number
auditable rather than approximate.
The factor types
Each node in the tree has a type. A product node's value is the product of its
children; the leaves are:
type | Meaning | Scale |
|---|---|---|
baseP | The candidate's base rate (prior) before any evidence β the proposition names the candidate. | A probability in [0, 1]. |
relatedPropositionLift | How much one piece of evidence (the proposition) multiplies the odds. May carry a prior block naming the predicted target's own attributes behind the lift β see below. | > 1 supports the candidate, < 1 argues against it, 1 is neutral. |
hitLinkPropositionLift | The lift of the link value that points at the returned hit β on a cross-table query, how much the linking field's own value mattered (e.g. { "product": 4 } in an impressions table that links to products). Common on cross-table predictions, and often the most interesting factor there. | Same scale as relatedPropositionLift: > 1 supports, < 1 argues against. |
hitPropositionLift | The aggregated lift of a proposition on the hit itself. Unlike the leaves above it carries its own nested factors array. | Same scale; read its factors for the breakdown. |
$group | A bundle of correlated features the engine folded into one theme, contributing a single redundancy-adjusted lift instead of several near-duplicate votes (group re-expression β see Tuning). | Reads as { "type": "β¦", "$group": [ β¦ ] }; its lift already accounts for the correlation among members. |
normalizer (exclusiveness, trueFalseExclusiveness, β¦) | Structural corrections that keep the candidate distribution normalized. | Multiplicative; usually near 1. |
calibration | A per-candidate confidence correction. The name says which mechanism produced it β rowCap, support-tempering(β¦), temperature(Ο=β¦), or a combination. Omitted entirely when it would be exactly 1.0. | Near 1.0 = well supported; well below 1.0 = the prediction rests on thin or overlapping evidence. config.calibrate can add a residual temperature. |
composition β mediation | A later inference stage replaced the direct probability β routing evidence through a linked sibling column. The factor is the ratio replacedP / directP, so the tree still multiplies to $p. | > 1 the stage raised the candidate, < 1 lowered it. |
This table lists the leaves you will actually meet on a prediction. It is
not exhaustive β the scoring format also emits arithmetic and text-scoring
nodes (product, sum, division, exp, log, tf, idf, token,
similarity, weightedAverage, baseLift, regression, neighborContext,
β¦). Write a $why renderer so an unrecognised type degrades to showing its
value rather than dropping the node.
Explaining through the target's own attributes
Available since: v2.8.0.
When the predicted target is a linked row (or a text value), the candidate's
own fields never appear as direct evidence β they enter through the
cross-table prior: each piece of evidence is correlated with the target table's
attributes, and candidates whose attributes match get their lift smoothed
toward that association. A relatedPropositionLift factor that was smoothed
this way now carries an optional prior block naming exactly which target
attributes did the work:
{
"type": "relatedPropositionLift",
"proposition": { "description": { "$has": "it" } },
"value": 2.79,
"prior": {
"type": "linkedPrior",
"factors": [
{ "proposition": { "processor.department": { "$has": "IT" } }, "value": 2.56 },
{ "proposition": { "$not": { "processor.role": { "$has": "Manager" } } }, "value": 1.14 }
]
}
}
Read it back: the evidence description:"it" lifts this processor by 2.79Γ,
and the prior behind that lift is the processor's own profile β invoices
mentioning it go to the IT department 2.56Γ more often than chance, and
this candidate is in IT. A $not proposition means the candidate lacks
the attribute and the shown lift is for that absence (e.g. not being a
Manager mildly helps here). Each prior factor's value is the co-occurrence
lift P(evidence & attribute) / (P(evidence) Γ P(attribute)) β above 1 the
pairing supports the candidate, below 1 it argues against.
Two things to know when rendering it:
- The block appears only on factors that actually used a cross-table (or
text-token) prior β typically predictions of a link field with
basedOn, and text-value predictions where the candidate value's own tokens play the same role. Factors without such a prior simply omit the field. - The prior factors describe the smoothing target; they are not extra
multiplicative terms. The parent factor's
valueis still the number that enters the$pproduct β thepriorblock explains where it leaned.
Highlight the evidence inside the predicted item
Available since: v2.8.0.
For a UI, the raw prior factors still need assembling. Add $highlight
to a link-target predict's select and each candidate comes back with its
own field values marked in place, ready to display β the same
target-attribute lifts, rendered by the field's analyzer:
{
"from": "assignments",
"where": { "role": "backend developer" },
"predict": "person",
"basedOn": ["tags"],
"select": ["$p", "$value",
{ "$highlight": { "posPreTag": "<mark>", "posPostTag": "</mark>",
"negPreTag": "<del>", "negPostTag": "</del>" } }]
}
Each hit's $highlight is an array of { "score", "field", "highlight" }
entries over the candidate's fields (the same shape search's $highlight
uses, so rendering code is shared). An attribute that supports the candidate
(prior lift > 1) gets the positive tags, one that argues against (lift < 1)
gets the negative tags β for the query above, a backend-tagged person renders
tags: <mark>backend</mark> while a frontend-tagged one renders
tags: <del>frontend</del>, making the reason for the ranking visible at a
glance. The default markup is <font color="green"> / <font color="red">;
the parametric form above sets your own, and "encoder": "text" skips HTML
escaping.
Only attributes the candidate actually has can be marked β the $not
absences remain visible in $why.prior but have nothing to highlight.
Candidate-side $highlight currently covers link targets; on a text-value
predict the v2 surface does not yet decompose the candidate's tokens this way
(the $why prior blocks on text targets appear on the v1 surface).
All of these β baseP, the normalizer group, every relatedPropositionLift and
hitLinkPropositionLift, and any
composition β are direct children of the top-level factors array. A composition
stage appends its factor next to the lifts rather than nesting them, so code that reads
$why.factors (for example to pull the highlight off each lift) sees the same shape
whether or not a stage fired.
Because group re-expression is on by default, a
factor may itself be a $group β the tree is still exact arithmetic (a group is
one node whose lift already accounts for the correlation among its members), it
just reads as a theme rather than a flat list. That is what keeps confidence
honest when a record has many overlapping features.
Probabilities, honestly
A few properties worth knowing:
$pis a probability, not a relevance score, and how well it is calibrated is something you can measure rather than take on trust:_evaluatereports calibration (ECE, reliability bins) on your own held-out rows. Do that before you set a threshold such as "auto-accept above 0.8". Where the error shows up as under-confidence, a threshold is conservative; where it shows up as over-confidence,config.calibratefits a residual temperature per query. The published measurements are on the Benchmarks pages.$pis calibrated relative to the data it learned from. Words the table has never seen carry no evidence either way, so in a text that is mostly unfamiliar, the few words the table does know decide the answer β and can make it confidently wrong. On a support corpus where back only ever appeared in "I want my money back", the ticket "I am not sure, can someone just call me back" predictsrefundat 0.98: back is the only discriminating word the table knows. Check$whyto see which evidence carried an answer, and route input that is far from your data to a person rather than trusting$palone.- Nothing is certain on finite data: a value never seen with the evidence gets a
small smoothed
$p, not zero (see Candidates, evidence, and$f). - A
$numeric/ quantity field contributes via its neighbourhood, so sparse numeric evidence is smoothed rather than memorized (see Schema Design). - Unsupported computed columns fail loud (a
4xx), never a silent empty value.
Tuning inference β config.ai
When known features are correlated β brand and model, or several tokens
that always co-occur β treating each as independent evidence over-weights the
overlap and inflates confidence. config.ai selects how the engine
re-expresses the known features before predicting, to correct for that.
v2 defaults to group. Group formation is on out of the box, because
real-world evidence is usually correlated and honest confidence on redundant
features is the common case. Set config.ai per query to change it:
{
"from": "products",
"where": { "name": "fullcorn rye bread", "tags": { "$has": "gluten" } },
"predict": "category",
"config": { "ai": "fast" }
}
config.ai | What it does |
|---|---|
fast (flat) | Plain naive Bayes β no re-expression. Cheapest; opt into this only when features are genuinely independent. |
v1 (and) | Peels clear feature combinations into joint terms (an "and" of co-occurring evidence). |
group (v2) (default) | Group formation β themes correlated features into a single group that contributes once. The engine default when config.ai is absent: it keeps confidence honest on correlated evidence and is best-or-tied across the tested corpora. v2 is an alias for this default. |
high | And + group formation. Additionally peels strong feature combinations into joint terms before grouping the residual. Opt into it for that extra step β it adds cost without beating group on the tested corpora, so it is not the default and is not aliased to v2. |
The default profile (group) applies group formation, which detects that
several features move together and folds them into one group that contributes a
single, calibrated piece of evidence, rather than letting each cast a near-duplicate
vote. This directly counters the over-confidence a record with many redundant
features would otherwise produce. config.ai: "high" adds an And step on top
(strong combinations become joint terms first); it is not the default because that
extra step adds cost without improving accuracy on the tested corpora. Group
formation composes with config.calibrate (the residual temperature fit); the two
are independent knobs. To turn re-expression off entirely, set config.ai: "fast".
An unknown profile name is rejected loudly (4xx), never silently ignored.
Related
- Schema Design (v2) Β· Query Operator Reference (v2) Β· v2 Introduction
- v1 Inference β the deeper treatment of priors and representation learning.