Inference (v2)

This guide explains the inference operations on the v2 (rep2) engine: predict, recommend, relate, ranking by context, and explainable results with $why. The statistical model is the same family as v1 inference, exposed through the simpler Query2 syntax.

The idea behind Aito is that the same from/where you'd use to filter data is also the evidence for a prediction. You don't train a model per question β€” you ask the database.

Predict

Point predict at a field and Aito ranks its values by probability given your where evidence:

{
  "from": "products",
  "where": { "name": { "$match": "rye bread" } },
  "predict": "category",
  "select": ["$p", "$value"]
}

Each hit has a $value (a candidate category) and $p (its probability). Results are ordered most-probable first.

Candidates, evidence, and $f

where is evidence, not a filter on the answer. A predict ranks the target's whole value set, and the evidence decides each value's probability. Which values are candidates depends on the target:

target columncandidates
a plain column (PaymentMethod, category)every value that occurs in the column
a link (GLCode β†’ glCodes)every row of the linked table β€” including a row no other row points to yet

So a predict with evidence still returns values that never occurred with that evidence. Select $f to see which ones did: $f is the number of rows that match the where and carry the value.

The one exception: a where on the predicted field itself, or on a path under it. {"GLCode": {"$or": [...]}} or {"GLCode.Department": "IT & Infrastructure"} while predicting GLCode is not evidence β€” it is a statement about the answer, and it restricts the candidate set to the values it admits. That is what scoping below is built on.

{
  "from": "invoices",
  "where": { "Processor": "Emily Davis" },
  "predict": "GLCode",
  "select": ["$value", "$p", "$f"],
  "limit": 4
}
{ "offset": 0, "total": 10, "hits": [
  { "$f": 20, "$p": 0.847, "$value": "F001" },
  { "$f": 0, "$p": 0.042, "$value": "R001" },
  { "$f": 0, "$p": 0.032, "$value": "E001" },
  { "$f": 0, "$p": 0.027, "$value": "E002" } ] }

Emily Davis's 20 invoices are all coded F001, so F001 has $f: 20. The other accounts have $f: 0 β€” never seen with this evidence β€” and still get a small probability rather than zero:

  • Why not zero? Every estimate is smoothed: an account's probability given the evidence is pulled toward how common that account is overall, and the pull weakens as the evidence gets more rows of its own. An account that is frequent elsewhere (R001) keeps more of the remaining probability than a rare one. Nothing is ever certain on finite data, and a $p of exactly 0 would claim that it is.
  • Why is that useful? A GL account added to glCodes today is a candidate immediately, at a small $p, and rises as invoices are coded to it β€” there is no cold-start gap to engineer around.

To return only the values seen with the evidence, add having:

{
  "from": "invoices",
  "where": { "Processor": "Emily Davis" },
  "predict": "GLCode",
  "select": ["$value", "$p", "$f"],
  "having": { "$f": { "$gte": 1 } }
}
{ "offset": 0, "total": 1, "hits": [
  { "$f": 20, "$p": 0.847, "$value": "F001" } ] }

having filters the ranked list after scoring: the remaining $p values are not renormalised, and $f is the only field it accepts.

Scoping candidates per customer

If one Aito database serves several customers (tenants), decide how each predict's candidates are scoped. A where term that is evidence will not do it, and neither will a nested from when the target is a link: the restriction scopes base rates and counts, but a link target's candidates stay every row of the linked table.

What works today:

  1. A where on the target's own path β€” the linked table carries the scoping column, and the query names it under the predicted field:

    {
      "from": "invoices",
      "where": { "Description": "monthly cloud hosting", "Processor.Department": "Finance" },
      "predict": "Processor",
      "select": ["$value", "$p"]
    }
    

    Processor.Department sits under the predicted field, so it scopes the candidate set rather than acting as evidence: on the demo this ranks the two employees in Finance, where the same predict without it ranks all ten. The rest of the where is ordinary evidence. In a multi-tenant schema the scoping column is the tenant β€” {"GLCode.tenant": "acme"} while predicting GLCode ranks only that tenant's accounts, including ones it has never used yet.

    Equality, $in, ranges and other operators, and $and/$or/$not all scope the same way, as long as every leaf of a combinator sits under the link path β€” a clause satisfiable through a term outside it ({"$or": [{"GLCode.tenant": "acme"}, {"TotalAmount": 10}]}) constrains nothing, because a row can satisfy it without the linked condition holding.

    This is the option to reach for first: it is one query, and it needs no per-tenant schema. It scopes the candidates only: the probabilities are still learned from every tenant's rows, so another tenant's bookings move this tenant's ranking. To keep them out as well, put the tenant into a nested from beside the link filter β€” see Multi-tenant isolation. Neither is access control, so both belong beside a server-side check that the caller may see that tenant, not instead of one.

    Available since: v2.8.3 for the operator and combinator forms, and v2.9.0 for all of them again β€” v2.8.4 regressed link-target scoping by an attribute of the target and ignored it.

  2. A per-tenant linked table β€” the target links to a table that holds only that tenant's values (for example one glCodes_<tenant> table per tenant, or one database per tenant). The candidates are scoped by construction, and a new value is a candidate as soon as its row exists. This is the design to choose when tenants must never see each other's values. The statistics still come from the rows the query sees, so on a shared invoice table pair it with a nested from as well.

  3. Two steps over one shared table, when the tenant is not a column on the linked table and only the co-occurrence history distinguishes the values.

    1. Ask for the tenant's own values: a predict whose where holds only the tenant, with "having": { "$f": { "$gte": 1 } }.
    2. Run the real predict β€” the invoice's evidence in where (and the tenant as a nested from, so the base rates are the tenant's) β€” and keep only the hits whose $value is in the set from step 1.

    Do not put the invoice's evidence into step 1: $f counts rows that match the whole where, so with real evidence beside the tenant it is 0 for most values and having returns nothing.

having shapes what a query returns; it is not an access control.

Across linked tables

Evidence can come from a linked table using a dotted path, and basedOn projects linked-table fields into the prediction context as extra evidence:

{
  "from": "invoices",
  "where": { "Description": "monthly cloud hosting", "Processor.Department": "IT & Infrastructure" },
  "basedOn": ["Department"],
  "predict": "GLCode",
  "select": ["$p", "$value"]
}

Cross-table predict / recommend / relate run end-to-end on v2, and cross-table priors (the dominant accuracy contributor) are on by default.

Evidence through the target's attributes. basedOn names columns of the linked table, and those attributes then carry evidence of their own: predicting GLCode with basedOn: ["Department"] lets a department the invoice's evidence implies lift every account belonging to it, including accounts this evidence has never co-occurred with. Left out, the default is deliberately narrow β€” only a linked column that a where term also names is projected β€” so the fields a prediction is allowed to draw on stay the ones the caller declared.

Scoping the candidate set. Apart from the target's own path (above), where supplies evidence, not an output domain β€” a predict ranks the field's global value set, with values never seen under your where in the tail at $f: 0. For a link target the value set is the linked table (every target row, matching the v1 engine), so an entity that exists but was never referenced in training is still a ranked candidate β€” scored by its cross-table priors and name evidence rather than co-occurrence. That is what lets a brand-new employee whose name appears in an invoice's text be predicted as its handler. When the tail must not show them at all (the multi-tenant case: a tenant-scoped predict must not list other tenants' values as low-probability alternatives), add the candidate filter:

{
  "from": "invoices",
  "where": { "ReceiverName": "Acme Grocery Store" },
  "predict": "Processor",
  "having": { "$f": { "$gte": 1 } }
}

having keeps only candidates that actually co-occur with the where ($f β‰₯ 1), applied after ranking and before offset/limit, with no renormalisation. On the demo that is the difference between ranking all ten employees and ranking the four who have actually processed this receiver's invoices. Only $f with one comparison operator is accepted β€” anything else is a loud 400.

Cost follows the candidate count, not just the row count. A predict scores every candidate against the evidence, so a link target over a large table is the expensive shape β€” not because the table has many rows, but because each row is a value to rank. Scoping with a where on the target's own path is therefore a latency lever as much as a correctness one: it cuts the pool before scoring, where an unscoped prediction pays for every row of the linked table on every request. On a large linked table that difference dominates the response time. having does not help here β€” it filters after ranking, so everything is still scored. Two rules of thumb follow: scope by construction where you can, and treat β€œhow many candidates does this query rank?” as the first question when a prediction is slower than you expect.

Available since: v2.8.0 for having, and for the linked-table candidate universe on the v2 engine (before it, v2 ranked only link values seen in training). You can also predict the link itself (e.g. predict processor and let the engine draw on the linked employee's attributes), and match into array/multi-value columns with $has (membership), so a tag list or a basket of products is first-class evidence.

Recommend

recommend is predict paired with a goal β€” rank candidates by how well they achieve a desired outcome, not just by similarity:

{
  "from": "impressions",
  "where": { "context.user": "larry" },
  "recommend": "product",
  "goal": { "purchase": true },
  "limit": 10
}

This ranks products by P(purchase | context.user=larry, product), i.e. what Larry is most likely to actually buy β€” not merely what's popular.

Rank by context

You can rank candidates by the probability of a condition evaluated against the context (from) table, using $context inside orderBy:

{
  "from": "impressions",
  "get": "product",
  "orderBy": { "$p": { "$context": { "purchase": true } } },
  "select": ["$value", "$p"]
}

"Rank products by P(purchase | product)". It reuses the recommend machinery β€” the context condition becomes the goal β€” so there's no separate scoring path.

Relate

relate returns statistical relations instead of rows β€” which features move a target, with lift and information-gain style statistics:

{
  "from": "impressions",
  "where": { "purchase": true },
  "relate": ["product.category"]
}

Use it to discover what's associated with an outcome.

Each hit reports the relation from several angles: lift and info (the smoothed strength), fs/ps (exact counts and within-population rates), and relation (the raw 2Γ—2 contingency the numbers derive from).

Date columns relate as ranges. A Date field's values are grouped into calendar buckets sized by the field's span β€” days under ~3 months, months up to ~4 years, years beyond β€” so a two-year order history relates as monthly ranges instead of one relation per distinct day. Each bucket's condition is the re-queryable range form, e.g. {"$and":[{"order_date":{"$gte":"2024-03-01"}},{"order_date":{"$lt":"2024-04-01"}}]}, which you can paste straight back into a where.

Available since: v2.8.0 for the relation block and the date-range grouping.

Explainable results β€” $why

Predictions you can't explain are hard to trust. Add $why to select and every ranked candidate comes back with the factor tree that produced its score. For a collection where color predicts label (red→a, blue→b):

{
  "from": "obs",
  "where": { "color": "red" },
  "predict": "label",
  "select": ["$p", "$value", "$why"]
}

returns, for the top candidate:

{
  "$p": 0.8,
  "$value": "a",
  "$why": {
    "type": "product",
    "factors": [
      { "type": "baseP", "value": 0.5, "proposition": { "label": { "$has": "a" } } },
      { "type": "relatedPropositionLift", "proposition": { "color": "red" }, "value": 1.6 }
    ]
  }
}

Read it back: the base rate for label=a is 0.5, the evidence color=red lifts it by 1.6Γ—, and 0.5 Γ— 1.6 β‰ˆ the predicted 0.8. The explanation isn't a post-hoc approximation β€” it's the actual arithmetic the engine did.

The example above is trimmed to the two factors that carry the story. A real tree also contains structural normalizer nodes and a calibration node, so the exact $p is a little below the 0.5 Γ— 1.6 you'd get from just those two β€” every node the engine multiplied is shown, which is what makes the number auditable rather than approximate.

The factor types

Each node in the tree has a type. A product node's value is the product of its children; the leaves are:

typeMeaningScale
basePThe candidate's base rate (prior) before any evidence β€” the proposition names the candidate.A probability in [0, 1].
relatedPropositionLiftHow much one piece of evidence (the proposition) multiplies the odds. May carry a prior block naming the predicted target's own attributes behind the lift β€” see below.> 1 supports the candidate, < 1 argues against it, 1 is neutral.
hitLinkPropositionLiftThe lift of the link value that points at the returned hit β€” on a cross-table query, how much the linking field's own value mattered (e.g. { "product": 4 } in an impressions table that links to products). Common on cross-table predictions, and often the most interesting factor there.Same scale as relatedPropositionLift: > 1 supports, < 1 argues against.
hitPropositionLiftThe aggregated lift of a proposition on the hit itself. Unlike the leaves above it carries its own nested factors array.Same scale; read its factors for the breakdown.
$groupA bundle of correlated features the engine folded into one theme, contributing a single redundancy-adjusted lift instead of several near-duplicate votes (group re-expression β€” see Tuning).Reads as { "type": "…", "$group": [ … ] }; its lift already accounts for the correlation among members.
normalizer (exclusiveness, trueFalseExclusiveness, …)Structural corrections that keep the candidate distribution normalized.Multiplicative; usually near 1.
calibrationA per-candidate confidence correction. The name says which mechanism produced it β€” rowCap, support-tempering(…), temperature(Ο„=…), or a combination. Omitted entirely when it would be exactly 1.0.Near 1.0 = well supported; well below 1.0 = the prediction rests on thin or overlapping evidence. config.calibrate can add a residual temperature.
composition β€” mediationA later inference stage replaced the direct probability β€” routing evidence through a linked sibling column. The factor is the ratio replacedP / directP, so the tree still multiplies to $p.> 1 the stage raised the candidate, < 1 lowered it.

This table lists the leaves you will actually meet on a prediction. It is not exhaustive β€” the scoring format also emits arithmetic and text-scoring nodes (product, sum, division, exp, log, tf, idf, token, similarity, weightedAverage, baseLift, regression, neighborContext, …). Write a $why renderer so an unrecognised type degrades to showing its value rather than dropping the node.

Explaining through the target's own attributes

Available since: v2.8.0.

When the predicted target is a linked row (or a text value), the candidate's own fields never appear as direct evidence β€” they enter through the cross-table prior: each piece of evidence is correlated with the target table's attributes, and candidates whose attributes match get their lift smoothed toward that association. A relatedPropositionLift factor that was smoothed this way now carries an optional prior block naming exactly which target attributes did the work:

{
  "type": "relatedPropositionLift",
  "proposition": { "description": { "$has": "it" } },
  "value": 2.79,
  "prior": {
    "type": "linkedPrior",
    "factors": [
      { "proposition": { "processor.department": { "$has": "IT" } }, "value": 2.56 },
      { "proposition": { "$not": { "processor.role": { "$has": "Manager" } } }, "value": 1.14 }
    ]
  }
}

Read it back: the evidence description:"it" lifts this processor by 2.79Γ—, and the prior behind that lift is the processor's own profile β€” invoices mentioning it go to the IT department 2.56Γ— more often than chance, and this candidate is in IT. A $not proposition means the candidate lacks the attribute and the shown lift is for that absence (e.g. not being a Manager mildly helps here). Each prior factor's value is the co-occurrence lift P(evidence & attribute) / (P(evidence) Γ— P(attribute)) β€” above 1 the pairing supports the candidate, below 1 it argues against.

Two things to know when rendering it:

  • The block appears only on factors that actually used a cross-table (or text-token) prior β€” typically predictions of a link field with basedOn, and text-value predictions where the candidate value's own tokens play the same role. Factors without such a prior simply omit the field.
  • The prior factors describe the smoothing target; they are not extra multiplicative terms. The parent factor's value is still the number that enters the $p product β€” the prior block explains where it leaned.

Highlight the evidence inside the predicted item

Available since: v2.8.0.

For a UI, the raw prior factors still need assembling. Add $highlight to a link-target predict's select and each candidate comes back with its own field values marked in place, ready to display β€” the same target-attribute lifts, rendered by the field's analyzer:

{
  "from": "assignments",
  "where": { "role": "backend developer" },
  "predict": "person",
  "basedOn": ["tags"],
  "select": ["$p", "$value",
    { "$highlight": { "posPreTag": "<mark>", "posPostTag": "</mark>",
                      "negPreTag": "<del>", "negPostTag": "</del>" } }]
}

Each hit's $highlight is an array of { "score", "field", "highlight" } entries over the candidate's fields (the same shape search's $highlight uses, so rendering code is shared). An attribute that supports the candidate (prior lift > 1) gets the positive tags, one that argues against (lift < 1) gets the negative tags β€” for the query above, a backend-tagged person renders tags: <mark>backend</mark> while a frontend-tagged one renders tags: <del>frontend</del>, making the reason for the ranking visible at a glance. The default markup is <font color="green"> / <font color="red">; the parametric form above sets your own, and "encoder": "text" skips HTML escaping.

Only attributes the candidate actually has can be marked β€” the $not absences remain visible in $why.prior but have nothing to highlight. Candidate-side $highlight currently covers link targets; on a text-value predict the v2 surface does not yet decompose the candidate's tokens this way (the $why prior blocks on text targets appear on the v1 surface).

All of these β€” baseP, the normalizer group, every relatedPropositionLift and hitLinkPropositionLift, and any composition β€” are direct children of the top-level factors array. A composition stage appends its factor next to the lifts rather than nesting them, so code that reads $why.factors (for example to pull the highlight off each lift) sees the same shape whether or not a stage fired.

Because group re-expression is on by default, a factor may itself be a $group β€” the tree is still exact arithmetic (a group is one node whose lift already accounts for the correlation among its members), it just reads as a theme rather than a flat list. That is what keeps confidence honest when a record has many overlapping features.

Probabilities, honestly

A few properties worth knowing:

  • $p is a probability, not a relevance score, and how well it is calibrated is something you can measure rather than take on trust: _evaluate reports calibration (ECE, reliability bins) on your own held-out rows. Do that before you set a threshold such as "auto-accept above 0.8". Where the error shows up as under-confidence, a threshold is conservative; where it shows up as over-confidence, config.calibrate fits a residual temperature per query. The published measurements are on the Benchmarks pages.
  • $p is calibrated relative to the data it learned from. Words the table has never seen carry no evidence either way, so in a text that is mostly unfamiliar, the few words the table does know decide the answer β€” and can make it confidently wrong. On a support corpus where back only ever appeared in "I want my money back", the ticket "I am not sure, can someone just call me back" predicts refund at 0.98: back is the only discriminating word the table knows. Check $why to see which evidence carried an answer, and route input that is far from your data to a person rather than trusting $p alone.
  • Nothing is certain on finite data: a value never seen with the evidence gets a small smoothed $p, not zero (see Candidates, evidence, and $f).
  • A $numeric / quantity field contributes via its neighbourhood, so sparse numeric evidence is smoothed rather than memorized (see Schema Design).
  • Unsupported computed columns fail loud (a 4xx), never a silent empty value.

Tuning inference β€” config.ai

When known features are correlated β€” brand and model, or several tokens that always co-occur β€” treating each as independent evidence over-weights the overlap and inflates confidence. config.ai selects how the engine re-expresses the known features before predicting, to correct for that.

v2 defaults to group. Group formation is on out of the box, because real-world evidence is usually correlated and honest confidence on redundant features is the common case. Set config.ai per query to change it:

{
  "from": "products",
  "where": { "name": "fullcorn rye bread", "tags": { "$has": "gluten" } },
  "predict": "category",
  "config": { "ai": "fast" }
}
config.aiWhat it does
fast (flat)Plain naive Bayes β€” no re-expression. Cheapest; opt into this only when features are genuinely independent.
v1 (and)Peels clear feature combinations into joint terms (an "and" of co-occurring evidence).
group (v2) (default)Group formation β€” themes correlated features into a single group that contributes once. The engine default when config.ai is absent: it keeps confidence honest on correlated evidence and is best-or-tied across the tested corpora. v2 is an alias for this default.
highAnd + group formation. Additionally peels strong feature combinations into joint terms before grouping the residual. Opt into it for that extra step β€” it adds cost without beating group on the tested corpora, so it is not the default and is not aliased to v2.

The default profile (group) applies group formation, which detects that several features move together and folds them into one group that contributes a single, calibrated piece of evidence, rather than letting each cast a near-duplicate vote. This directly counters the over-confidence a record with many redundant features would otherwise produce. config.ai: "high" adds an And step on top (strong combinations become joint terms first); it is not the default because that extra step adds cost without improving accuracy on the tested corpora. Group formation composes with config.calibrate (the residual temperature fit); the two are independent knobs. To turn re-expression off entirely, set config.ai: "fast". An unknown profile name is rejected loudly (4xx), never silently ignored.