Inference (v2, Beta)

This guide explains the inference operations on the v2 (rep2) engine: predict, recommend, relate, ranking by context, and explainable results with $why. v2 is in beta; the statistical model is the same family as v1 inference, exposed through the simpler Query2 syntax.

The idea behind Aito is that the same from/where you'd use to filter data is also the evidence for a prediction. You don't train a model per question β€” you ask the database.

Predict

Point predict at a field and Aito ranks its values by probability given your where evidence:

{
  "from": "products",
  "where": { "name": { "$match": "rye bread" } },
  "predict": "category",
  "select": ["$p", "$value"]
}

Each hit has a $value (a candidate category) and $p (its probability). Results are ordered most-probable first.

Across linked tables

Evidence can come from a linked table using a dotted path, and basedOn projects linked-table fields into the prediction context as extra evidence:

{
  "from": "invoices",
  "where": { "Description": "monthly cloud hosting", "Processor.Department": "IT" },
  "basedOn": ["Department"],
  "predict": "GLCode",
  "select": ["$p", "$value"]
}

Cross-table predict / recommend / relate run end-to-end on v2, and cross-table priors (the dominant accuracy contributor) are on by default. You can also predict the link itself (e.g. predict processor and let the engine draw on the linked employee's attributes), and match into array/multi-value columns with $has (membership), so a tag list or a basket of products is first-class evidence.

Recommend

recommend is predict paired with a goal β€” rank candidates by how well they achieve a desired outcome, not just by similarity:

{
  "from": "impressions",
  "where": { "user": "john" },
  "recommend": "product",
  "goal": { "purchase": true },
  "limit": 10
}

This ranks products by P(purchase | user=john, product), i.e. what John is most likely to actually buy β€” not merely what's popular.

Rank by context

You can rank candidates by the probability of a condition evaluated against the context (from) table, using $context inside orderBy:

{
  "from": "impressions",
  "get": "product",
  "orderBy": { "$p": { "$context": { "purchase": true } } },
  "select": ["$value", "$p"]
}

"Rank products by P(purchase | product)". It reuses the recommend machinery β€” the context condition becomes the goal β€” so there's no separate scoring path.

Relate

relate returns statistical relations instead of rows β€” which features move a target, with lift and information-gain style statistics:

{
  "from": "impressions",
  "where": { "purchase": true },
  "relate": ["product.category"]
}

Use it to discover what's associated with an outcome.

Explainable results β€” $why

Predictions you can't explain are hard to trust. Add $why to select and every ranked candidate comes back with the factor tree that produced its score. For a collection where color predicts label (red→a, blue→b):

{
  "from": "obs",
  "where": { "color": "red" },
  "predict": "label",
  "select": ["$p", "$value", "$why"]
}

returns, for the top candidate:

{
  "$p": 0.8,
  "$value": "a",
  "$why": {
    "type": "product",
    "factors": [
      { "type": "baseP", "value": 0.5, "proposition": { "label": { "$has": "a" } } },
      { "type": "relatedPropositionLift", "proposition": { "color": "red" }, "value": 1.6 }
    ]
  }
}

Read it back: the base rate for label=a is 0.5, the evidence color=red lifts it by 1.6Γ—, and 0.5 Γ— 1.6 β‰ˆ the predicted 0.8. The explanation isn't a post-hoc approximation β€” it's the actual arithmetic the engine did.

The example above is trimmed to the two factors that carry the story. A real tree also contains structural normalizer nodes and a calibration node, so the exact $p is a little below the 0.5 Γ— 1.6 you'd get from just those two β€” every node the engine multiplied is shown, which is what makes the number auditable rather than approximate.

The factor types

Each node in the tree has a type. A product node's value is the product of its children; the leaves are:

typeMeaningScale
basePThe candidate's base rate (prior) before any evidence β€” the proposition names the candidate.A probability in [0, 1].
relatedPropositionLiftHow much one piece of evidence (the proposition) multiplies the odds.> 1 supports the candidate, < 1 argues against it, 1 is neutral.
hitLinkPropositionLiftThe lift of the link value that points at the returned hit β€” on a cross-table query, how much the linking field's own value mattered (e.g. { "product": 4 } in an impressions table that links to products). Common on cross-table predictions, and often the most interesting factor there.Same scale as relatedPropositionLift: > 1 supports, < 1 argues against.
hitPropositionLiftThe aggregated lift of a proposition on the hit itself. Unlike the leaves above it carries its own nested factors array.Same scale; read its factors for the breakdown.
$groupA bundle of correlated features the engine folded into one theme, contributing a single redundancy-adjusted lift instead of several near-duplicate votes (group re-expression β€” see Tuning).Reads as { "type": "…", "$group": [ … ] }; its lift already accounts for the correlation among members.
normalizer (exclusiveness, trueFalseExclusiveness, …)Structural corrections that keep the candidate distribution normalized.Multiplicative; usually near 1.
calibrationA per-candidate confidence correction. The name says which mechanism produced it β€” rowCap, support-tempering(…), temperature(Ο„=…), or a combination. Omitted entirely when it would be exactly 1.0.Near 1.0 = well supported; well below 1.0 = the prediction rests on thin or overlapping evidence. config.calibrate can add a residual temperature.
composition β€” mediation, nameBoostA later inference stage replaced the direct probability β€” routing evidence through a linked sibling column (mediation), or matching an identity/name token (nameBoost). The factor is the ratio replacedP / directP, so the tree still multiplies to $p.> 1 the stage raised the candidate, < 1 lowered it.

This table lists the leaves you will actually meet on a prediction. It is not exhaustive β€” the scoring format also emits arithmetic and text-scoring nodes (product, sum, division, exp, log, tf, idf, token, similarity, weightedAverage, baseLift, regression, neighborContext, …). Write a $why renderer so an unrecognised type degrades to showing its value rather than dropping the node.

All of these β€” baseP, the normalizer group, every relatedPropositionLift and hitLinkPropositionLift, and any composition β€” are direct children of the top-level factors array. A composition stage appends its factor next to the lifts rather than nesting them, so code that reads $why.factors (for example to pull the highlight off each lift) sees the same shape whether or not a stage fired.

Because group re-expression is on by default, a factor may itself be a $group β€” the tree is still exact arithmetic (a group is one node whose lift already accounts for the correlation among its members), it just reads as a theme rather than a flat list. That is what keeps confidence honest when a record has many overlapping features.

Probabilities, honestly

A few properties worth knowing:

  • Scores are calibrated probabilities, not arbitrary relevance numbers β€” a $p of 0.8 means roughly an 80% chance. Calibration is kept honest by support-tempering (on by default): confidence is tempered by the effective statistical support of the evidence, visible as a support-tempering(auto) factor in $why; config.calibrate can additionally fit a residual temperature. See the calibration factor.
  • A $numeric / quantity field contributes via its neighbourhood, so sparse numeric evidence is smoothed rather than memorized (see Schema Design).
  • Unsupported computed columns fail loud (a 4xx), never a silent empty value.

Tuning inference β€” config.ai

When known features are correlated β€” brand and model, or several tokens that always co-occur β€” treating each as independent evidence over-weights the overlap and inflates confidence. config.ai selects how the engine re-expresses the known features before predicting, to correct for that.

v2 defaults to group. Group formation is on out of the box, because real-world evidence is usually correlated and honest confidence on redundant features is the common case. Set config.ai per query to change it:

{
  "from": "products",
  "where": { "brand": "acme", "line": "pro", "color": "red" },
  "predict": "category",
  "config": { "ai": "fast" }
}
config.aiWhat it does
fast (flat)Plain naive Bayes β€” no re-expression. Cheapest; opt into this only when features are genuinely independent.
v1 (and)Peels clear feature combinations into joint terms (an "and" of co-occurring evidence).
group (v2) (default)Group formation β€” themes correlated features into a single group that contributes once. The engine default when config.ai is absent: it keeps confidence honest on correlated evidence and is best-or-tied across the tested corpora. v2 is an alias for this default.
highAnd + group formation. Additionally peels strong feature combinations into joint terms before grouping the residual. Opt into it for that extra step β€” it adds cost without beating group on the tested corpora, so it is not the default and is not aliased to v2.

The default profile (group) applies group formation, which detects that several features move together and folds them into one group that contributes a single, calibrated piece of evidence, rather than letting each cast a near-duplicate vote. This directly counters the over-confidence a record with many redundant features would otherwise produce. config.ai: "high" adds an And step on top (strong combinations become joint terms first); it is not the default because that extra step adds cost without improving accuracy on the tested corpora. Group formation composes with config.calibrate (the residual temperature fit); the two are independent knobs. To turn re-expression off entirely, set config.ai: "fast". An unknown profile name is rejected loudly (4xx), never silently ignored.