Evaluation Guide (v2)

The _evaluate endpoint measures how well a query performs on your own data, with no separate testing infrastructure β€” on a collection a predict or estimate query; on a classic table also a recommend, relate, match, or generic query (see Which evaluations run where). It splits the data into train and test sets, hides the test rows from Aito, runs the query for each test row, and reports accuracy, ranking, and probability-quality metrics.

POST /api/v2/_evaluate dispatches by the from table's engine: collection (rep2) tables run the native v2 evaluate; classic (rep1) tables keep the v1 evaluate pipeline. This guide covers the collection behaviour.

Which evaluations run where

On a collection, _evaluate measures a predict or an estimate query. The other kinds are evaluated only on a classic table ("type": "table"), through the v1 pipeline:

query kindon a collectionon a table ("type": "table")
predict (incl. a per-member .$feature target)yesyes
estimate β€” regression metrics mae, rmse, r2yes (since v2.11.2)yes
recommend, relate, match, similaritynot yetyes

A kind that is not available on a collection is refused with a message saying so and pointing here β€” it is not answered with an empty or partial result.

Scoring a numeric target: an estimate evaluation. Give estimate instead of predict, with the evidence as $get bindings (or literals). Each held-out row is estimated from the others by the same estimator _estimate uses ("model": knn, the default, or regression), and scored against the value it actually has:

{
  "test": { "$index": { "$mod": [5, 0] } },
  "evaluate": {
    "from": "sales",
    "where": { "store": { "$get": "store" }, "week": { "$get": "week" } },
    "estimate": "units"
  },
  "select": ["n", "mae", "rmse", "r2", "baselineMae", "maeGain"]
}

The report has the table engine's regression metrics: mae, rmse, mse, r2, mape, medianAbsoluteError, the error percentiles p50–p99, and the baseline β€” every row estimated as the training mean β€” as baselineMae / baselineRmse, with maeGain / rmseGain (baseline minus model; positive means the model helps) and their relative …Improvement. Evidence is by equality: a $get or a literal; an operator in where is refused with a message saying so. A test row the model has no estimate for (no training row shares its evidence) is left out of n, and the response message says how many.

Available since: v2.11.2. Up to v2.11.1 an estimate evaluation on a collection was refused ("not yet available on collections").

Available since: v2.11.0 for the refusal message. Up to v2.10.3 a recommend (or estimate, relate, match, similarity) evaluation on a collection failed with evaluate: missing 'evaluate.predict', which named a missing field rather than the real limitation.

How evaluation works

An evaluate query has two parts:

  • evaluate β€” the query to measure: on a collection, a _predict or _estimate body; on a table, also the body you would send to _recommend, _relate, _match, or _query (see Which evaluations run where).
  • test (or testSource) β€” which rows to test on.

Which rows train depends on the selector you used β€” the same rule on both API generations, so the two engines run the same experiment for the same request:

  • test (and the v1 default sample): HOLD-OUT. The training data is everything except the test rows β€” they are hidden by a transient masked delete, so no test row can leak into its own prediction. This is the generalisation measurement, and the right default.
  • testSource: RESUBSTITUTION. The training data is the whole table, test rows included. This is deliberate, not a leak: scoring rows the model has seen measures overfitting (compare it against a held-out run of the same query), and it is the basis of the calibration temperature fit. If you want a hold-out with a windowed test set, express the window with test operators (e.g. $index+$mod, or a field cutoff) instead of testSource.

trainSamples in the response always tells you which experiment ran: it equals the full table size under resubstitution, and total βˆ’ testSamples under hold-out.

For each test row, Aito runs the query against the training data and compares the result against the row's known answer. The metrics aggregate those per-row outcomes.

Available since: v2.8.0 for the testSource resubstitution alignment β€” before it, v2 always held the test rows out, so a testSource evaluate reported a different (held-out) accuracy than the same request on v1.

{
  "test": { "$index": { "$mod": [5, 0] } },
  "evaluate": {
    "from": "products",
    "where": { "name": { "$get": "name" } },
    "predict": "category"
  },
  "select": ["accuracy", "baseAccuracy", "meanRank", "n"]
}

Choosing the test set

Evaluate runs one query per test row, so the runtime grows with the size of the test set. On collection tables the selectors are:

  • $index β€” exact rows by position. { "$index": 0 }, or several under $or ({ "$or": [ { "$index": 0 }, { "$index": 1 }, { "$index": 2 } ] }). Useful for pinning a reproducible spot-check.

  • $index + $mod β€” a deterministic fold (recommended). { "$index": { "$mod": [k, r] } } selects every row whose position is ≑ r (mod k) β€” a reproducible 1/k slice with no overlap between folds. This is the leak-free way to do k-fold cross-validation on a collection: fold r is the test set, the other kβˆ’1 folds train.

    "test": { "$index": { "$mod": [10, 0] } }
    

    runs a deterministic 10% hold-out; sweep r from 0..9 for full 10-fold CV.

  • testSource β€” a windowed or separate test set. Use where (field conditions and comparison operators) to filter, limit/offset to take a window, and from to draw the test set from a different table:

    "testSource": { "from": "holdout", "limit": 200 }
    "testSource": { "where": { "month": { "$lte": "2024-07" } } }
    
  • Compound selectors β€” $and / $or mixing positions and fields. An $index selector and a field condition can be combined, and nested:

    "test": { "$and": [ { "$index": { "$mod": [3, 0] } }, { "promo": true } ] }
    

    selects the rows that are both in the fold and on promotion; $or selects either. The $index inside is still the row's position in the table, not its position among the rows the field condition kept.

    Available since: v2.9.0. On v2.8.4 and earlier, a test selector mixing $index with a field condition was refused with a 400.

  • Operator conditions. Both the test/testSource where and the evaluate where accept comparison operators β€” e.g. a time cutoff { "month": { "$lte": "2024-07" } } (String columns compare lexicographically) β€” alongside $get bindings.

  • $get through a link. A binding may read the test row's LINKED value β€” { "product.pet_type": { "$get": "product.pet_type" } } gives the model the product's category as evidence for a prediction about the order. One link hop is supported; a deeper path is refused by name.

    Available since: v2.10.0. On v2.9.2 and earlier a $get whose source was a link path bound nothing β€” silently. Every test row was then scored with no evidence at all, so the evaluate reported the model performing at exactly its own baseAccuracy and nothing said why. A $get naming a source that resolves nowhere (a typo, a dropped column) is now a 400 for the same reason: scoring every row blind must not look like a weak model.

A baseAccuracy of 0.0 is about your SPLIT, not your model. The baseline names one value and is scored on the test rows; 0.0 means no test row carries it. That is what a test set drawn from a different part of the population than the train rows produces β€” and testSource with a limit takes the first n rows, which is not a random sample when your export arrives ordered by the column you are predicting. accuracyGain is then the accuracy itself, which reads as a spectacular model. The response explains it in message; to get a representative split, select the test rows with $sample, $hash or $index $mod instead of a positional limit.

Available since: v2.10.0 for the explanation, and for the guarantee that the baseline never answers a value absent from the training population. On v2.9.2 and earlier the baseline was scored from the full table's value frequencies even for values the hold-out had removed, so it could be right about test rows whose class it had never been trained on.

No implicit default test set. Unlike v1 tables, a collection evaluate that omits both test and testSource is a loud 400, not a silent 100-row sample β€” an evaluation you didn't scope should fail rather than quietly measure something you didn't intend. The sampling selectors $sample (including the {n, of, seed} form) and $hash (with $mod) work on collections, alongside $index $mod for a deterministic fraction.

Evaluating one tenant β€” a nested from

Available since: v2.11.3. On v2.11.2 and earlier a nested evaluate.from is refused with a 400 ("must be a table name string"); use a per-tenant table or view there.

On a collection that holds many customers, evaluate.from takes the same nested from a _predict does, so a tenant's model is measured on that tenant alone:

{
  "test": { "fold": 0 },
  "evaluate": {
    "from": { "from": "invoices", "where": { "tenant": "A" } },
    "where": { "vendor": { "$get": "vendor" }, "description": { "$get": "description" } },
    "predict": "gl"
  },
  "select": ["accuracy", "baseAccuracy", "n", "trainSamples"]
}

The test rows are drawn from the tenant's rows only, and the model trains on the rest of that tenant. No other tenant's row is tested on or trained on, so the numbers equal the same evaluate over a table holding only that tenant's rows. $index and $sample count positions within the tenant's rows, as they would in that table. Putting the tenant into evaluate.where instead measures a different model: where is evidence, so every tenant's rows still shape the predictions (see Multi-tenant isolation).

Reading the metrics

Select the metrics you want with the top-level select field (omit select to get all of the scalar metrics below). They fall into three groups β€” did it get the answer right, were the probabilities good, and is the confidence calibrated.

Discrimination & ranking β€” did it pick the right answer?

MetricMeaning
n / testSamplesNumber of test rows actually evaluated
trainSamplesAverage number of training rows per prediction
accuracyShare of test rows where the top result was correct
baseAccuracyAccuracy of the model given no evidence, scored on the same test rows β€” the baseline to beat. Your query's constant filters still apply, so the baseline faces the same population; only the per-row $get bindings are dropped. The baseline can only answer a value the training population actually carries, and when it scores 0.0 the response's message says why
baseSamplesThe rows baseAccuracy was measured over: equal to trainSamples for an unscoped evaluate, smaller when the query's where carries constant conditions β€” so the baseline can be checked
accuracyGainaccuracy - baseAccuracy
meanRankAverage position of the correct answer, 0-based (0 = always top). Same convention as v1, so the two generations are comparable.
mrrMean reciprocal rank β€” mean(1 / rank), the standard ranking metric (higher is better)
coverageFraction of test rows whose true answer was among the candidates at all
error1 - accuracy
baseError1 βˆ’ the share of the most common value among the training rows in scope β€” the error of always answering that value, measured on the training population. It feeds the probability-quality baseline. It is not 1 βˆ’ baseAccuracy: baseAccuracy is measured on the test rows, so on a small test set the two can differ (for example baseAccuracy 0.52 over 25 test rows beside baseError 0.513 over 76 training rows)
baseMeanRank / rankGainThe base-rate model's meanRank, and baseMeanRank βˆ’ meanRank (higher is better)
featuresAverage number of evidence features per prediction

Probability quality β€” proper scoring rules. These grade the whole probability, not just the top pick, over every test row (a truth the model rated β‰ˆ0 is penalised, not skipped):

MetricMeaning
logLossCross-entropy in nats β€” mean(βˆ’ln p(truth)). Lower is better. The standard "log loss."
brierScoreMean squared error of the top prediction's confidence. Lower is better.
geomMeanLikelihoodGeometric-mean probability assigned to the truth = exp(βˆ’logLoss) = 1 / perplexity. In [0, 1], higher is better.
logLossSkill1 βˆ’ logLoss / baseLogLoss β€” the fraction of the base-rate model's uncertainty removed. Higher is better (0 = no better than guessing the base rate).

Calibration β€” is a stated 80% actually right 80% of the time?

MetricMeaning
eceExpected Calibration Error β€” the gap between confidence and accuracy, averaged over confidence bins. Lower is better (0 = perfectly calibrated).
reliabilityOpt-in array: one row per confidence bin (confidence, accuracy, n) β€” the reliability diagram ece summarises.

Always compare accuracy against baseAccuracy (a model that beats the base rate is learning something), then read logLoss/geomMeanLikelihood for probability quality and ece for calibration. With the group default (below), predictions on correlated evidence are less over-confident β€” that shows up as a lower ece and better logLoss, not necessarily higher accuracy.

Legacy / information-theoretic names

Aito has measured prediction quality in bits over a prior since long before the Python-tooling vocabulary settled β€” those names are still available as exact aliases:

Legacy (bits)= ModernNote
mxe (mean cross-entropy)logLoss / ln 2Same quantity, base-2 instead of nats
hthe base-rate model's mxeEntropy of the target (bits): the cross-entropy of guessing from the value frequencies alone, i.e. with no evidence
informationGainlogLossSkill, in bitsBits saved over the base-rate model (the Kononenko–Bratko information score): h βˆ’ mxe
baseGmp / geomMeanLiftβ€”The base-rate model's geomMeanP, and geomMeanP / baseGmp
geomMeanPβ‰ˆ geomMeanLikelihoodCoverage-conditioned: averaged only over rows where the truth was a candidate, so it can slightly flatter the model β€” prefer geomMeanLikelihood / logLoss, which count every row

Evaluating a per-member target (.$feature)

For an array or set column β€” tags, categories, a basket β€” the prediction is over the column's members, spelled predict: "tags.$feature" exactly as in _predict. Evaluate scores it the way v1 scores a feature target:

  • One case per test row. A row with no members (an empty or missing cell) is skipped, not counted as a miss.
  • The truth is the row's set of members. The row is accurate when the top prediction is any of them.
  • correct and the rank are the best-ranked true member, so meanRank reads "how far down the list was the first right tag".
  • baseAccuracy asks the same question of the no-evidence baseline, and the base-rate probability of a member is its share of the training rows that have any member.
{
  "test": { "$index": { "$mod": [3, 0] } },
  "evaluate": {
    "from": "products",
    "where": { "name": { "$get": "name" } },
    "predict": "tags.$feature"
  },
  "select": ["n", "accuracy", "baseAccuracy", "meanRank"]
}

Two things are refused rather than quietly reinterpreted. .$feature on a single-valued column is a 400 that names the column and tells you to predict it directly. config.calibrate is not applied to a feature target β€” the temperature fit needs a single true value per row, and a row with several members has none β€” and the response message says so.

Available since: v2.9.0. On v2.8.4 and earlier, predict: "tags.$feature" was a 400 claiming no test row carried the field.

Per-case output

The metrics summarise; the cases let you look at the rows behind them. Three selects return per-row entries. None is included by default, so you ask for them by name:

SelectContains
casesEvery evaluated test row
accurateCasesOnly the rows whose top prediction was correct
errorCasesOnly the rows whose top prediction was wrong
{
  "test": { "$index": { "$mod": [5, 0] } },
  "evaluate": {
    "from": "products",
    "where": { "name": { "$get": "name" } },
    "predict": "category"
  },
  "select": ["accuracy", "errorCases"]
}

Each entry has the same shape. One of the errorCases from the sandbox:

{
  "offset": 5,
  "testCase": { "name": "Pirkka sugar 1 kg", "category": "106" },
  "accurate": false,
  "top": { "$value": "115", "$p": 0.155 },
  "correct": { "$value": "106", "$p": 0.023, "rank": 10 }
}
FieldMeaning
offsetThe row's position among the evaluated test rows, from 0
testCaseThe fields the query read from the row: every $get-bound field, plus the target
accurateWhether the top prediction matched the truth
topThe top prediction, as $value and $p. Absent when the query returned no candidates
correctThe truth as the model ranked it: $value, $p, and its 0-based rank. Absent when the truth was not a candidate at all β€” those rows count against coverage

For a .$feature target, testCase carries the row's members as a JSON array, and correct is the best-ranked of them.

Differences from v1 cases. The shape is close to the v1 evaluate cases, with three differences to know when porting a script:

  • testCase is not the whole row. v1 echoes every field of the test row; v2 echoes only the bound $get fields and the target. Fetch the row itself if you need its other fields.
  • top and correct carry the raw $value. For a link target that is the key, not the linked row that v1 returns.
  • correct.rank is v2-only. v1 reports the rank in aggregate (meanRank) but not per case.

Inference config inside evaluate

The evaluate object accepts the same config as a predict β€” and it is applied, never silently ignored:

{
  "test": { "$index": { "$mod": [5, 0] } },
  "evaluate": {
    "from": "purchases",
    "where": { "supplier": { "$get": "supplier" } },
    "predict": "cost_center",
    "config": { "ai": "high", "calibrate": true }
  }
}
  • config.ai selects the inference profile for every per-row predict. v2 defaults to group (group formation), so an evaluate with no config.ai already reflects the honest-confidence behaviour of a default v2 predict β€” see Inference (v2) β€” Tuning. Set "fast" to measure plain naive Bayes, or "high" for and + group.
  • config.calibrate fits a temperature on the training rows only (the fit can never see the held-out test rows) and applies it to the evaluated predictions. When applied, the response reports the fitted tau; when it cannot be applied (e.g. too little data for the fit), the response says exactly why in message. There is no silent no-op.

Time limit and partial results

Every evaluate query has a time budget, maxTime, in seconds (default 300, capped at 3600). If the budget runs out, evaluate returns the results computed so far and marks the response:

{
  "truncated": true,
  "requested": 1000,
  "message": "Evaluation hit the 300s time limit after 240 of 1000 test rows; ...",
  "n": 240,
  "accuracy": 0.71
}

Truncation is never silent β€” the truncated, requested, and message fields are always included when it happens, even if your select does not list them. If you hit the limit, narrow the test set (a smaller $mod fraction or limit) or raise maxTime.

Worked example: 5-fold cross-validation

Run a deterministic 5-fold sweep and average the accuracy across folds:

{
  "test": { "$index": { "$mod": [5, 0] } },
  "evaluate": {
    "from": "messages",
    "where": { "message": { "$get": "message" } },
    "predict": "operation"
  },
  "select": ["n", "trainSamples", "accuracy", "baseAccuracy", "meanRank", "logLoss", "ece"]
}

Repeat with [5, 1], [5, 2], [5, 3], [5, 4] and average β€” every row is tested exactly once, and no test row ever trains its own prediction. accuracy vs baseAccuracy tells you how much the message content helps; logLoss grades the probabilities and ece tells you whether the confidence is calibrated.