Quality Monitoring: Measure Predictions Before and After You Ship
An automated decision is only as good as its measured accuracy on your own
data, and that accuracy changes as the data changes. _evaluate measures a
prediction the way you would run it: it holds rows out, predicts each one from
the rest, and reports accuracy, ranking and calibration against a baseline.
Because it is an ordinary API call, the same measurement can run in CI, on a
schedule, or before a threshold is changed.
The queries below measure one decision on the live v2 sandbox:
which employee processes an incoming invoice (invoices.Processor), predicted
from the invoice's Description and ProductName. The table has 101 invoices.
Figures in the responses are rounded.
A baseline measurement
Hold out every fifth invoice and predict its processor from the other rows:
{
"test": { "$index": { "$mod": [5, 0] } },
"evaluate": {
"from": "invoices",
"where": {
"Description": { "$get": "Description" },
"ProductName": { "$get": "ProductName" }
},
"predict": "Processor"
},
"select": ["n", "trainSamples", "accuracy", "baseAccuracy", "accuracyGain", "meanRank", "logLoss", "ece"]
}
{ "kind": "evaluation", "data": {
"n": 21, "trainSamples": 80, "accuracy": 0.952, "baseAccuracy": 0.476,
"accuracyGain": 0.476, "meanRank": 0.143, "logLoss": 0.453, "ece": 0.071 } }
20 of the 21 held-out invoices are routed correctly (95.2%), against 47.6% for
always answering the most common processor. meanRank is 0-based, so 0.143
means the right answer is nearly always first. Read accuracy against
baseAccuracy: a prediction that does not beat the base rate is not learning
anything.
The $get bindings copy each test row's own values into the query, so the
measurement runs exactly the query you would run in production.
Look at the errors
Aggregate numbers hide the cases that matter. errorCases lists the test rows
the prediction got wrong:
{
"test": { "$index": { "$mod": [5, 0] } },
"evaluate": {
"from": "invoices",
"where": {
"Description": { "$get": "Description" },
"ProductName": { "$get": "ProductName" }
},
"predict": "Processor"
},
"select": ["accuracy", "errorCases"]
}
{ "kind": "evaluation", "data": { "accuracy": 0.952, "errorCases": [ {
"offset": 20,
"testCase": { "Description": "Purchase of Cloud Services (AWS) for IT department",
"ProductName": "Cloud Services (AWS)", "Processor": "Evelyn Carter" },
"accurate": false,
"top": { "$value": "Carol White", "$p": 0.835 },
"correct": { "$value": "Evelyn Carter", "$p": 0.0003, "rank": 3 } } ] } }
The one miss is an AWS invoice processed by Evelyn Carter, where the
prediction said Carol White at 0.835. Both are IT Managers in employees, and
the invoice text does not say which of them handles it. That is a question for
the process owner, not a tuning problem, and it is visible only because the
case is listed.
Check that the confidence means something
If $p is used as an auto-approval threshold, a stated 90% must be right about
90% of the time. ece summarises the gap; reliability lists it per
confidence bin:
{
"test": { "$index": { "$mod": [5, 0] } },
"evaluate": {
"from": "invoices",
"where": {
"Description": { "$get": "Description" },
"ProductName": { "$get": "ProductName" }
},
"predict": "Processor"
},
"select": ["ece", "reliability"]
}
{ "kind": "evaluation", "data": { "ece": 0.071, "reliability": [
{ "bin": 0.8, "confidence": 0.835, "accuracy": 0.667, "n": 3 },
{ "bin": 0.9, "confidence": 0.945, "accuracy": 1.0, "n": 18 } ] } }
The 18 predictions made at about 94.5% confidence were all right; the three made at about 83.5% were right twice. Three rows are too few to conclude anything about that bin, and the output says how many rows each bin holds so that you can tell.
Test on newer data
A random split tells you how the prediction does on data like the past. A time split tells you how it does on data that arrived later, which is the situation in production. Hold out the invoices from July onwards:
{
"test": { "InvoiceDate": { "$gte": "2024-07-01" } },
"evaluate": {
"from": "invoices",
"where": {
"Description": { "$get": "Description" },
"ProductName": { "$get": "ProductName" }
},
"predict": "Processor"
},
"select": ["n", "trainSamples", "accuracy", "baseAccuracy", "meanRank"]
}
{ "kind": "evaluation", "data": {
"n": 33, "trainSamples": 68, "accuracy": 1.0, "baseAccuracy": 0.424, "meanRank": 0.0 } }
Trained on the 68 older invoices, the prediction routes all 33 newer ones
correctly. InvoiceDate is a String column, so $gte compares it
lexicographically, which is date order for ISO dates.
The same measurement in SQL:
SELECT * FROM evaluate('invoices', 'Processor', given => 'Description, ProductName',
test => 'InvoiceDate >= ''2024-07-01''')
Compare configurations before switching
config inside evaluate is applied to every prediction, so an A/B of two
inference profiles is two calls. With the plain naive Bayes profile:
{
"test": { "$index": { "$mod": [5, 0] } },
"evaluate": {
"from": "invoices",
"where": {
"Description": { "$get": "Description" },
"ProductName": { "$get": "ProductName" }
},
"predict": "Processor",
"config": { "ai": "fast" }
},
"select": ["n", "accuracy", "logLoss", "ece"]
}
{ "kind": "evaluation", "data": { "n": 21, "accuracy": 0.952, "logLoss": 0.448, "ece": 0.031 } }
Accuracy is the same as with the default group profile, and ece is lower
(0.031 against 0.071). On 21 test rows that difference is small evidence;
repeat it on the other folds ([5, 1] to [5, 4]) before acting on it.
Monitoring
Each call above is a read-only request, so monitoring is a scheduled job:
- run the baseline measurement on a fixed fold, and alert when
accuracyoraccuracyGainfalls below the level you accepted at launch; - run the time split with a moving cutoff, so the test set is always the most recent data;
- keep
errorCasesfrom each run, so a drop in accuracy comes with the rows that caused it.
Why Aito for this
- The measured query is the production query. There is no separate training pipeline whose behaviour can differ from what is served.
- Baselines are built in.
baseAccuracyandlogLossSkillsay whether the evidence adds anything over the base rate. - Failures are loud. An unknown metric or profile is a
400naming the alternatives, not a silently empty report.
Related: Evaluation (v2) Β· Automated GL coding Β· Inference (v2)