Expense categorization (preliminary)

Task: given a purchase invoice β€” its vendor, tax rate, amount, line count and free-text line description β€” predict the expense category (GL account) it should be booked to (48 categories). This is the everyday accounts-payable automation problem: every finance team codes invoices to a chart of accounts, and most of the work is repetitive.

The realistic wrinkle is that the account is determined by what was bought, not who sent it. Big generalist vendors sell across many accounts, so a model that just memorizes "vendor β†’ its most common category" barely clears the ~20% vendor-lookup oracle. Doing better means reading the line-item text β€” which identifies the invoice type, and the type determines the account β€” modelling the residual ambiguity, and, on top of irreducible label noise, saying how sure it is.

The data. A synthetic but realistically-structured corpus (benchmarks/gen_corpus.py, seed 41): 10,000 training invoices, a 2,000-row held-out test set, 48 categories, a ~3% base rate (most-common class), and 10% label noise (so ~90% is the accuracy ceiling and calibration is meaningful). Categories are emitted from a latent invoice-type topic model: each type sprays a bag of description keywords, so the signal accumulates across the text rather than sitting in any single feature. One corpus is consumed identically by every method β€” Aito reads the normalized tables and follows the vendor β†’ industry link itself; the baselines get the same join denormalized into a flat feature table β€” so the comparison is fair by construction.

Methods. Aito with both engines (v1 / rep1 and v2 / rep2), a tuned scikit-learn Random Forest, LightGBM, and FLAML AutoML β€” all scored on the identical test set with identical metric definitions. Every method emits the same accuracy + calibration + speed schema; reproduction lives in benchmarks/. An LLM + RAG baseline (gpt-5-mini, retrieval over invoice history) is a slot that still needs re-running on this corpus β€” see below.

Top-1 accuracy, top-3, and rank

MethodTop-1Top-3Mean rank of true class
Aito v179.8%90.1%3.3
Aito v279.8%89.8%3.4
Random Forest73.7%90.6%3.4
LightGBM76.1%89.8%3.4
FLAML AutoML70.2%89.3%3.5
LLM (gpt-5-mini + RAG)β€”β€”β€”

Vendor-lookup oracle β‰ˆ 20% top-1; base rate β‰ˆ 3%; 10% label noise caps accuracy near 90%. 2,000 held-out rows, which puts a 95% interval of about Β±2 points on each cell. (The LLM row is a pending slot β€” see below.)

Provenance. dataset expense-categorization (benchmarks/gen_corpus.py, seed 41) Β· scale 10,000 train / 2,000 test, 48 categories Β· Aito from booktest ExpenseCategorizationEvaluation (v1 = rep1, v2 = rep2) Β· RF / LightGBM / FLAML AutoML from the Python baselines in benchmarks/baselines. Every result figure is a build-time metric token resolved from a committed metrics.json β€” nothing in the tables is hand-typed.

Read honestly β€” and on this hold-out, top-1 does separate the methods:

  • Both Aito engines lead, and the lead is measurable. Every cell below is scored on the same rows, so the comparison is paired: what decides it is the rows where two methods disagree, not whether their intervals overlap. Against LightGBM, the strongest baseline, Aito v1 wins 185 rows that LightGBM misses and loses 110 (+3.7 pt, p < 0.0001); against Random Forest it is +6.2 pt (p < 0.0001) and against FLAML AutoML +9.7 pt (p < 0.0001). Every method lands 3–4Γ— above the ~20% vendor oracle β€” they are all reading the line-item text, which is where the account signal lives β€” but the two that read it through learned per-feature evidence read it better.
  • The two Aito engines are a tie here, and we say so. v1 and v2 differ by +0.0 pt, on 49 rows against 49: p = 1, which is a tie at this sample size. An earlier revision of this page showed v2 four points behind v1; that gap was a 200-row sample, not an engine difference, and it disappears on a hold-out ten times the size.
  • Aito v1 has the best mean rank (3.3), with Aito v2 second (3.4), but RF keeps the best top-3 (90.6% vs Aito v1's 90.1%): where the ensemble is unsure it still spreads plausible mass across nearby categories.
  • LightGBM stays close on top-1 but with confidence it can't back up β€” that is the calibration story below.

Calibration: does the confidence mean anything?

Accuracy alone hides whether a model knows when it's unsure β€” which is what a finance team needs to decide "auto-post" vs "route to a human". We report the expected calibration error (ECE, lower is better), the top-1 Brier score (lower is better), and the calibration gap (mean confidence βˆ’ accuracy; >0 means overconfident, <0 underconfident).

MethodECEBrierCalibration gap
Aito v10.0610.1420.054
Aito v20.0480.1410.009
Random Forest0.2510.242-0.251
LightGBM0.1750.1910.175
FLAML AutoML0.0810.1880.067
LLM (gpt-5-mini + RAG)β€”β€”β€”

This is where the methods separate:

  • Aito is well-calibrated straight out of the model β€” no temperature scaling or post-hoc step. v2 has the smallest calibration gap of any method (gap 0.009, ECE 0.048); v1 runs mildly hot (gap 0.054, ECE 0.061) while tying v2 on top-1.
  • Aito v2 also has the lowest ECE and Brier score of any method (ECE 0.048, Brier 0.141). FLAML AutoML is the best-calibrated baseline (ECE 0.081, gap 0.067) β€” but it only gets there after a budgeted per-target hyperparameter search (see training time); the tree defaults don't.
  • Random Forest is badly underconfident (gap -0.251, the highest ECE at 0.251): its averaged-tree votes hedge, so it under-claims on the calls it gets right.
  • LightGBM is overconfident (gap 0.175, ECE 0.175): its softmax outputs are sharply peaked β€” mean confidence 93.6% β€” so it asserts high confidence on predictions that are right less often. On a task with real ambiguity, that is the dangerous failure mode.

Why it matters: AP automation runs on a confidence threshold β€” auto-post above X, route the rest to a human. A well-calibrated model auto-posts its high-confidence slice safely; an overconfident model auto-posts its confident-but-wrong predictions straight into the ledger. Calibration, not raw top-1, is what makes the accuracy usable β€” and it is where Aito (with no training step) and AutoML (with a long one) pull ahead of the tree defaults.

The number that isn't accuracy: training time

MethodTime to absorb new data (update the model)
Aito (v1 or v2)0 s β€” incremental; new rows are queryable immediately
Random Forest2 s (full refit)
LightGBM221 s (full refit)
FLAML AutoML83 s (budgeted search)

Every tree/AutoML baseline needs a retraining pipeline to absorb new invoices; Aito absorbs them on insert, with a $why explanation behind every prediction. On AP coding β€” where a new vendor or a re-mapped account shows up weekly β€” "0 s and explainable" is a large part of the real total cost. Note the best-calibrated baseline, AutoML, is also the most expensive to keep current.

Predict latency

The mirror image of training time. Aito does its learning at predict time (no model to pre-fit), so a single prediction costs tens of milliseconds; the tree models, once trained, infer in well under a millisecond.

MethodPredict (ms, mean)p90 (ms)Training
Aito v117.424.40 s
Aito v221.431.30 s
Random Forest0.080.082 s
LightGBM0.240.24221 s
FLAML AutoML0.010.0183 s
LLM (gpt-5-mini + RAG)β€”β€”β€”

So the honest read is: for a single field, a pre-trained tree serves faster β€” Aito's predict-time inference is tens of ms. What Aito buys with that is no training pipeline and any field predictable on demand β€” which is exactly what the next number is about.

Coding a whole invoice: multi-target latency (rep1 vs rep2)

Real AP coding isn't one prediction β€” a clerk codes several fields per invoice: the expense category, a cost center, an approver, a handler. A tree stack needs a separately-trained model per field; Aito predicts all four from the same data with no per-target training. This is where the per-invoice latency β€” and the rep1 vs rep2 engine cost β€” actually lands. approver/handler are the interesting ones: org-scoped, high-cardinality employee links, exactly the cross-table target rep2's link priors target.

FieldKindrep1 top-1rep2 top-1rep1 (ms)rep2 (ms)
categorycategory link79.8%79.8%14 ms16 ms
cost_centercategorical50.5%45.3%10 ms13 ms
approveremployee link (org-scoped)61.0%57.6%13 ms27 ms
handleremployee link (org-scoped)68.3%81.6%12 ms27 ms
total / invoiceall four fields49 ms83 ms

Handler moved in v2.11.0. Before this release, when one table had two links to the same target β€” approver and handler, both to employees β€” a prediction through one link could read the other link's cached rows. With that fixed, the v2 engine's handler accuracy rose by about 29 points. A leakage check with the employee attributes shuffled within each organisation falls back to 61%, so the gain is the attributes' real signal, not a leak.

The total is the per-invoice cost of coding all four fields on one engine, with 0 s training either way. On the org-scoped employee links, rep2's link-prior machinery earns its extra latency on handler: it lifts it to 81.6% (rep2) from 68.3% (rep1), while approver stays roughly level (57.6% rep2 vs 61.0% rep1) β€” see the GL-coding benchmark for the rank story at depth.

How much data must I label?

The training-size sweep that used to sit here is now the section's main scaling page, because it is the one with external baselines in it:

β†’ How much data do I need? β€” top-1 by training-set size from 1 to 1 000 000 rows, this same corpus, against Random Forest, LightGBM and FLAML AutoML. At N=1 000 Aito v2 leads the field by roughly seven points with no training step; by N=100 000 LightGBM is ahead. The neighbouring question β€” how the engine behaves as the database grows to 10M rows β€” is the capacity ladder.

LLM baseline β€” pending on this corpus

The realistic LLM deployment for this task is retrieval-augmented: for each invoice, retrieve the most similar historical invoices (same vendor, similar text) and put them in the prompt, so gpt-5-mini copies the account from the vendor's retrieved history rather than reasoning about opaque category codes. A raw few-shot prompt (no retrieval) has nothing to reason about and scores near the floor; RAG is what makes an LLM competitive here.

A prior RAG run reached parity with the tree/Aito cluster on the previous corpus vintage, but it has not been re-run on the current line-item corpus (it needs Azure credentials and costs ~15 s per prediction), so the LLM column reads β€” above. The archived prior result is kept at benchmarks/results/stale-oldcorpus/llm.json; benchmarks/baselines/run_llm.py --retrieval rag regenerates it against this corpus.

How to read this

  • Aito v1 and v2 tie at the top of the top-1 table (79.8% each), and both lead every baseline β€” LightGBM, RF and AutoML β€” by a paired margin that separates at this sample size (p < 0.0001 against LightGBM, the closest). The tree defaults pay in calibration (RF underconfident, LightGBM overconfident); Aito is best or near-best on calibration with 0 s training and a $why behind every prediction; AutoML matches Aito's calibration only after a per-target search.
  • Statistics: n=2 000 test (95% CI β‰ˆ Β±2pp), single synthetic seed. Read the direction β€” which method handles the ambiguity and stays calibrated β€” as the durable signal; treat individual cells as point estimates.
  • The baselines are given the vendorβ†’industry join for free, so this isolates modelling + calibration + training cost, not who saw which column.
  • We still report where Aito trails: RF keeps the best top-3, and a pre-trained tree serves a single field orders of magnitude faster than Aito's predict-time inference.

Generated. Every figure in the comparison tables is interpolated at docs-build time from benchmarks/results/*.json β€” Aito's from the ExpenseCategorizationEvaluation booktest, the baselines' from the Python run_*.py scripts β€” merged into a metrics.json by the docs generator. They update whenever the evaluation is re-baselined and can't silently drift. Dataset facts (48 classes, ~3% base rate, 10% label noise) are stated from the corpus generator; an unfilled method slot (the LLM here) renders as β€”. Nothing in the tables is hand-typed.