Automated GL coding & invoice routing (preliminary)
Task: given an invoice's text (sender, product, description) and its company, predict three things a finance team would otherwise fill in by hand:
- Processor โ which employee should handle it
- Acceptor โ who approves it
- GL code โ the accounting category (32 codes)
Processor and acceptor are company-scoped cross-table link predictions (invoice โ employee). The corpus is Zipf-distributed, so invoice volume concentrates on the large companies: a test invoice's company has ~157 employees on average (the candidate pool), of which only ~9 are plausible acceptors in that company's history. That high-cardinality, company-scoped link is exactly the kind of target plain bag-of-features models handle poorly. Aito predicts each field independently from the invoice text and company alone โ the targets are never inputs.
Methods. Aito is measured with both engines โ v1 (rep1 / TableDb,
/api/v1) and v2 (rep2 / CollectionDb, /api/v2) โ from the
InvoiceRoutingEvaluation booktest. Baselines are a tuned scikit-learn Random
Forest, LightGBM (with a company-local class-index encoding) and FLAML
AutoML, all scored on the identical 2,000-row held-out test set with
company-scoped candidate masking for the link targets; reproduction scripts and
method notes live in benchmarks/invoice-routing/.
Top-1 accuracy (10k training rows, 2,000-row test)
| Target | RF | LightGBM | FLAML | Aito v1 | Aito v2 |
|---|---|---|---|---|---|
| GL code (32 classes) | 48.9% | 35.0% | 65.7% | 66.8% | 68.7% |
| Processor (high-cardinality link) | 8.0% | 2.3% | 19.4% | 16.2% | 23.9% |
| Acceptor (link) | 28.6% | 42.4% | 56.2% | 46.2% | 61.4% |
Provenance. dataset invoice-routing (InvoiceData synthetic, Zipf) ยท scale
10k train / 2,000 test ยท
Aito from booktest InvoiceRoutingEvaluation compare-* (v1 = rep1,
v2 = rep2) ยท RF / LightGBM / FLAML from the Python baselines in
benchmarks/invoice-routing. All five methods score the same held-out
rows โ the docs build fails if the baselines disagree on the hold-out size.
Every figure is a build-time metric token resolved from a committed
metrics.json โ nothing here is hand-typed.
What the Aito queries declare. The tree models receive the linked
employee and GL-code columns as input features, so the Aito queries name the
same evidence with basedOn: ["department", "role"] for processor and
acceptor, ["department", "name"] for GL code. A GL code's name describes
the invoice text it has to match, so it is evidence; an employee's name is
not named, because a person's name predicting who handles an invoice is the
kind of correlation a routing decision should not inherit unasked.
The hold-out was 200 rows until 2026-09-19. At that size a 95% interval near 70% accuracy is about ยฑ6.4 points โ wider than most differences this table reports, so the page could not resolve its own comparisons. It is now 2,000 rows (ยฑ1.6โ2.2 points). The larger sample first overturned one claim โ on 2 000 rows FLAML led processor โ and that loss was diagnosed to a defect in how the engine scored link candidates it had never seen (see below); with the fix v2 leads all three again.
Aito v2's 95% intervals on the enlarged hold-out:
| Target | Aito v2 | 95% interval | FLAML |
|---|---|---|---|
| GL code | 68.7% | [66.6%, 70.7%] | 65.7% |
| Processor | 23.9% | [22.1%, 25.9%] | 19.4% |
| Acceptor | 61.4% | [59.2%, 63.5%] | 56.2% |
Read honestly โ Aito v2 leads all three targets, and the margins differ:
- GL code โ Aito v2 wins by ~3pp (68.7% vs FLAML's 65.7%), with Aito v1 (66.8%) level with FLAML, and the tree models beaten decisively (RF 48.9% / LightGBM 35.0%). The field-level priors carry the low-cardinality target well โ with no training. On the enlarged hold-out the margin is now real: v2's interval ([66.6%, 70.7%]) excludes FLAML's 65.7%.
- Processor โ the hard cross-table target, where the tree models nearly
collapse (RF 8.0%,
LightGBM's local-index encoding
2.3%).
Aito v2 (23.9%, [22.1%, 25.9%])
leads FLAML AutoML
(19.4%) by
~4.6pp, with
Aito v1 (16.2%) third. This
row has a history worth knowing: on the 200-row hold-out v2 appeared to lead;
on 2 000 rows it lost to FLAML by 2.4 points with non-overlapping intervals;
and that loss turned out to be a scoring defect, not a modelling limit. An
employee who had never processed an invoice was given a larger base rate
than one who had processed five (the never-seen candidate's rate was taken
over the company's rows, the observed one's over the whole table), and the
absence of each other department was credited as evidence on top of the
presence of the candidate's own. 85.7% of top-1 picks were employees with no
training rows. With both corrected the same rows score
23.9% โ measured, since
2026-09-20, with the query naming its linked evidence. Until then the
identity stage read
employees.namewhether a query asked for it or not, so these books declared nothing and nobody could see the difference; once that stage began obeyingbasedOn, a query that named the columns the tree models are handed scored 5.5 points above one that named none. v2's interval excludes FLAML's score, so the lead is resolved at this sample size โ and note what it costs FLAML to get there: a per-target hyperparameter search (418 s of training) against Aito's none. - Acceptor โ the clearest Aito win, and a reversal of this page's earlier baseline (where FLAML led and Aito trailed badly). Aito v2 reaches 61.4%, clearly ahead of FLAML (56.2%) and LightGBM (42.4%) with non-overlapping intervals (v2 [59.2%, 63.5%]); v1 (46.2%) still trails, which is the engine gap, not the task ceiling. From text + company alone the acceptor remains genuinely under-determined, so a ceiling below 100% is inherent โ but within that ceiling v2 now converts its rank advantage (below) into top-1.
What changed: the engine's scoring redesign โ mediation-aware link scoring plus property-based inference over the linked row's fields โ closed a gap where FLAML had led both link targets; the 2 000-row hold-out then exposed a remaining defect in how never-seen link candidates were scored, and its fix put processor back ahead. Aito v2 now leads acceptor and GL code and processor, all three with intervals that exclude FLAML. FLAML AutoML remains the strongest baseline โ closest on every row โ and what it costs to get there (a per-target hyperparameter search) is the training-time row.
v1 vs v2: the rank story
Mean rank of the true candidate shows where the accuracy comes from. On the link-heavy targets, v2's cross-table link priors rank the true candidate consistently higher than v1 โ and both engines' ranks improved further under the scoring redesign:
| Target | v1 mean rank | v2 mean rank |
|---|---|---|
| Acceptor | 2.1 | 1.0 |
| Processor | 109.7 | 22.3 |
| GL code | 3.9 | 2.6 |
Provenance. mean rank of the true candidate (lower is better) from the
same InvoiceRoutingEvaluation compare-* snapshots. Acceptor/processor are
ranked within the invoice's company (~157 employees on average); GL code within
the 32 codes.
For acceptor, v2 puts the correct employee at rank 1.0 on average โ near the top of the ~9 plausible acceptors โ with v1 behind at 2.1 (both a large improvement over this page's earlier baseline, where v1 ranked the true acceptor in the mid-20s). For processor, the gap between engines is still wide: v2 reaches rank 22.3 out of the ~157-employee company pool, ~5ร better than v1's 109.7. Earlier baselines showed this rank advantage without a top-1 advantage โ the signal was there but the final tie-break wasn't. The scoring redesign converted the ranks into top-1 wins; processor's mean rank says there is still headroom left on the hardest target.
The number that isn't accuracy: training time
| Method | Time to absorb new data (update the model) |
|---|---|
| Aito (v1 or v2) | 0 s โ incremental; new rows are queryable immediately |
| Random Forest | 3 s (full refit) |
| LightGBM | 1145 s (full refit) |
| FLAML AutoML | 418 s (budgeted search, per target) |
Provenance. trainSeconds from the Python baseline result files
(benchmarks/invoice-routing/results/*.json); Aito's 0 s is architectural โ
rows are queryable on insert with no fit step.
This is the part a top-1 table hides. Every baseline needs a retraining
pipeline to absorb new invoices; Aito absorbs them on insert, with a $why
explanation behind every prediction. FLAML, the closest baseline, pays
~418 s of search
per target and must be re-run when the data changes โ and on this baseline
it still trails Aito v2 on all three targets. On a workload where routings
change weekly, "0 s, explainable, and at-or-above the tuned baselines" is the
whole trade.
How to read this
- Aito v2 leads all three targets โ with no training and full
$whyexplanations. Acceptor is the largest win (~5pp over FLAML), then processor (~4.6pp) and GL code (~3pp); all three v2 intervals exclude FLAML. Processor was a 2.4pp loss until a link-candidate scoring defect was found through this very table and fixed. - The jump is the scoring redesign, not a new corpus: mediation-aware link scoring plus property-based inference over the linked row's fields is what moved the link targets (earlier baselines of this page had FLAML ahead on both). The same redesign also lifted the rank story on both engines.
- Acceptor top-1 remains partly intrinsic: text + company under-determines who approves, so a ceiling below 100% is inherent, not just a model gap.
- Statistics: n=2,000 test set, 95% Wilson intervals published per cell (ยฑ1.6โ2.2pp here), single synthetic-data seed. The hold-out was 200 rows until 2026-09-19, where the interval (ยฑ6.4pp) was wider than most differences the table reported; every comparison above is now resolvable by its own sample. One seed is still one seed โ read the direction as the durable signal.
- Same corpus family, larger scale: the speed benchmark runs this same InvoiceData corpus family across a 1k โ 10M scaling ladder, measuring accuracy, rank and latency at each point โ so if the two pages show different accuracy, it's scale and engine tuning, not a different dataset. The baselines on this page exist only at this 10k point; they have not been run at 1M or 10M, so nothing on either page claims that the tree/AutoML methods plateau while Aito keeps climbing. That cross-method comparison at scale is still to be done.
Generated. Aito's figures are interpolated at docs-build time from the
InvoiceRoutingEvaluation compare-* booktest snapshots, and the RF /
LightGBM / FLAML figures from the committed Python baseline result files
(benchmarks/invoice-routing/results/*.json) โ both via metrics.json files
the docs generator produces. They update whenever the evaluation is
re-baselined and can't silently drift; nothing on this page is hand-typed.
Multi-scale accuracy for Aito now lives on the
corpus-size ladder
(1k โ 10M); running the baselines at those scales is still a follow-up.