Automated GL coding & invoice routing (preliminary)

Task: given an invoice's text (sender, product, description) and its company, predict three things a finance team would otherwise fill in by hand:

  1. Processor โ€” which employee should handle it
  2. Acceptor โ€” who approves it
  3. GL code โ€” the accounting category (32 codes)

Processor and acceptor are company-scoped cross-table link predictions (invoice โ†’ employee). The corpus is Zipf-distributed, so invoice volume concentrates on the large companies: a test invoice's company has ~157 employees on average (the candidate pool), of which only ~9 are plausible acceptors in that company's history. That high-cardinality, company-scoped link is exactly the kind of target plain bag-of-features models handle poorly. Aito predicts each field independently from the invoice text and company alone โ€” the targets are never inputs.

Methods. Aito is measured with both engines โ€” v1 (rep1 / TableDb, /api/v1) and v2 (rep2 / CollectionDb, /api/v2) โ€” from the InvoiceRoutingEvaluation booktest. Baselines are a tuned scikit-learn Random Forest, LightGBM (with a company-local class-index encoding) and FLAML AutoML, all scored on the identical 2,000-row held-out test set with company-scoped candidate masking for the link targets; reproduction scripts and method notes live in benchmarks/invoice-routing/.

Top-1 accuracy (10k training rows, 2,000-row test)

TargetRFLightGBMFLAMLAito v1Aito v2
GL code (32 classes)48.9%35.0%65.7%66.8%68.7%
Processor (high-cardinality link)8.0%2.3%19.4%16.2%23.9%
Acceptor (link)28.6%42.4%56.2%46.2%61.4%

Provenance. dataset invoice-routing (InvoiceData synthetic, Zipf) ยท scale 10k train / 2,000 test ยท Aito from booktest InvoiceRoutingEvaluation compare-* (v1 = rep1, v2 = rep2) ยท RF / LightGBM / FLAML from the Python baselines in benchmarks/invoice-routing. All five methods score the same held-out rows โ€” the docs build fails if the baselines disagree on the hold-out size. Every figure is a build-time metric token resolved from a committed metrics.json โ€” nothing here is hand-typed.

What the Aito queries declare. The tree models receive the linked employee and GL-code columns as input features, so the Aito queries name the same evidence with basedOn: ["department", "role"] for processor and acceptor, ["department", "name"] for GL code. A GL code's name describes the invoice text it has to match, so it is evidence; an employee's name is not named, because a person's name predicting who handles an invoice is the kind of correlation a routing decision should not inherit unasked.

The hold-out was 200 rows until 2026-09-19. At that size a 95% interval near 70% accuracy is about ยฑ6.4 points โ€” wider than most differences this table reports, so the page could not resolve its own comparisons. It is now 2,000 rows (ยฑ1.6โ€“2.2 points). The larger sample first overturned one claim โ€” on 2 000 rows FLAML led processor โ€” and that loss was diagnosed to a defect in how the engine scored link candidates it had never seen (see below); with the fix v2 leads all three again.

Aito v2's 95% intervals on the enlarged hold-out:

TargetAito v295% intervalFLAML
GL code68.7%[66.6%, 70.7%]65.7%
Processor23.9%[22.1%, 25.9%]19.4%
Acceptor61.4%[59.2%, 63.5%]56.2%

Read honestly โ€” Aito v2 leads all three targets, and the margins differ:

  • GL code โ€” Aito v2 wins by ~3pp (68.7% vs FLAML's 65.7%), with Aito v1 (66.8%) level with FLAML, and the tree models beaten decisively (RF 48.9% / LightGBM 35.0%). The field-level priors carry the low-cardinality target well โ€” with no training. On the enlarged hold-out the margin is now real: v2's interval ([66.6%, 70.7%]) excludes FLAML's 65.7%.
  • Processor โ€” the hard cross-table target, where the tree models nearly collapse (RF 8.0%, LightGBM's local-index encoding 2.3%). Aito v2 (23.9%, [22.1%, 25.9%]) leads FLAML AutoML (19.4%) by ~4.6pp, with Aito v1 (16.2%) third. This row has a history worth knowing: on the 200-row hold-out v2 appeared to lead; on 2 000 rows it lost to FLAML by 2.4 points with non-overlapping intervals; and that loss turned out to be a scoring defect, not a modelling limit. An employee who had never processed an invoice was given a larger base rate than one who had processed five (the never-seen candidate's rate was taken over the company's rows, the observed one's over the whole table), and the absence of each other department was credited as evidence on top of the presence of the candidate's own. 85.7% of top-1 picks were employees with no training rows. With both corrected the same rows score 23.9% โ€” measured, since 2026-09-20, with the query naming its linked evidence. Until then the identity stage read employees.name whether a query asked for it or not, so these books declared nothing and nobody could see the difference; once that stage began obeying basedOn, a query that named the columns the tree models are handed scored 5.5 points above one that named none. v2's interval excludes FLAML's score, so the lead is resolved at this sample size โ€” and note what it costs FLAML to get there: a per-target hyperparameter search (418 s of training) against Aito's none.
  • Acceptor โ€” the clearest Aito win, and a reversal of this page's earlier baseline (where FLAML led and Aito trailed badly). Aito v2 reaches 61.4%, clearly ahead of FLAML (56.2%) and LightGBM (42.4%) with non-overlapping intervals (v2 [59.2%, 63.5%]); v1 (46.2%) still trails, which is the engine gap, not the task ceiling. From text + company alone the acceptor remains genuinely under-determined, so a ceiling below 100% is inherent โ€” but within that ceiling v2 now converts its rank advantage (below) into top-1.

What changed: the engine's scoring redesign โ€” mediation-aware link scoring plus property-based inference over the linked row's fields โ€” closed a gap where FLAML had led both link targets; the 2 000-row hold-out then exposed a remaining defect in how never-seen link candidates were scored, and its fix put processor back ahead. Aito v2 now leads acceptor and GL code and processor, all three with intervals that exclude FLAML. FLAML AutoML remains the strongest baseline โ€” closest on every row โ€” and what it costs to get there (a per-target hyperparameter search) is the training-time row.

v1 vs v2: the rank story

Mean rank of the true candidate shows where the accuracy comes from. On the link-heavy targets, v2's cross-table link priors rank the true candidate consistently higher than v1 โ€” and both engines' ranks improved further under the scoring redesign:

Targetv1 mean rankv2 mean rank
Acceptor2.11.0
Processor109.722.3
GL code3.92.6

Provenance. mean rank of the true candidate (lower is better) from the same InvoiceRoutingEvaluation compare-* snapshots. Acceptor/processor are ranked within the invoice's company (~157 employees on average); GL code within the 32 codes.

For acceptor, v2 puts the correct employee at rank 1.0 on average โ€” near the top of the ~9 plausible acceptors โ€” with v1 behind at 2.1 (both a large improvement over this page's earlier baseline, where v1 ranked the true acceptor in the mid-20s). For processor, the gap between engines is still wide: v2 reaches rank 22.3 out of the ~157-employee company pool, ~5ร— better than v1's 109.7. Earlier baselines showed this rank advantage without a top-1 advantage โ€” the signal was there but the final tie-break wasn't. The scoring redesign converted the ranks into top-1 wins; processor's mean rank says there is still headroom left on the hardest target.

The number that isn't accuracy: training time

MethodTime to absorb new data (update the model)
Aito (v1 or v2)0 s โ€” incremental; new rows are queryable immediately
Random Forest3 s (full refit)
LightGBM1145 s (full refit)
FLAML AutoML418 s (budgeted search, per target)

Provenance. trainSeconds from the Python baseline result files (benchmarks/invoice-routing/results/*.json); Aito's 0 s is architectural โ€” rows are queryable on insert with no fit step.

This is the part a top-1 table hides. Every baseline needs a retraining pipeline to absorb new invoices; Aito absorbs them on insert, with a $why explanation behind every prediction. FLAML, the closest baseline, pays ~418 s of search per target and must be re-run when the data changes โ€” and on this baseline it still trails Aito v2 on all three targets. On a workload where routings change weekly, "0 s, explainable, and at-or-above the tuned baselines" is the whole trade.

How to read this

  • Aito v2 leads all three targets โ€” with no training and full $why explanations. Acceptor is the largest win (~5pp over FLAML), then processor (~4.6pp) and GL code (~3pp); all three v2 intervals exclude FLAML. Processor was a 2.4pp loss until a link-candidate scoring defect was found through this very table and fixed.
  • The jump is the scoring redesign, not a new corpus: mediation-aware link scoring plus property-based inference over the linked row's fields is what moved the link targets (earlier baselines of this page had FLAML ahead on both). The same redesign also lifted the rank story on both engines.
  • Acceptor top-1 remains partly intrinsic: text + company under-determines who approves, so a ceiling below 100% is inherent, not just a model gap.
  • Statistics: n=2,000 test set, 95% Wilson intervals published per cell (ยฑ1.6โ€“2.2pp here), single synthetic-data seed. The hold-out was 200 rows until 2026-09-19, where the interval (ยฑ6.4pp) was wider than most differences the table reported; every comparison above is now resolvable by its own sample. One seed is still one seed โ€” read the direction as the durable signal.
  • Same corpus family, larger scale: the speed benchmark runs this same InvoiceData corpus family across a 1k โ†’ 10M scaling ladder, measuring accuracy, rank and latency at each point โ€” so if the two pages show different accuracy, it's scale and engine tuning, not a different dataset. The baselines on this page exist only at this 10k point; they have not been run at 1M or 10M, so nothing on either page claims that the tree/AutoML methods plateau while Aito keeps climbing. That cross-method comparison at scale is still to be done.

Generated. Aito's figures are interpolated at docs-build time from the InvoiceRoutingEvaluation compare-* booktest snapshots, and the RF / LightGBM / FLAML figures from the committed Python baseline result files (benchmarks/invoice-routing/results/*.json) โ€” both via metrics.json files the docs generator produces. They update whenever the evaluation is re-baselined and can't silently drift; nothing on this page is hand-typed. Multi-scale accuracy for Aito now lives on the corpus-size ladder (1k โ†’ 10M); running the baselines at those scales is still a follow-up.