Invoice line β product matching (preliminary)
Task: an ERP receives a purchase-invoice line reading
983668 Fazer Konfektyr 1kg, 24 kpl from a supplier. The catalogue holds
thousands of SKUs, a dozen of them Fazer Konfektyr variants. Which SKU is this
line?
This is matching, not classification, and the difference is what the benchmark is about:
- the answer set is a live table that gains and loses rows every week, so there is no fixed label space to train on;
- every candidate has its own attributes (name, supplier, unit of measure, price) that are themselves evidence;
- the same product is billed under a different wording by every supplier, so the line text alone often cannot separate siblings.
The industry default is a lexical index (BM25 / Elasticsearch) over the
catalogue, with a rules layer bolted on for the article numbers. Aito answers
it as a cross-table link prediction β predict: "sku" on the invoice-line
table β and that single query reads three kinds of evidence at once:
- catalogue text β what a search index reads;
- billing history β every wording this SKU has been invoiced under before, and by whom;
- the line's structured fields β supplier, unit of measure, unit price, quantity β each with its own learned lift.
Accuracy against BM25
A search baseline is only honest if it reads the same evidence. BM25 is therefore run in three strengths: over the catalogue text only; over the training wordings only (each SKU as the bag of lines it was billed on); and over both, which is the fair one and the row to compare against.
| Method | Top-1 | Top-5 |
|---|---|---|
| BM25 β catalogue text only | 64.4% | 91.2% |
| BM25 β matched training wordings only | 88.8% | 98.4% |
| BM25 β both (the fair baseline) | 89.2% | 98.8% |
Aito β one _predict query, no training | 83.6% | 95.6% |
Provenance. dataset erp-line-matching β the aito-erp-demo aurora
catalogue (generated by a published, seed-stable generator) Β· task: invoice
line β catalogue SKU, a cross-table link target Β· engine v2 (rep2 /
CollectionDb) Β· hold-out by line Β· config.ai = "and" Β·
n = 250 held-out lines for the
head-to-head. Produced by the ErpMatchingBenchmark demoQueryCost booktest
case, which refuses to publish at all if the sample is short or the machine
was loaded. Every figure here is a build-time metric token resolved from that
run's committed metrics.json β nothing on this page is hand-typed.
The catalogue-only row is the headline. A lexical index over the catalogue β the thing most matching systems actually ship β reaches 64.4%. Aito reaches 83.6% on the same lines and the same candidates. The gap is what reading the billing history is worth, and it is the largest effect on this page.
The fair row is the honest one. Give BM25 the history too β index every past wording against its SKU β and it closes that gap and passes Aito: 89.2% against Aito's 83.6%. On this row the fair lexical baseline wins, and the reason is a property of the query rather than of the engine.
The demo sends basedOn ["supplier"]. basedOn states which of a linked row's
fields a query may use as evidence, and the demo names the supplier β not the
product NAME, which is the thing an invoice line is actually matching against.
The fair BM25 arm indexes those names. So this row compares a baseline that
reads product names with an engine that has not been asked to, and until
2026-09-19 the difference was invisible: the identity stage read the names
regardless of what a query declared. It now obeys the declaration
(why), and the demo's query lost
the evidence it had never asked for.
Over the full 2000-line sweep the same corpus answers what a matched query does. Naming the product name reaches 91.5% β above the fair baseline β against 86.5% naming nothing and 85.0% for the supplier-only form the demo sends. Read the head-to-head row as a measure of the demo's query, which is under-declared for its own task, and this row as what the engine does when told what the evidence is.
At top-5 the fair baseline is ahead (98.8% vs 95.6%) β Aito converts more of its shortlist into a correct first answer, BM25 gets the right SKU into the shortlist slightly more often. If your workflow shows a human five candidates, that difference matters more than the top-1 row.
So the claim this page supports is not "Aito is a better ranker than BM25". It is: BM25 needs the history-indexing layer built, maintained and rebuilt; Aito's is the database.
Where the two disagree
Top-1 agreement against the fair BM25 baseline, over the same 250 lines:
| lines | |
|---|---|
| both right | 198 |
| Aito right, BM25 wrong | 11 |
| BM25 right, Aito wrong | 25 |
| neither | 16 |
The two methods agree on the overwhelming majority and are close overall, but they are not interchangeable: they fail on different lines. Aito's exclusive wins are lines where the wording is the supplier's own and only the history can resolve it; BM25's are lines where the wording quotes the catalogue and popularity has pulled Aito toward a better-attested sibling. The disagreeing set is small enough here that the split between those two columns is suggestive, not decisive β the regime story is what the per-band analysis in the booktest is for.
What a win and a loss look like
Counts say how often each method wins. They do not say which half of your workload lands where β so here are the actual lines, taken from the measured set in order rather than chosen, with each SKU's training billing count beside it. Warmth is the variable the whole comparison turns on, so it is printed rather than described.
Both right β the bulk of the work. The supplier quotes enough of the catalogue name that the text settles it. Neither method is doing anything clever, and on a workload made only of these lines a search index is the right tool.
| invoice line | correct SKU | Aito picked | fair BM25 picked |
|---|---|---|---|
bauhaus pora bit sarja 50mm v2 β Ranta Tukku | Bauhaus Drill Bit Set 50mm v2 (Riga) (32x billed) | correct Β· p=0.801 | correct |
Marimekko Thro 150x Natu β Hansa Trading Oy | Marimekko Throw 150x200 Natural (Tallinn) (56x billed) | correct Β· p=0.996 | correct |
maali tela 10cm musta v2 β LΓ€nsi Tukku Oy | Paint Roller 10cm Black v2 (Turku) (81x billed) | correct Β· p=0.871 | correct |
Aito right, BM25 wrong β what the history buys. The wording names a family and the catalogue holds several members; the text cannot choose between them. What chooses is who is billing, at what unit and price β fields a lexical index has no way to use, because they are not text to match against. Look at what BM25 answers instead: never something unrelated, always a near-identical variant β a different Mk, a different pack size, the same product from a different warehouse. The text genuinely does not separate them, and the fair baseline has nothing left to break the tie with.
| invoice line | correct SKU | Aito picked | fair BM25 picked |
|---|---|---|---|
tyyny 50cm valkoinen mk2 hd kouvola standard β Pohjoinen Kauppa Oy | Cushion 50cm White Mk2 HD (Kouvola) (106x billed) | correct Β· p=0.439 | Cushion 50cm White HD (Oulu) (28x billed) |
tv 55" qled v2 β Kukkatukku Salo Oy | TV 55" QLED v2 (Turku) (52x billed) | correct Β· p=0.354 | TV 55" QLED Graphite v2 (Oulu) (37x billed) |
juusto 250g mk2 4-pack β Meridian Supply | Cheese 250g Mk2 4-pack (Oulu) (35x billed) | correct Β· p=0.989 | Cheese 250g Mk2 4-pack (Turku) (88x billed) |
fazer konfektyr yogurt 6-pack β Oulun Erikoistukku | Fazer Konfektyr Yogurt 6-pack (Oulu) (61x billed) | correct Β· p=0.977 | Fazer Konfektyr Yogurt Mk3 6-pack (Turku) (84x billed) |
ripsivΓ€ri 250ml β Kukkatukku Salo Oy | Mascara 250ml (Turku) (56x billed) | correct Β· p=0.486 | Mascara 250ml (Tallinn) (33x billed) |
Aito wrong, BM25 right β where we lose. Two different failures sit in this cell, and they need different fixes. Most are warm-candidate domination: the text names a thinly-billed SKU and Aito answers with a far more familiar sibling, confidently. The rest are attribute confusions between two comparably-billed variants, where a colour or a pack size got dropped β the opposite of a popularity problem.
| invoice line | correct SKU | Aito picked | fair BM25 picked |
|---|---|---|---|
drill bit set 25mm β Ranta Tukku | Drill Bit Set 25mm (Riga) (28x billed) | Tikkurila Drill Bit Set 25mm Yellow 10pk (Riga) (40x billed) Β· p=0.218 | correct |
Bread Mk3 6-pack tallinn standard β Vellamo Wholesale | Bread Mk3 6-pack (Tallinn) (10x billed) | Valio Oy Bread 6-pack (Kouvola) (105x billed) Β· p=0.279 | correct |
verkkokauppa.com tv 55" oled musta pro β LΓ€nsi Tukku Oy | Verkkokauppa.com TV 55" OLED Black Pro (Turku) (79x billed) | TV 55" OLED Black 10pk (Riga) (43x billed) Β· p=0.406 | correct |
maito 1l mk4 6-pack β Aurinko Import Oy | Milk 1L Mk4 6-pack (Riga) (24x billed) | Milk 1L Mk4 (Tallinn) (40x billed) Β· p=0.348 | correct |
Berner Oy Stor Box 1L Mk2 β Frontier Goods Ltd | Berner Oy Storage Box 1L Mk2 (Tallinn) (58x billed) | Berner Oy Storage Box 750ml (Tallinn) (52x billed) Β· p=0.441 | correct |
Neither β the task's own floor. Truncated and abbreviated wordings over a
catalogue of near-identical variants. Both methods often land on the same
wrong SKU, which is the tell: the line does not contain what would separate
them. Lines like these are why the workflow needs a human reviewer and why
$why earns its latency β no ranking change will fix a line that does not say
which warehouse it means.
| invoice line | correct SKU | Aito picked | fair BM25 picked |
|---|---|---|---|
marimekko dress l beige β Pohjoinen Kauppa Oy | Marimekko Dress L Beige (Oulu) (63x billed) | Marimekko Dress L Beige HD (Tampere) (31x billed) Β· p=0.286 | Marimekko Dress L Beige HD (Tampere) (31x billed) |
Tikkurila Dril Bit Set 25mm Yell 10pk riga budget β Pohjanmaan Tukku Oy | Tikkurila Drill Bit Set 25mm Yellow 10pk (Riga) (40x billed) | Tikkurila Drill Bit Set 25mm Yellow HD (Oulu) (63x billed) Β· p=0.983 | Tikkurila Drill Bit Set 25mm Yellow HD (Oulu) (63x billed) |
Marimekko Cand Vani Blac tampere premium β Frontier Goods Ltd | Marimekko Candle Vanilla Black (Tampere) (32x billed) | Marimekko Candle Vanilla v2 (Tallinn) (56x billed) Β· p=0.251 | Marimekko Candle Vanilla v2 (Tallinn) (56x billed) |
Where Aito is still weak
Two things on this corpus, stated plainly.
Warm-candidate domination. There are
30
lines that the catalogue-only index β a model with no history at all β gets
right and Aito gets wrong. In most of them the text names a thinly-billed SKU
and Aito answers with a sibling it has seen ten to twenty times as often
(T-shirt S Navy Pro, billed 5 times, loses to T-shirt S Navy, billed 86 β
one row of the worked-example tables the same booktest emits, which list every
line in this cell with both billing counts).
It is a real defect and the one we are working on: a correct SKU seen five
times should not lose to a sibling seen five hundred when the text names the
first one. The remainder of that cell is a different problem β confusions
between two comparably-billed variants, where a colour or a pack size is
what got dropped, and occasionally a pair whose catalogue names are identical
apart from a warehouse the invoice line never mentions.
No history, no advantage. Every gain on this page comes from the billing history. On public entity-matching corpora where the gold pairs have no transaction history at all β two product catalogues to be reconciled, nothing ever bought β Aito has no history to read and measures well behind BM25. Those runs are not in this page's metrics pipeline and no figures are quoted for them, but the direction is not in doubt and the reason is structural, not a tuning gap. If your matching problem is a one-off reconciliation of two static lists, use a search engine. If it is a recurring flow where the same things get matched over and over and people correct the results, that is the case this page measures.
Latency
The demo's exact query β five where conditions, $why on, limit 5 β
against the full catalogue:
| regime | ms |
|---|---|
| first request, cold caches | 5788 ms |
| still warming β mean of the first 25 requests | 467 ms |
| steady state β mean of the last 25 | 295 ms |
| mean over all 250, ramp included | 342 ms |
| p50 / p90 over all 250 | 317 ms / 469 ms |
There are three regimes here and only one of them is the number to plan with. The run is one open database with one view shared by every query, over the same 250 distinct lines as the accuracy tables β so it warms as it goes, and the table says so instead of averaging it away:
- The first request pays 5788 ms while the caches fill. That is what you will see if you open a fresh database and send one query β more than an order of magnitude above the steady state, so a deployment that matters should warm up rather than let a user meet it.
- Warming takes longer than one request. The first 25 average 467 ms, against 295 ms for the last 25.
- 295 ms is the serving figure β the regime a live deployment runs in, and the one to size capacity against. The mean over all 250 is higher (342 ms) precisely because it still carries the ramp; it is reported here rather than quietly dropped.
These are single-node, single-threaded figures with $why on β asking for the
explanation is part of the cost.
Provenance. Same run as the accuracy tables, on a machine the case checks was quiet β the publication refuses above a 1-minute load threshold sampled during the latency window, because contention moves these numbers by more than anything this table reports.
For implementers
If you are building this on Aito, the shape below is the one that produced every number above.
Schema. The catalogue is one collection; the transaction collection links to it. That link is what makes this a prediction rather than a search.
PUT /api/v2/schema/products
{
"type": "collection",
"columns": {
"sku": { "type": "String" },
"name": { "type": "Text", "analyzer": "english" },
"supplier": { "type": "String" },
"category": { "type": "String" },
"unit_of_measure": { "type": "String" },
"unit_price": { "type": "Decimal" }
}
}
PUT /api/v2/schema/invoice_lines
{
"type": "collection",
"columns": {
"description": { "type": "Text", "analyzer": "english" },
"billing_supplier": { "type": "String" },
"unit_of_measure": { "type": "String" },
"unit_price_eur": { "type": "Decimal" },
"quantity": { "type": "Int" },
"sku": { "type": "String", "link": "products.sku" }
}
}
The query. Put the whole incoming line in where β the free text and the
structured fields β and predict the link column.
POST /api/v2/_predict
{
"from": "invoice_lines",
"where": {
"description": "983668 Fazer Konfektyr 1kg, 24 kpl",
"billing_supplier": "Baltic Trade House",
"unit_of_measure": "kg",
"unit_price_eur": 12.40,
"quantity": 24
},
"predict": "sku",
"config": { "ai": "and" },
"select": ["$p", "$value", "$why"],
"limit": 5
}
Four things that matter, each measured:
- Feed every field, not just the text. The structured fields carry real lift β supplier especially, because who bills you narrows the catalogue before a single token is read.
config.ai: "and"treats co-occurring tokens as joint evidence rather than independent votes. It is what makes a multi-token wording behave like a phrase.- Do not add the catalogue name as candidate evidence. The obvious tuning
move is
basedOn: ["name"], so the product's own name scores against the line. On this corpus it costs about ten points of top-1 (table below). With nobasedOn, every linked sub-feature is already used; naming one narrows the evidence rather than adding to it. - Write the corrections back. Every accepted match becomes a training row the next query reads, with no reindex and no fit step. That is the loop the whole advantage rests on β a deployment that never writes back its corrections is the catalogue-only row, not the Aito row.
Measured over the full 2000-line hold-out:
| query setting | top-1 |
|---|---|
default (no basedOn) | 86.5% |
basedOn: ["supplier"] β what the demo sends | 85.0% |
basedOn: ["name"] | 91.5% |
Warm the database before you serve from it. The first request pays 5788 ms and the ramp runs past the first two dozen β see latency and the warming strategies guide. This is a deployment detail, not a tuning one, and it is the most common reason a first trial feels slower than this page.
Ask for $why. Each hit carries the factors that produced its
probability β which tokens fired, which structured field lifted it, the base
rate it started from. It costs latency (the numbers above include it) and it is
what makes a human reviewer's correction an informed one rather than a coin
flip.
Read honestly
- The comparison that flatters us is the wrong one. Against a catalogue-only search index the margin is large; against BM25 given the same history it is small. The second is the real comparison, and the second is what the page leads with.
- The advantage is history, not ranking. Take the history away and the advantage goes with it β see above. This benchmark measures a recurring flow with feedback, which is the case Aito is for; it does not claim anything about static list reconciliation.
- Warm candidates still beat thin correct ones on most of the 30 lines a history-free index gets right and we do not. Open defect, not a tuning knob.
- The corpus is generated, one seed. It comes from the
aito-erp-demogenerator, which models real supplier behaviour β each supplier writes the same product its own way, some quote the catalogue, some don't β but it is not a customer's data, and a generator cannot surprise you the way real wording does. n = 250 for the head-to-head (the query-setting table runs the full 2000-line hold-out). At that size a point or two is inside the noise; read the direction and the size of the catalogue-only gap, not the last digit. - The numbers are produced, not typed. The booktest behind them refuses to publish a short run, or a run on a loaded machine, so a truncated diagnostic pass cannot quietly overwrite what this page shows.