Support ticket triage with an agent (preliminary)
Task: a support ticket arrives. Five decisions follow: the product, the category, the priority, the resolution and the first step. A ticket counts as right only when all five are right.
Five arms answer the same tickets: Aito alone, an LLM alone, an LLM with
retrieved context (LLM + RAG), Aito's shortlist and $why handed to an LLM, and
Aito with an LLM on the steps Aito is unsure of. The LLM is gpt-5-mini
throughout. The benchmark runs live in the agent demo,
and its code and recorded runs are in
aito-agent-demo.
The held-out tickets were never loaded
The 300 tickets the arms are scored on are the newest in the log, and they were never loaded into Aito: Aito predicts on tickets it has not seen. The fixture is templated, though, so 166 of them share their exact text with a ticket that was loaded. Their rows were never loaded, but their wording was.
So the result is split by wording, and the row to read is the one with new wording: the 134 held-out tickets whose text appears nowhere in the loaded data.
All five decisions right
| Arm | New wording (n = 134) | Seen wording (n = 166) | All held-out (n = 300) |
|---|---|---|---|
| Aito only | 61.9% | 73.5% | 68.3% |
| LLM + RAG | 63.4% | 58.4% | 60.7% |
Aito shortlist + $why โ LLM | 63.4% | 62.7% | 63.0% |
| Aito + LLM on unsure steps | 59.7% | 60.8% | 60.3% |
| LLM alone | 26.9% | 18.1% | 22.0% |
On new wording, Aito and the LLM arms that get context are level. Aito only reaches 61.9% (95% interval 53.5%โ69.7%) and LLM + RAG 63.4%; a paired McNemar test on the tickets where they disagree gives p = 0.87, so this is a tie, not a win either way. The same holds for the other two arms with context.
Aito's lead on all 300 comes from the seen wording. Over all held-out tickets Aito only leads (68.3% vs 60.7%), and the split shows where that lead comes from: tickets whose exact text Aito has already seen. We publish the new-wording row as the result.
The LLM alone is far behind on either split (26.9% on new wording). It is the one arm that answers without the company's history.
What this benchmark does not show: Aito ahead of LLM + RAG on unseen text. What it does show is that a database query with no model call matches an LLM with retrieval on the same decisions, with no tokens spent on the tickets it answers.
Read honestly
- Synthetic data. The ticket log is generated, with planted effects, so it measures whether each arm finds causes that are known to be there. Real ticket logs are messier.
- One seed, one model. gpt-5-mini for every LLM arm, one generated log.
- The split came after a first look. The new-wording figures for Aito were
seen in the 2026-09-30 audit before
novel_text.pyfixed the split; the split's definition was committed before the other arms were re-scored. - No new calls. The split re-scores the recorded runs; no model or Aito call was repeated.
Provenance. Every figure on this page is a build-time token resolved from
core/docs/metrics/support-triage-novel-text.json, which is copied from
scripts/support_fixture/results/novel_text.json in
aito-agent-demo at commit
f3c3fb00 (PR #31).