Aito is a predictive database. You load your data into it like an ordinary database, and then you query the
unknown the way you query the known. You get a predicted value, a calibrated probability you can set a
threshold on, and the evidence behind it. There is no model to train.
Use Aito when
The decision repeats, and your own history knows the answer. Which GL account an invoice goes to,
who approves it, which category a ticket belongs to, which product a customer buys next. The answer
lives in your records, not in public knowledge.
You need to know how sure the answer is. You want to automate the confident cases, ask a person in
the middle, and step back when unsure. Aito's probabilities are calibrated: expected calibration error
0.048 on expense categorisation (benchmark), and 0.018 and 0.020 on the
public banking77 and CLINC150 intent sets (harness and results).
The data is relational, sparse or changing. Linked tables, text fields, new customers and new
vendors arriving every day. A correction you write counts in the next query, with no retraining step.
You need many predictions, not one. A prediction behind every field, list and search in an
application, without a model or pipeline per use case or per customer.
Matching that repeats against your own labelled history. For example, an invoice line worded the
way a product was billed before: the history is the signal.
An agent needs grounded tools. An LLM agent can call Aito for the likely answer and its confidence,
or to shortlist options before it reasons. Measured: about 70 ms per decision with no LLM tokens
(results), and a tool shortlist that cut an agent's prompt about 17-fold
while improving its pick (results).
Don't use Aito when (or not yet)
You need open-ended language understanding or generation. Use a language model. On public intent
benchmarks, an LLM with retrieval (RAG) is more accurate than Aito
(results). Aito answers faster with no LLM
tokens, and its confidence is calibrated, but it does not beat RAG on accuracy there.
You have one large, flat, stable dataset and one decision to optimise for years. A trained model
(for example LightGBM) wins on accuracy above roughly 100,000 rows of flat data
(benchmark).
Raw full-text search at very large scale is the whole job. A dedicated search engine is faster
(for example Elasticsearch, 4-5 ms vs Aito's 6-8 ms on our queries,
benchmark).
You need approximate nearest-neighbour search over a very large unfiltered vector collection. Aito's
vector search is exact: fast when filtered or at moderate size, slower than an ANN index across
everything.
Matching records with no shared history to a catalogue (current state). A search engine (BM25)
currently ranks better for this. An engine fix is in progress, so this is a current state, not a
permanent one.
The case is genuinely new, with no history behind it._predict will give a low probability, which is
honest but not useful. That is the language model's job.
Quick decision checklist
Is there history of this decision in my data? No → not Aito.
Do I need a confidence I can act on? Yes → Aito fits.
Is it one decision on big, flat, stable data? Yes → consider a trained model.
Is it about reading or writing free text? Yes → a language model, possibly with Aito as its tool.
Things to know before you build
Data separation between your customers: use a separate instance or collection per customer. Do not
rely on a customer id in a query condition, and for now do not rely on a population-restricting query
(a nested from) either, to keep customers' data apart.
Check the evidence on every answer: each result carries $p and $why. For "which X fits this
record", prefer _predict. When you rank linked items with _recommend, check that the top candidates
have supporting history before you act on them.
Interfaces: REST API, Python SDK, and SQL over the Postgres wire protocol: a SELECT subset, where
JOINs follow declared links only. A join on a plain column is refused today (SQLSTATE 0A000), so run that
in a tool like DuckDB; plain-column joins are planned.