What's new
What changed in each Aito release, from your side: new capabilities, improvements and fixes. Written per release โ not a commit log.
Which build am I on? GET /version on your instance reports the build it
is running. Docs for a specific release live at /docs/api/<version>/; this
page and /docs/api/ describe the current release, and /docs/api/edge/
tracks development ahead of it โ so a feature listed here is available on the
current release unless an entry says otherwise.
v2.11.4
A correctness release for tables with deletes and multi-tenant queries. A
search or listing after a delete shows exactly the live rows, deletes and
upserts write exactly the rows they matched (also under concurrent writes),
concurrent writes no longer fail on another commit's cleanup, a nested from
keeps its population exact, and filters through several links restrict a
linked-row prediction. Relate honours orderBy, and SQL honours
NULLS FIRST/LAST.
New
- Count links directly.
$lengthover a bare$refspath counts the referring rows (a node's degree), andorderBycan sort by a$length. See graphs.
Improved
- More accurate matching of text pairs. When none of a text pair's calibration fits is acceptable, predictions no longer fall back to the most confident curve, which made entity matching between differently worded catalogues much less accurate.
- Relate honours
orderBy: hits sort by the named member (a bare key sorts descending), and a sorted answer whose candidate cap dropped values carries arelate.order_by_partialwarning. Ties among equal-score relate hits order deterministically by value. See query reference. - Reading rows of a table with deletes not yet optimized is fast again: a per-row read no longer slows down with every delete. On 1M rows it took ~65 ms per row after 100 deletes, and ran out of memory after ~1,000; it is now a few microseconds whatever the number of deletes.
- Joins through links are much faster on the first query: it no longer
builds the whole link index. A cold
select โฆ limit 5through a link at 10M rows goes from ~12-23 s to under 0.5 s, and the first filter on a linked field from ~22 s to ~3 s. A linked field's lookup is built once instead of twice, and over a declared link it reuses the lookup built when the data was written. - A nested
fromis faster on repeated queries: its population is built as a table of the kept rows and cached by content. See multi-tenant. - SQL
information_schemaanswers as relations:WHERE(=,<>,LIKE/ILIKE,IN,NOT IN,IS NULL),ORDER BYandLIMIT/OFFSETare honoured, views are reported asVIEW,data_typeis a type name, and the PostgreSQL standard columns are present, so tools such as dbt can read the catalog. Some query shapes still answer wrong; see Known issues. See SQL. x-knowledgecolumns (experimental): every$whyfactor is a proposition you can send back as awhere($phrase,$symbol), a word outside its phrases is read on its own rows, and phrase learning is faster. See schema design.
Changed
- A nested
fromthat selects no rows no longer fails with400 request.invalid:_query,_search,_recommend,_matchand relate$patternsanswer200with no hits and aquery.empty_populationwarning, and_predictand_evaluateanswer400 query.empty_population. See common errors. - A relate
selectthat names no relate member is a400, and anorderByrelate cannot honour is a warning rather than an error. - A filter through two or more links that cannot be evaluated on a
linked-row prediction now answers
400naming the path, instead of being ignored.
Fixed
- A nested
fromgives the same answers as a table holding only its rows: a$refswalk no longer omits population rows, evidence from several fields no longer drifts with other tenants' rows, and a link-path condition in itswhere(such as"company_id.industry": "accounting") no longer fails with400. - After deleting rows, a search or listing shows exactly the live rows, also
before the table is optimized: no deleted row is listed, the last row is no
longer left out, and a
whereon it finds it. - Concurrent deletes and updates on one table write exactly the rows they
matched: v2
_delete, SQLDELETE,UPDATEandINSERT โฆ ON CONFLICT DO UPDATE SET. Before, with another delete in flight, they could delete or overwrite the neighbouring rows. A write affects the rows that matched when it was requested and still match when it commits (PostgreSQL's rule for keyed tables), and the reported count is the rows actually written. A large write that keeps losing to concurrent writes, or a pgwireCOMMITthat cannot be replayed, now fails fast with SQLSTATE40001so the client retries, instead of stalling the table. - Under concurrent writes, a write could fail with 'physical file does not
exist' when another commit's cleanup removed its not-yet-committed files.
Cleanup now keeps unreferenced files for 30 minutes before removing them, so
the disk space of deleted or replaced data is reclaimed up to 30 minutes later.
Operators can set the window with
-Daito.gc.graceMillis=<ms>. - An
on_conflict=updateupsert after a delete replaces its own row and keeps its neighbours: v2batchand keyedimportwith?on_conflict=update, and SQLINSERT โฆ ON CONFLICT (key) DO UPDATE. Before, it could delete the next row and keep the upserted row's old version. - A filter through two or more links restricts a linked-row prediction
(
_predict,_recommend,_match), for examplewhere: {"product.supplier.country": "FI"}. - A nested
fromkeeps its population exact withAITO_WRITE_DEFER_MERGEenabled,$patternsover a nestedfromcounts only the population's rows, and_evaluateover a nestedfromtrains on every remaining row after a delete. - SQL:
ORDER BY โฆ NULLS FIRST | LASTis honoured; the NULL group sorts as in PostgreSQL (last ascending, first descending);GROUP BYor a value-ordered get over a nullable column no longer fails;evaluate(โฆ, where => c)evaluates over the populationc, as documented; a catalogLIKEhonours the backslash escape. See SQL. - A view can copy a
Vectorcolumn. See multi-tenant. - Relate no longer offers another row's value on a nullable column or on a collection with deleted rows not yet optimized.
Known issues
- A long write (a large
_delete, e.g. 100k rows, or a ~1,000-row keyed upsert into a large table) on a table that receives steady concurrent inserts can fail with HTTP 500 after exhausting its commit retries. Workaround: pause writes to that table, or split the operation into smaller batches. - On collections, selecting or ordering by a linked row's field (
select,orderBy) can fail with HTTP 500 ('has no resolvable link') or show another row's value when the link's target key does not exist, the target row was deleted (until the table is optimized), or the field is empty (null).$vectorSimilaritythrough a link to a row with an empty vector can score with another row's vector, which changes sorting. Prediction probabilities and explanations are unaffected; stored data is unaffected. Fixed in the next release. - After deleting rows, and until the table is optimized,
_relatestatistics (including$patterns) can be computed incorrectly. Workaround: run optimize on the table after deletes before relying on_relate. - Some SQL
information_schemaquery shapes (LIKE โฆ ESCAPE,||concatenation, quoted or catalog-qualified table names, joins betweeninformation_schemaviews) can return empty or wrong results instead of an error. Fixed in the next release. - After deleting the last rows of a table, and until it is optimized,
_relateon a Set field, annlp-analyzed text field or a linked field (field.attribute) can fail with an error (reported as 400 or 500). Workaround: run optimize on the table after deleting.
v2.11.3
A correctness release. Filters on classic tables keep every condition,
predictions on collections read the right rows when a column holds nulls or
rows were deleted, and a nested from keeps a tenant's population exact after
a delete. You can now evaluate and match within one tenant, and /import
recognises free text and accepts CSV.
New
- Evaluate and match within one population.
_evaluateand_matchaccept a nestedfrom(for example{"from": "invoices", "where": {"tenant": "A"}}): test rows are drawn from that population and the model trains on the rest of it. See evaluation and multi-tenant. /importrecognises free text and accepts CSV. A string column that reads like prose is inferred asText(short codes, labels and ids stayString),content-type: text/csvwith a header row is accepted, and the response warns about prose in a column you declaredString. See collections.
Improved
- Every field you list in
basedOncontributes evidence. Relation discovery keeps a share for each declared field, so one field with many relations no longer crowds another out of the scoring. See improving inference.
Changed
_estimate,_aggregateand_similarityrefuse a nestedfromwith501and name the limitation (they answered400 missing 'from'). Their statistics are not restricted to a population, so they refuse rather than answer over every row. See multi-tenant.
Fixed
- Filters on classic tables keep every condition. On a classic table, a
v2 query whose
wherecombined an equality with a comparison in one object ({"tag": "blue", "n": {"$lt": 0}}) ignored the comparison; and a comparison could answer with the result of an earlier comparison on a different column of the same table. Both now apply exactly. _evaluatewith per-row comparison thresholds ($getinside$lt/$gtโฆ) on classic tables now evaluates each row against its own threshold (previously every row used the first row's), no longer warns that thewherehas no$getbinding, and reports a baseline without the per-row condition. See evaluation.- Predictions on collections with nullable columns or pending deletes read the
right rows. On a collection, a nullable column that holds nulls, or a table
with deleted rows not yet optimized, could feed predict,
$whyand text (BM25) scoring with another row's value for that column, so per-row statistics were computed from shifted values. Affected: collections ("type": "collection") with nullable columns containing nulls since v2.2.0, and collections with deleted-but-not-optimized rows since v2.1.47. Classic tables are not affected. No action is needed after upgrading. - A nested
fromafter a delete holds exactly its population. On a collection with recent deletes, a nestedfromcould return the wrong rows, and_evaluatecould hold out rows that had been deleted. - A prediction offers only values that live rows hold. A value held only by
deleted rows (until the collection was optimized), or only by rows outside a
nested
from, could be offered as a candidate;_evaluateno longer offers a label that occurs only in its held-out rows. - Base rates after a delete count the rows that remain, however the collection is split into segments.
- Scoping a link-target prediction with a
whereon the target's attribute (for example"approver.customer_id": X) no longer reverses the ranking of candidates the evidence has no discovered relation to. This was the remaining case noted in v2.11.2.
Known issues
- Inside a nested from, a $refs walk can omit population rows, and two-feature evidence can shift probabilities slightly. Single-feature and plain where conditions are exact. Per-tenant collections are unaffected.
- A link-path condition in the where of a query over a nested from returns 400 on collections.
- After deleting rows, and until the table is optimized, a search without
selectcan leave out the table's last row (awhereon that row may return nothing), and a listing with nowherecan still show the deleted row. Rows are not shown to the wrong tenant, and values are not mixed between rows. Stored data is unaffected. Workaround: add an explicitselectto the query, or run optimize on the table. Fixed in a following release. - On collections, a prediction's
selectcan show a value for a linked row's empty (null) field, and$vectorSimilaritythrough a link to a row with an empty vector can score with another row's vector, which changes sorting. Prediction probabilities and explanations are unaffected; stored data is unaffected. Fixed in a following release. - With
AITO_WRITE_DEFER_MERGEenabled (off by default), a nestedfromwithoutselectcan return rows outside itswhere, on any table, with or without deletes. Workaround: add an explicitselect, or put the filter in a top-levelwhere. Fixed in the next release. - SQL:
NULLS FIRST/NULLS LASTin ORDER BY is ignored (the default null ordering is applied). Ininformation_schema, ORDER BY and IN filters oncolumnsare ignored, views are reported with table_type 'TABLE', andcolumns.data_typereturns a numeric JDBC type code instead of a type name. Fixed in a following release. _evaluateover a nestedfrom(new in this release) can leave one row out of training after rows are deleted from the table, until the table is optimized; the reported accuracy can then differ slightly. Workaround: optimize the table before evaluating. Fixed in the next release._relatewith$patternsover a nestedfromcan compute its statistics over the whole table, including other tenants' rows, rather than the nested population, once the table has been optimized or with AITO_WRITE_DEFER_MERGE enabled. Keep each tenant's data in its own table or instance. Fixed in a following release.- After deleting rows, and until the table is optimized,
_relatestatistics (including$patterns) can be computed incorrectly and may still include the deleted rows. Workaround: run optimize on the table after deletes before relying on_relate. - Concurrent
_deleterequests on the same table can delete rows other than the requested ones and leave the requested rows in place. Concurrent SQLDELETE/UPDATEstatements over the PostgreSQL interface share the same pattern. Until fixed, send deletes and SQL updates to a table one at a time. - After deleting rows from a table, an
on_conflict=updateupsert (batch, import, or SQLINSERT โฆ ON CONFLICT DO UPDATEover the PostgreSQL interface) can delete a neighbouring row and keep the old version of the upserted row. Workaround: run optimize on the table after deleting and before upserting. Fixed in the next release. - A
_predict,_recommendor_matchthat predicts a linked row and filters through two or more links (e.g.where: {"product.supplier.country": "FI"}) can ignore that filter and answer over all previously seen candidates, without an error. Filters through a single link are not affected. Fixed in the next release.
Docs
- A new multi-tenant guide answers, per query type, whether a query scoped to one tenant can be influenced by another tenant's rows, and gives the patterns that keep them out.
- A new support triage benchmark page.
v2.11.2
A reliability and first-hour release. Collections with Vector columns are stored correctly, API errors say what to fix (a missing key, a key without the right, a missing content type, a missing table), link paths are checked before a query runs, and self-links traverse. You can evaluate estimates on collections, and the free Docker image handles analyzer settings and API keys more safely.
If you have collections with Vector columns (including nullable ones) written before v2.11.2, reload them. Earlier builds could store a Vector column in a way that later made some queries fail with an error. Re-insert the rows, or recreate the collection and load it again. Optimizing does not repair it. Collections without Vector columns are not affected.
New
- Evaluate estimates on collections.
_evaluatewithestimatenow runs on collections and reports MAE, RMSE, Rยฒ, MAPE and error percentiles, each next to a baseline of always predicting the average, so you can see how much the model beats it. Recommend evaluation on collections is not available yet. See evaluation. GET /api/v2/schemadescribes Vector columns (dimensions, similarity and embedder), so the schema you read back can be PUT again to recreate the collection. Over the PostgreSQL wire protocol a Vector column isvector(n). See vector search.- Self-links traverse. A path through a link from a table to itself (such as
prev.user) now follows the link. See graphs.
Improved
$notdeletes on collections, and v2_deletetakes$ne,$in,$andand$or. See collections.- A
wherevalue of the wrong type for its column now returns a warning (where.type_mismatch) alongside the answer, so a mistyped filter is visible.
Changed
- API key errors are consistent on every route and method.
401means fix your key: it is missing, or not a key of this database.403means the key is real but lacks the right, which is almost always a read-only key used for a write (this was401up to v2.11.1). A missing key on a POST was405up to v2.11.1. See common errors. - A POST without
content-type: application/jsonanswers415, with a message naming the header to send (it was405up to v2.11.1). See common errors. DELETE /api/v1/schemadeletes tables only. Collections are kept, and the delete is refused if a kept collection links into a table it would delete.- The free Docker image is based on Ubuntu (Temurin 17) instead of Alpine. Existing data volumes keep working.
- The free Docker image shows API keys in full only once, on the boot that
generates them; afterwards they are masked. Read them with
docker exec <container> cat /io/state/.aito-api-keys, which now also holds keys you set yourself. See Docker. - An empty
READ_WRITE_APIKEYorAPIKEYnow stops the free image with an explanation, instead of silently using different keys. - The free image's startup log says when API keys changed since the last start, and states what telemetry is sent.
Fixed
- Collections with Vector columns could be stored so that some later queries failed with an error. New writes are stored correctly; see the reload note above.
- Ranking within a scoped candidate set is closer to the unscoped ranking. A
wherecondition on a target's attribute (for example"approver.customer_id": X) no longer down-weights evidence based on the narrowed set, and_evaluatewith such a condition reports a real baseline instead of 0. One remaining case, a candidate with no discovered relation to the evidence, is not fixed yet. - A link path that cannot be resolved is refused, with an error naming the
hop, instead of being ignored.
_relateover a copied column's path on a v1-engine table returns its relations (it answered with none). - A v1 query on a missing table names it and lists the tables that exist. An older database that has not been migrated says it must be migrated.
GET /api/v1/schemaworks when a collection has a Vector column; a Vector column on a v1 table is refused with400; an unknown operator or column in a v2_deleteis400.- On the free Docker image, a Text column declared with an analyzer object
(for example
{"type": "language", "language": "english", "useDefaultStopWords": true}, or a delimiter or n-gram analyzer) was refused as "unknown analyzer"; it now works as on every other build. - The free image's compose file uses the same volume, port and binding as the Docker page.
Docs
- The clustering page is now clearly a design preview: clustering is under development and cannot be enabled in current releases.
- Every v2 example in the docs was re-run against a live server, and the ones that were wrong are fixed.
- The performance page says which write mode its write-path figures were measured in.
v2.11.1
The first free Docker image built from the current engine. If you are upgrading from the 1.0.x image, read the authentication change first.
Changed
- Authentication is on by default in the free image. The 1.0.x images
answered requests without a key. From v2.11.1, the image generates an API key
pair on first boot, stored in
/io/state/.aito-api-keys, and every request needs anx-api-keyheader. Clients written against 1.0.x must send one. See Upgrading from 1.0.1.
Fixed
- The free Docker image starts when a licence key is set. In free images
before v2.11.1, setting
AITO_LICENSE_KEYleft the container running but never serving: it logged the licence check result and then waited indefinitely, so neither the HTTP nor the SQL port answered. Images started without a key were not affected. - If the licence check cannot complete, the server now starts in free mode after
AITO_LICENSE_STARTUP_TIMEOUTseconds (default 60) instead of waiting indefinitely.
v2.11.0
A correctness release for linked data and SQL. Several queries that used to
return a confident wrong answer now return the right one or a clear refusal:
sorting and selecting through links, _relate with more than one where
condition, SQL joins and GROUP BY, and v2 queries on v1-engine tables.
Predictions that route to people are fairer to new candidates and read each
link's own rows. Link-target predictions and graph filters are faster, and
the v1 streaming insert takes loads of any size again.
New
- A Python SDK page. Python SDK covers
pip install aitoai, the v2 client thataitoai1.0 makes the default, querying, errors and warnings, and loading data. The quickstart shows the same call from Python. - More worked examples. Seven new v2 use cases (recommendations, smart
search, autocomplete, autofill, product analytics, quality monitoring,
price and demand), and the playground now covers every read
endpoint:
_relate,$why,_evaluate, graphs, and common mistakes with the error each one returns. See use cases.
Improved
- Link-target predictions are faster. An unscoped prediction of a linked field, such as a processor chosen from every employee, answers in about half the time on a 1M-row table. The first such prediction after a restart or a write also starts faster, and scoped predictions on large candidate sets cost less.
- Filtered
$refsconditions are much faster. A$refs โฆ $existswith a condition object such as{"relation": "is-a"}runs about 19ร faster on large graphs, with identical answers. See graphs. $predictionssays its confidence is in-sample. A response that selectsfield.$predictionsnow carries aselect.predictions_in_samplewarning: each row is still part of the data its own prediction is read from, so the confidences read higher than held-out accuracy. Use_evaluatefor a number you can quote. See the query reference.- Request size limits answer with 413. An ordinary request body is limited
to 10 MB, and a larger one gets
413with a message that names the limit, where it used to get a500or a closed connection. See limits.
Changed
- Regression, by design: a candidate with no rows is ranked neutrally. When a prediction ranks linked rows (a new approver, a new hire), a candidate that appears in no row used to inherit the negative evidence of other candidates, including ones the query's scope had excluded, so it ranked below every observed candidate. It now carries no evidence either way and ranks by its base rate. The cost, measured on this release: on the invoice-routing benchmark, the acceptor mean rank on the v1 engine goes from 2.04 to 2.15. Top-1 accuracy is unchanged.
- SQL
INNER JOINdrops rows whose link does not resolve. A join used to keep rows whose foreign key isNULLor names a missing row, like aLEFT JOIN, which inflatedcount,sumandLIMITresults. Counts over such joins go down. A dotted path withoutJOINkeeps the link-projection behaviour. See JOIN along a link. - A SQL
JOINthat does not follow the declared link is refused. If theONcompares a column that is not a link, a link to another table, or a column other than the link's key, the query now fails with SQLSTATE0A000and shows the correctON. It used to return another table's data. - A SQL self-join is refused.
โฆ FROM employees a JOIN employees b ON โฆused to pair each row with itself. It now fails with SQLSTATE0A000, so clients such as DuckDB andpostgres_fdwcan fall back and do the join themselves. See the SQL reference. - A SQL literal whose type cannot match its column is refused.
WHERE price = 'abc'(also through a joined column) used to return 0 rows silently. See the SQL guide. - v2 queries on a v1-engine table refuse a path they cannot follow. On a
type: "table"table, awherepath of two or more links, or a misspelt path, now answers400naming what failed. It used to be ignored, and the query returned every row. A table follows one link; a path through two links needs a collection. See graphs. - Thin text evidence falls back to its words. When the rows containing all the
words of a
Textvalue in awhereare too few to measure, a prediction now uses what the separate words say instead of the base rate. This can move$pby a few hundredths onTextcolumns; the top answer was unchanged in our benchmarks. _evaluateon a collection explains what it cannot evaluate. On a collection,_evaluaterunspredict. Asking forestimate,recommend,relate,matchorsimilaritynow answers a400that says so and gives the workaround. See evaluation.
Fixed
- Two links to the same table each read their own rows. When one table has
two links to the same target (for example
approverandhandler, both toemployees), a prediction through one link could read the other link's rows. On the expense benchmark, handler top-1 on the v2 engine goes from 52.6% to 81.6%. With the employee attributes shuffled as a leakage check, it falls back to 61.25%. _relateand$patternstreat a multi-keywhereas one condition.{"field": โฆ, "customer_id": โฆ}is now the conjunction, as in_query,_predictand v1. Before, each key was related separately, so lifts and frequencies came from one key's rows, and a value could appear twice with different lifts.- Sorting keeps rows without a value. An
orderByon a nullable column or through a nullable or dangling link used to drop the rows that have no value, in JSON and in SQLORDER BY. They are now returned, listed last in both ascending and descending order. See sorting. - Selecting through a dangling link returns
null. A link that names a missing row used to fail the whole query with a500. _relateon a linked field relates only what co-occurs. It used to list every value of the linked field, even ones that never appear with thewhere, and awherethat matched nothing still related everything. It now behaves as it does on the table's own fields.$existsthrough a link that names a missing row. A row whose required link points at a row that does not exist is now treated like aNULLlink.$exists: trueused to count it.$notthrough a link on a v1-engine table. Through a nullable or dangling link,$notno longer matches rows that have no linked value._predicton a nullable link that holdsNULLs. It answered500as soon as a row actually used the null.$lengthand$distinctLengthignore empty values. Over a$refsprojection of a nullable attribute, an empty value was counted: teams {red,NULL, red} counted 3 and 2 distinct, not 2 and 1. See graphs.- SQL
GROUP BYreturns theNULLgroup. Rows whose group key isNULLor unresolved were left out of the result. They now form their own group. See the SQL guide. $innext to another operator on the same field works in JSON and SQL.WHERE k IN (โฆ) AND k IS NOT NULLused to fail withUnknown operation: $in.- SQL table aliases keep the table name's case.
SELECT p.name FROM Products pused to be read as a link path and fail. basedOn$patternswith a stopword. Awheretext value containing a word the analyzer removes, such as the or for, used to fail the whole prediction. Such words are now skipped.- The v1 streaming insert streams again.
POST /api/v1/data/{table}/streamhad been limited to 8 MiB. It now takes a load of any size, chunked or withContent-Length. See limits. - v1
_estimateand_aggregateanswer in the v1 shape on v2 tables. Through/api/v1they returned the v2 envelope. The v1_aggregateconditional count{"$f": {โฆ}}and the v1_relatearray form (relate: [{โฆ}]) now work too. A v1_relatewithwhere: {"$exists": "product"}related nothing, and now relates the product's fields as v1 does. See the v1 reference.
Also in this release: an x-knowledge column now uses the phrases it learned
as evidence in predict, and $why shows them as $phrase. x-knowledge is
experimental and may change between releases. See
schema design.
v2.10.3
Fixes for linked data: a misspelt path is now refused instead of quietly
answering, a condition that crosses two links works on graphs whose keys are
named after their entity, a predict reads a linked field through the right
link, and $why explains a same-member $refs filter.
Fixed
- A condition two links away follows the link's declared key. Inside a
$refsfilter,{"$exists": {"contact_id.role": "CTO"}}crosses a second link. On a graph whose keys are named after their entity (contact_id, notid) the condition was dropped, and the filter answered every row. It now follows the key the link declares, and a row whose link is empty satisfies no condition through it. See graphs. - Each link reads its own linked rows. When two links in a database
point at tables that share a key-column name and a field name, a
_predictreading that field through one link (for examplebasedOn: ["name"]on a linked target) could read it through the other. Each link now reads the rows of the table it names. $whyexplains a same-member$refsfilter. A query whosewhereheld a same-member condition such as{"$refs.edges.subject": {"$exists": {"relation": "is-a", "target": "CTO"}}}answered 400 as soon as it selected$why. It now explains the condition like any other. The refusal of an unsupported same-member shape also reads as a sentence again, where it had failed with "invalid escape". See graphs.$patternsacross a nullable link. Mining patterns across a link whose value is empty on some rows, or names a missing row, answered 500. Those rows now simply contribute no linked value. See the knowledge-graph use case.
Changed
- A misspelt path is refused. A
whereon a dotted path that does not resolve (a misspelt field at any hop, a misspelt column inside a$refsfilter, or a path through a self-link) now answers 400 naming the failing segment, instead of 200 with an empty or unconditioned result. See Common errors.
v2.10.2
Two knowledge-graph additions โ counting independent sources, and mining
patterns across a link โ plus two fixes: loading data into v2-engine tables
through the v1 API, and _relate properties that name a whole text value
through a link.
Added
$distinctLengthโ weighted corroboration.$lengthover a$refspath counts the referring rows, so ten edges filed by one source read as ten corroborations;$distinctLengthcounts the distinct members โ the number of independent sources behind a claim. SQL:aito.distinct_cardinality(col). See graphs.$patternsmines across a link path.{"$patterns": ["body", "cust.seg"]}now mines a linked field alongside the table's own, so a mined rule can span the join. See the knowledge-graph use case.
Fixed
- A v1 bulk load into a v2-engine table works. Loading rows through
POST /api/v1/data/{table}/stream, or through a file upload (/api/v1/data/{table}/file), into a table on the v2 engine answered 500 and loaded nothing, and a file upload's job ended as failed. Both now load the rows. This is the path the Python SDK'supload_fileand the console's CSV upload use. See the v1 reference. _relateaccepts a whole text value named through a link. A$propsentry such as"product.name": "Pirkka banana", wherenameis a text field on the linked table, was refused as "no rows carry" that value. It is now accepted.
Changed
- A v1 bulk load into a v2 table with a primary key or identity column is
refused. It answers 400 instead of loading, because that path cannot enforce
the key; a file upload is refused at its initiate call, before anything is
uploaded. Load such a table with
POST /api/v2/data/{table}/batch, which does. See the v1 reference.
v2.10.1
No engine changes. A documentation release: the responses printed in the quickstart and the sandbox walkthrough are now GENERATED from a test run against the sandbox's own data, so what a page shows is what the sandbox returns. Several pages contradicted themselves or each other, and those are corrected.
Changed
- Printed responses are generated, not pasted. Every response shown in the quickstart and the sandbox walkthrough is captured by a test that loads the sandbox's data, runs each documented call, and lifts the answers onto the page. They reproduce the live sandbox to the last digit. A pasted response drifts the moment the engine's numbers move; a generated one cannot, and the page fails its own test instead.
- The calibration claim says what was measured. It is scoped to its
evidence โ 26 held-out invoices, where predictions made at 82% confidence
were right 100% of the time โ rather than stated as a general property, and
the inference guide now tells you to measure calibration on
your own data with
_evaluate. The page also said support-tempering was on by default; it is not, and there is no such node in$why.
Fixed
- The
_evaluateexample inllms.txtworks verbatim. It returned a400. Copy-paste now runs. - An unknown
wherefield is documented where it bites. The contract โ andhavingโ appear in the query parameters and in a new inference section, Candidates, evidence, and$f, which states the candidate-universe rule, what$fcounts, how thin candidates are smoothed, and what happens on a cold start. - The benchmarks overview contradicted itself. Invoice routing is measured
on a 2,000-row hold-out; one place said 200. A literal
95%%rendered as written. See benchmarks. _evaluate's metrics are defined.h,baseSamples,baseGmp,geomMeanLift,rankGainandfeatureswere returned but never explained. See evaluation.- Links that went to the wrong place. The jobs link now resolves to limits, and the quota links are page-qualified.
- Every internal link resolves. 32 broken links are fixed โ stale operation paths, relative v1 links, and links that ignored the documentation's base path โ and the docs build now checks every internal link before publishing.
- The GL-coding use case prints real responses. It is generated from the
same test run as the quickstart; the previous page named the wrong approver
and printed a
meanRankno evaluation can produce. baseErroris documented correctly. It is 1 minus the majority class's share of the TRAINING rows, not 1 minusbaseAccuracy(which is measured on the test rows), so the two need not add up to 1. See evaluation.- The
_evaluatereference example shows today's response. It had been captured from the older engine and showed that engine's shape. - The sandbox GL example uses evidence that matches. It queried a description no row had, so it showed only the prior.
New
- A multi-tenant pattern in schema design. How to serve
many tenants from one shared schema. A nested
fromscopes base rates and a plain field's candidates to one tenant, but not a linked target's candidates; the guide gives the two patterns that do work for those โ a per-tenant linked table, or a tenant-scopedhavingstep before the predict. If tenants must never see each other's values, use the per-tenant table:havingshapes what a query returns and is not access control. config.ailistsv2. The alias existed and was not documented.
v2.10.0
Migration fixes, and one behaviour change to check: a multi-word $has on a
Text column is now refused rather than answered. The v1 endpoints answer over
storage migrated to the v2 engine โ five of them used to fail โ several
ranking defects that produced a confident wrong answer are fixed, and
deleting rows now takes effect on predictions immediately rather than at the
next optimize.
Changed
- A multi-word
$hason a Text column is now refused.$hasnames ONE analyzed token; a value that analyses to several โ{"name": {"$has": "BASF Oy"}}โ returns a400naming the operator that means what you asked, instead of an answer. On classic tables that request previously returned no rows, silently, so an application filtering this way got an empty result where it now gets an error. Single-token$hasis unchanged, and still analyzes the value, so{"$has": "BASF"}finds the stored tokenbasf. Write instead: the bare value for the whole value (exact equality),$matchfor all of its words,$searchfor phrases,ORand negation. If you wrapped a plain value in$hasto work around the filter behaviour of an earlier release, that wrapper is the shape that now fails โ send the bare value. See the query reference.
Fixed
- An unknown
wherefield inside$and/$or/$notmatched everything. A condition on a column that does not exist โ a typo, or a field a schema change removed โ was refused bare but silently DROPPED inside a combinator, leaving the surrounding operator unconstrained.{"$or": [{"typo": "x"}]}returned every row of the table where the same condition on its own returned none. A query whose filter was meant to narrow a population therefore returned the whole of it, with no error. The unknown field is now resolved wherever it appears, and warns. See the query reference. /api/v1xanswers in v1's shapes, not the raw v2 body. The endpoint exists so a v1 client whose storage moved to the v2 engine keeps parsing what it always parsed, and it was serving v2 bodies instead:_predictreturned$valuewhere a v1 client readsfieldandfeature,_recommendreturned$valuefor a flattened item record,_relatereturned{field: value}and a numericinfowhere v1 reads{field: {"$has": value}}and aninfoobject, and_evaluatewrapped its metrics. All four now match/api/v1. See the v1 reference.- The v1 endpoints answer over storage migrated to the v2 engine.
_evaluate,_match,_similarity,_estimateand_aggregatereturned400 failed to open '<table>'โ naming a table that visibly exists โ against a table whose storage had been migrated. Unmodified v1 clients now keep working after a migration, which is what the migration promises. See the v1 reference. _relateover migrated storage answers v1's own spellings. A v1 request usingrelate: "field"(a bare field name) or a nestedfromfailed with a400; and where it did answer,relatedandconditioncame back as{"field": value}where a v1 client reads{"field": {"$has": value}}, so the relation could not be named from the response._evaluatescored every test row from the prior when its evidence was bound through a link.{"$get": "product.category"}silently bound nothing, so the model reported an accuracy sitting exactly on its own baseline. An unresolvable path is now refused by name instead. See evaluation.- A nullable column's ABSENCE counts as evidence again. Where a column only applies to some rows, "no value" is often the signal โ and after moving to the v2 engine a model reading that column scored at its own baseline instead. On a 38 000-row demo predicting a customer segment from one nullable column, the v1 engine scored 0.745 and the v2 engine 0.475, its base rate; the two now agree at 0.745. Evidence that reads a value the row does not have states the absence, rather than contributing nothing.
- Conditioning on absence no longer fails.
{"col": {"$exists": false}}counted the right rows as a filter but returned a400naming an internal bitset width when used as evidence for a_predictโ whenever the column's last row happened to carry a value. Both uses now answer. - An exclusion on the predicted field survives
basedOn.{"product": {"$not": x}}sent withbasedOnkept the excluded candidate's slot and its probability and dropped only its value, so the response carried a hit with a$pand no$valueโ at rank 1. - A link candidate with no training rows no longer outranks one that was observed. Predicting a link target under a filter that narrows the population โ an employee within one company, say โ ranked never-seen candidates above candidates with real history, because the two were scored against different populations.
- A multi-word text value is scored at the strength of the whole match. Each word was scored at one word's strength and the duplicates then collapsed, so a four-word product name that matched exactly was ranked as weakly as a one-word coincidence.
- A deleted row no longer counts towards a prediction. Deleting rows
removed them from every query โ a filter returned none of them, and counts
excluded them โ but not from what the model had learned, until the database
was next optimized. So a value whose rows had all been deleted could still be
offered as a prediction, at the probability it had before the delete: asked
for a recommendation, Aito could name a product you had removed. With
evidence pointing at it, it could do so confidently. A delete now takes
effect immediately and means the same thing whether or not the database has
been optimized since. This also corrects
_evaluate, which holds its test rows out by deleting them and was therefore scoring a model against a population that still contained every held-out answer.
Improved
/docs/api/is now the v2 documentation. The v1 reference moved to/docs/api/v1/and stays fully supported; existing deep links into the old root resolve to their v1 section. See API v2.
v2.9.2
Two query bugs and two speedups on the plain database surface. A range
filter now answers on the v1 engine โ it used to return a server error โ
and relate on a text column counts the rows it says it counts. The
capacity and head-to-head benchmark figures are re-measured.
Fixed
- A range filter on the v1 engine returned a server error. A two-sided
window โ
{"id": {"$gt": 51200, "$lt": 52200}}, a SQLBETWEEN, or any two comparisons on the same column โ failed with an internalbit & operation requires same sized bit setsinstead of answering. One-sided comparisons were always correct, and the v2 engine was never affected. See the query reference. relateand$patternson a Text column reported the wrong support. A mined rule shown as one whole value was counted over every row sharing its words, so a rule displayed as a company with 245 invoices could report 331 โ the rows of that company and of every other name containing the same words. Text conditions now replay to exactly the rows their support counted. See the query reference.
New
- Choose which index of a text column you relate โ
.$distinctand.$token. A Text column carries both its distinct whole values and its analyzed tokens;relate,$patternsand$propscan now name either (product.name.$distinct,product.name.$token).$patternsoffers both pools by default and ranks them together, so a specific whole-value rule is available where it is the better one. Asking a column for an index it does not have is refused by name. See the query reference.
Improved
- Range and comparison filters no longer cost more as the table grows.
$lt$lte$gt$gteon numbers, strings and timestamps, plus$startsWith,$modand the timestamp component operators, allocated a bitset per matching value โ so a comparison keeping many of a column's distinct values grew quadratically with the table. At 102 400 rows a 1 000-wide window went from 378 ms to 20 ms. - Relational aggregates resolve their fields once per query.
select: ["$count", "$avg:amount", "$min:id", "$max:id"]under a filter went from 111 ms to 2.5 ms at 102 400 rows on the v2 engine. See the query reference.
Changed
- The capacity page reports the memory a server process uses, not the benchmark harness's. The old figures were read from a JVM that also held the test corpus, and were wrong in both directions. At 10 000 000 rows v2 holds the database in 213 MB of heap (published: 1 720 MB) and v1 in 5 290 MB (published: 2 603 MB) โ so a 10M expense database does not fit a v1 engine in a 4 GB heap. See capacity.
- The Elasticsearch head-to-head is re-measured. The previous table timed
one query shape at a time, so the first shape measured absorbed the whole
stack's warm-up and the last looked fastest โ which is why it showed
filter+textbeating both of its own constituents. The shapes are now interleaved on an idle machine. See performance.
v2.9.1
No engine changes. The scaling figures are re-measured on an optimized database โ the configuration a deployment actually runs โ and now lead with the warm median rather than a mean that averaged the first requests after a restart into steady state.
Changed
- The scaling tables measure an optimized database, and lead with warm latency. Two things moved together. The measurement now runs against an optimized database rather than an as-ingested one, and each row leads with the warm median, with the warm p90, the mean including the cold start, and the first query beside it. At 10 000 000 rows v1 reads about 210 ms at the warm median (417 ms p90, 355 ms mean including the cold start) and v2 about 279 ms (823 ms, 601 ms). The v2.9.0 page reported the as-ingested database with the mean: 853 ms median and 1 638 ms mean for v1 at the same scale. Those as-ingested figures stay on the page as the comparison, so the cost of leaving a database uncompacted is still visible. The engine is unchanged from v2.9.0. See scaling.
- The capacity page says plainly that an as-ingested database answers more slowly than an optimized one, and points at the warm-up breakdown. See capacity.
v2.9.0
The v2 API is in beta. Recommendations now use their query and your
basedOn in earnest, and relate and _evaluate reach per member and per
word. One change needs action: in a filter, a bare value on a Text column
now matches the whole value exactly โ use $match to filter by words.
Changed
- A bare value on a Text column filters by the exact value. In a filter โ
_query,_search,_aggregate, a SQL=, a delete โ{"name": "BASF Oy"}now selects the rows whosenameis exactlyBASF Oy(case-sensitive), matching v1. It used to match the words in any order, so it also returned โ and a delete also removed โ "BASF Construction Chemicals Finland Oy". To filter by words, use$match. As evidence in_predict,_recommendor_match, or when ordering by$p, a bare Text value still counts word by word. See the query reference.
New
relatea linked text field word by word.relate: ["product.name.$feature"]returns one relation per word of the linked product's name, as that column's analyzer produces it. See the query reference._evaluatemeasures per-member predictions and combined test selectors.predict: "tags.$feature"evaluates a multi-label prediction member by member, and atestselector can combine$indexwith a field condition. See evaluation.- Two public benchmarks, each with its data, its baselines and its numbers: smart search & recommendations and invoice line โ product matching.
Improved
- Recommendations use their query and
basedOn._recommendnow scores every candidate against its context exactly, where most candidates used to be ranked by how likely they were to meet the goal regardless of the query. On our smart-search benchmark, a recommendation from query words alone went from nDCG@10 0.12 โ about random โ to 0.45, level with BM25 over the catalogue, and reaches 0.53โ0.55 withbasedOn. A For You shelf ranked from the shopper's profile went from 0.11 to 0.18, and 0.24 withbasedOn. Rankings will differ from v2.8.4. - A text where through a link resolves a new word in about half a millisecond, where it took about 180 ms.
- A v1 match against a link target of over 100 000 rows answers in about 2 s instead of about 12 s.
Fixed
- A link-target prediction scoped by an attribute of the target honours that scope again. A regression in v2.8.4.
basedOnapplies to_recommend. v2.8.4 accepted it and ignored it.relateon a set field through a link relates each member, as it does on the field itself, instead of each whole set.- A text where through a link counts each word as evidence, as the same where on the table's own text column does.
- Statistics caches no longer serve a result computed for different data.
Two caches were keyed on only part of what they read, so a query could answer
200with statistics from other data; they now key on all of it.
Also in this release: an experimental x-knowledge column type and its
$surprisal operator. Experimental surfaces carry no compatibility promise and
may change between releases โ see schema design.
v2.8.4
The v2 API is in beta. A latency-and-correctness release: the largest single
cost of a scoped link-target prediction is gone, a linked column now supports
every operator its direct counterpart does, and six ways a query could return
the wrong rows or candidates with a 200 are fixed.
New
- Every where-operator works through a link.
$gt/$gte/$lt/$lte,$mod,$numeric,$startsWithand$searchapply to a field reached by its dotted path ({"part.price": {"$gte": 10}}) exactly as to the column itself โ the operator vocabulary is a property of the field's type, not of how you reached it. A bare value,$oror$inagainst a linked Set/Array column is membership, as on the column, and a column value operation selects through the link ("select": ["part.desc.$tokenCount"]). See fields through a link.
Improved
- A scoped link-target prediction answers in a fraction of the time. When a
prediction is scoped by a where on a linked path (
invoice.customer: X), the engine walked every candidate value of the link โ every seen and every unseen target row โ to test which ones the scope admits; it now derives the candidates from the linked rows the scope admits. The candidate universe of a link target is also built once per database state rather than once per request, and the per-candidate evidence selection does its grouping once. On a 16 000-candidate scope a fresh request went from about 0.8 s to about 0.6 s; on a few-thousand-candidate scope it answers in about 0.3 s. This particular change does not affect the answers. See improving inference. - Matching is faster and more accurate. When ranking candidates, the engine
built a pool of relations twice the size it could keep, refined all of them,
and then discarded the surplus. It now refines only what it keeps. On our ERP
matching benchmark the realistic
basedOnarms gain up to ten points of top-1 accuracy (for example 62.5% to 72.5% matching on name and supplier, mean reciprocal rank 0.758 to 0.817) and the sweep runs in about half the time; an expense-categorization run drops from about 157 ms to about 108 ms per query, and a 10 000-row invoice match from about 840 ms to about 320 ms per probe. Because the pool now holds different relations, individual rankings and the factor weights reported by$whycan shift slightly in either direction. See improving inference.
Fixed
- A
$not,$andor$oron a linked path now scopes the candidates and the evidence, not only the rows.{"product.id": {"$not": 0}}restricted which rows were counted but left the candidate pool and the evidence drawn from the whole linked table โ a200with plausible content and the wrong candidates. Every combinator over a linked path is now applied to both. - A
gethonours a where on the get field itself. An autocomplete query such asget: "name"with{"name": {"$startsWith": "mil"}}returned candidates from the column's whole distinct set, so unrelated values were offered beside the matching ones. The candidate universe now respects a same-field where the way it respects every other where. - Set-, array- and text-valued cells read consistently on the remaining paths, and a bit-set difference refuses a width mismatch instead of adopting the operand's width โ the cause of a negation that could answer with zero rows.
- A field reached through a null link is null. Through a nullable link
with no target on that row,
$match,$hasand$existscounted the row as present โ it read the linked table's first row โ so{"partN.name": {"$exists": true}}answered every row. Such a row now matches no operator, is excluded by$not, projects asnullinselect, and is selected by{"partN.name": {"$exists": false}}: the same rules as a nullable column read directly. - A modified link target is visible through the link at once. After
_modifychanged a row in a linked table, a path through that link could still read the row's previous values; the link now always resolves to the live row. - Excluding an item from a recommendation no longer offers it. A statement
about the item being inferred (
{"product": {"$not": 5}}while recommendingproduct) is a filter on the answer, never evidence for it; an excluded item โ including one the model had not seen โ is absent from the candidates instead of returned at rank 1. Statements about contextual data remain inference evidence, as before.
v2.8.3
The v2 API is in beta. A correctness-and-latency release: two silent
wrong-answer fixes the demo apps reported, a bit-operation bug that could
return plausible-but-wrong predictions, a broad prediction-latency pass, and a
deliberate narrowing of what basedOn is allowed to reach.
Changed
basedOnnow means exactly the fields it names. A field is paired when a known names it โ the engine no longer infers extra Text fields of a linked table whenever your query happened to carry any Text known. That inferred reach helped some shapes and hurt others, but the reason for removing it is that abasedOnis a statement of which evidence is permitted, and silently reaching past it made that statement advisory. If a prediction relied on the implicit reach, name those fields explicitly. Text evidence is now weighted by a fitted calibration rather than a fixed assumption. See improving inference.
Improved
- Prediction answers faster, across the board. A joined representation now
answers a position from the side that owns it (memoised per position set), a
linked text column seeks its token postings instead of re-expanding them,
numeric value priors are computed when read rather than once per candidate,
linked
selects resolve lazily, and the rotating cache evicts in constant time instead of scanning every slot on insert. The gains are largest on link-target predictions over wide linked tables.
Fixed
$noton a legacytablereturned nothing at all. On atype: "table"table queried through v2,{"field": {"$not": value}}returned zero rows instead of the complement โ a200with an empty body, indistinguishable from "no match". That is the basket-exclusion shape ($andof$not), so a legacy-backed database could serve no recommendations rather than reporting an error. Two independent causes were found and fixed: a numeric lookup that missed even when the value was present, and the negation itself._recommendignored a filter on a linked field. A$matchon a linked column constrained_searchcorrectly but not_recommend's candidate pool, which was drawn from the whole linked table โ200 OK, plausible content, wrong candidates. Every linked where-operator is now applied to the pool.- A bit-operation fast path could return a plausible wrong answer. When two
inputs were segmented differently, the dense fast path in
orcombined the wrong chunks: a shorter input raised an index error, a longer one returned an incorrect result with no exception. It is reachable from the predict path, so some predictions could be quietly wrong. Inputs are now aligned before they are combined. - Set-, array- and text-valued columns are read consistently. Membership was re-derived independently in several places, each matching only the shapes its author had in mind โ a column reads back as an array, a sequence, a set or a Java collection depending on codec and storage engine. Seven silent degradations are fixed and the derivation now lives in one place.
- The v2 documentation generator emits the wording its own artifact carries, so the published reference and the generated source no longer drift apart.
v2.8.2
The v2 API is in beta. One breaking change to _match's hit shape (below),
the contract fixes behind three defects our demo apps hit while migrating, and
a large cut in link-target prediction latency.
Changed
_matchhits carry$valuealone. The v1 spellingsfeatureandfieldare no longer returned by default:featurewas a verbatim copy of$value, andfieldwas always your own request'smatchparameter, so neither added anything. Every ranked-value op โ_predict,_recommend,_matchโ now reads identically, which is what the$valuerename was for. Porting a v1 body that reads them?select: ["feature"]andselect: ["field"]still return them, unchanged. See the v2 introduction.
New
- Rank candidates by an aggregate.
select: ["$f"]and{ "$sum": { "$context": โฆ } }no longer require the query to also carry a matchingorderBy, andorderBy: { "$sum": { "$context": โฆ } }is now accepted. So "which candidate values contribute most to the goal" is one query, instead of pulling the whole candidate set and sorting it client-side. See the query reference.
Improved
- Link-target predictions answer several times faster on a fresh request.
Column-only representations read rows directly, a query runs one prediction
instead of one per candidate column, and the
$whyexplanation is built only for the rows that ask for it. - A write costs less afterwards. Class counts and linked-attribute buckets are now kept per segment, so appending data re-counts only the segment that changed rather than the whole collection โ the first query after a write is much closer to the warm one.
- Steadier memory on long-lived instances. Caches keyed by data lineage are bounded by the generations still retained rather than by an entry count, so they follow what the instance is actually holding.
- An attribute counts once per candidate. When several pieces of evidence pointed at the same attribute of a candidate, it could be credited more than once and inflate that candidate's probability. It is now counted once.
Fixed
- Per-candidate aggregates across a link returned zero.
$fand{ "$sum" | "$mean": { "$context": โฆ } }over agetwhose path traverses a link โproduct.name,context.weekโ reported0for every candidate, with a200and a well-formed body. Direct fields and the link field itself were correct, so a chart could show a confident flat line at zero for a product with hundreds of purchases. They now resolve the candidate's rows through the field that owns them. relate: { "$props": โฆ }accepts set-valued properties. A property held in aSetcolumn (product.tags) was refused with a message saying no rows carried the value โ while rows plainly did, because membership and equality are different questions about the same data. Set-valued properties, and a list of values for one property, both work now.- An empty
$andor$orclause list. Building a filter list that turns out empty โ{ "id": { "$and": [] } }from an empty basket โ failed the request. An empty$andnow constrains nothing and an empty$orselects nothing, which is what each means and what v1 already did. - A union view's refusal names the branch again. When an output column resolves in no branch, the error names each branch and the columns it tried, rather than only reporting that every projection was absent.
- The v1 type reference no longer prints internal class names. A few
anonymous formats reached the served pages as a Java class name and an object
hash instead of their type, in the
AnyValueandRegressionExplanationreferences.
v2.8.1
The v2 API is in beta. A follow-up to v2.8.0: query latency on large
collections, better link-target predictions, and a set of correctness fixes
around evidence and aggregation โ plus one new relate form.
New
- Relate an entity's own property values.
relate: { "$props": โฆ }asks what distinguishes a row's own attributes against thewhere, so "what is unusual about this customer" is a single query rather than one per field. See the query reference (and the SQL spelling in the SQL reference). config.aiapplies on v1_recommend,_matchand_search. Where it genuinely cannot apply โ_relateand_similarityโ the request is now refused with a 400 instead of being accepted and ignored, so a tuning setting never silently does nothing. See improving inference.
Improved
- Large collections answer much faster. Broad selects, page fetches and
orderByon a plain field no longer pay a per-query floor; ordering reads the index's own value order, and computed fields are evaluated over the filtered universe rather than the whole collection. The biggest gains are on a scopedorderByover a computed field, which was the slowest shape at scale. - Leaner prediction. Per-state ordinal maps, segment-lifetime column reuse and primitive cursors in the bit loops cut allocation churn on the predict path, so repeated inference on a warm collection holds a smaller footprint.
- Better link-target predictions. A link candidate is now scored through its target's attributes under a field-pair prior, so a candidate is judged by what it actually looks like rather than only by how often it has been seen.
- Name matching fades with history. The name-identity boost decays as a row accumulates history instead of being discounted by query coverage, so a well-established entity is ranked on its record rather than its name.
Fixed
- A
whereon the predict target restricts candidates, not evidence. Naming the target inwherenarrowed the candidate set and was counted as evidence for it, which skewed the probabilities; it now only restricts. - Deleted rows no longer form a group. A deleted row could still produce its
own
GROUP BYgroup;$fand$contextaggregates now intersect the live rows. whyanswers on a large collection. The KNN explanation is bounded, so_estimatewithwhyno longer fails as the collection grows, and it reads evidence rows from the collection rather than the field.basedOnon a text name no longer regresses link-target predict.- Env branches survive a binary-format upgrade. After v2.8.0's format bump an env branch failed plain reads with a raw internal error until it was repaired by hand. Branches now migrate on first use, and a read against a not-yet migrated state returns a clear error naming the remedy instead of an internal failure. See environments and common errors.
- Documentation rendering. Table captions rendered as literal markup on 19 pages, and a dark page backdrop made overflowing text unreadable.
v2.8.0
The v2 API is in beta. This release is mostly speed and memory at scale โ text postings and column indexes are now prepared at write time and read straight from memory-mapped storage instead of being decoded onto the heap โ alongside a wider SQL surface and explanations that point at the evidence inside a predicted item.
New
- Derive a table from a query.
INSERT โฆ SELECTandCREATE TABLE โฆ AS SELECTbuild a table in the database from a query over existing ones, so a derived dataset no longer round-trips through your client. See the SQL guide. evaluate()andestimate()as SQL table functions. Measure accuracy and get a prediction's estimate from SQL, withtest/testLimit/train/resubstitutionarguments to control the split. See the SQL reference.- SQL answers in PostgreSQL result format.
/api/v2/_sqlcan answer with rows and column metadata in the pg result shape instead of the hits shape, so a Postgres client renders it directly. See the SQL reference. - Multi-column
GROUP BYand multi-keyORDER BY. Group by joint tuples and order by several keys, with Postgres's default null placement per key. See the SQL reference. havingโ filter candidates on the server. Filter ranked results on$fafter ranking and beforeoffset/limit, so an "at least N occurrences" cut happens in the engine instead of in your client. See the query reference.- Explain a predicted link or text target through its own attributes. When
the prediction is a linked row,
$whynow shows the candidate's own fields as the prior that carried it, rather than leaving that path unexplained. See the inference guide. $highlighton a predicted target. Add it to a link-target predict'sselectand each candidate comes back with its own field values marked where the evidence sits โ ready to display. See the query reference.- Structured JSON logs.
AITO_LOG_FORMAT=jsonwrites one JSON object per line to stdout, with context such asdatabaseinlined as a field, so a self-hosted instance ships into your existing log pipeline with no parsing rules. See running it on your own infrastructure. - Union views are served live. A stored union view now answers from its
sources once one of them changes, so it cannot go stale;
_refreshbecame an optimisation rather than a correctness step. See schema design.
Improved
- Faster and much lighter at scale. Columns and token postings are persisted in read-ready form, so every column type opens in constant time and reads stay memory-mapped rather than decoding whole blobs onto the heap. Large states use substantially less query-time memory and answer faster.
- Better link predictions. Predicting a linked field now draws candidates from the linked table rather than only the values seen in training, so a valid target that never appeared in the training rows can still be predicted and ranked. See the inference guide.
- The first query after a write. A
recommendissued right after a write no longer pays to refit the prior before answering. - Ranking on catalogue-shaped data. Name-based identity matching is now bounded by the evidence that supports it instead of a fixed internal ceiling, which improves ranking where many rows share similar names.
- Quieter error logs. Client-input errors are no longer logged at ERROR with stack traces, so an operator's error channel reflects server problems. See common errors.
Fixed
- One
evaluatetrain/test contract across engines. An undefinedtrainnow means the complement of the test set on both engines, andtestSourcealigns with it โ so the same request reports the same accuracy whichever engine backs the table. See evaluation. - Baselines measure the right population. An evaluation's baseline is the model with no evidence, scoped by the query's own filters rather than a wider pool; filtered evaluations previously reported an optimistic baseline.
estimateover a nullable numeric column estimates over the non-null subset instead of failing.- Explanations on temporal columns. A KNN
$whyover a v2 collection's temporal columns no longer errors. - v1 schema
GETrenders every scalar type reachable from v2 instead of failing with a 500. relateandsimilarity.relatereported inverted frequencies, andsimilaritydiscarded partial evidence; both fixed.- Deleted and superseded rows. A deleted row now releases its primary key, and a superseded row is no longer offered twice as a prediction candidate.
- Merged
Array/Setcolumns. A width mismatch between the merge writer and reader could decode stored values into wrong ones; fixed. - SQL that real tools emit. Comments, positional
INSERT, and aFROM-lessSELECTnow parse. See the SQL reference. - pgwire. A duplicate column name keeps Postgres's own answer instead of
being renamed, and a data error reports
22P02/22P04rather than a syntax error. See the SQL reference.
v2.7.0
The v2 API is in beta. This release unifies the v2 response contract across
storage engines, extends where filtering (Date ranges and let fields), and
adds on-prem observability with a Prometheus metrics endpoint.
New
- Prometheus metrics endpoint.
GET /metricsexposes request counts, a request-latency histogram, and JVM/disk gauges in Prometheus text format โ per database โ so a self-hosted Aito scrapes straight into Prometheus and Grafana. See monitoring your instance. - Filter on
Dateranges.wherenow accepts range comparisons onDatecolumns, so "orders in the last quarter" is expressible directly. See the query reference. - Filter on
letfields. A field defined withletcan now be used inwhereโ equality,$in, and same-field$ormembership โ so a value you derive filters the same query that defines it. See the query reference.
Improved
- One response contract across engines. A v2 endpoint now returns the same
response shape regardless of which storage engine backs the table โ one
selectvocabulary, one similarity-score key, a consistent errorkind, and a response-time header โ so client code no longer special-cases the engine. - Honest HTTP error codes. Capability and input errors are classified into
proper 4xx codes with a machine-readable
kind, instead of returning prose in the error field. See common errors. - Faster v2 queries. Leaner bitset intersections, single-pass posting-list walks, and segment-aligned per-known tails cut work on large states; the AND-prior evidence gate is now on by default.
- On-prem operations. An operations runbook and a
restore-statecommand with a backup/restore drill for self-hosted deployments, and production no longer logs at DEBUG by default. See clusters & operations.
Fixed
- Namespaces and mounted sub-envs no longer 500. SQL DDL/DML and v2 schema/data/query addressed at a namespace or a mounted sub-env returned a 500 (an internal cast error); they are now handled correctly.
_evaluatehonoursbasedOn._evaluatesilently ignoredbasedOn, which voided any measurement run through it against a based-on database; fixed.- Filters no longer dropped under a tenant scope.
recommendandrelateover a v1-backed table could silently dropwherefilters when run under a tenant scope; fixed. - Nullable columns. A nullable column could leak one row's value onto rows that have none; fixed.
- Empty-tokenising text filters. A text
wherewhose value tokenised to nothing matched the whole table instead of nothing; fixed. - SQL over pgwire. Column names are now identical across both transports, and notices are delivered.
v2.6.1
The v2 API is in beta. This release is mostly stabilization โ SQL/pgwire correctness and lower memory at scale โ with a few additions to the SQL surface.
New
- Prediction as a SQL table function.
SELECT * FROM predict('invoices', 'category', given => 'vendor = ''Acme''')(andpredictions(โฆ, k => 5)) runs a hypothetical prediction for a row you don't have yet โ the SQL spelling of what_predictdoes over JSON. See the SQL reference. - List/Set member edits in
_modify.update'ssetaccepts{"tags": {"$add": ["sale"], "$remove": ["draft"]}}to add or remove members of a list/Setcolumn per matching row โ additive and safe under concurrent updates. - SQL table aliases.
FROM customers AS candFROM customers cnow parse, so BI tools, ORMs and query builders that alias tables (and both sides of a JOIN) work as written.
Improved
- Faithful transactions over pgwire. A multi-statement request is now one
implicit transaction on the extended protocol too (pgjdbc and most drivers), so a
failed statement no longer leaves earlier ones committed.
DISCARD ALLnow fully resets connection state โ important for connection poolers. - Honest error codes. Capability gaps return
0A000(feature not supported) instead of42601(syntax error), so federating clients (Metabase, DuckDB, postgres_fdw, SQLAlchemy) fall back gracefully instead of reporting your query as broken. Clients can also introspect available functions via a generatedpg_proc. - A prediction column is bounded by default.
SELECT predictions(col) FROM tnow defaults toLIMIT 10instead of running one inference per row over the whole table โ matchingrecommend/relate/searchand the JSON API. - Lower v2 memory at scale. Several unbounded rep2 caches are now bounded and a per-query mask is leaner, continuing v2.6.0's 10M-scale reductions.
Fixed
- Correct results in large multi-segment states. Fixed a case where a state with more than 64 segments could silently drop recommendation/prediction evidence.
- Three SQL engine bugs found by a new conformance suite, including
WHERE price < 5 OR price > 50no longer erroring.
v2.6.0
The v2 API is in beta. This release adds new query surface and continues hardening it.
New
- SQL is a full query surface now. Rank candidate values with
match(โฆ), mine frequent patterns withpatterns(โฆ), run a relevance search withaito.search(โฆ)(carrying$why,$highlightand$matches), and find nearest vectors withORDER BY col <-> '[โฆ]' LIMIT kor theaito.knn/aito.nnconditions. See the SQL reference. - Vector predicates for prediction โ
aito.clusterandaito.semanticasWHEREconditions, andvector(n)columns to hold embeddings. - TIMESTAMP columns over SQL. Declare and range-filter timestamps, do
now()+ interval arithmetic in aWHERE, read components (EXTRACT,date_part, day-of-week in both spellings), and bucket withdate_trunc. See the SQL reference. - Define your schema in SQL โ
CREATE TABLEdeclaring links and analysed text, andCREATE VIEWover the query functions. - Real SQL transactions โ
BEGIN/COMMIT/ROLLBACKare buffered and merged on commit, with working savepoints for partial rollback. - Aito's functions live in an
aitoschema, keeping the public namespace clean while the key concepts stay reachable unprefixed. select: [field.$predictions]returns a field's ranked predictions inline in a v2 query.
Beta stabilization
Hardening across the v2 engine and its SQL surface:
- More accurate predictions by default. Name-based identity boosting is now off by default โ on a 10M-row benchmark it was costing ~30 points of top-1 accuracy by letting a name match take over the score. Predictions that relied on the old default will change, and should improve; the boost is still available and is now bounded, so it can add signal without dominating. See improving inference.
- Much lower v2 memory at scale. Large v2 predictions allocate far less per call, and the internal firing cache now has a real size budget with eviction โ a 10M-row workload no longer balloons memory.
- No more 15-minute hangs on out-of-memory. A fatal out-of-memory error mid-request used to stall the connection until the request timeout and then return an unusable body; it now fails fast and cleanly.
- Assorted correctness fixes. A
Textcolumn added to an already-populated table now feeds a$textview; and over the SQL wire, a data query that mentions a catalog name is answered from your data (not the catalog), a partial rollback no longer discards the whole transaction,CREATE VIEWvalidates its functions up front instead of degrading silently,COLLATEis honoured, andGET /versionreports the real version.
v2.5.3
New
- SQL over the Postgres wire protocol is on by default, and the Docker image
now ships with API keys configured. Connect
psql, JDBC/ODBC or psycopg straight at Aito โ see the SQL guide. - TLS is required for the SQL listener on a shared port, and SQL sessions are bounded per connection.
Improved
- A read-only API key can no longer open a read-write SQL session โ the SQL surface now honours key scope the same way the REST API does.
- Each database is authorized by its own keys, not the server's.
Fixed
$whyexplanations keep their highlights when scoring composes several factors โ highlight output no longer nests one level too deep.- A refused SQL listener no longer takes the server down with it; a SQL session no longer outlives the database it was opened against.
v2.5.2
Fixed
$whylifts and highlights are reported at the level clients expect, so$highlight/$matchesresults render correctly again.
v2.5.1
New
- Hybrid search: combine BM25 text relevance with vector similarity in one
ranked query โ
$vectorSimilarityand the self-calibrating$vectorIdfblend with$p. See vector search. /api/v1runs on the v2 engine for the query family, so v1 clients get v2 performance without changing a line.- The vector dimension cap is lifted (previously ~1017 dimensions).
Improved
selectaccepts$highlight/$matchesas aliases for$whyhighlights.
Fixed
- Cross-tenant isolation:
$and/$notover a linked field no longer returns rows from another tenant's data. _matchon a text column now predicts per token, matching v1 behaviour._estimaterestores v1 per-field attribution inwhy, and defaults to AdjustedKNN as v1 does.- Several v1โv2 response-shape differences corrected (
_predict,_recommend,_relate,_match), so a v1 client sees the shape it expects.