What's new

What changed in each Aito release, from your side: new capabilities, improvements and fixes. Written per release โ€” not a commit log.

Which build am I on? GET /version on your instance reports the build it is running. Docs for a specific release live at /docs/api/<version>/; this page and /docs/api/ describe the current release, and /docs/api/edge/ tracks development ahead of it โ€” so a feature listed here is available on the current release unless an entry says otherwise.

v2.11.4

A correctness release for tables with deletes and multi-tenant queries. A search or listing after a delete shows exactly the live rows, deletes and upserts write exactly the rows they matched (also under concurrent writes), concurrent writes no longer fail on another commit's cleanup, a nested from keeps its population exact, and filters through several links restrict a linked-row prediction. Relate honours orderBy, and SQL honours NULLS FIRST/LAST.

New

  • Count links directly. $length over a bare $refs path counts the referring rows (a node's degree), and orderBy can sort by a $length. See graphs.

Improved

  • More accurate matching of text pairs. When none of a text pair's calibration fits is acceptable, predictions no longer fall back to the most confident curve, which made entity matching between differently worded catalogues much less accurate.
  • Relate honours orderBy: hits sort by the named member (a bare key sorts descending), and a sorted answer whose candidate cap dropped values carries a relate.order_by_partial warning. Ties among equal-score relate hits order deterministically by value. See query reference.
  • Reading rows of a table with deletes not yet optimized is fast again: a per-row read no longer slows down with every delete. On 1M rows it took ~65 ms per row after 100 deletes, and ran out of memory after ~1,000; it is now a few microseconds whatever the number of deletes.
  • Joins through links are much faster on the first query: it no longer builds the whole link index. A cold select โ€ฆ limit 5 through a link at 10M rows goes from ~12-23 s to under 0.5 s, and the first filter on a linked field from ~22 s to ~3 s. A linked field's lookup is built once instead of twice, and over a declared link it reuses the lookup built when the data was written.
  • A nested from is faster on repeated queries: its population is built as a table of the kept rows and cached by content. See multi-tenant.
  • SQL information_schema answers as relations: WHERE (=, <>, LIKE/ILIKE, IN, NOT IN, IS NULL), ORDER BY and LIMIT/OFFSET are honoured, views are reported as VIEW, data_type is a type name, and the PostgreSQL standard columns are present, so tools such as dbt can read the catalog. Some query shapes still answer wrong; see Known issues. See SQL.
  • x-knowledge columns (experimental): every $why factor is a proposition you can send back as a where ($phrase, $symbol), a word outside its phrases is read on its own rows, and phrase learning is faster. See schema design.

Changed

  • A nested from that selects no rows no longer fails with 400 request.invalid: _query, _search, _recommend, _match and relate $patterns answer 200 with no hits and a query.empty_population warning, and _predict and _evaluate answer 400 query.empty_population. See common errors.
  • A relate select that names no relate member is a 400, and an orderBy relate cannot honour is a warning rather than an error.
  • A filter through two or more links that cannot be evaluated on a linked-row prediction now answers 400 naming the path, instead of being ignored.

Fixed

  • A nested from gives the same answers as a table holding only its rows: a $refs walk no longer omits population rows, evidence from several fields no longer drifts with other tenants' rows, and a link-path condition in its where (such as "company_id.industry": "accounting") no longer fails with 400.
  • After deleting rows, a search or listing shows exactly the live rows, also before the table is optimized: no deleted row is listed, the last row is no longer left out, and a where on it finds it.
  • Concurrent deletes and updates on one table write exactly the rows they matched: v2 _delete, SQL DELETE, UPDATE and INSERT โ€ฆ ON CONFLICT DO UPDATE SET. Before, with another delete in flight, they could delete or overwrite the neighbouring rows. A write affects the rows that matched when it was requested and still match when it commits (PostgreSQL's rule for keyed tables), and the reported count is the rows actually written. A large write that keeps losing to concurrent writes, or a pgwire COMMIT that cannot be replayed, now fails fast with SQLSTATE 40001 so the client retries, instead of stalling the table.
  • Under concurrent writes, a write could fail with 'physical file does not exist' when another commit's cleanup removed its not-yet-committed files. Cleanup now keeps unreferenced files for 30 minutes before removing them, so the disk space of deleted or replaced data is reclaimed up to 30 minutes later. Operators can set the window with -Daito.gc.graceMillis=<ms>.
  • An on_conflict=update upsert after a delete replaces its own row and keeps its neighbours: v2 batch and keyed import with ?on_conflict=update, and SQL INSERT โ€ฆ ON CONFLICT (key) DO UPDATE. Before, it could delete the next row and keep the upserted row's old version.
  • A filter through two or more links restricts a linked-row prediction (_predict, _recommend, _match), for example where: {"product.supplier.country": "FI"}.
  • A nested from keeps its population exact with AITO_WRITE_DEFER_MERGE enabled, $patterns over a nested from counts only the population's rows, and _evaluate over a nested from trains on every remaining row after a delete.
  • SQL: ORDER BY โ€ฆ NULLS FIRST | LAST is honoured; the NULL group sorts as in PostgreSQL (last ascending, first descending); GROUP BY or a value-ordered get over a nullable column no longer fails; evaluate(โ€ฆ, where => c) evaluates over the population c, as documented; a catalog LIKE honours the backslash escape. See SQL.
  • A view can copy a Vector column. See multi-tenant.
  • Relate no longer offers another row's value on a nullable column or on a collection with deleted rows not yet optimized.

Known issues

  • A long write (a large _delete, e.g. 100k rows, or a ~1,000-row keyed upsert into a large table) on a table that receives steady concurrent inserts can fail with HTTP 500 after exhausting its commit retries. Workaround: pause writes to that table, or split the operation into smaller batches.
  • On collections, selecting or ordering by a linked row's field (select, orderBy) can fail with HTTP 500 ('has no resolvable link') or show another row's value when the link's target key does not exist, the target row was deleted (until the table is optimized), or the field is empty (null). $vectorSimilarity through a link to a row with an empty vector can score with another row's vector, which changes sorting. Prediction probabilities and explanations are unaffected; stored data is unaffected. Fixed in the next release.
  • After deleting rows, and until the table is optimized, _relate statistics (including $patterns) can be computed incorrectly. Workaround: run optimize on the table after deletes before relying on _relate.
  • Some SQL information_schema query shapes (LIKE โ€ฆ ESCAPE, || concatenation, quoted or catalog-qualified table names, joins between information_schema views) can return empty or wrong results instead of an error. Fixed in the next release.
  • After deleting the last rows of a table, and until it is optimized, _relate on a Set field, an nlp-analyzed text field or a linked field (field.attribute) can fail with an error (reported as 400 or 500). Workaround: run optimize on the table after deleting.

v2.11.3

A correctness release. Filters on classic tables keep every condition, predictions on collections read the right rows when a column holds nulls or rows were deleted, and a nested from keeps a tenant's population exact after a delete. You can now evaluate and match within one tenant, and /import recognises free text and accepts CSV.

New

  • Evaluate and match within one population. _evaluate and _match accept a nested from (for example {"from": "invoices", "where": {"tenant": "A"}}): test rows are drawn from that population and the model trains on the rest of it. See evaluation and multi-tenant.
  • /import recognises free text and accepts CSV. A string column that reads like prose is inferred as Text (short codes, labels and ids stay String), content-type: text/csv with a header row is accepted, and the response warns about prose in a column you declared String. See collections.

Improved

  • Every field you list in basedOn contributes evidence. Relation discovery keeps a share for each declared field, so one field with many relations no longer crowds another out of the scoring. See improving inference.

Changed

  • _estimate, _aggregate and _similarity refuse a nested from with 501 and name the limitation (they answered 400 missing 'from'). Their statistics are not restricted to a population, so they refuse rather than answer over every row. See multi-tenant.

Fixed

  • Filters on classic tables keep every condition. On a classic table, a v2 query whose where combined an equality with a comparison in one object ({"tag": "blue", "n": {"$lt": 0}}) ignored the comparison; and a comparison could answer with the result of an earlier comparison on a different column of the same table. Both now apply exactly.
  • _evaluate with per-row comparison thresholds ($get inside $lt/$gtโ€ฆ) on classic tables now evaluates each row against its own threshold (previously every row used the first row's), no longer warns that the where has no $get binding, and reports a baseline without the per-row condition. See evaluation.
  • Predictions on collections with nullable columns or pending deletes read the right rows. On a collection, a nullable column that holds nulls, or a table with deleted rows not yet optimized, could feed predict, $why and text (BM25) scoring with another row's value for that column, so per-row statistics were computed from shifted values. Affected: collections ("type": "collection") with nullable columns containing nulls since v2.2.0, and collections with deleted-but-not-optimized rows since v2.1.47. Classic tables are not affected. No action is needed after upgrading.
  • A nested from after a delete holds exactly its population. On a collection with recent deletes, a nested from could return the wrong rows, and _evaluate could hold out rows that had been deleted.
  • A prediction offers only values that live rows hold. A value held only by deleted rows (until the collection was optimized), or only by rows outside a nested from, could be offered as a candidate; _evaluate no longer offers a label that occurs only in its held-out rows.
  • Base rates after a delete count the rows that remain, however the collection is split into segments.
  • Scoping a link-target prediction with a where on the target's attribute (for example "approver.customer_id": X) no longer reverses the ranking of candidates the evidence has no discovered relation to. This was the remaining case noted in v2.11.2.

Known issues

  • Inside a nested from, a $refs walk can omit population rows, and two-feature evidence can shift probabilities slightly. Single-feature and plain where conditions are exact. Per-tenant collections are unaffected.
  • A link-path condition in the where of a query over a nested from returns 400 on collections.
  • After deleting rows, and until the table is optimized, a search without select can leave out the table's last row (a where on that row may return nothing), and a listing with no where can still show the deleted row. Rows are not shown to the wrong tenant, and values are not mixed between rows. Stored data is unaffected. Workaround: add an explicit select to the query, or run optimize on the table. Fixed in a following release.
  • On collections, a prediction's select can show a value for a linked row's empty (null) field, and $vectorSimilarity through a link to a row with an empty vector can score with another row's vector, which changes sorting. Prediction probabilities and explanations are unaffected; stored data is unaffected. Fixed in a following release.
  • With AITO_WRITE_DEFER_MERGE enabled (off by default), a nested from without select can return rows outside its where, on any table, with or without deletes. Workaround: add an explicit select, or put the filter in a top-level where. Fixed in the next release.
  • SQL: NULLS FIRST / NULLS LAST in ORDER BY is ignored (the default null ordering is applied). In information_schema, ORDER BY and IN filters on columns are ignored, views are reported with table_type 'TABLE', and columns.data_type returns a numeric JDBC type code instead of a type name. Fixed in a following release.
  • _evaluate over a nested from (new in this release) can leave one row out of training after rows are deleted from the table, until the table is optimized; the reported accuracy can then differ slightly. Workaround: optimize the table before evaluating. Fixed in the next release.
  • _relate with $patterns over a nested from can compute its statistics over the whole table, including other tenants' rows, rather than the nested population, once the table has been optimized or with AITO_WRITE_DEFER_MERGE enabled. Keep each tenant's data in its own table or instance. Fixed in a following release.
  • After deleting rows, and until the table is optimized, _relate statistics (including $patterns) can be computed incorrectly and may still include the deleted rows. Workaround: run optimize on the table after deletes before relying on _relate.
  • Concurrent _delete requests on the same table can delete rows other than the requested ones and leave the requested rows in place. Concurrent SQL DELETE/UPDATE statements over the PostgreSQL interface share the same pattern. Until fixed, send deletes and SQL updates to a table one at a time.
  • After deleting rows from a table, an on_conflict=update upsert (batch, import, or SQL INSERT โ€ฆ ON CONFLICT DO UPDATE over the PostgreSQL interface) can delete a neighbouring row and keep the old version of the upserted row. Workaround: run optimize on the table after deleting and before upserting. Fixed in the next release.
  • A _predict, _recommend or _match that predicts a linked row and filters through two or more links (e.g. where: {"product.supplier.country": "FI"}) can ignore that filter and answer over all previously seen candidates, without an error. Filters through a single link are not affected. Fixed in the next release.

Docs

  • A new multi-tenant guide answers, per query type, whether a query scoped to one tenant can be influenced by another tenant's rows, and gives the patterns that keep them out.
  • A new support triage benchmark page.

v2.11.2

A reliability and first-hour release. Collections with Vector columns are stored correctly, API errors say what to fix (a missing key, a key without the right, a missing content type, a missing table), link paths are checked before a query runs, and self-links traverse. You can evaluate estimates on collections, and the free Docker image handles analyzer settings and API keys more safely.

If you have collections with Vector columns (including nullable ones) written before v2.11.2, reload them. Earlier builds could store a Vector column in a way that later made some queries fail with an error. Re-insert the rows, or recreate the collection and load it again. Optimizing does not repair it. Collections without Vector columns are not affected.

New

  • Evaluate estimates on collections. _evaluate with estimate now runs on collections and reports MAE, RMSE, Rยฒ, MAPE and error percentiles, each next to a baseline of always predicting the average, so you can see how much the model beats it. Recommend evaluation on collections is not available yet. See evaluation.
  • GET /api/v2/schema describes Vector columns (dimensions, similarity and embedder), so the schema you read back can be PUT again to recreate the collection. Over the PostgreSQL wire protocol a Vector column is vector(n). See vector search.
  • Self-links traverse. A path through a link from a table to itself (such as prev.user) now follows the link. See graphs.

Improved

  • $not deletes on collections, and v2 _delete takes $ne, $in, $and and $or. See collections.
  • A where value of the wrong type for its column now returns a warning (where.type_mismatch) alongside the answer, so a mistyped filter is visible.

Changed

  • API key errors are consistent on every route and method. 401 means fix your key: it is missing, or not a key of this database. 403 means the key is real but lacks the right, which is almost always a read-only key used for a write (this was 401 up to v2.11.1). A missing key on a POST was 405 up to v2.11.1. See common errors.
  • A POST without content-type: application/json answers 415, with a message naming the header to send (it was 405 up to v2.11.1). See common errors.
  • DELETE /api/v1/schema deletes tables only. Collections are kept, and the delete is refused if a kept collection links into a table it would delete.
  • The free Docker image is based on Ubuntu (Temurin 17) instead of Alpine. Existing data volumes keep working.
  • The free Docker image shows API keys in full only once, on the boot that generates them; afterwards they are masked. Read them with docker exec <container> cat /io/state/.aito-api-keys, which now also holds keys you set yourself. See Docker.
  • An empty READ_WRITE_APIKEY or APIKEY now stops the free image with an explanation, instead of silently using different keys.
  • The free image's startup log says when API keys changed since the last start, and states what telemetry is sent.

Fixed

  • Collections with Vector columns could be stored so that some later queries failed with an error. New writes are stored correctly; see the reload note above.
  • Ranking within a scoped candidate set is closer to the unscoped ranking. A where condition on a target's attribute (for example "approver.customer_id": X) no longer down-weights evidence based on the narrowed set, and _evaluate with such a condition reports a real baseline instead of 0. One remaining case, a candidate with no discovered relation to the evidence, is not fixed yet.
  • A link path that cannot be resolved is refused, with an error naming the hop, instead of being ignored. _relate over a copied column's path on a v1-engine table returns its relations (it answered with none).
  • A v1 query on a missing table names it and lists the tables that exist. An older database that has not been migrated says it must be migrated.
  • GET /api/v1/schema works when a collection has a Vector column; a Vector column on a v1 table is refused with 400; an unknown operator or column in a v2 _delete is 400.
  • On the free Docker image, a Text column declared with an analyzer object (for example {"type": "language", "language": "english", "useDefaultStopWords": true}, or a delimiter or n-gram analyzer) was refused as "unknown analyzer"; it now works as on every other build.
  • The free image's compose file uses the same volume, port and binding as the Docker page.

Docs

  • The clustering page is now clearly a design preview: clustering is under development and cannot be enabled in current releases.
  • Every v2 example in the docs was re-run against a live server, and the ones that were wrong are fixed.
  • The performance page says which write mode its write-path figures were measured in.

v2.11.1

The first free Docker image built from the current engine. If you are upgrading from the 1.0.x image, read the authentication change first.

Changed

  • Authentication is on by default in the free image. The 1.0.x images answered requests without a key. From v2.11.1, the image generates an API key pair on first boot, stored in /io/state/.aito-api-keys, and every request needs an x-api-key header. Clients written against 1.0.x must send one. See Upgrading from 1.0.1.

Fixed

  • The free Docker image starts when a licence key is set. In free images before v2.11.1, setting AITO_LICENSE_KEY left the container running but never serving: it logged the licence check result and then waited indefinitely, so neither the HTTP nor the SQL port answered. Images started without a key were not affected.
  • If the licence check cannot complete, the server now starts in free mode after AITO_LICENSE_STARTUP_TIMEOUT seconds (default 60) instead of waiting indefinitely.

v2.11.0

A correctness release for linked data and SQL. Several queries that used to return a confident wrong answer now return the right one or a clear refusal: sorting and selecting through links, _relate with more than one where condition, SQL joins and GROUP BY, and v2 queries on v1-engine tables. Predictions that route to people are fairer to new candidates and read each link's own rows. Link-target predictions and graph filters are faster, and the v1 streaming insert takes loads of any size again.

New

  • A Python SDK page. Python SDK covers pip install aitoai, the v2 client that aitoai 1.0 makes the default, querying, errors and warnings, and loading data. The quickstart shows the same call from Python.
  • More worked examples. Seven new v2 use cases (recommendations, smart search, autocomplete, autofill, product analytics, quality monitoring, price and demand), and the playground now covers every read endpoint: _relate, $why, _evaluate, graphs, and common mistakes with the error each one returns. See use cases.

Improved

  • Link-target predictions are faster. An unscoped prediction of a linked field, such as a processor chosen from every employee, answers in about half the time on a 1M-row table. The first such prediction after a restart or a write also starts faster, and scoped predictions on large candidate sets cost less.
  • Filtered $refs conditions are much faster. A $refs โ€ฆ $exists with a condition object such as {"relation": "is-a"} runs about 19ร— faster on large graphs, with identical answers. See graphs.
  • $predictions says its confidence is in-sample. A response that selects field.$predictions now carries a select.predictions_in_sample warning: each row is still part of the data its own prediction is read from, so the confidences read higher than held-out accuracy. Use _evaluate for a number you can quote. See the query reference.
  • Request size limits answer with 413. An ordinary request body is limited to 10 MB, and a larger one gets 413 with a message that names the limit, where it used to get a 500 or a closed connection. See limits.

Changed

  • Regression, by design: a candidate with no rows is ranked neutrally. When a prediction ranks linked rows (a new approver, a new hire), a candidate that appears in no row used to inherit the negative evidence of other candidates, including ones the query's scope had excluded, so it ranked below every observed candidate. It now carries no evidence either way and ranks by its base rate. The cost, measured on this release: on the invoice-routing benchmark, the acceptor mean rank on the v1 engine goes from 2.04 to 2.15. Top-1 accuracy is unchanged.
  • SQL INNER JOIN drops rows whose link does not resolve. A join used to keep rows whose foreign key is NULL or names a missing row, like a LEFT JOIN, which inflated count, sum and LIMIT results. Counts over such joins go down. A dotted path without JOIN keeps the link-projection behaviour. See JOIN along a link.
  • A SQL JOIN that does not follow the declared link is refused. If the ON compares a column that is not a link, a link to another table, or a column other than the link's key, the query now fails with SQLSTATE 0A000 and shows the correct ON. It used to return another table's data.
  • A SQL self-join is refused. โ€ฆ FROM employees a JOIN employees b ON โ€ฆ used to pair each row with itself. It now fails with SQLSTATE 0A000, so clients such as DuckDB and postgres_fdw can fall back and do the join themselves. See the SQL reference.
  • A SQL literal whose type cannot match its column is refused. WHERE price = 'abc' (also through a joined column) used to return 0 rows silently. See the SQL guide.
  • v2 queries on a v1-engine table refuse a path they cannot follow. On a type: "table" table, a where path of two or more links, or a misspelt path, now answers 400 naming what failed. It used to be ignored, and the query returned every row. A table follows one link; a path through two links needs a collection. See graphs.
  • Thin text evidence falls back to its words. When the rows containing all the words of a Text value in a where are too few to measure, a prediction now uses what the separate words say instead of the base rate. This can move $p by a few hundredths on Text columns; the top answer was unchanged in our benchmarks.
  • _evaluate on a collection explains what it cannot evaluate. On a collection, _evaluate runs predict. Asking for estimate, recommend, relate, match or similarity now answers a 400 that says so and gives the workaround. See evaluation.

Fixed

  • Two links to the same table each read their own rows. When one table has two links to the same target (for example approver and handler, both to employees), a prediction through one link could read the other link's rows. On the expense benchmark, handler top-1 on the v2 engine goes from 52.6% to 81.6%. With the employee attributes shuffled as a leakage check, it falls back to 61.25%.
  • _relate and $patterns treat a multi-key where as one condition. {"field": โ€ฆ, "customer_id": โ€ฆ} is now the conjunction, as in _query, _predict and v1. Before, each key was related separately, so lifts and frequencies came from one key's rows, and a value could appear twice with different lifts.
  • Sorting keeps rows without a value. An orderBy on a nullable column or through a nullable or dangling link used to drop the rows that have no value, in JSON and in SQL ORDER BY. They are now returned, listed last in both ascending and descending order. See sorting.
  • Selecting through a dangling link returns null. A link that names a missing row used to fail the whole query with a 500.
  • _relate on a linked field relates only what co-occurs. It used to list every value of the linked field, even ones that never appear with the where, and a where that matched nothing still related everything. It now behaves as it does on the table's own fields.
  • $exists through a link that names a missing row. A row whose required link points at a row that does not exist is now treated like a NULL link. $exists: true used to count it.
  • $not through a link on a v1-engine table. Through a nullable or dangling link, $not no longer matches rows that have no linked value.
  • _predict on a nullable link that holds NULLs. It answered 500 as soon as a row actually used the null.
  • $length and $distinctLength ignore empty values. Over a $refs projection of a nullable attribute, an empty value was counted: teams {red, NULL, red} counted 3 and 2 distinct, not 2 and 1. See graphs.
  • SQL GROUP BY returns the NULL group. Rows whose group key is NULL or unresolved were left out of the result. They now form their own group. See the SQL guide.
  • $in next to another operator on the same field works in JSON and SQL. WHERE k IN (โ€ฆ) AND k IS NOT NULL used to fail with Unknown operation: $in.
  • SQL table aliases keep the table name's case. SELECT p.name FROM Products p used to be read as a link path and fail.
  • basedOn $patterns with a stopword. A where text value containing a word the analyzer removes, such as the or for, used to fail the whole prediction. Such words are now skipped.
  • The v1 streaming insert streams again. POST /api/v1/data/{table}/stream had been limited to 8 MiB. It now takes a load of any size, chunked or with Content-Length. See limits.
  • v1 _estimate and _aggregate answer in the v1 shape on v2 tables. Through /api/v1 they returned the v2 envelope. The v1 _aggregate conditional count {"$f": {โ€ฆ}} and the v1 _relate array form (relate: [{โ€ฆ}]) now work too. A v1 _relate with where: {"$exists": "product"} related nothing, and now relates the product's fields as v1 does. See the v1 reference.

Also in this release: an x-knowledge column now uses the phrases it learned as evidence in predict, and $why shows them as $phrase. x-knowledge is experimental and may change between releases. See schema design.

v2.10.3

Fixes for linked data: a misspelt path is now refused instead of quietly answering, a condition that crosses two links works on graphs whose keys are named after their entity, a predict reads a linked field through the right link, and $why explains a same-member $refs filter.

Fixed

  • A condition two links away follows the link's declared key. Inside a $refs filter, {"$exists": {"contact_id.role": "CTO"}} crosses a second link. On a graph whose keys are named after their entity (contact_id, not id) the condition was dropped, and the filter answered every row. It now follows the key the link declares, and a row whose link is empty satisfies no condition through it. See graphs.
  • Each link reads its own linked rows. When two links in a database point at tables that share a key-column name and a field name, a _predict reading that field through one link (for example basedOn: ["name"] on a linked target) could read it through the other. Each link now reads the rows of the table it names.
  • $why explains a same-member $refs filter. A query whose where held a same-member condition such as {"$refs.edges.subject": {"$exists": {"relation": "is-a", "target": "CTO"}}} answered 400 as soon as it selected $why. It now explains the condition like any other. The refusal of an unsupported same-member shape also reads as a sentence again, where it had failed with "invalid escape". See graphs.
  • $patterns across a nullable link. Mining patterns across a link whose value is empty on some rows, or names a missing row, answered 500. Those rows now simply contribute no linked value. See the knowledge-graph use case.

Changed

  • A misspelt path is refused. A where on a dotted path that does not resolve (a misspelt field at any hop, a misspelt column inside a $refs filter, or a path through a self-link) now answers 400 naming the failing segment, instead of 200 with an empty or unconditioned result. See Common errors.

v2.10.2

Two knowledge-graph additions โ€” counting independent sources, and mining patterns across a link โ€” plus two fixes: loading data into v2-engine tables through the v1 API, and _relate properties that name a whole text value through a link.

Added

  • $distinctLength โ€” weighted corroboration. $length over a $refs path counts the referring rows, so ten edges filed by one source read as ten corroborations; $distinctLength counts the distinct members โ€” the number of independent sources behind a claim. SQL: aito.distinct_cardinality(col). See graphs.
  • $patterns mines across a link path. {"$patterns": ["body", "cust.seg"]} now mines a linked field alongside the table's own, so a mined rule can span the join. See the knowledge-graph use case.

Fixed

  • A v1 bulk load into a v2-engine table works. Loading rows through POST /api/v1/data/{table}/stream, or through a file upload (/api/v1/data/{table}/file), into a table on the v2 engine answered 500 and loaded nothing, and a file upload's job ended as failed. Both now load the rows. This is the path the Python SDK's upload_file and the console's CSV upload use. See the v1 reference.
  • _relate accepts a whole text value named through a link. A $props entry such as "product.name": "Pirkka banana", where name is a text field on the linked table, was refused as "no rows carry" that value. It is now accepted.

Changed

  • A v1 bulk load into a v2 table with a primary key or identity column is refused. It answers 400 instead of loading, because that path cannot enforce the key; a file upload is refused at its initiate call, before anything is uploaded. Load such a table with POST /api/v2/data/{table}/batch, which does. See the v1 reference.

v2.10.1

No engine changes. A documentation release: the responses printed in the quickstart and the sandbox walkthrough are now GENERATED from a test run against the sandbox's own data, so what a page shows is what the sandbox returns. Several pages contradicted themselves or each other, and those are corrected.

Changed

  • Printed responses are generated, not pasted. Every response shown in the quickstart and the sandbox walkthrough is captured by a test that loads the sandbox's data, runs each documented call, and lifts the answers onto the page. They reproduce the live sandbox to the last digit. A pasted response drifts the moment the engine's numbers move; a generated one cannot, and the page fails its own test instead.
  • The calibration claim says what was measured. It is scoped to its evidence โ€” 26 held-out invoices, where predictions made at 82% confidence were right 100% of the time โ€” rather than stated as a general property, and the inference guide now tells you to measure calibration on your own data with _evaluate. The page also said support-tempering was on by default; it is not, and there is no such node in $why.

Fixed

  • The _evaluate example in llms.txt works verbatim. It returned a 400. Copy-paste now runs.
  • An unknown where field is documented where it bites. The contract โ€” and having โ€” appear in the query parameters and in a new inference section, Candidates, evidence, and $f, which states the candidate-universe rule, what $f counts, how thin candidates are smoothed, and what happens on a cold start.
  • The benchmarks overview contradicted itself. Invoice routing is measured on a 2,000-row hold-out; one place said 200. A literal 95%% rendered as written. See benchmarks.
  • _evaluate's metrics are defined. h, baseSamples, baseGmp, geomMeanLift, rankGain and features were returned but never explained. See evaluation.
  • Links that went to the wrong place. The jobs link now resolves to limits, and the quota links are page-qualified.
  • Every internal link resolves. 32 broken links are fixed โ€” stale operation paths, relative v1 links, and links that ignored the documentation's base path โ€” and the docs build now checks every internal link before publishing.
  • The GL-coding use case prints real responses. It is generated from the same test run as the quickstart; the previous page named the wrong approver and printed a meanRank no evaluation can produce.
  • baseError is documented correctly. It is 1 minus the majority class's share of the TRAINING rows, not 1 minus baseAccuracy (which is measured on the test rows), so the two need not add up to 1. See evaluation.
  • The _evaluate reference example shows today's response. It had been captured from the older engine and showed that engine's shape.
  • The sandbox GL example uses evidence that matches. It queried a description no row had, so it showed only the prior.

New

  • A multi-tenant pattern in schema design. How to serve many tenants from one shared schema. A nested from scopes base rates and a plain field's candidates to one tenant, but not a linked target's candidates; the guide gives the two patterns that do work for those โ€” a per-tenant linked table, or a tenant-scoped having step before the predict. If tenants must never see each other's values, use the per-tenant table: having shapes what a query returns and is not access control.
  • config.ai lists v2. The alias existed and was not documented.

v2.10.0

Migration fixes, and one behaviour change to check: a multi-word $has on a Text column is now refused rather than answered. The v1 endpoints answer over storage migrated to the v2 engine โ€” five of them used to fail โ€” several ranking defects that produced a confident wrong answer are fixed, and deleting rows now takes effect on predictions immediately rather than at the next optimize.

Changed

  • A multi-word $has on a Text column is now refused. $has names ONE analyzed token; a value that analyses to several โ€” {"name": {"$has": "BASF Oy"}} โ€” returns a 400 naming the operator that means what you asked, instead of an answer. On classic tables that request previously returned no rows, silently, so an application filtering this way got an empty result where it now gets an error. Single-token $has is unchanged, and still analyzes the value, so {"$has": "BASF"} finds the stored token basf. Write instead: the bare value for the whole value (exact equality), $match for all of its words, $search for phrases, OR and negation. If you wrapped a plain value in $has to work around the filter behaviour of an earlier release, that wrapper is the shape that now fails โ€” send the bare value. See the query reference.

Fixed

  • An unknown where field inside $and/$or/$not matched everything. A condition on a column that does not exist โ€” a typo, or a field a schema change removed โ€” was refused bare but silently DROPPED inside a combinator, leaving the surrounding operator unconstrained. {"$or": [{"typo": "x"}]} returned every row of the table where the same condition on its own returned none. A query whose filter was meant to narrow a population therefore returned the whole of it, with no error. The unknown field is now resolved wherever it appears, and warns. See the query reference.
  • /api/v1x answers in v1's shapes, not the raw v2 body. The endpoint exists so a v1 client whose storage moved to the v2 engine keeps parsing what it always parsed, and it was serving v2 bodies instead: _predict returned $value where a v1 client reads field and feature, _recommend returned $value for a flattened item record, _relate returned {field: value} and a numeric info where v1 reads {field: {"$has": value}} and an info object, and _evaluate wrapped its metrics. All four now match /api/v1. See the v1 reference.
  • The v1 endpoints answer over storage migrated to the v2 engine. _evaluate, _match, _similarity, _estimate and _aggregate returned 400 failed to open '<table>' โ€” naming a table that visibly exists โ€” against a table whose storage had been migrated. Unmodified v1 clients now keep working after a migration, which is what the migration promises. See the v1 reference.
  • _relate over migrated storage answers v1's own spellings. A v1 request using relate: "field" (a bare field name) or a nested from failed with a 400; and where it did answer, related and condition came back as {"field": value} where a v1 client reads {"field": {"$has": value}}, so the relation could not be named from the response.
  • _evaluate scored every test row from the prior when its evidence was bound through a link. {"$get": "product.category"} silently bound nothing, so the model reported an accuracy sitting exactly on its own baseline. An unresolvable path is now refused by name instead. See evaluation.
  • A nullable column's ABSENCE counts as evidence again. Where a column only applies to some rows, "no value" is often the signal โ€” and after moving to the v2 engine a model reading that column scored at its own baseline instead. On a 38 000-row demo predicting a customer segment from one nullable column, the v1 engine scored 0.745 and the v2 engine 0.475, its base rate; the two now agree at 0.745. Evidence that reads a value the row does not have states the absence, rather than contributing nothing.
  • Conditioning on absence no longer fails. {"col": {"$exists": false}} counted the right rows as a filter but returned a 400 naming an internal bitset width when used as evidence for a _predict โ€” whenever the column's last row happened to carry a value. Both uses now answer.
  • An exclusion on the predicted field survives basedOn. {"product": {"$not": x}} sent with basedOn kept the excluded candidate's slot and its probability and dropped only its value, so the response carried a hit with a $p and no $value โ€” at rank 1.
  • A link candidate with no training rows no longer outranks one that was observed. Predicting a link target under a filter that narrows the population โ€” an employee within one company, say โ€” ranked never-seen candidates above candidates with real history, because the two were scored against different populations.
  • A multi-word text value is scored at the strength of the whole match. Each word was scored at one word's strength and the duplicates then collapsed, so a four-word product name that matched exactly was ranked as weakly as a one-word coincidence.
  • A deleted row no longer counts towards a prediction. Deleting rows removed them from every query โ€” a filter returned none of them, and counts excluded them โ€” but not from what the model had learned, until the database was next optimized. So a value whose rows had all been deleted could still be offered as a prediction, at the probability it had before the delete: asked for a recommendation, Aito could name a product you had removed. With evidence pointing at it, it could do so confidently. A delete now takes effect immediately and means the same thing whether or not the database has been optimized since. This also corrects _evaluate, which holds its test rows out by deleting them and was therefore scoring a model against a population that still contained every held-out answer.

Improved

  • /docs/api/ is now the v2 documentation. The v1 reference moved to /docs/api/v1/ and stays fully supported; existing deep links into the old root resolve to their v1 section. See API v2.

v2.9.2

Two query bugs and two speedups on the plain database surface. A range filter now answers on the v1 engine โ€” it used to return a server error โ€” and relate on a text column counts the rows it says it counts. The capacity and head-to-head benchmark figures are re-measured.

Fixed

  • A range filter on the v1 engine returned a server error. A two-sided window โ€” {"id": {"$gt": 51200, "$lt": 52200}}, a SQL BETWEEN, or any two comparisons on the same column โ€” failed with an internal bit & operation requires same sized bit sets instead of answering. One-sided comparisons were always correct, and the v2 engine was never affected. See the query reference.
  • relate and $patterns on a Text column reported the wrong support. A mined rule shown as one whole value was counted over every row sharing its words, so a rule displayed as a company with 245 invoices could report 331 โ€” the rows of that company and of every other name containing the same words. Text conditions now replay to exactly the rows their support counted. See the query reference.

New

  • Choose which index of a text column you relate โ€” .$distinct and .$token. A Text column carries both its distinct whole values and its analyzed tokens; relate, $patterns and $props can now name either (product.name.$distinct, product.name.$token). $patterns offers both pools by default and ranks them together, so a specific whole-value rule is available where it is the better one. Asking a column for an index it does not have is refused by name. See the query reference.

Improved

  • Range and comparison filters no longer cost more as the table grows. $lt $lte $gt $gte on numbers, strings and timestamps, plus $startsWith, $mod and the timestamp component operators, allocated a bitset per matching value โ€” so a comparison keeping many of a column's distinct values grew quadratically with the table. At 102 400 rows a 1 000-wide window went from 378 ms to 20 ms.
  • Relational aggregates resolve their fields once per query. select: ["$count", "$avg:amount", "$min:id", "$max:id"] under a filter went from 111 ms to 2.5 ms at 102 400 rows on the v2 engine. See the query reference.

Changed

  • The capacity page reports the memory a server process uses, not the benchmark harness's. The old figures were read from a JVM that also held the test corpus, and were wrong in both directions. At 10 000 000 rows v2 holds the database in 213 MB of heap (published: 1 720 MB) and v1 in 5 290 MB (published: 2 603 MB) โ€” so a 10M expense database does not fit a v1 engine in a 4 GB heap. See capacity.
  • The Elasticsearch head-to-head is re-measured. The previous table timed one query shape at a time, so the first shape measured absorbed the whole stack's warm-up and the last looked fastest โ€” which is why it showed filter+text beating both of its own constituents. The shapes are now interleaved on an idle machine. See performance.

v2.9.1

No engine changes. The scaling figures are re-measured on an optimized database โ€” the configuration a deployment actually runs โ€” and now lead with the warm median rather than a mean that averaged the first requests after a restart into steady state.

Changed

  • The scaling tables measure an optimized database, and lead with warm latency. Two things moved together. The measurement now runs against an optimized database rather than an as-ingested one, and each row leads with the warm median, with the warm p90, the mean including the cold start, and the first query beside it. At 10 000 000 rows v1 reads about 210 ms at the warm median (417 ms p90, 355 ms mean including the cold start) and v2 about 279 ms (823 ms, 601 ms). The v2.9.0 page reported the as-ingested database with the mean: 853 ms median and 1 638 ms mean for v1 at the same scale. Those as-ingested figures stay on the page as the comparison, so the cost of leaving a database uncompacted is still visible. The engine is unchanged from v2.9.0. See scaling.
  • The capacity page says plainly that an as-ingested database answers more slowly than an optimized one, and points at the warm-up breakdown. See capacity.

v2.9.0

The v2 API is in beta. Recommendations now use their query and your basedOn in earnest, and relate and _evaluate reach per member and per word. One change needs action: in a filter, a bare value on a Text column now matches the whole value exactly โ€” use $match to filter by words.

Changed

  • A bare value on a Text column filters by the exact value. In a filter โ€” _query, _search, _aggregate, a SQL =, a delete โ€” {"name": "BASF Oy"} now selects the rows whose name is exactly BASF Oy (case-sensitive), matching v1. It used to match the words in any order, so it also returned โ€” and a delete also removed โ€” "BASF Construction Chemicals Finland Oy". To filter by words, use $match. As evidence in _predict, _recommend or _match, or when ordering by $p, a bare Text value still counts word by word. See the query reference.

New

  • relate a linked text field word by word. relate: ["product.name.$feature"] returns one relation per word of the linked product's name, as that column's analyzer produces it. See the query reference.
  • _evaluate measures per-member predictions and combined test selectors. predict: "tags.$feature" evaluates a multi-label prediction member by member, and a test selector can combine $index with a field condition. See evaluation.
  • Two public benchmarks, each with its data, its baselines and its numbers: smart search & recommendations and invoice line โ†’ product matching.

Improved

  • Recommendations use their query and basedOn. _recommend now scores every candidate against its context exactly, where most candidates used to be ranked by how likely they were to meet the goal regardless of the query. On our smart-search benchmark, a recommendation from query words alone went from nDCG@10 0.12 โ€” about random โ€” to 0.45, level with BM25 over the catalogue, and reaches 0.53โ€“0.55 with basedOn. A For You shelf ranked from the shopper's profile went from 0.11 to 0.18, and 0.24 with basedOn. Rankings will differ from v2.8.4.
  • A text where through a link resolves a new word in about half a millisecond, where it took about 180 ms.
  • A v1 match against a link target of over 100 000 rows answers in about 2 s instead of about 12 s.

Fixed

  • A link-target prediction scoped by an attribute of the target honours that scope again. A regression in v2.8.4.
  • basedOn applies to _recommend. v2.8.4 accepted it and ignored it.
  • relate on a set field through a link relates each member, as it does on the field itself, instead of each whole set.
  • A text where through a link counts each word as evidence, as the same where on the table's own text column does.
  • Statistics caches no longer serve a result computed for different data. Two caches were keyed on only part of what they read, so a query could answer 200 with statistics from other data; they now key on all of it.

Also in this release: an experimental x-knowledge column type and its $surprisal operator. Experimental surfaces carry no compatibility promise and may change between releases โ€” see schema design.

v2.8.4

The v2 API is in beta. A latency-and-correctness release: the largest single cost of a scoped link-target prediction is gone, a linked column now supports every operator its direct counterpart does, and six ways a query could return the wrong rows or candidates with a 200 are fixed.

New

  • Every where-operator works through a link. $gt/$gte/$lt/$lte, $mod, $numeric, $startsWith and $search apply to a field reached by its dotted path ({"part.price": {"$gte": 10}}) exactly as to the column itself โ€” the operator vocabulary is a property of the field's type, not of how you reached it. A bare value, $or or $in against a linked Set/Array column is membership, as on the column, and a column value operation selects through the link ("select": ["part.desc.$tokenCount"]). See fields through a link.

Improved

  • A scoped link-target prediction answers in a fraction of the time. When a prediction is scoped by a where on a linked path (invoice.customer: X), the engine walked every candidate value of the link โ€” every seen and every unseen target row โ€” to test which ones the scope admits; it now derives the candidates from the linked rows the scope admits. The candidate universe of a link target is also built once per database state rather than once per request, and the per-candidate evidence selection does its grouping once. On a 16 000-candidate scope a fresh request went from about 0.8 s to about 0.6 s; on a few-thousand-candidate scope it answers in about 0.3 s. This particular change does not affect the answers. See improving inference.
  • Matching is faster and more accurate. When ranking candidates, the engine built a pool of relations twice the size it could keep, refined all of them, and then discarded the surplus. It now refines only what it keeps. On our ERP matching benchmark the realistic basedOn arms gain up to ten points of top-1 accuracy (for example 62.5% to 72.5% matching on name and supplier, mean reciprocal rank 0.758 to 0.817) and the sweep runs in about half the time; an expense-categorization run drops from about 157 ms to about 108 ms per query, and a 10 000-row invoice match from about 840 ms to about 320 ms per probe. Because the pool now holds different relations, individual rankings and the factor weights reported by $why can shift slightly in either direction. See improving inference.

Fixed

  • A $not, $and or $or on a linked path now scopes the candidates and the evidence, not only the rows. {"product.id": {"$not": 0}} restricted which rows were counted but left the candidate pool and the evidence drawn from the whole linked table โ€” a 200 with plausible content and the wrong candidates. Every combinator over a linked path is now applied to both.
  • A get honours a where on the get field itself. An autocomplete query such as get: "name" with {"name": {"$startsWith": "mil"}} returned candidates from the column's whole distinct set, so unrelated values were offered beside the matching ones. The candidate universe now respects a same-field where the way it respects every other where.
  • Set-, array- and text-valued cells read consistently on the remaining paths, and a bit-set difference refuses a width mismatch instead of adopting the operand's width โ€” the cause of a negation that could answer with zero rows.
  • A field reached through a null link is null. Through a nullable link with no target on that row, $match, $has and $exists counted the row as present โ€” it read the linked table's first row โ€” so {"partN.name": {"$exists": true}} answered every row. Such a row now matches no operator, is excluded by $not, projects as null in select, and is selected by {"partN.name": {"$exists": false}}: the same rules as a nullable column read directly.
  • A modified link target is visible through the link at once. After _modify changed a row in a linked table, a path through that link could still read the row's previous values; the link now always resolves to the live row.
  • Excluding an item from a recommendation no longer offers it. A statement about the item being inferred ({"product": {"$not": 5}} while recommending product) is a filter on the answer, never evidence for it; an excluded item โ€” including one the model had not seen โ€” is absent from the candidates instead of returned at rank 1. Statements about contextual data remain inference evidence, as before.

v2.8.3

The v2 API is in beta. A correctness-and-latency release: two silent wrong-answer fixes the demo apps reported, a bit-operation bug that could return plausible-but-wrong predictions, a broad prediction-latency pass, and a deliberate narrowing of what basedOn is allowed to reach.

Changed

  • basedOn now means exactly the fields it names. A field is paired when a known names it โ€” the engine no longer infers extra Text fields of a linked table whenever your query happened to carry any Text known. That inferred reach helped some shapes and hurt others, but the reason for removing it is that a basedOn is a statement of which evidence is permitted, and silently reaching past it made that statement advisory. If a prediction relied on the implicit reach, name those fields explicitly. Text evidence is now weighted by a fitted calibration rather than a fixed assumption. See improving inference.

Improved

  • Prediction answers faster, across the board. A joined representation now answers a position from the side that owns it (memoised per position set), a linked text column seeks its token postings instead of re-expanding them, numeric value priors are computed when read rather than once per candidate, linked selects resolve lazily, and the rotating cache evicts in constant time instead of scanning every slot on insert. The gains are largest on link-target predictions over wide linked tables.

Fixed

  • $not on a legacy table returned nothing at all. On a type: "table" table queried through v2, {"field": {"$not": value}} returned zero rows instead of the complement โ€” a 200 with an empty body, indistinguishable from "no match". That is the basket-exclusion shape ($and of $not), so a legacy-backed database could serve no recommendations rather than reporting an error. Two independent causes were found and fixed: a numeric lookup that missed even when the value was present, and the negation itself.
  • _recommend ignored a filter on a linked field. A $match on a linked column constrained _search correctly but not _recommend's candidate pool, which was drawn from the whole linked table โ€” 200 OK, plausible content, wrong candidates. Every linked where-operator is now applied to the pool.
  • A bit-operation fast path could return a plausible wrong answer. When two inputs were segmented differently, the dense fast path in or combined the wrong chunks: a shorter input raised an index error, a longer one returned an incorrect result with no exception. It is reachable from the predict path, so some predictions could be quietly wrong. Inputs are now aligned before they are combined.
  • Set-, array- and text-valued columns are read consistently. Membership was re-derived independently in several places, each matching only the shapes its author had in mind โ€” a column reads back as an array, a sequence, a set or a Java collection depending on codec and storage engine. Seven silent degradations are fixed and the derivation now lives in one place.
  • The v2 documentation generator emits the wording its own artifact carries, so the published reference and the generated source no longer drift apart.

v2.8.2

The v2 API is in beta. One breaking change to _match's hit shape (below), the contract fixes behind three defects our demo apps hit while migrating, and a large cut in link-target prediction latency.

Changed

  • _match hits carry $value alone. The v1 spellings feature and field are no longer returned by default: feature was a verbatim copy of $value, and field was always your own request's match parameter, so neither added anything. Every ranked-value op โ€” _predict, _recommend, _match โ€” now reads identically, which is what the $value rename was for. Porting a v1 body that reads them? select: ["feature"] and select: ["field"] still return them, unchanged. See the v2 introduction.

New

  • Rank candidates by an aggregate. select: ["$f"] and { "$sum": { "$context": โ€ฆ } } no longer require the query to also carry a matching orderBy, and orderBy: { "$sum": { "$context": โ€ฆ } } is now accepted. So "which candidate values contribute most to the goal" is one query, instead of pulling the whole candidate set and sorting it client-side. See the query reference.

Improved

  • Link-target predictions answer several times faster on a fresh request. Column-only representations read rows directly, a query runs one prediction instead of one per candidate column, and the $why explanation is built only for the rows that ask for it.
  • A write costs less afterwards. Class counts and linked-attribute buckets are now kept per segment, so appending data re-counts only the segment that changed rather than the whole collection โ€” the first query after a write is much closer to the warm one.
  • Steadier memory on long-lived instances. Caches keyed by data lineage are bounded by the generations still retained rather than by an entry count, so they follow what the instance is actually holding.
  • An attribute counts once per candidate. When several pieces of evidence pointed at the same attribute of a candidate, it could be credited more than once and inflate that candidate's probability. It is now counted once.

Fixed

  • Per-candidate aggregates across a link returned zero. $f and { "$sum" | "$mean": { "$context": โ€ฆ } } over a get whose path traverses a link โ€” product.name, context.week โ€” reported 0 for every candidate, with a 200 and a well-formed body. Direct fields and the link field itself were correct, so a chart could show a confident flat line at zero for a product with hundreds of purchases. They now resolve the candidate's rows through the field that owns them.
  • relate: { "$props": โ€ฆ } accepts set-valued properties. A property held in a Set column (product.tags) was refused with a message saying no rows carried the value โ€” while rows plainly did, because membership and equality are different questions about the same data. Set-valued properties, and a list of values for one property, both work now.
  • An empty $and or $or clause list. Building a filter list that turns out empty โ€” { "id": { "$and": [] } } from an empty basket โ€” failed the request. An empty $and now constrains nothing and an empty $or selects nothing, which is what each means and what v1 already did.
  • A union view's refusal names the branch again. When an output column resolves in no branch, the error names each branch and the columns it tried, rather than only reporting that every projection was absent.
  • The v1 type reference no longer prints internal class names. A few anonymous formats reached the served pages as a Java class name and an object hash instead of their type, in the AnyValue and RegressionExplanation references.

v2.8.1

The v2 API is in beta. A follow-up to v2.8.0: query latency on large collections, better link-target predictions, and a set of correctness fixes around evidence and aggregation โ€” plus one new relate form.

New

  • Relate an entity's own property values. relate: { "$props": โ€ฆ } asks what distinguishes a row's own attributes against the where, so "what is unusual about this customer" is a single query rather than one per field. See the query reference (and the SQL spelling in the SQL reference).
  • config.ai applies on v1 _recommend, _match and _search. Where it genuinely cannot apply โ€” _relate and _similarity โ€” the request is now refused with a 400 instead of being accepted and ignored, so a tuning setting never silently does nothing. See improving inference.

Improved

  • Large collections answer much faster. Broad selects, page fetches and orderBy on a plain field no longer pay a per-query floor; ordering reads the index's own value order, and computed fields are evaluated over the filtered universe rather than the whole collection. The biggest gains are on a scoped orderBy over a computed field, which was the slowest shape at scale.
  • Leaner prediction. Per-state ordinal maps, segment-lifetime column reuse and primitive cursors in the bit loops cut allocation churn on the predict path, so repeated inference on a warm collection holds a smaller footprint.
  • Better link-target predictions. A link candidate is now scored through its target's attributes under a field-pair prior, so a candidate is judged by what it actually looks like rather than only by how often it has been seen.
  • Name matching fades with history. The name-identity boost decays as a row accumulates history instead of being discounted by query coverage, so a well-established entity is ranked on its record rather than its name.

Fixed

  • A where on the predict target restricts candidates, not evidence. Naming the target in where narrowed the candidate set and was counted as evidence for it, which skewed the probabilities; it now only restricts.
  • Deleted rows no longer form a group. A deleted row could still produce its own GROUP BY group; $f and $context aggregates now intersect the live rows.
  • why answers on a large collection. The KNN explanation is bounded, so _estimate with why no longer fails as the collection grows, and it reads evidence rows from the collection rather than the field.
  • basedOn on a text name no longer regresses link-target predict.
  • Env branches survive a binary-format upgrade. After v2.8.0's format bump an env branch failed plain reads with a raw internal error until it was repaired by hand. Branches now migrate on first use, and a read against a not-yet migrated state returns a clear error naming the remedy instead of an internal failure. See environments and common errors.
  • Documentation rendering. Table captions rendered as literal markup on 19 pages, and a dark page backdrop made overflowing text unreadable.

v2.8.0

The v2 API is in beta. This release is mostly speed and memory at scale โ€” text postings and column indexes are now prepared at write time and read straight from memory-mapped storage instead of being decoded onto the heap โ€” alongside a wider SQL surface and explanations that point at the evidence inside a predicted item.

New

  • Derive a table from a query. INSERT โ€ฆ SELECT and CREATE TABLE โ€ฆ AS SELECT build a table in the database from a query over existing ones, so a derived dataset no longer round-trips through your client. See the SQL guide.
  • evaluate() and estimate() as SQL table functions. Measure accuracy and get a prediction's estimate from SQL, with test / testLimit / train / resubstitution arguments to control the split. See the SQL reference.
  • SQL answers in PostgreSQL result format. /api/v2/_sql can answer with rows and column metadata in the pg result shape instead of the hits shape, so a Postgres client renders it directly. See the SQL reference.
  • Multi-column GROUP BY and multi-key ORDER BY. Group by joint tuples and order by several keys, with Postgres's default null placement per key. See the SQL reference.
  • having โ€” filter candidates on the server. Filter ranked results on $f after ranking and before offset/limit, so an "at least N occurrences" cut happens in the engine instead of in your client. See the query reference.
  • Explain a predicted link or text target through its own attributes. When the prediction is a linked row, $why now shows the candidate's own fields as the prior that carried it, rather than leaving that path unexplained. See the inference guide.
  • $highlight on a predicted target. Add it to a link-target predict's select and each candidate comes back with its own field values marked where the evidence sits โ€” ready to display. See the query reference.
  • Structured JSON logs. AITO_LOG_FORMAT=json writes one JSON object per line to stdout, with context such as database inlined as a field, so a self-hosted instance ships into your existing log pipeline with no parsing rules. See running it on your own infrastructure.
  • Union views are served live. A stored union view now answers from its sources once one of them changes, so it cannot go stale; _refresh became an optimisation rather than a correctness step. See schema design.

Improved

  • Faster and much lighter at scale. Columns and token postings are persisted in read-ready form, so every column type opens in constant time and reads stay memory-mapped rather than decoding whole blobs onto the heap. Large states use substantially less query-time memory and answer faster.
  • Better link predictions. Predicting a linked field now draws candidates from the linked table rather than only the values seen in training, so a valid target that never appeared in the training rows can still be predicted and ranked. See the inference guide.
  • The first query after a write. A recommend issued right after a write no longer pays to refit the prior before answering.
  • Ranking on catalogue-shaped data. Name-based identity matching is now bounded by the evidence that supports it instead of a fixed internal ceiling, which improves ranking where many rows share similar names.
  • Quieter error logs. Client-input errors are no longer logged at ERROR with stack traces, so an operator's error channel reflects server problems. See common errors.

Fixed

  • One evaluate train/test contract across engines. An undefined train now means the complement of the test set on both engines, and testSource aligns with it โ€” so the same request reports the same accuracy whichever engine backs the table. See evaluation.
  • Baselines measure the right population. An evaluation's baseline is the model with no evidence, scoped by the query's own filters rather than a wider pool; filtered evaluations previously reported an optimistic baseline.
  • estimate over a nullable numeric column estimates over the non-null subset instead of failing.
  • Explanations on temporal columns. A KNN $why over a v2 collection's temporal columns no longer errors.
  • v1 schema GET renders every scalar type reachable from v2 instead of failing with a 500.
  • relate and similarity. relate reported inverted frequencies, and similarity discarded partial evidence; both fixed.
  • Deleted and superseded rows. A deleted row now releases its primary key, and a superseded row is no longer offered twice as a prediction candidate.
  • Merged Array/Set columns. A width mismatch between the merge writer and reader could decode stored values into wrong ones; fixed.
  • SQL that real tools emit. Comments, positional INSERT, and a FROM-less SELECT now parse. See the SQL reference.
  • pgwire. A duplicate column name keeps Postgres's own answer instead of being renamed, and a data error reports 22P02 / 22P04 rather than a syntax error. See the SQL reference.

v2.7.0

The v2 API is in beta. This release unifies the v2 response contract across storage engines, extends where filtering (Date ranges and let fields), and adds on-prem observability with a Prometheus metrics endpoint.

New

  • Prometheus metrics endpoint. GET /metrics exposes request counts, a request-latency histogram, and JVM/disk gauges in Prometheus text format โ€” per database โ€” so a self-hosted Aito scrapes straight into Prometheus and Grafana. See monitoring your instance.
  • Filter on Date ranges. where now accepts range comparisons on Date columns, so "orders in the last quarter" is expressible directly. See the query reference.
  • Filter on let fields. A field defined with let can now be used in where โ€” equality, $in, and same-field $or membership โ€” so a value you derive filters the same query that defines it. See the query reference.

Improved

  • One response contract across engines. A v2 endpoint now returns the same response shape regardless of which storage engine backs the table โ€” one select vocabulary, one similarity-score key, a consistent error kind, and a response-time header โ€” so client code no longer special-cases the engine.
  • Honest HTTP error codes. Capability and input errors are classified into proper 4xx codes with a machine-readable kind, instead of returning prose in the error field. See common errors.
  • Faster v2 queries. Leaner bitset intersections, single-pass posting-list walks, and segment-aligned per-known tails cut work on large states; the AND-prior evidence gate is now on by default.
  • On-prem operations. An operations runbook and a restore-state command with a backup/restore drill for self-hosted deployments, and production no longer logs at DEBUG by default. See clusters & operations.

Fixed

  • Namespaces and mounted sub-envs no longer 500. SQL DDL/DML and v2 schema/data/query addressed at a namespace or a mounted sub-env returned a 500 (an internal cast error); they are now handled correctly.
  • _evaluate honours basedOn. _evaluate silently ignored basedOn, which voided any measurement run through it against a based-on database; fixed.
  • Filters no longer dropped under a tenant scope. recommend and relate over a v1-backed table could silently drop where filters when run under a tenant scope; fixed.
  • Nullable columns. A nullable column could leak one row's value onto rows that have none; fixed.
  • Empty-tokenising text filters. A text where whose value tokenised to nothing matched the whole table instead of nothing; fixed.
  • SQL over pgwire. Column names are now identical across both transports, and notices are delivered.

v2.6.1

The v2 API is in beta. This release is mostly stabilization โ€” SQL/pgwire correctness and lower memory at scale โ€” with a few additions to the SQL surface.

New

  • Prediction as a SQL table function. SELECT * FROM predict('invoices', 'category', given => 'vendor = ''Acme''') (and predictions(โ€ฆ, k => 5)) runs a hypothetical prediction for a row you don't have yet โ€” the SQL spelling of what _predict does over JSON. See the SQL reference.
  • List/Set member edits in _modify. update's set accepts {"tags": {"$add": ["sale"], "$remove": ["draft"]}} to add or remove members of a list/Set column per matching row โ€” additive and safe under concurrent updates.
  • SQL table aliases. FROM customers AS c and FROM customers c now parse, so BI tools, ORMs and query builders that alias tables (and both sides of a JOIN) work as written.

Improved

  • Faithful transactions over pgwire. A multi-statement request is now one implicit transaction on the extended protocol too (pgjdbc and most drivers), so a failed statement no longer leaves earlier ones committed. DISCARD ALL now fully resets connection state โ€” important for connection poolers.
  • Honest error codes. Capability gaps return 0A000 (feature not supported) instead of 42601 (syntax error), so federating clients (Metabase, DuckDB, postgres_fdw, SQLAlchemy) fall back gracefully instead of reporting your query as broken. Clients can also introspect available functions via a generated pg_proc.
  • A prediction column is bounded by default. SELECT predictions(col) FROM t now defaults to LIMIT 10 instead of running one inference per row over the whole table โ€” matching recommend/relate/search and the JSON API.
  • Lower v2 memory at scale. Several unbounded rep2 caches are now bounded and a per-query mask is leaner, continuing v2.6.0's 10M-scale reductions.

Fixed

  • Correct results in large multi-segment states. Fixed a case where a state with more than 64 segments could silently drop recommendation/prediction evidence.
  • Three SQL engine bugs found by a new conformance suite, including WHERE price < 5 OR price > 50 no longer erroring.

v2.6.0

The v2 API is in beta. This release adds new query surface and continues hardening it.

New

  • SQL is a full query surface now. Rank candidate values with match(โ€ฆ), mine frequent patterns with patterns(โ€ฆ), run a relevance search with aito.search(โ€ฆ) (carrying $why, $highlight and $matches), and find nearest vectors with ORDER BY col <-> '[โ€ฆ]' LIMIT k or the aito.knn / aito.nn conditions. See the SQL reference.
  • Vector predicates for prediction โ€” aito.cluster and aito.semantic as WHERE conditions, and vector(n) columns to hold embeddings.
  • TIMESTAMP columns over SQL. Declare and range-filter timestamps, do now() + interval arithmetic in a WHERE, read components (EXTRACT, date_part, day-of-week in both spellings), and bucket with date_trunc. See the SQL reference.
  • Define your schema in SQL โ€” CREATE TABLE declaring links and analysed text, and CREATE VIEW over the query functions.
  • Real SQL transactions โ€” BEGIN / COMMIT / ROLLBACK are buffered and merged on commit, with working savepoints for partial rollback.
  • Aito's functions live in an aito schema, keeping the public namespace clean while the key concepts stay reachable unprefixed.
  • select: [field.$predictions] returns a field's ranked predictions inline in a v2 query.

Beta stabilization

Hardening across the v2 engine and its SQL surface:

  • More accurate predictions by default. Name-based identity boosting is now off by default โ€” on a 10M-row benchmark it was costing ~30 points of top-1 accuracy by letting a name match take over the score. Predictions that relied on the old default will change, and should improve; the boost is still available and is now bounded, so it can add signal without dominating. See improving inference.
  • Much lower v2 memory at scale. Large v2 predictions allocate far less per call, and the internal firing cache now has a real size budget with eviction โ€” a 10M-row workload no longer balloons memory.
  • No more 15-minute hangs on out-of-memory. A fatal out-of-memory error mid-request used to stall the connection until the request timeout and then return an unusable body; it now fails fast and cleanly.
  • Assorted correctness fixes. A Text column added to an already-populated table now feeds a $text view; and over the SQL wire, a data query that mentions a catalog name is answered from your data (not the catalog), a partial rollback no longer discards the whole transaction, CREATE VIEW validates its functions up front instead of degrading silently, COLLATE is honoured, and GET /version reports the real version.

v2.5.3

New

  • SQL over the Postgres wire protocol is on by default, and the Docker image now ships with API keys configured. Connect psql, JDBC/ODBC or psycopg straight at Aito โ€” see the SQL guide.
  • TLS is required for the SQL listener on a shared port, and SQL sessions are bounded per connection.

Improved

  • A read-only API key can no longer open a read-write SQL session โ€” the SQL surface now honours key scope the same way the REST API does.
  • Each database is authorized by its own keys, not the server's.

Fixed

  • $why explanations keep their highlights when scoring composes several factors โ€” highlight output no longer nests one level too deep.
  • A refused SQL listener no longer takes the server down with it; a SQL session no longer outlives the database it was opened against.

v2.5.2

Fixed

  • $why lifts and highlights are reported at the level clients expect, so $highlight / $matches results render correctly again.

v2.5.1

New

  • Hybrid search: combine BM25 text relevance with vector similarity in one ranked query โ€” $vectorSimilarity and the self-calibrating $vectorIdf blend with $p. See vector search.
  • /api/v1 runs on the v2 engine for the query family, so v1 clients get v2 performance without changing a line.
  • The vector dimension cap is lifted (previously ~1017 dimensions).

Improved

  • select accepts $highlight / $matches as aliases for $why highlights.

Fixed

  • Cross-tenant isolation: $and / $not over a linked field no longer returns rows from another tenant's data.
  • _match on a text column now predicts per token, matching v1 behaviour.
  • _estimate restores v1 per-field attribution in why, and defaults to AdjustedKNN as v1 does.
  • Several v1โ†’v2 response-shape differences corrected (_predict, _recommend, _relate, _match), so a v1 client sees the shape it expects.