Multi-tenant isolation (v2)
One Aito collection can hold many customers' rows โ an accounting SaaS with
every client's invoices in one invoices table, say. This page answers one
question for each query type: when a query is scoped to tenant A, can
tenant B's rows change A's answer? It then gives the patterns that keep them
out.
Everything here is checked by a test that builds two databases that differ only in tenant B's rows and requires tenant A's answer to be byte-identical in both. The tables below are that test's result.
Isolation here is about statistics: which rows a model learns from. It is not access control. Check on your server that the caller may act for the tenant they name, and send the tenant scope from there โ never from the client.
where is evidence, not a scope
The mistake to avoid first: putting the tenant into the where of an inference
query.
{ "from": "invoices",
"where": { "tenant": "A", "vendor": "Acme" },
"predict": "gl" }
In a _predict, _recommend, _relate or _evaluate, the where is
evidence: what you know about the row you are asking about. tenant: "A"
becomes one condition among others. The base rates, the vendor's lift and
every other statistic are still learned from all tenants' rows, so tenant
B's bookings for Acme move tenant A's answer. The test confirms it: with the
tenant in where, every model-based query differs once B's rows exist.
To learn only from tenant A, restrict the population with a nested from
instead.
The nested from scopes the population
{ "from": { "from": "invoices", "where": { "tenant": "A" } },
"where": { "vendor": "Acme" },
"predict": "gl" }
The inner where masks every other tenant's rows out of the table the query
sees. The outer where is evidence as usual. The model is then built from A's
rows alone: base rates, lifts, $why factors and the candidate values of a
plain column. See
Query reference โ Population restriction.
A link target also needs a candidate filter
When the predicted field is a link, its candidates are the rows of the
linked table, whichever rows the nested from keeps. If every tenant links to
one shared accounts table, a nested-from predict of account still ranks
tenant B's accounts, at a small probability. Their names, at least, reach
tenant A.
Add a filter on the link target's own tenant column:
{ "from": { "from": "invoices", "where": { "tenant": "A" } },
"where": { "vendor": "Acme", "account.tenant": "A" },
"predict": "account" }
account.tenant sits under the predicted field, so it narrows the candidate
set instead of acting as evidence
(Inference โ Scoping candidates per customer).
You need both, because each does one half:
| the query has | candidates | statistics |
|---|---|---|
a link filter account.tenant only | A's accounts | every tenant's rows |
a nested from only | every tenant's accounts | A's rows |
a nested from and the link filter | A's accounts | A's rows |
Only the last row is isolated. The alternative is a per-tenant linked table (one
accounts_<tenant> per tenant), which scopes the candidates by construction.
What is isolated today
"Isolated" below means that tenant A's answer is byte-identical whether or not tenant B's rows exist. That holds with B's rows in one segment with A's, in their own segment, and in mixed segments.
| query | scoped by a nested from (+ link filter) | tenant in where |
|---|---|---|
_query / _search (rows) | isolated | isolated (a filter here) |
_predict, plain column | isolated, seen and unseen evidence values, text evidence | not isolated |
_predict, link column | isolated with the link filter; not without it | not isolated |
$why | isolated | not isolated |
_recommend | isolated | not isolated |
$nearest / $knn | refused: 501 | isolated (a filter here) |
_evaluate | isolated: test and training rows both come from the population | not isolated |
_match | isolated; a link target needs the link filter, as for _predict | โ |
_estimate, _aggregate, _similarity | refused: 501 | โ |
_relate (field list) | refused: 501 | not isolated |
_relate with $patterns | isolated (support counts and lifts over the population) | not isolated |
The refusals fail loudly, never with a quiet answer computed over every tenant. What to use for those today:
- Vector search (
$nearest,$knn) is a ranking over rows, not a model. So the tenant condition in itswhereis a filter:{"tenant": "A", "$nearest": {โฆ}}returns only A's rows, identical to a table holding only A. - Plain
_relate,_estimate,_aggregateand_similarityhave no isolated form on a shared collection (_relatewith$patternsdoes: see the table). For these, use one of the physical layouts below.
Available since: v2.11.3 for the _evaluate and _match
rows and the 501 refusals. On v2.11.2 and earlier _evaluate and _match
refuse a nested from with a 400, and the other three answer
400 missing 'from'; measure a tenant's model there with a per-tenant table
or view.
Physical isolation: a per-tenant table or view
A table that holds only tenant A's rows is isolated by construction, for every query type. There are two ways to get one.
A table per tenant (or a database per tenant) is the simplest, and it
serves the query types above that a nested from cannot.
A SQL view per tenant keeps one shared table for writes and materialises each tenant's rows as a collection:
CREATE VIEW invoices_tenant_a AS
SELECT id, vendor, description, gl, account, approved
FROM invoices WHERE tenant = 'A';
A view is a real collection, so every query runs on it unchanged:
"from": "invoices_tenant_a". The test checks that _predict (seen and unseen
vendor), $why and _recommend on the view are byte-identical to a table
holding only A's rows. A link predict with a link filter and basedOn on the
view equals the nested-from answer.
What to know about views:
- A write to the source rebuilds the view. A view is materialised, not a
stored
SELECT: a refresh skips an unchanged source, and after any write it rebuilds the whole view. Views suit read-heavy tenants; a table per tenant suits write-heavy ones. - List the columns if the source has a
Vectorcolumn. A view that copies aVectorcolumn (includingSELECT *over such a table) is refused atCREATE VIEWfor now. Leave the vector column out of the view. - Create views over the Postgres wire protocol. The REST
_sqlendpoint runs read-only statements only (see SQL โ Views). - A view still links to the shared linked table, so a link predict on it needs the link filter from above.
Pooling with consenting tenants
A tenant with little history can borrow from others that agree to share. Give
the nested from a set of tenants:
{ "from": { "from": "invoices", "where": { "tenant": { "$in": ["A", "C"] } } },
"where": { "tenant": "A", "vendor": "Initech" },
"predict": "gl" }
Tenants outside the set have no influence: the answer is byte-identical with or
without tenant B's rows, in both the $in and the $or spelling. Inside the
set, the tenant in the outer where is evidence, so A's own history weighs
more than C's without shutting C out.
That weighting has a crossover. On a synthetic test, tenant A booked vendor Zeta to its own account, 8000, and 300 pooled rows from C booked Zeta to 9000. A's own coding became the top answer only at about 64 own rows. Below that, the pooled answer won. Pool while a tenant is new. Once its own history is substantial, weigh whether to switch it to its own population. The crossover depends on your data, so measure it on your own evaluation set.
Performance
A nested from is implemented as a per-query transient mask of the other
tenants' rows. Measured on one 128 000-row collection of 255 tenants, a scoped
_predict took 1.0โ2.3 s (median) against 0.2โ0.5 s for the same predict
without the scope โ roughly a 1 s floor per query, largely independent of the
tenant's size. How the cost scales with the total row count is under
benchmark. For large multi-tenant deployments, a per-tenant collection (a
table or a view) is the recommended layout.