Relationships, Links & Graphs (v2)
In Aito, related rows are connected by links, not copied together by joins. A
link is a foreign key a column declares ("link": "products.id"). Once declared,
you address a linked row's fields by a dotted path (product.name) and
project, filter, sort, and predict through it β there is no row-multiplying
cartesian product. A join is a path projection: nothing is copied.
That one mechanism has four spellings, and they all lower to the same link-path projection in the engine:
| Spelling | Where | What it does |
|---|---|---|
Dotted path β product.name | any where / select / orderBy | follow a forward link and read a field |
SQL JOIN β¦ ON | _sql | the same forward link, in SQL syntax |
| Link-join view | Schema Design | a saved link exposed as a navigable column |
Inline join from | Query Reference | an ad-hoc link declared in the query |
This guide covers the forward direction (links and joins), the inverse
direction ($refs β reading a row from the rows that point at it), and the
knowledge-graph patterns those two unlock.
Links are joins
Say orders has a product column linking to products.id. To read each
order's product name and category, follow the link with a dotted path β no join
clause, no duplicated rows:
{
"from": "orders",
"select": ["id", "product.name", "product.category"]
}
product.name resolves each order's linked products row and reads its name.
The identical result in SQL is a join along that link:
SELECT o.id, p.name, p.category
FROM orders o JOIN products p ON o.product = p.id
Both lower to the same projection β the SQL JOIN is a thin skin over the link
path, not a cartesian product (see SQL & Postgres β Joins). Filter
and sort through links the same way (where { "product.category": "dairy" },
orderBy "product.price"). To expose a link as a first-class navigable column
without repeating the path in every query, save it as a link-join view (see
Schema Design β Link-join views).
Knowledge graphs: store flat, present grouped
Everything above is the forward direction. The rest of this guide reads a link
the other way β and that is what turns a set of links into a queryable graph.
Aito models graphs without a graph database: a graph is just two collections
and a back-link β entities, the edges between them, and the $refs operator
that reads an edge from the entity it points at. The payoff is a knowledge graph
that is probabilistic, explainable, and incrementally writable β predict
missing edges, classify a node from its neighborhood, and ground every answer in
the edges that support it, all without training a model.
Store the graph as flat triples in an edges collection, with the nodes in an
entities collection:
{
"entities": {
"type": "collection",
"columns": {
"id": { "type": "String" },
"kind": { "type": "String" },
"name": { "type": "Text" }
}
},
"edges": {
"type": "collection",
"columns": {
"id": { "type": "Int" },
"subject": { "type": "String", "link": "entities.id" },
"relation": { "type": "String" },
"target": { "type": "String", "link": "entities.id" },
"source": { "type": "String" }
}
}
}
One row per edge (subject βrelationβ target). source records where the edge
came from (for provenance). This "store flat, present grouped" model is the whole
design β no adjacency lists, no nested sets, no graph type. A relation is just a
string, so the vocabulary is open: add new relation kinds without a schema change.
Traversing forward
A forward link goes from a row to the row it points at. From edges you read
the endpoints and their attributes with a dotted path:
{
"from": "edges",
"where": { "subject": "acme", "relation": "uses" },
"select": ["target", "target.kind"]
}
target.kind resolves each edge's target entity and reads its kind β a join,
expressed as a path.
Inverse links
The new piece is going the other way β from an entity to the edges that point
at it. $refs.edges.subject on an entity is the set of edges whose subject
is that entity: its neighborhood. (See the Query Reference
for the full $refs surface.)
Does this entity have any edges?
{
"from": "entities",
"where": { "$refs.edges.subject": { "$exists": true } }
}
Which entities have an engaged_with edge? β filter the referring edges:
{
"from": "entities",
"where": { "$refs.edges.subject": { "$exists": { "relation": "engaged_with" } } }
}
Project the neighborhood β each entity's relations and targets as an array:
{
"from": "entities",
"where": { "kind": "prospect" },
"select": [
"name",
{ "relations": "$refs.edges.subject.relation" },
{ "targets": "$refs.edges.subject.target" }
]
}
// β { "name": "Alice", "relations": ["engaged_with", "works_at"], "targets": ["analytics", "acme"] }
Nothing is copied onto the entity row β the neighborhood is derived from the
back-links. A query reads the same whether you had stored the edges on the row or
discovered them via $refs.
Multi-hop paths
Forward links chain to any depth. A dotted path follows as many links as your
schema has, each hop resolving the row the previous one points at. From an impression
to its context, to that context's user, to the user's tags is context.user.tags:
two links and an attribute, in one path. The same path works in where, in select,
and as evidence for a prediction; on collections there is no depth cap on a forward
chain. On tables (type: "table") paths go one hop; basedOn names only the
predicted link target's own fields (see Link depth).
Self-links chain like any other link. A column declared "link": "this.id" points at
another row of the same table, such as a visit's previous visit (prev). A path can go
through it and on: the visits whose previous visit was by a club member are
{
"from": "visits",
"where": { "prev.user.tags": { "$has": "club-member" } },
"limit": 5
}
Available since: v2.11.2. On v2.11.1 and earlier a path through a
this.<key> self-link is refused with a 400 naming the link; filter on the
self-link's own value instead.
$refs composes with forward paths. Filter referring edges by a linked
attribute of their target (one hop past the edge):
{
"from": "entities",
"where": { "$refs.edges.subject": { "$exists": { "target.kind": "product" } } }
}
Or resolve a forward link first and then its inverse links (<link>.$refs.β¦).
On an emails collection whose sender links to an entity:
{
"from": "emails",
"where": { "sender.$refs.edges.subject": { "$exists": true } }
}
That last one returns emails whose sender has at least one edge.
What is not there yet: paths of unknown length. Every path above has a length you
write down. Reachability (can B be reached from A at any depth), shortest paths, and
"everything below this node" in a hierarchy need a traversal whose depth is not known
in advance, and there is no native recursive path query yet. Today the caller composes
them across queries: an agent loops, one query per step, until the frontier is empty.
For graph algorithms such as PageRank or community detection, export to a graph
database. Two narrower limits on $refs filters are listed under Limits.
Typed edges
To ask for one edge that matches several attributes at once β e.g. an
is-a β CTO edge, not "any is-a edge and (separately) any CTO edge" β give $exists
an object of conditions. They all hold on the same referring edge:
{
"from": "entities",
"where": { "$refs.edges.subject": { "$exists": { "relation": "is-a", "target": "CTO" } } }
}
This is the query typed graphs need: it matches an entity only if it has a single
edge that is both is-a and points at CTO. $has works in place of $exists, and
this is the same shape a stored member-link set uses ({"items": {"$has": {β¦}}}) β so
"same-member" reads identically whether the edges are discovered via $refs or stored
as a set.
Predicting over the graph
Two prediction modes fall out of the same storage.
Node classification β predict an entity attribute from its neighborhood. Put a
$refs condition in the where as evidence, then predict:
{
"from": "entities",
"where": { "$refs.edges.subject": { "$exists": { "relation": "engaged_with" } } },
"predict": "segment"
}
The reverse-linked edges act as evidence with no data copied onto the entity β
Aito weighs "has an engagement" toward the segment, with a $why explanation.
Edge prediction β predict the missing endpoint of a relation. Because an edge
is a row with linked endpoints, predicting target given subject + relation is
link prediction, and recommend over the target link ranks candidate endpoints:
{
"from": "edges",
"where": { "subject.role": "tech", "relation": "uses" },
"predict": "target"
}
Condition on the subject's attributes (here subject.role), not just its id:
that is what lets prediction generalize to a new entity that has no edges yet.
Edge-level prediction is the newer of the two modes.
Generalize over the candidates with basedOn. By default the ranked target
candidates are distinguished partly by their identity, so a target entity seen in only
a few edges gets weak signal and one never linked before gets none. basedOn names the
candidate target's own attributes, so ranking generalizes over those attributes
instead of the raw id:
{
"from": "edges",
"where": { "subject.role": "tech", "relation": "conversion" },
"predict": "target",
"basedOn": ["kind", "industry"]
}
The fields in basedOn are attributes of the target entity table (the thing
being predicted) β read it as rank candidate targets by their kind and industry, so
a target with no prior edges is still predictable from what kind of entity it is. The
two generalizations pull in opposite directions and live in different places: condition
on the known side (the subject) through the where (subject.role above);
generalize over the predicted side (the target candidates) through basedOn.
basedOn applies only when the prediction target is a link.
Point at a related entity β $examine. Conditioning on subject.role means you
already know the role. When you only have a reference to the entity, $examine looks
its attributes up for you: name the link, the entity's id (at), and which attributes to
condition on. Aito reads that row and expands the attributes into evidence:
{
"from": "deals",
"where": { "account": { "$examine": { "at": "newco", "basedOn": ["industry", "size"] } } },
"predict": "outcome"
}
This conditions the prediction on newco's industry and size β so it works even if
newco appears in no deals yet: the model generalizes through the entity's
attributes, not its unseen id. $examine resolves to the same account.industry: β¦
condition you could write by hand; it just spares you the lookup. (Requires basedOn
today; at may be a bare id or { "id": β¦ }.)
Getting text into a graph
Aito reasons over a graph; it does not extract one from raw text. Turning unstructured text into structured relationships is a separate, upstream step β and a near-certain one: the facts are stated in the text, so extraction should recover them with high confidence, not guess them. Use the right tool for it:
- an LLM β prompt it to emit
(subject, relation, object)triples from each document; - an NLP pipeline β named-entity recognition + relation extraction (e.g. spaCy, or a hosted NLP service);
- a specialized / neural relation-extraction model where accuracy matters most.
Write the resulting triples into an edges collection, and Aito takes over: it predicts
the missing edges, classifies a node from its neighborhood, and grounds every answer
with $why β incrementally, and without training a model.
The division of labor is the point. Extraction turns text into facts (high-confidence,
done upstream). Aito turns facts into predictions and explanations β probabilistic,
calibrated, and over the graph you've built. Don't use Aito's predict as an extractor:
guessing a stated fact from loose token correlations is a low-confidence answer to a task
that should be near-certain.
Provenance and explanations
source on each edge is provenance: filter by it, and count corroboration as the
number of sources for a (subject, relation, target) β no stored confidence, the
support is the signal. Retraction is deleting the sourcing edge. Combined with
$why on a prediction (which names the contributing evidence), an agent can ground
an answer in which edges, from what source, drove it.
Count the sources, not the rows. $distinctLength is $length's
weighted-corroboration sibling: it counts the distinct members of a projection, so
one talkative crawler filing the same claim ten times reads as one corroboration
rather than ten.
Available since: v2.10.2. On v2.10.1 and earlier, $distinctLength is not a
known operator; $length (the raw count) has been available since v2.7.0.
{
"from": "claims",
"select": [
"id", "subject", "relation", "target",
{ "evidenceRows": { "$length": "$refs.evidence.claim.source" } },
{ "sources": { "$distinctLength": "$refs.evidence.claim.source" } }
],
"orderBy": { "$desc": "sources" }
}
Both read the same $refs projection β which preserves duplicates deliberately, so
that the two counts can differ. On a stored set, whose members are unique by
construction, they agree. In SQL the pair is cardinality(col) (Postgres's own) and
aito.distinct_cardinality(col) (Aito's, since Postgres has no per-row equivalent).
Worked example
entities (companies, prospects, products, projects), edges (engaged_with,
uses, sponsors, works_at), and emails (linked to a sender) is enough to:
- classify β is a prospect engaged?
where {$refs.edges.subject: {$exists: {relation: engaged_with}}}; - prioritize β predict an email's priority from whether its sender is engaged
(
sender.$refs.edgesβ¦as evidence); - associate β from an entity, project its edges' targets to find related companies, products, and projects.
This is exactly the shape an agent maintaining a CRM-style knowledge graph needs: retrieve a neighborhood, predict the next best action, and cite the supporting edges.
Stored member-link sets (baskets)
The other way to relate rows is to store the links as a set on the row β a
Set/Array of foreign keys β instead of discovering them via $refs. A basket of
products, a person's tags, a document's authors:
{
"baskets": {
"type": "collection",
"columns": {
"id": { "type": "Int" },
"items": { "type": "Int[]", "link": "products.id" },
"label": { "type": "String" }
}
}
}
Each member is a real products row, so you query the set's linked attributes with a
dotted path β and the query reads the same as the $refs surface, because a
member-link set and a discovered $refs set are the same thing.
Match a member by its attributes β products.name is projected across the members
and matched by token, so a shared word finds different brands:
{
"from": "baskets",
"where": { "items.name": { "$has": "banana" } }
}
This matches a basket whose products include "pirkka banana" or "chiquita banana" β
the token banana is shared across brands (query a token, not the whole phrase).
select ["items.name"] still returns the whole names.
Same-member conjunction β for one member that matches several attributes at once
(not "some member is A and some other is B"), give $has an object:
{
"from": "people",
"where": { "items": { "$has": { "a": "fire", "b": "fighter" } } }
}
Matches a person who has a single item that is both fire and fighter β the same
operator as the $refs typed-edge query above.
Predict a member β predict "items.$feature" scores each candidate member
independently (non-exclusive): "given this basket, which product is most likely also
present" β the link-prediction shape for stored sets.
Limits as of v2.11.0
Known gaps in the engine as of v2.11.0, each with its workaround. Each line goes when its fix ships.
-
Operators inside a same-member object. In
{"$exists": {β¦}}/{"$has": {β¦}}only equality works ({"relation": "is-a"}); an operator such as{"amount": {"$gt": 1000}}is refused. Operators there are planned. -
Sorting by a
$refspath. AnorderBynaming a raw$refs.β¦path is ignored (rows come back unsorted). To sort by a count, select the value under an alias and sort by the alias:"select": [{"n": {"$length": "$refs.edges.subject.relation"}}], "orderBy": {"$desc": "n"}.Available since: v2.11.4 to sort by the count directly β
"orderBy": {"$desc": {"$length": "$refs.edges.subject"}}. On v2.11.3 and earlier$lengthis refused insideorderBy; use the alias above. -
$lengthon the bare link path, up to v2.11.3. On v2.11.3 and earlier$lengthneeds an attribute of the referring rows ($refs.edges.subject.relation); the bare link path ($refs.edges.subject) is refused with "field not found". A newer build counts the referring rows themselves, so the bare path is the degree. -
Self-links declared as
this.<key>, up to v2.11.0. On v2.11.0 and earlier a link declared"link": "this.id"is refused with a400in awhere, aselectand a$refsfilter; a newer build traverses it (see Multi-hop paths). A self-link that names its own table ("link": "companies.company_id") traverses like any other link on every version. -
Nested
$refsin a filter. A$refsinside another$refscondition inwhereis refused; the same nested path works as aselectprojection. Split the filter across two queries. -
One forward hop inside a same-member object. In
{"$exists": {β¦}}a condition may follow one forward link past the referring row (target.kind); a longer chain there (context.visit.user) is refused. Outside$exists, forward chains have no depth limit. -
$examinedoes not accept a nullable link. On a link declared"nullable": true,$examineis refused with "is not a link". A required link works whatever its target key is named (employees.Nameas well asid).
What's not yet ergonomic
$refs makes graphs usable today; a few advanced conveniences are still to come
(tracked in the predictive-graphs roadmap):
- Predict about existing rows in one query β "select these edges, predict each one's
target." The near-term ergonomic is a projected prediction column,select: ["target.$predictions"], where the inference context for each row is the row itself (no separatewhere-for-prediction). Today you do this per row by conditioning on the known side in thewhere(subject.role: β¦, see Generalize over the candidates above). Tracked as priority 1 in the predictive-graphs roadmap. basedOna neighborhood β predicting an attribute based on the whole set of reverse-linked edges ($refs.edges.subject) as one named input, rather than awherecondition, is not yet wired; thewhere-evidence form under Predicting over the graph covers the capability meanwhile.- Paths of unknown length β fixed-length paths are built (forward chains of any
depth, plus
$refs; see Multi-hop paths). Reachability (A β*β B), shortest paths and hierarchy walks are composed by the caller across queries (an agent loops). There is no native recursive path query yet. - Graph aggregates β done (weighted corroboration available since
v2.10.2). Degree, neighbour counts and basket size come from
$length, a per-row count of an array/set field (a$refsprojection or a stored member-link set):{ "from": "entities", "select": ["id", { "degree": { "$length": "$refs.edges.subject.relation" } }] }gives each node's out-degree. (It counts the projected referring edges, so name any edge attribute β e.g.relation; since v2.11.4 it also takes the bare$refs.edges.subject.) Weighted corroboration β the number of distinctsources backing an edge β is$distinctLengthover the same projection; see Provenance and explanations above.