Knowledge Graphs β Node Classification and Link Prediction
A knowledge graph is usually two products: a graph database to store the edges,
and a model to predict the missing ones. Aito v2 is one. Edges are ordinary rows
in a collection, $refs walks them backwards, and the same inference that
predicts a column predicts an edge β so classification, link prediction and the
explanation all come from the storage you already loaded.
No training step, no embedding refresh, and no second system to keep in sync: write a triple and it is queryable in the next request.
The shape
Two collections are enough. Nodes carry attributes; edges link node to node and record where the claim came from.
PUT /api/v2/schema/entities
{ "type": "collection",
"columns": {
"id": { "type": "String" },
"kind": { "type": "String" },
"name": { "type": "Text", "analyzer": "english" } } }
PUT /api/v2/schema/edges
{ "type": "collection",
"columns": {
"subject": { "type": "String", "link": "entities.id" },
"relation": { "type": "String" },
"target": { "type": "String", "link": "entities.id" },
"source": { "type": "String" } } }
subject and target are links, so a query can walk an edge forwards
(target.kind) without a join. $refs.edges.subject walks it backwards: from a
node to the edges that point out of it.
Node classification β label a node from its neighborhood
What a node is connected to says what it is. Classify an entity by conditioning on its edges and predicting the attribute:
{
"from": "entities",
"where": { "$refs.edges.subject": { "$exists": { "relation": "sponsors" } } },
"predict": "kind",
"select": ["$value", "$p", "$why"],
"limit": 5
}
The $exists form is a same-member condition: it matches a node that has
one edge satisfying every clause together, not a node with a sponsors edge
somewhere and an unrelated second edge. That distinction is what makes typed-edge
questions answerable rather than approximately answerable.
$why names the edges that drove the label, so a reviewer can see the evidence
rather than a score.
Link prediction β the missing edge
Predicting an edge is predicting its target column. Because target is a link,
the candidates are entities and each one is ranked by calibrated probability:
{
"from": "edges",
"where": { "subject.kind": "company", "relation": "uses" },
"predict": "target",
"select": ["$value", "$p", "$why"],
"limit": 5
}
This generalises from the subject's attributes β "companies like this one use analytics" β which is what makes it answer for a node with no edges yet, the case a graph-embedding model has to be retrained to handle.
Condition on the subject's properties rather than its id when the entity is new; an id the graph has never seen carries no signal on its own.
Degree and corroboration β the graph's own statistics
$length counts an array/set field per row, and a $refs projection is such a
field, so out-degree is a projection away:
{
"from": "entities",
"select": ["id", { "degree": { "$length": "$refs.edges.subject.relation" } }],
"orderBy": { "$desc": "degree" },
"limit": 10
}
Available since: v2.10.2 for $distinctLength (below) and for
$patterns over a link path (see Cross-table pattern mining). On v2.10.1 and earlier the
first returns an unknown-operator error and the second rejects the linked field.
Everything else on this page works on v2.10.0.
Provenance needs a different count. Ten edges filed by one crawler are one corroboration, not ten, so weight by distinct sources:
{
"from": "claims",
"select": [
"id",
{ "evidenceRows": { "$length": "$refs.evidence.claim.source" } },
{ "sources": { "$distinctLength": "$refs.evidence.claim.source" } }
],
"orderBy": { "$desc": "sources" },
"limit": 10
}
There is no stored confidence score to maintain: the support is the signal, and retraction is deleting the sourcing edge.
Cross-table pattern mining
relate with $patterns mines co-occurring conjunctions, and it reads link paths
β so the mined rule can span the join, which is where the surprising ones live:
Available since: v2.10.2. On v2.10.1 and earlier a linked field in the
$patterns list is rejected with no such field 'customer.segment'.
{
"from": "notes",
"where": { "customer.churn": true },
"relate": { "$patterns": ["body", "customer.segment"] },
"select": ["related", "condition", "lift", "n"],
"limit": 5
}
A hit like "body mentions thermal and customer.segment = construction β
churn, lift 2.4" is a rule no single-table mine can state.
Why this is not a graph database
| Graph DB + model | Aito v2 | |
|---|---|---|
| Storage | Native graph store | Ordinary collections with links |
| Missing edges | Train an embedding model, refresh it | predict the target column |
| New node | Retrain or cold-start | Generalises from its attributes immediately |
| Explanation | Similarity score | $why factor tree over the actual edges |
| Write path | Load, then re-embed | Row is queryable in the next request |
| Confidence | Stored weight to maintain | Corroboration counted from sources |
The trade is real and worth stating: Aito has no native recursive path query, so
N-hop reachability (A β*β B) is composed by the caller across queries. Forward
paths chain to any depth, with one reverse $refs hop on top. For traversal-heavy workloads β shortest path, community detection β
a graph database is the right tool. For predicting and explaining edges over
data you already have, this is one system instead of two.
Next
- Graphs guide β the full
$refssurface, provenance and worked examples - Query reference β
$exists,$length,$distinctLength - Inference β how
predictand$whywork