Predictive Knowledge Graph

Prediction Over Your Graph

A graph database answers what is connected. Aito answers what is probably true about the connections you have not seen yet: the missing edge, the label on an unlabelled node, the counterparty that does not fit. Every answer is a calibrated probability with the edges that drove it, and it reflects the next write with no retraining.


What it does

  • Link prediction, including for nodes with no edges yet. Predict the missing endpoint of a relation. Condition on the node's attributes and a brand-new node gets an answer from what kind of node it is, not from a history it does not have.
  • Node classification, with the reason. Label a node from its neighbourhood. $why names the edge conditions that moved the answer, so a person can check it and an agent can cite it.
  • Suspicious edges. Predict who the counterparty should be. When the actual counterparty gets a low probability, the edge is flagged, with the fields that expected someone else.
  • Grounding and corroboration for GraphRAG. Fetch an entity, its neighbourhood and the evidence behind it in one query. $why says which edges drove an answer, and $distinctLength counts how many independent sources back a claim.

Relationships are ordinary links. Aito follows them forward with dotted paths, as deep as your schema goes, reads them backwards with $refs, and combines both in one query, in the same store as your facts, text, and vectors.

A Company AI query card: deals are filtered by whether their account has a contact with the role CTO, reached backwards through $refs, and won is predicted. The answer is 36% probability of winning, and $why lists the condition company_id.$refs.contacts.company_id has role=CTO with a lift of times 1.38.

Evidence read through a link, in Company AI, an open reference application, on its synthetic demo data. Whether a CTO is on file exists nowhere on the deal: the query goes forward to the account and back to its people, and $why names that condition and its lift.


Link prediction

Model your data as entities and edges, one row per subject, relation, target. Predicting target given subject and relation is link prediction:

POST /api/v2/_predict
{
  "from": "edges",
  "where": { "subject.role": "tech", "relation": "uses" },
  "predict": "target",
  "basedOn": ["kind", "industry"]
}

Conditioning on subject.role rather than a subject id is what lets a new node with no edges get a prediction. basedOn does the same on the other side: it ranks candidate targets by their own attributes, so a target that has never been linked is still a candidate.


Node classification

Classify an entity from its neighbourhood, here a segment for anyone who has engaged with something:

POST /api/v2/_predict
{
  "from": "entities",
  "where": { "$refs.edges.subject": { "$exists": { "relation": "engaged_with" } } },
  "predict": "segment",
  "select": ["$value", "$p", "$why"]
}

The reverse-linked edges act as evidence with no data copied onto the entity, and $why names the edge condition behind the answer.

A Company AI query card: companies are filtered by whether a contact with the role CFO links to them, read backwards through $refs, and industry is predicted. The answer ranks accounting at 45%, then erp at 19%, ecommerce at 16%, analytics at 8% and consultancy at 6%.

The same in Company AI: what kind of company this is, judged only by who works there. Synthetic demo data.


A counterparty that does not fit

Fraud screening on a graph usually means searching for suspicious patterns. Aito asks a more direct question: given everything else about this edge, who would we expect on the other end?

POST /api/v2/_predict
{
  "from": "payments",
  "where": { "payer.segment": "retail", "category": "consulting" },
  "predict": "payee",
  "select": ["$value", "$p", "$why"]
}

If the payee actually on the payment sits far down that list with a low $p, the payment is flagged, and $why shows which fields pointed to someone else. It is a probability to threshold, not a rule to maintain.


Grounding for GraphRAG and agents

An agent answering from a knowledge graph needs two things beyond retrieval: which edges support the answer, and how many independent sources stand behind them. $why gives the first. $distinctLength gives the second, counting distinct sources rather than rows, so one source filing the same claim ten times reads as one corroboration, not ten:

POST /api/v2/_query
{
  "from": "claims",
  "select": [
    "id", "subject", "relation", "target",
    { "sources": { "$distinctLength": "$refs.evidence.claim.source" } }
  ]
}

Writing a fact is adding an edge, and retracting it is deleting one. The next query reflects it.


What it is not

Aito is a predictive database that reads links both ways, not a graph traversal engine. Any path you can name, you can query: forward links chain to any depth, and $refs reads a link backwards. What it does not do is search for paths whose length you do not know in advance, such as reachability, shortest paths, or everything below a node in a hierarchy. Today that is a bounded agent loop, one query per step. There are no graph algorithms such as PageRank or community detection either. If those are the core of your workload, run them in a graph database, fed from Aito over SQL. Where Aito fits is the other half: predicting the edge a graph database does not have yet, and saying why.


Get Started

Start for free → Model your entities and edges and predict over them.

Read the graph docs → The full operator surface: links, $refs, node classification, link prediction, and provenance.