Personalised Search Autocomplete
A search box that completes the same prefix the same way for everyone is a lookup table. A useful one ranks completions by what this user is likely to type, which means learning from the searches people have actually made. In Aito v2 that is a prediction over the query log: the typed prefix limits the candidates, and everything else you know about the session is evidence.
The queries below run against the live v2 sandbox. Its
contexts table is a grocery shop's session log: one row per page view, with
the queryPhrase the shopper had typed and a user link to users.
Probabilities in the responses are rounded.
Popular completions for a prefix
Without personalisation, a completion is the share of past searches that
start with the prefix. Open queryPhrase as candidates with get and filter
by the prefix:
{
"from": "contexts",
"where": { "queryPhrase": { "$startsWith": "m" } },
"get": "queryPhrase",
"orderBy": "$p",
"select": ["$value", "$p"],
"limit": 5
}
{ "total": 8, "hits": [
{ "$value": "milk", "$p": 0.218 },
{ "$value": "meat", "$p": 0.206 },
{ "$value": "macaroni", "$p": 0.200 },
{ "$value": "myllyn paras", "$p": 0.182 },
{ "$value": "maxi", "$p": 0.082 } ] }
Eight distinct searches in the log start with "m", and the ranking is their frequency.
Personalised completions
Add the user. With predict, a condition on the predicted field itself
restricts the candidates (only phrases starting with "m" can be answered),
while the user condition is evidence that re-weights them:
{
"from": "contexts",
"where": { "user": "larry", "queryPhrase": { "$startsWith": "m" } },
"predict": "queryPhrase",
"select": ["$value", "$p"],
"limit": 5
}
{ "total": 8, "hits": [
{ "$value": "meat", "$p": 0.412 },
{ "$value": "milk", "$p": 0.153 },
{ "$value": "macaroni", "$p": 0.133 },
{ "$value": "myllyn paras", "$p": 0.132 },
{ "$value": "maxi", "$p": 0.066 } ] }
Larry has searched for "meat" twice, so it moves to the top. For "veronica"
the same query puts "macaroni" first (0.335), ahead of "milk" and "meat". The
other candidates keep a share of the probability, so a user with little
history still gets sensible completions rather than an empty list.
The same query in SQL, where LIKE 'm%' is the prefix condition:
SELECT * FROM predictions('contexts', 'queryPhrase',
given => 'user = ''larry'' AND queryPhrase LIKE ''m%''', k => 5)
Why this completion
Add $why to see what moved the ranking:
{
"from": "contexts",
"where": { "user": "larry", "queryPhrase": { "$startsWith": "m" } },
"predict": "queryPhrase",
"select": ["$value", "$p", "$why"],
"limit": 1
}
For "meat" the factor tree has three parts: the base rate of "meat" over the
whole log, a normaliser that rescales it against the other candidates, and a lift of
1.86 from user: "larry". That last factor is the personalisation, stated as
a number you can show or log.
Suggestions before the first keystroke
An empty search box can already suggest something. Exclude the empty phrase (page views without a search) and let the user be the only evidence:
{
"from": "contexts",
"where": { "user": "larry", "queryPhrase": { "$not": "" } },
"predict": "queryPhrase",
"select": ["$value", "$p"],
"limit": 5
}
For Larry the top suggestions are "pirkka", "bread", "banana", "vegetable" and "paper". These probabilities are small (the largest is 0.076), because 117 different phrases compete; show them as suggestions, not as a confident guess.
Why Aito for this
- The log is the model. Every search written to
contextscounts in the next request. There is no suggestion index to rebuild. - Any evidence, same query. The user is one condition. The weekday, the
current basket or the page type are more conditions in the same
where. - Explainable ranking.
$whyseparates popularity from personalisation.
Related: Smart search Β· Query Reference Β· Inference Β· Playground