Personalised Search Autocomplete

A search box that completes the same prefix the same way for everyone is a lookup table. A useful one ranks completions by what this user is likely to type, which means learning from the searches people have actually made. In Aito v2 that is a prediction over the query log: the typed prefix limits the candidates, and everything else you know about the session is evidence.

The queries below run against the live v2 sandbox. Its contexts table is a grocery shop's session log: one row per page view, with the queryPhrase the shopper had typed and a user link to users. Probabilities in the responses are rounded.

Without personalisation, a completion is the share of past searches that start with the prefix. Open queryPhrase as candidates with get and filter by the prefix:

{
  "from": "contexts",
  "where": { "queryPhrase": { "$startsWith": "m" } },
  "get": "queryPhrase",
  "orderBy": "$p",
  "select": ["$value", "$p"],
  "limit": 5
}
{ "total": 8, "hits": [
  { "$value": "milk", "$p": 0.218 },
  { "$value": "meat", "$p": 0.206 },
  { "$value": "macaroni", "$p": 0.200 },
  { "$value": "myllyn paras", "$p": 0.182 },
  { "$value": "maxi", "$p": 0.082 } ] }

Eight distinct searches in the log start with "m", and the ranking is their frequency.

Personalised completions

Add the user. With predict, a condition on the predicted field itself restricts the candidates (only phrases starting with "m" can be answered), while the user condition is evidence that re-weights them:

{
  "from": "contexts",
  "where": { "user": "larry", "queryPhrase": { "$startsWith": "m" } },
  "predict": "queryPhrase",
  "select": ["$value", "$p"],
  "limit": 5
}
{ "total": 8, "hits": [
  { "$value": "meat", "$p": 0.412 },
  { "$value": "milk", "$p": 0.153 },
  { "$value": "macaroni", "$p": 0.133 },
  { "$value": "myllyn paras", "$p": 0.132 },
  { "$value": "maxi", "$p": 0.066 } ] }

Larry has searched for "meat" twice, so it moves to the top. For "veronica" the same query puts "macaroni" first (0.335), ahead of "milk" and "meat". The other candidates keep a share of the probability, so a user with little history still gets sensible completions rather than an empty list.

The same query in SQL, where LIKE 'm%' is the prefix condition:

SELECT * FROM predictions('contexts', 'queryPhrase',
  given => 'user = ''larry'' AND queryPhrase LIKE ''m%''', k => 5)

Why this completion

Add $why to see what moved the ranking:

{
  "from": "contexts",
  "where": { "user": "larry", "queryPhrase": { "$startsWith": "m" } },
  "predict": "queryPhrase",
  "select": ["$value", "$p", "$why"],
  "limit": 1
}

For "meat" the factor tree has three parts: the base rate of "meat" over the whole log, a normaliser that rescales it against the other candidates, and a lift of 1.86 from user: "larry". That last factor is the personalisation, stated as a number you can show or log.

Suggestions before the first keystroke

An empty search box can already suggest something. Exclude the empty phrase (page views without a search) and let the user be the only evidence:

{
  "from": "contexts",
  "where": { "user": "larry", "queryPhrase": { "$not": "" } },
  "predict": "queryPhrase",
  "select": ["$value", "$p"],
  "limit": 5
}

For Larry the top suggestions are "pirkka", "bread", "banana", "vegetable" and "paper". These probabilities are small (the largest is 0.076), because 117 different phrases compete; show them as suggestions, not as a confident guess.

Why Aito for this

  • The log is the model. Every search written to contexts counts in the next request. There is no suggestion index to rebuild.
  • Any evidence, same query. The user is one condition. The weekday, the current basket or the page type are more conditions in the same where.
  • Explainable ranking. $why separates popularity from personalisation.

Related: Smart search Β· Query Reference Β· Inference Β· Playground

← All v2 use cases