Smart Search: Relevance Plus Personalisation

Text relevance tells you which products match the words. It cannot tell you which of the matching products this shopper will buy. A shop that sells both regular and lactose-free milk returns both for "milk", and the order should depend on who is asking. Aito v2 answers both halves over the same data: BM25 ranks the matches, and a recommendation restricted to those matches ranks them by the probability of a purchase given the user.

The queries below run against the live v2 sandbox: impressions records which product was shown in which context and whether it was bought (purchase), and context.user reaches the shopper through two links. Scores and probabilities in the responses are rounded.

Plain relevance: the same for everyone

The search query type ranks by BM25:

{
  "from": "products",
  "search": { "field": "name", "text": "milk" },
  "select": ["name", "$score"],
  "limit": 4
}
{ "total": 6, "hits": [
  { "name": "Valio semi-skimmed milk 1l", "$score": 1.966 },
  { "name": "Pirkka Finnish nonfat milk 1l", "$score": 1.966 },
  { "name": "Pirkka Finnish semi-skimmed milk 1l", "$score": 1.825 },
  { "name": "Fazer Sininen milk chocolate slab 200g", "$score": 1.825 } ] }

Six product names contain "milk". BM25 favours shorter names, so the chocolate slab scores as well as a carton of milk, and nothing in the ranking depends on the shopper.

Personalised: rank the matches by the chance of a purchase

Move the text condition into a recommendation. A condition on a field of the recommended link (product.name) restricts the candidates to the matching products, and context.user is evidence:

{
  "from": "impressions",
  "where": {
    "context.user": "larry",
    "product.name": { "$match": "milk" }
  },
  "recommend": "product",
  "goal": { "purchase": true },
  "select": ["name", "$p"],
  "limit": 5
}
{ "total": 6, "hits": [
  { "name": "Pirkka lactose-free semi-skimmed milk drink 1l", "$p": 0.092 },
  { "name": "Valio eilaβ„’ Lactose-free semi-skimmed milk drink 1l", "$p": 0.061 },
  { "name": "Fazer Sininen milk chocolate slab 200g", "$p": 0.031 },
  { "name": "Valio semi-skimmed milk 1l", "$p": 0.025 },
  { "name": "Pirkka Finnish semi-skimmed milk 1l", "$p": 0.025 } ] }

For Larry, the two lactose-free drinks come first, and each is at least twice as likely to be bought as any regular milk. For "veronica" the six products are close together (between 0.018 and 0.026): her history does not separate them, and the ranking says so instead of inventing a preference.

$p is the probability that the product is bought when shown to this user, so it can be compared across searches and thresholded.

The same recommendation in SQL, which returns the product ids with p:

SELECT * FROM recommend('impressions', 'product', goal => 'purchase = true',
  given => 'context.user = ''larry'' AND product.name @@ ''milk''', k => 5)

Blend relevance and purchase probability

When relevance should still count, multiply the two. $similarity is the BM25 relevance of the $match condition, expressed as a lift, and { "$p": { "$context": … } } is the purchase probability per candidate:

{
  "from": "impressions",
  "where": {
    "context.user": "larry",
    "product.name": { "$match": "milk" }
  },
  "get": "product",
  "orderBy": {
    "$multiply": ["$similarity", { "$p": { "$context": { "purchase": true } } }]
  },
  "select": ["name", "$score"],
  "limit": 5
}

The top two for Larry are again the lactose-free drinks. In this form the $match does not remove the other products from the candidates: a product that does not match scores 0 and sorts after every match, so page no further than the number of matches.

Show why a result matched

A search result can mark the matched words for the UI. The parametric $highlight takes your own tags:

{
  "from": "products",
  "search": { "field": "name", "text": "lactose-free milk" },
  "select": ["name", { "$highlight": { "posPreTag": "<mark>", "posPostTag": "</mark>" } }],
  "limit": 3
}
{ "total": 2, "hits": [
  { "name": "Pirkka lactose-free semi-skimmed milk drink 1l",
    "$highlight": [ { "field": "name",
      "highlight": "Pirkka <mark>lactose</mark>-<mark>free</mark> semi-skimmed <mark>milk</mark> drink 1l" } ] },
  { "name": "Valio eilaβ„’ Lactose-free semi-skimmed milk drink 1l",
    "$highlight": [ { "field": "name",
      "highlight": "Valio eila&trade; <mark>Lactose</mark>-<mark>free</mark> semi-skimmed <mark>milk</mark> drink 1l" } ] } ] }

$matches returns the same information as data: each matched token, its character positions, and its contribution to the score. In SQL both are columns of aito.search:

SELECT name, highlight FROM aito.search('products', 'name', 'lactose-free milk',
  startSel => '<mark>', stopSel => '</mark>', k => 3)

Why Aito for this

  • One store for search and behaviour. The text index and the purchase history are the same database, so personalisation needs no second system and no sync job.
  • Candidates you control. The text condition decides what can be returned; the user decides the order.
  • Honest when there is no signal. A user whose history does not separate the matches gets probabilities that are close together, not a confident guess.

Related: Full-text search Β· Autocomplete Β· Recommendations Β· Query Reference

← All v2 use cases