Predictive Cart Autofill

Many shoppers buy much the same things every week. Autofill proposes that basket in one click: predict which products this shopper is likely to buy on this visit, keep the ones above a threshold, and let the shopper remove what they do not want. In Aito v2 this is a multi-label prediction over past visits, with each product scored independently.

The queries below run against the live v2 sandbox. Each row of visits is one shopping trip, with its user, weekday and purchases, an array of product ids. Probabilities in the responses are rounded.

Predict the basket

purchases is an array, so predict its members with .$feature. Each product gets its own probability of being in the basket, and the probabilities do not sum to one:

{
  "from": "visits",
  "where": { "user": "larry" },
  "predict": "purchases.$feature",
  "select": ["$value", "$p"],
  "limit": 5
}
{ "total": 42, "hits": [
  { "$value": "6410405093677", "$p": 0.417 },
  { "$value": "2000818700008", "$p": 0.392 },
  { "$value": "6411401028373", "$p": 0.370 },
  { "$value": "6407870071224", "$p": 0.360 },
  { "$value": "6437002001454", "$p": 0.348 } ] }

For Larry the top five are iceberg salad, Pirkka bananas, Tyrkisk Peber chocolate, ham sausage and VAASAN rye bread. purchases holds product ids rather than links, so the names come from products.

Add the context of the visit

The weekday is more evidence in the same where:

{
  "from": "visits",
  "where": { "user": "larry", "weekday": "Saturday" },
  "predict": "purchases.$feature",
  "select": ["$value", "$p"],
  "limit": 5
}
{ "total": 42, "hits": [
  { "$value": "6410405093677", "$p": 0.443 },
  { "$value": "6411401028373", "$p": 0.431 },
  { "$value": "2000818700008", "$p": 0.406 },
  { "$value": "6437002001454", "$p": 0.401 },
  { "$value": "6410405216120", "$p": 0.391 } ] }

Larry shops mostly on Saturdays and Sundays, and on a Saturday the lactose-free milk enters the top five. With a threshold of 0.4, the Saturday prediction fills four products, where the prediction without a weekday fills one.

Complete a basket that has been started

Once the shopper has added something, condition on it:

{
  "from": "visits",
  "where": { "purchases": { "$has": "6411300000494" } },
  "predict": "purchases.$feature",
  "select": ["$value", "$p"],
  "limit": 5
}
{ "total": 42, "hits": [
  { "$value": "6411300000494", "$p": 1.000 },
  { "$value": "2000818700008", "$p": 0.510 },
  { "$value": "6414880021620", "$p": 0.376 },
  { "$value": "2000503600002", "$p": 0.361 },
  { "$value": "6410405093677", "$p": 0.319 } ] }

The product in the condition (Juhla Mokka coffee) comes back first with a probability of 1.0, because every visit that matches the condition contains it. Drop the items that are already in the basket on the client side. After the coffee come Pirkka bananas, the weekend newspaper, Chiquita bananas and iceberg salad.

Measure before you autofill

A wrong autofill costs the shopper a click per unwanted item, so measure the prediction on held-out visits before choosing a threshold:

{
  "test": { "$index": { "$mod": [5, 0] } },
  "evaluate": {
    "from": "visits",
    "where": { "user": { "$get": "user" } },
    "predict": "purchases.$feature"
  },
  "select": ["n", "accuracy", "baseAccuracy", "meanRank", "logLoss", "baseLogLoss", "logLossSkill"]
}
{ "kind": "evaluation", "data": {
  "n": 147, "accuracy": 0.449, "baseAccuracy": 0.442, "meanRank": 1.993,
  "logLoss": 1.131, "baseLogLoss": 1.053, "logLossSkill": -0.074 } }

On this demo dataset, the user alone hardly improves on popularity: the top product is in the basket on 44.9% of the 147 held-out visits, against 44.2% for always proposing the most common products, and logLossSkill is slightly negative. That is a finding, not a failure of the query: on this data, autofill from the user's identity is no better than a popularity list. Test other evidence the same way, such as the weekday or the month, and ship the version that beats the baseline on your data.

Why Aito for this

  • Multi-label by name. purchases.$feature scores every product on its own, which is what a basket needs.
  • Evidence is additive. The user, the weekday and the started basket are conditions in one where; add or drop them without changing the query shape.
  • Measured in the same API. _evaluate runs the same query on held-out rows and reports it against the baseline.

Related: Recommendations Β· Multi-label tagging Β· Evaluation (v2)

← All v2 use cases