Predictive Cart Autofill
Many shoppers buy much the same things every week. Autofill proposes that basket in one click: predict which products this shopper is likely to buy on this visit, keep the ones above a threshold, and let the shopper remove what they do not want. In Aito v2 this is a multi-label prediction over past visits, with each product scored independently.
The queries below run against the live v2 sandbox. Each row of
visits is one shopping trip, with its user, weekday and purchases, an
array of product ids. Probabilities in the responses are rounded.
Predict the basket
purchases is an array, so predict its members with .$feature. Each product
gets its own probability of being in the basket, and the probabilities do not
sum to one:
{
"from": "visits",
"where": { "user": "larry" },
"predict": "purchases.$feature",
"select": ["$value", "$p"],
"limit": 5
}
{ "total": 42, "hits": [
{ "$value": "6410405093677", "$p": 0.417 },
{ "$value": "2000818700008", "$p": 0.392 },
{ "$value": "6411401028373", "$p": 0.370 },
{ "$value": "6407870071224", "$p": 0.360 },
{ "$value": "6437002001454", "$p": 0.348 } ] }
For Larry the top five are iceberg salad, Pirkka bananas, Tyrkisk Peber
chocolate, ham sausage and VAASAN rye bread. purchases holds product ids
rather than links, so the names come from products.
Add the context of the visit
The weekday is more evidence in the same where:
{
"from": "visits",
"where": { "user": "larry", "weekday": "Saturday" },
"predict": "purchases.$feature",
"select": ["$value", "$p"],
"limit": 5
}
{ "total": 42, "hits": [
{ "$value": "6410405093677", "$p": 0.443 },
{ "$value": "6411401028373", "$p": 0.431 },
{ "$value": "2000818700008", "$p": 0.406 },
{ "$value": "6437002001454", "$p": 0.401 },
{ "$value": "6410405216120", "$p": 0.391 } ] }
Larry shops mostly on Saturdays and Sundays, and on a Saturday the lactose-free milk enters the top five. With a threshold of 0.4, the Saturday prediction fills four products, where the prediction without a weekday fills one.
Complete a basket that has been started
Once the shopper has added something, condition on it:
{
"from": "visits",
"where": { "purchases": { "$has": "6411300000494" } },
"predict": "purchases.$feature",
"select": ["$value", "$p"],
"limit": 5
}
{ "total": 42, "hits": [
{ "$value": "6411300000494", "$p": 1.000 },
{ "$value": "2000818700008", "$p": 0.510 },
{ "$value": "6414880021620", "$p": 0.376 },
{ "$value": "2000503600002", "$p": 0.361 },
{ "$value": "6410405093677", "$p": 0.319 } ] }
The product in the condition (Juhla Mokka coffee) comes back first with a probability of 1.0, because every visit that matches the condition contains it. Drop the items that are already in the basket on the client side. After the coffee come Pirkka bananas, the weekend newspaper, Chiquita bananas and iceberg salad.
Measure before you autofill
A wrong autofill costs the shopper a click per unwanted item, so measure the prediction on held-out visits before choosing a threshold:
{
"test": { "$index": { "$mod": [5, 0] } },
"evaluate": {
"from": "visits",
"where": { "user": { "$get": "user" } },
"predict": "purchases.$feature"
},
"select": ["n", "accuracy", "baseAccuracy", "meanRank", "logLoss", "baseLogLoss", "logLossSkill"]
}
{ "kind": "evaluation", "data": {
"n": 147, "accuracy": 0.449, "baseAccuracy": 0.442, "meanRank": 1.993,
"logLoss": 1.131, "baseLogLoss": 1.053, "logLossSkill": -0.074 } }
On this demo dataset, the user alone hardly improves on popularity: the top
product is in the basket on 44.9% of the 147 held-out visits, against 44.2% for
always proposing the most common products, and logLossSkill is slightly
negative. That is a finding, not a failure of the query: on this data, autofill
from the user's identity is no better than a popularity list. Test other
evidence the same way, such as the weekday or the month, and ship the version
that beats the baseline on your data.
Why Aito for this
- Multi-label by name.
purchases.$featurescores every product on its own, which is what a basket needs. - Evidence is additive. The user, the weekday and the started basket are
conditions in one
where; add or drop them without changing the query shape. - Measured in the same API.
_evaluateruns the same query on held-out rows and reports it against the baseline.
Related: Recommendations Β· Multi-label tagging Β· Evaluation (v2)