Guide

Most recommendation engines retrieve. They do not recommend.

A recommendation engine should pick the product that moves this shopper toward their next purchase and your goal. Most engines do something simpler: they find the item most similar to the last click and show it. That is search with a different name.

  • Clerk.io: recommendations draw 7% of traffic and produce 26% of revenue
  • Similar-product engines optimize the click, which reinforces old habits
  • Next-basket prediction asks what the shopper buys next, and when
In short

Google's 2016 paper on YouTube recommendations described a two-tower model: embed users and items in one space, measure similarity, return the closest items. Nearly every personalization vendor copied the first stage. It retrieves. It does not decide.

Retrieval answers "what is similar to what they did?" and optimizes for the click. The agent answers "what will this person buy next, and which action gets them there?" and optimizes for the goal you choose: a sale now, a bigger basket, or value over the year.

The gap has been measured. At IJCAI 2015, Theocharous, Thomas and Ghavamzadeh showed a policy with a lower click rate producing more revenue and higher lifetime value, and the ACM RecSys 2023 tutorial on lifetime value says why: systems optimize clicks, ratings and dwell time. Next-basket prediction is one of the agent's 16 named algorithms, and it is the one retrieval cannot copy.

What retrieval does

Covington, Adams and Sargin at Google published "Deep Neural Networks for YouTube Recommendations" at ACM RecSys 2016. The system has two stages: candidate generation, which narrows billions of videos to hundreds by similarity, and ranking, which orders those hundreds with richer features. The paper is explicit about the split.

The method, step by step: embed products by what appears together in sessions, carts and orders; embed each shopper by their history; measure the distance between shopper and product; return the closest items. It learns that people who did X tend to do Y. It never learns that this shopper should see Z because Z leads to the next order.

Most ecommerce vendors implement the first stage and call it recommendation. The ranking stage, where a value model could live, is missing or optimizes for engagement. Clerk.io's own figures show how much is at stake: recommendations draw 7% of ecommerce traffic but produce 24% of orders and 26% of revenue.

Three questions, three different engines

What is similar?

Retrieval. Method: similarity in an embedding space. Optimizes the probability of a click. Result: more of the same. Discount buyers get more discounts.

Which order?

Ranking. Method: learning to rank with more features. Optimizes a mix of clicks, time on page and add-to-cart. Result: a better order of the same similar items.

Which action?

The agent. Method: predict each shopper's next basket and value, then pick the product, message, channel and moment that move them toward the goal. Result: the most valuable next step, which may be a product outside the shopper's pattern.

The research community has measured the gap. At IJCAI 2015, Theocharous, Thomas and Ghavamzadeh showed that of two ad recommendation policies, the one with the lower click rate produced more revenue and a higher lifetime value. Ferraro and colleagues' 2023 systematic review in Expert Systems with Applications found that most recommender research still optimizes engagement rather than economic value. The ACM RecSys 2023 tutorial on customer lifetime value calls the shift from clicks to value a change of paradigm.

The loop that shrinks baskets

A shopper buys a discounted item. The engine places them next to other discount buyers. It shows them more discounts, because those have the highest similarity. They click, which confirms the prediction, so the engine moves them deeper into the discount cluster. The mid-range product that would have opened a new category is never shown.

The engine's accuracy goes up. The customer's value goes down. That is the failure mode of optimizing the click: the system gets better at predicting what the shopper will click and worse for the store.

The opposite is visible in a store's own orders. In one online grocery, 30% of orders include a recommended item and 68% of recommendation revenue comes from a single cart-page block that reminds each shopper of their usual items (case study). In a DIY store, orders placed after a recommendation were nearly 4 times the store's average (case study).

Next-basket prediction

The most valuable question in ecommerce is not "what will this shopper click next?" It is "what will their next basket contain, and when will they buy it?" Next-basket prediction models the purchase sequence over time: buying cycles, category expansion, seasonality, replenishment.

That one prediction changes everything downstream. The browse-abandonment email goes out three days before the predicted purchase window, not 24 hours after the visit. The products in it fit the next basket, not the last click. The channel is the one this shopper answers at this point in their cycle. A shopper with a large predicted basket sees a different offer from one with a small one.

Retrieval has no concept of time, no concept of value and no concept of what comes next. It has a similarity score and a catalog. The agent has one memory of each shopper and a goal you set; the how it decides page shows what it does with them.

Five tests for your engine

Does it know what a customer is worth? Can it recommend something the shopper has never browsed? Do the email picks differ from the on-site picks? Does a one-time buyer get the same logic as your best customer? Can you measure whether a recommendation changed what a customer bought later, not just whether they clicked?

Three or more answers of no means the engine retrieves. In a 50/50 test at a kitchen appliance store, the agent produced 2.3 times as many add-to-carts per visitor, significant at 95%. In a motorbike marketplace, viewers who saw a recommendation booked test rides at 6.5 times the rate (case study).

Questions about recommendation engines

What is the difference between retrieval and recommendation?

Retrieval finds the items most similar to a shopper's history by measuring distance in an embedding space. Recommendation worth the name picks the product that moves this shopper toward their next purchase and your goal. Most engines sold as recommendation do retrieval.

Why do similar-product engines shrink baskets?

They optimize the click. Shoppers click what resembles what they already bought, so discount buyers see more discounts and the mid-range product that would open a new category never appears. The engine gets more accurate while the customer gets less valuable.

What is next-basket prediction?

A model of each shopper's purchase sequence over time that predicts what their next basket will contain and when they will buy it. It is one of the agent's 16 named algorithms and the one a similarity engine cannot copy.

Which paper started the two-tower recommendation model?

"Deep Neural Networks for YouTube Recommendations" by Covington, Adams and Sargin at Google, published at ACM RecSys 2016. It describes two stages, candidate generation and ranking; most vendors implemented only the first.

See which of last week's visitors you missed

Thirty minutes, with your store and ad accounts open. Then 30 days free. A holdout group decides: if the agent doesn't add orders in 30 days, you don't pay.

Book 30 minutes
GeorgiYavorBoryanaBoyanMariaIsaacNikoletaYou'll talk to one of us.