Part 4: Recommender Backtesting

Learn how backtesting, shadow mode, and A/B testing reduce risk and validate recommendation systems before full production rollout.

Key Takeaways:

Offline Metrics Are Not Enough
Backtesting Exposes Risk Before Launch
Real-World Testing Builds Trust

The model looked better offline

The ML team trains a new ranker. Offline metrics improve. Precision@k is up. NDCG is up. The validation curves look clean. Everyone is excited.

Then the A/B test launches.

CTR is flat. Conversion drops for new users. Latency increases. A few high-margin categories disappear from the top slots. Customer support receives complaints that recommendations feel repetitive. Product leaders ask the obvious question:

“Why didn’t we catch this before launch?”

Sometimes the answer is that the offline metric was too narrow. Sometimes the logged data was biased by the old recommender. Sometimes the model was evaluated on clicked items but never tested against the full decision context. Sometimes the system changed more than the model: a new retrieval source, a feature freshness change, a re-ranking rule, or a serving fallback.

This is why recommendation teams need backtesting, not just offline model validation.
‍

Offline evaluation, backtesting, and A/B testing are not the same thing

These terms are often blended together. They should be separated.

A mature recommender program uses all of them. Offline metrics help iterate. Backtesting catches obvious behavior problems. Shadow mode catches production issues. A/B testing proves real impact.

Skipping backtesting is like deploying a trading strategy because it looked good on a static classification metric.
‍

What backtesting means for recommendations

Backtesting asks: Given historical recommendation requests, what would the new system have shown?
‍

A backtest reconstructs the decision context as closely as possible:

The goal is not to pretend we can perfectly know what users would have done, we cannot do that. The goal is to understand whether the new system behaves plausibly and safely before users see it.
‍

The first backtesting question: “What changed?”

A recommendation system can change in many ways.


A practical backtesting report

A useful report might include:

The case studies matter. A spreadsheet of metrics can hide obvious product issues. Looking at twenty replayed examples often reveals problems faster than tuning another offline metric.
‍

The hard part: logged data came from the old system

Historical data is not neutral. Users only interacted with what the old system chose to show them. That means offline evaluation can reward models that imitate the old recommender rather than models that would make better decisions.

This is the core challenge behind off-policy evaluation. OPE methods try to estimate how a new policy would perform using data collected by a different policy. They can be powerful, but they rely on assumptions. Recent work on offline recommender evaluation emphasizes that unobserved confounding can create severely biased estimates and that the unconfoundedness assumption is not directly testable.
‍

What to log today so backtesting works tomorrow

Many companies cannot backtest because they did not log enough information.

At minimum, log:

Teams often log only the final recommendation and the click. That is not enough. If you do not log candidates, scores, features, rules, and context, you cannot explain why the system behaved differently.
‍

Backtesting diversity and personalization

Diving deeper into diversity of a recommender system results, there are three important questions:

  1. Given one user, are recommendations changing throughout the day, week, or month?
  2. Given two different users, are they receiving different recommendations?
  3. Which features boost popular items versus personalized items?

Backtesting is a good place to answer these.

A new model can improve average ranking metrics while making recommendations more repetitive for heavy users. Or it can increase diversity while hurting relevance for sparse users. Backtesting helps catch these patterns before the A/B test.
‍

Shadow mode: the bridge to production

After offline backtesting, run the new system in shadow mode.

In shadow mode, the new system receives live requests and produces recommendations, but users still see the old system. This tests:

Shadow mode is especially important for real-time recommenders because production traffic has edge cases that historical datasets miss: missing sessions, bot-like behavior, new catalog items, traffic spikes, cache misses, and malformed events.

Designing the A/B test

Backtesting reduces risk, but A/B testing measures user response.

A good A/B test should define:

The hardest part is choosing metrics that match the product goal. A system that improves click-through but reduces long-term satisfaction is not necessarily a win. Netflix has written about optimizing for long-term member satisfaction rather than only immediate engagement, and that principle applies broadly.
‍

A simple rollout policy

This is slower than shipping a model from a notebook. But it is faster than recovering from a bad recommender launch.
‍

Trust Is Built Through Real-World Testing

A recommender should earn trust by surviving a sequence of increasingly realistic tests.

Offline evaluation asks whether the model learned a pattern. Backtesting asks whether the system would have made reasonable past decisions. Shadow mode asks whether the system can run live. A/B testing asks whether users actually benefit.

Production personalization needs all four.
‍

Explore the Full Series

This article is part of our four-part series on production recommendation systems. Explore Part 1, Part 2, and Part 3.

‍

Research references

Forward Deployed Engineering.
In Your Environment.
In Your Time Zone.

Inside your standups, architecture decisions, workflows, and production environments
We build on your stack, for your business
1,000’s of AI & Data engagements across complex production environments
Discuss Your Challenge
Start Building

Continue Reading

Part 3: Better Recommendations
Accuracy Alone Is Not Enough
Part 2: Production Recommenders
Build the Full Ranking Funnel
Part 1: Real-Time Recommendations
Freshness Drives Relevance