The model looked better offline
The ML team trains a new ranker. Offline metrics improve. Precision@k is up. NDCG is up. The validation curves look clean. Everyone is excited.
Then the A/B test launches.
CTR is flat. Conversion drops for new users. Latency increases. A few high-margin categories disappear from the top slots. Customer support receives complaints that recommendations feel repetitive. Product leaders ask the obvious question:
“Why didn’t we catch this before launch?”
Sometimes the answer is that the offline metric was too narrow. Sometimes the logged data was biased by the old recommender. Sometimes the model was evaluated on clicked items but never tested against the full decision context. Sometimes the system changed more than the model: a new retrieval source, a feature freshness change, a re-ranking rule, or a serving fallback.
This is why recommendation teams need backtesting, not just offline model validation.
Offline evaluation, backtesting, and A/B testing are not the same thing
These terms are often blended together. They should be separated.

A mature recommender program uses all of them. Offline metrics help iterate. Backtesting catches obvious behavior problems. Shadow mode catches production issues. A/B testing proves real impact.
Skipping backtesting is like deploying a trading strategy because it looked good on a static classification metric.
What backtesting means for recommendations
Backtesting asks: Given historical recommendation requests, what would the new system have shown?
A backtest reconstructs the decision context as closely as possible:

The goal is not to pretend we can perfectly know what users would have done, we cannot do that. The goal is to understand whether the new system behaves plausibly and safely before users see it.
The first backtesting question: “What changed?”
A recommendation system can change in many ways.

A practical backtesting report
A useful report might include:

The case studies matter. A spreadsheet of metrics can hide obvious product issues. Looking at twenty replayed examples often reveals problems faster than tuning another offline metric.
The hard part: logged data came from the old system
Historical data is not neutral. Users only interacted with what the old system chose to show them. That means offline evaluation can reward models that imitate the old recommender rather than models that would make better decisions.
This is the core challenge behind off-policy evaluation. OPE methods try to estimate how a new policy would perform using data collected by a different policy. They can be powerful, but they rely on assumptions. Recent work on offline recommender evaluation emphasizes that unobserved confounding can create severely biased estimates and that the unconfoundedness assumption is not directly testable.
What to log today so backtesting works tomorrow
Many companies cannot backtest because they did not log enough information.
At minimum, log:

Teams often log only the final recommendation and the click. That is not enough. If you do not log candidates, scores, features, rules, and context, you cannot explain why the system behaved differently.
Backtesting diversity and personalization
Diving deeper into diversity of a recommender system results, there are three important questions:
- Given one user, are recommendations changing throughout the day, week, or month?
- Given two different users, are they receiving different recommendations?
- Which features boost popular items versus personalized items?
Backtesting is a good place to answer these.

A new model can improve average ranking metrics while making recommendations more repetitive for heavy users. Or it can increase diversity while hurting relevance for sparse users. Backtesting helps catch these patterns before the A/B test.
Shadow mode: the bridge to production
After offline backtesting, run the new system in shadow mode.
In shadow mode, the new system receives live requests and produces recommendations, but users still see the old system. This tests:

Shadow mode is especially important for real-time recommenders because production traffic has edge cases that historical datasets miss: missing sessions, bot-like behavior, new catalog items, traffic spikes, cache misses, and malformed events.
Designing the A/B test
Backtesting reduces risk, but A/B testing measures user response.
A good A/B test should define:

The hardest part is choosing metrics that match the product goal. A system that improves click-through but reduces long-term satisfaction is not necessarily a win. Netflix has written about optimizing for long-term member satisfaction rather than only immediate engagement, and that principle applies broadly.
A simple rollout policy

This is slower than shipping a model from a notebook. But it is faster than recovering from a bad recommender launch.
Trust Is Built Through Real-World Testing
A recommender should earn trust by surviving a sequence of increasingly realistic tests.
Offline evaluation asks whether the model learned a pattern. Backtesting asks whether the system would have made reasonable past decisions. Shadow mode asks whether the system can run live. A/B testing asks whether users actually benefit.
Production personalization needs all four.
Explore the Full Series
This article is part of our four-part series on production recommendation systems. Explore Part 1, Part 2, and Part 3.
Research references
- Google Developers — Recommendation systems overview (last updated 2025-08-25): https://developers.google.com/machine-learning/recommendation/overview/types
- Google Research — Deep Neural Networks for YouTube Recommendations: https://research.google/pubs/deep-neural-networks-for-youtube-recommendations/
- Meta Engineering — Scaling the Instagram Explore recommendations system: https://engineering.fb.com/2023/08/09/ml-applications/scaling-instagram-explore-recommendations-system/
- LinkedIn Engineering — Concourse: Generating Personalized Content Notifications in Near-Real-Time: https://engineering.linkedin.com/content/engineering/en-us/blog/2018/05/concourse--generating-personalized-content-notifications-in-near
- arXiv — Counterfactually Evaluating Explanations in Recommender Systems: https://arxiv.org/abs/2203.01310
- arXiv — Offline Recommender System Evaluation under Unobserved Confounding: https://arxiv.org/abs/2309.04222
- arXiv — Demystifying Sequential Recommendations: Counterfactual Explanations via Genetic Algorithms: https://arxiv.org/abs/2508.03606
- arXiv — FedFlex: Federated Learning for Diverse Netflix Recommendations: https://arxiv.org/abs/2507.21115


