The playlist problem
A music app learns that you like one artist. It recommends that artist again. Then a similar artist. Then another similar artist. The songs are all plausible. The click model is not wrong. You might even listen for a while.
But after a week, the app starts to feel small.
This is one of the paradoxes of recommender systems: a system can be accurate and still create a bad experience. It can predict the next click while reducing discovery. It can increase short-term engagement while narrowing the user’s world. It can personalize so aggressively that two users see different recommendations, but each individual user sees the same kind of thing again and again.
That is why mature recommendation systems need diversity and explainability as first-class concerns and not just "post-launch polish".
This is also why “accuracy” is an incomplete word for these systems. A recommender is trying to be useful, trustworthy, fresh, varied, and aligned with product goals.
Diversity is not one metric
When teams say “we need more diversity,” they often mean different things.

A recommendation list can look diverse in one dimension and monotonous in another. Ten movies from ten different genres may all be old. Ten products from ten brands may all be at the same price point. Ten dating profiles may look diverse demographically but be nearly identical according to user behavior patterns. Ten articles may be from different publishers but have the same political framing.
So the first step is to name the dimension.
Diversity should be measured at multiple levels
A useful monitoring dashboard should be able to show CTR and conversion, but also whether the system is becoming narrower.

Those are production questions. They require logging impressions, item attributes, user segments, retrieval sources, ranker scores, and final positions. Without that data, diversity becomes a subjective complaint instead of an operating metric.
Popularity is not the enemy. Unchecked popularity is.
Popular items are popular for a reason. A system that refuses to recommend popular items will usually perform poorly. The problem is when popularity becomes the default answer for every uncertainty.
This happens because popular items have more data. They are easier to learn from, easier to validate offline, and safer as fallbacks. Over time, that creates a feedback loop:

The result is a “rich get richer” system. It may look good in aggregate metrics while quietly reducing personalization and catalog health.
A more mature system separates popularity from relevance:

This is where re-ranking and exploration matter. The system can keep popular items where they are useful while reserving controlled room for personalized, novel, or long-tail candidates.
Explainability: “why did we recommend this?”
Explainability has at least two audiences.
The first audience is the user. A user-facing explanation might say:
- “Because you watched…”
- “Popular near you”
- “Similar to items in your cart”
- “New from a creator you follow”
- “Trending in your network”
- “Recommended because you searched for…”
The second audience is the operator: product managers, merchandisers, editors, data scientists, support teams, and compliance reviewers. Their questions are more precise:

A practical explanation record
One lightweight pattern is to log an explanation record for each served slate.
This needs to preserve enough information to answer operational questions, so it's not necessary to expose raw sensitive features or proprietary model internals.
The strongest systems make explanations available at multiple levels:

Counterfactual thinking: “what would have changed the recommendation?”
Explainability becomes more useful when it supports counterfactual questions.
Instead of only asking “why did this happen?” ask:
- What minimal change in the user’s recent history would have changed the top recommendation?
- If the user had not clicked item X, would item Y still be recommended?
- If inventory availability changed, which items would replace the current slate?
- If we remove the popularity feature, does the slate become more personalized or just worse?
- If we increase diversity weight, which categories enter the top 10?
Recent research on sequential recommendation explanations has focused on exactly this kind of question: what changes in a user’s interaction sequence would lead to different recommendations? Other work evaluates explanations by measuring their counterfactual impact on recommendations.
The production value is clear. Counterfactual tools help teams debug ranking behavior without relying only on intuition.
Diversity and explainability belong together
Diversity without explainability can feel arbitrary. Explainability without diversity can reveal that the system is narrow.
Together, they let teams answer richer questions:

These questions make personalization auditable.
How to improve diversity without breaking relevance
The right mechanism depends on the failure mode.

The key is to avoid treating diversity as a universal slider. Too much diversity can make the product feel random. Too little makes it feel stale. Different surfaces need different tolerances.
A search results page may need tight relevance. A homepage carousel may tolerate more discovery. A “because you watched” row should stay semantically coherent. A “new for you” row should intentionally explore.
What to monitor before and after a diversity change

Personalization Must Be Explainable
The worst recommender is not always the one that is obviously wrong. Sometimes it is the one that is narrowly right.
A mature personalization system should know how to answer four questions:
- Why did this user receive this recommendation?
- Would another user receive something different?
- Would this same user receive something different tomorrow?
- What happens if we change the features, constraints, or diversity weights?
When a team can answer those questions, personalization becomes more than prediction. It becomes a system the organization can understand, trust, and improve.
Explore the Full Series
This article is part of our four-part series on production recommendation systems. Explore Part 1, Part 2, and Part 4.
Research references
- Google Developers — Recommendation systems overview (last updated 2025-08-25): https://developers.google.com/machine-learning/recommendation/overview/types
- Google Research — Deep Neural Networks for YouTube Recommendations: https://research.google/pubs/deep-neural-networks-for-youtube-recommendations/
- Meta Engineering — Scaling the Instagram Explore recommendations system: https://engineering.fb.com/2023/08/09/ml-applications/scaling-instagram-explore-recommendations-system/
- LinkedIn Engineering — Concourse: Generating Personalized Content Notifications in Near-Real-Time: https://engineering.linkedin.com/content/engineering/en-us/blog/2018/05/concourse--generating-personalized-content-notifications-in-near
- arXiv — Counterfactually Evaluating Explanations in Recommender Systems: https://arxiv.org/abs/2203.01310
- arXiv — Offline Recommender System Evaluation under Unobserved Confounding: https://arxiv.org/abs/2309.04222
- arXiv — Demystifying Sequential Recommendations: Counterfactual Explanations via Genetic Algorithms: https://arxiv.org/abs/2508.03606
- arXiv — FedFlex: Federated Learning for Diverse Netflix Recommendations: https://arxiv.org/abs/2507.21115


