Part 2: Production Recommenders

See how production recommenders move from millions of items to a relevant, reliable, and measurable personalized experience.

Key Takeaways:

Retrieval Finds the Right Candidates
Ranking Adds Precision and Context
Re-Ranking Makes Results Production-Ready

The impossible request

A product manager asks for a simple thing:

“Show each user the ten best items.”

The catalog has fifty million items. The page needs to load in under 200 milliseconds. Some items are out of stock. Some are too similar to each other. Some are legally ineligible in certain regions. Marketing wants campaign exposure. The trust and safety team has filters. Finance wants margin-aware ranking. The ML team wants to use a richer model, but the serving team is worried about latency.

The phrase “the ten best items” has quietly become a distributed systems problem.

This is why modern recommender systems are usually staged. Google’s recommendation-system overview describes a common architecture with candidate generation, scoring, and re-ranking. Meta’s Instagram Explore system uses a multi-stage funnel with retrieval, first-stage ranking, second-stage ranking, and final re-ranking. YouTube’s well-known production paper also frames recommendations as a large-scale candidate-generation and ranking problem.

The pattern exists because no single model can do everything well at once. Retrieval needs scale. Ranking needs precision. Re-ranking needs product judgment. Serving needs reliability. Experimentation needs measurement.

The rest of this post walks through the system as it actually tends to look in production.
‍

The production funnel

Each stage has a different job. The mistake is trying to make one stage solve another stage’s problem.


Stage 1: Retrieval is about not missing good options

Retrieval answers the question: from a huge universe, which items are worth considering?

It needs high recall under strict latency and cost limits, not perfect ordering. If retrieval drops the right item too early, the ranker never gets a chance to fix it.

Common retrieval sources include:

A strong retrieval layer usually combines multiple sources. Meta’s Instagram Explore write-up is explicit about this: retrieval sources can be heuristic or ML-based, real-time or pre-generated, and the system mixes different source types with tunable weights.
‍

Why two-tower retrieval shows up everywhere

Two-tower models are popular because they split the world into two encoders:

The item tower can precompute item embeddings and store them in a vector index. At request time, the system computes or retrieves a user/session embedding and performs approximate nearest-neighbor search.

This works well when the catalog is large because the system avoids scoring every item with a heavy model. It also creates a clean separation between retrieval and ranking. Retrieval gets plausible candidates quickly; ranking can spend more compute on a smaller set.º

But two-tower retrieval has limits. The dot product or similarity score is usually too simple to capture every business constraint, cross-feature interaction, and product nuance. That is why it is usually the beginning of the funnel, not the whole funnel.
‍

Stage 2: Ranking is where the system gets more opinionated

The ranker answers a narrower question: given a few hundred or thousand candidates, which ones are best for this user, context, and product surface?

Because the candidate set is smaller, the model can use richer features:

The ranker may optimize for click probability, conversion probability, watch time, expected revenue, retention, satisfaction, or a weighted mix of objectives.

A simple click model can be useful as a baseline. But if it becomes the only objective, the system can learn bad habits: clickbait, repetition, popularity bias, short-term engagement loops, or over-personalization. This is why large production systems increasingly use multi-objective ranking and guardrail metrics.

The ranker should answer: “what should this product choose to show, given user value, business value, and long-term trust?”
‍

Stage 3: Re-ranking is where the product becomes real

A pure ranking model might put ten nearly identical items at the top. Technically, they scored highest. Product-wise, the slate is bad.

Re-ranking turns model scores into an experience. It may apply:

This layer is often underestimated because it looks less glamorous than the model. In reality, re-ranking is where many of the product and business requirements live.

It also makes the system more explainable to operators. If the final slate is poor, you can ask: did retrieval miss variety? Did the ranker over-score a category? Did re-ranking over-constrain the slate? Did a business rule dominate the model?

Without these stage boundaries, debugging becomes guesswork.
‍

Stage 4: Serving is where good notebooks go to fail

A ranking model can look excellent offline and still fail in production.

Production serving needs:

The serving layer is also where batch, nearline, and real-time design choices become visible. Some data is precomputed. Some is fetched from online stores. Some is computed from the request. Some is cached. Some is deliberately omitted because it is too expensive or unstable.

LinkedIn’s Concourse example is useful here. Their near-real-time scoring platform had to support fanout, feature decoration, and model-based scoring at very high scale. The blog discusses local RocksDB stores, partitioning by recipient, Kafka topics, and feature data refresh strategies.
‍

A production-ready view of the stack

A recommender is cross-functional by nature. If ownership is unclear, the system will drift: data breaks, features change, experiments become untrusted, and nobody knows why a recommendation appeared.
‍

The architecture should preserve debuggability

The more complex the model, the more important it is to keep the system inspectable.

A good recommender should let the team answer questions like:

This does not mean every user-facing explanation must expose internal model details. It means internal teams need enough traceability to operate the system.

Recommendation systems are feedback loops. When they fail quietly, they can train on their own mistakes.
‍

A practical build sequence

Teams often want to build the whole stack at once. That is risky. A better sequence is:

This sequence makes the system useful before it becomes sophisticated.

‍

Explore the Full Series

This article is part of our four-part series on production recommendation systems. Explore Part 1, Part 3, and Part 4.

‍

Research references

Forward Deployed Engineering.
In Your Environment.
In Your Time Zone.

Inside your standups, architecture decisions, workflows, and production environments
We build on your stack, for your business
1,000’s of AI & Data engagements across complex production environments
Discuss Your Challenge
Start Building

Continue Reading

Part 4: Recommender Backtesting
Test Before You A/B Test
Part 3: Better Recommendations
Accuracy Alone Is Not Enough
Part 1: Real-Time Recommendations
Freshness Drives Relevance