Most production recommenders are cascades. Candidate generators feed a pre-ranker, which feeds a heavy ranker built on hundreds of engineered features. Yandex's Sona Technical Report describes a different design. Sona is a generative AI model that brings candidate generation and ranking into a single system, replacing the multiple stages typically used in recommendation pipelines. Yandex tested the model in a seven-day live production experiment on its smart speakers. In an online A/B test, it replaced more than 15 candidate generators, the pre-ranking stage, and the ranking stage with one served transformer.
What Problem Does Sona Solve?
Cascades split one decision across separately trained models. Each stage optimizes its own objective, and the ranker only sees what upstream stages let through. Yandex's previous stack on the Yandex Music surface consumed hundreds of features, including signals from Argus, Yandex's earlier recommender transformer. Sona puts candidate generation and ranking around one shared user representation. The encoder reads the listener's history once per request. A decoder generates candidates. A Ranking Module scores them against the same encoder states. No component uses hand-engineered features. Inputs are logged event fields (track ID, artist ID, duration, likes, played time, surface flags) and learned Semantic IDs.
On Yandex smart speakers, playback can begin without the user first selecting an artist, genre, or mood. The research team describes this as a pure-recommendation setting.
How Sona's Architecture Works
Sona is built around a shared encoder and a decoder, with a Ranking Module attached to the encoder's output. The encoder processes the user's interaction history as a sequence, producing a compact user representation. The decoder generates candidate items autoregressively from that representation, and the Ranking Module scores each candidate against the same encoder states. This design eliminates the feature-engineering pipeline that cascades require, replacing hand-crafted features with learned Semantic IDs and raw logged event fields.
Semantic IDs are learned discrete codes assigned to items, similar to tokenization in language models. They allow the model to generalize across items and to reason about item similarity without explicit feature engineering. The decoder generates candidates by predicting Semantic ID sequences, and the Ranking Module then re-scores those candidates using the shared user representation.
Production Results
In the seven-day live A/B test on Yandex smart speakers, Sona replaced more than 15 candidate generators, the pre-ranking stage, and the ranking stage with a single served transformer. Yandex reports that the system was served in production, demonstrating that a single generative model can absorb the functions of an entire recommendation cascade.
The experiment validates a broader shift in recommender systems: from multi-stage, feature-heavy pipelines toward unified generative architectures that share a single user representation across generation and ranking.
2026 Context
As of 2026, the recommendation landscape is increasingly moving toward generative and transformer-based approaches. Large-scale platforms have begun exploring unified models that replace traditional cascades, driven by advances in sequence modeling, Semantic ID learning, and efficient serving of transformer models at low latency. Yandex's Sona is one of the first documented production deployments of a single generative model replacing a full recommendation cascade, and its A/B test results provide evidence that this architecture can work at scale in a live environment.
Why It Matters
Sona's design points to a future where recommendation systems are simpler to maintain, easier to train end-to-end, and more consistent in their objectives. By eliminating hand-engineered features and multiple independently optimized stages, Sona reduces the risk of objective mismatch and information loss between stages. The approach also aligns recommendation more closely with recent advances in generative AI, where a single model handles multiple tasks through shared representations.
via MarkTechPost
