Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade
MarkTechPost Asif Razzaq
Yandex built Sona, one model that does the whole recommendation job on smart speakers. It replaced a long cascade in testing and lifted listening, likes, and repeat commands.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Yandex’s Sona report is a shot across the bow for the old recommendation cascade. Instead of one model finding candidates, another filtering them, and a third making the final call, Sona folds the whole job into a single generative system. Yandex says it tested that switch in a seven-day live experiment on its smart speakers, and the new setup replaced more than 15 candidate generators plus the pre-ranker and ranker.
The cleaner part of the idea is also the bluntest: no hand-engineered features. Sona works from logged event fields such as track ID, artist ID, duration, likes, played time, and surface flags, along with learned Semantic IDs. For Yandex Music, that matters because playback on smart speakers can start without anyone first picking an artist, genre, or mood. The team calls that a pure-recommendation setting, which is a polite way of saying the model has to do the hard part on its own.
Under the hood, Sona leans on a shared user representation. The encoder reads listening history once, the decoder generates candidates, and a separate Ranking Module scores them against the same encoder state. The history handling is split: the most recent 2,048 events get heavier attention, while older events go through a lighter path. Yandex says that keeps most of the quality of full attention while cutting inference cost roughly in half.
The item side is equally engineered, just not by humans hand-crafting features. Each track is turned into three discrete codes. A frozen multimodal LLM reads the first 90 seconds of audio plus title, artists, and tags, then a four-layer refinement transformer lines those signals up with listening behavior. Residual K-means then compresses the result into three codebooks of 32,000 entries each. Yandex says that setup beat a CLMR audio baseline, with Recall@1000 rising from 0.8111 to 0.8524.
What makes this more than a neat architecture diagram is the live result. In the seven-day A/B test, on 15% of randomly selected users per split, Sona lifted active users by 4.53%, total listening time by 6.30%, likes by 11.42%, repeat commands by 17.99%, and deeply engaged users by 7.37%. Those gains sit on top of earlier improvements, and on active users the uplift was 2.35 times the +1.93% improvement Argus had previously delivered on the same surface.
My take — AI-written commentary, not fact-checked reporting
This is the sort of thing recommendation systems have been inching toward for years: fewer stitched-together stages, fewer hand-built feature piles, more one-model control. The nice irony is that the big win here is also the boring one — less plumbing, fewer excuses. Silicon Valley keeps calling this magic; Yandex just seems to have built a better funnel.
Read more about this at: MarkTechPost