✦ STORY

Netflix Algorithm Improvements Shifted Viewing Away From Its Biggest Hits

Published

A couple watching netflix


▣ DATA BRIEF · STREAMING ALGORITHMS

☀ Key Stats

◆ Improvements to Netflix’s recommendation system reduced the concentration of title plays by 1.2%, shifting viewing away from the most popular titles toward moderately popular ones.

◆ Recommendation concentration itself fell by 5.7%, according to the study’s title-level concentration measure.

◆ The experiment involved 8,559,252 Netflix subscribers over a 60-day period from February to April 2025.

◆ Users receiving the updated recommendation system played 1.2% more distinct titles than users kept on the older system.

◆ Overall plays increased 0.62%, while total viewing hours rose 0.37%.

◆ The number of days on which users played something increased 0.21%.

◆ The share of first-in-session plays originating from recommendations increased 0.5%, while plays per visit rose 0.47%.

◆ Users completed 1.3% more titles under several watch-duration measures and gave 0.57% more positive ratings.

◆ The experimental treatment combined 12 major algorithmic changes, so the results cannot be attributed to any single recommendation update.


Continue reading ↓


Improvements to Netflix’s recommendation system reduced the concentration of viewing across titles by 1.2%, while shifting plays away from the platform’s biggest hits toward moderately popular content, according to a new research paper.

The experiment covered 8,559,252 subscribers over 60 days and compared an older, frozen recommendation system with Netflix’s evolving production system.

The updated system also increased the number of distinct titles users played by 1.2%, overall plays by 0.62% and viewing hours by 0.37%.

The results suggest that better personalization did not simply push viewers deeper into already dominant hits or toward extremely obscure titles.

◎ HEADLINE NUMBER

1.2% less concentrated

Title plays became less concentrated after users received the newer recommendation system.

The experiment covered more than 8.5 million subscribers

The researchers studied a Netflix experiment conducted over 60 days between February and April 2025.

One group remained on the recommendation system that existed before the experiment, while the other experienced the production system as Netflix continued releasing improvements.

The treatment included 12 major changes spanning areas such as homepage ranking, new model features, changes to model architecture and different weighting of prediction targets.

The user interface itself did not change, allowing the researchers to focus on changes generated by the recommendation technology.

Important distinction: The experiment measures the combined effect of a package of recommendation improvements. It does not show which of the 12 individual changes produced each result.

Recommendations became 5.7% less concentrated

The largest concentration effect appeared before viewers even pressed play.

The researchers found that the concentration of titles appearing in recommendations fell by 5.7%.

They measured concentration using the Herfindahl-Hirschman Index, a metric that rises when activity is concentrated among a smaller number of titles.

The decline means the improved system distributed recommendation exposure across a broader range of content.

Most of the additional exposure went to what the researchers call the middle-tail, rather than the least popular part of the catalog.

↳ HOW THE STUDY DIVIDED NETFLIX TITLES

Bottom 50%

Long-tail titles

Next 45%

Middle-tail titles

Top 5%

“Superstar” titles

The researchers classified titles according to their position in the combined play distribution across the treatment and control groups.

The long-tail showed little change, while the middle 45% gained recommendation share at the expense of the top 5%.

Viewers followed the shift toward middle-ranking titles

The change in recommendations translated into a similar change in what subscribers actually played.

Viewing moved away from the superstar group and toward middle-tail titles, while the long-tail again changed little.

As a result, the concentration index for plays fell by 1.2%.

The result runs against the idea that increasingly sophisticated recommendation systems necessarily funnel attention toward both mega-hits and highly niche content while hollowing out the middle.

↔ WHERE VIEWING SHIFTED

Superstars ↓
Middle-tail ↑
Long-tail ≈ little change

The authors’ explanation is that very popular titles already generate enough viewing data for recommendation systems to identify likely audiences fairly well.

Moderately popular titles have smaller audiences but still generate enough information for improved algorithms to become better at matching them to individual users.

The most niche titles may still have too little interaction data for incremental improvements to make as much difference.

Users watched 1.2% more distinct titles

The broader distribution of viewing did not come at the expense of overall engagement during the experiment.

Instead, all four of the paper’s main engagement measures increased.

+1.2%

Distinct titles played

+0.62%

Overall plays

+0.37%

Viewing hours

+0.21%

Days with a play

The percentages are small in absolute terms, but the study’s unusually large experimental sample allows the researchers to estimate them precisely.

The findings also show that greater content diversity and higher engagement moved in the same direction during this particular experiment.

↳ USER BEHAVIOR · RECOMMENDATION RELIANCE

More plays started from Netflix recommendations

The researchers also examined how users reached the first title they played in a viewing session.

The share originating from Netflix recommendations increased by 0.5%, while the number of plays per visit rose 0.47%.

A ranking measure based on how high the played title appeared on the homepage increased by 0.1%.

Together, those measures indicate that users relied slightly more heavily on the recommendation system after the upgrades.

Users also found more titles they kept watching

Higher play counts alone would not necessarily mean the recommended titles were better matches.

The researchers therefore examined how long viewers continued watching titles and whether they gave them positive ratings.

Users completed 1.3% more titles at the study’s 15-, 30- and 45-minute viewing thresholds and provided 0.57% more positive thumbs.

That does not mean the average newly played title was unambiguously better.

Some measures of match quality as a share of all plays declined slightly, which the authors say can occur when better recommendations encourage users to start additional, more marginal titles they otherwise would not have watched.

What the result supports: Users found and played more titles they watched for substantial periods. The experiment does not establish that every additional recommendation was a higher-quality match.


✦ Why it matters ✦

Recommendation systems influence which movies, television shows, music, products and news stories people encounter on large digital platforms.

A long-running concern is that personalization can reinforce already popular products because those products generate more interaction data for algorithms to learn from.

The Netflix experiment suggests that this effect can change as recommendation technology improves.

In this case, better recommendations shifted attention away from the biggest hits, but they did not push much additional viewing into the most obscure half of the catalog.

The main beneficiaries were titles in the middle: popular enough to produce meaningful viewing data but not so dominant that the older system could already identify their audiences easily.

That distinction could matter for streaming-platform content strategy because it suggests the potential return from moderately popular titles may increase as personalization gets better.

The evidence comes from one platform, however, and does not establish that recommendation upgrades on music, shopping, social media or other services would redistribute consumption in the same way.

ⓘ How to read the findings

The experiment involved 8,559,252 Netflix subscribers over a 60-day period from February to April 2025. The paper does not provide a geographic breakdown of the experimental sample, so the results should not be described as representing a particular country.

The treatment was not one isolated algorithm change. It combined 12 major recommendation-system updates, meaning the study identifies the effect of the package rather than the contribution of each component.

The control group continued using the pre-experiment recommendation system while the production treatment received improvements released during the test. The interface itself remained unchanged.

The 5.7% recommendation-concentration decline and 1.2% play-concentration decline are relative changes in the Herfindahl-Hirschman Index. They do not mean that 5.7% or 1.2% of recommendations or plays disappeared.

The researchers define “superstar” titles as the top 5% of the play distribution, “middle-tail” as the next 45% and “long-tail” as the bottom 50%. Those are analytical categories created for the study rather than Netflix consumer labels.

The engagement effects — including 0.37% more viewing hours, 0.62% more plays and 1.2% more distinct titles — are relative percentage changes compared with the control group, not percentage-point changes.

The findings describe the effect of incremental recommendation improvements over roughly two months. They do not measure what would happen if Netflix removed personalization entirely or replaced its current system with a much older algorithm.

The study measures plays, viewing time, recommendation exposure and platform feedback. It does not measure broader viewer welfare, subscription retention, content-production costs or whether users considered the overall catalog objectively better.

All six authors were employed by Netflix during the project. Some also list affiliations with Northwestern University’s Kellogg School of Management, Cornell University or Cornell Tech.

The paper was posted online in August 2026 and should be treated as preliminary research. A peer-reviewed journal publication of this specific paper was not identified at the time of writing.

Author

◆ ◆ ◆

✦ Keep exploring

More From This Topic

◆ ◆ ◆



Discover more from StatsJournalist.com

Subscribe now to keep reading and get access to the full archive.

Continue reading