Metrics aren’t neutral: they reflect the interests of whoever built them

Ever wondered why audience metrics don’t seem to fully capture the truth, and what the long-term, macro effects are on news media? This is the issue that Sophie Chauvet investigated during her PhD research, and a topic she shared on stage at our recent Audiencers’ Festival in Paris. Metrics are misunderstood, and following them blindly may have high costs for journalists and the democratic health of media landscapes. Opening the black box however may offer clues as to how to really seize their potential. It lies in seeing them for what they really are: not a defined reality, but indicators of power dynamics at the heart of internet economics, mostly defined by statistics, and with very real risks for media diversity and journalistic autonomy.

Key takeaways to copy & share
> Audience metrics are not a neutral reading of what readers want. They also encode the interests of whoever built the model. That is the conclusion of a PhD study of the data monetisation ecosystem, drawing on hundreds of white papers, events and interviews across ad tech, platforms, measurement and analytics companies

> Measurement was never a picture of reality. Nielsen and Mediametrie treated the audience as a convention, a figure agreed between media and advertisers in order to predict revenue. The internet's promise that clicks were finally direct feedback from the public is what makes the confusion possible

> Platforms do not just supply the numbers, they collect them back. Running across every site they are installed on gives them an overview no publisher has, which they use to optimise their own advertising models. Those models are rarely audited, and the entropy in the system (bots, spam, unaudited incentives to inflate) is not theirs to disclose

> Statistical prediction needs scale, so the system rewards whoever was already big. Large outlets negotiate directly with platforms over which metrics matter, and those models then ship as the free analytics smaller outlets use. Four consequences follow: winner-take-all monetisation, hierarchies that are hard to challenge, consolidation incentives, and volume rewarded over original reporting

> The sharpest line in the research: adopt the metrics platforms provide, and your strategy becomes symmetrical to platforms' interests. The argument is not to abandon metrics but to treat them as indicators rather than objectives, and to know what territory is being mapped before you choose them

The paradox: Are metrics bullsh*t?  

In the last 20 years, newsrooms have seen the ubiquitous rise of audience analytics. These came with a promise: gauging who and how many people consumed news online, and how to better optimise resources to ensure maximum efficiency and profitability. 

This promise went way beyond simply knowing audiences better: it implied re-structuring newsrooms, integrating data science expertise, and welcoming the new rules of digital content distribution in editorial strategies, often hidden behind simplified, easy-to-use yet inscrutable dashboards. 

At the same time, as clickbait content rose with the hope of it funding more serious, investigative and high democratic impact news, journalistic identities were shaken to their core: what does it mean to be a journalist in service of the people, if what the people want – represented by these numbers, is to be entertained? 

And yet, 20 years after analytics first appeared in newsrooms, even digital-native news outlets that made the choice to fully abide to metrics, such as BuzzFeed and Vice Media, don’t seem to be doing so well, while others have more successfully ridden the digital transition wave. How can we explain this paradox? Some would argue, metrics are bullshit.

With this rumor in mind, and a passion for journalism’s service to society, I started my research in 2018 to figure out what was at stake. I wanted to understand the paradox of metrics, why they were so elusive, and what and how their broader effects might harm or benefit news outlets. 

What I have found is that metrics are neither bullshit, nor a catch-all solution to news outlets’ jeopardised financial situation, but rather complex tools deeply re-shaping media landscapes according to both brute and hardly changeable statistical parameters, and more whimsical, mostly platforms-led power dynamics that derive from the statistical authority inscribed in the black box of metrics. 

Opening the black box: tools and questions 

What I have found, in a nutshell, is that the bigger the size, the bigger the wins. But to find out about this simplified insight, I first had to take a step back from newsrooms, find the right tools, and ask key questions in order to open the black box of analytics the best I could.  

First, where do metrics come from? A genealogical approach proved to be deeply informative.  

Second, who profits from metrics? Following the money, and identifying who defines them, helped identify the main technical and political dynamics at their core. That’s where the macro approach became fruitful. 

Third, how realistic are metrics? There, I investigated what happens when a quantified measurement captures a social process, and what gets erased.  

After analysing hundreds of white papers, events, and interviews with practitioners across the data monetisation ecosystem, including ad tech, platforms, measurement and analytics companies, there was a single metaphor that helped to understand the stakes of metrics, and why they may negatively re-shape media landscapes: confusing the map for the territory, or how metrics tend to be self-fulfilling prophecies, thus impacting reality based on biased premises. 

The crux of metrics: Self-fulfilling prophecies, or why we confuse the map for the territory

Confusing the map for the territory is an expression borrowed from Alfred Korzybski, which is often used in quantification studies to describe the nefarious impacts of ill-used measurement. It means thinking the measurement model associated with an object is as real as the object. In the context of metrics, this helps to understand why there might be some confusion as to whether clicks mainly represent the audience’s interest, when in fact, they also represent the interests of those who built the model of metrics.  

Let’s break this down further, by answering the first question: where do metrics come from? Here, the history of audience measurement, as told among others by Cécile Méadel, Philip Napoli, and Josiane Jouët, is particularly enlightening. To practitioners in companies such as Médiamétrie and Nielsen, who historically developed studies on audience measurement on television and radio thanks to their Audimat and Peoplemeter, the audience was understood to be more of a convention than a reality, a means to gauge advertising revenue, with measurements mostly agreed upon by the different media involved.  

As such, measurements hold a kernel of truth, but they also serve a function of coordination between players in order to synchronise, predict and project financial benefits, with measurement containing an added layer of filters based on the parties’ main interests, which are also audited in order to ensure some degree of fairness within the ecosystem.  

With the rise of the internet in the late 90s, the first analytics companies, such as Urchin in 1998 that was acquired by Google, emerged with the promise of more realistic measurement. The internet was sold at the time as a new medium where clicks served as the new proxies for direct and real-time feedback from the public, through personal computers.  

This promised to help bypass the often heavy, costly, and slow studies of more traditional audience measurement. This promise was then thought to be fulfilled exponentially with the rise of social media and their troves of personal data waiting to be collected at a scale never seen before.  

This helps explain why the public is sometimes conflated with analytics, or why the map is confused for the territory, as analytics tools entered the newsroom in the 2010s, a field traditional measurement companies hardly ever dared to enter so as to protect journalistic integrity from advertising’s interests.  

Who profits from confusing metrics? 

The rise of analytics in newsrooms coincided with the rise of platforms and real-time data collection at scale. Here, the answer to the second question, who profits from metrics, is answered. Indeed, one aspect that contributed to making platforms so powerful is that their models rely partially not only on providing real time data on website use, but also on collecting this data back to improve their own models, by maintaining a more global overview on all the websites they are implemented on.  

With this broader perspective on audience flows, they are able to optimise advertising models which are key to their financial monopoly over the internet. These models however are black-boxed, as platforms are rarely audited.  

This has severe consequences on the internet ecosystem. For example, recent trials have proved that Google have abused this position of power over data flows by choosing specific, opaque algorithms that prioritised them over news outlets in an undue fashion, leading the latter to lose millions in revenue.  

But beyond the potential for fraud, these black boxes, this network of metrics used by many newsrooms, also contain the different ties of coordination between a slew of internet players that monetise data across the ecosystem.  

This, added to the complexity of the technical layers that sustain online consumption, and the un-audited incentives for inflating numbers, not to mention spam and bots, contribute to a phenomenon of entropy, which help explain why metrics do not always make sense. Disclosing this entropy is not in the interest of platforms. 

Overall in that regard, metrics, when provided by bigger, non-journalistic players, could be regarded not as direct public feedback, but rather as a reflection of bundled interests of the technical infrastructure of the internet, and the political economic negotiations that happen often behind closed doors. 

Metrics have therefore kept their function of coordinating the interests of the players involved, but without much transparency, and with an added layer of realism which is heavily advertised by many internet players, whether in online advertising, or via analytics companies.  

The benefits of scale: Supercharging the Goodhart’s law  

Now there is perhaps one factor that is the most significant when it comes to further understanding why metrics are confused for reality, and why they tend to become self-fulfilling prophecies with deep impacts on media landscapes: scale. This is due both to the fact that metrics are made of statistics, and because online audience flows are dominated by platforms.  

First, the statistical nature of online data flows highlights the importance of scale in monetising data online. This is caused by the fact that most predictions online are based on probabilities, and their associated deterministic patterns.  

This statistical technique helps to observe long term audience consumption patterns and how to optimise the most fine-tuned, cost-efficient audience trajectories according to a specific click architecture. In order to fulfil the promise that predictions become true, scale is a sine qua non condition in order to decipher what historical patterns are the most profitable.  

A corollary of this is that this deterministic parameter tends to cement historical repetitions, and to further benefit media players who existed first, as both benefits of scale tend to generate cumulative effects, and to provide a reputation for reliable profit which is highly valuable when it comes to organising and predicting budgets.  

This one main parameter is what allows platforms, whose main objective is to accumulate data, to maintain a stronghold on media landscapes, by adopting this meta-deterministic perspective on data patterns, and therefore, holding a monopolistic, black-boxed overview of what performs best online. 

By using probabilities and providing metrics as a map, platforms may supercharge the effect of self fulfilling prophecies online, through metrics that often tend to become objectives rather than simple indicators. 

This mechanism is called the Goodhart law. When metrics are used as objectives, they cease to be good metrics. In the case of metrics in news, it means that using scale in order to fulfil the deterministic parameter of finding truth through statistics tends to create self-fulfilling prophecies. By incentivising scale, volume, clicks, and page views, there are some negative macro effects that come as a consequence of mis understanding the power of metrics.  

Macro impact #1: The winner-take-all effect  

First, there is a winner-take all effect. The bigger the media, in terms of volume of clicks and data, the bigger the possibilities of monetising the data fruitfully. This winner-take-all effect is similar to the one observed in the business model of platforms, where scale advantages tend to become cemented and supercharged once a certain volume is reached. The bigger the media in terms of data volume, whether through subscriptions or through the accumulation of page views, the bigger the monetisation opportunities. The benefits of owning historical data tend to be cumulative as the predictions are more reliable and more precise.  

Furthermore, metrics models provided by platforms are often based on those bigger media outlets, as platforms are naturally more interested in those large pools of data that benefit their own meta-deterministic models of data analysis. In that sense, adopting the metrics provided by platforms tends to mean, that one’s strategy becomes symmetrical to platforms’ interests.  

Macro impact #2: Media hierarchies  

A second macro effect of this model is that it tends to also favour media hierarchies that are difficult to challenge. Bigger outlets remain at the top in terms of traffic and monetisation, because platforms tend to model their metrics and software according to those levels of scale and volume, which is what is of interest to them. Bigger outlets are also those who benefit more from in-person negotiations with platforms as to what metrics matter, and those models tend to be those that are then provided in free versions of analytics software that are more accessible to smaller outlets.  

Furthermore bigger outlets with more financial resources are more able to tweak analytics parameters, whereas smaller outlets tend to have more ready-made, black-boxed tools that may not be that adaptable to their own needs. Those outlets who have higher volumes of data also tend to be favoured in negotiations with other players of the data monetisation ecosystem, such as ad tech companies, as they benefit from a solid monetisation reputation.  

This model is not necessarily new, as traditional measurement companies also used similar hierarchical, rankings-based models, where media outlets with more resources and established historical deterministic patterns, have access to more precise and efficient measurement. 

In this digital analytics model therefore, new players have fewer chances to compete in terms of data monetisation because they lack the historical reputation and the reliability of accumulated historical patterns that unlock trust in predictions and funding opportunities. Therefore, the parameters of determinism, scale, and self-fulfilling prophecies are revealed by the simple fact that measurement and metrics are based on statistics and projections.  

Macro impact #3: Media concentration  

The data monetisation benefits provided by scale include a third significant macro impact on media landscape: media concentration incentives. In order to compete in this quantified, rankings-based system, some media outlets have adopted strategies to access those scale benefits. By creating alliances, accumulating media outlets, and re-joining databases, more volume can be reached. In that sense, media conglomerates can be incentivised to optimise online advertising and data monetisation opportunities, with a potential risk for media diversity.  

Macro impact #4: Devaluing journalism 

The final macro consequence that derives from the volume incentive, which is perhaps the most significant of all, is the devaluation of journalistic standards. From a strictly volume-oriented data monetisation perspective, it makes sense to produce more content which translates into more clicks, at the expense of journalistic standards. While clicks and webpages can easily be inflated from a quantitative perspective, the labour that tends to come with quality, original reporting, which often takes time, tends to be disregarded, from the perspective of this quantified financial model.  

In that sense, favouring financial value to access scale benefits tends to empty journalism from its values, and the value of its labour. Some strategies adopted by some media outlets include quickly producing articles at scale, by translating them into different languages, copy-pasting news wire reports, and coordinating with trending online keywords that do not necessarily stem from organic audience behaviours or from journalists’ flair.  

Journalists have sometimes conceded, when interviewed for this research, that they felt they had to twist their instincts in order to ensure content production that may have felt industrial rather than the byproduct of careful listening of their public’s needs. On the brighter side, another strategy also includes the attempt to quantify the value of journalism’s democratic mission, by promoting the quality of the involvement with the public as a selling point.  

Other media employees, involved more in advertising monetisation rather than direct editorial output, have also mentioned how adopting these keyword-tweaking strategies, informed by listening to more global internet trends provided by platforms, tend to be fruitful. Small tweaks, multiplied by thousands of webpages, thanks to benefits of scale, result in interesting profits.  

However, these strategies also tend to benefit platforms’ meta-deterministic needs, with consequences on how social issues covered by journalists are phrased online. With certain keywords being chosen over others, in order to fulfil the internet’s echo-chamber of what gets attention or not, meaning, what is observed statistically as echoing what was said in the past or by others, platforms mechanically end up having quite a say as to how social issues might be framed. 

In the era of AI, this editing power based on historical patterns might be supercharged, as suggestions not only contain keywords but sometimes entire sentences. Mechanically, this also runs the risk of framing issues in a more stereotypical way.  

Re-seizing the power of metrics 

Metrics are rarely neutral tools. They benefit some and devalue others, according to decisions that are black boxed behind their deceptive simplicity. They might contain a kernel of truth, but they often reflect the interests of those who manufactured them. In the context of news, they might be one reflection of the public’s clicks, but this reflection is also heavily mediated by the infrastructure of those who are able to have an overview of the internet’s secret flows and complex behind-the-door negotiations.  

Theoretically, metrics should be seen as indicators, not as objectives. In practice however, they tend to re shape media landscapes in a way that is detrimental to media diversity, journalistic autonomy, and may contribute to journalists’ loss of meaning.  

Following a map provided by non-journalistic players may lead journalism into platforms’ projected territory, one that is biased and mostly driven by scale, where truth is defined according to statistical parameters and their own interests. That’s the price to pay when meaningful, values-oriented work is translated into simplified quantified measurements.  

But I would argue that the potential of metrics to really benefit news remains unexplored. Metrics, when seized with the awareness of what they can and cannot do, can be powerful tools to visualise injustice and quantify the reality of invisible social issues.  

When the self-fulfilling prophecy aspect of measurement is mitigated, and when metrics are seen as simple indicators, they can be used to jolt actions that favour values rather than value. Knowing the limits of metrics, and being aware of what territory is being mapped when choosing them, is therefore key to secure journalism’s interests and social mission.  

This project has received funding from the European Union’s Horizon 2020 research and innovation programme under the Marie Skłodowska-Curie grant agreement No 765140 – based at Samsa and with the University of Toulouse.