Posted on

Audience measurement that can operate with less personal data

A website owner usually wants answers to fairly boring questions.

How many people visited?

Which pages were popular?

Did anyone click the new navigation link?

Did traffic come from search, a newsletter, or another site?

None of those questions automatically requires building a durable advertising identity for every visitor.

Measurement and profiling are different jobs

Analytics becomes much more invasive when the unit of analysis changes from what happened on this site to what this person does across many sites and devices.

A publisher can often measure page views, broad traffic sources, session counts, device classes, or conversion totals without trying to recognize the same person everywhere else on the web.

France’s data-protection authority, CNIL, provides a useful concrete model. Its guidance allows certain audience-measurement trackers to qualify for a consent exemption only under restrictive conditions: the purpose must stay limited to audience measurement or A/B testing, the data must not be cross-checked with unrelated customer files or visits to other sites, the tracker must remain scoped to a single publisher, IP addresses must be truncated, and tracker lifetimes are limited. See CNIL’s guidance on audience measurement.

That is not the only possible privacy-preserving design.

It is useful because it shows the engineering principle clearly: collect enough to answer the measurement question, but do not quietly turn analytics into an identity business.

Aggregate answers lose some detail

Collecting less has tradeoffs.

A system that refuses to create persistent user histories may be worse at answering questions such as:

  • Did the same person return six months later?
  • Which advertisements did this exact user see before subscribing?
  • How does one person’s behavior compare across several unrelated properties?

Those can be commercially useful questions.

They are also the questions that require more durable identity.

A publisher therefore has to separate what it genuinely needs from what is merely interesting because technology makes it possible.

More data is not automatically better measurement

Individual-level histories can create their own errors.

People clear cookies. Families share devices. One person uses several browsers. Privacy tools isolate identifiers. Automated traffic contaminates logs. Cross-device systems make probabilistic matches that may be wrong.

A giant profile can look precise while containing bad joins.

Aggregate measurement has uncertainty too, but at least the uncertainty is closer to the question being asked.

If the question is How many times was this article read?, a system does not necessarily need to answer Who else lives with this reader?

That distinction matters at the end of the Surveillance Economy section.

The choice is not between perfect analytics and total blindness.

There is a large middle ground where websites can measure their own performance without insisting on remembering everybody everywhere.

Posted on

The difference between synthetic supply and actual human consumption

The internet can contain a mountain of content nobody climbs.

Generative systems make that distinction increasingly important. Producing one million pages is now an engineering problem. Getting one million people to read them remains an attention problem.

Those are not the same market.

Publication counts measure supply

Suppose an automated network publishes 100,000 articles, comments, product pages, or social posts in a day.

That tells us something real about the supply side of the internet. The material exists. Servers store it. Search crawlers may fetch it. APIs may distribute it. Other bots may quote it.

But the publication count cannot tell us whether people consumed any of it.

Human consumption needs different evidence: unique human visitors, watch time, reading time, survey data, validated engagement, subscriptions, purchases, comments from identifiable people, or other measures connecting content to actual attention.

Even engagement counts require care. A 2024 paper on misinformation showed a broader version of this problem: observed engagement does not necessarily map cleanly to underlying consumer preference because producer strategies influence what users encounter and how they respond. See The distorting effects of producer strategies: Why engagement does not reveal consumer preferences for misinformation.

The lesson travels beyond misinformation. Supply, exposure, engagement, and preference are four different things.

Synthetic abundance can be mostly self-contained

Imagine 10,000 generated pages created for search engines. Crawlers request them. Monitoring systems check them. automated accounts repost links to them. Analytics registers traffic.

A graph of machine activity may look enormous before a human ever arrives.

That does not make the pages irrelevant. Some may eventually reach people. The point is that machine production can grow much faster than human attention.

This creates a strange version of Dead Internet Theory. The Web can become increasingly synthetic by volume without becoming equally synthetic in what people actually spend their time consuming.

The opposite can also happen: a relatively small amount of synthetic material can receive enormous human attention if recommendation systems amplify it.

The useful question is where the humans enter the chain

A careful study should identify which quantity it measured.

How many items were published? How many were delivered to feeds? How many impressions were generated? How many accounts interacted? How many of those accounts were likely human? How long did people actually engage?

Each step narrows the claim.

“Half of the available content is synthetic” would not mean “half of what people read is synthetic.” “Half of traffic is automated” would not mean “half of attention belongs to bots.”

The internet can manufacture supply almost without limit.

Human attention remains stubbornly finite.