Posted on

Email open tracking and the uncertainty of its measurements

An email service can report that you opened a message without ever seeing your eyes.

Usually it saw an image request.

Email-open tracking commonly works by placing a tiny remotely hosted image inside an HTML message. Mailchimp describes its own implementation this way: when open tracking is enabled, it embeds a tiny invisible graphic in the email. When the graphic is requested, the service records an open. See Mailchimp’s explanation of open tracking.

That is a useful measurement.

It is not a perfect description of human behavior.

A loaded image is not the same thing as a read message

Consider the chain:

  1. The message arrives.
  2. The mail client loads remote content.
  3. The tracking image is requested.
  4. The sender records an open.

The fourth step is real.

The assumption is that step two happened because a person opened and read the message.

Sometimes it did.

Sometimes software loaded the image for privacy, caching, security scanning, previewing, or other automated reasons.

Apple’s Mail Privacy Protection is a major example. Apple says the feature hides a user’s IP address when remote email content is fetched, preventing senders from using that IP to determine location or connect it to other online activity. See Apple’s Mail Privacy Protection support documentation.

Mailchimp explicitly warns that bot or proxy activity such as Apple’s Mail Privacy Protection can falsely inflate open metrics. See Mailchimp’s open and click rate documentation.

False negatives happen too

The uncertainty runs in both directions.

A person can genuinely read an email while their client blocks remote images.

A plain-text reader may never request the tracking image. A privacy tool may strip it. A proxy may cache one copy and reuse it. Network failures can interrupt loading.

So a missing pixel request does not prove that nobody read the message either.

That leaves marketers with a measurement that is useful in aggregate but fuzzier at the individual level than the word open suggests.

Better language produces better conclusions

A tracking system may know:

The remote resource associated with this message was requested.

From that, it may infer:

The email was opened.

What it usually cannot prove is:

A specific human carefully read and understood the message at 9:14 AM.

That distinction matters because open data can drive automation. A recipient may be segmented, retargeted, scored as engaged, or sent a follow-up based on the signal.

The Surveillance Economy does not only collect data.

It builds decisions on top of measurements whose confidence is easy to forget.

Sometimes the machine knows that an image loaded.

Then the dashboard upgrades that fact into a human action.

Posted on

The difference between campaign reach and persuasive effect

A million impressions is not a million changed minds.

That sounds obvious.

Online influence research forgets it surprisingly often.

Platforms can measure views, impressions, followers, shares, clicks, and estimated reach. Those numbers describe exposure.

Persuasion is a different claim.

To show persuasion, a researcher needs evidence that attitudes, beliefs, choices, or behavior changed because of the campaign.

Reach is easier to count

Suppose a covert network produces 40,000 posts that appear in front of five million accounts.

Those are substantial distribution numbers.

They can tell us that the operation existed, produced material at scale, and reached an audience.

They cannot tell us how many people believed it.

Some users may have ignored it.

Some may have already agreed.

Some may have mocked it, blocked it, or forgotten it five seconds later.

A share can even come from a critic.

The counter still goes up.

A documented example shows the gap

A 2023 Nature Communications study linked longitudinal survey responses from U.S. Twitter users with their exposure to accounts associated with Russia’s Internet Research Agency during the 2016 U.S. election.

The researchers found that exposure was highly concentrated: one percent of users accounted for 70 percent of exposures. They also found that IRA content was outweighed by domestic political and news content in participants’ feeds. Most importantly, across the outcomes they examined, they found no evidence of a meaningful relationship between exposure to the Russian influence campaign and changes in attitudes, polarization, or voting behavior. See the study in Nature Communications.

That result does not prove the campaign had zero effect on every person.

It shows why reach statistics alone cannot establish persuasive effect.

Persuasion requires a different research design

Useful evidence can include:

  • randomized experiments,
  • before-and-after surveys,
  • longitudinal panel studies,
  • behavioral data tied to exposure,
  • credible comparison groups,
  • natural experiments with carefully stated limitations.

Even then, causation can be difficult because people choose what they read and whom they follow.

A campaign may disproportionately reach people already sympathetic to its message. High engagement can therefore reflect selection rather than persuasion.

Big numbers make seductive headlines

“Campaign reached 20 million users” is concrete.

“Campaign changed an unknown number of opinions under conditions we cannot fully identify” is less exciting.

The second sentence may be more accurate.

Manufactured Consensus research should therefore keep at least three measurements separate:

How much material was produced?

How many people encountered it?

What changed because they encountered it?

Those questions require different evidence.

A campaign can fail to persuade and still be worth studying.

It can waste attention, distort a discussion, create confusion, provoke journalists, or simply demonstrate an attempt at manipulation.

But an attempt to influence is not proof of successful influence.

Reach tells us how far the message traveled.

Persuasion asks what happened when it arrived.

Posted on

Regional platforms omitted from familiar maps of internet activity

The internet you know is partly a geographic accident.

Ask an American to name the major social platforms and the list will probably include YouTube, Facebook, Instagram, TikTok, Reddit, X, WhatsApp, and perhaps a few newer services.

That list can feel like a map of the internet.

It is not.

Regional platforms can be enormous while remaining peripheral to observers elsewhere. DataReportal’s Digital 2026: Japan report, drawing on LINE’s published business data, estimated that LINE had about 99 million monthly active users in Japan in late 2025. That was equivalent to more than 90% of the country’s internet users. See Digital 2026: Japan.

In South Korea, KakaoTalk occupied a similarly dominant position. DataReportal’s 2026 South Korea report, using Kakao’s earnings data, placed KakaoTalk at about 49.1 million monthly active users in the country. See Digital 2026: South Korea.

Those are not obscure side communities.

They are major parts of everyday online life.

Familiar platform lists create familiar blind spots

Suppose somebody tries to answer the question: Where are people talking online now?

If the investigation samples only globally familiar American platforms, the answer will already be biased before the first post is counted.

The same problem applies beyond social media.

Countries and regions can develop their own messaging apps, forums, video services, marketplaces, blogging systems, gaming communities, and news aggregators. Some become culturally dominant without becoming equally visible in English-language technology coverage.

That can make local internet activity look strangely absent to outsiders.

Regional evidence breaks universal claims

A strong claim about the entire internet needs evidence that spans more than one country’s favorite websites.

Researchers should ask which services dominate the population being studied, how those platforms are used, whether their content is public or private, and whether language differences make outside measurement difficult.

Even user counts require caution. A monthly active user is not the same thing as a public author, and messaging activity is not the same thing as searchable web publishing.

But regional platforms matter precisely because they complicate the story that humans simply stopped communicating online.

Some activity moved into places an English-speaking observer rarely visits.

Dead Internet Theory becomes much less persuasive when the map forgets entire countries.

Before declaring a city empty, make sure it is actually on the map.

Posted on

Measuring how many distinct domains a person encounters in a month

If the internet feels like the same fifteen websites over and over, that impression can be measured.

One simple experiment is to count how many distinct domains a person actually visits during a month.

Browsers already keep enough local history to make that possible. Firefox, for example, stores browsing history as part of a user’s profile data and can preserve it through backup and sync features. See Mozilla’s current documentation on Firefox browsing data and backups.

The interesting part is not extracting the data.

It is deciding what to count.

Define the unit before counting it

Does news.example.com count separately from example.com?

What about docs.example.com and shop.example.com?

A reasonable approach is to reduce URLs to their registrable domain, so multiple pages and subdomains belonging to the same site count once. The exact rule should be documented because different definitions can produce very different totals.

Then define a visit.

A page that opened automatically in the background should probably not count the same way as a page the user actively selected. Redirects, authentication domains, CDNs, payment processors, embedded trackers, and browser-internal pages can inflate the number without representing meaningful destinations.

A privacy-conscious analysis can stay entirely local. The browser history never needs to leave the machine. A script can extract only the hostname and date, discard full URLs and query strings, reduce hosts to domains, calculate totals, then delete the working copy.

One number is useful but incomplete

Suppose a person encountered 312 distinct domains in one month.

That sounds diverse until you learn that 80 percent of their page visits happened on five of them.

Or perhaps they visited only 90 domains but spent substantial time reading dozens of independent sites.

Useful companion measures include:

  • total distinct domains;
  • visits per domain;
  • percentage of activity concentrated in the top 5 or top 10;
  • number of domains visited only once;
  • how many were reached directly, from search, or through social links;
  • how many were independent sites versus major platforms.

Even then, the numbers do not measure quality.

Ten excellent specialist sites may be better than 1,000 random junk domains.

The value of the exercise is narrower.

Dead Internet Theory often describes a subjective sensation: the web seems smaller, more repetitive, more centralized.

Browsing history can turn part of that sensation into a testable question.

Maybe the web became smaller.

Maybe your route through it did.

Posted on

Shared accounts that complicate the idea of one user, one person

The internet encourages a convenient fiction: one account equals one person.

Sometimes it does. Often it does not.

A restaurant account may be handled by whoever is working that week. A newsroom account can be used by several editors. A company support profile may be staffed around the clock by different employees. Families share logins. Clubs, nonprofits, bands, open-source projects, and political organizations all maintain public identities that outlive whichever individual happens to be typing.

That makes account-level measurement a poor shortcut for counting people.

Platforms explicitly support multi-person identities

This is not an obscure edge case. Meta’s Facebook documentation says a Page can give multiple trusted people access to create content, answer messages, respond to comments, manage ads, and perform other tasks as the Page.

To the public, the Page is one identity.

Behind it may be five humans with different schedules, writing styles, locations, devices, and habits.

A behavior-based classifier watching only the public stream could interpret those changes in several ways. Posting may appear around the clock. Tone may shift abruptly. One person may write long conversational replies while another mostly posts links. A third may schedule promotional messages in advance.

The account looks inconsistent because it is not one person.

Shared identity creates strange measurement artifacts

Suppose a study counts 10,000 accounts and treats them as 10,000 users. Some accounts belong to one person. Some people own several accounts. Some accounts are operated by teams. Some are bots. Some are organizations using automation plus human staff.

The total number of actual people cannot be recovered by simply counting rows in the account table.

Shared accounts can also confuse attempts to detect automation. A corporate profile may publish on a precise schedule because software queues posts, then switch to unmistakably human conversation when an employee replies. One analyst calls it a bot. Another calls it human. Both may be looking at real parts of the same operating model.

This is one reason the phrase “fake population” needs careful handling. The account may not correspond to a single human, but that does not make it fake. A library, newspaper, game studio, or community project is a real social actor even if no individual human maps one-to-one onto its username.

Identity online is often organizational

The better question is what kind of entity the account represents and how its activity is produced.

Is it one person? A rotating staff? A human using scheduling tools? A bot controlled by a team? An organization publishing official statements? A compromised profile? Those categories are more useful than forcing everything into “human account” or “bot account.”

Dead Internet Theory is strongest when it notices that online identity is becoming synthetic, automated, and difficult to verify. But difficulty does not justify assuming every non-personal account is artificial.

Sometimes one username hides a machine.

Sometimes it hides the night shift.

Posted on

Helpful automation as a confounder in bot prevalence estimates

The word bot often arrives carrying guilt before the evidence does.

That is understandable. Spam bots, credential-stuffing tools, fake engagement networks, and automated scams are real. But the same technical category also contains search crawlers, uptime monitors, feed fetchers, archive crawlers, API clients, accessibility services, and scripts doing routine work for actual people.

If a study counts automation without separating those roles, “bot prevalence” can become a much scarier number than “deceptive automation prevalence.”

Some bots are infrastructure

Google openly documents that Googlebot automatically requests pages so Google Search can discover and index them. Google also describes crawling more generally as automated software used to discover and understand pages across the web.

Those requests are machine-generated. They belong in a traffic report about automation. But their purpose is not to imitate a human participant.

The same is true for a service checking a site every minute to see whether it is offline. An RSS reader fetching a feed on behalf of a subscriber is automated. So is a script downloading public weather data every hour. A preservation crawler saving a website before it disappears can be extremely aggressive compared with ordinary browsing and still be doing something useful.

This creates a measurement problem.

Intent and function matter

A security provider may classify automated traffic according to whether it appears benign or malicious. A social-network researcher might instead care whether an account presents itself as a person. A publisher studying comment spam cares about automated posting. A server administrator may care only about load.

All four can use the word “bot” while measuring different things.

Imperva’s recent bot reports illustrate the scale issue. Its 2026 report says automated requests accounted for more than half of observed web traffic in 2025. The report also distinguishes malicious automation from the broader automated total.

That distinction is essential. The headline number is not a count of fake humans.

Useful automation can distort simple prevalence claims

Imagine a small technical website. Human readers visit it 10,000 times in a month. Search crawlers, monitoring systems, AI retrieval tools, and archive crawlers together make 20,000 requests.

A traffic-level measurement could correctly report that most requests were automated.

A reader-level claim that “most of the site’s audience was fake” would be unsupported.

This is one reason Dead Internet Theory needs narrower categories than human versus bot. Some automation is adversarial. Some is deceptive. Some is commercial. Some is maintenance. Some is preservation. Some is a human using a tool to avoid doing repetitive work by hand.

The interesting question is not merely how much automation exists.

It is what the automation is doing, who operates it, whether it represents itself honestly, and whether it affects what people believe they are interacting with.

Without those distinctions, a search crawler and a fake grassroots account wind up in the same bucket. Technically they are both automated. Socially, they are not remotely the same phenomenon.

Posted on

Page requests, accounts, and people as incompatible measures of internet population

How many “users” are on the internet?

Before answering, somebody has to define user.

A server log counts requests. A social network counts accounts. A survey counts people. Those numbers can all be correct while describing completely different populations.

That distinction matters whenever somebody tries to estimate how much online activity is human, synthetic, or abandoned.

One person can become many records

A single person can own several social accounts, use multiple browsers and devices, run scripts, operate a business page, and make thousands of page requests in one day.

The reverse also happens. One account may represent a company, family, newsroom, project, or team rather than one individual. A bot may control many accounts. An organization may generate millions of automated requests without possessing anything resembling millions of “users.”

Pew Research Center described this problem directly when discussing social-media data: researchers often struggle to translate accounts into people because an account may be a person, a duplicate account, or an automated account. Its discussion of social-media research methods also notes a broader sampling problem: platforms that are easiest to study are not necessarily representative of everyone online.

Even basic profile counts can multiply one human identity. Pew found as far back as 2009 that more than half of adult social-network users in its survey had two or more profiles, usually spread across different services. The exact percentage is historical, but the measurement problem never went away.

Requests are even farther from people

Network traffic is another layer removed.

One page load can trigger dozens of requests for images, JavaScript, fonts, APIs, ads, and analytics. A crawler can generate millions of requests while representing one automated system. A mobile app may poll an API repeatedly while its owner does nothing.

So a statement like “bots generated more requests than humans” cannot be converted into “there are more bots than people.” It is comparing events to actors.

Likewise, “there are 500 million accounts” does not establish 500 million distinct people. Accounts can be duplicated, dormant, shared, organizational, compromised, automated, or abandoned.

A useful estimate names the unit

Good population claims are boringly specific.

Instead of saying “half the internet is bots,” a careful study might say that a defined percentage of HTTP requests observed by a security provider were classified as automated during a particular year. Or that a percentage of sampled accounts on one platform showed behavior consistent with automation. Or that survey respondents reported using a service.

Those are different statements because they answer different questions.

Dead Internet Theory discussions often become confusing at exactly this point. A number describing traffic gets compared with a number describing accounts, then both get interpreted as a number describing people.

The arithmetic may be flawless. The categories are not.

Before asking whether the internet is populated by humans or machines, the first useful question is simpler: What exactly are we counting?