The easiest part of the internet to study is not necessarily the most representative part.
Researchers often work with whatever data a platform exposes. That can be completely reasonable. The mistake comes later, when a finding about one service, one API, or one slice of users quietly grows into a claim about “the internet.”
A study of public Twitter posts never measured private Facebook groups, email, Reddit, Discord, independent forums, YouTube comments, personal blogs, game chats, Nostr, or millions of websites that do not expose comparable data.
The sample may be huge and still be narrow.
Platform choice is already a filter
Pew Research Center has discussed this problem directly. In a 2018 Q&A about social-media research, its researchers noted that Twitter attracted academic study partly because much of its data was public and available, while Facebook had a larger population but exposed less of its activity to researchers.
That creates a selection effect before anyone runs a model.
Researchers naturally gravitate toward data they can collect. If bot behavior is unusually visible on a platform with an open API, that platform may become overrepresented in the literature compared with more closed services.
Then there is sampling inside the platform itself.
A 2014 paper titled “When is it Biased? Assessing the Representativeness of Twitter’s Streaming API” examined Twitter’s then-available streaming sample and found periods where sampled hashtag trends diverged from the platform’s fuller activity. The important lesson is broader than that old API: access mechanisms can introduce their own bias.
Ten million posts can still be the wrong ten million
Large datasets feel authoritative because the numbers are enormous.
But sample size does not repair a bad sampling frame.
Imagine collecting 20 million comments from a platform popular with cryptocurrency traders, marketers, and automated alert accounts. You may estimate automation very accurately for that dataset. It would be reckless to assume the same prevalence on a private parenting forum, a university mailing list, or a small hobby Discord.
Different communities have different incentives for automation. Finance attracts trading and price bots. Gaming communities may contain moderation and stat bots. News platforms attract link-posting systems. Small private groups may have almost none.
Good claims keep their borders
A careful paper says what it sampled: which platform, which dates, which languages, which account types, which API, and what was excluded.
It then limits the conclusion accordingly.
“This detector classified 18 percent of sampled public accounts on Platform X during Period Y as likely automated” is a claim someone can inspect.
“Eighteen percent of people on the internet are bots” is a different claim entirely.
This matters for Dead Internet Theory because the theory is global by nature. It talks about the internet as a whole while much of the available evidence comes from a handful of measurable platforms.
Those studies can reveal real synthetic populations.
They just cannot turn one city block into a census of the planet.
