Posted on

How sampled platforms distort estimates of the whole internet

The easiest part of the internet to study is not necessarily the most representative part.

Researchers often work with whatever data a platform exposes. That can be completely reasonable. The mistake comes later, when a finding about one service, one API, or one slice of users quietly grows into a claim about “the internet.”

A study of public Twitter posts never measured private Facebook groups, email, Reddit, Discord, independent forums, YouTube comments, personal blogs, game chats, Nostr, or millions of websites that do not expose comparable data.

The sample may be huge and still be narrow.

Platform choice is already a filter

Pew Research Center has discussed this problem directly. In a 2018 Q&A about social-media research, its researchers noted that Twitter attracted academic study partly because much of its data was public and available, while Facebook had a larger population but exposed less of its activity to researchers.

That creates a selection effect before anyone runs a model.

Researchers naturally gravitate toward data they can collect. If bot behavior is unusually visible on a platform with an open API, that platform may become overrepresented in the literature compared with more closed services.

Then there is sampling inside the platform itself.

A 2014 paper titled “When is it Biased? Assessing the Representativeness of Twitter’s Streaming API” examined Twitter’s then-available streaming sample and found periods where sampled hashtag trends diverged from the platform’s fuller activity. The important lesson is broader than that old API: access mechanisms can introduce their own bias.

Ten million posts can still be the wrong ten million

Large datasets feel authoritative because the numbers are enormous.

But sample size does not repair a bad sampling frame.

Imagine collecting 20 million comments from a platform popular with cryptocurrency traders, marketers, and automated alert accounts. You may estimate automation very accurately for that dataset. It would be reckless to assume the same prevalence on a private parenting forum, a university mailing list, or a small hobby Discord.

Different communities have different incentives for automation. Finance attracts trading and price bots. Gaming communities may contain moderation and stat bots. News platforms attract link-posting systems. Small private groups may have almost none.

Good claims keep their borders

A careful paper says what it sampled: which platform, which dates, which languages, which account types, which API, and what was excluded.

It then limits the conclusion accordingly.

“This detector classified 18 percent of sampled public accounts on Platform X during Period Y as likely automated” is a claim someone can inspect.

“Eighteen percent of people on the internet are bots” is a different claim entirely.

This matters for Dead Internet Theory because the theory is global by nature. It talks about the internet as a whole while much of the available evidence comes from a handful of measurable platforms.

Those studies can reveal real synthetic populations.

They just cannot turn one city block into a census of the planet.

Posted on

Measuring web loss without confusing absence from an archive with nonexistence

There is a tempting but dangerous equation in web archaeology:

No Wayback capture = the page did not exist.

That conclusion feels reasonable because web archives are enormous. It is also wrong.

An archive can fail to contain a page for many reasons that have nothing to do with whether the page once existed. A crawler may never have discovered it. The site may have blocked crawling. The page may have required a form submission or login. The resource may have been streaming media, a database result, or dynamically generated content that the crawler could not preserve correctly.

Archives are samples, not omniscient recordings

The Library of Congress describes web archiving as a process that begins with selected seed URLs and follows links according to collection scope. Its own technical guidance acknowledges that current tools cannot capture all web content, specifically naming difficult categories such as streaming media, deep-web content, databases, and multimedia-rich experiences.

That means archive coverage is shaped by selection and technology before researchers ever begin measuring loss.

A direct measurement of web decay shows why definitions matter. Pew Research Center’s 2024 study of disappearing online content sampled pages from Common Crawl and then tested whether those URLs were still accessible on the live web. It found 38 percent of sampled pages from 2013 inaccessible by 2023. That is a measurement of live accessibility, not a claim that 38 percent had vanished from every archive and mirror.

An absent or poorly timed archive capture is therefore evidence about the archive’s holdings. It is not automatically evidence about the historical Web.

Define what “lost” means before counting it

A study of web loss needs a starting population. One useful method is to begin with URLs known to have existed because they appear in an old directory, published bibliography, crawl dataset, sitemap, software package, or contemporaneous list.

Researchers can then ask separate questions:

Is the URL live today? Does it redirect? Does the same document survive at another URL? Is there a capture in one archive? In several archives? Does an independent copy survive in a PDF, mirror, repository, CD-ROM, or email attachment?

Those questions produce different kinds of loss.

A dead original URL is link loss. A missing page at its original domain may still have a complete archived copy. A page absent from the Wayback Machine may survive in another web archive. A document may have moved without a redirect. And a genuinely vanished work may leave only quotations or screenshots.

Collapsing all of those conditions into “gone” produces a dramatic number and a weak study. Pew’s 2024 study, for example, found that 38 percent of its sampled 2013 pages were no longer accessible a decade later. When the Internet Archive later rechecked Pew’s broader dataset against the Wayback Machine, it showed that many URLs dead on the live web still had archived copies, while also warning that it had not checked every smaller web archive. “Dead live URL,” “absent from Wayback,” and “vanished everywhere” are different measurements.

Careful claims are narrower and more useful

The strongest language matches the evidence: “No capture was found in the archives searched” rather than “the page was never archived anywhere.” “The original URL no longer resolves” rather than “the document was deleted.” “The page was observed live in 2012 and dead in 2017” rather than inventing an exact disappearance date between those observations.

This restraint matters to Dead Internet Theory because a real phenomenon—massive web decay—does not become stronger when archive gaps are counted as proof of disappearance.

The old Web has lost an extraordinary amount of material. Measuring how much requires resisting the urge to turn every blank spot in the archive into a gravestone.