Posted on

Human audits of randomly sampled public discussions

If you go looking only for creepy bot-like conversations, you will find a creepy bot-like internet.

That is not a measurement.

One way to test claims about synthetic conversation is much less dramatic: choose discussions according to a sampling rule decided in advance, then have human reviewers inspect them without selecting only the suspicious ones.

The boring part is the useful part.

Start with a sample that did not already know the answer

A reasonable audit might define a platform, date range, language, discussion type, and method for randomly selecting threads or comments. The sampling process should include quiet, ordinary, messy conversations as well as obvious spam.

Otherwise the researcher is measuring the contents of a folder labeled “weird stuff I noticed,” not the population of the platform.

Human reviewers can then classify observable characteristics: obvious commercial spam, disclosed automation, copied text, coherent human conversation, unknown or ambiguous authorship, and so on.

The word unknown is important.

A reviewer cannot reliably prove that a polished comment came from a human simply because it sounds natural. Nor can repetitive language alone prove automation.

Research on human recognition of social bots illustrates the problem. A 2024 experimental study asked people to identify bots on the VKontakte social network and found that human labeling itself can be difficult enough to undermine the idea of perfect “ground truth.” See Experimental Evaluation: Can Humans Recognise Social Media Bots?.

Disagreement is data

Suppose three reviewers inspect the same account. One calls it automated, one calls it human, and one marks it uncertain.

Throwing away the disagreement would make the final number look cleaner while hiding the most important fact: the evidence was ambiguous.

A useful audit should report how reviewers were instructed, whether they worked independently, how often they agreed, which categories caused disagreement, and how uncertain cases affected the final estimate.

Researchers can use statistical measures of inter-rater agreement, but the plain-language interpretation matters too. “Reviewers agreed on 92 percent of cases” tells a different story from “half the accounts could not be classified confidently.”

A sample answers a bounded question

Even a careful human audit does not establish how much of “the internet” is synthetic.

It can estimate what appeared in a defined sample from a defined platform under defined conditions. Different languages, communities, dates, recommendation systems, and access states may produce different populations.

That limitation is not a weakness. It is what makes the claim testable.

Dead Internet Theory becomes harder to evaluate when every strange screenshot is treated as representative and every normal conversation is dismissed as an exception.

Random sampling reverses that habit.

Do not ask the internet to show you something spooky.

Ask a sample what is actually there, and leave room for the honest answer: sometimes we cannot tell.