Posted on

False positives when people are classified as bots

A bot detector does not look inside an account and find a tiny robot certificate.

It infers.

Systems may examine posting frequency, timing, repeated URLs, follower patterns, account age, client software, text similarity, network connections, or combinations of many signals. Those patterns can be useful. They can also describe real people.

Some humans post at strange hours. Some obsessively repeat links. Newsrooms schedule material. Fan accounts behave mechanically. Customer-service workers answer from templates. Activists coordinate campaigns. Power users can produce enough activity to look less human than an actual spam script trying very hard to look normal.

That is the false-positive problem.

A score is not ground truth

Automated classifiers are usually probabilistic or heuristic. A high score may mean “this account resembles examples classified as automated,” not “we proved software controls this account.”

The limits have been debated in the research literature. A 2022 paper, “Investigating the Validity of Botometer-based Social Bot Studies”, examined studies that used the popular Botometer system to estimate bot prevalence and argued that some methodologies produced serious false-positive problems. The authors manually inspected accounts labeled as bots in published work and challenged the reliability of using detector scores as direct population counts.

That paper is a critique, not proof that every bot detector is useless. Different tools, datasets, thresholds, and research designs can perform differently. Its value here is narrower: classification error is real enough that prevalence estimates need validation rather than blind trust in a score.

Platforms admit the same basic problem operationally. X’s help documentation on suspended accounts says most suspensions target spammy or fake behavior, but also explicitly acknowledges that real people’s accounts are sometimes suspended by mistake and provides an appeal process.

False positives distort more than one account

If a system incorrectly labels 5 percent of human accounts as automated, that can badly distort an estimate when the true bot population is small. The error becomes especially important when researchers turn individual classifications into sweeping claims such as “one third of the conversation was bots.”

The correct response is not to abandon automated detection. Large datasets often make manual classification impossible. The response is to report thresholds, uncertainty, validation methods, known error rates, and the exact population being sampled.

Manual inspection also has limits. Humans can misclassify sophisticated bots, satire accounts, coordinated teams, and users writing in unfamiliar languages. There is no magical human eyeball that solves everything.

Strange behavior is evidence, not identity

This distinction matters for Dead Internet Theory because a suspicious-looking account can easily become an anecdote supporting a much larger conclusion.

Maybe the account really is automated. Maybe it is a person using scheduling tools. Maybe it is a teenager posting fifty times in an hour. Maybe it is a customer-support employee working from macros. Maybe the account was compromised.

Behavioral signals can justify investigation. They do not automatically establish authorship.

The more dramatic the prevalence claim, the more important that becomes. If we are trying to measure synthetic humanity, real humans accidentally counted as machines are not statistical debris. They are exactly the classification error the study is supposed to control.