Posted on

What evidence could falsify a claim that most online conversation is synthetic

“Most online conversation is synthetic” sounds testable until somebody asks what result would make you stop believing it.

That question is not rhetorical. It is the difference between a claim that can be investigated and a suspicion that absorbs every possible outcome.

A useful version might be: More than 50 percent of public conversational posts on Platform X during Month Y were generated and published without meaningful human authorship.

Now the terms have edges. There is a platform, a time period, a unit of analysis, a threshold, and a definition of synthetic authorship.

A strong claim needs a possible losing condition

Karl Popper’s idea of falsifiability is often oversimplified, but the basic principle is useful here: a claim should be capable of conflicting with possible observations. The Stanford Encyclopedia of Philosophy’s discussion of falsifiability explains that a statement becomes empirically testable when conceivable observations could count against it.

For the example above, a carefully designed study could draw a representative sample of public posts from the defined platform and period, classify them using several forms of evidence, validate the classification against known accounts or manual review, and estimate the synthetic share with uncertainty bounds.

If the best-supported estimate came out at 12 percent, with even the upper confidence limit nowhere near 50 percent, that would be strong evidence against the claim for that defined population.

It would not prove that all other platforms are mostly human. It would falsify the claim as stated.

Definitions have to stay put after the result

This is where broad internet theories can become slippery.

Suppose a study finds little evidence of synthetic comments. Someone can reply that the bots are actually in likes. A study checks likes; the claim moves to search results. Researchers examine search; the claim shifts to private messages. Evidence about one period is dismissed because the “real takeover” happened later.

Any one of those new questions might be worth studying. But continually changing the claim after contrary evidence prevents the original claim from ever losing.

That is not stronger skepticism. It is a measurement problem wearing armor.

Anecdotes can motivate a study, not settle it

A thread containing fifty obvious chatbot replies proves that fifty obvious chatbot replies existed. A strange profile with a generated face proves something interesting about that profile. A spam wave can show that one community was flooded.

None of those observations establishes the share of synthetic conversation across a platform, much less across the entire internet.

Prevalence is a population question. It needs sampling, definitions, error estimates, and scope.

The same standard cuts both ways. A handful of vivid human conversations cannot disprove large-scale synthetic activity either. Human examples are anecdotes too.

Dead Internet Theory becomes more interesting when its claims are allowed to fail.

Define what “most,” “online conversation,” and “synthetic” mean. State where and when the claim applies. Decide in advance what evidence would count against it.

Then go looking.

If no conceivable result could change the conclusion, the conclusion was never being measured in the first place.

Posted on

How sampled platforms distort estimates of the whole internet

The easiest part of the internet to study is not necessarily the most representative part.

Researchers often work with whatever data a platform exposes. That can be completely reasonable. The mistake comes later, when a finding about one service, one API, or one slice of users quietly grows into a claim about “the internet.”

A study of public Twitter posts never measured private Facebook groups, email, Reddit, Discord, independent forums, YouTube comments, personal blogs, game chats, Nostr, or millions of websites that do not expose comparable data.

The sample may be huge and still be narrow.

Platform choice is already a filter

Pew Research Center has discussed this problem directly. In a 2018 Q&A about social-media research, its researchers noted that Twitter attracted academic study partly because much of its data was public and available, while Facebook had a larger population but exposed less of its activity to researchers.

That creates a selection effect before anyone runs a model.

Researchers naturally gravitate toward data they can collect. If bot behavior is unusually visible on a platform with an open API, that platform may become overrepresented in the literature compared with more closed services.

Then there is sampling inside the platform itself.

A 2014 paper titled “When is it Biased? Assessing the Representativeness of Twitter’s Streaming API” examined Twitter’s then-available streaming sample and found periods where sampled hashtag trends diverged from the platform’s fuller activity. The important lesson is broader than that old API: access mechanisms can introduce their own bias.

Ten million posts can still be the wrong ten million

Large datasets feel authoritative because the numbers are enormous.

But sample size does not repair a bad sampling frame.

Imagine collecting 20 million comments from a platform popular with cryptocurrency traders, marketers, and automated alert accounts. You may estimate automation very accurately for that dataset. It would be reckless to assume the same prevalence on a private parenting forum, a university mailing list, or a small hobby Discord.

Different communities have different incentives for automation. Finance attracts trading and price bots. Gaming communities may contain moderation and stat bots. News platforms attract link-posting systems. Small private groups may have almost none.

Good claims keep their borders

A careful paper says what it sampled: which platform, which dates, which languages, which account types, which API, and what was excluded.

It then limits the conclusion accordingly.

“This detector classified 18 percent of sampled public accounts on Platform X during Period Y as likely automated” is a claim someone can inspect.

“Eighteen percent of people on the internet are bots” is a different claim entirely.

This matters for Dead Internet Theory because the theory is global by nature. It talks about the internet as a whole while much of the available evidence comes from a handful of measurable platforms.

Those studies can reveal real synthetic populations.

They just cannot turn one city block into a census of the planet.

Posted on

Human-operated automation and the limits of a human-or-bot binary

A surprising amount of online activity lives in the awkward middle between “human” and “bot.”

A person writes ten posts on Sunday and schedules them for the week. A script watches an RSS feed and publishes links automatically, but the owner personally answers every reply. A customer-service system drafts responses that employees approve. A social account reposts material by rule until a human steps in during breaking news.

Which of those is the bot?

The binary starts failing as soon as humans use automation as a tool rather than surrendering the account completely.

Automation comes in degrees

At one end is ordinary manual use: a person decides what to say and presses the button each time.

Then come scheduling tools, templates, macros, automatic cross-posting, feed-driven publishing, generated drafts, moderation assistants, and scripts that perform narrow actions. Human control can remain substantial even while a large percentage of visible events are technically generated by software.

At the other end are autonomous systems that select content, generate text, decide when to publish, interact with users, and continue operating without routine human approval.

Those are different arrangements, but a detector watching timestamps and posting patterns may compress them into the same label.

Researchers have used the word cyborg for this middle category. A 2024 study, “Cyborgs for strategic communication on social media”, analyzed more than 3.1 million Twitter users from datasets related to the 2020 coronavirus pandemic and U.S. election. The authors described cyborg accounts as hybrids combining automated scripts with manual participation and identified accounts whose bot/human classifications changed across time windows.

The important point is not the exact number of cyborgs in that study. It is that mixed operation is measurable enough to deserve its own category.

A human may operate something that behaves like a bot

Consider a one-person news service. Software monitors twenty feeds, posts headlines automatically, and alerts the owner when someone replies. The owner then reads the conversation and responds personally.

Calling the account fully human ignores the automated publishing. Calling it fully bot ignores the human editorial decisions and conversation.

The same ambiguity appears with AI assistance. If a person asks a model for a draft, rewrites half of it, checks the facts, and publishes under their own name, the final text has machine involvement without being autonomously authored. If the system generates and publishes 5,000 posts without review, that is a very different production process.

Better labels describe control

Useful analysis can ask who selects the objective, who creates the content, who decides when to publish, whether a human reviews outputs, and whether the system can act independently after launch.

Those questions produce a spectrum of control rather than a theatrical courtroom verdict: HUMAN or BOT.

That matters for Dead Internet Theory because synthetic activity does not need to replace humans completely. Humans can multiply themselves through automation. One operator can create the visible output that once required a staff. A team can supervise hundreds of automated identities. A normal user can automate repetitive tasks without attempting to deceive anybody.

The future internet may be difficult to count precisely because the important unit is no longer “person or machine.”

Increasingly, it is person with machine.

Posted on

Shared accounts that complicate the idea of one user, one person

The internet encourages a convenient fiction: one account equals one person.

Sometimes it does. Often it does not.

A restaurant account may be handled by whoever is working that week. A newsroom account can be used by several editors. A company support profile may be staffed around the clock by different employees. Families share logins. Clubs, nonprofits, bands, open-source projects, and political organizations all maintain public identities that outlive whichever individual happens to be typing.

That makes account-level measurement a poor shortcut for counting people.

Platforms explicitly support multi-person identities

This is not an obscure edge case. Meta’s Facebook documentation says a Page can give multiple trusted people access to create content, answer messages, respond to comments, manage ads, and perform other tasks as the Page.

To the public, the Page is one identity.

Behind it may be five humans with different schedules, writing styles, locations, devices, and habits.

A behavior-based classifier watching only the public stream could interpret those changes in several ways. Posting may appear around the clock. Tone may shift abruptly. One person may write long conversational replies while another mostly posts links. A third may schedule promotional messages in advance.

The account looks inconsistent because it is not one person.

Shared identity creates strange measurement artifacts

Suppose a study counts 10,000 accounts and treats them as 10,000 users. Some accounts belong to one person. Some people own several accounts. Some accounts are operated by teams. Some are bots. Some are organizations using automation plus human staff.

The total number of actual people cannot be recovered by simply counting rows in the account table.

Shared accounts can also confuse attempts to detect automation. A corporate profile may publish on a precise schedule because software queues posts, then switch to unmistakably human conversation when an employee replies. One analyst calls it a bot. Another calls it human. Both may be looking at real parts of the same operating model.

This is one reason the phrase “fake population” needs careful handling. The account may not correspond to a single human, but that does not make it fake. A library, newspaper, game studio, or community project is a real social actor even if no individual human maps one-to-one onto its username.

Identity online is often organizational

The better question is what kind of entity the account represents and how its activity is produced.

Is it one person? A rotating staff? A human using scheduling tools? A bot controlled by a team? An organization publishing official statements? A compromised profile? Those categories are more useful than forcing everything into “human account” or “bot account.”

Dead Internet Theory is strongest when it notices that online identity is becoming synthetic, automated, and difficult to verify. But difficulty does not justify assuming every non-personal account is artificial.

Sometimes one username hides a machine.

Sometimes it hides the night shift.

Posted on

False positives when people are classified as bots

A bot detector does not look inside an account and find a tiny robot certificate.

It infers.

Systems may examine posting frequency, timing, repeated URLs, follower patterns, account age, client software, text similarity, network connections, or combinations of many signals. Those patterns can be useful. They can also describe real people.

Some humans post at strange hours. Some obsessively repeat links. Newsrooms schedule material. Fan accounts behave mechanically. Customer-service workers answer from templates. Activists coordinate campaigns. Power users can produce enough activity to look less human than an actual spam script trying very hard to look normal.

That is the false-positive problem.

A score is not ground truth

Automated classifiers are usually probabilistic or heuristic. A high score may mean “this account resembles examples classified as automated,” not “we proved software controls this account.”

The limits have been debated in the research literature. A 2022 paper, “Investigating the Validity of Botometer-based Social Bot Studies”, examined studies that used the popular Botometer system to estimate bot prevalence and argued that some methodologies produced serious false-positive problems. The authors manually inspected accounts labeled as bots in published work and challenged the reliability of using detector scores as direct population counts.

That paper is a critique, not proof that every bot detector is useless. Different tools, datasets, thresholds, and research designs can perform differently. Its value here is narrower: classification error is real enough that prevalence estimates need validation rather than blind trust in a score.

Platforms admit the same basic problem operationally. X’s help documentation on suspended accounts says most suspensions target spammy or fake behavior, but also explicitly acknowledges that real people’s accounts are sometimes suspended by mistake and provides an appeal process.

False positives distort more than one account

If a system incorrectly labels 5 percent of human accounts as automated, that can badly distort an estimate when the true bot population is small. The error becomes especially important when researchers turn individual classifications into sweeping claims such as “one third of the conversation was bots.”

The correct response is not to abandon automated detection. Large datasets often make manual classification impossible. The response is to report thresholds, uncertainty, validation methods, known error rates, and the exact population being sampled.

Manual inspection also has limits. Humans can misclassify sophisticated bots, satire accounts, coordinated teams, and users writing in unfamiliar languages. There is no magical human eyeball that solves everything.

Strange behavior is evidence, not identity

This distinction matters for Dead Internet Theory because a suspicious-looking account can easily become an anecdote supporting a much larger conclusion.

Maybe the account really is automated. Maybe it is a person using scheduling tools. Maybe it is a teenager posting fifty times in an hour. Maybe it is a customer-support employee working from macros. Maybe the account was compromised.

Behavioral signals can justify investigation. They do not automatically establish authorship.

The more dramatic the prevalence claim, the more important that becomes. If we are trying to measure synthetic humanity, real humans accidentally counted as machines are not statistical debris. They are exactly the classification error the study is supposed to control.

Posted on

Helpful automation as a confounder in bot prevalence estimates

The word bot often arrives carrying guilt before the evidence does.

That is understandable. Spam bots, credential-stuffing tools, fake engagement networks, and automated scams are real. But the same technical category also contains search crawlers, uptime monitors, feed fetchers, archive crawlers, API clients, accessibility services, and scripts doing routine work for actual people.

If a study counts automation without separating those roles, “bot prevalence” can become a much scarier number than “deceptive automation prevalence.”

Some bots are infrastructure

Google openly documents that Googlebot automatically requests pages so Google Search can discover and index them. Google also describes crawling more generally as automated software used to discover and understand pages across the web.

Those requests are machine-generated. They belong in a traffic report about automation. But their purpose is not to imitate a human participant.

The same is true for a service checking a site every minute to see whether it is offline. An RSS reader fetching a feed on behalf of a subscriber is automated. So is a script downloading public weather data every hour. A preservation crawler saving a website before it disappears can be extremely aggressive compared with ordinary browsing and still be doing something useful.

This creates a measurement problem.

Intent and function matter

A security provider may classify automated traffic according to whether it appears benign or malicious. A social-network researcher might instead care whether an account presents itself as a person. A publisher studying comment spam cares about automated posting. A server administrator may care only about load.

All four can use the word “bot” while measuring different things.

Imperva’s recent bot reports illustrate the scale issue. Its 2026 report says automated requests accounted for more than half of observed web traffic in 2025. The report also distinguishes malicious automation from the broader automated total.

That distinction is essential. The headline number is not a count of fake humans.

Useful automation can distort simple prevalence claims

Imagine a small technical website. Human readers visit it 10,000 times in a month. Search crawlers, monitoring systems, AI retrieval tools, and archive crawlers together make 20,000 requests.

A traffic-level measurement could correctly report that most requests were automated.

A reader-level claim that “most of the site’s audience was fake” would be unsupported.

This is one reason Dead Internet Theory needs narrower categories than human versus bot. Some automation is adversarial. Some is deceptive. Some is commercial. Some is maintenance. Some is preservation. Some is a human using a tool to avoid doing repetitive work by hand.

The interesting question is not merely how much automation exists.

It is what the automation is doing, who operates it, whether it represents itself honestly, and whether it affects what people believe they are interacting with.

Without those distinctions, a search crawler and a fake grassroots account wind up in the same bucket. Technically they are both automated. Socially, they are not remotely the same phenomenon.

Posted on

Page requests, accounts, and people as incompatible measures of internet population

How many “users” are on the internet?

Before answering, somebody has to define user.

A server log counts requests. A social network counts accounts. A survey counts people. Those numbers can all be correct while describing completely different populations.

That distinction matters whenever somebody tries to estimate how much online activity is human, synthetic, or abandoned.

One person can become many records

A single person can own several social accounts, use multiple browsers and devices, run scripts, operate a business page, and make thousands of page requests in one day.

The reverse also happens. One account may represent a company, family, newsroom, project, or team rather than one individual. A bot may control many accounts. An organization may generate millions of automated requests without possessing anything resembling millions of “users.”

Pew Research Center described this problem directly when discussing social-media data: researchers often struggle to translate accounts into people because an account may be a person, a duplicate account, or an automated account. Its discussion of social-media research methods also notes a broader sampling problem: platforms that are easiest to study are not necessarily representative of everyone online.

Even basic profile counts can multiply one human identity. Pew found as far back as 2009 that more than half of adult social-network users in its survey had two or more profiles, usually spread across different services. The exact percentage is historical, but the measurement problem never went away.

Requests are even farther from people

Network traffic is another layer removed.

One page load can trigger dozens of requests for images, JavaScript, fonts, APIs, ads, and analytics. A crawler can generate millions of requests while representing one automated system. A mobile app may poll an API repeatedly while its owner does nothing.

So a statement like “bots generated more requests than humans” cannot be converted into “there are more bots than people.” It is comparing events to actors.

Likewise, “there are 500 million accounts” does not establish 500 million distinct people. Accounts can be duplicated, dormant, shared, organizational, compromised, automated, or abandoned.

A useful estimate names the unit

Good population claims are boringly specific.

Instead of saying “half the internet is bots,” a careful study might say that a defined percentage of HTTP requests observed by a security provider were classified as automated during a particular year. Or that a percentage of sampled accounts on one platform showed behavior consistent with automation. Or that survey respondents reported using a service.

Those are different statements because they answer different questions.

Dead Internet Theory discussions often become confusing at exactly this point. A number describing traffic gets compared with a number describing accounts, then both get interpreted as a number describing people.

The arithmetic may be flawless. The categories are not.

Before asking whether the internet is populated by humans or machines, the first useful question is simpler: What exactly are we counting?

Posted on

Bot traffic versus bot-authored public conversation

A statistic such as “53 percent of web traffic is automated” sounds like it should tell us how much of the internet is made by bots.

It does not.

It tells us something important, but narrower: machines are responsible for a large share of requests reaching websites and applications. That is a measurement of traffic, not a census of who wrote the visible conversation.

Imperva’s 2026 Bad Bot Report says automated systems generated more than 53 percent of observed web traffic in 2025. That is a remarkable number. It is also easy to misuse.

A request is not a comment

Web traffic measurements usually count HTTP requests. A crawler fetching an article creates a request. A monitoring service checking whether a page is alive creates a request. A scraper downloading prices creates requests. An API client retrieving data creates requests. An attacker testing credentials creates many requests very quickly.

None of those activities necessarily writes anything that another person will read.

Google describes Googlebot as the crawler used by Google Search. It visits pages automatically so they can be indexed. Those visits are bot traffic in the literal sense, but Googlebot is not sitting in the comments pretending to be your uncle.

The mismatch also works in the other direction. One automated account can publish thousands of posts while producing only a modest fraction of a platform’s total network traffic. Meanwhile a single human opening a modern page can trigger requests for HTML, images, scripts, fonts, analytics, advertisements, APIs, and background updates.

The units simply do not map cleanly.

Measuring synthetic conversation requires conversation data

If the question is “What share of public discussion is machine-authored?” the sample has to contain public discussion: posts, replies, comments, messages, or other defined conversational units.

Researchers then need a defensible method for deciding which items are automated. That may involve account behavior, posting tools, content provenance, network patterns, manual review, disclosures, or known ground-truth accounts. Each method has uncertainty and false positives, but at least it measures the thing being claimed.

A traffic report cannot substitute for that work.

This distinction matters for Dead Internet Theory because traffic statistics are often used as if they prove a stronger proposition: if most requests are automated, then most apparent human activity must also be automated. That conclusion does not follow.

The automated web is already enormous. Search crawlers, security scanners, monitoring systems, AI agents, scrapers, spam tools, and malicious bots really do talk to servers all day without a person clicking anything.

That fact is worth studying on its own.

But a machine requesting a webpage and a machine impersonating a person in a conversation are two different events. Counting the first cannot tell us how common the second is.

Posted on

The origins of Dead Internet Theory and the claims bundled into it

Dead Internet Theory is easier to discuss once it stops being treated as one claim.

The name suggests a single proposition: the internet is “dead.” In practice, the theory bundles together several very different ideas. Some are measurable. Some are plausible but difficult to quantify. Others require evidence of coordination or intent that traffic statistics cannot provide.

Separating them is more interesting than either swallowing the whole theory or dismissing everything attached to it.

The theory emerged from a feeling before it became a label

The exact prehistory is fuzzy because the idea circulated through anonymous and semi-anonymous online communities. Earlier discussions on places such as Wizardchan and 4chan are repeatedly cited as precursors. The text that gave the modern theory a recognizable form appeared in 2021 on Agora Road’s Macintosh Cafe in a thread titled “Dead Internet Theory: Most Of The Internet Is Fake”.

The post gathered existing suspicions into a larger story: the internet felt less human and less varied than it once had; bot activity was widespread; algorithms amplified repetitive material; and much apparent online life might be artificial. Stronger versions added claims of deliberate manipulation by governments, corporations, or other powerful actors.

Later in 2021, Kaitlyn Tiffany’s Atlantic article “Maybe You Missed It, but the Internet ‘Died’ Five Years Ago” brought the obscure theory to a much wider audience.

The bundled claims are not equivalent

At least five questions tend to get mixed together:

Is a large share of web traffic automated? Are many public accounts automated? Is a large share of visible content machine-generated? Do recommendation systems make the Web feel more repetitive by concentrating attention? And is this artificial activity centrally coordinated to manipulate the public?

Evidence for one does not establish the others.

For example, Imperva’s 2026 Bad Bot Report says automated requests accounted for more than half of observed web traffic in 2025. That is significant evidence that machines generate an enormous amount of network activity.

It does not mean more than half of blog posts, forum comments, social-media users, or people are bots. Search crawlers, monitoring tools, AI agents, scrapers, malicious automation, API clients, and many other systems all generate requests without pretending to be human authors.

Traffic is not population. Population is not authorship.

Some parts can be tested better than others

Researchers can measure bot-like behavior on a particular platform, sample account networks, analyze known automated traffic, track content duplication, or estimate the prevalence of generated material under a defined methodology. Those studies still have false positives and sampling limits, but at least the claims can be operationalized.

Claims about a coordinated hidden system controlling most online discourse require a different kind of evidence: identified actors, infrastructure, documents, financial relationships, technical links, or reproducible observations demonstrating coordination. A graph showing lots of automated requests cannot supply that missing step.

This is where Dead Internet Theory becomes useful as a study subject rather than a conclusion.

The modern internet undeniably contains bots, synthetic media, engagement manipulation, algorithmic repetition, abandoned human spaces, and industrial-scale automated traffic. Those are real phenomena worth measuring.

Whether they add up to a “dead internet” is not one question. It is a stack of questions wearing the same trench coat.