Posted on

Autonomous agents browsing and purchasing on behalf of humans

The next bot visiting an online store may be there because a human genuinely wants to buy something.

That sounds contradictory only if every automated visitor is classified as fake traffic.

Agentic commerce is built around the opposite idea: a person gives software an objective and permission to act, then the software searches, compares, negotiates, or purchases on that person’s behalf.

Visa’s current Intelligent Commerce documentation describes agents that can search for products and make purchases using credentials and controls tied to the user’s authenticated instructions. Mastercard has similarly described autonomous agentic commerce in which software can find and buy products within limits set by the consumer.

That traffic is automated.

The demand behind it may be entirely human.

Delegation breaks the old traffic categories

Suppose someone tells an agent:

Find a replacement water filter compatible with this refrigerator, spend less than $45, and order it from a reputable seller.

The agent may visit search systems, merchant catalogs, product pages, compatibility documents, review pages, checkout endpoints, and payment services. A human could have produced dozens of those requests manually. Instead software produces them in a few minutes.

From the server’s perspective, the activity may look like bot traffic.

From the shopper’s perspective, it is delegated labor.

Calling it an independent “user” would be misleading because the agent does not have its own need for a refrigerator filter. Calling the demand fake would also be misleading because a real person authorized a real purchase.

Intent and actor become separate measurements

This distinction creates a major problem for future internet statistics.

A website may receive fewer direct human page views while facilitating more human economic activity. One person’s agent could browse hundreds of pages. A family might use several specialized agents. A single agent might compare twenty sellers before producing one purchase.

The number of machine requests can therefore grow while the number of underlying people stays constant.

Payment networks are already treating this as an authentication problem. Visa and Mastercard have developed mechanisms intended to help merchants distinguish authorized shopping agents from malicious bots and to bind agent actions to consumer instructions.

That distinction will matter far beyond payments.

Dead Internet Theory often treats machine activity as evidence that human activity has disappeared. Autonomous agents introduce a third category: machine activity carrying human intent.

The browser may not be a person anymore.

But somebody may still be shopping through it.

Posted on

Synthetic players used to populate multiplayer environments

A multiplayer match can be full without being full of people.

That is not necessarily fraud. Sometimes it is game design.

In September 2019, Epic Games announced that Fortnite would add bots to Battle Royale matchmaking. The company said the bots would behave similarly to normal players and help create a smoother path for players developing their skills. As players improved, Epic said they would encounter fewer bots, and competitive playlists would exclude them. See Epic’s original Fortnite Matchmaking Update.

That is a wonderfully clean example because the synthetic players were not evidence of a secret fake population. They were an openly designed part of the game.

A bot can serve the player rather than deceive the player

Multiplayer games have several reasons to insert nonhuman participants.

They can reduce queue times when there are not enough suitable players available. They can make beginner matches less punishing. They can fill abandoned slots, provide practice targets, keep older modes playable, or tune the difficulty of a session.

In those cases the bot has a job inside the game system.

But the presence of bots changes what population metrics mean.

A player may enter a 100-character match and reasonably feel surrounded by activity. That does not establish that 99 other humans were simultaneously available in that skill bracket, region, platform pool, and mode.

The software can supply some of the missing population itself.

Identification changes the experience

Whether players care depends partly on disclosure and expectations.

A labeled practice bot is ordinary game machinery. A bot intentionally designed to be indistinguishable from a human opponent can create a different reaction, especially if players believe matchmaking statistics represent a large active community.

Even then, identifying synthetic players from behavior alone is difficult. Weak human players behave strangely. Experienced players experiment. Network problems produce robotic movement. Human opponents can have names generated by the game. A suspiciously easy victory is not forensic evidence.

Developers, telemetry, explicit labels, or documented matchmaking rules provide much stronger evidence.

This makes multiplayer games an unusually useful Dead Internet Theory case study.

A synthetic population can be real in the sense that it genuinely exists inside the environment, affects outcomes, occupies slots, and reacts to players. Yet it is not a human community.

The interesting question is therefore not whether the match is “fake.”

It is what kind of population the game is presenting—and whether the player is expected to know the difference.

Posted on

Simulated livestream audiences and the appearance of shared presence

Livestreaming sells more than video.

It sells the feeling that other people are there with you right now.

The viewer counter rises. Chat moves quickly. Emotes flood the screen. Someone reacts to a joke before you finish laughing. Even when everyone is physically scattered, the stream feels like a room.

That sense of shared presence makes audience simulation unusually powerful.

Twitch explicitly defines view-botting as artificially inflating a live viewer count with illegitimate scripts or tools. Its guidance also notes that view-botting can be accompanied by chat bots intended to imitate interaction between the streamer and viewers. See Twitch’s How to Handle Viewership Botting and Fake Engagement.

The platform distinguishes this from legitimate traffic arriving through hosting, embeds, or promotion. That distinction matters because a number on a screen does not explain what produced the number.

A concurrent connection is not the same as attention

Suppose a stream shows 2,000 viewers.

That could represent 2,000 people actively watching. It could include people with the stream running in a background tab, legitimate embedded players, automated monitoring, or artificial clients designed specifically to increase the counter.

Those categories all create network activity, but they do not represent the same social reality.

Chat can be even more deceptive because text looks intentional. A scripted account can post emotes, greetings, repeated praise, or simple reactions timed to resemble audience participation. A visitor entering the stream may infer that a large, active community already exists.

This is social proof with a heartbeat animation.

Real audiences leave more than counts

No single metric proves genuine engagement.

Human audiences produce uneven behavior: conversations that persist across streams, recognizable regulars, subscriptions, moderation history, callbacks, questions, jokes, disagreements, and participation that does not move in suspiciously synchronized blocks.

Even those signals can be automated, so investigators usually need patterns across time rather than one screenshot of a viewer counter.

This does not mean every unexpectedly popular stream is botted. Twitch itself warns against confusing artificial inflation with legitimate sources of sudden traffic.

The narrower lesson is that apparent co-presence can be manufactured.

A livestream with a crowded counter and noisy chat may genuinely be a packed digital room. It may also be a quieter room surrounded by software making chair noises.

Dead Internet Theory often asks whether anybody is really online.

Livestream botting turns that abstract question into something you can watch happen in real time.

Posted on

Generated comments used to seed an otherwise empty discussion

An empty comment section tells a visitor something.

Maybe nobody cared. Maybe nobody has arrived yet. Maybe the community is small. Whatever the explanation, zero replies are honest information about the current state of the conversation.

That makes emptiness tempting to fix.

A platform can generate a starter comment, a suggested question, or an apparent reaction so the next visitor does not feel like the first person walking into an empty room. Used transparently, that can be a design tool. Used without disclosure, it becomes simulated participation.

A 2026 Scientific Reports experiment on generative AI in social-media discussions tested interventions including AI-written conversation starters, comment assistance, feedback, and reply suggestions. Some tools increased activity, but the researchers also found tradeoffs involving perceived quality and authenticity. See The impact of generative AI on social media: an experimental study.

The important word there is experimental. Participants were studying AI-assisted discussion, not being tricked into believing synthetic participants were ordinary members of a live community.

Seeding changes what the room appears to contain

A genuine conversation starter can be simple: the platform itself posts a clearly labeled prompt asking visitors what they think.

A deceptive version looks different. Several apparently ordinary users arrive first. One asks a question. Another agrees. A third adds a mild disagreement. The exchange exists primarily to make later visitors believe people are already present and engaged.

That appearance matters because human beings use visible participation as social evidence. A thread with comments feels safer to enter than one with none. A product with discussion looks more noticed. A new community with chatter looks more established.

Generated comments can therefore manufacture not just text but social proof.

Disclosure changes the meaning

There is nothing inherently wrong with a machine starting a conversation.

A clearly labeled bot can welcome new users, post daily questions, summarize prior discussion, or keep a support forum organized. Nobody needs to mistake it for a stranger who independently wandered in and cared enough to comment.

The problem is the false inference created when synthetic comments are dressed as ordinary participation.

Ten generated comments do not equal ten interested people. They may represent one operator trying to overcome the cold-start problem of an empty community.

That distinction fits Dead Internet Theory almost perfectly. A page can look socially occupied while containing very little spontaneous human activity.

The useful question is not merely, “Was this comment generated?”

It is, “What does this comment ask me to believe about who is actually here?”

Posted on

Repeated model phrasing as a source of apparent cultural sameness

Sometimes the modern web feels as though thousands of unrelated pages hired the same copy editor.

The words recur. The paragraph rhythms recur. The same polite transitions and tidy three-part explanations appear in product pages, newsletters, essays, documentation, and corporate announcements that have no obvious connection to one another.

Large language models offer one plausible mechanism for some of that sameness.

A 2025 paper in the Proceedings of COLING examined unusually rapid changes in scientific English and identified a set of words whose increased frequency was consistent with large-language-model use. The most famous example was “delve,” alongside words such as “intricate” and “underscore.” See Why Does ChatGPT “Delve” So Much?.

The interesting part is not that one word supposedly exposes a robot. It does not.

Shared generators can create shared habits

Language models do not choose every valid phrase with equal probability. Training, fine-tuning, system prompts, preference optimization, and common user instructions can all push output toward certain constructions.

Then publishing scale amplifies the bias.

If one person uses a model to draft one article, a recurring phrase is trivia. If thousands of publishers use closely related systems to generate millions of pages, small stylistic preferences can become visible across the web.

This produces a strange cultural effect: sites built by unrelated organizations can begin to sound related.

The result resembles standardization. Introductions converge. Explanations arrive in similar shapes. Conclusions become reassuring. Certain adjectives become fashionable almost overnight.

Style is evidence of influence, not proof of authorship

This is where detection gets slippery.

Humans imitate fashionable language. Editors impose house styles. Search optimization rewards structures that competitors then copy. Templates produce repeated phrasing. A human who reads a great deal of AI-assisted prose may begin using some of the same language without ever pressing a generate button.

So finding “delve” on a page does not establish that a machine wrote it. Neither does a tidy em dash, a three-item list, or a sentence beginning with “It is important to note.”

The defensible claim is broader: widespread use of similar generative systems can introduce measurable lexical patterns into published language.

That matters to Dead Internet Theory because people often describe the modern web as feeling strangely homogeneous before they can explain why.

Some of that feeling may come from business consolidation, search optimization, platform templates, and shared cultural trends. Some may come from millions of writers using the same handful of models as invisible collaborators.

The web does not have to be fake to become stylistically flatter.

It only needs enough people—and machines—to keep reaching for the same words.

Posted on

Synthetic content entering future model-training collections

Once machine-generated material is published on the open web, future data collectors may not know where it came from.

A generated article can be indexed, quoted, reposted, translated, scraped, summarized, and eventually included in a training collection. The next model may then learn partly from the output of earlier models.

That creates an obvious feedback-loop question: what happens when generative systems increasingly learn from synthetic material produced by generative systems?

One influential answer came from a 2024 Nature paper, AI models collapse when trained on recursively generated data. The researchers showed that indiscriminate recursive training on model-generated data can progressively distort the learned distribution. Rare parts of the original data disappear first, and later generations can become increasingly narrow.

The phrase model collapse came to summarize that risk.

Unfortunately, the slogan is simpler than the experiment.

Synthetic data is not one substance

The paper did not show that one AI-written webpage contaminates a future model like plutonium dropped into a reservoir.

Its results depend on the training setup, how synthetic samples are generated, how much original data is retained, and how later generations are selected. In one of the paper’s language-model experiments, preserving a sample of original training data substantially reduced the degradation compared with replacing the data recursively.

Synthetic data is also deliberately useful in many machine-learning systems. Researchers generate examples to fill rare categories, build question-answer datasets, test safety behavior, or create training cases that would be expensive to label by hand.

So the useful distinction is not human data good, synthetic data poison.

It is between controlled synthetic data with known provenance and uncontrolled synthetic material that enters a collection as though it were an independent sample of the world.

The web makes provenance difficult

This is where the open internet becomes messy.

A crawler may see a thousand pages without knowing whether they represent a thousand human authors, fifty content farms using the same model, translated copies of one generated article, or output recursively derived from other generated output.

The apparent size of the corpus can therefore grow faster than its independent informational diversity.

That is the Dead Internet Theory connection worth taking seriously. The danger is not mystical AI inbreeding. It is a data-accounting problem.

If future models are trained on web-scale collections, identifying source quality, duplication, synthetic provenance, and genuinely independent human material becomes increasingly important.

The web has always contained copies. Generative systems simply make it possible to manufacture those copies, variations, and derivatives at industrial speed.

Future models will not merely need more data.

They will need to know what kind of data they are looking at.

Posted on

AI-generated citations pointing to nonexistent publications

A fake citation can wear a very convincing suit.

It may have plausible authors, a plausible journal, a plausible year, a plausible title, and formatting that looks exactly like a reference copied from an academic paper. The weakness appears only when somebody tries to find the source.

That failure mode has been measured rather than merely complained about. A 2023 Scientific Reports study examined 636 citations generated in short literature reviews by GPT-3.5 and GPT-4. In that specific experiment, 55 percent of the GPT-3.5 citations and 18 percent of the GPT-4 citations referred to works the researchers could not verify as having actually been published. See Fabrication and errors in the bibliographic citations generated by ChatGPT.

Those numbers should not be treated as timeless error rates for every later model. They describe particular models, prompts, and tasks. What they demonstrate is the mechanism: fluent text generation can produce the shape of scholarship without the underlying publication.

Formatting is cheap; existence is the hard part

Bibliographic references are unusually easy for language models to imitate because they are highly patterned.

Author names. Year. Article title. Journal. Volume. Issue. Pages. DOI.

A model can assemble those pieces into something statistically convincing even when no database record sits underneath them. Real journals and real researchers can even be combined into an imaginary paper, which makes casual inspection less useful.

The mistake becomes dangerous when later writers copy the fabricated reference without checking it. A false citation can migrate from an AI answer into a blog post, report, student paper, reference manager, or another model’s training material. Repetition then makes the nonexistent work look increasingly established.

Verification means following the reference

The practical test is boring and effective: try to locate the original publication.

Search the journal or publisher. Resolve the DOI. Check Crossref, PubMed, a library catalog, or the relevant scholarly database. Confirm that the title, authors, year, and publication details actually match.

Even a real paper can be misrepresented, so verification should not stop at proving the paper exists. The source also has to support the claim being attached to it.

This is where synthetic citations fit Dead Internet Theory unusually well. The problem is not simply that a machine made a mistake. It is that the web can acquire references to intellectual objects that never existed, and those references can then circulate like ordinary scholarship.

A citation is supposed to point backward to evidence.

When the pointer leads nowhere, the polished formatting is just scenery.

Posted on

Generated answers to machine-generated questions

A question-and-answer page usually implies that somebody wanted to know something.

That assumption is no longer safe.

Modern language systems can generate a question, generate an answer to that question, and repeat the process thousands or millions of times without any person ever expressing the underlying curiosity. Technically, that can be useful. Publicly, it can create the appearance of demand where none existed.

Synthetic question-and-answer generation has been studied for years as a machine-learning technique. In 2020, researchers showed that models could create synthetic question-and-answer pairs at scale and use them to train question-answering systems. Their work, Training Question Answering Models From Synthetic Data, treated the generated material as training data rather than evidence that real people had asked those questions.

That distinction is the whole issue.

A synthetic dataset is honest about what it is

In machine-learning research, automatically generated questions can reduce the cost of manually labeling datasets. The purpose is explicit: manufacture examples so a model can practice connecting questions with answers.

A public knowledge site works differently.

Readers usually infer that a question represents some form of real information demand. Somebody encountered a problem, wondered about a subject, or needed clarification. The answer exists because the question existed first.

When both sides are generated, that causal chain disappears.

The page may still contain useful information. A machine-generated question such as “How does a checksum detect file corruption?” can receive a perfectly good generated answer. But the existence of the page tells us nothing about whether anybody actually asked it, searched for it, or needed it.

The loop can manufacture its own reason for existing

The problem becomes stranger when content systems use generated questions primarily to justify generated answers.

One system identifies a topic gap. Another generates plausible questions. A model answers them. Pages are published. Search engines discover the pages. Later systems scrape those pages as examples of what people discuss online.

At that point the web contains a conversation whose demand and supply were both synthetic.

That does not automatically make the content worthless. Synthetic Q&A is useful enough that researchers continue to build carefully validated datasets around it. The important test is whether the material helps actual readers and whether its origin is represented honestly.

A generated FAQ based on a real manual can be useful. A million autogenerated Q&A pages built only because a publishing system discovered a keyword gap are something else.

The difference is not whether a machine wrote the question.

It is whether the question serves a human information need—or merely gives another machine something to answer.

Posted on

Automated accounts that recycle one another’s material

Ten accounts repeating the same thing can look like ten sources.

That is one of the simplest ways automation can distort the apparent size of an online conversation. A bot does not need to invent anything. It can copy a caption, repost a link, slightly alter a sentence, or relay material from another automated account. After enough hops, the visible network looks busy even though very little independent creation occurred.

Researchers have documented this pattern in commercial social-media environments. A 2018 study of SoundCloud activity examined more than 12 million comments and found highly active suspicious accounts that posted repetitive comments, frequently reposted existing content, and contributed relatively little original material. The authors used comment uniqueness and network behavior as clues when distinguishing likely bots and semi-automated accounts from ordinary users. See Social bots in a commercial context — A case study on SoundCloud.

Circulation is not creation

Reposting has legitimate uses. Human communities share jokes, announcements, songs, emergency information, and news links constantly. Automated accounts can also perform useful redistribution: mirroring updates, relaying weather alerts, or syndicating posts to another platform.

The measurement problem begins when repeated circulation is interpreted as independent authorship or independent agreement.

Imagine one account posts a sentence. Twenty automated accounts copy it. Another hundred accounts encounter those copies instead of the original. A researcher who counts only visible posts might record 121 pieces of activity. A reader may perceive widespread agreement. Yet the intellectual source may still be one person, one script, or one upstream feed.

Attribution gets weaker as the material moves. Usernames change. Links disappear. Screenshots replace original posts. Small rewrites break exact-text matching. Eventually a recycled statement may look native to the account currently carrying it.

Repetition can manufacture apparent consensus

This is why source tracing matters more than raw counts.

A useful investigation asks whether accounts are posting independently, whether they share identical URLs or wording, whether their timing is synchronized, and whether the chain leads back to a common source. Recent research on coordinated reposting on Bluesky has used precisely this kind of timing and shared-content analysis to distinguish ordinary repost behavior from suspicious coordination.

None of this means repeated material is automatically bot-generated. Humans copy one another too. Fan communities, customer-service teams, volunteer campaigns, and newsrooms all reuse language.

The narrower point is easier to establish: many visible copies do not imply many independent origins.

Dead Internet Theory often treats repetition as evidence that nobody real is speaking. That conclusion goes too far. Repetition can be human, automated, coordinated, accidental, or mixed.

But when automated accounts recycle one another’s material, the internet can appear to contain more voices than it contains sources. That distinction matters.

Posted on

Bots replying to other bots in public comment threads

A comment thread with fifty replies looks busier than a comment thread with five.

That visual shortcut works only if “reply” is being used as a rough proxy for human attention. Once automated accounts can trigger other automated accounts, the relationship breaks down.

A bot can post. Another bot can detect the post and answer. The first bot can react to the answer. A moderation bot can add a warning. A link-preview bot can fetch the URL. A promotional account can repost the exchange somewhere else. The platform records activity at every step even if no human is watching in real time.

This is not merely theoretical.

A 2023 study in EPJ Data Science examined an experimental ecosystem of six social bots on Twitter. The researchers explicitly tracked bot-bot, bot-human, and human-bot interactions. They even used a mediator to manage timing partly because the bots could otherwise interact with one another simultaneously and risk violating platform spam rules. See Emergent local structures in an ecosystem of social bots and humans on Twitter.

The experiment found bot-to-bot exchanges were more mechanistic and less diverse than interactions involving humans. That distinction is exactly what raw activity counts hide.

A reply is an event, not proof of a person

To establish that a public exchange is genuinely bot-to-bot, researchers need more than the conversation looking repetitive.

Useful evidence can include known account ownership, source code, API behavior, posting schedules, platform labels, controlled experiments, or technical logs showing that automated systems generated both sides. Behavioral classifiers can help identify likely automation, but probability is not the same thing as proof.

That matters because humans repeat themselves too. People schedule posts. Communities use templates. Customer-service workers paste canned responses. An odd conversation is not automatically a machine conversation.

Volume can exceed attention

The deeper Dead Internet Theory question is what happens when machines generate activity primarily in response to other machines.

A thousand replies might represent a thousand people. They might represent one human operator controlling several tools. They might represent a handful of automated agents bouncing triggers around a network. The visible counter alone cannot tell us which population exists underneath it.

This does not mean public comment sections are secretly composed mostly of bots. That is a much larger claim requiring platform-specific evidence.

It means something narrower and measurable: online conversation volume and human participation are no longer interchangeable quantities.

The internet can be busy without anybody being particularly busy using it.