Posted on

Whether a community can remain useful with clearly identified bot participants

A community does not become fake merely because some of its useful members are software.

That is an important boundary for Dead Internet Theory.

Bots can certainly flood discussions, imitate people, manufacture consensus, and make a platform look more populated than it is. But automation also performs work that human communities deliberately ask it to perform: fixing links, fighting vandalism, posting weather alerts, archiving discussions, enforcing routine moderation rules, or notifying people when something changes.

The useful distinction is not bot versus human.

It is what the bot is doing, whether people know what it is, and who remains responsible for it.

Wikipedia is an obvious counterexample

Wikipedia has used bots for more than two decades. Its bot policy requires automated processes to be useful, approved, operated responsibly, and generally run through separate accounts that clearly identify their automated status.

The Wikimedia global bot policy explicitly requires bot accounts to identify themselves and ties them to operators who can answer for their behavior. Bots perform repetitive jobs such as fixing redirects, updating data, repairing links, and reverting vandalism.

That is automation embedded inside a human-governed community rather than automation pretending to be the community.

The difference is disclosure and control.

Useful bots have boundaries

A weather bot that posts an alert when a threshold is crossed has a narrow job. A moderation bot that marks obvious spam has a defined rule set. An archive bot that preserves disappearing links provides infrastructure.

Problems grow when the role becomes ambiguous.

If a bot begins participating in arguments while presenting itself as an ordinary member, people no longer know whether social feedback represents another human. If it generates thousands of comments, its volume may crowd out actual participants. If nobody can appeal its decisions, efficiency replaces governance.

Good community automation therefore needs limits: clear identification, a defined task, rate controls, an operator, logs where appropriate, and a way for humans to challenge or stop it.

Usefulness is measurable

Instead of asking whether a community contains bots, ask what happens because the bots are there.

Do they reduce repetitive labor? Preserve information? Help moderators respond faster? Give users accurate notifications? Or do they inflate activity counts, dominate discussion, and mislead people about human participation?

Those are observable outcomes.

A community with ten thousand hidden fake personas may be deeply synthetic even if humans designed them all. A community with fifty clearly labeled maintenance bots may remain overwhelmingly human in the ways that matter.

That makes a useful ending point for Synthetic Humanity.

The presence of machines is not the definition of a dead internet.

The more revealing question is whether humans still understand the system, govern it, and remain the reason the community exists.

Posted on

Who is accountable when an autonomous account publishes false information

A bot cannot answer the editor’s phone.

That becomes important when an autonomous account publishes something false.

The system may have selected the topic, generated the wording, and posted the message without a person approving that exact sentence. But the publication still sits inside a chain of human and organizational decisions: somebody created or configured the account, somebody chose its permissions, somebody operates the service behind it, and somebody usually retains the power to stop it.

That chain is where practical editorial responsibility begins.

Find the operator before blaming the character

Autonomous systems can make online identities feel like actors in their own right. A named AI persona posts regularly, replies to users, and develops a recognizable voice. When it says something false, the easiest sentence is “the AI made a mistake.”

That describes the immediate mechanism. It does not identify who can repair the damage.

Useful questions are more concrete:

Who owns the account? Who supplied the instructions and data? Who chose to allow automatic publishing? Which platform hosts it? Is there a human operator who can see its output? Who can delete a post, publish a correction, change the prompt, disable a tool, or suspend the account?

Those roles may belong to different organizations.

A model provider may supply the underlying system while a publisher controls the deployment. A third-party automation service may schedule the posts. The social platform controls distribution and enforcement. The operator decides whether the account continues running after errors appear.

Automation does not eliminate the correction problem

Traditional publishing developed boring machinery for errors: corrections pages, editor contacts, retractions, version histories, complaints, and identifiable publishers.

Autonomous publishing still needs those functions.

A system that can publish continuously but cannot reliably receive a correction request is not more independent. It is less accountable.

The practical standard is therefore simple even when the legal questions vary by jurisdiction: somebody should be visibly responsible for the automated account’s operation, reachable when something goes wrong, and capable of correcting or stopping it.

Legal liability can depend on facts, contracts, location, platform rules, and the nature of the harm. That is a separate question from the editorial one.

Autonomy changes the workflow, not the existence of an operator

There may eventually be long chains of agents in which one system researches, another writes, another verifies, and another publishes. The individual false sentence could emerge from interactions no person predicted.

That makes logging and supervision more important, not less.

If nobody can reconstruct why an account published a claim, the system has created an accountability hole.

Dead Internet Theory often imagines a web full of machine voices with nobody behind them.

Technically, some voices may operate for long periods without live human attention.

But when a false claim needs correction, the interesting question is not whether the bot has a conscience.

It is who gave it the microphone and who still knows where the off switch is.

Posted on

The difference between synthetic supply and actual human consumption

The internet can contain a mountain of content nobody climbs.

Generative systems make that distinction increasingly important. Producing one million pages is now an engineering problem. Getting one million people to read them remains an attention problem.

Those are not the same market.

Publication counts measure supply

Suppose an automated network publishes 100,000 articles, comments, product pages, or social posts in a day.

That tells us something real about the supply side of the internet. The material exists. Servers store it. Search crawlers may fetch it. APIs may distribute it. Other bots may quote it.

But the publication count cannot tell us whether people consumed any of it.

Human consumption needs different evidence: unique human visitors, watch time, reading time, survey data, validated engagement, subscriptions, purchases, comments from identifiable people, or other measures connecting content to actual attention.

Even engagement counts require care. A 2024 paper on misinformation showed a broader version of this problem: observed engagement does not necessarily map cleanly to underlying consumer preference because producer strategies influence what users encounter and how they respond. See The distorting effects of producer strategies: Why engagement does not reveal consumer preferences for misinformation.

The lesson travels beyond misinformation. Supply, exposure, engagement, and preference are four different things.

Synthetic abundance can be mostly self-contained

Imagine 10,000 generated pages created for search engines. Crawlers request them. Monitoring systems check them. automated accounts repost links to them. Analytics registers traffic.

A graph of machine activity may look enormous before a human ever arrives.

That does not make the pages irrelevant. Some may eventually reach people. The point is that machine production can grow much faster than human attention.

This creates a strange version of Dead Internet Theory. The Web can become increasingly synthetic by volume without becoming equally synthetic in what people actually spend their time consuming.

The opposite can also happen: a relatively small amount of synthetic material can receive enormous human attention if recommendation systems amplify it.

The useful question is where the humans enter the chain

A careful study should identify which quantity it measured.

How many items were published? How many were delivered to feeds? How many impressions were generated? How many accounts interacted? How many of those accounts were likely human? How long did people actually engage?

Each step narrows the claim.

“Half of the available content is synthetic” would not mean “half of what people read is synthetic.” “Half of traffic is automated” would not mean “half of attention belongs to bots.”

The internet can manufacture supply almost without limit.

Human attention remains stubbornly finite.

Posted on

Comparing logged-in experience with a platform’s public-facing activity

The same platform can look alive in one browser and nearly empty in another.

Open a service while logged out and you may see trending posts, public profiles, search-indexable pages, or a carefully selected landing experience. Log in and the platform may replace that view with a personalized feed built from follows, history, location, subscriptions, moderation settings, and recommendations.

Neither screen is necessarily fake.

They are different windows into the same system.

Access conditions change what gets observed

A public-facing page is often designed for discovery. It may emphasize material that performs well with anonymous visitors or search engines. Some content may be hidden because it requires an account, membership, age confirmation, or community approval.

The authenticated experience has the opposite problem: it may be so personalized that one user’s feed says little about what another person sees.

Google’s own Search documentation gives a useful general example of this problem. It notes that results can differ because of time, context, location, and personalization, and that not every difference comes from personalized history. See Why your Google Search results differ from others.

Social platforms add still more variables.

A comparison needs controlled conditions

If someone wants to compare the public and logged-in experience, the test should hold as much constant as possible.

Use the same date and time. Record the location and language. Compare equivalent URLs or queries. Note whether the logged-in account follows anyone, has years of history, or was just created. Capture what is unavailable in one mode rather than silently treating missing content as nonexistent.

Even then, the comparison is a snapshot.

Recommendation systems change. Trending topics move. Experiments may place users in different interface groups. Moderation rules may vary by region.

That is why screenshots saying “look how dead this site is logged out” or “look how active my feed is logged in” are observations, not population estimates.

The two views answer different questions

The public view can tell us what the platform exposes to outsiders.

The authenticated view can tell us what a particular account is offered after the system knows more about it.

Neither automatically reveals how many humans are participating, how much activity is automated, or what the average member experiences.

For Dead Internet Theory, this distinction matters because the feeling that a platform is empty or synthetic can be heavily shaped by which surface someone is standing on.

Before deciding that the town is abandoned, it helps to check whether you are looking through the front window, the employee entrance, or somebody else’s personalized television.

Posted on

Human audits of randomly sampled public discussions

If you go looking only for creepy bot-like conversations, you will find a creepy bot-like internet.

That is not a measurement.

One way to test claims about synthetic conversation is much less dramatic: choose discussions according to a sampling rule decided in advance, then have human reviewers inspect them without selecting only the suspicious ones.

The boring part is the useful part.

Start with a sample that did not already know the answer

A reasonable audit might define a platform, date range, language, discussion type, and method for randomly selecting threads or comments. The sampling process should include quiet, ordinary, messy conversations as well as obvious spam.

Otherwise the researcher is measuring the contents of a folder labeled “weird stuff I noticed,” not the population of the platform.

Human reviewers can then classify observable characteristics: obvious commercial spam, disclosed automation, copied text, coherent human conversation, unknown or ambiguous authorship, and so on.

The word unknown is important.

A reviewer cannot reliably prove that a polished comment came from a human simply because it sounds natural. Nor can repetitive language alone prove automation.

Research on human recognition of social bots illustrates the problem. A 2024 experimental study asked people to identify bots on the VKontakte social network and found that human labeling itself can be difficult enough to undermine the idea of perfect “ground truth.” See Experimental Evaluation: Can Humans Recognise Social Media Bots?.

Disagreement is data

Suppose three reviewers inspect the same account. One calls it automated, one calls it human, and one marks it uncertain.

Throwing away the disagreement would make the final number look cleaner while hiding the most important fact: the evidence was ambiguous.

A useful audit should report how reviewers were instructed, whether they worked independently, how often they agreed, which categories caused disagreement, and how uncertain cases affected the final estimate.

Researchers can use statistical measures of inter-rater agreement, but the plain-language interpretation matters too. “Reviewers agreed on 92 percent of cases” tells a different story from “half the accounts could not be classified confidently.”

A sample answers a bounded question

Even a careful human audit does not establish how much of “the internet” is synthetic.

It can estimate what appeared in a defined sample from a defined platform under defined conditions. Different languages, communities, dates, recommendation systems, and access states may produce different populations.

That limitation is not a weakness. It is what makes the claim testable.

Dead Internet Theory becomes harder to evaluate when every strange screenshot is treated as representative and every normal conversation is dismissed as an exception.

Random sampling reverses that habit.

Do not ask the internet to show you something spooky.

Ask a sample what is actually there, and leave room for the honest answer: sometimes we cannot tell.

Posted on

Behavioral timing as a clue to automation and its limits

A person sleeps. A bot does not have to.

That makes timing one of the most tempting clues in bot detection. An account posts every ten minutes for three days. Replies appear within seconds at all hours. Hundreds of messages arrive with machine-like regularity. The pattern looks less like ordinary human behavior and more like a scheduler or script.

Sometimes that is exactly what it is.

Researchers studying social bots have long used temporal features such as posting frequency, intervals between actions, bursts of activity, and round-the-clock operation as part of larger detection systems. But timing becomes unreliable when it is treated as a standalone test.

Humans automate their clocks too

A perfectly real person can schedule posts in advance.

Newsrooms queue headlines. Businesses schedule marketing messages. creators prepare posts for different time zones. Customer-service teams hand an account from one shift to another. A family or organization may share one login. Someone working nights can look suspicious to a model trained around daytime routines.

Automation can also deliberately imitate human timing by introducing random delays, quiet periods, and varied schedules.

The result is an arms race in which simple rules become less useful precisely because both humans and bots can violate them.

A 2020 PLOS ONE study on automatic bot detection demonstrated the broader problem. Researchers tested Botometer scores against verified human and bot datasets and found classification errors, including false positives and false negatives. The authors warned that automated scores can vary and should not be treated as ground truth. See The false positive problem of automatic bot detection in social science research.

Timing features are subject to the same caution.

Patterns become stronger when they agree with other evidence

Suppose an account posts exactly once every sixty seconds around the clock.

That is suspicious.

Now add identical phrasing, API-generated metadata, synchronized behavior with hundreds of related accounts, known automation software, or an operator who openly identifies the account as a bot. The case becomes much stronger.

By contrast, irregular human-looking timing does not prove a person is present. Modern automation can produce randomness cheaply.

Behavioral evidence works best as a bundle: timing, content similarity, network structure, account history, technical metadata, and known ownership reinforcing one another.

For Dead Internet Theory, this prevents an easy mistake. The internet contains plenty of behavior that looks mechanical. Some of it is mechanical. Some of it is humans using tools. Some of it is humans behaving repetitively because humans are extremely capable of doing repetitive things.

A clock can point toward automation.

It cannot identify the hand that wound it.

Posted on

Why AI-text detectors struggle to establish authorship

An AI-text detector does not interview the author.

It examines the text.

That distinction matters because most detectors are trying to infer origin from statistical patterns: predictability, vocabulary, sentence structure, variation, or other features associated with examples of human and machine writing. They can estimate whether a passage resembles material produced by a model. That is not the same thing as possessing evidence of who actually wrote it.

The weakness becomes obvious when human writing shares the same statistical traits.

A 2023 study published in Patterns tested seven GPT detectors against essays written by non-native English speakers. The researchers reported a high false-positive rate, with many human-written TOEFL essays classified as AI-generated. They linked part of the problem to lower linguistic variability and predictability in the writing. See GPT detectors are biased against non-native English writers.

That does not mean every detector always performs badly. A separate 2024 study using carefully sampled GRE writing data found that purpose-built detectors could achieve strong performance without the same observed bias. The disagreement is useful: detector performance depends heavily on the data, model, task, and population being tested.

Text changes after generation

Authorship also becomes harder to infer when documents are edited.

A person can heavily revise machine-generated text. A model can rewrite human text. A student can run prose through grammar software. An editor can simplify an article. A translator can alter vocabulary and sentence structure. A detector then sees the final statistical surface, not the sequence of decisions that produced it.

The models themselves also change. A detector tuned to one generation of language models may be less reliable against newer systems or against models deliberately prompted to vary their style.

This makes a percentage such as “82% AI” dangerously easy to overinterpret.

It is a classification score produced under assumptions. It is not a timestamped writing session, revision history, prompt log, document provenance record, or confession.

Authorship needs stronger evidence

If the question truly matters, stronger evidence can include version history, drafts, notes, source material, editing records, system logs, or direct testimony that can be checked against the document’s development.

A detector may contribute one clue. It should not magically turn statistical resemblance into identity evidence.

This is particularly important for Dead Internet Theory because visual and linguistic sameness can encourage people to label anything bland, repetitive, or polished as synthetic.

Sometimes they will be right.

Sometimes a boring human wrote a boring sentence.

The detector can estimate patterns. Authorship is a historical claim about how a document came into existence. Those are different jobs.

Posted on

Watermark removal through ordinary editing and format conversion

A missing watermark does not prove that an image was made by a human.

That sounds obvious, but it matters because watermarking is often discussed as though it permanently divides synthetic media from everything else. In practice, digital files are copied, resized, recompressed, screenshotted, exported, and converted between formats constantly. Some provenance signals survive those operations. Others do not.

The first distinction is between visible watermarks and embedded information.

A visible mark is part of the pixels: a logo, label, or pattern intentionally placed where a viewer can see it. Embedded provenance can instead live in metadata or a cryptographically signed manifest associated with the file. Systems such as the C2PA specification are designed to record information about an asset’s origin and editing history.

Neither approach is indestructible.

Ordinary processing can erase evidence

A screenshot creates a new file from what was displayed on the screen. The screenshot may visually reproduce the original image almost perfectly while omitting the original file’s metadata.

Exporting from an image editor can do something similar. A JPEG may be converted to PNG. A platform may resize the image and create its own copy. A messaging service may recompress it. Metadata can disappear even when nobody intended to hide anything.

Visible marks have different weaknesses. Cropping may remove a mark near an edge. Resizing can make a subtle mark unreadable. Heavy compression can interfere with invisible signal-based watermarks. A screenshot generally preserves a visible watermark because the pixels are copied, but it may destroy metadata attached to the original file.

This is one reason provenance systems increasingly discuss more than one recovery method. C2PA, for example, supports the idea of durable credentials that can use mechanisms such as content fingerprints or watermarks to reconnect altered media with provenance stored elsewhere.

Absence is weak evidence

Suppose an image arrives without any AI label or provenance metadata.

That tells you what is present in the file you received. It does not reliably tell you what was present three transformations earlier.

The image may have originated in a camera and lost metadata. It may have been generated by a model and lost metadata. It may have been edited by a person, processed by a platform, downloaded, and screenshotted several times.

The same problem exists in reverse. A surviving provenance credential can provide useful evidence about origin and editing history, but it does not make the depicted claim true. Authentic provenance and factual accuracy are separate questions.

Watermarks therefore work best as evidence that can survive a chain of handling, not as a magical binary test.

The internet is an enormous file-conversion machine. Any detection system that assumes media will remain in its pristine original container is going to meet reality very quickly.

Posted on

Provenance labels and their survival through reposting

A provenance label is useful only while it remains attached to the thing it describes.

That sounds obvious until media starts moving through the internet.

An image leaves a camera, enters an editor, gets uploaded to a social network, downloaded, resized, screenshotted, pasted into a message, reposted by another account, and compressed again. Somewhere in that journey, the information explaining where the image came from may disappear.

The Coalition for Content Provenance and Authenticity is trying to make that chain more durable through the C2PA standard and Content Credentials. The current C2PA technical specification defines a system for attaching cryptographically verifiable provenance information to digital assets: who or what created them, which tools modified them, and how their history developed.

That is stronger than a plain text label saying “AI generated.”

It is still not magic glue.

The file and its history can become separated

Ordinary internet workflows routinely create new files.

A platform may resize a photograph. A messaging app may recompress it. A user may take a screenshot. An editor may export the picture into another format. A social network may strip metadata from the downloadable rendition even if it inspected that metadata during upload.

Once the distributed copy no longer carries the original manifest, a viewer may see the media without seeing its provenance.

C2PA explicitly anticipates this problem. Its specification includes soft bindings, such as content fingerprints or invisible watermarks, that can help reconnect a modified asset with provenance stored elsewhere. The standard describes a “Durable Content Credential” as one that uses such mechanisms so provenance can potentially be recovered even after embedded metadata is lost.

That is an important design choice because exact-file identity is fragile on the modern web.

Provenance is evidence, not a truth machine

Even an intact credential has limits.

It can provide tamper-evident claims about an asset’s history and the entities that signed those claims. It does not prove that every statement depicted in the media is factually true. A perfectly authentic photograph can have a misleading caption. A real camera can photograph a staged scene. An editor can truthfully disclose an AI-generated asset whose underlying claim is still nonsense.

Provenance answers questions such as where did this come from and what happened to it?

Truth requires additional evidence.

Reposting is the durability test

For provenance systems to matter at internet scale, information has to survive beyond the first platform that understands it.

That means compatible tools, durable bindings, visible labels, repositories that can recover manifests, and platforms willing to preserve or reconstruct provenance when they transform media.

Otherwise a credential may work beautifully at the point of creation and vanish three reposts later—exactly when somebody encounters the content without its original context.

Dead Internet Theory is partly a problem of uncertain origins. Synthetic media makes that uncertainty harder.

Provenance systems offer a practical response, but their real test is not whether a label can be attached.

It is whether the label can survive the internet.