Posted on

Why AI-text detectors struggle to establish authorship

An AI-text detector does not interview the author.

It examines the text.

That distinction matters because most detectors are trying to infer origin from statistical patterns: predictability, vocabulary, sentence structure, variation, or other features associated with examples of human and machine writing. They can estimate whether a passage resembles material produced by a model. That is not the same thing as possessing evidence of who actually wrote it.

The weakness becomes obvious when human writing shares the same statistical traits.

A 2023 study published in Patterns tested seven GPT detectors against essays written by non-native English speakers. The researchers reported a high false-positive rate, with many human-written TOEFL essays classified as AI-generated. They linked part of the problem to lower linguistic variability and predictability in the writing. See GPT detectors are biased against non-native English writers.

That does not mean every detector always performs badly. A separate 2024 study using carefully sampled GRE writing data found that purpose-built detectors could achieve strong performance without the same observed bias. The disagreement is useful: detector performance depends heavily on the data, model, task, and population being tested.

Text changes after generation

Authorship also becomes harder to infer when documents are edited.

A person can heavily revise machine-generated text. A model can rewrite human text. A student can run prose through grammar software. An editor can simplify an article. A translator can alter vocabulary and sentence structure. A detector then sees the final statistical surface, not the sequence of decisions that produced it.

The models themselves also change. A detector tuned to one generation of language models may be less reliable against newer systems or against models deliberately prompted to vary their style.

This makes a percentage such as “82% AI” dangerously easy to overinterpret.

It is a classification score produced under assumptions. It is not a timestamped writing session, revision history, prompt log, document provenance record, or confession.

Authorship needs stronger evidence

If the question truly matters, stronger evidence can include version history, drafts, notes, source material, editing records, system logs, or direct testimony that can be checked against the document’s development.

A detector may contribute one clue. It should not magically turn statistical resemblance into identity evidence.

This is particularly important for Dead Internet Theory because visual and linguistic sameness can encourage people to label anything bland, repetitive, or polished as synthetic.

Sometimes they will be right.

Sometimes a boring human wrote a boring sentence.

The detector can estimate patterns. Authorship is a historical claim about how a document came into existence. Those are different jobs.

Posted on

Automated translations that create the appearance of multilingual authorship

A website published in ten languages can look like a much larger publishing operation than a website published in one.

Historically, multilingual editions often implied translators, regional editors, separate desks, or local writers. Machine translation changes that arithmetic. One original article can now appear in several languages almost instantly.

That can be useful. It can also create the appearance of many voices where there is still only one source.

SmartNews gave a straightforward example in 2026 when it launched AI-powered translation for news content in Spanish and Chinese. The feature translates headlines and full articles on demand for readers rather than pretending that separate Spanish- and Chinese-language reporters produced the underlying story. See the company’s announcement of its AI-powered translation feature.

That distinction is worth preserving.

Translation expands reach, not reporting

Suppose an English-language article is automatically translated into Spanish, Japanese, German, and Portuguese.

The publication now has five readable versions. It does not have five independently reported stories.

All five still depend on the same interviews, documents, omissions, factual errors, editorial decisions, and assumptions present in the original. If the source article is wrong, translation can distribute the mistake efficiently across languages.

Automated translation can add its own errors too: mistranslated names, technical terms, idioms, units, quotations, and culturally specific language. Human review may reduce those problems, but the important question for authorship remains separate from translation quality.

Who wrote the original? Who translated it? Was the translation machine-generated, human-edited, or independently adapted for another audience?

Translation, adaptation, and original authorship are different things

A translation tries to preserve the source in another language.

An adaptation may change examples, context, measurements, references, and explanations for a new audience.

Original reporting gathers new evidence.

Those three activities can all produce pages that look like fresh articles in a content-management system, especially when each version has its own URL and local-language headline.

Clear attribution keeps the difference visible. A page can identify the original author, link to the source edition, and note that the version was automatically translated or reviewed by a human translator.

Machine translation is not fake multilingualism. It genuinely increases the number of people who can read a piece of information.

But twenty translated pages are not twenty independent witnesses.

They are one source speaking through twenty linguistic mirrors.