Posted on

Automated translations that create the appearance of multilingual authorship

A website published in ten languages can look like a much larger publishing operation than a website published in one.

Historically, multilingual editions often implied translators, regional editors, separate desks, or local writers. Machine translation changes that arithmetic. One original article can now appear in several languages almost instantly.

That can be useful. It can also create the appearance of many voices where there is still only one source.

SmartNews gave a straightforward example in 2026 when it launched AI-powered translation for news content in Spanish and Chinese. The feature translates headlines and full articles on demand for readers rather than pretending that separate Spanish- and Chinese-language reporters produced the underlying story. See the company’s announcement of its AI-powered translation feature.

That distinction is worth preserving.

Translation expands reach, not reporting

Suppose an English-language article is automatically translated into Spanish, Japanese, German, and Portuguese.

The publication now has five readable versions. It does not have five independently reported stories.

All five still depend on the same interviews, documents, omissions, factual errors, editorial decisions, and assumptions present in the original. If the source article is wrong, translation can distribute the mistake efficiently across languages.

Automated translation can add its own errors too: mistranslated names, technical terms, idioms, units, quotations, and culturally specific language. Human review may reduce those problems, but the important question for authorship remains separate from translation quality.

Who wrote the original? Who translated it? Was the translation machine-generated, human-edited, or independently adapted for another audience?

Translation, adaptation, and original authorship are different things

A translation tries to preserve the source in another language.

An adaptation may change examples, context, measurements, references, and explanations for a new audience.

Original reporting gathers new evidence.

Those three activities can all produce pages that look like fresh articles in a content-management system, especially when each version has its own URL and local-language headline.

Clear attribution keeps the difference visible. A page can identify the original author, link to the source edition, and note that the version was automatically translated or reviewed by a human translator.

Machine translation is not fake multilingualism. It genuinely increases the number of people who can read a piece of information.

But twenty translated pages are not twenty independent witnesses.

They are one source speaking through twenty linguistic mirrors.

Posted on

Generated video presenters and simulated newsroom authority

Television news has trained viewers to read a familiar set of signals very quickly.

A person sits behind a desk. Graphics move over one shoulder. The voice is steady. The lighting is professional. The presenter looks directly into the camera and speaks with the calm certainty of somebody who belongs in a newsroom.

Generative video can reproduce most of that surface without requiring a human presenter at all.

Xinhua News Agency demonstrated the idea publicly in 2018 when it introduced what it called an AI news anchor, developed with Sogou. Xinhua described the system as having the image, voice, facial expressions, and actions of a real person and said it could read supplied texts continuously across its website and social platforms. See Xinhua’s announcement of its AI news anchor.

The important word there is supplied.

Presentation is not reporting

An anchor can read a sentence without knowing whether the sentence is true.

That is already true for human presenters to some extent: newsroom reporting is a collective process involving reporters, editors, producers, photographers, researchers, and sources. The face on screen is not necessarily the person who gathered the evidence.

A synthetic presenter makes that separation impossible to ignore.

The avatar can look authoritative while contributing nothing to the reporting process. It may have conducted no interview, examined no document, visited no location, and made no editorial judgment. It is an interface between a script and an audience.

That can be perfectly legitimate if the publisher is clear about it.

Somebody still has to own the sentence

The useful question is not “Is the anchor real?” It is “Who stands behind what the anchor is saying?”

A trustworthy operation should still have identifiable editorial responsibility: a publication, newsroom, editor, producer, reporter, or documented source chain. Viewers should be able to distinguish a generated presentation layer from the process that produced the underlying claims.

This becomes especially important when synthetic presenters are cheap enough to create for thousands of channels, languages, niches, and localities. A person-shaped interface can give a tiny automated publishing operation the visual weight of a television network.

That appearance is not itself evidence of fraud. It is evidence that old visual shortcuts for judging institutional scale have become less reliable.

A suit, a desk, and eye contact used to be expensive.

Now they can be rendered.

Posted on

Synthetic historical photographs circulating as ordinary archival images

Historical photographs feel different from illustrations because they appear to contain a physical trace of an event.

A camera was somewhere. Light hit film or a sensor. Somebody pressed a shutter. Even when a photograph is staged, cropped, selectively captioned, or poorly interpreted, there is usually an original object with a date, creator, negative, publication history, or archival chain that researchers can investigate.

A generated “historical photograph” can imitate the surface of all of that without any of it existing.

The problem becomes serious after the image leaves the generator.

A caption can manufacture an archive

In 2024, AAP FactCheck investigated a black-and-white Facebook image claiming to show Henry Ford sitting in his first automobile in 1896. The image looked unusually crisp and historical enough to attract attention, but it was AI-generated and did not accurately depict Ford or his quadricycle. AAP found it was one of multiple synthetic “historical” pictures being circulated by the same kind of page. See AI used to generate fake history photos.

The dangerous step is not merely generating the picture. It is attaching a confident caption.

“Henry Ford, 1896” changes an illustration into an apparent record. Repost it a few times, crop away the original context, and the image can begin appearing in collections, blogs, slideshows, family-history pages, and search results as if it came from an archive.

At that point visual inspection is a weak defense. Film grain, scratches, period clothing, lens softness, and damaged borders can all be synthesized too.

Provenance matters more than vibes

For a supposed archival photograph, useful questions are boring and powerful: Which archive holds it? Is there a catalog number? Who is the photographer? When was the image first published? Is there a negative, contact sheet, newspaper reproduction, or accession record? Can another institution independently identify it?

That chain is provenance.

The absence of provenance does not automatically prove an image is fake. Plenty of genuine family photographs survive with almost no metadata. But the less context an image has, the weaker the historical claim should become.

Synthetic imagery makes this standard more important because “looks old” no longer implies “was created in the past.”

The internet has always mislabeled photographs. Generative systems add a new category: photographs of events that never passed through a camera at all.

History needs more than sepia.

Posted on

AI-written book summaries without reliable source grounding

A summary has one job that matters more than elegance: remain faithful to the thing being summarized.

That sounds obvious, but generative systems make it easy to produce a polished account of a book without showing where any particular claim came from. A paragraph can sound exactly like literary criticism while quietly inserting an event, motive, quotation, or theme that is not actually in the text.

The problem is not unique to books. Researchers studying abstractive summarization have documented a general phenomenon usually called hallucination: generated summaries can contain information that is not supported by the source. An ACL 2022 paper, Hallucinated but Factual!, examined precisely this problem and distinguished unsupported additions that happen to be true from additions that are not.

For a book summary, even a true addition can still be misleading if the task is supposed to describe what the book itself says.

A summary is a claim about a source

Suppose a generated summary says a novel is primarily about guilt after war. Maybe that is a reasonable interpretation. Maybe the book never addresses war at all and the model has blended it with another title. The sentence alone does not tell you which happened.

The same problem appears with nonfiction. A summary may attribute an argument to an author because the argument is common in books on the subject, not because it appears in that particular book.

This is why source grounding matters.

A better workflow gives the summarizer the actual text, chapter excerpts, notes, or a reliable edition and preserves a route back to the source. Important claims can be tied to chapter numbers, page ranges, quotations, or at least clearly identified sections. The more specific the summary becomes, the more useful those anchors are.

Confidence is not traceability

Readers often use summaries because they have not read the original. That makes the summary unusually powerful: there may be no immediate contradiction available in the reader’s own memory.

A generated synopsis can therefore create a strange secondhand literature in which thousands of people know what a book supposedly says without anyone checking the book.

AI can be helpful here. It can compress chapters, compare sections, extract recurring concepts, or turn notes into a study guide. But the useful version is grounded in the source rather than merely sounding like somebody who read it.

A book summary should leave a trail back to the book.

Otherwise the internet gains one more confident description of a source that nobody in the publishing chain can prove was actually followed.

Posted on

Machine-produced software tutorials and unverifiable execution claims

A software tutorial makes a quiet promise: these steps work.

That promise is stronger than ordinary explanatory prose. If an article tells you to install a package, edit a configuration file, run three commands, and expect a service to start on port 8080, the reader assumes somebody has tested that sequence or at least verified it against the relevant documentation.

Generated tutorials can imitate that confidence without executing anything.

The result is a particularly dangerous form of plausible nonsense because code has syntax. A command can look exactly right while naming a package that does not exist, using an obsolete flag, referring to a file path from another operating system, or producing an output the command could never return.

The terminal is a better fact-checker than prose

Package-name hallucinations are a concrete example. Recent research on “slopsquatting” has examined cases where coding models invent realistic package names. If a user blindly follows an installation command for a nonexistent package, that mistake can become a security problem if somebody later registers the invented name with malicious code. A 2026 paper on package-name hallucinations describes this as a supply-chain risk and tests methods for checking package existence before installation. See Names Can Hurt.

The larger lesson is simpler than the security scenario: software instructions need execution evidence.

A trustworthy tutorial can say which versions were tested, what operating system was used, what dependencies were installed, and what output appeared. Better still, the example can include a repository, fixture, test, container, or script that another person can run.

That turns “this should work” into something reproducible.

Generated code is not automatically untrustworthy

Machine assistance can be genuinely useful for explaining APIs, drafting examples, converting commands between shells, or filling in repetitive setup steps. The problem is not that a model touched the article. The problem is publishing an execution claim that nobody verified.

There is a big difference between:

pip install some-package

and:

“Tested on Python 3.13 on Ubuntu 26.04; this command was run successfully on September 17, 2026.”

The second statement is evidence about an event.

A generated tutorial can be perfectly correct. But correctness should be demonstrated through execution, not inferred from how convincingly the markdown code block is formatted.

Software does us one favor that many other subjects do not: it usually lets us test the claim.

Run the commands.

Posted on

Synthetic recipe collections and the absence of kitchen testing

A recipe is one of the easiest forms of writing to fake convincingly.

It has a predictable structure: ingredients, quantities, steps, temperature, cooking time, serving suggestion. A language model has seen enough of that pattern to generate something that looks like a recipe almost instantly.

What the page cannot tell you is whether anybody actually cooked it.

That difference matters more than it might seem. Kitchen testing checks relationships that fluent text cannot verify by itself: whether the dough is too wet, whether the sauce splits, whether the stated temperature burns the food, whether the timing is realistic, whether the proportions produce twelve servings or three, and whether an unusual ingredient combination is safe.

Plausible instructions are not the same as tested instructions

A striking example appeared in 2023 when New Zealand supermarket Pak ‘n Save released its Savey Meal-bot, an AI tool intended to suggest meals from ingredients users had on hand. When people began entering household products rather than normal groceries, the system generated dangerous outputs, including an “aromatic water mix” that would create chlorine gas. The supermarket’s warning stated that generated recipes were not reviewed by a human and were not guaranteed to be suitable for consumption. The incident was reported in detail by The Guardian.

That case involved deliberately adversarial inputs, so it should not be treated as proof that ordinary generated recipes routinely become chemical weapons. It demonstrates something narrower and more useful: the system could produce authoritative-looking culinary instructions without having any physical process behind them to catch the nonsense.

A kitchen would have caught it immediately.

Testing is evidence

Traditional recipe development often looks inefficient compared with text generation because somebody cooks the thing repeatedly. Ingredients get weighed. Oven temperatures get adjusted. Instructions get rewritten when a supposedly obvious step turns out not to be obvious.

That labor produces information.

A publisher can certainly use AI to brainstorm variations, rewrite instructions, scale quantities, or organize existing recipes. But a collection of generated recipes should not quietly inherit the authority of a tested cookbook if nobody has tested the results.

Useful disclosure can be simple: “AI-generated suggestion, not kitchen-tested,” or “Developed with AI assistance and tested by our kitchen.” Those two statements describe very different products.

Synthetic recipe collections are interesting because they expose a larger problem with generated content. The page can reproduce the form of expertise while skipping the physical act that originally created the expertise.

For food writing, the missing evidence is sitting in a pan.

Posted on

AI-generated travel guides without firsthand travel

Travel writing carries an implication that ordinary reference writing does not: somebody has been there.

A guide recommending a quiet cafe near a station, warning that a museum takes longer than expected, or suggesting which neighborhood feels dead after 9 p.m. sounds like accumulated observation. Traditional travel guides can still be wrong or outdated, but their authority usually comes from some combination of reporting, local contributors, editorial research, and firsthand experience.

A generated itinerary can reproduce the shape of that authority without any visit having occurred.

That does not make every AI travel recommendation useless. It changes what kind of evidence the reader should expect.

Plausibility is cheap; local verification is not

A model can combine names of attractions, restaurants, transit routes, and neighborhoods into an itinerary that reads naturally. The weak point is often not grammar but freshness and physical reality.

A 2026 study in the Journal of Consumer Behaviour examined hallucinations in AI travel planning and specifically distinguished errors such as a nonexistent restaurant from factual inaccuracies such as bad opening hours. The researchers found that hallucinations reduced perceived accuracy and, through that, usefulness and trust. They also emphasized that a bad recommendation becomes much more salient when a traveler actually reaches a place and discovers that it is closed or does not exist. See the study on AI hallucinations in tourism.

That is exactly the problem with synthetic travel expertise. A sentence can be statistically plausible while the front door is permanently locked.

A useful generated guide needs visible grounding

A machine-assisted travel guide becomes more trustworthy when its claims can be traced to current sources: official attraction hours, transit agencies, hotel policies, restaurant websites, recent local reporting, reservation systems, maps, or clearly dated reviews.

It also helps to distinguish categories of claim. “The Louvre is in Paris” is a stable fact. “Go Tuesday morning because the line is usually short” is experiential and time-sensitive. “This neighborhood is safe and lively after midnight” is an even broader judgment that may depend on whose experience is being represented.

Those should not all be written with the same confidence.

A good automated itinerary can save time by organizing known information. It can even suggest combinations a traveler would not have considered. But it should not borrow the voice of the seasoned traveler while hiding that nobody actually stood on the corner, rode the bus, ate the meal, or found the locked gate.

Travel advice does not become firsthand knowledge merely because it is written in the first person.

Posted on

Template-generated sports and financial reports versus generative invention

An automatically written earnings report and a chatbot improvising a news story may both be called “AI journalism,” but technically they can be very different systems.

That distinction matters because the failure modes are different.

The Associated Press began automating large numbers of corporate earnings briefs more than a decade ago using structured data from Zacks Investment Research and software from Automated Insights. AP described a system in which company data flowed into a story structure created with its editors. The organization said automation let it expand from roughly 300 manually written earnings stories per quarter to thousands of short reports while journalists spent more time on analysis and original reporting. AP also labeled the automated stories and described the source data behind them. See AP’s explanation of its automated earnings reporting.

The important part is not that a machine typed the sentences. It is that the range of possible sentences was constrained by structured inputs and a known reporting template.

A template transforms data

A structured system might receive revenue, earnings per share, analyst expectations, and year-over-year changes. Rules decide which facts matter and how they are expressed. If revenue rises 12 percent, the software does not need to invent a reason. It can simply report the number.

That still requires quality control. Bad source data produces bad stories. A broken rule can mislabel a gain as a loss. A missing field can produce awkward output. But the relationship between source and sentence is relatively inspectable.

Open-ended generative systems add another layer of uncertainty. They can summarize supplied material, but they can also produce details not explicitly present in the source, blend background knowledge into the answer, or invent connective explanations that sound reasonable.

That is a different problem from templating.

Automation is not one category

A box score turned into a five-paragraph game recap is mostly a transformation problem. A system asked, “Why did the team lose?” is being asked to interpret. A financial template stating that revenue fell is different from a model explaining why revenue fell without reporting, interviews, or an earnings-call transcript.

The word “automated” hides those distinctions.

Readers do not need every publication to expose its software stack, but they do benefit from knowing what kind of process produced the article. Was it generated directly from structured data? Was a human editor involved? Were claims sourced from documents? Did the system infer explanations that nobody reported?

Routine automation can be extremely useful precisely because it is narrow. The problem begins when the authority of a data-driven report is carried over to prose that has much more freedom to invent.

A machine filling a template and a machine making up the next sentence are not the same thing.

Posted on

Machine-written product descriptions across vast retail catalogs

A catalog with twenty products can be written by hand. A catalog with twenty thousand products creates a different incentive.

Every item needs a title, description, features, materials, compatibility notes, dimensions, care instructions, and search-friendly language. That is exactly the kind of repetitive work automated text generation can accelerate. The danger begins when the system stops rephrasing supplied facts and starts filling gaps with plausible ones.

Shopify’s own documentation for its AI product-description feature, Shopify Magic, makes the distinction unusually clear. Merchants can provide a title and a few keywords, then generate a complete description. But Shopify also warns that generated copy can introduce product benefits or facts that the merchant never supplied, including details borrowed from similar products. Its guidance says merchants remain responsible for the accuracy of what they publish and should review generated text closely. See Shopify’s documentation on automatically generating product descriptions.

That warning gets more important as the catalog grows.

Plausible specifications are still invented specifications

A language model is very good at knowing what a product description usually sounds like. If the item is a jacket, the copy may naturally mention weather resistance. If it is a cable, the model may invent compatibility language. If it is a kitchen tool, it may add claims about dishwasher safety or materials because those details are common in similar listings.

The prose can sound more complete than the underlying record.

That creates a subtle reversal. Instead of the description being a readable version of verified product data, the description becomes a source of new claims that somebody now has to investigate after the fact.

The Federal Trade Commission’s general advertising guidance is boring but useful here: advertisers are responsible for express and implied claims, and material claims need a reasonable basis. Automation does not transfer that responsibility to the model.

The source record has to remain authoritative

The safest workflow is simple. Structured product data comes first: manufacturer specifications, measured dimensions, tested compatibility, ingredients, materials, warranty terms, and other facts that can be checked. Generated copy can then reorganize those facts into readable prose.

When the generated description adds something not present in the source record, it should be treated as an unverified suggestion, not as a discovered fact.

This matters because scale changes the consequences. One invented sentence on one listing is a correction. One invented attribute propagated across ten thousand SKUs becomes a catalog-level data problem.

Machine-written product descriptions are not inherently deceptive. They are a publishing tool. The problem begins when the smoothness of the language hides the difference between information supplied by the merchant and information guessed by the machine.

A huge catalog can be automated. Responsibility cannot.

Posted on

Automatically generated local-news pages without local reporting

A page can look local without anyone local having touched it.

Put a town name in the masthead, add weather, crime summaries, school-board keywords, sports results, and a few photographs, and the site begins to resemble a small newsroom. Search engines may surface individual pages to residents who never see enough of the publication to ask who actually works there.

Generative systems make that appearance cheap to reproduce across hundreds or thousands of places.

NewsGuard began documenting this broader phenomenon in 2023, when it identified 49 news and information sites that appeared to be mostly or entirely generated by AI. Its report, “Rise of the Newsbots”, described sites using generated material to imitate ordinary news publishing at content-farm scale. By 2024, NewsGuard was also warning about networks of sites presenting themselves as local news operations despite limited or opaque reporting infrastructure.

The important issue is not simply whether software wrote the sentences.

It is whether anybody reported the story.

Local reporting begins before the article exists

A reporter attending a council meeting can hear what was left out of the press release. They can ask a follow-up question, recognize a familiar name, call the school superintendent, check a court file, notice that the mayor avoided an answer, or learn from a resident that the official version of events makes no sense.

A generated summary can process the documents produced by that world. It does not automatically gather the missing evidence.

This is the difference between writing and reporting.

A system can rewrite a sheriff’s press release into smooth prose in seconds. It can summarize a city agenda or turn a sports box score into a readable article. Those uses may be perfectly legitimate if the source and automation are clear.

But a page assembled from upstream material should not be confused with an independent local newsroom merely because it contains the town’s name.

The byline should lead somewhere

Readers evaluating a local-news site can ask basic questions that remain surprisingly powerful.

Who owns the publication? Are reporters named? Do those reporters have histories of original work? Does the site link to source documents? Are local people quoted directly? Are corrections published? Is there a physical or institutional presence in the community? When a story makes an original factual claim, can you tell how the publication learned it?

None of these tests proves quality by itself. A tiny one-person newsroom may have almost no infrastructure. A strong local blogger may work from a kitchen table. Conversely, a polished corporate site can employ real reporters and still produce bad journalism.

The point is to look for evidence of newsgathering, not visual polish.

Synthetic news can inherit real reporting invisibly

There is another wrinkle. Generated local-news pages often summarize information originally gathered by somebody else: a newspaper, public broadcaster, government agency, sports reporter, or community member.

That means an automated site can appear productive while depending on a shrinking layer of humans doing the expensive work underneath it.

For Dead Internet Theory, this is a more useful concern than simply counting AI-written articles.

The web can look densely populated with local publications while the number of people actually attending meetings, making calls, checking records, and witnessing events declines.

A town can have more pages about itself and less journalism at the same time.