Posted on

Broken download mirrors for freeware and shareware

A freeware program used to live in several places at once. The author’s site held the current version; a university FTP server, a download portal, or a hobbyist’s mirror held copies. Shareware grew up on that redundancy. When one host died or ran out of bandwidth, another copy kept the program findable, and no single outage could silence it.

Mirrors break in two directions

A mirror can vanish, which is the visible failure. Domains lapse, university accounts get pruned, portals purge older versions during redesigns, and the download page everyone linked to starts returning 404s. The other failure is quieter: a mirror stays up and stops being a faithful copy.

In late 2011, Nmap developer Gordon Lyon, who writes as Fyodor, discovered that CNET’s Download.com was wrapping Nmap inside a portal-made installer that offered browser toolbars and changed browser settings. Lyon’s original warning to Nmap users argued that users were being handed CNET’s executable while believing they were downloading Nmap’s own installer. Krebs on Security independently tested the wrapper and found that several antivirus engines flagged CNET’s installer. Wireshark got similar treatment until its project director sent a cease-and-desist. CNET later apologized and said open-source packages would no longer be bundled. The episode showed why a mirror can remain online and still stop being a faithful copy of what an author released.

What a faithful copy looks like

Verification starts with the author. If the developer still runs a download page with checksums, compare the candidate file’s hash against the published value; that alone rejects a doctored copy. Official mirror lists, which shareware authors once maintained by hand, name the servers the developer vouches for. When the author has disappeared, Wayback Machine captures of the download page and the Internet Archive’s Software Library are the best surviving evidence of exactly which file was released.

That is also where recovery stops. A mirror records what it shelved, not what it was supposed to be. If the author’s site and release notes are gone, and nobody captured the official file alongside its hash, the portal copy is all that remains — proof that a program was hosted somewhere, not proof of what it was. Mirrors solved availability, never identity. When a copy is all that survives, certainty about what was actually released is the first casualty.

Posted on

Podcast episodes stranded behind dead media-hosting accounts

Podcasting has always been two things bolted together: an RSS feed that describes episodes, and a hosting account that stores the audio files the feed points to. The arrangement works invisibly until one half stops being paid for or gets switched off. Then the episodes in that account vanish the way a dead domain vanishes — except the show’s listings sometimes stick around, pointing at nothing.

In March 2024, true-crime host Kaigan Carrie’s show Evolving Prisons, hosted through Spotify for Podcasters, landed an Outstanding Indie Podcast nomination at the True Crime Awards. Less than 48 hours later her account was flagged for “suspicious payments activity” and, within five minutes, deleted. Because the RSS feed disappeared at the same time, Podnews reported that the show began to be dropped from other platforms, including Apple Podcasts. Spotify later restored the audio and, after Podnews followed up, reimbursed lost revenue and emailed former subscribers. What it could not restore was the audience. “I’ve lost every single person that was ‘following’ my podcast,” Carrie told Podnews — those followers, in every other app, were not coming back.

The feed is the show

Each item in a podcast feed carries a title, a description, a publish date, and an enclosure: a URL pointing at the audio file, which lives on the hosting service’s servers, not inside the feed. Directory apps like Apple Podcasts and Spotify mostly just fetch that feed and play what it points to.

The feed is therefore not the audio itself — it is the wiring that tells podcast apps what exists and where to fetch it. Break it and every app depending on it can fail at once. During a 2019 FeedPress outage, many podcast episodes became unreachable for hours because a shared RSS dependency went down. The outage was temporary; the dependency is structural.

What a dead account removes

When a hosting account dies, the damage is never just one file. A host bundles the audio, per-episode metadata, artwork, analytics, and feed. Simplecast’s help documentation warns that deleting a show permanently purges the podcast details, audio files, analytics, and RSS feed and recommends moving the feed and downloading files before deletion. Cancelled subscriptions, lapsed free accounts, and closed-down hosts take the whole bundle with them, leaving episode links to return nothing but 404s.

The cascade

That one deletion propagates. Downloads already queued in apps fail for months, old links across the web — in blog posts, archives, and show pages — all point at the same dead URL, and because nobody controls the expired account, nothing redirects anywhere. Episode listings may survive in the Wayback Machine or in a directory that polled the feed before it died, so the show can still look present while playing nothing.

What survives

Recovery starts with feed metadata. A captured RSS feed preserves the full episode index — titles, dates, descriptions, and the audio URLs — so the record of what existed and when is usually not the hard part. The hard part is the audio itself. It survives only where someone copied it: in a listener’s download folder, on a backup drive, or on archive.org, which hosts podcast episodes people upload deliberately. Tools that bulk-download podcast enclosures exist precisely because a feed can outlive the availability of the files it once listed.

That asymmetry is the useful lesson. Feed metadata rebuilds an index of what was lost, but an index is not an episode. The sound is gone unless someone, somewhere, pressed download.

Posted on

Video embeds whose original uploads no longer exist

An embedded video looks as though it lives on the page. Under a news article the play button sits inside the story, and nothing advertises that the moving pictures are stored somewhere else. The page holds only a thin player. It asks YouTube for a file by an eleven-character ID, and YouTube answers. That arrangement spares the page owner the storage and bandwidth, and it stays invisible until the answer stops coming.

An embed is a pointer, not a copy

An embed is a citation with a play button. The page stores no video; it points at youtube.com/embed/<ID>, and the player fetches everything else from servers you do not control. That makes the evidence fragile in a specific way. When the upload is deleted, set to private, blocked in a region, or removed along with the account that posted it, the embed’s request fails and the player reports “Video unavailable.” The surrounding article still renders perfectly. Nothing tells the author that one of its windows shut, and nothing tells the reader what the footage showed.

The upload disappears; the page stays

The case that shows this best is a Vice News documentary that was quietly retired. “Inside Saudi Crown Prince’s Ruthless Quest for Power” appeared on the Vice News YouTube channel on June 19, 2023 and drew more than three-quarters of a million views before Vice set it to private within four days, as The Intercept reported. The original link became unavailable, and embeds that depended on it stopped functioning. The surviving record came from reporting, metadata, and copies made while the video was still accessible—not from the embed itself. Without an independent copy, a dead player can leave behind little more than a title and an ID.

The problem is not exotic. Human Rights Watch rechecked the videos and photos it had linked to in its reports since 2007 and found that 619 of 5,396 items, roughly 11 percent, had vanished, frequently behind the literal words “Video unavailable.” Mnemonic, which preserves conflict footage, documents the same erosion at scale: about 21 percent of the 1.7 million YouTube videos in its Syrian Archive were already gone by mid-2020, and about 14 percent of the 444,000 in its Yemeni Archive. Even away from war, the churn is steady — of 105 million public videos that the Archive Team indexed in 2009–2010, 4 million were already deleted by August 2010. The uploads that die are often exactly the ones that matter most.

What survives, and what does not

Recovery works within narrow limits. A page capture can save the title, uploader, date, and description, while specialized crawls can sometimes save the video itself. Current Archive-It guidance says YouTube videos and embeds can be archived, but collection is intermittent and replay sometimes requires special handling. A transcript or caption track preserves what was said, never everything that was shown. Copies survive because someone fetched them on time — Mnemonic holds uploads that the platforms later took down, which is why that footage still exists in any form. Reuploads appear, but they carry no chain of custody, so they cannot prove what the original actually was.

An embed is a loan, and the lender answers only while it wants to. The deleted upload and the vanished description are gone regardless of how many pages quoted them. If a clip matters, download it, and with it the ID, the channel, the date, and a hash of the file; keep the transcript too. An embed is a convenience, not a record.

Posted on

Image-hosting failure and the destruction of illustrated forum tutorials

A tutorial is a stack of image links

A forum post that explains how to fix something rarely works as plain text. The words say “remove the trim, then the connector,” but the real instruction is in the photographs: which screw, which clip, which cable runs where. For years, the easiest way to attach those pictures was pasting a link to a free image host. The host absorbed the storage and bandwidth, the forum stayed cheap to run, and the tutorial’s illustrations sat on hardware nobody in the thread controlled.

The Imgur purge

That arrangement became visibly fragile in 2023. Imgur announced that, starting May 15, it would focus on removing old, unused, and inactive content that was not tied to an account. Importantly, Imgur clarified that this did not mean every anonymous upload would be deleted: an image had to be anonymous, unused, and unviewed before it would be considered under that rule. The narrower policy still exposed the underlying problem—an old tutorial could depend on an image host whose retention rules the forum never controlled.

The TinyPic precedent

It was not the first such funeral. TinyPic, another no-account image host, announced in 2019 that declining usage and advertising revenue made the free service unsustainable. Its shutdown notice, preserved by Archive Team, disabled third-party viewing in late August and shut the servers down in September after a short download window. Forums that had leaned on TinyPic watched old illustrated threads lose their pictures. A service can promise personal downloads; it cannot reattach a generation of link tags scattered across posts.

A repair you can’t follow

The loss is not nostalgia. Text remains searchable; the reader simply cannot act on it. A photo of a specific circuit trace or a strut mount tells the newcomer what “this one bolt” means; without it, the sentence is guesswork. Communities stay online and keep posting while their older half becomes a gallery of broken-image icons, present, searchable, and useless for the job at hand.

What a rescue actually recovers

There are partial salvages. Web archives may preserve pages while their images still load, and communities can make their own emergency copies. When Imgur announced its 2023 policy change, OpenStreetMap contributors documented a backup effort covering Imgur images referenced by their ecosystem. A local copy and a re-uploaded file can restore a tutorial somebody cared about.

But a pile of rescued images is not a tutorial. The Archive Team’s saved files have no owner to re-host them, and reattaching URLs means editing old posts one by one, which almost nobody does. Recovery works only when someone with a copy decides to rebuild the thread. Everything else is gone with the hosting company’s ledger. The cheap fix was never free; it was a loan, and the lender came to collect.

Posted on

Free ISP webspace as a disappearing publishing infrastructure

For most of the web’s first two decades you did not have to pay a hosting bill to publish. Internet providers and free hosts gave customers space under addresses built from somebody else’s domain: members.aol.com, home.att.net, geocities.com/~username. The personal page came bundled with the connection. That bargain turned ordinary households into publishers, and it quietly decided who owned the result — everyone stood on rented infrastructure, and the terms were set by a company, not by any rule of the web.

Publishing was bundled with the connection

AOL Hometown is the cleanest example. It bundled personal web hosting with AOL membership and gave nontechnical users point-and-click ways to publish pages under AOL-controlled addresses. Archive Team’s history of AOL Hometown records the service as a major hosted-web community that ultimately disappeared with little preservation. The result was a genuine publishing infrastructure: family pages, fan archives, memorials, hobby references, and personal histories, bound together by guestbooks, hit counters, and reciprocal links — exactly the parts archives are worst at keeping.

When the provider retires the hosting

AOL shut Hometown on October 31, 2008. A surviving copy of AOL’s September 2008 shutdown email told members that their content would no longer be available after that date and urged them to save it before deletion. Yahoo closed GeoCities the following year. What was lost in both cases was large: the pages themselves, the direct links, and everything dynamic — guestbook posts, counters, CGIs — that a static snapshot cannot hold.

What survives is partial. For GeoCities, the Internet Archive ran dedicated crawls before the shutdown, while volunteers also raced to mirror large parts of the service. AOL Hometown never received a comparable rescue and is logged by Archive Team as essentially lost.

A saved copy is not a working address

The limits of recovery show up in the address bar. An archived page can be displayed, but an old link still points at a host that no longer exists. GeoCities URLs, AOL Hometown addresses, and their incoming links from forums, signatures, directories, and search results all stop resolving at once. Nobody preserved the redirects, because the platforms never built them. Preservation stops short of putting the page back where readers expect to find it.

Keeping pages and links alive

A writer who wants both content and address to survive has to take on what the free hosting never promised. Keep the actual files — download them before the deadline. Move to hosting you control, ideally a domain you own. Where pages move, set redirects from old URLs, or leave a pointer, so the links built around the page can follow. That is the whole lesson of the webspace era: treat free publishing space as borrowed, keep the copy, and plan the move before the provider announces the closing.

Posted on

Personal web pages lost when university accounts expire

For decades, almost every college with a Unix server handed students and staff a homepage. Somewhere under a people. or users. subdomain sat a ~username directory that could hold anything publishable: a résumé, lecture notes, a lab’s working papers, a thesis in progress, a Web ring for a hobby. Blogs and small projects grew there because the hosting was free and the address looked respectable, which is why other pages linked to them and course lists cited them. The university paid for the bandwidth. The account holder treated the space as theirs. Neither side treated it as permanent, and that mismatch eventually shows up as a dead link.

The arrangement only worked while the affiliation lasted.

When the account expires

University accounts are scoped to enrollment or employment. When the affiliation ends, so can the web account. Dartmouth’s policy for expiring website accounts is unusually explicit: sites on its host service expire 60 days after the owner leaves unless they are transferred elsewhere or reassigned to a current Dartmouth member.

A documented shutdown shows what this looks like. In June 2024, Missouri State University announced that personal and course websites on people.missouristate.edu and courses.missouristate.edu would be discontinued on October 1, 2024. Sites and files left on the servers would be deleted; the announcement said no action was needed from anyone who no longer used a site, only that it would disappear. It pointed account holders at commercial hosts like GoDaddy and Bluehost, and Web Strategy noted it could not help anyone move a site it did not build. Course material was steered to the university’s learning-management system; official pages moved to its content platform. The personal pages had no landing spot. Today the old URLs return 404. Years of faculty lab pages, class notes, and student projects have vanished from the university’s own servers.

What survives and what does not

Recovery is partial and accidental. The Internet Archive still serves copies of some Missouri State pages captured by routine crawling years before the shutdown; a 2015 capture of one faculty research page is still readable today. But the Wayback Machine holds only what someone happened to crawl, frozen at the moment of capture. Unlinked pages, recent changes, and sites nobody indexed are absent. Whole sets of files — especially the PDFs and images such pages accumulated — survive only if a crawler happened to reach them. An archive copy is a snapshot, not a way to restore the account.

A few schools offer a real exit instead. Dartmouth lets departing account holders transfer a site to Reclaim Hosting or hand ownership to a current student, staff member, or faculty member. Its WordPress documentation also warns users to export while their NetID is still active, because the standard export does not contain every theme, plugin, or image file. These options share a condition: the account holder has to act before institutional access disappears. The default everywhere is deletion.

A personal page on a university account is rented space, not ownership. It sits on servers the author does not control, tied to a status the author does not keep. When the status lapses, the page lapses with it. Nobody has to decide to erase this part of web history; it is an operational consequence. The record of what the internet forgot grows one expired account at a time.

Posted on

Domain expiration and the loss of an author’s original address

For two years, dykewrite.com was a fixed point on the web. DykeWrite, an e-zine and webring for dyke and lesbian culture founded in 2002, published essays and pulled member sites together under one shared name. The Wayback Machine’s last capture of the original is dated June 13, 2004. Soon after, the page went blank and the domain was taken up by someone else. The address had stopped meaning what its writers had built into it.

A domain anchors what links point at

A domain looks like property; in practice it is a registration that must keep being renewed. Let it lapse far enough and another registrant can acquire the name. The new registrant has no claim on your old writing, but they control the address that other pages still associate with it.

The real value sits in the connections, not the page. Citations, webring listings, a blogroll entry, an email address in a bio — all point at the name. When the holder changes, the author keeps the writing and loses the location where the web agreed to find it. A reader who follows an old link lands on a parking page, an ad farm, or nothing.

What one lapsed name cost

DykeWrite is a small, documented case of exactly that. The revival that rebuilt the site on Neocities in 2022 records the history: the original launched a rebuilt community system on June 1, 2004, and the archival trail reaches only to mid-June before the page went blank and a domain squatter took the name over. The important fact is not who owns the domain now; it is that the original authors lost control of the address their readers knew.

The loss ran deeper than a homepage. That final system had two weeks of life before captures stopped, so the interactive parts of the site — the reason the rebuild happened — were never meaningfully crawled. The static pages survive as copies; the community had to move to a new name because the old one no longer pointed anywhere.

Content survives; the address does not

Archives keep the writing, not the location. The captured pages live under web.archive.org paths, wrapped in timestamps, as in this June 13, 2004 capture. The content is readable; the address is not the one anyone cited or linked to. The revival abandoned the old name, moved to Neocities, and states plainly that it is not affiliated with the original authors. The identity traveled; the original address could not.

Recovery is opportunistic in general. In 2009, McCown, Marshall, and Nelson published a study in Communications of the ACM that surveyed people who had lost websites and rebuilt them from the Internet Archive, Google, and search-engine caches. Reconstruction returned whatever had been crawled while the site was alive; dynamic sections and rarely visited pages often did not come back.

The address is the part that needs protection

The lesson is cheap to apply. Renew a domain on autopay and keep a calendar note anyway, hold a full copy of the content somewhere you control, and archive the site while it is live rather than after it lapses. An archive is a library of what a site said. The location where it said it belongs to whoever keeps paying the lease.

Posted on

Shortened URLs as a preservation dependency

A shortened URL is not a pointer to a page. It is a key to somebody else’s lookup table. The address carries no destination at all, only a code that tells the browser whom to ask. Whether an answer ever arrives depends on an intermediary you never chose: the shortening service must still exist, still resolve, and still hold the entry.

A second dependency you never see

goo.gl launched in December 2009 and spread wherever a link had to fit in a sentence: tweets, footnotes, printed books. In 2019 Google stopped accepting new URLs but kept resolving old ones, saying existing links would keep working. That promise had a shelf life. On July 18, 2024, Google announced that goo.gl URLs would stop redirecting after August 25, 2025. It later changed the plan. In an August 2025 update, Google said links that had shown activity in late 2024 would continue to work; links already marked for deactivation would stop redirecting.

Which goo.gl links still work today therefore depends on whether they were popular during one specific recent window. The losers return 404s. The address did not rot; the operator’s decision to keep answering did.

What the archives kept

Recovery depended on capturing the mappings while they still worked. By 2025 the Wayback Machine already held large numbers of goo.gl redirects, and Archive Team ran a dedicated goo.gl preservation project before the August shutdown date. An archived redirect can preserve the original destination even after Google’s live lookup stops answering.

The conditions are strict. A link nobody captured leaves no trace, and a 404 says little about what it once pointed to. After the shutdown goo.gl even returns two flavors of 404: a Google-branded page for codes that used to resolve, a Firebase-branded page for codes that never existed. That is a tombstone, not a map.

Recovery returns an address, not a page

Finding the original URL rebuilds only half of the link. Michael Nelson documented exactly that failure with goo.gl/0R8XX6, one of 26 shortened URLs in a 2017 survey of self-driving-car datasets. The short URL stopped redirecting in 2025, and its Carnegie Mellon destination no longer resolved either. Recovering the citation required archives of both the redirect and the destination page.

A shortened URL is a preservation dependency: part of the address lives on infrastructure you do not own. TinyURL has outlived many of its targets, and bit.ly may outlive more. The point is not that any one service will fail, but that several already have, and that each failure moves the cost of rescue onto archives and readers. When a link matters, shorten the convenience, not the citation: keep the full URL and the content itself, and treat the short code as a fetch address rather than a record of what exists.

Posted on

Content drift: working links whose evidence has changed

The link works; the evidence does not

A URL that still loads is no guarantee that the page behind it still supports what you linked. The request succeeds, the browser renders, and the fact you relied on is gone. Researchers call this content drift: a resource keeps its address while its content changes until it no longer represents what someone cited.

Broken links are honest; drift is not

A dead link announces its failure. A 404 or an expired domain tells you something is missing. Content drift never announces itself. The click succeeds, the page looks normal, and nothing warns you that the sentence you came to check was rewritten a day, a month, or a decade ago.

The stockpile page that changed overnight

One documented example comes from the early pandemic. The Strategic National Stockpile page described the stockpile as supporting “state, local, tribal, and territorial responders.” Archived versions from April 3 and April 4, 2020 show the wording changing immediately after public controversy over how the stockpile was supposed to serve states. A Los Alamos web-preservation report uses the two captures as a concrete example of content drift: the live URL survived while the policy language at that URL changed.

How often it happens

Drift is common. In 2016, Shawn Jones and colleagues published Scholarly Context Adrift, examining web references drawn from millions of scholarly articles. Among references for which a representative archived copy could be compared with the live resource, more than three quarters had changed.

Dated records are the workaround

That is why dated captures and version records matter. The Wayback Machine and the Memento protocol are built around the same basic need: reaching a version of a web resource as it existed at a particular time. Recovery has limits. Most references in that 2016 study had no usable archived copy near the publication date, and an archive only preserves what someone captured, whenever they captured it.

When a URL still works, the instinct is to trust it. The fix is a timestamp: save a copy the day you read it, note the date, and link to the capture alongside the address. Drift is silent by default; the way to hear it is to keep a record of what you saw.

Posted on

Link rot in the references of published research

Read a paper from a decade ago and you will eventually follow a footnote to the web. Sometimes the page is still there. Often it is not.

How research references rot

When a database, a government report, or a project page vanishes, every paper that cited it points into the void. That is link rot: the URL stops resolving and returns a 404 or an error. Its quieter cousin is content drift, where the address still works but the page now shows something different. A journalism site can keep the URL for an old article while silently replacing the material behind it — the link works, and the citation is still wrong. Researchers Martin Klein and Herbert Van de Sompel, with colleagues, grouped both problems under the single term reference rot.

Their 2014 study in PLOS ONE checked over one million references to web resources pulled from more than 3.5 million science, technology, and medicine articles published between 1997 and 2012. It found reference rot in roughly one in five articles overall; among articles containing web references, the share rose to about seven in ten.

The law decays the same way. Jonathan Zittrain, Kendra Albert, and Lawrence Lessig studied legal citations and found that more than 70 percent of URLs in the law-journal sample and 50 percent of URLs in U.S. Supreme Court opinions no longer pointed to the originally cited information.

Why missing evidence matters

A citation has a job: let a later reader find the material a claim rests on and check whether the claim is fair. Break that path and the paragraph stands on an assertion no one can inspect. Replication becomes guesswork, and a researcher who wants to reuse the dataset linked in a methods section finds it went down with the lab page that hosted it.

The subtler damage comes from pages that appear healthy. Automated checks catch plain 404s, but a page returning a normal status code can still be a custom error page, a redirect to a homepage, or a rewritten article that no longer contains the cited fact. A dead link at least announces its failure; drift does not. It quietly converts a reproducible citation into an unsupported one, in a legal opinion as readily as in a scientific paper.

What persistent identifiers and archives can preserve

Two tools stretch the shelf life of a citation. The first is a persistent identifier such as a DOI or a handle. It keeps a stable name for a resource that can change address. But an identifier is a name, not a copy. If the publisher’s servers disappear, the DOI still resolves — to an error.

The second tool is an archived copy. Web archives such as the Internet Archive’s Wayback Machine keep snapshots later readers can retrieve; the Memento protocol exists specifically to find a snapshot from around the date a paper appeared. The 2014 report behind Perma.cc came out of Harvard’s legal citation study: instead of hoping a cited page stays up, authors archive it up front and link to a preserved copy held across a distributed network of libraries. Journals increasingly archive their own supplementary material, easiest to capture while the author’s browser is still pointing at it.

These measures do not make citations immortal. Archives choose what to capture and can miss heavily dynamic or login-gated pages. A snapshot may catch the page as it briefly looked, not as the author actually saw it. An archive is only as permanent as its funding.

What survives is decided early. Citing a DOI, archiving a copy while the page is still live, and preferring sources with existing archival coverage all cost little at writing time. The alternative is a footnote that spends its life pointing at a hole. The internet forgets on a schedule; citations that remember to archive can leave a record that outlasts the page they cite.