Posted on

Public records lost through database replacement rather than deliberate deletion

Not every missing public record was deleted. Sometimes it fell between two databases.

That distinction is less dramatic than a purge and often more useful. Government systems are replaced for the same reasons any large information system is replaced: aging software, rising maintenance costs, security requirements, vendor changes, new workflows, and years of customizations that have turned the original design into something nobody particularly wants to touch anymore.

The danger appears during migration. A new system may not have the same fields, identifiers, attachments, search behavior, or public interface as the old one. Data can survive internally while becoming much harder for the public to locate.

The 2023 retirement of FOIAonline provides a documented example. FOIAonline had served as a shared Freedom of Information Act request and records platform for multiple federal agencies. EPA decided to decommission it on September 30, 2023, and participating agencies moved toward replacement case-management systems.

The National Archives’ Office of Government Information Services documented the transition in detail. Its FOIAonline decommissioning notice said partner-agency data would be migrated to replacement systems, while public access to the FOIAonline application itself would end on September 30, 2023. In a March briefing, EPA officials urged registered users to download data they wanted readily available and acknowledged that some agencies could experience temporary service gaps while open cases moved.

Migration changes more than storage

A public database is partly the records and partly the machinery for finding them.

Suppose the old system exposes request numbers, agency names, dates, dispositions, attachments, requester correspondence, and full-text search. A replacement might retain the official case record while exposing fewer fields publicly. Another system might import the metadata but not preserve old attachment URLs. Search filters can change. Stable identifiers can be replaced. Previously public material can require a different route to obtain.

None of that necessarily means the underlying record was intentionally destroyed.

FOIAonline makes the point neatly because the data and the public interface had different afterlives. A 2024 FOIA Advisory Committee discussion noted that FOIAonline had served as a centralized repository containing eleven years of released records and that public-interest groups had copied material before shutdown; MuckRock and the Project on Government Oversight preserved nearly 34,000 documents. Meanwhile, agencies migrated their own case data into separate replacement systems.

Compare before calling something lost

A careful investigation starts with exports or snapshots of the old system. Record counts can be compared. Fields can be mapped. Attachments can be sampled. Old identifiers can be searched in the replacement. Web archives can show what the public interface exposed before shutdown.

The result may reveal several categories rather than one: records successfully migrated, records retained but exposed through a different search interface, fields or identifiers that no longer map cleanly, attachments moved to new URLs, and genuinely missing material. EPA, for example, now says its replacement FOIAXpress reading room contains 1.8 million records previously released through FOIAonline. That is evidence of substantial continuity, not evidence that every old public record vanished.

That is less satisfying than a single accusation, but it is how information systems actually fail.

Database replacement can create access loss without anyone pressing a delete button. Sometimes the records migrate successfully while the old centralized search, identifiers, links, or context disappear. FOIAonline is useful precisely because its shutdown produced a mixed result rather than a melodramatic one: a dead public platform, migrated agency data, recreated repositories, and independently rescued documents.

Posted on

Local newspaper closures and inaccessible reporting archives

When a local newspaper dies, people usually talk about what the town will lose tomorrow: city-hall coverage, school-board meetings, obituaries, court reporting, sports, and the thousand small facts nobody else bothers to write down.

There is another problem hiding behind the shutdown. The newspaper may also have been the only practical doorway into what happened yesterday.

A century of reporting can exist in several incompatible forms at once: bound paper, microfilm, a proprietary article database, newsroom photographs, PDF page images, a website CMS, and born-digital material that never appeared in print. Closing the publication does not automatically move all of that into a public archive.

The Rocky Mountain News shows both the danger and what a serious rescue can accomplish. The Denver paper closed in 2009 after nearly 150 years. The Denver Public Library acquired major parts of the Rocky’s archive, including clipping files, photographs, nearly 300,000 born-digital images, and the newspaper’s full run on microfilm.

That is a preservation success, but it was not automatic.

A newspaper website is only one layer

A closed newspaper’s old domain may remain online, redirect elsewhere, become an aggregator, or disappear. Even if the website survives, it may not contain everything the newsroom published.

The library’s separate Rocky Mountain News Digital Collections now expose more than 300,000 born-digital photographs from the 1990s through 2009. Access to the newspaper’s history nevertheless remains split among digital images, microfilm, bound issues, clipping files, negatives, and other archival systems. There is no single magic folder called the newspaper.

That fragmentation matters to researchers. A dead article URL may be recoverable through microfilm but no longer searchable on the open web. A photograph may survive without its original caption page. A commercial database may contain full text but require a library card or subscription. A clipping file can preserve local context that search engines never indexed.

Libraries can preserve what ownership changes break

Local libraries, historical societies, universities, and state newspaper programs often become the long-term custodians because their incentives are different from those of a failing publisher. Their job is not to make next quarter’s archive pay for itself. It is to keep the record usable.

That work still has limits. Copyright can restrict online access. Digitizing millions of pages takes money and years. Optical character recognition is imperfect. Born-digital interactive projects may be much harder to preserve than printed stories.

The Library of Congress treats that born-digital layer as a preservation problem in its own right. Its newspaper division web-archives news sites specifically to preserve ephemeral born-digital material, while noting that technical and resource limits prevent perfect capture. Once a newsroom and its servers are gone, the web-only parts may have no paper equivalent waiting underneath.

A newspaper closure therefore creates two deadlines. One concerns who will report the next council meeting. The other concerns whether anyone will still be able to find the council meeting from twenty years ago.

Posted on

Government website transitions and missing historical publications

Government websites change for ordinary reasons: a new administration arrives, an agency reorganizes, a contractor replaces the CMS, or an office decides that twelve years of nested directories would look better as one enormous search box.

The historical consequence can be less ordinary. Reports that were once directly linked may move to a new repository, receive new filenames, lose old landing pages, or disappear from navigation entirely. From the outside, relocation and deletion can look exactly the same: yesterday’s URL returns 404.

That problem is serious enough that preservation institutions created the End of Term Web Archive. The project began in 2008 and has collected U.S. federal websites around the 2008, 2012, 2016, 2020, and 2024 presidential transitions. The archive exists because an administrative transition is also a web transition.

A missing URL is only the first clue

When a government publication vanishes from its old address, the first task is not to declare it erased. Agencies often move documents between content systems, publications databases, records repositories, and redesigned websites.

Useful checks include searching the exact title, report number, filename, quoted phrases, agency publication catalog, and the current domain. A PDF may survive under a completely different path. An agency may have moved older material into an archival section without keeping redirects from the original URLs.

The White House provides an unusually clear example because the domain itself passes to the incoming administration. NARA’s archived presidential websites page preserves successive administrations separately and warns that archived sites are frozen historical records whose broken internal or external links are not repaired. A publication can therefore survive in the presidential archive even though the current whitehouse.gov no longer exposes it at the old address.

Archives help answer what changed

The End of Term collections are especially useful because they capture websites before and during known periods of change. They include HTML, images, PDFs, spreadsheets, multimedia, and other files stored in web-archive formats. NARA’s broader web-records guidance likewise treats long-term preservation of federal web content as part of maintaining the historical record.

That still does not prove every missing file was deliberately removed. Web crawlers miss things. Some databases require forms or scripts that are difficult to capture. Files may have been outside the crawl scope, blocked, generated dynamically, or linked only from pages the crawler never reached.

This distinction matters. “The current website no longer exposes this document at its old URL” is an observable fact. “The government deleted the document from existence” is a much stronger claim and often a false one.

Government web transitions create real historical gaps, but careful preservation lets us separate redesign damage from actual disappearance. Without that comparison layer, a routine CMS migration can look like a purge, while a genuine loss can hide behind the bland language of a website refresh.

Posted on

Vanished API documentation for discontinued products

An old program can contain a perfectly readable API call that nobody can confidently explain anymore.

That is what makes vanished developer documentation different from ordinary link rot. The code may survive. The function name may survive. Somebody may even have a working binary. What disappears is the contract around the interface: accepted parameters, return values, authentication rules, rate limits, version differences, error behavior, and examples showing how the designers expected it to be used.

API documentation can disappear by policy, not accident. Square’s API lifecycle documentation says retired APIs are removed from current developer documentation on or after retirement, although older API-reference versions may still retain information. That is a rational maintenance policy for a live developer portal and a preservation problem for anyone trying to understand old integrations years later.

Code does not contain the whole interface

Suppose an old application calls a method named people.get. The source can show the endpoint and the fields the programmer happened to request. It may not tell you which fields were optional, which scopes authorized them, what happened when a profile was private, whether pagination existed, or which behavior changed between API versions.

Examples are equally important. Good API documentation records idioms that are not obvious from a method signature. It may warn that a value is deprecated, explain an ordering guarantee, define a timestamp format, or show the correct sequence for a multi-step operation.

A graceful retirement shows what should survive. When Wikimedia shut down its experimental API Portal in June 2026, it did not simply erase the documentation. The shutdown announcement said documentation would move to other Wikimedia technical-documentation sites while old API routes were deprecated gradually.

Version context is part of the archive

Wikimedia’s historical API Portal record goes further: it points readers to Wayback Machine snapshots and a downloadable database archive of the retired documentation wiki. That preserves not just today’s replacement instructions but evidence of how the old interface was described while it existed.

Useful preservation should include reference pages, guides, sample code, changelogs, migration notices, schema definitions, SDK documentation, and dates. The date matters because API documentation is often continuously edited. A capture from one version may describe behavior that was false in another.

The Wayback Machine can rescue documentation after a vendor removes it, but dynamic documentation sites can replay badly. Search boxes, JavaScript navigation, generated code samples, or version selectors may not survive a simple crawl. Saving static exports, schemas, SDK packages, example repositories, and migration notices alongside web captures gives future readers more than one route back in.

There is also a limit to what documentation can recover. Once the actual service is gone, a reference manual cannot reproduce undocumented server behavior, data, or authentication infrastructure. It can explain the machine, not resurrect it.

Still, that explanation matters. When API documentation disappears, software history loses the difference between “this old code is strange” and “this old code was correctly following a contract that no longer exists.”

Posted on

Lost firmware files and the repairability of discontinued hardware

A discontinued router can have perfectly good flash memory, power regulation, radios, and Ethernet ports and still become difficult to repair because one small binary file vanished from a support site.

Firmware sits in an awkward place between hardware and software. It is software, so manufacturers can remove it from a website with no visible change to the physical product. But it is also part of the device’s operating state. A repair may require exactly the right image for a model, hardware revision, region, or carrier variant.

That makes old firmware surprisingly important.

A 2024 thread on Sierra Wireless’s own community forum gives a mundane example. An owner of an EM7565 modem had upgraded to newer firmware and then wanted to restore SWI9X50C_01.08.04.00, which he reported was no longer available through the normal download route. That was not an isolated misunderstanding: in an earlier thread, a Sierra Wireless representative explained that the public source carried only the current or most recent firmware packages and directed requests for older approved versions through commercial channels.

That is not exotic data loss. It is ordinary product support aging out.

The exact version can matter

Firmware is rarely interchangeable merely because the device name matches. Vendors release different builds for hardware revisions, countries, carriers, flash layouts, or radio configurations. A technically valid image for the wrong variant can fail, refuse to install, disable features, or in the worst case leave a device unable to boot normally.

Manufacturers also have legitimate reasons to stop distributing a particular build: security flaws, component changes, carrier certification, or defects may make an old image unsuitable for newly manufactured units. That is exactly why an archive needs context rather than treating every recovered binary as safe firmware.

A useful firmware archive needs provenance

Saving a .bin file is only the beginning. An archive should also record the manufacturer, exact model, hardware revision, region, version number, release date, original filename, checksums, and preferably the vendor’s release notes and flashing instructions.

Checksums help establish that a file has not changed. Provenance matters because a random firmware image from a forum attachment may be authentic, modified, corrupted, or intended for a subtly different device. Without that surrounding information, preservation can turn into roulette with a soldered flash chip.

There are good counterexamples. D-Link still exposes directories containing older DIR-300 firmware releases, making multiple historical versions independently downloadable instead of presenting only the newest build.

Discontinued hardware does not automatically become useless. Sometimes the missing service manual, driver, or firmware image is the only thing separating repairable equipment from e-waste. The hardware may survive twenty years. Its recovery file can disappear on an ordinary Tuesday.

Posted on

Package removal and the reproducibility of old software projects

An old software project can survive in GitHub, compile nowhere, and still look perfectly preserved to anyone browsing the source.

The missing piece is often a dependency.

Modern projects rarely contain every component they need. Instead, they name packages hosted elsewhere and ask a package manager to download the exact versions during installation or build time. If one of those versions disappears, the project may become unreproducible even though its own repository is untouched.

The most famous demonstration arrived in March 2016, when a developer unpublished a collection of packages from npm, including the tiny left-pad module. npm’s own postmortem, “kik, left-pad, and npm”, acknowledged that unrestricted unpublishing had allowed one removal to disrupt dependent software and promised policy changes.

A lockfile can describe a missing thing

Lockfiles help reproducibility by recording precise dependency versions instead of accepting whatever happens to be current. That is useful, but a lockfile is not an archive.

If package-lock.json, yarn.lock, Cargo.lock, or another lockfile says a project needs version 1.2.3 of a package, the build is only reproducible if version 1.2.3 can still be obtained. The lockfile can tell you exactly which brick is missing. It cannot manufacture the brick.

The left-pad event made that distinction painfully obvious. A vast dependency graph had been built on the assumption that registry entries would remain fetchable. When one small package vanished, the failure propagated far beyond the original author.

npm responded by tightening removal rules. Its follow-up on changes to the unpublish policy made dependency impact part of the decision. The current npm unpublish policy continues that approach: older packages can be removed only under limited conditions, including having no dependents in the public registry.

Reproducibility requires preserving dependencies too

There are several ways to reduce the risk. Organizations can maintain registry mirrors or caches. Projects can vendor critical dependencies into their own source tree when licensing permits. Build systems can use artifact repositories that retain the exact packages used for releases. Preservation projects can archive both source repositories and the package ecosystems around them.

Even that is not perfect. Old dependencies may require obsolete compilers, operating systems, certificate chains, or package-manager behavior. A package file may survive while the service needed to resolve it does not.

That is why preserving software is harder than saving source code. The program may be a small island connected by hundreds of bridges to other projects. If enough of those bridges disappear, the island is still visible. You just cannot get there anymore.

Posted on

Code-hosting shutdowns and the survival of repository history

A dead code-hosting site can leave behind a perfectly usable copy of the software and still lose half the story.

The obvious thing to save is the source tree: the files that compile into the program. But a repository can also contain years of commits, branches, tags, issue reports, wiki pages, release downloads, contributor names, timestamps, and discussions explaining why some ugly-looking line of code exists in the first place. That surrounding history is often what makes an old project understandable.

Google Code is a useful example because its shutdown was unusually well preserved. The hosting service was turned down in early 2016, but Google left behind the Google Code Archive, a read-only collection that reports more than 1.4 million projects, 1.5 million downloads, and 12.6 million issues.

That is the good version of a shutdown.

A repository is a timeline, not a folder

If all you preserve is the latest source snapshot, you can still read the program. What you cannot do is reliably reconstruct how it got there.

A Git, Mercurial, or Subversion repository may reveal when a feature appeared, who changed it, which release branch contained a fix, and whether a suspicious-looking file was temporary or intentional. Tags can connect code to a published release. Commit messages can explain design choices that were never documented anywhere else.

Then there are the things version control does not normally contain. Bug trackers may include reproduction steps and hardware details. Wiki pages may be the only installation manual. Release downloads can contain compiled binaries, sample data, or tools that were never checked into source control.

Google’s archive schema makes the separation unusually explicit: project metadata, wiki listings, issue summaries, commit summaries, downloads, individual issues, and source archives are stored as different objects. That is the preservation lesson. Copying the repository alone does not automatically save everything the hosting platform knew about the project.

Migration needs more than git clone

For Git projects, a mirror clone can preserve branches, tags, and repository objects. It does not preserve GitHub issues, pull requests, project boards, release notes, or externally hosted downloads unless those are exported separately. Other hosting systems have their own combinations of repository data and platform-specific records.

A complete rescue therefore needs an inventory. Save the repository. Export issue trackers and wikis. Download release artifacts. Preserve project descriptions and links. Record the original URL structure when possible so old references can be mapped to the surviving copy.

Even Google Code’s unusually thorough archive has limits. It is read-only, and resurrecting a project elsewhere still requires someone to interpret the preserved pieces and rebuild a working development environment.

When a code host closes, the question is not merely whether the source survived. The better question is how much of the project’s memory survived with it.

Posted on

CMS replacements that discard old revision histories

A content-management migration can look perfect from the front page and still destroy part of the historical record.

The reason is simple: the version you see today is not always the only version that ever mattered. A long-running page may have years of revisions behind it — corrections, changed claims, deleted paragraphs, updated links, different authors, and revision notes explaining why something changed. When a site moves from one CMS to another, those earlier states are easy to leave behind because the migration team is usually focused on getting the current page online.

Drupal makes the problem unusually visible because revisions are explicit records rather than invisible editor state. Drupal’s revision documentation shows that revisions can carry timestamps, authors, log messages, and complete earlier content states. Its migration system separately documents migration of node revisions, including the need to map revision identifiers. Moving the current article and moving its history are therefore separate jobs.

That distinction matters outside Drupal too. A replacement CMS may import title, body, author, and publication date while ignoring the old system’s revision tables. The visible site survives. The evidence of how it became that site does not.

What revision history can tell us

Revision history is useful when researchers need to answer questions the current page cannot answer. When was a statement added? Was a number corrected after publication? Did a product page once promise a feature that later disappeared? Did an agency quietly rewrite guidance rather than publish a new document?

Without revisions, the surviving page becomes a flattened final state. It is still useful, but it has lost chronology.

That history matters because reverting a Drupal page does not erase the later revisions; the system creates another revision while retaining the older record. A migration that flattens everything to one current page throws away chronology the source CMS intentionally preserved.

A proper export needs more than HTML

Preserving revision history requires exporting the relationships between the versions as well as the text itself. Useful migration data can include revision IDs, timestamps, authorship, status, moderation state, log messages, attachments, and which revision was current at a given time.

If the old system is still running, this is mostly an engineering problem. Once the database has been retired, overwritten, or discarded, recovery becomes much harder. A Wayback Machine capture may preserve several public versions of a page, but it will not recreate every internal revision the CMS once knew about.

That is the quiet danger of CMS replacement. Nothing has to look broken. The new site can be faster, cleaner, and technically successful while an invisible layer of its history disappears during launch weekend.

Posted on

Forum migrations that break post identifiers and citation links

A forum thread’s address is its identity: how a search result returns you to the right post, how an aging guide still links to the answer, how a footnote points at a documented dispute. When a community migrates to new forum software, that address becomes a question of housekeeping — and communities frequently answer it badly. The posts survive the move. The boring little URL in front of them often does not.

A post’s address is an identifier

Forums identify threads and posts with internal IDs, but every platform exposes those IDs through a different URL scheme. A migration preserves continuity only if the old addresses can still be mapped to the new records. Discourse’s migration documentation treats this as a first-class problem: its permalink table maps old URL paths to new destinations and answers with permanent redirects.

Even with careful import tooling, identifiers can drift across multiple generations of forum software. A XenForo migration case documents exactly that problem: a forum that had passed through vBulletin 3, 4, and 5 needed to preserve mappings among several generations of IDs because older links otherwise landed on the wrong threads.

What broken references cost, and how continuity survives

A failed permalink costs more than one reader’s bad click. The forum loses its accumulated citations: threads quoted in reviews, linked from repair guides, or referenced in research. Search engines and bookmarks keep asking for addresses the live site no longer understands. Historical web-archive captures are a different matter: if the old URL was captured before the migration, that archived version can still remain available under its original address even when the live site now returns an error.

The fix is boring but effective: migrate the data and keep the addresses. Standard practice writes a mapping from old URLs to new ones and answers with a 301 for content that exists and a 410 for threads deleted along the way, so search engines treat both as permanent. Posts’ internal links get rewritten to the new addresses so they no longer depend on the old domain resolving. Recovery has limits. Rebuilding one post’s view of a thread requires the full URL map, and deleted content stops at the 410 — a clear marker that something existed, not a copy of it.

A forum is not a pile of posts; it is a pile of posts and the addresses people use to return to them. A migration that discards its old addresses turns a decade of discussion into a string of 404s while the content sits a few keystrokes away, unreachable. Keeping the numbers around is usually cheap. Forgetting them is permanent.