Posted on

Form-driven databases that cannot be preserved by following links

A web crawler starts at the entry page and follows every link it can reach. Content that exists only in answer to a search question falls outside that walk: the database has no link to offer, so it produces a page only when a human types a query.

Records that exist only when asked

In 2001 Michael Bergman used the term deep Web for material exposed through searchable databases and estimated at the time that it dwarfed the conventionally crawlable web. The exact scale estimate aged badly; the underlying mechanical point did not. A database may reveal records only after it receives a query.

Every search box is that kind of database. A crawler never sees what the box reveals, because the records have no addresses until someone asks for them.

The catalogue behind the form

The Haddon catalogue shows the cost. Built by Marcus Banks in Oxford with UK Economic and Social Research Council funding, it documented about 1,000 pre-war ethnographic films, searchable online from 1996.

The hardware aged. By 2005 the database engine no longer ran on current operating systems, funding had ended, and the catalogue went dark. The Oxford-led Gone Dark study documented the case: web archiving had not captured the searchable database, but the underlying data had been preserved offline, making later revival possible.

The same gap in the archives

The same study’s Kwetu case shows the pattern in web archives: the Wayback Machine holds the front pages and images, but “the search function does not work and no access to anything behind the search paywall is available.” The database behind Kwetu.net’s search portal, over one million manuscripts, stayed out of every crawl and lives on only in former owners’ hands.

Exports and documented queries

What preserves a form-driven database is not a crawl of its search page but an export of its contents, or ordinary links that give records stable addresses. The UK National Archives calls material reachable only through forms, pick lists, or search boxes not “machine reachable” and recommends static links or downloadable alternatives.

Documented queries capture a fraction. Each saved query is one row of a larger table; fifty useful questions preserve fifty answers, not the five thousand records behind the form. A database dump preserves everything, but it needs the operator’s cooperation, which a crawler never gets.

The form is a door, not the archive. What survives a database’s death is whatever was copied to disk while the door was still open.