Posted on

Ancient Web: StarLing DB Built a CGI Laboratory for Comparing Languages

StarLing DB looks less like a modern website than a control panel for somebody trying to reverse-engineer human language.

Visit the StarLing database server

Its pages expose a collection of linguistic databases through old-fashioned CGI forms: text boxes, checkboxes, field selectors and query options. The interface is dense because the underlying data is dense.

A global typological database can be searched by language, alternate name, dialect, location, population, classification, dictionary and grammar sources, consonant system, syllable structure, tones, stress, noun number, noun classes, gender, demonstratives, pronouns, syntax, ergativity, noun incorporation and preposition/postposition behavior.

That is not a navigation menu. It is a research instrument.

The database still feels like the database

StarLing was developed around comparative and historical linguistics. Individual datasets include etymological material for language families as well as dictionaries and typological resources.

The Altaic etymology interface, for example, exposes fields for reconstructed Proto-Altaic forms, meanings and correspondences in Turkic, Mongolian, Tungusic-Manchu, Korean and Japanese.

Whether a particular long-range reconstruction is accepted by every linguist is a separate scholarly question. What matters from a Web-history perspective is that the site exposes the underlying records instead of flattening them into a narrative article.

The Global Lexicostatistical Database adds another layer. Its documentation explains that data are morphologically segmented to support manual and automated analysis, and that entries are distributed in multiple forms: searchable online data, print-ready PDFs and editable tables.

The older download page is especially evocative. Etymological databases appear as downloadable packages with names such as altaic.exe, ie.exe, cauc.exe, sintib.exe and drav.exe, many dated 2005. Dictionaries are distributed the same way.

This is the scientific Web before every database became an API product with an account tier.

The CGI server even displays counters for pages generated by its scripts. One current query page reports more than 2.4 million generated pages.

That number is not meaningful as traffic. It is meaningful as evidence of a site architecture that has been answering structured queries for a very long time.

For digital archaeologists, StarLing is useful because both the knowledge and the machinery remain visible. You can see the fields researchers cared about, the file formats they distributed, the database categories they used and the awkward but direct interface that connected a browser to the underlying records.

A redesign could make it prettier.

It would have to work hard not to make it less informative.

Explore the StarLing linguistic databases