Posted on

Ancient Web: StarLing DB Built a CGI Laboratory for Comparing Languages

StarLing DB looks less like a modern website than a control panel for somebody trying to reverse-engineer human language.

Visit the StarLing database server

Its pages expose a collection of linguistic databases through old-fashioned CGI forms: text boxes, checkboxes, field selectors and query options. The interface is dense because the underlying data is dense.

A global typological database can be searched by language, alternate name, dialect, location, population, classification, dictionary and grammar sources, consonant system, syllable structure, tones, stress, noun number, noun classes, gender, demonstratives, pronouns, syntax, ergativity, noun incorporation and preposition/postposition behavior.

That is not a navigation menu. It is a research instrument.

The database still feels like the database

StarLing was developed around comparative and historical linguistics. Individual datasets include etymological material for language families as well as dictionaries and typological resources.

The Altaic etymology interface, for example, exposes fields for reconstructed Proto-Altaic forms, meanings and correspondences in Turkic, Mongolian, Tungusic-Manchu, Korean and Japanese.

Whether a particular long-range reconstruction is accepted by every linguist is a separate scholarly question. What matters from a Web-history perspective is that the site exposes the underlying records instead of flattening them into a narrative article.

The Global Lexicostatistical Database adds another layer. Its documentation explains that data are morphologically segmented to support manual and automated analysis, and that entries are distributed in multiple forms: searchable online data, print-ready PDFs and editable tables.

The older download page is especially evocative. Etymological databases appear as downloadable packages with names such as altaic.exe, ie.exe, cauc.exe, sintib.exe and drav.exe, many dated 2005. Dictionaries are distributed the same way.

This is the scientific Web before every database became an API product with an account tier.

The CGI server even displays counters for pages generated by its scripts. One current query page reports more than 2.4 million generated pages.

That number is not meaningful as traffic. It is meaningful as evidence of a site architecture that has been answering structured queries for a very long time.

For digital archaeologists, StarLing is useful because both the knowledge and the machinery remain visible. You can see the fields researchers cared about, the file formats they distributed, the database categories they used and the awkward but direct interface that connected a browser to the underlying records.

A redesign could make it prettier.

It would have to work hard not to make it less informative.

Explore the StarLing linguistic databases

Posted on

Ancient Web: Zhongwen.com Turned 4,000 Chinese Characters into Genealogy Trees

The CacheRat source list landed on a small Zhongwen.com page that transliterates English given names into Chinese. The rabbit hole behind it is much bigger.

Visit Zhongwen.com’s English-name page

Zhongwen.com was built by Rick Harbaugh around a project he called Chinese character genealogies: diagrams showing how characters are related through their component parts.

The site’s web edition dates to the 1990s, and Harbaugh’s Chinese Characters: A Genealogy and Dictionary was self-published in 1998 before publication was later taken over by Yale University Press.

The central idea is computationally elegant.

Traditional Chinese dictionaries often organize characters through a fixed list of section headings commonly called radicals. Harbaugh’s system instead uses computerized cross-referencing to connect characters through all of their meaningful components.

The site says more than 4,000 characters can be arranged into these graphical family trees.

A dictionary that behaves like a graph

Harbaugh’s explanation is that every component of a character can itself be treated as a character or historical component, allowing the reader to trace relationships backward toward a relatively small set of roots.

That makes the site useful in two ways.

First, a student who recognizes part of a complicated character has another route to finding it. Second, the structure makes relationships visible rather than reducing every character to an isolated dictionary entry.

The web edition also preserves classical reading material, character FAQs, etymology pages and other language resources. Some pages proudly recommend Netscape Navigator 3.0 or higher, which is the kind of sentence that now functions as its own timestamp.

The project is also unusually transparent about its limits.

Harbaugh notes that the genealogies are based mainly on traditional etymologies derived from ancient seal-script evidence and do not represent the final word in modern historical linguistics.

That caveat makes the project more useful, not less. It tells the reader what kind of map this is.

The little English-name transliteration page is still handy, but it is almost incidental compared with the deeper accomplishment: a 1990s website using hypertext and computer cross-referencing to make a writing system’s internal structure visible.

That is exactly the sort of thing early web enthusiasts were good at doing—building a specialized tool because the medium suddenly made it possible.

Explore Zhongwen.com starting from the original name-transliteration page