Posted on

Ancient Web: Zompist Tried to Prove English Spelling Has Rules

English spelling has spent centuries earning its reputation as a practical joke, and Mark Rosenfelder’s old Zompist page responds with the deeply Internet-era reaction: fine, let’s model the damn thing.

Visit Zompist’s English spelling page

The page, dated 2000, is titled “Hou tu pranownse Inglish.” Its argument is not that English spelling is elegant. It is that the system contains far more regularity than people usually give it credit for.

Rosenfelder lays out sound values, dialect assumptions, letter combinations, vowel rules, consonant rules, exceptions, historical leftovers, and the famous troublemakers such as gh and ough.

He also takes a swing at George Bernard Shaw’s famous “ghoti” joke, arguing that the proposed pronunciation only works by ignoring where English spelling rules actually permit those letter values.

The page then does something that makes it especially interesting as an old Web artifact: it turns the argument into a computational test.

Make the computer pronounce it

Rosenfelder assembled a sample lexicon of more than 5,000 English words and a set of rules for his Sound Change Applier. According to the page, the rules generated 59 percent of pronunciations perfectly and 85 percent either perfectly or with only relatively minor errors.

That does not make English spelling sane.

It does demonstrate that “English spelling is completely random” is not a particularly useful explanation.

The page is also full of old-browser archaeology. It discusses Unicode support as something the reader’s browser may or may not handle correctly and explains notation choices partly in terms of what HTML could reliably display.

That mixture is classic early specialist Web publishing: a serious subject, a personal voice, hand-built technical notation, downloadable data, and the assumption that an interested reader will happily scroll for a very long time.

Modern search results tend to chop questions like this into isolated answers: why is ough weird, why is knight spelled that way, why does c have two sounds?

Rosenfelder tried to put the machinery in one place.

CacheRat’s 1,967 Ancient Web Domains research list includes pages like this because old personal sites often preserve the full argument instead of optimizing each paragraph into a separate answer box.

English spelling is still a mess.

At least this mess comes with documentation.

Read the full spelling system

Posted on

Ancient Web: Gernot Katzer’s Spice Pages Index 10,000 Spice Names in Dozens of Languages

A spice website becomes something else entirely when the alphabetic index passes ten thousand names.

Visit Gernot Katzer’s Spice Pages

Gernot Katzer’s site documents 117 spice plants, with emphasis on ethnic cuisines, especially in Asia. Individual entries mix culinary use with history, chemistry, botany, photographs and the etymology of plant and spice names.

Then the indexing system goes completely off the rails in the best possible way.

The site says its large alphabetic index contains more than 10,000 spice names in roughly eighty languages. Separate indices handle Greek, Cyrillic, Hebrew, Arabic, Indic scripts, Chinese and Japanese characters, Thai and Lao, Vietnamese, Tibetan, Korean, Georgian, Armenian and other writing systems.

A food site built like a reference database

The pages can be browsed by English name, botanical family, geography, plant part or spice mixture.

That means cumin is not merely a recipe ingredient. It becomes a plant, a chemical profile, a name with linguistic history, a regional cooking tradition and a node in a much larger taxonomic system.

The site covers common kitchen staples and less familiar material: ajwain, cubeb pepper, fingerroot, grains of paradise, long pepper, pandanus, silphion, Sichuan pepper, zedoary and dozens more.

Katzer also ties the project to travel. His pages reference journeys through India, Sri Lanka, Bangladesh and Nepal and the foods encountered there.

This is exactly the sort of specialist information architecture that search engines benefit from but rarely reproduce. A generic article can explain what cardamom tastes like. Katzer’s site lets somebody approach cardamom from botany, language, geography or cuisine and keep moving sideways through connected subjects.

This site was found in CacheRat’s broader Ancient Web research corpus.

Explore the 1,967 Ancient Web Domains research list

The Spice Pages are what happens when a kitchen question gets answered by someone who also wants the Latin name, the Sanskrit name, the molecule and the migration route.

Return to Gernot Katzer’s Spice Pages