Posted on

Language coverage and the discoverability of smaller linguistic communities

The web does not become equally searchable merely because two pages are both public.

Language changes the problem.

A search engine needs enough documents, links, query history, language understanding, spelling models, translation resources, and relevance evidence to retrieve useful material. Those resources are not distributed evenly among the world’s languages.

Google openly describes some of this unevenness. Its language-selection documentation says Search may return material from another language when it considers that useful, including cases where there is not enough information in the language used for the query. Google also offers translated search results, but that specific translation feature is currently available only for a defined list of languages.

That does not mean other languages are absent from Google’s index. It means different search and translation capabilities have different coverage.

Smaller languages have a retrieval problem as well as a publishing problem

A community may publish excellent local material and still be difficult for outsiders to find.

Academic work on cross-lingual information retrieval has documented the technical gap directly. A 2022 paper, Pivot Through English: Reliably Answering Multilingual Questions without Document Retrieval, notes that open-retrieval question answering in lower-resource languages has historically lagged behind English partly because non-English document retrieval itself is harder.

That distinction is important.

If a search in English produces many results while the equivalent search in a smaller language produces only a few, it does not follow that the smaller linguistic community has little to say. The difference may involve fewer indexed documents, weaker cross-language matching, different query terminology, fewer inbound links, less machine-readable material, or simply less training and evaluation data for the retrieval system.

English is a terrible universal control group

Researchers studying the apparent diversity of the web can accidentally measure English-language infrastructure and mistake it for the internet.

A fair comparison should consider how much material is actually published in each language, whether equivalent queries exist, whether the script is handled correctly, whether search systems recognize morphology and spelling variation, and whether translation is introducing or hiding results.

Cross-language search can also create a strange asymmetry. A locally important page may be obvious to native speakers using local terminology but practically invisible to an English speaker searching for an English translation of the same concept.

The page is alive. The community is alive. The search path between them and the outside observer is weak.

That is an important form of Algorithmic Reality.

A search engine can make the web feel overwhelmingly English, or overwhelmingly dominated by a few high-resource languages, without deleting a single page written anywhere else.

Sometimes the missing internet is speaking perfectly clearly.

The discovery system just does not speak back as well.