Posted on

Domain concentration across ordinary search results

The web can contain a million pages while a search result shows you ten.

That compression is unavoidable. Search engines have to rank. The interesting question is which domains keep surviving the compression.

If a handful of large sites appear repeatedly across ordinary queries, users can experience the web as far more concentrated than the underlying collection of websites actually is.

That is where Algorithmic Reality — The Internet You Are Allowed to See begins.

Concentration is visible across many queries, not just one

One search results page is a weak sample.

A query for a specific company should reasonably return that company’s site several times. A technical query may be dominated by the official documentation. A breaking-news query may favor a small group of publishers because they have current reporting.

The more useful test is horizontal: run many queries in a defined category, record the domains that appear, then ask how much of the available result space is occupied by the same publishers.

A 2024 audit of Google Search news results across Brazil, the United Kingdom, and the United States analyzed more than 220,000 results and reported substantial concentration among a limited group of outlets. The researchers specifically used concentration measures such as the Herfindahl-Hirschman Index and Gini coefficient rather than judging diversity from a few screenshots. See Auditing Google’s Search Algorithm: Measuring News Diversity Across Brazil, the UK, and the US.

Google itself has acknowledged the problem category for years. In a 2012 search-quality update it described a change called Domain Crowding intended to surface a more diverse set of domains when too many results came from the same site. Search systems have continued to use site-diversity mechanisms since then.

The existence of such mechanisms tells us something important: relevance and source diversity are not automatically the same objective.

Concentration is not automatically bad

A domain appearing repeatedly may deserve to appear.

Official documentation can be better than ten scraped copies. A specialist site may dominate a narrow subject because it genuinely has the strongest material. A local query may reasonably favor a few authoritative local sources.

That is why domain concentration alone cannot measure search quality.

It also cannot tell us how diverse the entire accessible web is. Search results are a ranked selection from an index, and the index itself is already a selection from the web.

What concentration does measure is exposure.

If five large domains collectively occupy half the first-page positions across thousands of queries, those domains receive repeated opportunities to be discovered while thousands of smaller sites receive none.

Discovery creates its own reality

This matters because most users do not inspect the whole index. They interact with the ranked surface.

A site can exist, remain technically accessible, and publish excellent material while being practically invisible to anyone who relies on ordinary search discovery.

That is a different kind of internet disappearance from deletion.

The page is still there.

The algorithmic city map simply stopped putting a road through its neighborhood.