Posted on

The discoverability of independent sites with little inbound linking

A useful website can be almost invisible simply because hardly anybody links to it.

That is not a mystical penalty. It follows from how web search discovers and evaluates pages.

Google’s link best-practices documentation says links help its systems find new pages to crawl and also act as signals for understanding relevance. Google’s broader reliability guidance says references from prominent sites can also contribute evidence that a source is trustworthy.

An isolated site therefore starts with two related disadvantages: fewer paths by which crawlers can discover it and fewer external references helping search systems understand where it fits.

Good content does not automatically create good connectivity

Consider a retired engineer who publishes a small site about an obsolete control system.

The pages may contain original schematics, repair notes, and information unavailable anywhere else. But the audience is tiny. Few modern sites discuss the hardware. The old forums that once linked to the material may be gone. The author may never have promoted the site.

The result is a page with high specialist value and low web connectivity.

That is very different from a low-quality page that receives few links because nobody finds it useful.

Search systems cannot perfectly distinguish those cases from first principles. They observe the evidence available to them.

Internal linking and sitemaps can help with discovery, but external links do more than expose a URL. They place the site inside the web’s graph of relationships.

Isolation can be measured separately from content quality

A useful investigation should avoid assuming that low ranking proves poor material.

Check whether the page is indexed. Count unique referring domains. Look at the age and relevance of those links. Compare the site’s technical accessibility with better-connected competitors. Search for exact phrases or the site’s name to see whether the engine can retrieve it when ambiguity is removed.

Then inspect the page itself.

Does it contain firsthand material? Original photographs? Technical data? Citations? Evidence of subject expertise? Information absent from the larger sites outranking it?

Those questions separate connectivity from content value.

This also explains why the independent web can feel smaller than it is. Large platforms constantly receive new links because people are already there. Small sites can publish into near silence.

The web remains technically decentralized, but discovery is strongly influenced by the network of references connecting one page to another.

A page with no incoming roads may still contain a museum.

You just have to know the dirt road exists.

Posted on

Search-result diversity beyond the first screen

A search engine can contain diversity that almost nobody sees.

The first screen is not the whole result set. It is the part most people treat as the result set.

That distinction matters because ranking systems intentionally compress enormous candidate pools into a very small visible surface. Google even describes a site diversity system intended to keep one domain from dominating too many of the top results in ordinary cases.

But “top results” is doing a lot of work there.

The sources visible after the first ten, twenty, or fifty positions may look quite different from the sources presented first.

Available diversity is not encountered diversity

A useful audit can record domains at multiple depths for the same query.

Perhaps the first screen contains national publishers, Wikipedia, Reddit, YouTube, and a few large commercial sites. Deeper results may introduce personal pages, regional organizations, specialist forums, academic PDFs, independent blogs, and old technical archives.

If so, the search index contains more source variety than the first screen suggests.

That does not mean users experience that variety.

A large Backlinko analysis of roughly four million Google results found a steep decline in click-through as ranking position fell and reported that only a small fraction of searchers clicked results on the second page. The exact percentage belongs to that dataset and period, not to every search forever, but the general behavioral point is hard to miss: deeper availability and actual exposure are very different things. See We Analyzed 4 Million Google Search Results.

Google also abandoned its experiment with continuous scrolling and returned to explicit result pagination in 2024, again placing a user action between the first batch of results and the next one.

Search depth changes the story you tell about the web

Suppose somebody searches ten technology questions and records only the first screen. They may conclude that the modern web is dominated almost entirely by a small group of platforms.

That observation may be accurate for first-screen exposure.

It does not establish that those platforms dominate every indexed result or every relevant page available farther down.

The opposite mistake is possible too. Finding fifty wonderful independent sites on page six does not prove ordinary users are discovering them.

A strong study should therefore report both things: the diversity that exists at increasing result depth and the probability that users actually reach those depths.

Dead Internet Theory often asks where the weird little sites went.

Sometimes they went nowhere.

They are still standing six blocks behind the billboard, while almost everyone turns around at the first intersection.

Posted on

Duplicate filtering and the visibility of original sources

The same article can exist at ten URLs without becoming ten independent pieces of knowledge.

Search engines know this, which is why they try to group duplicates instead of filling result pages with copies.

Google calls this process canonicalization. Its canonicalization documentation explains that when several pages contain the same or very similar primary content, Google clusters them and chooses one representative URL as the canonical version. That version is usually the one shown in ordinary search results.

For users, this is mostly good housekeeping.

Without duplicate filtering, one syndicated article could occupy half a results page.

The interesting problem is what happens when the search system’s chosen representative is not the page that actually originated the material.

Copies can become easier to see than sources

Syndication is normal on the web. News articles are republished by partners. Press releases appear on hundreds of sites. Product data travels through retailers. Blog posts are mirrored. Scrapers copy pages without permission.

The search engine has to decide which versions belong together and which one should represent the cluster.

Publishers can provide hints using redirects, sitemaps, and rel="canonical", but Google explicitly describes canonical selection as something its systems ultimately determine. Its current canonicalization troubleshooting guide even notes that Google may sometimes choose a different canonical than the site owner prefers.

Syndicated content is especially awkward. Google’s guidance says that publishers who do not want partner copies appearing in Search should generally have those partners block indexing rather than relying on canonical tags alone.

That is a strong clue that “original source wins automatically” is not a safe assumption.

Duplicate filtering changes visibility, not history

If a more prominent copy becomes the version users encounter, the original article has not ceased to exist.

Its role has changed from visible destination to hidden member of a duplicate cluster.

This matters when reconstructing where information came from. A search result may point to the largest distributor, the technically cleanest version, or the URL the system judged most useful—not necessarily the page with the earliest publication timestamp.

Investigating an origin therefore requires more than accepting the first ranked copy. Compare dates, bylines, attribution, canonical tags, archives, syndication notices, and links back to upstream material.

Duplicate filtering solves a real search problem. Nobody wants twelve identical wire stories before the thirteenth result.

But the cleaner result page can conceal the messy genealogy underneath it.

The internet you see may contain one visible document.

Behind it may be a whole family tree of copies—and the ancestor is not guaranteed to be standing at the front.

Posted on

Language coverage and the discoverability of smaller linguistic communities

The web does not become equally searchable merely because two pages are both public.

Language changes the problem.

A search engine needs enough documents, links, query history, language understanding, spelling models, translation resources, and relevance evidence to retrieve useful material. Those resources are not distributed evenly among the world’s languages.

Google openly describes some of this unevenness. Its language-selection documentation says Search may return material from another language when it considers that useful, including cases where there is not enough information in the language used for the query. Google also offers translated search results, but that specific translation feature is currently available only for a defined list of languages.

That does not mean other languages are absent from Google’s index. It means different search and translation capabilities have different coverage.

Smaller languages have a retrieval problem as well as a publishing problem

A community may publish excellent local material and still be difficult for outsiders to find.

Academic work on cross-lingual information retrieval has documented the technical gap directly. A 2022 paper, Pivot Through English: Reliably Answering Multilingual Questions without Document Retrieval, notes that open-retrieval question answering in lower-resource languages has historically lagged behind English partly because non-English document retrieval itself is harder.

That distinction is important.

If a search in English produces many results while the equivalent search in a smaller language produces only a few, it does not follow that the smaller linguistic community has little to say. The difference may involve fewer indexed documents, weaker cross-language matching, different query terminology, fewer inbound links, less machine-readable material, or simply less training and evaluation data for the retrieval system.

English is a terrible universal control group

Researchers studying the apparent diversity of the web can accidentally measure English-language infrastructure and mistake it for the internet.

A fair comparison should consider how much material is actually published in each language, whether equivalent queries exist, whether the script is handled correctly, whether search systems recognize morphology and spelling variation, and whether translation is introducing or hiding results.

Cross-language search can also create a strange asymmetry. A locally important page may be obvious to native speakers using local terminology but practically invisible to an English speaker searching for an English translation of the same concept.

The page is alive. The community is alive. The search path between them and the outside observer is weak.

That is an important form of Algorithmic Reality.

A search engine can make the web feel overwhelmingly English, or overwhelmingly dominated by a few high-resource languages, without deleting a single page written anywhere else.

Sometimes the missing internet is speaking perfectly clearly.

The discovery system just does not speak back as well.

Posted on

Search localization and different views of the same subject

Two people can type the same words into the same search engine and receive different maps of the web.

Sometimes that is the entire point.

A search for pizza, weather, tax attorney, or hardware store is nearly useless without geographic context. Search systems therefore use location and other contextual signals to decide what “relevant” means for the person asking.

Google’s Search Help documentation says results can vary according to location, language, device type, recent searches, and personalization. Its guide to location in Search says Google always estimates a general search area and can use more precise location when permission is available.

That means there is no single permanent result page for many queries.

Geography changes meaning before it changes ranking

Suppose two users search for football.

One may be in Missouri. Another may be in Manchester. The word itself sits inside different local contexts before ranking even begins.

The same problem appears with ordinary services. Google’s own documentation uses local searches as examples because somebody looking for bicycle repair in Paris should not receive exactly the same result set as somebody looking in Hong Kong.

News can vary too. A regional source may be more relevant near the event it covers. Businesses have service areas. Regulations differ by jurisdiction. Languages and spelling vary across borders.

Localization therefore is not automatically evidence that a search engine is manipulating reality in some sinister personalized bubble. Often it is simply solving an ambiguous query with available context.

It still changes the internet people perceive

The important Dead Internet Theory question is methodological.

If one researcher searches from Chicago and declares, “These are the websites Google shows for this topic,” that claim is incomplete. A researcher in London, Delhi, or São Paulo may see a meaningfully different set.

A proper comparison should control or record location, language settings, device, login state, time, and query wording. Country and region settings can be changed deliberately to compare result sets.

Even then, variation does not prove bias. It establishes that the search experience is conditional.

That matters when people argue that some source has “disappeared from Google.” It may be buried everywhere. It may appear strongly in one region and weakly in another. It may surface only for users whose language settings match the page.

The web underneath can be identical while the visible layer differs from person to person.

Algorithmic Reality is not merely a ranking of pages.

It is a ranking of pages for you, here, now.

Posted on

Authority signals and the advantage of established publishers

A famous publisher can be wrong and an obscure hobbyist can be the world’s best source on one strange little subject.

Search engines still need a way to rank both of them.

That forces search systems to use proxies for relevance, usefulness, and reliability. Some are specific to the page. Others emerge from the wider web around it: links, references, reputation, topical history, and signals that other people treat the source as worth consulting.

Google’s Reliable results on Search documentation gives a simple example. If other prominent sites link to or refer to a piece of content, that can suggest that the source is reliable. Google’s ranking systems guide also describes link-analysis systems, including the modern descendants of PageRank, along with systems intended to surface more reliable and authoritative information.

None of that translates into a simple “big site bonus.”

It does create an accumulation problem.

Established publishers have history to spend

A long-running publication may have millions of inbound links, recognizable authors, years of citations, structured archives, stable URLs, and a large audience that continuously creates new references to its work.

A new independent site starts with almost none of that.

Even if both publish equally useful pages today, they arrive at the ranking system with very different histories.

That advantage can be deserved. Institutions that repeatedly publish accurate material should not be forced to prove themselves from zero on every query. Strong reputation signals help suppress spam, impersonation, disposable content farms, and pages created yesterday to exploit today’s search demand.

But proxies have edges.

A retired engineer may maintain the best documentation for an obsolete machine. A collector may have photographed a component that no museum has cataloged. A regional historian may know a subject ignored by national publishers. These sources can be exceptionally valuable while possessing very little conventional web authority.

Authority is not the same as expertise on every page

Google itself notes that ranking operates substantially at the page level and that good site-wide signals do not guarantee every page will rank highly.

That distinction matters.

The existence of authority systems does not prove that a large publisher automatically beats a specialist. Nor does a small site’s poor ranking prove that its content was judged incorrect. Ranking combines many signals, and the exact weighting is not public.

The useful observation is narrower: established publishers have more opportunities to accumulate the signals search systems can observe.

That affects the internet people experience.

If discovery repeatedly favors sources with large existing reputations, the web can appear more institutionally concentrated than the underlying supply of knowledge really is.

The specialist page may still exist.

It just arrives at the race without forty thousand people already pointing at it.

Posted on

Freshness preferences and the burial of still-useful older pages

A ten-year-old page about yesterday’s earthquake is probably not what you need.

A ten-year-old page explaining how a thirty-year-old machine works might be exactly what you need.

Search systems have to tell the difference.

Google describes freshness systems designed for queries where newer information is expected. Its own examples include a newly released movie, where recent reviews are useful, and an earthquake, where breaking information may suddenly become more relevant than evergreen preparedness pages.

That is sensible. The trouble starts when people turn a conditional preference for freshness into a universal rule that “Google hates old pages.”

Some knowledge has an expiration date

Prices change. Elections happen. Software versions are replaced. Laws are amended. A restaurant closes. A hurricane moves inland.

For those subjects, publication date and update date can be useful signals because the underlying world changed.

Other subjects are stubbornly durable.

An old repair manual may describe a discontinued radio better than anything published last week. A twenty-year-old technical essay may still be the clearest explanation of a protocol that has not materially changed. A personal history written by somebody who was there can become more valuable with age rather than less.

Search therefore cannot simply sort the web by date.

Google says its freshness systems operate where a query appears to deserve fresh results. Its broader ranking systems still consider many other signals involving relevance, usefulness, reliability, links, and context.

Old pages can still become hard to find

Even when age is not an automatic penalty, older useful pages can lose visibility indirectly.

Newer pages may accumulate stronger links. A large publisher may rewrite an old subject in a format that better matches current search language. An old page may use obsolete terminology. Its site may have degraded technically. Its internal links may disappear after redesigns. Competitors may simply publish something better.

That makes diagnosis important.

If an old page ranks poorly, compare queries where freshness matters with queries where it does not. Check whether newer competing pages actually contain newer facts. Examine whether the old page is indexed, linked, technically accessible, and still aligned with the language people use to search for the subject.

Merely changing a date is not evidence that a document became more useful.

The useful question is not “How old is this page?”

It is “Did anything about the answer become old?”

Algorithmic Reality can make yesterday feel disproportionately large because recent material is often easier to surface and easier to produce. But the older web is not obsolete merely because its timestamp looks archaeological.

Sometimes the best answer really is sitting under a date that scares the ranking-conscious.

Posted on

Crawl budgets and the visibility of small websites

“Google did not crawl my page” and “Google ran out of crawl budget for my tiny website” are not the same claim.

Search crawlers do have finite resources. They cannot fetch every URL continuously, and they have to avoid hammering a server until it falls over. Google describes a site’s crawl budget as the set of URLs its systems can and want to crawl, based mainly on crawl capacity and crawl demand.

But Google’s current crawl-budget documentation is surprisingly blunt about who should worry about it. The advanced guide is aimed mainly at sites with roughly a million or more changing pages, sites with tens of thousands of rapidly changing pages, or sites showing large numbers of URLs as discovered but not indexed.

For an ordinary small website, “crawl budget” can become an impressive-sounding diagnosis for a much simpler problem.

Small sites can still be missed

A small site does not need to exhaust some giant quota to have pages overlooked or refreshed slowly.

Googlebot primarily discovers URLs through links from pages it already knows, along with sitemaps and other discovery mechanisms. If an article is buried several levels deep, linked only through JavaScript Google cannot reliably interpret, omitted from navigation, or effectively orphaned, discovery can be slow.

Server behavior matters too. Google’s crawling troubleshooting guide says crawling can be reduced when a site responds slowly, returns server errors, rate-limits requests, or otherwise signals that it cannot comfortably handle more traffic.

And being crawled still does not guarantee being indexed.

That last distinction matters. Crawling asks, did the search engine fetch it? Indexing asks, did the search engine retain it as a searchable document? Ranking asks, where does that indexed page appear for a query?

Three different gates.

Crawl attention is not ranking authority

A small site can be crawled perfectly and rank nowhere useful. It can also rank well for a narrow query despite being crawled far less often than a major news site.

Popularity and update frequency influence crawl demand because frequently changing or important URLs may need to be revisited more often. That does not mean crawl frequency itself is a simple ranking score.

For a small independent site, a useful investigation starts with boring evidence: is the URL linked internally, present in a sitemap, reachable without errors, fetched by the crawler, indexed, and relevant to an actual query?

Only then does “crawl budget” become a useful explanation rather than an SEO ghost story.

The web does ration crawler attention. But for small sites, the more common problem is often not that the search engine has no time left.

It is that the page has not given the search system a strong enough path, reason, or signal to return.

Posted on

Search-index coverage versus the size of the accessible web

A web page can be publicly accessible and still be absent from a search engine’s index.

That distinction sounds technical until somebody tries to use search results as a census of the internet.

Search engines do not begin with every page on the web and then decide how to rank them. They first have to discover URLs, crawl them, process what they find, decide which versions are duplicates, and determine whether a page belongs in the index at all. Only after that can an indexed page compete for a search result.

Google’s own documentation is explicit about the limits. Its guide to how Search works says Googlebot does not crawl every page it discovers and that indexing is not guaranteed. Search Console’s Page indexing report documentation goes even further: Google does not guarantee that all pages everywhere will make it into its index.

Accessible is not the same as indexed

A page can return a perfectly ordinary 200 OK response in a browser and remain outside the index for several reasons.

The crawler may not know the URL exists. The page may be weakly linked. It may duplicate another page. A site may accidentally block crawling or indexing. Google may crawl the page and still decide not to index it. A URL may simply be waiting in the discovery queue.

Search Console even distinguishes between Discovered – currently not indexed and Crawled – currently not indexed. In the first case Google knows the URL but has not fetched it yet. In the second, it fetched the page but did not add it to the searchable index.

Neither state means the page is unavailable on the web.

Index size is not web size

This matters for Dead Internet Theory because search engines are often used as informal measuring instruments.

If a search for some obscure subject produces only twelve useful pages, several explanations are possible. Perhaps only twelve useful pages exist. Perhaps more exist but are poorly linked. Perhaps some are not indexed. Perhaps the search engine clustered similar material, ranked other pages higher, or failed to interpret the terminology used by a small community.

A result count therefore measures something closer to what a particular search system currently exposes for a query than the total amount of relevant material online.

Even Search Console’s site-level numbers only describe URLs Google knows about. A completely unknown URL is missing from both the indexed and non-indexed totals.

That is why estimating the size of the web from search indexes is slippery. The crawler’s frontier, the index, and the publicly accessible web are three overlapping but different things.

A search engine can be enormous without being complete.

The map can contain billions of roads and still leave towns off it.

Posted on

Domain concentration across ordinary search results

The web can contain a million pages while a search result shows you ten.

That compression is unavoidable. Search engines have to rank. The interesting question is which domains keep surviving the compression.

If a handful of large sites appear repeatedly across ordinary queries, users can experience the web as far more concentrated than the underlying collection of websites actually is.

That is where Algorithmic Reality — The Internet You Are Allowed to See begins.

Concentration is visible across many queries, not just one

One search results page is a weak sample.

A query for a specific company should reasonably return that company’s site several times. A technical query may be dominated by the official documentation. A breaking-news query may favor a small group of publishers because they have current reporting.

The more useful test is horizontal: run many queries in a defined category, record the domains that appear, then ask how much of the available result space is occupied by the same publishers.

A 2024 audit of Google Search news results across Brazil, the United Kingdom, and the United States analyzed more than 220,000 results and reported substantial concentration among a limited group of outlets. The researchers specifically used concentration measures such as the Herfindahl-Hirschman Index and Gini coefficient rather than judging diversity from a few screenshots. See Auditing Google’s Search Algorithm: Measuring News Diversity Across Brazil, the UK, and the US.

Google itself has acknowledged the problem category for years. In a 2012 search-quality update it described a change called Domain Crowding intended to surface a more diverse set of domains when too many results came from the same site. Search systems have continued to use site-diversity mechanisms since then.

The existence of such mechanisms tells us something important: relevance and source diversity are not automatically the same objective.

Concentration is not automatically bad

A domain appearing repeatedly may deserve to appear.

Official documentation can be better than ten scraped copies. A specialist site may dominate a narrow subject because it genuinely has the strongest material. A local query may reasonably favor a few authoritative local sources.

That is why domain concentration alone cannot measure search quality.

It also cannot tell us how diverse the entire accessible web is. Search results are a ranked selection from an index, and the index itself is already a selection from the web.

What concentration does measure is exposure.

If five large domains collectively occupy half the first-page positions across thousands of queries, those domains receive repeated opportunities to be discovered while thousands of smaller sites receive none.

Discovery creates its own reality

This matters because most users do not inspect the whole index. They interact with the ranked surface.

A site can exist, remain technically accessible, and publish excellent material while being practically invisible to anyone who relies on ordinary search discovery.

That is a different kind of internet disappearance from deletion.

The page is still there.

The algorithmic city map simply stopped putting a road through its neighborhood.