Posted on

The discovery problem for communities that deliberately remain unlisted

Not every community wants to be found.

That sounds almost incompatible with the modern internet, where visibility is usually treated as the prize. More search traffic. More followers. More recommendations. More reach.

But some groups optimize for the opposite.

Discord’s current access controls make this explicit. A server can be configured as Invite Only, Apply to Join, or Discoverable. Invite Only is the default privacy setting, while public discovery is a separate choice with its own requirements. See Discord’s Server Member Applications documentation.

A group that stays invite-only may therefore be functioning exactly as designed when search produces nothing.

Discovery can travel through trust instead of indexing

People find unlisted communities through friends, coworkers, creators, event organizers, newsletters, private messages, conferences, neighboring groups, or somebody simply saying, you should be in this server.

This is slower than public search.

It can also be much more selective.

The person sharing the invitation provides a kind of informal recommendation. They know the group and they know the prospective member. That social filter can reduce spam, raids, unwanted surveillance, and people joining only to disrupt the space.

The cost is obvious: people without the right connections may never know the group exists.

Invisibility can be intentional infrastructure

Facebook provides an even stronger version through hidden private groups. Its help documentation says hidden groups that a person has not joined do not appear in search results. See Facebook’s guidance on joining groups.

That means the absence is not a ranking failure.

The platform is obeying the group’s privacy design.

This distinction matters when measuring the health of online communities. If a researcher searches for groups about a niche hobby and finds only three, the correct conclusion is not automatically that only three communities exist.

Some may be public but poorly ranked.

Some may be private but discoverable.

Some may be intentionally unlisted.

Those are different states.

The invisible web can be populated on purpose

Dead Internet Theory often interprets hard-to-find human activity as evidence that human activity has declined.

Sometimes that is true.

But some of the internet’s invisibility is chosen by the people using it.

The group is not lost.

It is not waiting for better SEO.

It put the curtains up itself.

Posted on

Human-edited directories as an alternative discovery system

Before search engines became the default map of the web, one common alternative was much simpler:

People made lists.

A human-edited directory organizes websites into categories chosen by editors rather than continuously ranking billions of pages for every query. That sounds primitive compared with modern search, but it changes the discovery problem in interesting ways.

Curlie is a surviving example. It describes itself as a human-edited directory run by volunteer editors. Its editorial guidelines are public, and its editors are instructed to select, evaluate, describe, and organize sites according to published criteria. Curlie’s editor information explains that editors apply to manage categories and review suggested sites.

The selection process is therefore opinionated, but not invisible.

Curation makes the selector legible

A search engine can rank thousands of candidates through signals most users never see.

A directory instead says, in effect: these sites were selected for this category.

The editor may still make mistakes. The category structure may be awkward. A useful site may be omitted. But the basic selection model is understandable, and Curlie’s public guidelines even describe what kinds of sites are generally included or excluded.

That transparency has value.

A reader exploring a directory can browse sideways through neighboring categories instead of only asking for a precise keyword. That makes directories useful for accidental discovery and for subjects where the user does not yet know the right search terms.

Humans have a crawl budget too

The weakness is scale.

Volunteer editors cannot inspect the whole web. Categories can become stale. Suggested sites can wait for review. Editors can become inactive. New subjects can grow faster than the directory structure adapts.

Curlie itself acknowledges backlogs and depends on volunteers to maintain categories. Its model trades automated breadth for human judgment.

That means a missing site proves very little.

It may have been rejected under the guidelines. Nobody may have suggested it. An editor may not have reached it yet. The relevant category may have little active maintenance.

Human curation therefore does not solve Algorithmic Reality by producing a perfectly neutral internet.

It produces a different kind of selection.

The important difference is that the selection rule is easier to inspect: named categories, public editorial policies, human review, and visible organizational choices.

Modern discovery is often framed as a choice between good algorithms and bad algorithms.

The older web reminds us that there is another option.

Sometimes the map can simply admit that somebody drew it.

Posted on

Music recommendations and the narrowing or widening of listening habits

A recommendation system can make your musical world smaller or larger using the same listening history.

If somebody repeatedly plays death metal, the system can respond by finding more death metal. That deepens a known preference.

It can also use the same history to recommend adjacent scenes, older influences, unfamiliar artists, another country’s version of the genre, or something structurally similar that the listener has never searched for.

Both outcomes are personalization.

Spotify describes its current Taste Profile as an interpretation of what a person likes based on what and how they listen. That profile helps shape Home recommendations and other personalized experiences. Spotify even lets users exclude tracks and playlists when a one-off listen would otherwise distort that profile. See Taste Profile and Spotify’s explanation of excluding tracks from it.

That control exists because listening behavior is not a perfect statement of identity.

Sometimes the children’s song is for the child.

Repetition and discovery are both design choices

Recommendation systems often face a tradeoff between exploitation and exploration.

Exploitation means recommending something close to what the system already knows works. Exploration means spending some recommendation space on uncertain material that may broaden the listener’s taste.

Spotify’s discovery products demonstrate both impulses. Discover Weekly uses listening history to personalize recommendations, while features such as Fresh Finds and editorial discovery playlists deliberately introduce less familiar material. In July 2026 Spotify described its weekly discovery playlists as tools for finding new releases, breakout tracks, and music beyond a listener’s existing rotation. See Spotify’s discovery-driven playlists.

So a personalized system is not automatically a musical filter bubble.

It can become one if similarity repeatedly wins over novelty.

Measure variety instead of guessing at it

Listening variety can be measured more carefully than asking whether recommendations “feel repetitive.”

Useful measures might include the number of distinct artists encountered, how often recommendations introduce artists never previously played, genre diversity, geographic diversity, catalog age, repeat rate, and the share of listening devoted to already-familiar music.

Even those metrics require interpretation. A person intentionally exploring one composer’s catalog may want less variety for a month. Another listener may explicitly want constant novelty.

The Algorithmic Reality issue is therefore not that a machine chooses songs.

Radio programmers, record stores, friends, DJs, critics, and record labels have always influenced what people hear.

The new difference is that the selector can continuously learn from the listener and rebuild the record shelf after every session.

Whether that shelf becomes a tunnel or a doorway depends on what the system is optimizing—and whether the listener can push back.

Posted on

Recommendation diversity and deliberate exposure to unfamiliar sources

A recommendation system that only shows you things very similar to what you already consumed can be extremely accurate and still make the internet feel tiny.

Accuracy is not the same thing as discovery.

Recommendation researchers therefore talk about properties such as diversity, novelty, serendipity, and catalog coverage alongside relevance. A system can deliberately spend some recommendation space on material that is less certain but potentially useful.

YouTube’s current discovery guidance says its systems look at what videos are often watched together and may identify videos viewers are likely to watch but have not been exposed to yet. See its Search and discovery tips.

That last phrase matters.

A recommendation does not have to be the statistically safest continuation of the user’s existing habits.

Diversity has more than one meaning

A feed can be diverse in topic but not source.

It can show politics, cooking, gaming, and science while all four come from the same handful of giant publishers. It can show many creators who all express roughly the same viewpoint. Or it can provide genuine source diversity while remaining tightly focused on one subject.

So “more diverse recommendations” needs a defined target.

Are we trying to increase unfamiliar creators? Less-popular items? Different viewpoints? Different languages? Different formats? A broader range of topics?

Those goals can conflict with one another.

Exploration costs certainty

The tradeoff is familiar in recommendation research: exploit what the system already knows works, or explore something less certain to learn more.

Too much exploitation creates repetition. Too much exploration produces a feed full of things the user does not want.

Popularity-bias research shows why this matters. A 2020 study on popularity debiasing in collaborative recommendation found that reducing popularity bias could improve qualities such as coverage and diversity with relatively small losses in accuracy in its experiments.

That does not mean every platform should maximize obscurity. Some popular material is popular because it is excellent.

The important point is that recommendation diversity can be an explicit design choice rather than an accidental by-product.

For a user, deliberate exploration also works outside the algorithm: visit a directory, search by a strange phrase, browse subscriptions chronologically, follow a link from a small site, or intentionally choose a source the default feed never surfaces.

Algorithmic Reality becomes less confining when either the system or the user occasionally asks a dangerous question:

What if the next thing is not more of the same?

Posted on

Search-result diversity beyond the first screen

A search engine can contain diversity that almost nobody sees.

The first screen is not the whole result set. It is the part most people treat as the result set.

That distinction matters because ranking systems intentionally compress enormous candidate pools into a very small visible surface. Google even describes a site diversity system intended to keep one domain from dominating too many of the top results in ordinary cases.

But “top results” is doing a lot of work there.

The sources visible after the first ten, twenty, or fifty positions may look quite different from the sources presented first.

Available diversity is not encountered diversity

A useful audit can record domains at multiple depths for the same query.

Perhaps the first screen contains national publishers, Wikipedia, Reddit, YouTube, and a few large commercial sites. Deeper results may introduce personal pages, regional organizations, specialist forums, academic PDFs, independent blogs, and old technical archives.

If so, the search index contains more source variety than the first screen suggests.

That does not mean users experience that variety.

A large Backlinko analysis of roughly four million Google results found a steep decline in click-through as ranking position fell and reported that only a small fraction of searchers clicked results on the second page. The exact percentage belongs to that dataset and period, not to every search forever, but the general behavioral point is hard to miss: deeper availability and actual exposure are very different things. See We Analyzed 4 Million Google Search Results.

Google also abandoned its experiment with continuous scrolling and returned to explicit result pagination in 2024, again placing a user action between the first batch of results and the next one.

Search depth changes the story you tell about the web

Suppose somebody searches ten technology questions and records only the first screen. They may conclude that the modern web is dominated almost entirely by a small group of platforms.

That observation may be accurate for first-screen exposure.

It does not establish that those platforms dominate every indexed result or every relevant page available farther down.

The opposite mistake is possible too. Finding fifty wonderful independent sites on page six does not prove ordinary users are discovering them.

A strong study should therefore report both things: the diversity that exists at increasing result depth and the probability that users actually reach those depths.

Dead Internet Theory often asks where the weird little sites went.

Sometimes they went nowhere.

They are still standing six blocks behind the billboard, while almost everyone turns around at the first intersection.

Posted on

Search-index coverage versus the size of the accessible web

A web page can be publicly accessible and still be absent from a search engine’s index.

That distinction sounds technical until somebody tries to use search results as a census of the internet.

Search engines do not begin with every page on the web and then decide how to rank them. They first have to discover URLs, crawl them, process what they find, decide which versions are duplicates, and determine whether a page belongs in the index at all. Only after that can an indexed page compete for a search result.

Google’s own documentation is explicit about the limits. Its guide to how Search works says Googlebot does not crawl every page it discovers and that indexing is not guaranteed. Search Console’s Page indexing report documentation goes even further: Google does not guarantee that all pages everywhere will make it into its index.

Accessible is not the same as indexed

A page can return a perfectly ordinary 200 OK response in a browser and remain outside the index for several reasons.

The crawler may not know the URL exists. The page may be weakly linked. It may duplicate another page. A site may accidentally block crawling or indexing. Google may crawl the page and still decide not to index it. A URL may simply be waiting in the discovery queue.

Search Console even distinguishes between Discovered – currently not indexed and Crawled – currently not indexed. In the first case Google knows the URL but has not fetched it yet. In the second, it fetched the page but did not add it to the searchable index.

Neither state means the page is unavailable on the web.

Index size is not web size

This matters for Dead Internet Theory because search engines are often used as informal measuring instruments.

If a search for some obscure subject produces only twelve useful pages, several explanations are possible. Perhaps only twelve useful pages exist. Perhaps more exist but are poorly linked. Perhaps some are not indexed. Perhaps the search engine clustered similar material, ranked other pages higher, or failed to interpret the terminology used by a small community.

A result count therefore measures something closer to what a particular search system currently exposes for a query than the total amount of relevant material online.

Even Search Console’s site-level numbers only describe URLs Google knows about. A completely unknown URL is missing from both the indexed and non-indexed totals.

That is why estimating the size of the web from search indexes is slippery. The crawler’s frontier, the index, and the publicly accessible web are three overlapping but different things.

A search engine can be enormous without being complete.

The map can contain billions of roads and still leave towns off it.

Posted on

Domain concentration across ordinary search results

The web can contain a million pages while a search result shows you ten.

That compression is unavoidable. Search engines have to rank. The interesting question is which domains keep surviving the compression.

If a handful of large sites appear repeatedly across ordinary queries, users can experience the web as far more concentrated than the underlying collection of websites actually is.

That is where Algorithmic Reality — The Internet You Are Allowed to See begins.

Concentration is visible across many queries, not just one

One search results page is a weak sample.

A query for a specific company should reasonably return that company’s site several times. A technical query may be dominated by the official documentation. A breaking-news query may favor a small group of publishers because they have current reporting.

The more useful test is horizontal: run many queries in a defined category, record the domains that appear, then ask how much of the available result space is occupied by the same publishers.

A 2024 audit of Google Search news results across Brazil, the United Kingdom, and the United States analyzed more than 220,000 results and reported substantial concentration among a limited group of outlets. The researchers specifically used concentration measures such as the Herfindahl-Hirschman Index and Gini coefficient rather than judging diversity from a few screenshots. See Auditing Google’s Search Algorithm: Measuring News Diversity Across Brazil, the UK, and the US.

Google itself has acknowledged the problem category for years. In a 2012 search-quality update it described a change called Domain Crowding intended to surface a more diverse set of domains when too many results came from the same site. Search systems have continued to use site-diversity mechanisms since then.

The existence of such mechanisms tells us something important: relevance and source diversity are not automatically the same objective.

Concentration is not automatically bad

A domain appearing repeatedly may deserve to appear.

Official documentation can be better than ten scraped copies. A specialist site may dominate a narrow subject because it genuinely has the strongest material. A local query may reasonably favor a few authoritative local sources.

That is why domain concentration alone cannot measure search quality.

It also cannot tell us how diverse the entire accessible web is. Search results are a ranked selection from an index, and the index itself is already a selection from the web.

What concentration does measure is exposure.

If five large domains collectively occupy half the first-page positions across thousands of queries, those domains receive repeated opportunities to be discovered while thousands of smaller sites receive none.

Discovery creates its own reality

This matters because most users do not inspect the whole index. They interact with the ranked surface.

A site can exist, remain technically accessible, and publish excellent material while being practically invisible to anyone who relies on ordinary search discovery.

That is a different kind of internet disappearance from deletion.

The page is still there.

The algorithmic city map simply stopped putting a road through its neighborhood.

Posted on

Ancient Web: Early Web Links Is Building a Card Catalog for the Personal Web

The old web did not only lose websites.

It lost ways of finding websites.

Visit Early Web Links

Early Web Links is a modern attempt to rebuild that missing layer. Its premise is blunt: before the internet went corporate, people made websites because they had something to share, and many of those sites are now difficult to discover through ordinary search.

So the site catalogs them by hand.

At the time of this article, the directory reports more than 12,000 sites and sorts them into the sort of categories a human actually browses: literature, programming, retro computing, zines, pets, collecting, astronomy, amateur radio, paranormal material, local communities, recipes and dozens more.

It also has a Random Site button, guestbook, site-of-the-week feature and a stream of new additions.

In other words, it behaves like a directory instead of apologizing for being one.

Discovery without pretending to know what you want

Search engines are excellent when the user has a query.

Directories solve a different problem.

They let somebody browse a neighborhood of ideas without already knowing the destination. That matters for personal websites because their value is often impossible to summarize into a high-volume keyword.

A Viennese zoologist’s giant shrew site, somebody’s tilde page, a homebrew electronics archive or a regional history page may be exactly what you wanted to find ten seconds after you discover that it exists.

Before that moment, there was no query.

Early Web Links leans into period aesthetics, but the useful part is underneath the Netscape buttons and under-construction graphics: human classification. Somebody looked at the site, decided what it was about and placed it near other things that might send a curious person sideways.

That is closer to a card catalog than an algorithmic feed.

CacheRat’s 1,967 Ancient Web Domains research list is obviously operating in the same archaeological territory. The difference is useful: CacheRat is a salvage list; Early Web Links is growing into a browsable directory.

Both exist because the modern web has an absurd amount of information and still manages to hide the interesting stuff.

The solution may not be smarter ranking.

Sometimes you just need shelves again.

Browse Early Web Links

Posted on

Ancient Web: Does Anyone Remember Websites? Asked the Right Question in 2017

In 2017 somebody asked a question that sounds increasingly less sarcastic every year:

Does anyone remember websites?

Read the original TTTThis page

The argument was simple. Search engines had become incredibly good at finding specific answers, but something got lost in the process: wandering into pages built by hobbyists, academics and technically inclined weirdos because they cared about something.

The page contrasts that older web with a commercial web where personalized sites are harder to surface among optimized pages built to rank, sell or convert.

It is not a claim that everything used to be better. Plenty of old sites were broken, ugly, abandoned or useless.

The point is narrower and better: searching is not the same thing as surfing.

Finding what you did not know to ask for

When the page hit Hacker News in November 2017, it drew hundreds of comments and nearly 900 points. The discussion immediately became a catalog of people remembering the exact experience the essay described: clicking through somebody’s handwritten pages and discovering useful material that no search query would have occurred to them to type.

That distinction still matters.

Modern search is optimized around intent. You ask for a thing and receive increasingly polished attempts at that thing.

The old personal web often worked sideways. You arrived looking for batteries and ended up reading a homemade test bench. You searched for one obscure technical term and discovered somebody’s entire research notebook. Links were not merely citations. They were trails.

That is why small directories, webrings, blogrolls and manually curated lists keep reappearing. They solve a different problem from search engines.

They help you discover something before you have language for it.

CacheRat’s 1,967 Ancient Web Domains research list is basically an industrialized version of that impulse: stop asking the web only for answers and start poking through its forgotten sheds again.

The original TTTThis page is flaky now, which is almost too appropriate.

The idea survived anyway.

Try the original page