Posted on

Distinguishing an undiscoverable web from a nonexistent web

A page that you cannot find is not necessarily a page that no longer exists.

That distinction is the difference between a dead web and an undiscoverable one.

Search engines expose only the portion of the web they have discovered, crawled, indexed, and decided to surface for a query. Independent systems can expose different material because they crawl and rank differently. Brave Search, for example, says it serves results from its own independent index, while Marginalia Search deliberately emphasizes non-commercial and independent sites with its own crawler and index software.

If the same query produces different obscure sites across different indexes, the missing material was not necessarily gone.

One map simply did not show it.

Test absence with more than one route

Suppose an old hobby topic appears to have vanished from the modern web.

A reasonable investigation might try:

  • ordinary search with several phrasings;
  • exact quoted phrases;
  • site-specific searches on likely hosts;
  • another search engine with a separate index;
  • curated directories or webrings;
  • old blogrolls and bookmark collections;
  • direct links from surviving related pages;
  • web archives if historical material is suspected.

These methods test different failure points.

An archive result may show that the page is genuinely dead today. A direct live link may show that the page exists but ranks poorly. A separate search index may reveal that one crawler found it while another did not.

Discovery failure can imitate disappearance

This matters because people experience the web primarily through interfaces.

If search stops surfacing independent forums, personal pages, niche blogs, and specialist archives, those sites can disappear from ordinary experience long before they disappear from servers.

From the user’s perspective, that feels like extinction.

Technically, it may be obscurity.

The distinction is not comforting in every case. A site nobody can discover may have almost the same practical reach as a deleted one.

But the remedies are different.

Deletion requires preservation or reconstruction.

Poor discovery requires better indexing, links, directories, feeds, cross-site recommendations, or deliberate exploration.

Dead Internet Theory needs this distinction

Claims that “there is nothing left out there” are difficult to evaluate if the only measuring instrument is the same ranked search interface being criticized.

A better test asks two separate questions:

Does the material still exist?

Can ordinary users still find it?

Those questions overlap, but they are not interchangeable.

The web can lose pages.

It can also lose roads.

Posted on

Search operators as tools for escaping default result selection

The default search box is not the only search box.

A few operators can force a search engine to reveal a very different slice of its index.

Google currently documents operators for exact phrases, restricting results to a site, excluding terms, and limiting results by date. Search Central also documents operators such as filetype: for finding specific document formats. See Refine Google searches and Google Search operators.

These are simple tools, but they matter because ordinary ranking tries to predict what is broadly most helpful.

Sometimes research requires something less broadly helpful and more specifically weird.

Change the question the ranking system receives

Suppose a search for an obscure piece of 1990s software mostly returns modern download sites and retrospective articles.

An exact quoted filename can suppress pages that merely discuss the subject. Adding filetype:pdf can expose old manuals. Restricting the query to a university or museum domain with site: can surface institutional archives. Excluding a dominant modern term with a minus sign can uncover older terminology.

For example:

"example.zip" -download

or:

site:edu "example software" filetype:pdf

The point is not the particular syntax. It is that a more constrained query asks the search engine to optimize inside a smaller box.

That can reveal material buried beneath the default interpretation of the topic.

Operators do not reveal a secret complete index

The escape hatch has limits.

Google warns that search operators are still constrained by indexing and retrieval systems. Its documentation for site: specifically says the results are not necessarily exhaustive and should not be used as a precise count of every indexed URL on a domain. See How to use the site: operator.

So operators can change selection without bypassing the search engine itself.

A page that was never discovered, never indexed, blocked from search, or removed from the index will not suddenly appear because the query got clever.

Operators are best understood as research controls.

They let the user reduce some of the ranking system’s freedom: search this domain, require this phrase, exclude this word, prefer this time window.

Default search answers the question it thinks you meant.

Operators are one way to make the argument more specific.

Posted on

StumbleUpon: the end of a shared culture of web wandering

StumbleUpon’s whole point was that you should not have to know what you wanted. Registered in late 2001 by four students at the University of Calgary, it grew into a browser toolbar with a single Stumble button. Pick a few categories — say, astronomy, photography, and bad puns — and each click opened a page chosen from what people like you had recommended. Camp described it in a 2007 interview as a hybrid that blended “collaborative human input with machine learning techniques, so users could discover great content they wouldn’t have thought to search for.”

Search assumed you knew the target

That was its bet against Google. Search was excellent when you knew exactly what you were looking for; StumbleUpon was for everything else. It borrowed the metaphor of channel surfing — surf directly between relevant content instead of searching, scanning, clicking, and going back. The service held nearly 500 topics, and your picks over time tuned the stream. Serendipity was a feature, not a side effect. Camp said he designed it so people could “stumble upon” sites recommended by like-minded people “[r]ather than presented with the most popular sites for a given keyword.”

Recommendations were a collective act

The engine ran on judgment. Thumbs-up on a page entered it into the database and shared it with others of similar taste; thumbs-down taught the system what to avoid. Users could also submit pages directly, add reviews, and friend the people whose taste they trusted. By 2007, Camp reported, 2 million members were stumbling about 5 million times a day and adding 16,000 new URLs daily. Every user was a curator, and the crowd’s taste — not an algorithm’s template — decided what surfaced.

The feed model replaced wandering

The slide was gradual. eBay bought the service for $75 million in 2007 and returned it to its founders in 2009. Then, chasing Pinterest and the social feed, StumbleUpon redesigned around an Activity stream, Trending pages, and a “StumbleDNA” profile that told you how much of you was tech. The wandering engine became another personalized feed, and traffic steadily declined.

The shutoff

In May 2018, Camp announced StumbleUpon would fold into Mix, a new discovery app from his studio Expa. Over the years it had delivered more than 40 million users personal content, serving up nearly 60 billion stumbles. On June 30, 2018, accounts were transitioned, and the site was gone. StumbleUpon.com still points toward Mix, which remains a web-discovery service built around finding and saving links from across the web. It is a descendant rather than a restoration of the old Stumble button and its public thumbs-up culture. What ended was the particular collective act: the human ratings that once helped steer where the rest of us wandered next.