Posted on

Geographic access restrictions and uneven archive coverage

The internet does not present the same face to everyone. A server decides what to send based on where the request appears to come from, and that treats human readers and crawlers alike: both arrive at a site from some country. What an archive holds is therefore shaped less by how hard it crawled and more by where its crawlers happened to sit.

Where you connect changes what you see

Geoblocking is server-side denial of access based on where a request appears to come from. A large 2018 measurement study tested major sites from vantage points in 177 countries and found CDN-based geoblocking in nearly all of them; in its broader sample, about 4.4 percent of domains used a CDN geoblocking feature in at least one country. The motive can be legal, commercial, sanctions-related, or operational rather than political.

That blur is worth remembering. The same study found that 9 percent of a widely used list of censored domains returned an ordinary block page in at least one country. Geographic restriction can look exactly like censorship, and vice versa, which is precisely why crawlers get confused.

A crawler’s vantage point becomes an archive’s bias

Any archive crawler has a geographic vantage point. If a site blocks that region, serves a regional substitute, or varies content by location, the archive may preserve the block page or the regional version rather than what visitors elsewhere saw.

Different archives also have different collection mandates and strengths. A 2013 study profiling twelve public web archives found that holdings varied enough by top-level domain and language that querying only one archive could miss captures held elsewhere. Earlier research likewise found substantial country-level imbalance in Internet Archive coverage. Geography is therefore one source of unevenness among several.

Regional captures add a second record

The practical fix is comparison, not faith in one universal archive. National archives, institutional collections, and other public web archives may hold different snapshots of the same site because they crawl different scopes at different times and from different infrastructure. The Memento protocol was designed to help systems discover versions across multiple archives, although its original public Time Travel aggregator was discontinued in 2025.

What no archive can rebuild

Recovery has hard limits. If every archive that might have encountered a page was served a block page, and no alternative collector preserved the real version, there is no hidden master copy to reconstruct later.

Treat coverage claims the way you treat the robots.txt logic: suspicious. Before concluding that old content is lost, check whether it looks different from inside the country where it lived. The record of the web is not one archive — it is the sum of every vantage point that happened to look, including yours.