Posted on

Sensitive inferences drawn from apparently ordinary browsing

A page about vitamins is not a medical record.

A page about churches is not a statement of faith.

A page about divorce is not proof of a divorce.

But enough ordinary-looking visits can still become a sensitive profile.

That is the difference between collected data and inferred data.

Ordinary signals can become intimate categories

The Federal Trade Commission’s 2024 report on major social-media and video-streaming companies warned that apparently ordinary interest categories can reveal or proxy much more sensitive information. The report specifically discussed risks when ad systems infer subjects such as sexual orientation, pregnancy, religion, political affiliation, or other sensitive characteristics. See A Look Behind the Screens.

The FTC had documented the same basic mechanism in its earlier data-broker study: brokers combined source records and behavioral information to create new consumer categories, including potentially sensitive segments involving health, age, ethnicity, income, religion, and other characteristics. See Data Brokers: A Call for Transparency and Accountability.

The important point is not that every visit produces a sensitive label.

It is that apparently mundane inputs can support an inference the person never supplied directly.

Inference is not confirmation

Suppose a device repeatedly visits:

  • an oncology information site,
  • a hospital parking page,
  • a cancer-support forum,
  • and a page about chemotherapy side effects.

A model may infer a cancer-related interest.

That inference might concern the device owner.

It might concern a spouse, parent, friend, client, student, research project, or news story.

The pattern can be informative without being definitive.

That uncertainty matters because a database label can look far more authoritative than the evidence that produced it.

Likely interested in X is not the same statement as X is a confirmed fact about this person.

Proxy categories can be sensitive too

A system does not need a field literally named religion to approximate religious affiliation.

It may have repeated visits to specific institutions, publications, holidays, products, events, or geographic locations.

Likewise, combinations of shopping, location, demographic, and browsing data can become proxies for income, pregnancy, health conditions, family status, or other traits.

The FTC’s 2024 report specifically noted that categories which are not sensitive on their own can sometimes combine into proxies for protected or sensitive characteristics.

That is a critical distinction.

Privacy is not only about whether a company collected a forbidden field.

It is also about what the company can derive from permitted ones.

Uncertain labels can still drive real decisions

An inference may influence which advertisement appears, which audience receives an offer, what content is recommended, or how a profile is valued by another system.

Even if the inference is wrong, the decision made from it is still real.

That is the strange asymmetry of the Surveillance Economy.

The machine can be uncertain about you and still act with confidence.