Posted on

The gap between anonymization claims and re-identification risk

Delete the name.

Delete the email address.

Delete the phone number.

The dataset may still describe somebody with uncomfortable precision.

That is the basic gap between removing obvious identifiers and making information genuinely difficult to link back to a person.

Identification can survive without a name field

Imagine a dataset containing:

  • age,
  • ZIP code,
  • workplace area,
  • timestamps,
  • repeated locations,
  • device type,
  • and purchase categories.

None of those fields has to say Leo Blanchette or Jane Smith.

But combinations can become distinctive.

A person who repeatedly appears at one residential location overnight and one workplace during weekdays may be easier to identify when another dataset contains addresses, employment records, or public information.

The second dataset supplies the missing label.

This is re-identification by linkage.

NIST treats de-identification as risk reduction, not a magic switch

The National Institute of Standards and Technology’s De-Identification of Personal Information explains that de-identification is intended to reduce privacy risk while preserving useful data, but also notes that researchers have shown some de-identified datasets can sometimes be re-identified.

NIST’s newer SP 800-188 guidance makes the same point operationally. Agencies are advised to evaluate the goals of release, the disclosure risks, the available de-identification methods, and the possibility of re-identification using auxiliary information.

That is a much more careful claim than:

We removed the names, therefore the data is anonymous.

Risk depends on what else exists

The same dataset can have different re-identification risk in different environments.

A record containing only an age range and broad region may be hard to link today.

Add exact timestamps, unusual travel patterns, or a publicly available social-media post describing the same event, and the picture changes.

This makes anonymity contextual.

The attacker—or researcher—does not have to work with the released dataset alone.

They can combine it with public records, commercial databases, breached information, maps, social posts, or another supposedly anonymous dataset.

Pseudonymous is not anonymous either

Replacing a name with an identifier such as user_847219 can be useful.

It stops every casual reader from seeing the person’s identity immediately.

But if a company maintains the lookup table connecting user_847219 to a real account, or if the identifier can be matched elsewhere, the record remains linkable.

That is pseudonymization, not disappearance.

The useful question is measurable risk

Good de-identification asks:

  • Which identifying attributes remain?
  • How unique are the combinations?
  • What outside datasets could be joined?
  • Who receives the data?
  • What technical and contractual limits exist?
  • Has anyone tested re-identification risk?

The Surveillance Economy benefits whenever anonymous is treated as a comforting adjective instead of a technical claim.

Removing the name is valuable.

It is not the same thing as removing the person.

Posted on

Data retention beyond the original reason for collection

A company can have a good reason to collect data today and a bad reason to keep it forever.

Those are separate questions.

A delivery service needs an address to deliver a package.

A fraud system may need transaction records long enough to investigate abuse.

A support desk may need logs while a problem is active.

The privacy issue changes when the original task ends but the information stays available indefinitely.

Storage creates future options

Data that no longer serves its original purpose can still be useful for something else.

Old location records may later support audience segmentation.

Old purchase histories may become training data.

Support logs may be mined for product analytics.

Account activity may be combined with newer records to create long-term behavioral profiles.

None of those later uses is guaranteed merely because the data was retained.

But retention makes them possible.

That is why retention policy is part of data governance rather than just storage housekeeping.

Current privacy rules increasingly connect retention to purpose

CalPrivacy’s current CCPA guidance says covered businesses must limit collection, use, and retention of personal information to purposes that are reasonably expected, compatible and disclosed, or separately agreed to, and that the activity must be reasonably necessary and proportionate. See CalPrivacy’s CCPA FAQ.

The agency’s enforcement guidance on data minimization makes the same principle explicit: retaining personal information is supposed to remain tied to the purposes for which it was collected or to another compatible disclosed purpose. See CalPrivacy Enforcement Advisory 2024-01.

Those rules are jurisdiction-specific.

The engineering principle is broader.

If nobody can explain why a dataset still exists, “we already have it” is not much of a retention policy.

A useful retention policy needs clocks, not poetry

Statements such as we retain information as long as necessary sound reassuring.

They become meaningful only when somebody has defined necessary.

A stronger policy identifies:

  • which dataset is retained,
  • for what purpose,
  • the retention period or decision rule,
  • what event starts the deletion clock,
  • what legal or security exceptions apply,
  • what happens to backups,
  • and who can approve a longer period.

Different data may need different clocks.

A chargeback record and a precise location history do not automatically deserve the same lifespan.

Old data attracts new explanations

The longer information survives, the more organizations, employees, products, and business models can change around it.

A dataset collected under one privacy policy can outlive the team that collected it.

A company can be acquired.

A product can be repurposed.

A vendor can change.

That does not mean old data is inevitably abused.

It means retention increases the number of future contexts in which the data might matter.

The Surveillance Economy does not only depend on collecting information.

It also depends on not throwing it away.

Posted on

Opt-out signals and the challenge of honoring them across intermediaries

A privacy signal can be one bit.

The supply chain behind it can contain dozens of companies.

That mismatch is the central problem with global opt-out mechanisms.

A browser can send a simple instruction such as:

Do not sell or share my personal information.

The difficult part is making sure that preference survives every handoff that follows.

The browser can express the choice once

Global Privacy Control, or GPC, is a browser-based signal designed to communicate an opt-out preference automatically to websites.

California regulators currently describe GPC as an opt-out preference signal that covered businesses must honor for applicable sale or sharing rights. In September 2025, California, Colorado, and Connecticut privacy regulators announced a coordinated sweep focused on businesses that may have failed to process GPC requests. See CalPrivacy’s enforcement announcement.

California also enacted the Opt Me Out Act in 2025, requiring browsers operating in California to provide built-in opt-out preference signals beginning in 2027. See CalPrivacy’s announcement.

That solves one usability problem.

The user does not need to hunt for a privacy link on every site.

One preference must cross many systems

Now imagine the website uses:

  • an analytics provider,
  • an ad exchange,
  • a demand-side platform,
  • a measurement vendor,
  • a data broker,
  • a server-side tag gateway,
  • and several downstream processors.

The original site receives the signal.

What happens next?

Does it suppress the relevant outbound event?

Does it attach an opt-out flag to downstream requests?

Do intermediaries understand the flag the same way?

Does a broker already holding a profile update its state?

Does the choice apply only to future sharing, or does it also affect retained data?

Those are implementation questions, not interface questions.

The ad industry itself recognizes the propagation problem

IAB Tech Lab’s current privacy standards portfolio explicitly focuses on communicating privacy preference signals through the digital advertising supply chain. Its Global Privacy Platform and related standards exist because a preference that stops at the visible website is not enough for a multi-party ecosystem. See IAB Tech Lab’s Privacy Standards.

That does not prove every intermediary honors every signal correctly.

It proves the industry understands that the signal needs a transport mechanism.

Compliance has to be tested in the data flow

A website can display Your privacy choices were saved while still sending data to outside services.

That does not automatically mean the transfer violates a law; some processing may remain permitted or necessary.

But the message cannot be evaluated from the button alone.

A real audit compares network behavior, server-side processing, vendor configuration, and downstream contracts before and after the preference is expressed.

The Surveillance Economy is full of settings that look local but govern distributed systems.

An opt-out is only as strong as the farthest system that is supposed to remember it.

Posted on

Data access requests as a way to inspect an assembled profile

The easiest way to misunderstand a data profile is to imagine it contains only things you personally typed into a form.

A real profile may be much stranger.

It can contain facts you volunteered, records acquired from other sources, device identifiers, inferred interests, household links, marketing segments, and internal labels you never saw.

That is why a data access request can be useful.

It turns an invisible profile into something inspectable.

Access can reveal the difference between supplied and inferred data

California’s current privacy guidance describes a consumer’s right to know what personal information covered businesses have collected and how they use and share it. See CalPrivacy’s CCPA FAQ.

Depending on the applicable law, company, and request, an access response may include categories or specific pieces of personal information, sources, purposes, or disclosures.

The most interesting comparison is often between what the person remembers supplying and what the company has assembled.

For example:

  • Volunteered: name, shipping address, email.
  • Observed: pages viewed, purchases, devices used.
  • Acquired: demographic or commercial attributes from another source.
  • Inferred: likely interests, household status, audience segment, predicted preference.

Those categories have different evidentiary meanings.

An inference is not automatically a fact merely because it appears in a company database.

The response can expose errors too

A profile might contain an outdated address, a device belonging to another household member, a purchase made as a gift, or an interest category inferred from one accidental click.

Without access, those errors can remain invisible while still influencing advertising, personalization, or other systems.

This connects directly to the data-broker problem. The Federal Trade Commission’s 2014 data broker report documented how brokers combine information from many sources into composite profiles and derive additional classifications from those records.

Access provides one way to inspect the result instead of merely guessing what the system knows.

An access response is not necessarily the whole backend

There are important limits.

Legal rights vary by jurisdiction. Exceptions can apply. Security-sensitive material may be withheld. A company may describe categories rather than expose every internal model. Data held by a separate company may require a separate request.

And a response from one organization does not reconstruct every copy that has already moved through the advertising or broker ecosystem.

So an access request should not be treated as a magical database dump.

It is evidence about a specific organization’s records and obligations at a specific time.

That is still valuable.

The Surveillance Economy is difficult to evaluate when every profile is hypothetical.

Access rights can turn at least part of the hypothesis into a document.

Posted on

Privacy policies that obscure the identities of data recipients

“We share information with trusted partners.”

That sentence may be true.

It may also tell you almost nothing.

Privacy policies often describe recipients by category: service providers, advertising partners, analytics companies, affiliates, vendors, processors, business partners, or other third parties.

Those labels can be useful.

They are not the same thing as knowing who actually receives the data.

A category explains a role, not an identity

Suppose a policy says location data may be shared with analytics and advertising partners.

The reader still does not know:

  • which companies,
  • how many companies,
  • whether the recipients change regularly,
  • whether those companies receive raw location or derived segments,
  • whether they can combine it with their own records,
  • whether they pass it onward.

The category gives the reader a general purpose.

It does not reconstruct the supply chain.

California’s privacy framework illustrates why both kinds of information matter. CalPrivacy describes a consumer’s right to know what personal information a covered business has collected and how it uses and shares that information. See CalPrivacy’s CCPA FAQ.

The practical value of that right depends heavily on how specifically the relationship can be described.

Vague language may be accurate and still hard to use

There are legitimate reasons policies use categories.

Vendors change. Large services may use hundreds of processors. Naming every infrastructure provider in the main policy can turn the document into a phone book that goes stale immediately.

The problem is not that categories exist.

The problem appears when the category is the only useful detail available about an important data flow.

“Business partners” can cover a lot of ground.

So can “service improvement.”

A reader trying to understand whether a location broker, ad exchange, cloud processor, fraud vendor, measurement company, or social platform receives the data may still be stuck.

Better transparency separates role from recipient

A more informative system can combine layers:

  • a plain-language explanation of the purpose,
  • the category of recipient,
  • a current vendor or subprocessors list,
  • the type of data each recipient gets,
  • and a change log when important relationships change.

That is more work than writing we may share data with partners.

It is also more useful.

The Federal Trade Commission’s long-running work on data brokers has repeatedly emphasized how difficult it can be for consumers to understand where information came from and where it travels once multiple intermediaries are involved. See the FTC’s Data Brokers: A Call for Transparency and Accountability.

The Surveillance Economy is a supply chain.

A privacy policy that names only categories may describe the boxes on the diagram while leaving all the arrows unlabeled.

Sometimes the most important privacy question is simply:

Who, exactly, got it?

Posted on

Location permissions requested for functions that do not obviously need them

A map asking for your location makes sense.

A flashlight asking for it raises a better question.

Permission systems are easiest to understand when the requested access maps directly onto the feature the person is trying to use.

Navigation needs location.

A camera needs camera access.

A voice recorder needs the microphone.

The interesting cases are the ones where the relationship is not obvious.

The scope matters as much as the permission name

“Location access” is not one thing.

Android currently distinguishes foreground from background location and approximate from precise location. Its developer documentation says apps should request only the type of location access critical to the user-facing feature, and that most use cases need location only while the user is actively engaging with the app. Background access is supposed to be reserved for cases where it is central to the function. See Android’s location-permission guide and background-location guidance.

That gives us a practical way to examine a permission request.

Do not ask merely:

Does this app use location?

Ask:

  • Does the feature need location at all?
  • Does it need precise location?
  • Does it need location continuously?
  • Does it need location when the app is closed?
  • Could a postal code or one-time location request do the job instead?

Those are materially different levels of access.

An unusual request is a clue, not proof of abuse

Suppose a shopping app asks for location.

That might support nearby-store inventory, local pickup, tax calculation, or fraud prevention.

Suppose a social app asks for location.

That might power nearby content or geotagging.

The request may have a legitimate explanation that is not obvious from the app icon.

Android’s own guidance even tells developers to inspect SDK dependencies because embedded libraries can themselves depend on location permissions.

So the presence of a permission is not enough to conclude that a company is secretly selling location histories.

A stronger investigation needs the actual data flow: network traffic, SDK documentation, privacy disclosures, regulator findings, or code showing where the location goes.

Minimization gives the question a measurable form

Android’s privacy documentation tells developers to minimize permission requests and prefer narrower methods where possible. Its newer Android 17 guidance continues that direction by emphasizing one-time and limited location access for common tasks that do not require permanent background tracking. See Android’s permission-minimization guidance.

That does not create a universal legal test.

It does create a useful engineering test:

Is the requested access proportional to the feature?

A weather app asking for approximate location while open is one thing.

The same app demanding precise background location forever is a different architecture.

The Surveillance Economy often hides in that difference between technically useful and actually necessary.

Posted on

Consent fatigue and the practical limits of repeated permission requests

A permission prompt can be meaningful.

The twentieth one before lunch may be less so.

Modern devices and websites ask people to make a constant stream of choices about cookies, notifications, location, microphones, cameras, contacts, advertising, personalization, tracking, analytics, and data sharing.

In theory, each prompt creates control.

In practice, attention is finite.

Repetition turns judgment into habit

A useful consent decision requires at least three things:

  • the person notices the request,
  • understands enough about what it means,
  • and cares enough in that moment to evaluate it.

Repeated prompts make all three harder.

A person trying to read an article may not want to become a part-time privacy lawyer first. Someone installing a transit app may simply need the bus schedule. If the interface repeatedly interrupts the task, the shortest path through the interruption becomes attractive.

That does not prove every acceptance is meaningless.

It does mean clicking Accept becomes a weaker signal of considered agreement when the environment trains people to click through routine interruptions.

Research presented at the FTC’s PrivacyCon has examined cookie-consent interfaces and found that design, timing, placement, and the mechanics of changing a decision affect how usable consent actually is. See “Okay, whatever”: An Evaluation of Cookie Consent Interfaces.

The title captures the problem rather well.

More prompts do not automatically create more control

Imagine a system that asks separately about:

  • analytics,
  • personalization,
  • advertising,
  • location,
  • notifications,
  • nearby devices,
  • background refresh,
  • and partner sharing.

Granularity can be valuable.

But if every feature generates a new interruption, the system can bury important choices inside a pile of trivial ones.

The person eventually learns a behavioral rule:

Make the popup go away.

That is not the same mental process as evaluating whether a company should collect precise background location for six months.

Reducing requests can improve the requests that remain

Android’s current privacy guidance tells developers to minimize permission requests and use alternatives when an app can perform a function without broad access. It also recommends requesting only the level of location access actually required. See Android’s permission-minimization guidance.

That design philosophy has a privacy benefit beyond reducing technical access.

It preserves attention for the moments when permission genuinely matters.

A camera app asking to use the camera is unsurprising.

A calculator asking for precise background location deserves more thought.

If every app asks for everything, the second case can begin to feel as routine as the first.

Consent fatigue is therefore not an argument against giving people choices.

It is an argument against treating human attention as infinite.

A permission system succeeds only when people can still tell which questions deserve an actual answer.

Posted on

Purposes bundled together inside a single data-use choice

One button can hide five decisions.

A privacy interface might ask whether you agree to personalized services.

That phrase can sound like one purpose.

Behind it may sit analytics, advertising, recommendation systems, audience measurement, data sharing, identity matching, fraud detection, and product research.

If all of those activities rise or fall with one switch, the user is not really deciding among purposes.

They are accepting a bundle.

A single choice can conceal unrelated uses

The problem is easiest to see when the visible feature and the secondary use are far apart.

A person may reasonably accept location so a weather app can show a local forecast.

That does not automatically mean they would make the same decision about using the location for advertising profiles.

Likewise, a customer may accept transaction processing while being less interested in behavioral advertising or cross-service personalization.

California’s current privacy guidance is useful here because it separates purposes conceptually. CalPrivacy says covered businesses must limit collection, use, and retention to purposes a consumer would reasonably expect, purposes compatible with those expectations and disclosures, or additional purposes the consumer actually agreed to. Collection and use must also be reasonably necessary and proportionate. See CalPrivacy’s CCPA FAQ.

That does not mean every multi-purpose setting is unlawful.

It does mean the purposes matter.

Bundling reduces information as well as choice

Suppose a privacy panel offers this:

Allow data use for improving services and personalized experiences.

That may sound harmless enough.

But a reader still needs to know what the system actually does:

  • Does “improving services” mean crash diagnostics?
  • Does “personalized” mean recommendations inside the app?
  • Does it mean targeted advertising across other services?
  • Are outside companies receiving the data?
  • Can analytics remain on while advertising stays off?

If the interface does not separate those questions, the user cannot express a more precise preference even if they understand the tradeoff.

The Federal Trade Commission’s report Bringing Dark Patterns to Light discusses interfaces that steer or obscure privacy choices, including designs that make meaningful refusal harder. Bundling is related but slightly different: the problem can exist even in a visually neutral interface if several distinct data uses are fused into one decision.

Granularity has costs too

There is an opposite failure mode.

A panel with 147 toggles can become useless homework.

Separating purposes does not require turning privacy into an aircraft cockpit.

The useful goal is understandable grouping: collection necessary for the service, optional measurement, optional personalization, optional advertising, optional sharing with outside parties, and other materially different uses.

The right amount of detail depends on the system.

But one principle survives the details:

A person cannot selectively consent to choices the interface never separates.

The Surveillance Economy often depends on turning many downstream uses into one convenient yes.

The clearer design asks which yes belongs to which purpose.

Posted on

Consent interfaces that make acceptance easier than refusal

A button can exist without being equally usable.

That is the problem with many consent interfaces.

The page may technically offer two outcomes:

Accept

and

Refuse.

But one choice can be a giant bright button on the first screen while the other requires opening settings, expanding categories, toggling switches, scrolling, and confirming again.

Both options exist.

They are not equally easy to exercise.

Effort is part of the choice

The Federal Trade Commission’s report Bringing Dark Patterns to Light describes privacy interfaces that highlight the option leading to more data collection while greying out, hiding, or adding extra steps to the option that limits collection. The report specifically discusses cookie interfaces where accepting is placed front and center while refusing or changing settings may require navigating additional screens. See Bringing Dark Patterns to Light.

A 2022 study presented at the FTC’s PrivacyCon evaluated cookie-consent interfaces and found that design choices such as blocking behavior, placement of controls, and the ability to later change a decision materially affect usability. See “Okay, whatever”: An Evaluation of Cookie Consent Interfaces.

The basic lesson is almost embarrassingly ordinary.

People are more likely to use the path that is easier to see and easier to finish.

Wording can steer without removing the option

Interface design does not need to hide the refusal button completely.

It can frame the choices differently.

One option might say:

Accept and continue

while the other says:

Manage preferences.

The first describes an outcome.

The second describes homework.

Color, size, placement, defaults, repeated prompts, and confusing category names can all affect how practical the alternatives feel.

In a 2024 international review involving the FTC and privacy and consumer-protection authorities, regulators reported that a majority of the 642 selected websites and apps examined used at least one potential dark pattern. The review did not conclude that every identified design violated the law, but it highlighted how interface interference and other techniques can steer users toward choices favorable to the business. See the FTC’s 2024 dark-pattern review announcement.

A visible option is not the same as a readily exercisable choice

This distinction matters because privacy discussions often collapse into a checkbox:

Was there a decline option? Yes or no?

That misses the interface around it.

A more useful audit asks:

  • How many clicks does acceptance require?
  • How many does refusal require?
  • Are both choices visible on the first screen?
  • Are they equally legible?
  • Are optional purposes preselected?
  • Does the interface keep asking after refusal?
  • Can the person later change the decision as easily as they made it?

Those are measurable properties.

They do not require guessing what the designer secretly intended.

Consent is an interaction, not a decorative legal layer

A system can have excellent disclosure text and still make the actual privacy-protective path unnecessarily difficult.

That is why consent quality cannot be judged only by whether the word privacy appears somewhere on the screen.

The Surveillance Economy often presents data collection as a choice.

The design of the choice determines how much that statement is worth.

A refusal button buried three screens deep is still a button.

It is also three screens deep.