Posted on

The supply chain behind a company a user never knowingly contacted

You do not need to visit a company’s website for that company to have data about you.

That sounds mysterious until the supply chain is drawn.

A person visits a familiar app or website.

The service participates in an advertising system.

An exchange sends a bid request.

A company observing that auction receives data.

That company may then combine, retain, sell, or analyze the information.

At no point did the user intentionally knock on its door.

Mobilewalla provides a documented route

In 2024, the Federal Trade Commission alleged that data broker Mobilewalla collected consumer data from real-time advertising bidding exchanges and third-party aggregators. The FTC said consumers often had no knowledge that Mobilewalla had obtained their information. See the FTC’s Mobilewalla enforcement announcement.

According to the FTC complaint, Mobilewalla collected and retained information contained in bid requests even when it did not win the auction. The agency alleged that, between January 2018 and June 2020, the company collected more than 500 million unique advertising identifiers paired with precise location data.

The FTC finalized the order in January 2025. See the final order announcement.

That gives us a supply chain with real documentation rather than speculation.

Each intermediary sees a different slice

The route can look roughly like this:

  1. A person uses a website or app.
  2. Advertising inventory becomes available.
  3. A real-time bidding exchange distributes a bid request.
  4. The request contains information useful for evaluating the ad opportunity.
  5. A bidder or data company receives that information.
  6. The recipient may combine it with other records or sell derived products downstream.

Not every advertising auction contains the same fields.

Not every participant retains the data.

Not every recipient uses it for profiling.

Those details have to be verified case by case.

The important architectural fact is that direct contact is not required.

The visible service is only the first hop

Consumers naturally think in terms of brands they recognize.

I gave this app my location.

I visited this website.

I used this store.

But modern data systems often operate through chains of processors, SDK vendors, ad exchanges, analytics companies, identity providers, brokers, and clients.

One company can therefore know about a person because another company generated the event and a third company passed it along.

That is why privacy policies full of phrases such as partners, service providers, and advertising companies can be difficult to evaluate without actual recipient names and technical evidence.

The Surveillance Economy becomes hardest to see at the point where the person and the data holder have never met.

The person knows the app.

The broker knows the identifier.

The supply chain introduces them without an introduction.

Posted on

Children’s activity profiles persisting into adult life

Children age faster than databases do.

A ten-year-old becomes a teenager.

A teenager becomes an adult.

The row in a database may simply remain a row.

That creates a basic privacy problem: information collected in childhood can survive beyond the context that made its collection seem reasonable.

Retention turns childhood activity into history

A child’s online activity may be recorded by educational tools, games, video services, connected toys, apps, websites, advertising systems, or family accounts.

A single event may be harmless.

Years of events can become a profile of interests, habits, devices, locations, purchases, learning activity, and inferred preferences.

The risk is not that every company necessarily keeps all of that forever.

The risk is that persistent systems need an explicit rule for when childhood data stops being useful enough to retain.

In January 2025, the Federal Trade Commission finalized changes to the U.S. Children’s Online Privacy Protection Rule. Among other changes, the updated rule says covered operators may retain children’s personal information only as long as reasonably necessary for the specific purpose for which it was collected and explicitly says it cannot be retained indefinitely. See the FTC’s 2025 COPPA rule announcement.

That requirement applies to covered services and children under the rule’s scope.

The broader design lesson is about expiration.

Age transitions are not automatic deletion events

A service may know a birth date.

That does not necessarily mean every connected analytics system, advertising partner, or historical dataset automatically receives an instruction saying:

This person is now older. Reconsider the old record.

A profile assembled when somebody was eleven can therefore remain technically linkable after the person becomes thirteen, sixteen, or eighteen unless the system deliberately changes how the data is treated.

Old labels can also become misleading.

Children change quickly. Interests, family circumstances, schools, devices, and identities evolve. A classification produced years earlier may persist long after its predictive value collapses.

The key questions are about lifecycle

A useful audit asks:

  • What childhood data is retained?
  • What purpose still requires it?
  • Is there a maximum retention period?
  • What happens when the user crosses an age threshold?
  • Are old advertising or inference labels deleted, reset, or merely carried forward?
  • Do downstream recipients receive the same deletion or age-transition signal?

The FTC’s revised COPPA rule also requires separate parental consent for certain disclosures related to targeted advertising, reinforcing the difference between providing a child-facing service and monetizing the resulting profile.

The Surveillance Economy usually talks about collection as a moment.

Allow or Don’t Allow.

Children’s data exposes the missing second half of the problem.

A collection decision lasts seconds.

A profile can last years.

If a system has no answer for when childhood ends inside the database, the database may keep remembering a person who no longer exists.

Posted on

Health-related browsing as a source of commercially valuable inference

Reading about a disease does not mean you have it.

You could be researching for a relative.

You could be writing an article.

You could be curious.

A tracking system may still find the visit valuable.

Browsing can reveal a probable concern without proving a condition

Health-related websites can expose unusually sensitive context simply through the pages a person visits.

A URL, page title, tracking event, search term, or advertising pixel may indicate that somebody viewed information about depression, fertility, cancer, addiction, medications, sexual health, or another medical subject.

The Federal Trade Commission’s technology staff has highlighted this problem directly. In a 2023 analysis of pixel tracking following the GoodRx and BetterHelp enforcement actions, the FTC explained that embedded tracking technologies can collect user activity and enable companies to analyze and infer information from it. See Lurking Beneath the Surface: Hidden Impacts of Pixel Tracking.

The signal is meaningful.

It is not necessarily true in the way a medical record is true.

That difference matters.

A visit can be commercially useful even when it is ambiguous

The FTC’s BetterHelp case alleged that the service shared personal information with advertising platforms for targeting and retargeting, including information connected to people who visited the service without necessarily becoming counseling customers. See the FTC’s BetterHelp enforcement announcement.

The GoodRx case went further into confirmed health information, with the FTC alleging that medication and health-condition data was shared with advertising companies and used to target users. See the FTC’s GoodRx case.

Those cases are not identical to merely reading a health article.

They show why the advertising value exists: health-related signals can be turned into targeting categories, audience lists, or assumptions about what a person might respond to.

Inference is where the danger becomes slippery

Suppose somebody reads three pages about insulin pumps.

An advertising system might infer diabetes-related interest.

That inference could be correct.

It could also describe a parent, nurse, student, journalist, investor, or person helping a friend.

Once an inferred label enters a profile, downstream systems may not preserve that uncertainty. A probabilistic guess can start behaving like a fact because a database field needs something to put in the row.

That is the important distinction for Dead Internet Theory research:

behavioral evidence is not the same thing as a confirmed personal attribute.

The Surveillance Economy does not need to know your diagnosis with clinical certainty.

It only needs to believe a health-related label is predictive enough to sell, target, rank, or segment.

A browser history can therefore become commercially valuable long before it becomes medically meaningful.

Posted on

Education platforms and the collection of student activity

A classroom platform cannot function without recording some student activity.

Assignments have to be submitted.

Progress has to be saved.

Teachers need to know who completed the work.

The privacy problem begins when data collected for school starts serving a different business.

Educational use and commercial use are not the same purpose

Digital learning systems can record logins, assignment submissions, quiz results, discussion participation, timestamps, device information, and other activity needed to operate the service.

That can support obvious educational functions: grading, feedback, attendance, troubleshooting, and progress reporting.

In 2023, the Federal Trade Commission obtained an order against education-technology provider Edmodo after alleging that it collected personal information from children and used persistent identifiers for advertising. The FTC said Edmodo improperly relied on schools and teachers for parts of its COPPA compliance while failing to give them the information needed to make those decisions. See the FTC’s Edmodo case page and enforcement announcement.

The useful distinction is not data collection versus no data collection.

It is collection for the educational service versus secondary use for something else.

Learning analytics can become a behavioral record

A single quiz score says little.

A long record can say much more.

Repeated logs may reveal when a student studies, which subjects are difficult, how quickly assignments are completed, which resources are opened, and how participation changes over time.

Those records may be useful to a teacher.

They can also become sensitive if they are retained indefinitely, combined across services, used for advertising, or exposed to people who do not need them for education.

That is why context matters.

A student may reasonably expect a math platform to remember completed math exercises.

That expectation does not automatically extend to building an advertising profile.

Families need the data relationship explained in plain language

A useful disclosure should make several things clear:

  • What student activity is recorded?
  • Which information is necessary to provide the educational feature?
  • Does the school control the account or does the vendor?
  • Are third-party services embedded?
  • Is data used for advertising, product development, analytics, or other secondary purposes?
  • When does the information get deleted after a class, account, or school relationship ends?

For children under 13 in the United States, COPPA adds specific legal requirements for covered services. Other student-privacy rules can also apply depending on the institution and jurisdiction.

But even before the law enters the room, the architectural question is simple.

The student entered the system to learn.

If the data later serves another purpose, that second purpose should not be smuggled inside the word education.

Posted on

Employer monitoring inside browser-based work tools

A green dot is not productivity.

Neither is a moving mouse.

Modern work tools can generate enormous amounts of activity data because so much work now happens through browsers, cloud dashboards, chat systems, ticket queues, document editors, and remote-management software.

That makes workplace measurement easy.

It does not make interpretation easy.

Work software can observe a lot

The UK Information Commissioner’s Office lists modern monitoring methods that include screenshots, webcam use, timekeeping, keystroke logging, productivity tools, internet activity, and location tracking. The ICO also notes that employers increasingly use analytics to infer worker performance and wellbeing. See the ICO’s guidance on monitoring workers.

Some of those tools have obvious legitimate purposes.

Security teams may need audit logs. Employers may need time records. Regulated businesses may have retention obligations. A support organization may need to know whether tickets are being answered.

The surveillance problem appears when a measurable signal becomes a substitute for the thing the employer actually cares about.

Activity is a proxy

Suppose one employee spends four hours reading specifications and solving a difficult problem, then writes twenty lines of code.

Another generates eight hours of constant keyboard and mouse activity while producing little useful work.

A crude activity score may prefer the second employee.

The software did not malfunction.

The metric did exactly what it was designed to do.

The mistake was treating observable motion as equivalent to valuable output.

The same problem appears with message counts, meeting hours, browser tabs, login duration, ticket volume, and response speed.

Those numbers can be useful in context.

They become dangerous when they are interpreted as direct measurements of effort, quality, judgment, creativity, or usefulness.

Workers need to know the boundaries

A monitoring system is easier to evaluate when workers can answer basic questions:

  • What is being recorded?
  • Is monitoring continuous or occasional?
  • Who can inspect the records?
  • Is the data used for security, payroll, performance evaluation, discipline, or several purposes?
  • How long is it retained?
  • Does monitoring extend to personal devices or off-hours activity?

The ICO’s guidance emphasizes lawfulness, fairness, transparency, necessity, and proportionality in workplace monitoring under UK data-protection law.

Those legal rules vary by jurisdiction.

The engineering lesson does not.

A browser-based workplace can become an unusually detailed observation environment because the work itself is already digital.

The important distinction is between recording signals that are genuinely needed and pretending every recorded signal is a trustworthy measure of the worker.

Sometimes the computer knows you clicked.

It still does not know whether the click was good work.

Posted on

Pay-for-privacy offers and the unequal ability to avoid surveillance

A privacy option can be perfectly visible and still be unaffordable.

That is the tension inside consent-or-pay models.

The user is offered two versions of a service:

  • accept certain personal-data processing, often for behavioral advertising, or
  • pay for an alternative that avoids some of that processing.

At first glance, this looks cleaner than a hidden tracker.

The trade is explicit.

The harder question is whether everybody has the same practical ability to choose.

Price changes the meaning of the choice

In April 2024, the European Data Protection Board issued an opinion on consent-or-pay models used by large online platforms. The EDPB said those platforms should provide a real choice and warned that, in most cases, presenting only two options—consent to behavioral-advertising processing or pay a fee—will not be enough by itself to demonstrate valid consent under the GDPR. See the EDPB’s summary and Opinion 08/2024.

The opinion is specifically about large online platforms under European data-protection law.

It should not be stretched into a universal rule for every subscription product everywhere.

But the underlying economic problem travels well.

If one person can easily pay and another cannot, the same privacy setting has a different practical cost for each of them.

“No ads” and “no tracking” are not the same product

A paid tier also needs to be examined carefully.

It may remove advertisements while retaining analytics.

It may disable behavioral targeting while still collecting security logs, service telemetry, account information, or measurements needed to operate the product.

It may reduce sharing with advertising partners without deleting historical profiles already held elsewhere.

So the phrase pay for privacy is too broad unless the service explains exactly what changes.

Useful questions include:

  • Which processing stops?
  • Which identifiers stop being created or shared?
  • Does historical data remain?
  • Are ads removed, or merely made non-personalized?
  • Does the paid version still use analytics or fraud-detection signals?

The label matters less than the data flow.

Privacy can become a luxury feature

The EDPB’s opinion explicitly warned against turning data protection into a premium feature available only to those willing or able to pay.

That does not mean every paid privacy feature is illegitimate.

A service has costs. Subscription revenue can fund a product that collects less advertising data. Some users may genuinely prefer that exchange.

The inequality appears when surveillance is effectively the default price for people who cannot afford the alternative.

The Surveillance Economy becomes especially visible at that moment.

Tracking is no longer hidden in a pixel.

It has a dollar value printed beside the button.

The interesting question is not merely whether the user was offered a choice.

It is what the choice actually costs, and what privacy the payment really buys.

Posted on

Contextual advertising compared with behavioral targeting

An advertisement for hiking boots on a hiking article does not need to know who you are.

It needs to know what the page is about.

That is the simplest form of contextual advertising.

Behavioral targeting works from a different question.

Instead of asking What is this page about?, it asks something closer to What does this person appear to be interested in?

Context can be enough

Google Ads describes contextual targeting as matching ads to content using topics, placements, keywords, language, and other signals about the page or placement. See Google Ads’ contextual targeting documentation.

In the clean conceptual version, a gardening page can show gardening ads because the page is about gardening.

The advertiser does not need a six-month browsing history to make that decision.

Behavioral advertising instead uses information associated with a person or device: prior browsing, app activity, searches, purchases, inferred interests, demographic assumptions, or membership in audience segments.

That requires some mechanism for remembering or recognizing the audience.

The labels can blur in real systems

This is where terminology becomes slippery.

A product may call part of its system contextual while also using other signals such as location, recent activity, account information, or campaign optimization.

Google’s own contextual-targeting documentation notes that several factors can contribute to ad selection, including content and, in some configurations, a visitor’s recent browsing history.

So contextual should not be treated as a magic privacy label.

The useful research question is narrower:

What information was actually used to choose this ad?

If the decision came only from the content and immediate context, the system needs far less personal history.

If it depended on a durable profile, it has crossed into behavioral targeting even if the page topic also mattered.

Less profiling does not guarantee better or worse advertising

Contextual advertising has obvious limits.

A page about cameras may be read by somebody who already bought one yesterday. A behavioral system may know that and choose something else.

On the other hand, behavioral profiles can be stale, wrong, or based on somebody else using the same device.

Performance therefore depends on the publisher, advertiser, product, measurement method, and implementation.

It is not responsible to declare one method universally superior.

The privacy difference is easier to state.

Contextual targeting can often operate with much less long-term personal data because the page supplies the targeting signal.

Behavioral targeting needs the person or device to supply it.

That is a significant architectural choice.

The Surveillance Economy often treats persistent identity as the natural price of advertising.

Contextual advertising is evidence that advertising can exist without making every reader carry a dossier into the auction.

Posted on

Audience measurement that can operate with less personal data

A website owner usually wants answers to fairly boring questions.

How many people visited?

Which pages were popular?

Did anyone click the new navigation link?

Did traffic come from search, a newsletter, or another site?

None of those questions automatically requires building a durable advertising identity for every visitor.

Measurement and profiling are different jobs

Analytics becomes much more invasive when the unit of analysis changes from what happened on this site to what this person does across many sites and devices.

A publisher can often measure page views, broad traffic sources, session counts, device classes, or conversion totals without trying to recognize the same person everywhere else on the web.

France’s data-protection authority, CNIL, provides a useful concrete model. Its guidance allows certain audience-measurement trackers to qualify for a consent exemption only under restrictive conditions: the purpose must stay limited to audience measurement or A/B testing, the data must not be cross-checked with unrelated customer files or visits to other sites, the tracker must remain scoped to a single publisher, IP addresses must be truncated, and tracker lifetimes are limited. See CNIL’s guidance on audience measurement.

That is not the only possible privacy-preserving design.

It is useful because it shows the engineering principle clearly: collect enough to answer the measurement question, but do not quietly turn analytics into an identity business.

Aggregate answers lose some detail

Collecting less has tradeoffs.

A system that refuses to create persistent user histories may be worse at answering questions such as:

  • Did the same person return six months later?
  • Which advertisements did this exact user see before subscribing?
  • How does one person’s behavior compare across several unrelated properties?

Those can be commercially useful questions.

They are also the questions that require more durable identity.

A publisher therefore has to separate what it genuinely needs from what is merely interesting because technology makes it possible.

More data is not automatically better measurement

Individual-level histories can create their own errors.

People clear cookies. Families share devices. One person uses several browsers. Privacy tools isolate identifiers. Automated traffic contaminates logs. Cross-device systems make probabilistic matches that may be wrong.

A giant profile can look precise while containing bad joins.

Aggregate measurement has uncertainty too, but at least the uncertainty is closer to the question being asked.

If the question is How many times was this article read?, a system does not necessarily need to answer Who else lives with this reader?

That distinction matters at the end of the Surveillance Economy section.

The choice is not between perfect analytics and total blindness.

There is a large middle ground where websites can measure their own performance without insisting on remembering everybody everywhere.

Posted on

Advertising fraud detection as a competing rationale for tracking

Not every tracking signal exists to sell you shoes.

Some exist because somebody is trying to steal the shoe advertiser’s money.

Digital advertising has a genuine fraud problem: bots can generate fake impressions, automated systems can click ads, fraudulent publishers can misrepresent inventory, and traffic can be manipulated to look human long enough to get paid.

Detecting that abuse requires observation.

The privacy question is what happens after the observation is collected.

Fraud detection needs signals

A fraud system may examine patterns such as:

  • unusually rapid clicks,
  • impossible or suspicious traffic volumes,
  • repeated activity from one device or network,
  • abnormal timing,
  • whether an ad was actually viewable,
  • whether a request came through an authorized advertising supply path,
  • signs that a browser or device is automated.

IAB Tech Lab’s current Security & Fraud work includes standards such as ads.txt, sellers.json, SupplyChain objects, and ads.cert, all intended to make the programmatic advertising supply chain harder to impersonate or manipulate.

Its Open Measurement SDK also exists partly to verify whether advertising impressions were genuinely viewable and measurable across apps, connected TV, and online video.

Those are legitimate protective functions.

The same observation can support several purposes

Now comes the difficult part.

A signal useful for detecting fraud can also be useful for identifying or profiling a user.

IP addresses can help spot impossible traffic patterns.

Device characteristics can distinguish a real phone from an emulator.

Behavioral timing can help identify automated clicking.

Persistent identifiers can help detect one entity creating thousands of supposedly independent events.

The fact that a signal is useful for security does not mean every later use is automatically justified by the security purpose.

That is the boundary worth inspecting.

The industry is already trying to separate anti-fraud from identity

IAB Tech Lab’s ads.cert roadmap is revealing here. Its work on authenticated devices describes a goal of allowing devices to attest that requests are legitimate in ways that help with invalid-traffic and anti-fraud efforts without enabling user tracking. See IAB Tech Lab’s ads.cert documentation.

Its newer ID-Less Solutions Guidance likewise discusses how advertising systems can perform functions such as measurement, frequency control, and fraud detection in environments where conventional identifiers are unavailable.

That matters because it separates two questions that are often lazily merged:

Do we need signals to detect fraud?

and

Do we need a persistent behavioral identity for that purpose?

The answer is not always the same.

Purpose limitation is the useful audit

A fraud-detection system should be evaluated by what it actually collects and what happens next.

Useful questions include:

  • Is the signal retained only as long as needed for fraud analysis?
  • Is it reused for ad targeting or audience enrichment?
  • Is it shared with unrelated recipients?
  • Can the same protective result be achieved with less identifying data?
  • Are aggregated or privacy-preserving signals available instead?

The existence of fraud is not a fictional excuse.

It is a real engineering problem.

But a real protective purpose does not turn every observation into a blank check.

The Surveillance Economy becomes harder to analyze when all tracking is treated as morally identical.

The better question is narrower:

What was this signal collected to accomplish, and did the system stop there?

Posted on

The gap between anonymization claims and re-identification risk

Delete the name.

Delete the email address.

Delete the phone number.

The dataset may still describe somebody with uncomfortable precision.

That is the basic gap between removing obvious identifiers and making information genuinely difficult to link back to a person.

Identification can survive without a name field

Imagine a dataset containing:

  • age,
  • ZIP code,
  • workplace area,
  • timestamps,
  • repeated locations,
  • device type,
  • and purchase categories.

None of those fields has to say Leo Blanchette or Jane Smith.

But combinations can become distinctive.

A person who repeatedly appears at one residential location overnight and one workplace during weekdays may be easier to identify when another dataset contains addresses, employment records, or public information.

The second dataset supplies the missing label.

This is re-identification by linkage.

NIST treats de-identification as risk reduction, not a magic switch

The National Institute of Standards and Technology’s De-Identification of Personal Information explains that de-identification is intended to reduce privacy risk while preserving useful data, but also notes that researchers have shown some de-identified datasets can sometimes be re-identified.

NIST’s newer SP 800-188 guidance makes the same point operationally. Agencies are advised to evaluate the goals of release, the disclosure risks, the available de-identification methods, and the possibility of re-identification using auxiliary information.

That is a much more careful claim than:

We removed the names, therefore the data is anonymous.

Risk depends on what else exists

The same dataset can have different re-identification risk in different environments.

A record containing only an age range and broad region may be hard to link today.

Add exact timestamps, unusual travel patterns, or a publicly available social-media post describing the same event, and the picture changes.

This makes anonymity contextual.

The attacker—or researcher—does not have to work with the released dataset alone.

They can combine it with public records, commercial databases, breached information, maps, social posts, or another supposedly anonymous dataset.

Pseudonymous is not anonymous either

Replacing a name with an identifier such as user_847219 can be useful.

It stops every casual reader from seeing the person’s identity immediately.

But if a company maintains the lookup table connecting user_847219 to a real account, or if the identifier can be matched elsewhere, the record remains linkable.

That is pseudonymization, not disappearance.

The useful question is measurable risk

Good de-identification asks:

  • Which identifying attributes remain?
  • How unique are the combinations?
  • What outside datasets could be joined?
  • Who receives the data?
  • What technical and contractual limits exist?
  • Has anyone tested re-identification risk?

The Surveillance Economy benefits whenever anonymous is treated as a comforting adjective instead of a technical claim.

Removing the name is valuable.

It is not the same thing as removing the person.