Posted on

People-search sites and the recurring work of removal

Removing yourself from a people-search site can feel like deleting a file.

It is often closer to pulling one weed while leaving the roots, seeds, and neighboring gardens alone.

People-search services compile information from public records, public profiles, commercial sources, and other data brokers. That means the visible profile is frequently not the original source.

One opt-out removes one copy

The Federal Trade Commission’s current consumer guidance explains that most people-search sites offer some form of opt-out process. A person may request that a site stop selling a report containing their address, phone number, relatives, or other information. See What To Know About People Search Sites That Sell Your Information.

That can be genuinely useful.

But the FTC also warns that opting out of one people-search site does not delete the underlying public records and may not remove information from relatives’ or associates’ profiles.

The source material still exists.

Other brokers may still possess copies.

Other lookup services may still publish their own versions.

Information can come back

The FTC notes another frustrating detail: if underlying public-record information changes, a person’s information can reappear for sale later.

That makes removal less like a permanent account deletion and more like maintenance.

A person can opt out today, move next year, trigger a new address record, and later discover a fresh profile assembled from updated sources.

The service may have honored the original removal request exactly as promised.

The ecosystem can still recreate the dossier from new inputs.

Verification can require more information

Removal itself can create an awkward privacy problem.

A site needs some way to distinguish the real subject of a profile from a stranger trying to alter somebody else’s record.

That may require a name, email address, phone number, profile URL, or another identifying detail.

The FTC specifically notes that opt-out processes may require information to verify identity.

So the person trying to reduce a data trail may have to provide data in order to prove which trail is theirs.

That does not make verification illegitimate.

It illustrates the asymmetry.

The broker already has the profile.

The individual has to find the broker, find the profile, learn the procedure, submit the request, complete verification, and later check whether the profile returned.

The burden scales badly

If ten services have ten copies, one person may face ten separate procedures.

If the information later propagates to new services, the process begins again.

Paid removal companies exist largely because repeated manual opt-outs are tedious enough to become a market of their own.

That is an unusual feature of the Surveillance Economy.

Collection is often automatic.

Removal is often manual.

Your data can move through brokers without you visiting their sites.

Your request to stop that movement usually cannot.

Posted on

Public records repackaged into searchable personal profiles

There is a difference between information being public and information being convenient.

A property record at a county office is public.

A marriage record may be public.

A professional license, court filing, voter-registration record, or prior address may be public.

Finding all of them used to require knowing where to look.

Aggregation changes that.

Scattered facts become one object

The Federal Trade Commission’s current consumer guidance on people-search services says these sites may collect information from federal, state, and local public records, public social-media profiles, and other data brokers, then compile the material into reports sold to users. See What To Know About People Search Sites That Sell Your Information.

The FTC lists examples including property records, driving records, voter-registration data, civil and criminal records, birth and marriage records, professional licenses, addresses, relatives, and other identifying information.

Each source can have a legitimate public purpose.

The privacy effect changes when they are recombined.

Instead of visiting five agencies, searching multiple databases, understanding jurisdiction, and manually deciding which John Smith is which, a person may type one name into one box.

That is not merely storage.

It is discoverability engineering.

Aggregation changes practical exposure

Suppose an old address exists in a county record from twelve years ago.

Technically, it was already public.

But a stranger would have needed to know the county, locate the right database, understand its search interface, distinguish the correct person, and connect the record to a current identity.

A people-search profile can perform much of that work in advance.

It may also place the old address beside a current phone number, relatives, approximate age, previous names, and newer addresses.

No single field had to become newly public.

The relationship between the fields became newly convenient.

Recombination can reveal more than the source records intended

Public records are usually created for specific administrative reasons.

Property records help document ownership and taxation.

Professional licenses document authorization to practice.

Court records document judicial proceedings.

Voter records support election administration.

When those records are copied into a commercial profile, the context changes from administration to lookup.

That can be useful for finding old friends, verifying identities, or researching a business contact.

It can also make stalking, harassment, unwanted contact, or casual snooping easier.

The FTC specifically notes that people facing stalking or domestic violence may have strong reasons not to want home addresses or family-member information made easily searchable.

“Already public” is not the whole privacy analysis

A common response to aggregation concerns is simple:

But this information was public anyway.

That is factually relevant.

It is not the end of the analysis.

Searchability, scale, linkage, indexing, copying cost, and access friction all affect practical privacy.

Ten dusty filing cabinets and one searchable dossier can contain the same facts while producing very different exposure.

The Surveillance Economy frequently creates value not by inventing new information, but by making old information much easier to connect.

Posted on

Cross-device matching and errors in household attribution

Your television does not necessarily belong to the person holding the phone beside it.

Advertising systems sometimes have to make that guess anyway.

Cross-device matching tries to connect activity from several internet-connected devices to the same person or household.

When the match is right, separate histories become one story.

When it is wrong, somebody else’s behavior can become part of yours.

Devices can be linked directly or inferred

The Federal Trade Commission’s staff report on cross-device tracking describes two broad approaches.

Deterministic matching uses stronger direct evidence, such as the same person logging into one account across multiple devices.

Probabilistic matching estimates relationships using signals that may include IP addresses, device information, location, browsing patterns, or other observations. See Cross-Device Tracking: An FTC Staff Report.

Both approaches can be useful.

They do not have the same error profile.

A shared login is evidence that the same account appeared on two devices.

A shared household network is evidence that two devices used the same connection.

Those statements are different.

Households are hostile environments for neat attribution

Imagine four people in one home.

They share:

  • a television,
  • a game console,
  • home Wi-Fi,
  • a family tablet,
  • streaming accounts,
  • and occasionally each other’s laptops.

One household member researches motorcycles.

Another shops for maternity clothes.

A teenager watches hours of gaming videos on the living-room TV.

A parent uses the same tablet to research retirement accounts.

A matching system that collapses those devices too aggressively can produce a fictional super-person who is simultaneously pregnant, retiring, buying a motorcycle, and speedrunning Elden Ring.

The database may be internally consistent.

The human it describes may not exist.

Shared IP addresses are not identities

A household router makes many devices appear to the outside world through one public IP address.

Workplaces, schools, libraries, hotels, apartment networks, mobile carriers, and VPNs can create even larger shared environments.

That makes network proximity useful as a signal but dangerous as a conclusion.

LiveRamp’s current identity-resolution documentation explicitly recognizes that shared device touchpoints can map to more than one persistent identity. See LiveRamp’s RampID identity-resolution documentation.

That is an important admission built into the machinery itself: one technical identifier does not always equal one human.

Attribution errors propagate

Once two devices are joined, later systems may inherit the connection.

Advertising measurement can credit the wrong device. Audience segments can absorb another person’s interests. Recommendations can become strange. A household-level campaign can be mistaken for person-level evidence.

This is not an argument that cross-device matching never works.

It is an argument that the resolution level matters.

Person.

Household.

Device.

Account.

Those are different objects.

The Surveillance Economy becomes unreliable when its databases quietly slide between them as though they were the same thing.

Posted on

Lookalike audiences and profiling people through other people’s behavior

You do not have to behave like a customer to be targeted like one.

Sometimes it is enough to resemble customers who came before you.

That is the logic behind lookalike, similar, and predictive audiences.

An advertiser starts with a source group: buyers, leads, subscribers, site visitors, or some other known audience.

The platform then looks for other people who share useful characteristics with that group.

The seed audience teaches the model what to search for

Google’s current Demand Gen documentation describes Lookalike segments as groups of people who share characteristics with members of an existing seed list. Advertisers can use customer lists, website visitors, app users, and other first-party sources to help find new customers who resemble people they already know. See Google Ads’ Demand Gen audiences overview.

LinkedIn uses a related concept called predictive audiences. Its current documentation says the system combines a source such as a contact list, conversion audience, lead form, or retargeting group with LinkedIn’s AI to generate a new audience predicted to perform similar actions. See LinkedIn’s predictive audience documentation.

The terminology changes.

The underlying move is similar.

Behavior from Group A helps decide who belongs in Group B.

The new person may never have volunteered the defining interest

Suppose a business uploads a list of customers who bought expensive camping equipment.

The advertising platform may identify patterns among those customers and find other users who statistically resemble them.

Those new users may never have visited the retailer.

They may never have searched for a tent.

They may never have said they enjoy camping.

Their eligibility can come from similarities discovered in data available to the platform.

That is why lookalike profiling is different from ordinary retargeting.

Retargeting says:

This person interacted with us before.

Lookalike modeling says:

This person resembles people who interacted with us before.

Resemblance is not intent

Statistical similarity can be commercially useful without being a personal truth.

Two people may share age, geography, browsing patterns, device usage, media habits, professional traits, or other attributes while wanting completely different things.

A model optimized for conversions does not need to prove that a new prospect has the same motive as the seed audience.

It only needs the resemblance to improve campaign performance often enough to be useful.

That distinction matters when interpreting the profile.

A person being placed into a lookalike audience does not establish that they hold the interests, needs, politics, health status, financial situation, or intentions of the seed group.

Profiling can propagate through other people

This is one of the stranger properties of modern advertising.

Your own behavior is not the only behavior that can shape what systems infer about you.

Other people’s conversion histories can become training examples that affect whether you are selected.

The Surveillance Economy therefore does not merely watch individuals.

It compares them.

A profile can be influenced by what you did.

A predictive profile can be influenced by what people who resemble you did.

Posted on

Sensitive inferences drawn from apparently ordinary browsing

A page about vitamins is not a medical record.

A page about churches is not a statement of faith.

A page about divorce is not proof of a divorce.

But enough ordinary-looking visits can still become a sensitive profile.

That is the difference between collected data and inferred data.

Ordinary signals can become intimate categories

The Federal Trade Commission’s 2024 report on major social-media and video-streaming companies warned that apparently ordinary interest categories can reveal or proxy much more sensitive information. The report specifically discussed risks when ad systems infer subjects such as sexual orientation, pregnancy, religion, political affiliation, or other sensitive characteristics. See A Look Behind the Screens.

The FTC had documented the same basic mechanism in its earlier data-broker study: brokers combined source records and behavioral information to create new consumer categories, including potentially sensitive segments involving health, age, ethnicity, income, religion, and other characteristics. See Data Brokers: A Call for Transparency and Accountability.

The important point is not that every visit produces a sensitive label.

It is that apparently mundane inputs can support an inference the person never supplied directly.

Inference is not confirmation

Suppose a device repeatedly visits:

  • an oncology information site,
  • a hospital parking page,
  • a cancer-support forum,
  • and a page about chemotherapy side effects.

A model may infer a cancer-related interest.

That inference might concern the device owner.

It might concern a spouse, parent, friend, client, student, research project, or news story.

The pattern can be informative without being definitive.

That uncertainty matters because a database label can look far more authoritative than the evidence that produced it.

Likely interested in X is not the same statement as X is a confirmed fact about this person.

Proxy categories can be sensitive too

A system does not need a field literally named religion to approximate religious affiliation.

It may have repeated visits to specific institutions, publications, holidays, products, events, or geographic locations.

Likewise, combinations of shopping, location, demographic, and browsing data can become proxies for income, pregnancy, health conditions, family status, or other traits.

The FTC’s 2024 report specifically noted that categories which are not sensitive on their own can sometimes combine into proxies for protected or sensitive characteristics.

That is a critical distinction.

Privacy is not only about whether a company collected a forbidden field.

It is also about what the company can derive from permitted ones.

Uncertain labels can still drive real decisions

An inference may influence which advertisement appears, which audience receives an offer, what content is recommended, or how a profile is valued by another system.

Even if the inference is wrong, the decision made from it is still real.

That is the strange asymmetry of the Surveillance Economy.

The machine can be uncertain about you and still act with confidence.

Posted on

Audience segments based on inferred interests rather than volunteered facts

A profile can say you are interested in something you never told anybody you liked.

That is not necessarily an error.

It may be the product working as designed.

Advertising systems routinely infer interests, habits, or purchase intent from behavior and then place users into audience segments.

The important word is infer.

Behavior becomes a label

Google’s current advertising documentation says Demand Gen audiences can include groups based on interests, habits, active research, demographic information, or prior interaction with a business. Google describes these audience categories as estimates and says its systems may classify people into groups such as sports fans, travelers, or people currently shopping for cars. See Google Ads’ Demand Gen audiences overview and About audience segments.

The label may therefore come from observed behavior rather than a form where the user checked:

I am currently shopping for a car.

A system might infer that interest from searches, videos, app activity, website visits, purchases, or other signals available to the platform.

The Federal Trade Commission’s 2024 report on large social-media and video-streaming services found that companies maintained user-interest information and used those interests primarily for targeted advertising. The report noted examples such as food, nightlife, parental-status-like categories, and shopping-interest segments. See A Look Behind the Screens.

An inference is useful precisely because it goes beyond volunteered data

If advertisers could target only facts people explicitly entered into profile forms, many useful commercial categories would be missing.

Someone researching tents, hiking boots, trail maps, and national parks may never click a button labeled Outdoor enthusiast.

A model can still make the inference.

That can make advertising more relevant.

It can also create labels the person never sees and never had an opportunity to correct.

The label can be wrong or stale

Behavior is ambiguous.

You may research diabetes for a relative.

You may shop for baby products for a coworker’s shower.

You may read luxury-car reviews because the engineering is interesting while having absolutely no intention of buying one.

You may spend a week researching divorce law for an article.

The resulting segment can mistake curiosity, work, gifts, research, or one-time events for stable personal interest.

And even a correct inference can expire.

A person who was shopping for a refrigerator last month probably does not want to be classified as a refrigerator enthusiast until retirement.

Segments change what the system decides to show

Once assigned, audience labels can affect ad eligibility, bidding, recommendations, campaign optimization, measurement, and other automated decisions.

That does not mean every segment produces an important consequence.

It means the system has turned behavior into a proposition about the person.

The Surveillance Economy does not only collect facts.

It manufactures new data from old data.

A browser history is one dataset.

What we think this person wants is another.

Posted on

Real-time advertising auctions and the spread of user information

An online ad can be sold in less time than it takes you to notice the empty rectangle where it will appear.

Before the winning ad arrives, information about the opportunity may already have traveled through an advertising auction.

That is the basic structure of real-time bidding, or RTB.

A bid request describes more than a rectangle

The IAB Tech Lab’s OpenRTB specification defines a standard way for an exchange or supply platform to ask bidders what they will pay for an advertising impression.

A bid request can contain information describing the site or app, device, user, advertising slot, auction rules, and optional audience or segment data. See the current OpenRTB specification and IAB Tech Lab’s OpenRTB overview.

Not every exchange sends every field.

Not every request contains personal information.

The important architectural point is that the auction request itself is a data-distribution event.

Multiple potential buyers may need enough information to decide whether the impression is valuable to them.

Only one may ultimately win.

Losing the auction does not mean never seeing the request

This distinction became unusually concrete in the Federal Trade Commission’s 2024 action against data broker Mobilewalla.

The FTC alleged that Mobilewalla collected and retained information from real-time bidding exchanges while participating in ad auctions, including information from bid requests even when the company did not win the advertisement. The agency alleged that the company accumulated hundreds of millions of advertising identifiers paired with precise location data and later used data for audience segmentation and other purposes. See the FTC’s Mobilewalla enforcement announcement.

Those are allegations in an enforcement action, not proof that every RTB bidder behaves that way.

But they demonstrate the structural issue clearly.

The information needed to evaluate an auction can be valuable even without the ad.

The actual payload matters

It is easy to describe RTB too dramatically.

A researcher should not assume that every auction contains a name, exact location, browsing history, or sensitive category.

The proper question is narrower:

What fields were actually sent, to which recipients, under which identifiers?

OpenRTB supports device and user context, but optional fields can be omitted, generalized, restricted, or transformed. Privacy rules, exchange policies, consent signals, browser restrictions, and seller configuration can all change the payload.

A packet capture, exchange documentation, contract, regulatory record, or bid-request sample is stronger evidence than merely observing that programmatic advertising exists on the page.

The auction creates a distribution problem

Traditional advertising sounds simple: a publisher shows an ad from an advertiser.

Programmatic advertising can involve publishers, supply-side platforms, exchanges, demand-side platforms, data providers, measurement companies, and other intermediaries.

The advertisement is the visible result.

The data path that produced it can be much wider.

That is what makes RTB important to the Surveillance Economy.

The auction is not just deciding which ad you will see.

It can also determine which companies get a chance to evaluate information about the person or device about to see it.

Posted on

Hashed email addresses as persistent advertising identifiers

A hashed email address looks wonderfully anonymous.

It is a long string of hexadecimal garbage.

That appearance can be misleading.

If two companies start with the same email address, normalize it the same way, and run the same hashing algorithm, they can produce the same hash.

Now the ugly string becomes a matching key.

Hashing hides the readable address, not necessarily the relationship

Google’s current Customer Match documentation instructs advertisers to normalize customer emails and hash them with SHA-256 before upload. Google then compares those hashed values against hashed account information to find matches. See Google Ads Data Manager’s Customer Match formatting guidance and How Google uses Customer Match data.

That is the important property.

Google does not need to reverse the hash into the original address in order to know that two records represent the same normalized email.

It only needs both sides to generate the same result.

Suppose:

leo@example.com

is normalized and hashed into:

4c...9f

An advertiser can upload 4c...9f.

A platform can independently hash its copy of leo@example.com and get 4c...9f too.

The plaintext address never has to travel in the matching file for the records to connect.

One-way does not mean anonymous

SHA-256 is a one-way cryptographic hash function. That is useful because the hash is not intended to be decrypted back into the original input.

But anonymity is a different question.

Email addresses come from a relatively structured and often guessable input space. More importantly, a party that already possesses the email address does not need to guess anything. It can simply hash its own copy and compare results.

LiveRamp’s current identity-resolution documentation explicitly accepts hashed email addresses as inputs for resolving records to persistent person- or household-level identifiers. See LiveRamp’s RampID identity-resolution documentation.

That is pseudonymization with matching utility intact.

The readable identifier is transformed.

The ability to connect records survives.

Persistent matching can outlive a cookie

Cookies can be cleared.

Browsers can partition storage.

Mobile advertising IDs can be reset or deleted.

An email address may remain stable for years.

If that email is repeatedly transformed into the same standardized hash, the hash can provide a durable bridge between customer databases, advertising systems, measurement tools, and identity-resolution services.

That does not mean every hashed email is shared broadly or used forever. Actual use depends on the service, contracts, retention rules, platform policies, and user choices.

But the technical lesson is simple.

Replacing a name with a deterministic code does not erase identity if everyone who matters knows how to produce the same code.

The Surveillance Economy frequently works by changing what an identifier looks like without changing what it can connect.

Posted on

Identity graphs that connect devices to people and households

A laptop does not know it lives with a television.

An advertising system may try to figure that out.

That is the job of an identity graph: connect identifiers that appear in different systems and decide which ones belong to the same person, household, or device cluster.

The graph might contain email addresses, phone numbers, postal addresses, cookies, mobile advertising IDs, IP addresses, connected-TV identifiers, customer IDs, and other persistent keys.

The point is not merely to store them.

The point is to connect them.

The graph turns fragments into relationships

LiveRamp’s current documentation describes identity resolution as connecting fragmented consumer touchpoints to a person- or household-based view. Its systems can resolve names, postal addresses, email addresses, phone numbers, cookies, mobile device IDs, IP addresses, connected-TV IDs, and other identifiers to persistent RampIDs. See LiveRamp’s identity-resolution documentation.

Experian describes the same general problem as matching people or households to devices and platforms using deterministic or probabilistic identity methods. See Experian’s identity-resolution guide.

That is a much larger object than a cookie.

A cookie identifies one browser context.

An identity graph tries to decide which browser belongs with which phone, which television, which email address, which customer record, and sometimes which household.

Some links are stronger than others

Not every edge in the graph has the same evidentiary quality.

A user logging into the same account on a phone and laptop creates a relatively strong connection.

An email address tied to a loyalty account can create another.

Other links may be inferred from patterns such as shared networks, repeated co-location, common household information, or other statistical signals.

The Federal Trade Commission’s 2017 report on cross-device tracking distinguished deterministic methods from probabilistic approaches that infer links between devices. See Cross-Device Tracking: An FTC Staff Report.

That difference matters.

A verified relationship says these two identifiers were directly connected by evidence.

A probabilistic relationship says these two identifiers appear likely to belong together.

Those are not interchangeable claims.

A bad edge contaminates the profile

Households are messy.

People share Wi-Fi. Children use parents’ tablets. Visitors connect phones to home networks. Couples share televisions. Old devices get sold. Work laptops travel home. Apartments change tenants.

If a graph incorrectly joins two people, activity from one can be attributed to the other.

That can affect advertising, measurement, recommendations, fraud models, or any later analysis built on the graph.

The larger lesson is that identity resolution does not merely collect more data.

It changes the unit of observation.

Instead of asking what one browser did, the system can try to ask what this person or this household did across many devices.

That is powerful when the links are right.

It is also why the links themselves deserve scrutiny.

The Surveillance Economy does not need every device to know your name.

It only needs a graph confident enough to connect the devices to something that does.

Posted on

Data brokers that assemble profiles from many unrelated sources

The unsettling thing about a data broker is that you may never have visited its website.

You may never have installed its app.

You may never have created an account.

The broker can still have a file about you.

That is possible because the business model begins where many of the previous Surveillance Economy articles end: somebody else already collected the fragments.

The FTC documented the assembly line years ago

The Federal Trade Commission’s major 2014 report on data brokers found that brokers obtained information from commercial sources, government records, publicly available sources, websites, and other data brokers. The report said the companies studied combined small pieces from many places into much more detailed composite profiles.

See the FTC’s Data Brokers: A Call for Transparency and Accountability and the agency’s summary of its findings.

The FTC found that brokers could collect purchase information, browsing activity, warranty registrations, public records, and other everyday data. It also found that brokers frequently bought information from other brokers, making the original source difficult for a consumer to reconstruct.

That structure remains important even as the specific companies and technologies change.

The value comes from combination

Suppose five separate systems know five separate things:

  • a retailer knows what you bought,
  • a public record contains an address,
  • an advertising system knows what device visited certain sites,
  • a location provider has observations tied to an advertising ID,
  • a people-search service has old phone numbers and relatives.

None of those datasets necessarily contains a complete person.

A broker’s value comes from matching them.

Names, addresses, emails, phone numbers, device IDs, household identifiers, probabilistic matches, and other linking fields can turn fragments into a composite.

Combination also combines errors

Profiles do not become true merely because they are large.

A broker can inherit an outdated address from one source, a mistaken household member from another, a purchase made for somebody else, or a device incorrectly assigned to the same person.

Once records are merged, the error can become part of a more authoritative-looking profile.

The FTC’s consumer guidance on people-search sites notes that these services may compile information from other brokers, public social-media profiles, and federal, state, and local public records. See What To Know About People Search Sites That Sell Your Information.

Direct consent can disappear several hops ago

A person may have knowingly given an address to a retailer or installed an app with location permission.

That does not mean the person understands every later broker-to-broker transfer.

The FTC’s 2024 X-Mode/Outlogic case is a concrete example of location data moving through an ecosystem that included third-party apps, SDKs, aggregators, and hundreds of clients. See the FTC’s final order announcement.

This is where the Surveillance Economy stops looking like one tracker following one browser.

The final profile may be assembled by a company the person has never heard of, from records created by companies that never saw the final profile.

No single source has to know everything.

The broker’s product is knowing how to put the pieces together.