Posted on

Cross-device matching and errors in household attribution

Your television does not necessarily belong to the person holding the phone beside it.

Advertising systems sometimes have to make that guess anyway.

Cross-device matching tries to connect activity from several internet-connected devices to the same person or household.

When the match is right, separate histories become one story.

When it is wrong, somebody else’s behavior can become part of yours.

Devices can be linked directly or inferred

The Federal Trade Commission’s staff report on cross-device tracking describes two broad approaches.

Deterministic matching uses stronger direct evidence, such as the same person logging into one account across multiple devices.

Probabilistic matching estimates relationships using signals that may include IP addresses, device information, location, browsing patterns, or other observations. See Cross-Device Tracking: An FTC Staff Report.

Both approaches can be useful.

They do not have the same error profile.

A shared login is evidence that the same account appeared on two devices.

A shared household network is evidence that two devices used the same connection.

Those statements are different.

Households are hostile environments for neat attribution

Imagine four people in one home.

They share:

  • a television,
  • a game console,
  • home Wi-Fi,
  • a family tablet,
  • streaming accounts,
  • and occasionally each other’s laptops.

One household member researches motorcycles.

Another shops for maternity clothes.

A teenager watches hours of gaming videos on the living-room TV.

A parent uses the same tablet to research retirement accounts.

A matching system that collapses those devices too aggressively can produce a fictional super-person who is simultaneously pregnant, retiring, buying a motorcycle, and speedrunning Elden Ring.

The database may be internally consistent.

The human it describes may not exist.

Shared IP addresses are not identities

A household router makes many devices appear to the outside world through one public IP address.

Workplaces, schools, libraries, hotels, apartment networks, mobile carriers, and VPNs can create even larger shared environments.

That makes network proximity useful as a signal but dangerous as a conclusion.

LiveRamp’s current identity-resolution documentation explicitly recognizes that shared device touchpoints can map to more than one persistent identity. See LiveRamp’s RampID identity-resolution documentation.

That is an important admission built into the machinery itself: one technical identifier does not always equal one human.

Attribution errors propagate

Once two devices are joined, later systems may inherit the connection.

Advertising measurement can credit the wrong device. Audience segments can absorb another person’s interests. Recommendations can become strange. A household-level campaign can be mistaken for person-level evidence.

This is not an argument that cross-device matching never works.

It is an argument that the resolution level matters.

Person.

Household.

Device.

Account.

Those are different objects.

The Surveillance Economy becomes unreliable when its databases quietly slide between them as though they were the same thing.

Posted on

AI search answers and which sources receive attribution

A citation beside an AI answer looks reassuringly familiar.

It resembles the old academic bargain: here is the claim, and here is the source that supports it.

AI search complicates that bargain because the answer may combine retrieved material, model-generated connective language, multiple pages, and information learned during training. The links shown beside the result are therefore not automatically sentence-by-sentence footnotes.

Google describes AI Overviews as snapshots assembled to help users understand information from a range of sources, with links for deeper exploration. In 2026 it also expanded Preferred Sources into AI Overviews and AI Mode, allowing users to make selected publishers stand out inside AI responses. See Google’s explanation of Preferred Sources in AI Search.

That alone tells us something important: source visibility inside an AI answer is a selection layer of its own.

The cited set is not just the ordinary top ten

A 2026 preprint examined 55,393 Google queries over forty days and compared AI Overview citations with ordinary first-page results. The researchers reported that nearly 30% of cited domains did not appear in the accompanying first-page organic results at all. See Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact.

That study is a preprint and should be read as research in progress, not as a permanent measurement of Google Search. But its result is useful: an AI answer can draw from a source-selection process that is not identical to the visible organic ranking beneath it.

The same paper decomposed responses into thousands of individual claims and reported that some claims were not supported by the pages cited alongside them. The most common problem was omission rather than a directly contradictory source.

In other words, a citation can be real while the connection between that citation and a particular sentence remains weak.

Trace the claim, not just the bibliography

The practical test is simple.

Pick a specific factual claim in the AI answer. Open the cited page. Find the relevant passage. Check whether the page actually supports the wording, scope, and certainty used in the answer.

If three sources are shown beside a paragraph, do not assume all three support every sentence in that paragraph.

Attribution still matters. It gives the reader somewhere to inspect, disagree, verify, and continue researching. That is far better than an answer with no visible origin at all.

But the presence of links should not be confused with transparent authorship.

Traditional search mainly ranked destinations. AI search increasingly constructs an answer first and then presents a selected trail back toward the web.

The trail is useful.

It is not the same thing as seeing the whole path the answer took.

Posted on

Automated accounts that recycle one another’s material

Ten accounts repeating the same thing can look like ten sources.

That is one of the simplest ways automation can distort the apparent size of an online conversation. A bot does not need to invent anything. It can copy a caption, repost a link, slightly alter a sentence, or relay material from another automated account. After enough hops, the visible network looks busy even though very little independent creation occurred.

Researchers have documented this pattern in commercial social-media environments. A 2018 study of SoundCloud activity examined more than 12 million comments and found highly active suspicious accounts that posted repetitive comments, frequently reposted existing content, and contributed relatively little original material. The authors used comment uniqueness and network behavior as clues when distinguishing likely bots and semi-automated accounts from ordinary users. See Social bots in a commercial context — A case study on SoundCloud.

Circulation is not creation

Reposting has legitimate uses. Human communities share jokes, announcements, songs, emergency information, and news links constantly. Automated accounts can also perform useful redistribution: mirroring updates, relaying weather alerts, or syndicating posts to another platform.

The measurement problem begins when repeated circulation is interpreted as independent authorship or independent agreement.

Imagine one account posts a sentence. Twenty automated accounts copy it. Another hundred accounts encounter those copies instead of the original. A researcher who counts only visible posts might record 121 pieces of activity. A reader may perceive widespread agreement. Yet the intellectual source may still be one person, one script, or one upstream feed.

Attribution gets weaker as the material moves. Usernames change. Links disappear. Screenshots replace original posts. Small rewrites break exact-text matching. Eventually a recycled statement may look native to the account currently carrying it.

Repetition can manufacture apparent consensus

This is why source tracing matters more than raw counts.

A useful investigation asks whether accounts are posting independently, whether they share identical URLs or wording, whether their timing is synchronized, and whether the chain leads back to a common source. Recent research on coordinated reposting on Bluesky has used precisely this kind of timing and shared-content analysis to distinguish ordinary repost behavior from suspicious coordination.

None of this means repeated material is automatically bot-generated. Humans copy one another too. Fan communities, customer-service teams, volunteer campaigns, and newsrooms all reuse language.

The narrower point is easier to establish: many visible copies do not imply many independent origins.

Dead Internet Theory often treats repetition as evidence that nobody real is speaking. That conclusion goes too far. Repetition can be human, automated, coordinated, accidental, or mixed.

But when automated accounts recycle one another’s material, the internet can appear to contain more voices than it contains sources. That distinction matters.