The Best Identity Data Providers

Get data for any location

Start your search

Last updated: September 2026

Identity Data: A Quick Guide to the Players and Technology

Identity data is an increasingly essential dataset given the fragmented digital ecosystem, where users engage across multiple devices, platforms, and environments. Identity data creates the infrastructure that ties together the signals of a single user, allowing for targeting, attribution, and personalization at scale.

That fragmentation continues to grow. The industry spent years preparing for one disruption, the end of the third-party cookie. What arrived instead was messier: partial identifiers, competing ID frameworks, and a growing share of impressions that carry no usable identifier at all. And with all of this, there’s been significant growth in synthetic data. 

What Is Identity Data?

Identity data comprises the various anonymous digital identifiers associated with an individual or household. These identifiers can include mobile advertising IDs (MAIDs), hashed emails (HEMs), cookies, IP addresses, usernames, and connected TV identifiers such as ACR-derived device IDs. Teams match and connect these identifiers into cohesive clusters to build identity graphs to understand which identifiers correspond to a single anonymous user or household profile.

A growing layer sits on top of these raw identifiers: universal ID frameworks like Unified ID 2.0, RampID, ID5's Universal ID, and Panorama ID, each attempting to provide a portable, privacy-conscious identifier that works across platforms. But most teams building identity infrastructure now plan for a multi-ID environment rather than betting on a single winner.

At its core, an identity graph functions as a database of identifiers and associated data points that collectively define a user's digital footprint. It plays a critical role in audience targeting, personalization, attribution, and measurement across the advertising and marketing ecosystem.

Deterministic vs. Probabilistic Linkages

Identity data includes a mix of deterministic and probabilistic linkages. Deterministic data originates from known, verified relationships, such as a user logging into a mobile app with their email address, which can tie a HEM directly to a MAID. This type of linkage is highly accurate but often lacks scale.

In contrast, probabilistic identity data is inferred from behavioral patterns, such as repeated dwell times at a specific IP address, suggesting that a device belongs to a particular household. These inferred connections trade some certainty for broader reach.

Identity graph providers often blend both approaches, using deterministic data to anchor their graphs and train their probabilistic models, then scaling reach through statistically modeled linkages across devices and environments. Machine learning has made this blending considerably more effective. Modern matching systems calculate the probability that records belong to the same person even when no field matches exactly, and can extract identity signals from unstructured sources. The tradeoff is unchanged. A model is only as good as the verified pairs used to train and validate it, which is why high-quality deterministic input data remains the constraint on graph accuracy.

How Unacast Powers Identity Graphs 

Unacast supplies data and data linkages essential to building identity graphs. As a location intelligence company, Unacast processes over 70-80 billion raw location signals daily from various providers to create a single, reliable dataset that fuels multiple products. Processing this data involves removing signals with incomplete data, merging them based on spatial and temporal proximity, and applying more robust analytics to each signal to further contextualize each individual signal.

Real-world behavior from Unacast’s location intelligence solutions helps make identity data actionable. A MAID on its own is just an anonymous identifier that tells you that a device exists, but not much else. Location intelligence adds critical behavioral context by revealing where that device goes, when, and how often. This allows companies to understand real-world habits, such as daily commutes, frequent store visits, or time spent at specific venues.

These patterns help tie devices to meaningful attributes that can then be used to link MAIDs to other identifiers, including HEMs and FLIPs. In essence, location data transforms static identifiers into dynamic user profiles.

Unacast's Data Linkages product enriches MAIDs with two powerful connectors:

  • MAID-to-HEM: Deterministic connections between mobile advertising IDs and hashed emails, enabling people-based targeting and CRM activation.
  • MAID-to-FLIP: Links a device to the IP addresses associated with its frequented locations, providing probabilistic insights into households or workplaces.

This enrichment data is highly valuable for identity graph builders seeking to increase the number of IDs per cluster, improve precision, or validate probabilistic models with known deterministic pairs.

Who Uses Identity Data and Why

Identity data is core infrastructure across several industries:

  • AdTech & MarTech: DSPs, SSPs, CDPs, and DMPs rely on identity graphs for targeting, segment extension, attribution, and cross-device measurement. Identity enables omnichannel messaging, personalization, and ROI tracking.
  • Retail & eCommerce: Brands use identity to link online and offline behavior, activate CRM segments in digital environments, and drive retention through personalization. Retail media networks in particular depend on resolved identity to prove that an ad drove an in-store visit.
  • Media & Streaming: Content platforms resolve identities to personalize viewing experiences, optimize subscriptions, and limit frequency across devices.
  • Data Marketplaces & Platforms: Data aggregators and resellers need identity data to enrich profiles, enhance reach, and standardize datasets. Increasingly, this collaboration happens inside data clean rooms, where two parties match on identity without either. 
  • Public Sector: identity data can be useful for the intelligence community, especially around fraud, anti-human trafficking, operational security, and other related use cases. 

Identity Resolution in Connected TV

CTV is the next frontier for identity resolution, but it operates somewhat differently. There were never cookies in CTV. Device identifiers do exist, but they are platform-specific and do not travel: Roku issues RIDA, Samsung TIFA, Amazon AFAI, Vizio VIDA. Each is stable inside its own ecosystem and meaningless outside it. A household with a Roku in the den and a Samsung in the bedroom presents as two unrelated devices, with nothing in the bid request to connect them.

What fills the gap is the IP address. In CTV, IP functions as the de facto household key. It is the one signal shared by every device on the same network, and the only practical bridge between a smart TV and the phone in the same room. Unlike on the open web, the major CTV platforms have shown little sign of restricting it.

IP alone is also noisy. Addresses churn on residential ISPs. Carrier-grade NAT places many subscribers behind a single address. Apartment buildings, offices, hotels, and shared Wi-Fi manufacture false households, while VPNs relocate a home entirely. An identity graph that treats one observed IP as proof of household membership inherits all of those errors and propagates them into targeting and attribution.

Behavioral corroboration is what narrows the error, and it is where MAID-to-FLIP applies directly. Instead of asserting a household from a single connection event, frequented-location data establishes which IP addresses a device is repeatedly associated with over time. That means the address where it dwells overnight, week after week, rather than the coffee shop it connected to once. This kind of connection is far more trustworthy and reliable, even as devices change. 

Automatic content recognition (ACR) data supplements this from the content side, revealing what a screen displayed across linear, VOD, and streaming. Important to understand here is that ACR describes the screen rather than the viewer. Connecting exposure to a person, and then to a store visit, a purchase, or a frequency cap that holds across TV and mobile, still depends on the household resolution layer underneath.

Identity Data Providers

These companies supply foundational data, such as MAIDs, hashed emails (HEMs), IP addresses, or behavioral signals that power identity resolution and help enrich identity graphs.

1. Unacast

Unacast provides high-quality location and identity data that links MAIDs to real-world behaviors like visitation patterns, enabling accurate enrichment with HEMs and frequently leveraged IPs (FLIPs). Our data fuels identity resolution across ad tech, retail, and analytics use cases by adding spatial and behavioral context to otherwise anonymous identifiers. Our identity data is optionally fused with location intelligence to provide higher confidence and additional data points for targeting and measurement purposes.

2. Factori

Factori specializes in mobile and digital identity enrichment, delivering curated datasets that enhance device identifiers with MAIDs and HEMs. Their offering supports identity resolution workflows across programmatic advertising, analytics, and attribution.

3. Truthset

Truthset focuses on validating the accuracy of identity data, particularly for demographic attributes associated with HEMs and MAIDs. Their scoring system helps marketers and platforms understand the quality and reliability of the identity data they're using, a function that has become more important as graphs lean harder on modeled linkages.

4. AtData

AtData approaches identity from the email side, offering email-centric intelligence, validation, and licensable identity signals. For teams building graphs anchored on hashed emails, AtData's data helps assess whether an email is active, valid, and associated with a real person before it becomes a node in the graph.

5. Samba TV

Samba TV supplies ACR data from smart TVs, providing viewership signals and connected TV identifiers. As CTV becomes a larger share of ad spend, ACR data has become one of the primary ways to extend an identity graph into the living room and connect household TV exposure to other devices.

Identity Graph Providers

These companies build and maintain large-scale identity graphs, the systems that connect various identifiers such as MAIDs, HEMs, cookies, IPs, and household-level data to model real people or devices across channels.

1. LiveRamp

LiveRamp created an identity graph built on RampID that connects online and offline identifiers across platforms for cross-device marketing and attribution. It is one of the most widely adopted identity graphs in the adtech and martech ecosystems, powering a variety of programmatic advertising tools and platforms.

2. Experian

Experian has identity resolution solutions that connect consumer data across channels using a persistent, household-based identity graph. They combine financial, demographic, and behavioral data with deterministic identifiers to provide marketers with accurate and scalable audience targeting.

3. Merkle (Merkury)

Merkury is Merkle's identity resolution platform that enables brands to own and operate their own private identity graph. It blends deterministic and probabilistic data to connect devices, emails, and offline data in a privacy-compliant, persistent identity solution.

4. TransUnion (TruAudience)

TransUnion's TruAudience marketing solutions, built on the identity assets it acquired with Neustar, maintain a persistent identity graph spanning online and offline identifiers. Its credit-bureau-grade referential data gives it unusually deep household and address-level coverage, which it uses to anchor deterministic matching.

5. ID5

ID5 operates a shared, neutral universal ID designed to let publishers and ad tech platforms recognize users in environments where third-party cookies are unavailable or unreliable. Rather than building a marketer-owned graph, ID5 sits at the infrastructure layer, providing a common identifier that many platforms can resolve against.

Identity Data: What to Expect Through 2027

The defining assumption of the last several years no longer holds. Google announced in April 2025 that it would keep third-party cookies in Chrome rather than deprecate them, and in October 2025 retired most of the Privacy Sandbox APIs, including Topics, Protected Audience, and Attribution Reporting, citing low adoption. Chrome still allows third-party cookies by default. The cookieless future that shaped six years of industry roadmaps never arrived.

Identity resolution remains just as important, but for different reasons than the industry expected.

Signal loss is now structural. The cause was never one browser's deprecation timeline. It is Apple's App Tracking Transparency, consent-gated environments, walled gardens, privacy-conscious browsers, and inventory that simply arrives without an identifier. Proximic by Comscore put the scale of it at 54% of mobile impressions and 36% of desktop impressions carrying no identifier, in its 2025 State of Programmatic report. That gap has no announced end date, which makes it a permanent engineering problem rather than a migration project.

Multiple ID frameworks will continue to coexist. UID2, RampID, ID5, and Panorama ID all persist with limited interoperability. Teams that planned to standardize on one identifier are instead building translation layers between several. Expect the operational burden of maintaining those mappings to grow through 2027.

Clean rooms are becoming standard infrastructure. Native capabilities in BigQuery, Snowflake, and Databricks have lowered the barrier considerably, and the holding companies have bought in directly. WPP acquired InfoSum and Publicis acquired Lotame, both in 2025. Matching quality inside a clean room is bounded by the quality of the identifiers each side brings to it.

Resolution is moving to real time. Batch-oriented graph refreshes are giving way to streaming architectures and on-demand APIs that resolve identity during a customer interaction rather than in a nightly job.

Regulatory requirements keep expanding. Twenty state consumer privacy laws are now in effect, with additional frameworks scheduled through 2027. Provenance, consent lineage, and the ability to document where a linkage came from are becoming procurement requirements rather than legal afterthoughts.

For technical teams building or improving identity graphs, all five point to the same conclusion. The quality and diversity of input data is everything. Models can extend reach, but they cannot manufacture ground truth. Unacast's role in enriching device data with real-world behavioral context and privacy-safe identifiers makes it a critical partner for the next generation of identity solutions.

Want to explore how MAIDs, HEMs, or FLIP enrichment can impact your project? Book a meeting with us today.

‍

More Blogs

Sort
No items found.

Book a Meeting

Meet with us and put Unacast’s data to the test.
bird's eye view of the city