← Back to blog

Measuring Product Data Quality: The Seventh Dimension

A wheel diagram of the seven dimensions of product data quality, with completeness, validity, consistency, uniqueness, accuracy, and timeliness surrounding a backpack product record, and relevance highlighted as the seventh.

Product data ownership sits across many functions within a business, but the directive is always the same: “Better product data.” It is an easy thing to say but a difficult thing to meaningfully act on, because you cannot improve what you do not measure. Most teams either do not measure product data quality at all, or they fail to get the full picture.

Product data quality is the degree to which a product record is complete, valid, consistent, and accurate for its category, such that every channel accepts it, and every system or shopper that reads it, human or AI, can trust it. In today’s world, the above is not enough. To have high quality content, a product must contain not just accurate, but also relevant data.

Most teams never get that far. They know their data needs work. They can feel the rework and the rejected listings, but the process can feel like a black box. If you are not measuring for the right things, every data project runs on instinct, and instinct does not scale to hundreds of thousands of SKUs across hundreds of categories.

The number one thing teams often do measure is fill rate, the percentage of fields that are not empty. This tells only a sliver of the story, and conveys nothing of quality. It is easy to compute, and it feels like progress. Unfortunately, it is also the number most likely to tell you your data is fine right up until a marketplace rejects half of it. And even if the listing goes up, it likely won’t be found. Measuring product data quality well takes a few more dimensions. We can borrow most of them from generic data-quality frameworks, but there is also one product-specific dimension that is key to successfully activating your catalog across all channels.

The first six dimensions

Generic data quality has a standard vocabulary, and most of it carries over to product data. Each one has a metric you can actually compute.

  1. Completeness. The share of required attributes that are populated.
  2. Validity. The share of values that match the format and type their field demands. A UPC that fails its check digit is invalid. A weight of “medium” is usually invalid. Identifiers (UPC, EAN, GTIN, etc.) are the highest-value place to measure validity because an invalid identifier means a rejected listing.
  3. Consistency. The share of records whose fields do not contradict each other. A title that says stainless steel next to a material attribute that says aluminum is a consistency failure, even though both fields are populated and each is well formed on its own.
  4. Uniqueness. The rate of unintended duplicates. Two records for the same product, or one identifier reused across two products, corrupts everything downstream, from inventory to recommendations.
  5. Accuracy. Whether a value is actually true for the product. This is difficult to measure, because it usually requires a source of truth to check against, and it is where category rules become essential.
  6. Timeliness. Whether the record reflects the current state of the product, its price, its availability, its specification, at the moment a channel reads it.

Measure only these six and you are already ahead of most teams. There is more to consider, though, and an attribute’s relevance to a search or a purchasing decision (not just its completeness or accuracy) further complicates things.

The dimension the frameworks miss: relevance

A record can be complete, valid, consistent, unique, accurate, and timely, but still fail the only test that matters at the moment of the sale. None of those things matter if they don’t tell the shopper what they came to find out.

This brings us to the seventh dimension, and it is the decisive one: relevance. It asks a different kind of question than the other six. The others ask whether the data is correct. Relevance asks whether anyone cares. Does this attribute, this fact, this phrasing answer something a shopper (human or AI) is actually trying to decide?

Consider a pair of leggings. A perfectly accurate record can list fiber content, care instructions, and country of origin. It can have a perfectly valid, consistent, and accurate PDP, yet never mention that they are high-waisted, have pockets, or hold their shape in a squat. The shopper is not searching for “82 percent polyamide.” They are searching for the benefits they care about. Here’s an example of a solid product listing:

Example product listing for high-waisted capri leggings showing fit and pocket attributes alongside price, color, size options, and product details

It does mention they are high-waisted and have pockets. But the content could even go a step further. How big are the pockets? Could you fit your phone in them? What’s the ratio of polyester to spandex, so a shopper understands the stretch? Can you tumble dry them?

Even a product that meets all standard definitions of complete can lack the relevant data to be found or properly indexed by AI agents. That’s because relevance is absent from the standard data-quality frameworks. Those frameworks were built to judge data against itself (its format, its internal logic, its source of truth) and all of those things are defined inside your own system. Relevance is defined outside of those systems, by the shopper, and it is the one dimension you cannot certify by looking only at your own catalog. A value is relevant only in its usefulness to a consumer.

What shoppers want and how they buy is an ever-changing target. The attributes that decided a purchase last season are not the ones that decide it this one. New use cases appear, new comparisons become the tiebreaker, search language changes over time, and AI assistants start asking for facts no template anticipated. A record that was perfectly relevant at launch can quickly go stale, not because a single value changed, but because the consumer (or agent) behavior changed around it. Relevance is the one dimension that decays while your data sits perfectly still. That is what makes it hard, and it is why relevance cannot be a project you finish. It has to be maintained and optimized on an ongoing basis.

That is why the measurement has to live in a pipeline, not an audit. The atronous data pipeline optimizes continuously. Every record is checked for identifier integrity, with a 100 percent format and check-digit pass rate and zero duplicates on delivered identifiers. Every attribute is validated against the constraint vocabulary for its category, across more than 400 categories, each enforced independently. Completeness is scored against what the category requires, not what a template asks for. And content can be scored and optimized against what shoppers and AI agents are actually looking for. The pipeline keeps content fresh as that demand shifts. A typical enterprise run surfaces 18 distinct categories of data quality issues and returns each one with its rationale, at a 98 percent or better success rate. Nothing is dropped in silence.

Why relevance is worth the work

Measuring product data quality properly is more work than reading a fill rate off a dashboard. Keeping it relevant is more work still, because it never quite finishes. It is worth it for three reasons.

  1. It tells you where you actually stand, per record and per category, so a data project can be aimed instead of guessed.
  2. It predicts the outcomes you care about, because the numbers that move it, category-required completeness and relevance, are the same numbers that move whether a product is found, matched to intent, and bought.
  3. It is becoming the difference between being found and being skipped. The reader of a product record is increasingly an AI agent that does not browse a page. It reads structured data, compares on it, and decides in a sentence whether your product answers the question. An attribute the shopper cares about, missing, is not a blank field. It is a reason to recommend someone else. And because what the agent asks for keeps changing, relevance is never fixed once. It is maintained.

Accurate and complete are table stakes, and they are not enough on their own. A record can be both accurate and complete and still answer a question nobody is asking. What separates quality content from merely correct content is whether it carries what the shopper wants to know, right now, in the language they are asking it. Since that target never stops moving, the only way to stay ahead of it is with a pipeline that keeps every record fresh, relevant, and optimized for human shoppers and AI agents alike, every time the question changes.

See your own data measured

The fastest way to know where your product data stands is to see a sample of it measured and returned. That is what the atronous Data Quality Assessment does. Send up to 50 SKUs out of your PIM or ERP, exactly as they live in your system, and we run them through the same pipeline our enterprise customers rely on. Within five business days you receive the sample back, generated and validated, along with the taxonomy and schema recommendations behind the work and a working session with a specialist to walk through what we found, what we added, and what it means for your catalog at scale.

This uses your own records to show exactly where the data needs work, and what better actually looks like.

Request your Data Quality Assessment.

Intelligence in every attribute.

Frequently asked questions

What is product data quality?

Product data quality is the degree to which a product record is complete, valid, consistent, and accurate for its category, such that every channel accepts it and every system or shopper that reads it, human or AI, can trust it. Correct alone is not enough: a high quality record also carries the data that is relevant to what shoppers and AI agents are actually trying to decide.

How do you measure product data quality?

Across seven dimensions. Six carry over from generic data-quality frameworks, each with a computable metric: completeness, validity, consistency, uniqueness, accuracy, and timeliness. The seventh, relevance, asks whether the record answers what a shopper or AI agent is actually trying to decide. Because that target keeps moving, relevance has to be measured continuously in a pipeline, not once in an audit.

Turn broken product data into verified listings.

Start with a conversation about your product data.