Ecommerce product data: Fields, feeds, and AI readiness
Learn what ecommerce product data includes, how to structure and validate it, and how to connect catalogs to feeds, structured data, and AI shopping.

Ecommerce product data answers a simple question: what exactly is this item, and can every system trust the answer? The record behind a product page also drives filters, search, inventory, feeds, marketplaces, and AI shopping. When it is incomplete or inconsistent, each surface tells a different story.
A durable product-data operation gives every item a canonical record, clear ownership, controlled values, and a reliable path to each destination. Here is the working model, from fields and identifiers to validation, feeds, structured data, and AI-shopping readiness.
What ecommerce product data means
Ecommerce product data is the set of facts, content, assets, relationships, and commercial signals used to describe, sell, fulfill, and distribute a product online. It includes the visible copy on a product detail page and the less visible fields that make a product searchable, comparable, and eligible for a channel.
A useful product record answers seven questions:
- What is it? Name, brand, category, description, and product type.
- Which item is it? SKU, GTIN, MPN, product ID, variant ID, and canonical URL.
- What is it like? Material, color, dimensions, ingredients, compatibility, fit, and other category attributes.
- Which version is it? Size, color, pack, configuration, bundle, or other variant relationship.
- Can I buy it now? Price, currency, availability, inventory, condition, and region.
- What will I receive? Images, video, included components, shipping, returns, warranty, and compliance details.
- Why should a system choose it? Use cases, benefits grounded in facts, reviews, relationships, and provenance.
Product data is the raw material for a product catalog. Product content is the customer-facing expression of that data. A product feed is a destination-specific output. A PIM is one possible system for governing the records. Keeping these terms separate prevents a feed or storefront from becoming an accidental source of truth.
Why the business role extends beyond the product page
Product data connects teams and surfaces that used to operate separately:
- Discovery: Search, facets, recommendations, category pages, shopping results, and AI prompts use attributes to match products to intent.
- Purchase confidence: Clear dimensions, fit, ingredients, compatibility, care, and policies answer the questions that create hesitation and returns.
- Operations: SKUs, variants, inventory, supplier details, and fulfillment fields support planning, picking, reporting, and support.
- Distribution: Marketplaces, retailers, ad platforms, affiliates, and social catalogs each consume a mapped version of the same product facts.
- Machine interpretation: Structured fields let software compare products without guessing from a long paragraph or an image.
Google's ecommerce documentation describes product data as attributes such as title, description, color, pricing, and availability. It recommends combining product structured data on product pages with direct data sharing through Merchant Center when appropriate. That pairing reflects the practical reality: a human-facing page and a machine-facing record have to agree.
A practical ecommerce product data model
Start with a shared core record, then extend it by category and destination. Apparel needs fit and fabric. Electronics need ports and compatibility. Beauty products need ingredients and warnings. Industrial parts need dimensions, tolerances, and replacement relationships.
| Layer | Example fields | Primary job |
|---|---|---|
| Identity | Product ID, SKU, GTIN, MPN, brand, canonical URL | Match, deduplicate, and update the right item |
| Description | Title, short description, long description, bullets, product type | Explain the product to people and systems |
| Taxonomy | Category path, product type, collection, channel category | Place the item in the right browse and feed context |
| Attributes | Size, color, material, dimensions, ingredients, fit, compatibility | Support filtering, comparison, recommendations, and use-case matching |
| Variants and relationships | Parent ID, variant ID, item group ID, bundle contents, accessories | Keep sellable versions and related products distinct |
| Commercial state | Price, sale price, currency, availability, inventory, condition | Represent the offer a shopper can actually buy |
| Logistics and policy | Weight, package dimensions, shipping, returns, warranty, restrictions | Set fulfillment expectations and channel eligibility |
| Media | Primary image, alternate images, video, manuals, alt text, image role | Show and support the product and its claims |
| Provenance | Source system, source URL, owner, confidence, last updated | Make values traceable and refreshable |
The model should distinguish facts from derived values. A manufacturer-confirmed material is different from an inferred style tag. A review summary is different from a specification. Keep that distinction visible so enrichment does not turn an assumption into a product promise.
An illustrative product record
This example shows one sellable variant of an apparel product. The values are illustrative. Replace them with approved data from your own sources.
{
"product_id": "northline-commuter-jacket",
"parent_id": "northline-commuter-jacket",
"variant_id": "northline-commuter-jacket-navy-m",
"sku": "NL-JKT-NVY-M",
"gtin": "0614141123452",
"brand": "Northline",
"title": "Northline Waterproof Commuter Jacket, Navy, Medium",
"category_path": ["Apparel", "Outerwear", "Rain Jackets"],
"attributes": {
"color": "navy",
"size": "M",
"fit": "relaxed",
"material": "recycled nylon",
"features": ["taped seams", "packable hood", "reflective trim"],
"use_cases": ["bike commuting", "travel", "light rain"]
},
"offer": {
"price": 148.00,
"currency": "USD",
"availability": "in_stock",
"condition": "new"
},
"media": {
"primary_image": "https://example.com/images/nl-jkt-nvy-m.jpg",
"image_role": "variant_front"
},
"relationships": {
"variant_options": {"color": "navy", "size": "M"},
"accessories": ["northline-packable-rain-pants"]
},
"provenance": {
"spec_source": "supplier-sheet-2026-05",
"inventory_source": "erp",
"last_verified": "2026-05-15"
}
}The parent identifies the product family. The variant identifies the sellable item with its own SKU, price, inventory, URL, and image when those differ. That structure lets a site show one product with options while a feed or order system still updates the exact item.
How product data moves from source to channel
Treat the architecture as a pipeline with a canonical record in the middle:
Supplier files, manufacturer pages, ERP, PIM, DAM, storefront, reviews → ingest → canonical product model → normalize → enrich → validate and approve → publish to PDPs, feeds, marketplaces, APIs, search indexes, structured data, and AI shopping surfaces → monitor and resync
Each source can remain authoritative for a different field. An ERP may own inventory and cost. A DAM may own approved media. A PIM may own localized descriptions and category attributes. A storefront may own the live URL and offer presentation. Supplier documentation may be the source for dimensions and materials. Record the owner and timestamp instead of blending values silently.
A PIM coordinates product-content workflow. An ERP manages operational and transactional data. A DAM manages files and rights. A feed system maps records to channel schemas. An API exposes records to software. A product data layer can connect these systems and provide a normalized, live output for AI commerce. A PIM and Catalog comparison explains why these layers can work beside each other rather than forcing one system to do every job.
The architecture is healthy when a team can change an approved fact once, see which outputs are affected, and trace a channel value back to its source.
The product-data workflow
1. Collect the source material
Inventory every place product facts live: supplier spreadsheets, PDFs, manufacturer pages, product detail pages, PIM and ERP exports, DAM assets, marketplace listings, reviews, and internal databases. Record format, owner, update cadence, coverage, and known gaps.
Collect the source, not just the extracted value. A dimension without its specification sheet or retrieval time is hard to audit. For external pages and documents, retain the source URL, retrieval timestamp, and extraction method.
2. Define the schema before enrichment
Create a core field list and category extensions. Mark each field as required, recommended, conditional, or prohibited for every important destination. Define accepted value sets, units, locale rules, and validation logic.
Keep a field dictionary. color, primary_color, and finish should represent distinct concepts, or be mapped intentionally. A shared dictionary makes source mappings reviewable and prevents every channel team from inventing its own names.
3. Normalize values and structure
Normalization makes equivalent facts comparable:
- Convert dimensions, weight, volume, and temperature to agreed units.
- Map categories to a controlled taxonomy and preserve the source category when useful.
- Normalize colors, sizes, materials, ingredients, conditions, and compatibility values.
- Standardize casing, punctuation, date formats, currency, and decimal rules.
- Split parent products from sellable variants and assign stable relationships.
- Deduplicate records and flag conflicting values rather than choosing silently.
Free text still has a place in descriptions. High-impact fields used for filters, feeds, and comparisons need typed values or controlled vocabularies.
4. Enrich from evidence
Enrichment fills gaps in a record: category-specific attributes, clearer titles, buyer-facing descriptions, use cases, compatibility, product relationships, image metadata, and channel fields. The input can come from approved supplier specs, manufacturer pages, product images, manuals, reviews, or internal systems.
AI can accelerate extraction, classification, and drafting. It should not invent a fabric blend, safety claim, voltage, ingredient, or compatibility fact. Require a source, confidence level, and review path for values that affect a purchase or compliance. Our guide to product data enrichment for AI commerce covers the difference between a nicer description and a usable structured record.
5. Validate before publishing
Run gates before a record reaches any destination:
- required-field completeness by category and channel;
- identifier format, uniqueness, and parent-child relationships;
- accepted values, units, currency, and locale;
- price, sale price, availability, and inventory logic;
- image URLs, variant-image matching, and media requirements;
- category and attribute compatibility;
- claims, restrictions, warnings, and policy fields;
- structured data and feed syntax.
Then validate the rendered destination. A valid source can still create a broken selector, wrong image, stale price, or hidden attribute on the product page.
6. Publish, monitor, and resync
A product record produces channel-specific outputs. Publish the right fields and format to each destination, then inspect acceptance, warnings, display, and freshness. Feed diagnostics, structured-data reports, zero-result searches, support questions, and product returns all reveal gaps the source audit missed.
Use event-driven or frequent updates for price, availability, and inventory. Use an approval workflow for slower-changing copy, attributes, and media. The right cadence depends on how quickly the field can change and how costly a stale value is.
For a deeper distribution model, see product data syndication.
Identifiers and variants are the join layer
Identifiers connect product data across systems. They deserve an explicit policy.
| Identifier | Scope | Use |
|---|---|---|
| SKU | Internal to the seller | Inventory, fulfillment, reporting, and internal updates |
| GTIN, UPC, or EAN | Shared trade-item identity when assigned | Marketplace matching, product feeds, and cross-retailer recognition |
| MPN | Manufacturer's model or part number | Product matching, parts catalogs, and technical commerce |
| Product ID | Your canonical record key | Database joins, APIs, and durable references |
| Variant ID or item group ID | Parent-child relationship | Group size, color, pack, and configuration options |
A SKU is not a GTIN. Do not reuse a SKU for a different sellable item, assign one GTIN to unrelated variants, or rely on a title to join records. Link to the GTIN definition when you document the policy for merchandisers and suppliers.
For variants, define which differences create a new sellable item. A navy medium jacket and a navy large jacket need separate inventory and usually separate SKUs. They can share a parent product and canonical family page. A two-pack and a single unit are different offers and need explicit pack or bundle data. A replacement filter compatible with a vacuum is a related product, not a color variant.
When a channel requires one row per item, send the sellable variant with its own ID, title or selected options, URL, image, price, availability, and identifiers. When a channel groups variants, send the stable parent or item-group relationship as well. Never make downstream systems infer whether two rows are options, bundles, or different products.
Feeds and structured data are outputs, not the catalog
A product feed is a file, stream, or API transfer built for a destination. It may be CSV, TSV, XML, JSON, or another channel-supported format. A feed can contain title, description, link, image, price, availability, brand, category, identifiers, variants, shipping, and custom fields. Google Merchant Center's product data specification is one example of a destination-specific contract.
The same canonical record can produce different outputs for Google Merchant Center, marketplaces, Meta catalogs, affiliate networks, internal search, and AI shopping surfaces. Each destination has different required fields, taxonomies, limits, accepted values, and update rules. Fix upstream when possible. A channel-specific rule should translate the record, not become a hidden second catalog.
Structured data is selected product information embedded in the product page with a machine-readable vocabulary such as Schema.org. Google's ecommerce structured-data guidance supports Product, Offer, ProductGroup, and related types for ecommerce. Google recommends product structured data on product pages, and its product structured-data documentation notes that merchant pages should keep markup in the initial HTML when possible for changing values such as price and availability. Markup must match what shoppers see. Structured data increases eligibility and understanding, it does not guarantee a rich result or a ranking.
Pair page markup with direct channel data where the destination supports it. Google's guidance explains that structured data can improve its understanding of price, discounts, and shipping, while Merchant Center feeds give larger or frequently changing catalogs more control over coverage and update timing. Test markup and monitor warnings after template changes. Our structured data glossary and product feed guide cover the implementation details without treating one output as the source of truth.
Make the record ready for AI shopping
AI shopping surfaces compare products against natural-language constraints. A shopper may ask for a navy rain jacket for bike commuting, a fragrance-free moisturizer for sensitive skin, or a USB-C monitor that works with a specific laptop. The record needs explicit facts that map to those constraints.
Prioritize five layers:
- Identity: Stable IDs, brand, category, product URL, and variant relationships.
- Decision attributes: Materials, dimensions, fit, ingredients, compatibility, use cases, included items, and constraints.
- Commercial state: Current price, currency, availability, seller, shipping, returns, and regional eligibility.
- Machine-readable delivery: Structured product pages, feeds, APIs, and typed product objects with consistent field names.
- Evidence and control: Source provenance, reviews or Q&A where available, policy data, update timestamps, and a way to block or correct an output.
OpenAI's current product-feed documentation illustrates the baseline for ChatGPT discovery: a stable item ID, title, factual description, URL, brand, seller, image, price, and availability, with richer fields for variants, attributes, shipping, returns, and reviews. The exact contract will vary by integration. The durable rule is to expose complete, current records instead of relying on a model to infer missing facts.
AI readiness is a data-quality discipline. Product data quality means records are accurate, complete, consistent, current, valid, unique, and useful across the places they appear. We provide a parallel product data layer, the foundation for a parallel storefront in AI commerce. We structure and enrich product data, publish live normalized product objects, keep changes synchronized, and measure how AI surfaces use the catalog. It complements a storefront or PIM, so teams can improve machine readability without replacing their existing commerce operation.
A practical implementation plan
Start with a narrow slice of the catalog and make the loop work end to end.
Phase 1: Choose a high-value scope
Select one category, one region, and the products or variants that matter most to revenue, returns, support volume, or channel eligibility. Include a few products with known feed or search problems. Define the destinations you will support first: storefront, Google Merchant Center, a marketplace, an API, or an AI shopping surface.
Phase 2: Assign ownership and design the model
Write the core schema, category extensions, field dictionary, identifier policy, source-of-truth map, and approval roles. Set a quality threshold for every destination. A product can be complete for the storefront and incomplete for a marketplace, so score completeness by channel.
Phase 3: Build the canonical pipeline
Connect sources, retain provenance, normalize values, resolve duplicates, group variants, and enrich missing fields from evidence. Keep transformations versioned. If you change a mapping, you should be able to explain which output records changed and why.
Phase 4: Publish and test the outputs
Render a sample of product pages. Inspect one record in each feed. Test structured data. Verify that the right variant image, price, availability, and URL appear together. Check channel diagnostics and trace failures upstream.
Phase 5: Operate the refresh loop
Monitor freshness lag, completeness, validation failures, duplicate rate, feed rejections, structured-data warnings, search gaps, and AI product visibility. Give every failure an owner and a response time. Expand to another category after the first loop is reliable.
A traditional PIM can remain the system for internal approvals, localization, and product-content governance. A feed platform can handle destination mappings. We built Catalog for teams that need a live, normalized product data layer for AI commerce beside those systems. The PIM versus Catalog comparison helps decide whether you need one layer or both.
Ecommerce product data audit checklist
Use this checklist for a representative sample, then automate the checks that repeat:
- Scope: Have you selected the categories, regions, channels, and high-impact products to audit?
- Completeness: Does each category have required fields for identity, attributes, media, commercial state, logistics, and policy?
- Accuracy: Can each important claim be traced to an approved source?
- Consistency: Are names, units, categories, colors, sizes, and materials normalized across PDPs, feeds, and marketplaces?
- Identifiers: Are SKUs unique, GTINs valid where assigned, MPNs mapped correctly, and canonical URLs stable?
- Variants: Are parent, variant, bundle, pack, and accessory relationships explicit? Does every sellable item have its own inventory and offer state?
- Freshness: How long does a price, inventory, or availability change take to reach each destination?
- Media: Does each variant use the right image and working URL, with required formats and useful alt text?
- Structured data: Does page markup match visible price, currency, availability, product identity, and variants? Is it present in the initial HTML where required?
- Feeds: Are destination fields, taxonomies, accepted values, and update schedules mapped and monitored?
- AI readiness: Can an AI shopping system answer who the product is for, what it is made of, which variant fits, what it costs, and whether it is available without guessing?
- Ownership: Does every failure route to a person or team with a documented fix path?
FAQs
What is ecommerce product data?
Ecommerce product data is the structured and unstructured information that describes, sells, fulfills, and distributes products online. It includes product facts, content, media, identifiers, variants, prices, availability, policies, and relationships.
What fields matter most?
Start with stable identity, title, category, product URL, price, availability, primary image, and variant relationships. Add category-specific decision fields such as material, dimensions, fit, ingredients, compatibility, care, and included items. The right field set depends on the product and its destinations.
Is ecommerce product data the same as a PIM?
No. Product data is the information itself. A PIM is a system for managing, enriching, approving, and distributing some of that information. A PIM can be the source of truth, while a product data layer or feed system handles downstream use cases.
What is the difference between a product feed and product data?
Product data is the broader record a business maintains. A product feed is a channel-specific file, stream, or API output created from that record. One canonical catalog can produce many feeds.
How do I prepare product data for AI shopping?
Make product facts explicit, typed, current, and traceable. Normalize identifiers and variants, expose price and availability, add category-specific attributes and policies, publish structured outputs, and monitor which products and variants are surfaced. Copy alone cannot supply missing product facts.
Build product data that travels
A product record should survive the trip from a supplier file to a product page, feed, marketplace, API, and AI shopping surface without changing meaning. That requires a canonical model, clear ownership, normalization, evidence-based enrichment, validation, and refresh rules.
We help brands build that product data layer for AI commerce. See how Catalog can structure and distribute your catalog so AI shopping systems can read live product facts while your existing storefront stays in place.
