Catalog raises $3M to build the product data layer for AI commerce. Read the announcement.
All posts
Product Discovery

Visual search technology for ecommerce: how it works

Learn how visual search technology turns images into product results, what ecommerce data it needs, how to implement it, and how to measure accuracy.

A shopper may have a photo of a chair, a screenshot of a jacket, or a product in front of them, but no useful name for it. A text box forces that shopper to translate what they see into words. Visual search technology makes the image the starting point.

For ecommerce teams, the hard part is not adding a camera icon. The hard part is returning the right product, variant, price, and availability after the image is uploaded. This guide explains how visual search works, where it helps product discovery, which image and catalog conditions it needs, and how to evaluate it without confusing visual similarity with a good match.

What is visual search technology?

Visual search technology lets a person search with a photo, screenshot, camera frame, or other image instead of typing a query. A computer vision system interprets the image, creates a representation of what it contains, and retrieves matching or related results from an index.

In an ecommerce setting, the result may be a product in your own catalog, a visually similar alternative, or a set of products that can be refined with text. The term also appears in psychology research to describe how people scan a scene. Here, visual search means image-led product discovery.

Visual search usually supports three different jobs:

JobWhat the shopper wantsTypical result
Exact product identification“What is this item?”The same product or SKU, when it is in the index
Similar-product discovery“Find products with this look.”A ranked set of visually similar products
Multimodal refinement“Find this look in brown and smaller.”Image matches narrowed by words, filters, or catalog attributes

Those jobs need different relevance rules. An exact lookup should protect product identity. A style search can favor shape, pattern, color, and material. A multimodal query must give the text instruction enough weight to change the result set.

How visual search works

Implementations vary, but most visual search systems have the same logical stages.

  1. Capture and select the target. The shopper uploads an image, takes a photo, pastes an image URL, or selects an object in a busy scene. Object detection, segmentation, a crop tool, or a user-drawn box can identify what to search.
  2. Create a visual representation. An image encoder turns pixels into a numerical representation, often called an embedding. Similar images land closer together in that representation space. A vision-language model can place text and images in a shared space so a phrase such as “brown velvet chair” can refine an image query.
  3. Retrieve candidates. The system searches an index of product-image representations, often with approximate nearest-neighbor retrieval. It returns a candidate set rather than making a final purchase decision.
  4. Apply catalog and business filters. Category, brand, region, price, stock, compatibility, size, and other structured fields remove candidates that look right but cannot satisfy the request.
  5. Rank and present results. The service combines visual similarity with exactness, text relevance, availability, variant logic, diversity, and sometimes behavioral signals. The interface should show the product record that matches the image, rather than an orphaned image URL.

Research on product visual-search embeddings describes learning image representations for content-based retrieval. Newer image-indexing systems also support image and text queries against one index. The architecture is the useful concept. You do not need to train a model from scratch to use it.

Visual search pipeline from an image query to filtered product results.

The catalog remains part of the retrieval system. An embedding can identify a close-looking shoe, but it cannot tell a shopper whether the shoe is available in size 9 unless the result joins to a current product record.

These terms overlap, which creates implementation confusion.

Search typeInputPrimary jobTypical ecommerce dependency
Text searchWords such as “waterproof trail shoes”Match language to product recordsTitles, descriptions, synonyms, attributes, and filters
Image searchA text query or an image used to find indexed images and pagesRetrieve images or pages that contain related imageryImage index, page context, metadata, and crawl coverage
Visual product searchA photo or screenshot used to find products or product familiesInterpret the visual object and return purchasable matchesReference images, embeddings, product identity, attributes, variants, price, and stock

Reverse image search can find the same image elsewhere on the web. Visual product search can find a product that looks like the image even when the pixels differ. An owned-store experience may combine both with text search, filters, and recommendations.

Visual search is useful when a shopper knows what an item looks like but does not know what it is called or how a catalog describes it.

Photograph-to-product lookup

A shopper photographs a lamp in a hotel, a backpack on a train, or a package on a shelf. The system can identify an exact product if the item and its reference images are in the index. If it is not, the experience should say that it found similar products rather than implying certainty.

Similar-product discovery

An out-of-stock product can lead to alternatives with a related silhouette, pattern, or material. This is a discovery use case, not a compatibility guarantee. A visually similar phone case, replacement filter, or machine part may still be functionally wrong. Compatibility attributes must remain a separate constraint.

Shop-the-look inspiration

A shopper uploads an outfit, room, or screenshot from social media. Object selection can isolate the chair, shoes, or bag, then return products with a related style. The shopper can move from inspiration to a product page without knowing the original brand.

Image plus text refinement

Images express appearance well. Text expresses constraints that an image cannot reliably prove. A shopper might upload a sofa and add “smaller, brown, under $1,000,” or upload a dress and add “long sleeve.” The search service needs structured fields for those constraints, and the interface should make it clear which part came from the image and which part came from the text.

Marketplaces can use a camera to start a listing or find related offers. A retailer can recognize packaging, barcodes, or a product family and then route the shopper to the right record. Industrial catalogs can use a photograph of a component as a starting point, while still requiring dimensions, standards, and compatibility checks before purchase.

Google describes a similar journey in its Lens shopping example: a shopper can photograph or upload an item, see product details across retailers, and refine some fashion and furniture searches with text. A platform example shows what the interaction can look like. It does not guarantee that every catalog receives the same exposure.

Product-image and catalog-data requirements

Visual search quality depends on two connected assets: the image index and the product data joined to each image. Better embeddings cannot repair missing variants, stale stock, or an image that does not clearly show the product.

Image requirements

  • Show the product clearly. Use a clean main image and additional angles that expose the shape, details, and scale. Avoid images where another product dominates the frame.
  • Cover realistic conditions. Include useful views for front, side, back, different orientations, lighting, and context. Google Cloud’s reference-image guidance describes three to eight representative images for one managed service. That is service-specific guidance, not a universal requirement.
  • Keep the product large enough to inspect. Follow the destination’s file, format, and dimension rules. For example, Google Merchant Center’s current guidance recommends images around 1,500 by 1,500 pixels or larger for best performance and describes a 500 by 500 pixel minimum for products beginning January 31, 2027. Those numbers are channel requirements and recommendations, not a visual-search ranking guarantee.
  • Map every image to an identifier. Store a stable product ID, variant ID, image role, and image URL together. Do not leave the index with an image that cannot resolve to a sellable record.
  • Keep image URLs stable and crawlable. A broken, blocked, or frequently changing URL creates gaps between the image index, product page, and downstream feeds.

Catalog-data requirements

RequirementFields to keep consistentWhy it matters
Product identityProduct ID, SKU, GTIN or other identifier, canonical URLJoins the visual result to the right product page and order data
Variant relationshipsParent ID, variant ID, color, size, material, pack or configurationStops one visual family from collapsing every size or color into one result
Visual attributesColor, pattern, finish, silhouette, material, texture, style, dimensionsGives filters and multimodal ranking facts that pixels cannot reliably supply
Taxonomy and brandCategory path, product type, brand, collectionNarrows candidate sets and makes results easier to explain
Commercial statePrice, currency, region, availability, inventory, sale statusPrevents the experience from sending shoppers to an unavailable or misleading offer
Freshness and provenanceSource, update time, feed status, approval stateMakes changes traceable when an image or product fact is wrong

The same discipline improves ordinary product data quality, product attributes, and ecommerce filters. Normalize values such as “navy,” “navy blue,” and “midnight” when they mean the same thing. Keep values such as “navy” and “black” distinct when they affect the buying decision.

Structured data and feeds matter when an external channel uses them. Google’s product structured data guidance says product markup can make details such as price, availability, ratings, and shipping eligible for richer Search, Images, and Lens experiences. Eligibility and display are not guaranteed. Your product page, feed, image index, and checkout should still agree even when a channel does not require markup.

Treat visual search as a retrieval product with a catalog operating model, rather than as a one-time model integration.

1. Start with one search job and category

Choose a category where appearance carries intent, such as apparel, furniture, beauty, decor, or visually distinctive parts. Define whether the first release is for exact identification, similar discovery, or image-plus-text refinement. A narrow job makes relevance labels and failure analysis possible.

2. Define what counts as a good result

Create a representative test set of customer photos, screenshots, crops, lighting conditions, and multi-object scenes. For each query, label the expected outcome:

  • exact product or acceptable variant;
  • acceptable similar product;
  • wrong category or irrelevant result;
  • no result because the catalog does not contain a suitable item.

Keep exact and similar labels separate. A model that returns stylish alternatives can look good on a similarity metric while failing an exact SKU lookup.

3. Choose the service boundary

You can use a managed visual or multimodal search service, a commerce-search platform with image support, or an internal retrieval service. An API is an integration boundary, not a quality strategy. Confirm current support, regions, limits, index-refresh behavior, image retention, and pricing before selecting a service. Google Cloud marks its older Vision API Product Search as maintenance mode. Its docs still illustrate product sets and reference images, while newer Image Warehouse documentation describes image and text similarity search with annotations. Microsoft’s Bing Search APIs, including its Visual Search API, were retired on August 11, 2025. Old tutorials are not implementation plans.

Keep text search and filters available. A camera-first feature should give shoppers a way to add words, select a target, correct the category, or continue with a normal query.

4. Build the index and its update path

Generate representations for the approved product images, attach product and variant metadata, and make catalog changes flow to the index. Add new products. Remove discontinued products. Re-embed changed imagery. Update availability and price through the path that serves results, not on a slower manual schedule. Monitor the lag from a source change to the result page.

For teams preparing this shared product data layer, Catalog structures product facts, normalizes attributes, keeps live data synchronized, and publishes machine-readable product objects to AI shopping surfaces. Catalog does not replace the visual retrieval service. It helps make the records that visual search returns consistent across the storefront, feeds, and AI commerce channels.

5. Measure offline relevance and online outcomes

Use both relevance tests and shopper behavior. Useful measures include:

MeasureWhat it tells you
Exact recall at 5 or 10Whether the known product appears in the first results
Similar-result precisionHow many shown alternatives are judged relevant by a trained reviewer
No-result rateWhere the query or catalog has no usable match
Refinement rateHow often shoppers add text, recrop, change category, or search again
Result click-through and add-to-cart rateWhether results lead to useful product exploration
Conversion, return, and support rateWhether the match holds up after the click and purchase
Latency, especially p95Whether the experience responds fast enough on real devices
Catalog coverage and freshness lagWhether the index contains the products and current state it claims to search

Review results by category, query type, device, image quality, and in-stock status. Do not compare raw similarity scores across unrelated queries without checking the service’s scoring behavior. Google’s Product Search guidance, for example, says result scores are useful for ordering within a query but are not calibrated for one universal threshold across queries.

Common failure modes and limits

Visual search has predictable failure modes. Design for them instead of hiding them behind confident labels.

FailureWhat goes wrongPractical guardrail
Lookalike, wrong productSimilar color or shape outranks the correct function or brandSeparate exact and similar modes; apply compatibility and category filters
Cluttered or occluded queryThe model reads the background, another object, or too little of the targetOffer object selection, cropping, and a text fallback
Thin image coverageOne front-facing packshot does not represent the item in real useAdd representative angles and context images, then test by query condition
Variant collapseA blue medium result resolves to a red large variant or parent productIndex parent-child relationships and return the chosen variant explicitly
Stale index or commerce dataA result is discontinued, unavailable, or shown at the wrong priceSync product changes and monitor freshness lag
Popularity biasLarge or frequently clicked products dominate visually relevant resultsSet category, diversity, and availability guardrails; audit tail products
Slow or costly retrievalHigh-resolution inference and large indexes make camera search frustrating or expensiveUse an appropriate embedding size, candidate limits, caching, and measured refresh schedules
Privacy and rights riskUploaded photos may contain people, private spaces, or images you cannot storeDocument retention and deletion, request consent where needed, and restrict access to query images
Unsupported implementation adviceA tutorial points to a retired API or an old modelCheck the provider’s current docs, release notes, regions, and limits before building

Visual similarity is evidence for discovery. It is not proof of identity, safety, fit, authenticity, or compatibility. Product records and human review still matter for high-risk categories.

FAQ

No. Visual search is a second entry point for shoppers who can show an item more easily than they can name it. Text search remains better for constraints such as size, compatibility, budget, delivery, and use case. The strongest ecommerce experiences combine image, text, filters, and a clear fallback.

Does visual search require AI?

Modern systems usually use computer vision and machine-learned image representations. Exact duplicate detection can use simpler image hashes, while product discovery typically needs a learned representation and a retrieval index. Generative AI is optional. The system still needs good product data whether the model generates language or only returns nearest neighbors.

What data does a visual search system need?

It needs query and reference images tied to stable product and variant IDs. It also needs categories, brand, normalized attributes, identifiers, canonical URLs, current price and availability, and a process for updating the index. Feeds or structured data may be required by an external shopping surface, while an owned-store implementation can use its own catalog API.

Is a visual search API required?

An API is common when a web or mobile interface sends an image to a managed service and receives product IDs. It is not the only deployment model. A commerce platform may provide the feature inside its search stack, or a team may host retrieval internally. Choose the boundary that lets you control catalog updates, privacy, latency, and evaluation.

How do I improve visual search results?

Start with clearer, representative images and complete image-to-variant mappings. Normalize color, material, shape, and category values. Keep product, price, availability, and URL data current. Then test exact and similar queries separately, review no-result and refinement sessions, and fix the catalog gaps that explain poor results.

Can visual search identify an exact product from any photo?

No. Exact identification depends on image quality, the object selected, reference-image coverage, catalog coverage, and the service’s ability to distinguish similar items. When certainty is low, label results as similar and give the shopper ways to refine the query.

Visual search works when the image query and the product record tell the same story. If your team needs help turning scattered product information into live, normalized data for AI commerce and product discovery, see how Catalog can help.