AI · Computer Vision · 2026
80,000 raw product photos, transformed into a searchable, self-describing catalog.

The client’s catalog was 80,000+ product images with barely any metadata — invisible to search, impossible to recommend from. We built a multimodal enrichment pipeline that makes the images describe themselves: OpenCLIP generates visual-similarity embeddings, BLIP-3 writes accurate captions automatically, and Segment Anything v2 isolates products from cluttered photography.
The enriched embeddings are indexed into ChromaDB and Vespa, unlocking the features shoppers expect from a modern storefront: semantic text search that understands intent, "find similar" visual discovery, and personalized recommendations — all computed offline against the full catalog, with no real-time inference constraints.




Let’s talk about what you’re shipping next.

Opening project
Semantic Commerce Search