Match a job Paths Subjects Questions Quizzes Pricing

Case Study: AI Product Photography for a Million-SKU Catalog

Design the batch pipeline that turns one seller phone photo into studio-quality listing and ad images — the segment-then-generate cascade that keeps product pixels sacred

Overview Read

Case Study: AI Product Photography for a Million-SKU Catalog

"Sellers upload one phone photo of a product; produce studio-quality, brand-consistent listing and ad images at catalog scale" is the interview prompt behind a feature every major marketplace has shipped some version of in the last few years: Amazon's ad-image generation tooling for sellers, Shopify Magic's product-photo editing, and standalone tools like Photoroom and Flair AI that exist almost entirely to solve this one problem. The business case is blunt. A traditional studio product shoot — a photographer, lighting, a set, retouching — costs on the order of tens of dollars per SKU and takes days of lead time to schedule and deliver; a marketplace with a million active SKUs and tens of thousands of new listings a day cannot staff its way out of that cost structure, and a seller who has to wait three days for photos before a listing goes live is a seller who lists somewhere else, or not at all. The product being designed here replaces the studio shoot with a phone photo and a pipeline: a seller snaps one picture of a product against their kitchen counter, and the system returns a clean white-background image for the listing, one or more lifestyle-scene images for the product page, and a set of pre-cropped ad formats sized for the marketplace's own placements and for the seller's paid social campaigns.

This is the fourth case study in this track's generative-vision sequence, and it is deliberately the batch-offline member of the block — a label worth stating up front because it is the single fact that most changes the shape of every step that follows. case-study-virtual-try-on's personal mode and case-study-generative-fill-inpainting-and-outpainting's editing surface are both interactive: a shopper or an editor is watching the screen, waiting on a result, and the whole system is built around a sub-five-second budget. Nothing here is watched in real time. A seller uploads a photo, walks away, and comes back later — minutes later, sometimes longer — to a finished listing. That single difference reopens design decisions those two cases had already closed: batching strategy, spot-instance economics, quality gating without a human in the request path, and what "scalable" even means change shape once nobody is staring at a spinner.

This subject follows the seven-step case-study framework ml-system-design-interview-framework and case-study-fraud-detection both use. It leans on generative-adversarial-networks-and-face-generation for the GAN-vs-diffusion comparison table Step 3's architecture argument is built on, on text-to-image-diffusion-models for diffusion and latent-space vocabulary, and — most directly — on case-study-generative-fill-inpainting-and-outpainting, whose mask-conditioning, crop-around-the-mask, and mask-distribution-engineering vocabulary this case reuses rather than re-derives: "background generation conditioned on the cutout" is this case's version of that subject's "inpainting as a constraint," just inverted (fill everything outside the product mask instead of inside an edit mask).


Pro content

Sign up free, then start a 14-day Pro trial — no card needed.

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.