Practice — Case Study: AI Product Photography for a Million-SKU Catalog (5 questions)
Arguing Against End-to-End Scene Regeneration Permalink →
A teammate proposes simplifying the product-photography pipeline: "We already have a strong diffusion model. Let's skip the separate segmentation stage — just feed it the seller's phone photo and a prompt like 'studio product photo, white background, this exact product,' and let it regenerate the whole scene in one pass. Fewer moving parts, one model to maintain."
- Using the product-fidelity requirement from Step 1, explain concretely why this proposal fails, independent of how good the diffusion model's output looks.
- Is this a problem the GAN-vs-diffusion comparison table (latency, quality, diversity, training stability, controllability) can resolve by picking a different generative architecture? Why or why not?
- Propose the minimal architectural change that fixes the proposal while keeping "one pass through a generative model" as a goal, and explain what guarantee it does and does not give you.
Share this question
Debugging a Product-Region Invariance Failure at Scale Permalink →
A week after a model update to the scene-generation stage, the product-region invariance pass rate drops from 99.7% to 94% across the whole catalog. Aesthetic scores and brand-consistency classifier scores both improved slightly over the same period. The regenerate loop is absorbing most of the failures, so no obvious spike in shopper complaints has appeared yet.
- Why is "no spike in shopper complaints yet" not reassuring here, and what should you actually check first?
- Propose two distinct plausible root causes for a fidelity-gate regression that arrives at the same time as an aesthetic-score improvement, and how you'd distinguish between them.
- The regenerate loop is currently absorbing the failures. Explain the hidden cost this creates even though nothing has shipped broken, and what metric would surface it.
Share this question