Practice — Case Study: Generative Fill — Inpainting, Outpainting and Object Removal (6 questions)
Arguing the Cascade Against a Single-Model Proposal Permalink →
A teammate on your Generative Fill team proposes simplifying the architecture: "We already have a solid mask-conditioned diffusion inpainting model for the generate-with-prompt path. Let's just use it for object removal too — pass an empty prompt when the user wants to remove something. One model, one service, less to maintain."
- Using the latency, quality, and diversity axes from the
GAN-vs-diffusion comparison table (
generative-adversarial-networks-and-face-generation), explain what this proposal gets wrong. - Removal requests are the majority of real Generative-Fill traffic. Quantify, at a back-of-envelope level, why this matters for the proposal's cost and latency impact, stating your assumptions.
- Is there any scenario in this product where routing a removal request through the diffusion path instead of the GAN path is actually the right call? If so, describe it; if not, explain why not.
Share this question
Fixing a Design That Diffuses the Whole Canvas Permalink →
A junior engineer's design doc for the generate-with-prompt path reads: "Encode the full user-uploaded image (up to 8000x8000 px) into the VAE's latent space, run the masked-diffusion U-Net over the entire latent, then decode the whole thing back to pixels." It works in their prototype on small test images.
- Explain concretely why this design will not meet the interactive latency budget once real users upload full-resolution photos, and why "it worked in the prototype" is misleading.
- Redesign the pipeline using the crop-around-the-mask principle. Be specific about what gets encoded/diffused and what doesn't.
- A user removes a single small logo (roughly 80x80 px) from an 8000x8000 px product photo. Compare, at a back-of-envelope level, the GPU cost of the original full-canvas design versus your redesign for this specific request.
Share this question