Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Fixing a Design That Diffuses the Whole Canvas

A junior engineer's design doc for the generate-with-prompt path reads: "Encode the full user-uploaded image (up to 8000x8000 px) into the VAE's latent space, run the masked-diffusion U-Net over the entire latent, then decode the whole thing back to pixels." It works in their prototype on small test images.

  1. Explain concretely why this design will not meet the interactive latency budget once real users upload full-resolution photos, and why "it worked in the prototype" is misleading.
  2. Redesign the pipeline using the crop-around-the-mask principle. Be specific about what gets encoded/diffused and what doesn't.
  3. A user removes a single small logo (roughly 80x80 px) from an 8000x8000 px product photo. Compare, at a back-of-envelope level, the GPU cost of the original full-canvas design versus your redesign for this specific request.
Solution

1. Why the full-canvas design fails at real scale

Diffusion inference cost scales with the resolution the U-Net actually processes — more latent pixels means more compute per denoising step, larger activation memory, and (past a point) inputs too large to fit a single GPU's memory without additional sharding. A small prototype image (say, 512x512 or 1024x1024) sits comfortably inside the budget that made the design "work"; an 8000x8000 real upload is roughly 60-250x more latent pixels than that prototype size, at which point the same pipeline is either far too slow for a 3-5 second interactive budget, or it simply doesn't fit in memory at all. The prototype's apparent success is misleading specifically because it never exercised the resolution regime real users will actually upload at — a classic case of a design validated only against inputs that don't stress the part of the system that actually breaks.

2. The crop-around-the-mask redesign

Instead of encoding the whole canvas, compute a bounded crop centered on the user's mask, sized to roughly the model's native working resolution (on the order of 512-1024 px) with a margin around the mask large enough to give the model real surrounding context to condition on. Only that crop is VAE-encoded, diffused, and decoded; the rest of the 8000x8000 canvas is never touched by the generative model at all. The generated crop is then blended back into the full-resolution canvas at its original position (Poisson or latent blending at the crop's boundary), so the final output is full resolution even though only a small region of it was ever run through the expensive model.

3. Back-of-envelope cost comparison for the small-logo removal

Under the full-canvas design, the U-Net processes latent pixels proportional to the entire 8000x8000 canvas regardless of how small the edit actually is — the cost is dominated by canvas size, not edit size. Under the crop-around-the-mask redesign, the 80x80 px logo plus a reasonable context margin easily fits inside a single small crop (well under the model's native working resolution), so the cost is close to the cheapest possible case the model can run — comparable to inpainting a small image outright, essentially independent of the 8000x8000 canvas it sits inside. Illustratively, if full-canvas diffusion at that resolution would cost on the order of tens of GPU-seconds (or simply fail to fit in memory), the cropped version costs on the order of a couple of GPU-seconds for the same edit — a difference of roughly one to two orders of magnitude for this specific request, and the gap only grows for edits on even larger canvases with similarly small masks.

Share this question

← Back to Case Study: Generative Fill — Inpainting, Outpainting and Object Removal practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.