Two Products, One Prompt: Scoping the On-Device and Cloud Paths
An interviewer says: "Design the feature that upscales and restores users' photos." A candidate immediately starts sketching a single RRDB-based GAN generator with an ESRGAN-style loss and asks what resolution to target.
- What is missing from the candidate's framing before any architecture discussion should start?
- Propose concrete assumption-table numbers (latency, upscale factor, compute location) for the two product points this prompt actually requires, and explain why a single set of numbers can't cover both.
- A teammate suggests including NVIDIA DLSS-style real-time game upscaling as a third product point to design in full. Should you? Explain what you'd say instead.
1. What's missing
The candidate has skipped straight to a model without clarifying which of several genuinely different products "super-resolution and restoration" refers to. A phone gallery upscale, a cloud photo restoration tool, an e-commerce zoom feature, and real-time game upscaling share a name but have almost nothing else in common — different latency budgets (milliseconds to minutes), different compute locations (on-device NPU vs. cloud GPU pool), different degradation severity (mild sensor noise vs. decades-old scratched film), and different cost structures (near-zero marginal cost vs. a real chargeable job). Picking one architecture before scoping which product point(s) are in play risks designing a system that's wrong for either of the two the prompt is actually asking about.
2. Concrete numbers for the two product points
On-device gallery upscale: 2-4x upscale factor, < 2 s end-to-end on a mid-tier phone NPU, tens-of-MB quantized model size ceiling, near-zero marginal cost per use, no human review, modest input degradation (standard phone camera noise and JPEG compression).
Cloud restoration: 8x+ upscale factor, seconds-to-minutes acceptable (async job, not interactive), no hard model-size ceiling, a real per-job GPU cost, heavy and varied real-world degradation (scratches, fading, unknown compression history), and a paid/pro tier that can support light human before/after review on a small fraction of jobs.
A single set of numbers can't cover both because the constraints are in direct tension: the on-device budget rules out anything but a small, fast, single-pass model, while the cloud budget's much heavier degradation and higher target quality (face restoration, scratch repair) genuinely benefit from a larger, slower model family — designing to satisfy both simultaneously would either make the on-device path too slow to ship or the cloud path too limited to be worth paying for.
3. Should DLSS-style real-time upscaling be a third fully-designed product point?
No — it should be named and explicitly scoped out, not designed in full, and the reason is itself worth stating: NVIDIA's public positioning of DLSS describes it as combining a trained upscaling network with motion vectors supplied by the game engine and temporal accumulation across frames, which means the model has access to privileged, per-pixel motion information a phone gallery app upscaling an already-captured JPEG simply doesn't have. It's also trained and shipped as a small model tightly co-designed with one specific rendering pipeline, not a general-purpose photo restorer. Designing it "in full" would mean designing a game-engine integration, which is a different system built around a different, much more constrained problem — the right move is to name it, explain the structural reason it's a different regime (engine-supplied motion vectors plus a narrow, specialized model), and keep the case focused on the two general-purpose product points the prompt is actually asking for.
Share this question