Advanced
Open
Pro
Debugging a U-Net That Ignores the Prompt
Part of the AI Engineer Interview path →
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
A U-Net-based diffusion model is generating realistic, high-quality images, but they consistently fail to match the text prompt — "a red car on a mountain road" produces a plausible car on a plausible road, but the car is frequently the wrong color and the setting is often generic rather than mountainous. Reconstruction/denoising quality itself (sharpness, realism) is not in question.
- Where in the U-Net architecture does text conditioning actually happen, and what are the queries, keys, and values at that point?
- Given the symptom described, propose two concrete places to look for the bug, tied to the mechanism from part 1.
- A teammate proposes fixing this by only injecting the text embedding once, at the very start of the network, instead of at every block. Evaluate this proposal.
Share this question