Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

Debugging a U-Net That Ignores the Prompt

A U-Net-based diffusion model is generating realistic, high-quality images, but they consistently fail to match the text prompt — "a red car on a mountain road" produces a plausible car on a plausible road, but the car is frequently the wrong color and the setting is often generic rather than mountainous. Reconstruction/denoising quality itself (sharpness, realism) is not in question.

  1. Where in the U-Net architecture does text conditioning actually happen, and what are the queries, keys, and values at that point?
  2. Given the symptom described, propose two concrete places to look for the bug, tied to the mechanism from part 1.
  3. A teammate proposes fixing this by only injecting the text embedding once, at the very start of the network, instead of at every block. Evaluate this proposal.

Share this question

← Back to Text-to-Image Generation with Diffusion Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.