Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Pro

CFG and DDIM: What Each One Actually Buys You

Your text-to-image system currently samples with 1,000 denoising steps and no classifier-free guidance. Product feedback says two things: generated images often don't closely match the prompt, and generation is too slow for the product's latency budget.

  1. Explain what classifier-free guidance (CFG) does mechanically, and why it requires a specific choice made during training, not just at sampling time.
  2. Explain what DDIM-style sampling changes, and why it doesn't require retraining the model.
  3. A colleague proposes fixing both problems by simply setting the CFG guidance scale very high. Explain the risk in this approach and how it interacts with the diffusion-vs-autoregressive "flexibility" trade-off from earlier in this subject.

Share this question

← Back to Text-to-Image Generation with Diffusion Models practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.