Advanced
Open
Pro
A Discriminator That's Too Good, Too Fast
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
A different training run shows a different symptom from the previous question: the discriminator's loss drops to near zero within the first few hundred steps and stays pinned there for the rest of training, but unlike before, you've already confirmed the team is using the non-saturating generator loss. The generator still barely improves.
- Why doesn't switching to the non-saturating loss fully solve this particular symptom?
- Explain, without deriving the underlying math, what WGAN's "critic estimating a distance" reframing changes relative to a standard classifier-style discriminator, and why that change helps here specifically.
- A colleague asks: "if WGAN fixes this, why does WGAN-GP exist at all — what was wrong with the original WGAN?" Answer their question.
Share this question