Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Debugging a Generator That Stopped Improving

You're training a face-generation GAN from scratch. Early in training, the discriminator's loss drops sharply and stays very low, while the generator's loss stays high and essentially flat — generated images remain obviously blurry and unrealistic for far longer than expected, with almost no visible improvement epoch over epoch.

  1. Explain the minimax objective's two competing terms in plain language, and state which one the generator is trying to minimize.
  2. Given the symptom described, what specific problem with the textbook minimax generator loss is the likely cause, and why?
  3. What is the standard fix, and why does it address the specific mechanism from part 2 rather than just being "a different loss that happens to work better"?
Solution

1. The two competing terms

The discriminator D wants to maximize E[log D(x)] + E[log(1 - D(G(z)))] — assign high probability-of-real to actual real images x and low probability-of-real to the generator's fakes G(z). The generator G wants to minimize that same expression, which in practice means minimizing the second term, E[log(1 - D(G(z)))] — it wants D(G(z)) to be as high as possible, i.e. it wants its fakes classified as real.

2. Diagnosing the likely cause

This matches the textbook minimax loss's known gradient-saturation problem. Early in training, G is weak and produces images D can easily reject, so D(G(z)) is close to 0 — and log(1 - D(G(z))), evaluated near D(G(z)) = 0, sits in a region of the function where its gradient with respect to G's parameters is nearly flat. G is trying to minimize a quantity that gives it almost no gradient signal exactly when it's furthest from good and needs the most guidance — which matches the described symptom precisely: a discriminator that's quickly become very confident (very low, stable discriminator loss) paired with a generator that isn't improving (flat generator loss), rather than the two networks trading off against each other as training normally looks.

3. The fix, and why it addresses the actual mechanism

The standard fix is the non-saturating generator loss: instead of minimizing log(1 - D(G(z))), G is trained to maximize log D(G(z))) directly. This isn't an arbitrary swap to "a loss that happens to train better" — it's specifically chosen because, evaluated at the same weak-generator region (D(G(z)) near 0), log D(G(z)) has a much steeper, more informative gradient than log(1 - D(G(z))) does at that point, even though both formulations push G in the same overall direction (fool D). The fix targets the exact saturation region diagnosed in part 2: it restores a usable gradient signal precisely when the generator is weakest, which is when the original formulation gave the least. If this training run is using the textbook loss rather than the non-saturating variant, switching is the first thing to try before reaching for architectural changes.

Share this question

← Back to Generative Adversarial Networks: From Minimax to StyleGAN practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.