Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Diagnosing a DreamBooth Model That Forgot Its Class

A team trains DreamBooth on 15 photos of a user, using the prompt "a photo of [V] person", with no prior preservation loss (reconstruction loss only). After training, two problems show up:

  • Generated images using [V] look almost identical to the exact poses and backgrounds in the 15 training photos, even when the prompt asks for a different setting.
  • Generating "a photo of a person" with no identifier now produces images that look suspiciously like the fine-tuned subject too.
  1. Name each failure mode and explain the mechanism behind it.
  2. Explain what class-specific prior preservation loss adds to training, and where its "generic class" training images come from.
  3. Write the combined training objective at the level of detail this subject uses, and explain what the weighting term controls.
Solution

1. The two failure modes

  • The first is overfitting: with only 15 training images and no counter-pressure, the model has enough capacity (it's fine-tuning the entire network) to memorize the exact poses, expressions and backgrounds in the training set rather than learning a generalizable representation of the subject that holds up in new poses and settings.
  • The second is catastrophic forgetting via language drift on the class noun: because "person" appears in every training prompt right next to the identifier, full-model fine-tuning pulls the general meaning of "person" toward this one particular subject. The bare class prompt, with no identifier at all, starts reflecting the fine-tuned subject because nothing during training pushed back against that drift.

2. What prior preservation adds, and where its data comes from

Prior preservation loss adds a second training signal that anchors the class noun's general meaning throughout fine-tuning. Its generic-class images are sampled from the prompt "a photo of a person" using the frozen, pre-fine-tuning model itself, generated once before training starts. Because these images come from the model's own prior — not new external data — training against them every step is literally reminding the model, continuously, what it already knew about the class before fine-tuning began, preventing that knowledge from being silently overwritten by 15 photos of one person.

3. The combined objective

L_total = L_reconstruction([V] person, user's photos)
        + λ · L_prior("a person", frozen model's own generic samples)

Both terms are the same diffusion noise-prediction MSE loss — the only difference is which prompt/image pairs feed each term. λ is the weighting factor that balances the two: too low, and the class term is too weak to prevent drift (reproducing the two failure modes above); too high, and the class term dominates enough to drown out the subject-specific signal, weakening how well [V] binds to the actual subject. Tuning λ is the practical lever for finding the point where the subject is faithfully learned without losing the class's general meaning.

Share this question

← Back to Personalizing Image Generation: DreamBooth, LoRA and Textual Inversion practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.