Diagnosing a DreamBooth Model That Forgot Its Class
A team trains DreamBooth on 15 photos of a user, using the prompt
"a photo of [V] person", with no prior preservation loss (reconstruction
loss only). After training, two problems show up:
- Generated images using
[V]look almost identical to the exact poses and backgrounds in the 15 training photos, even when the prompt asks for a different setting. - Generating
"a photo of a person"with no identifier now produces images that look suspiciously like the fine-tuned subject too.
- Name each failure mode and explain the mechanism behind it.
- Explain what class-specific prior preservation loss adds to training, and where its "generic class" training images come from.
- Write the combined training objective at the level of detail this subject uses, and explain what the weighting term controls.
1. The two failure modes
- The first is overfitting: with only 15 training images and no counter-pressure, the model has enough capacity (it's fine-tuning the entire network) to memorize the exact poses, expressions and backgrounds in the training set rather than learning a generalizable representation of the subject that holds up in new poses and settings.
- The second is catastrophic forgetting via language drift on the
class noun: because
"person"appears in every training prompt right next to the identifier, full-model fine-tuning pulls the general meaning of"person"toward this one particular subject. The bare class prompt, with no identifier at all, starts reflecting the fine-tuned subject because nothing during training pushed back against that drift.
2. What prior preservation adds, and where its data comes from
Prior preservation loss adds a second training signal that anchors the
class noun's general meaning throughout fine-tuning. Its generic-class
images are sampled from the prompt "a photo of a person" using the
frozen, pre-fine-tuning model itself, generated once before training
starts. Because these images come from the model's own prior — not new
external data — training against them every step is literally reminding
the model, continuously, what it already knew about the class before
fine-tuning began, preventing that knowledge from being silently
overwritten by 15 photos of one person.
3. The combined objective
L_total = L_reconstruction([V] person, user's photos)
+ λ · L_prior("a person", frozen model's own generic samples)
Both terms are the same diffusion noise-prediction MSE loss — the only
difference is which prompt/image pairs feed each term. λ is the
weighting factor that balances the two: too low, and the class term is
too weak to prevent drift (reproducing the two failure modes above); too
high, and the class term dominates enough to drown out the
subject-specific signal, weakening how well [V] binds to the actual
subject. Tuning λ is the practical lever for finding the point where
the subject is faithfully learned without losing the class's general
meaning.
Share this question