Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

Choosing an Identifier Token for DreamBooth

A teammate is implementing DreamBooth for the headshot product and proposes two candidate identifiers for the fine-tuning prompt "a photo of [V] person":

  • Option A: use the common word "human" as [V].
  • Option B: use a randomly generated string, "zqxpj42", as [V].
  1. Explain specifically why each option is likely to fail, in mechanistic terms (not just "it's a bad idea").
  2. What property should the actual identifier have, and what's the canonical example from the DreamBooth paper?
  3. Why does the identifier get paired with a class noun ("person") rather than used alone?
Solution

1. Why each option fails

  • Option A ("human"): this is an existing, heavily-used word with a strong, well-populated meaning already baked into the pretrained model from millions of training images. Fine-tuning the entire network on 15 photos of one subject under this word forces the model to reconcile two conflicting meanings for the same token — "this word means thousands of different people" and "this word now specifically means my one subject." The result is language drift: the word's general meaning shifts toward the fine-tuned subject, degrading unrelated generations that happen to use the word "human" elsewhere, not just prompts that deliberately invoke the identifier.
  • Option B ("zqxpj42"): the intuitive fix — pick something novel — fails for a subtler reason. The text encoder doesn't process raw characters; it processes subword tokens from a fixed vocabulary. A "random" string like this frequently decomposes into several existing subword tokens, each carrying its own pre-existing associations from pretraining. The identifier isn't actually novel in embedding space — it's an unusual combination of tokens that each still pull the representation in their own pretrained directions, which weakens how cleanly the model can bind the identifier to only the new subject.

2. The right property, and the canonical example

The identifier should be rare (infrequently used across the pretraining corpus, so it carries a weak, largely unentangled prior) and short — ideally a single token, so there's nothing to decompose and re-entangle. The DreamBooth paper's canonical example is "sks": rare enough in ordinary English to have almost no pre-existing meaning to fight against, short enough to bind as one clean unit.

3. Why pair the identifier with a class noun

Pairing gives full fine-tuning a place to anchor the general concept (the base model already understands "person" well) while the identifier absorbs what's specific to this one subject. It's also what makes class- specific prior preservation loss possible: the same class noun used in the identifier prompt is also used, without the identifier, in the prior-preservation prompt against the frozen model's own generic samples — the shared class noun is the hook the prior term uses to keep the class's general meaning anchored while the identifier specializes.

Share this question

← Back to Personalizing Image Generation: DreamBooth, LoRA and Textual Inversion practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.