Choosing an Identifier Token for DreamBooth
A teammate is implementing DreamBooth for the headshot product and proposes
two candidate identifiers for the fine-tuning prompt "a photo of [V] person":
- Option A: use the common word
"human"as[V]. - Option B: use a randomly generated string,
"zqxpj42", as[V].
- Explain specifically why each option is likely to fail, in mechanistic terms (not just "it's a bad idea").
- What property should the actual identifier have, and what's the canonical example from the DreamBooth paper?
- Why does the identifier get paired with a class noun (
"person") rather than used alone?
1. Why each option fails
- Option A (
"human"): this is an existing, heavily-used word with a strong, well-populated meaning already baked into the pretrained model from millions of training images. Fine-tuning the entire network on 15 photos of one subject under this word forces the model to reconcile two conflicting meanings for the same token — "this word means thousands of different people" and "this word now specifically means my one subject." The result is language drift: the word's general meaning shifts toward the fine-tuned subject, degrading unrelated generations that happen to use the word"human"elsewhere, not just prompts that deliberately invoke the identifier. - Option B (
"zqxpj42"): the intuitive fix — pick something novel — fails for a subtler reason. The text encoder doesn't process raw characters; it processes subword tokens from a fixed vocabulary. A "random" string like this frequently decomposes into several existing subword tokens, each carrying its own pre-existing associations from pretraining. The identifier isn't actually novel in embedding space — it's an unusual combination of tokens that each still pull the representation in their own pretrained directions, which weakens how cleanly the model can bind the identifier to only the new subject.
2. The right property, and the canonical example
The identifier should be rare (infrequently used across the
pretraining corpus, so it carries a weak, largely unentangled prior) and
short — ideally a single token, so there's nothing to decompose and
re-entangle. The DreamBooth paper's canonical example is "sks": rare
enough in ordinary English to have almost no pre-existing meaning to
fight against, short enough to bind as one clean unit.
3. Why pair the identifier with a class noun
Pairing gives full fine-tuning a place to anchor the general concept (the base model already understands "person" well) while the identifier absorbs what's specific to this one subject. It's also what makes class- specific prior preservation loss possible: the same class noun used in the identifier prompt is also used, without the identifier, in the prior-preservation prompt against the frozen model's own generic samples — the shared class noun is the hook the prior term uses to keep the class's general meaning anchored while the identifier specializes.
Share this question