Advanced
Open
Pro
LoRA's Parameter Math for a Diffusion U-Net
Part of the AI Engineer Interview path →
Part of the ML System Design Interview path →
Part of the Generative Vision & Image AI System Design path →
Your team is deciding whether to fine-tune the full diffusion U-Net or
use a LoRA adapter for a headshot personalization feature. One of the
U-Net's cross-attention projection matrices is 2048 x 2048.
- Compute the number of trainable parameters for a full fine-tune of just this one matrix, versus a rank-16 LoRA adapter on the same matrix. What percentage of the full count does the LoRA adapter use?
- State the general parameter-count formula LoRA uses, and explain in one sentence why it's cheap.
- This subject says LoRA is "the same trick" as
fine-tuning-sft-lora-rlhf-dpoteaches for LLMs. What exactly stays the same between the two, and what's different?
Share this question