Match a job Paths Subjects Questions Quizzes Pricing
Advanced Open Free

What Temporal Layers Add to an Image-Only U-Net

Your team's video generator is a subject-2-style U-Net with cross- attention for text conditioning, applied frame-by-frame to a video latent. Generated clips have good per-frame quality but visibly flickering, inconsistent motion — an object's shape or color shifts noticeably between adjacent frames even when the prompt describes smooth, continuous motion.

  1. Explain, mechanistically, why a subject-2-style U-Net (2D convolutions + cross-attention) produces exactly this failure mode.
  2. Describe temporal attention and temporal (3D) convolution, and explain what each one specifically fixes.
  3. Where do these new layers get added relative to the U-Net's existing structure, and what stays unchanged?
Solution

1. Why the described U-Net produces flickering

Subject 2's U-Net layers — 2D convolutions and cross-attention — all operate within a single frame. 2D convolutions only ever see a local spatial neighborhood inside one frame's latent; cross-attention lets an image feature attend to the text embedding, not to features in other frames. Nothing in that design gives the network any mechanism to compare or coordinate what it's producing in frame N with what it produced in frame N-1 or N+1 — each frame is effectively denoised as if it were an independent image that merely happens to be part of a sequence, which is exactly the failure mode described: good per-frame quality, no cross-frame coordination.

2. What temporal attention and temporal convolution add

  • Temporal attention reuses the same attention mechanism as cross-attention, redirected: instead of an image feature attending to a text embedding, a feature at a given spatial position in one frame attends to the corresponding (and nearby) features in other frames. This is what lets the network learn "how does this region evolve smoothly across frames," directly fixing the flickering/inconsistency symptom by giving each frame's generation visibility into neighboring frames.
  • Temporal (3D) convolution extends the convolution operation with a third axis: instead of a kernel sliding across only height and width within one frame, it slides across a small window of consecutive frames as well as across space, capturing local motion patterns directly via the convolutional inductive bias, the same way 2D convolutions capture local spatial patterns.

Both address the same root problem — no cross-frame information flow — via two different mechanisms (attention-based long-range lookup vs. convolution-based local motion capture).

3. Where the new layers go

Temporal attention and temporal convolution are interleaved into the existing downsampling and upsampling blocks, alongside the 2D convolutions and cross-attention subject 2 already covers — they are added layers, not a replacement for the existing structure. The U-Net's core shape (a sequence of downsampling blocks, a bottleneck, a symmetric sequence of upsampling blocks) and its existing spatial and text-conditioning capabilities are unchanged; the network simply gains an additional capability (temporal coherence) it didn't have before.

Share this question

← Back to Text-to-Video Generation practice

We use cookies for product analytics to improve OmniAtlas. See our Privacy Policy.