Objective
Four smooth 17-frame training trajectories sample continuous focal states. They are independent of the three held-out trajectories below.
Reference, the original all-plane baseline, and the selected focal-aware method under the same initialization, capacity and 1,000-iteration budget.
Four smooth 17-frame training trajectories sample continuous focal states. They are independent of the three held-out trajectories below.
Nine discrete targets cannot see errors that appear only between planes.
The loss matches focal motion and acceleration instead of applying CVVDP directly as a training loss.
It emphasizes perceptually important spatial and temporal residuals.
The nine-plane stack remains stable while the final 200 iterations optimize exact binary forward patterns.
Novel-video metrics average three unseen 125-frame trajectories; plane metrics average nine discrete focal planes.
| Metric | 1K baseline | Focal-Aware | Change |
|---|---|---|---|
| Mean novel-trajectory CVVDP JOD ↑ | 7.2209 | 8.0793 | +0.8584 |
| Mean novel saliency PSNR ↑ | 25.1111 dB | 26.6563 dB | +1.5452 dB |
| Mean novel saliency SSIM ↑ | 0.57085 | 0.67023 | +0.09938 |
| Mean novel full-frame PSNR ↑ | 26.0652 dB | 26.4263 dB | +0.3611 dB |
| Mean temporal-delta residual ↓ | 0.001494 | 0.001410 | −5.59% |
| Mean PSNR — 9 planes ↑ | 25.3873 dB | 25.3869 dB | −0.0004 dB |
| Mean SSIM — 9 planes ↑ | 0.51701 | 0.52452 | +0.00751 |
| Mean LPIPS — 9 planes ↓ | 0.50758 | 0.50061 | −0.00697 |
| Worst-plane PSNR ↑ | 24.8259 dB | 24.4988 dB | −0.3271 dB |
Each video is 5 seconds at 25 fps. The red circle guides gaze; the upper-right plot shows the continuous focal coordinate.
Center-biased scan across salient Dragon and Bunny regions.

Depth-weighted saliency emphasizes near Dragon/Bunny content.

Depth-gradient weighting crosses foreground/background boundaries.
