Single-objective perceptual losses for HDR phase-only CGH

A controlled search over 16 pure objectives and three difficulty axes. Final setting: 8 binary frames, 2×2 tied SLM logits, native ROI, original propagation and physical iris, 300 iterations, deterministic hard reconstruction.

Headline: reducing temporal frames alone makes L2-friendly speckle; narrowing the iris makes the task easier by low-pass denoising. Tying 2×2 SLM pixels creates the intended spatial resource bottleneck. Dense HDR encodings improve JOD modestly; pure structural metrics do not.

Choosing a genuinely harder task

1/2-frame mean L2 JOD: 7.933 / 8.815; native 4-frame: 9.355. A 0.75/0.5 iris unexpectedly raises L2 to 9.453/9.392, so it was rejected as a difficulty manipulation. A 2×2 tied SLM gives 7.772 at 4 frames and 8.099 at 8 frames.

Two batches of eight single losses

Batch A — pure structure: amplitude MS-SSIM, NLPD, MS-GMSD, HaarPSI; PU21-L1, PU21-MS-SSIM, PU21-NLPD; fixed exposure-stack MS-SSIM. These optimize stably but lose radiometric fidelity and JOD.

Batch B — dense perceptual encodings: sRGB-L2, amplitude Charbonnier, PU21-L2, PU21-Charbonnier, PQ-L2, μ-law(5000)-L2, HLG-L2, exposure-stack L2. Each is one transformed residual—not A+B and not an LR-only variant.

Dense objectiveBest LR (2-frame)2-frame JODSpatial-screen JOD
sRGB L20.18.6238.082
Exposure-stack L20.18.4968.111
PU21-L20.28.4698.148
PQ-L20.28.4418.153
Amplitude Charbonnier0.18.428
HLG-L20.18.414
PU21-Charbonnier0.18.332
μ-law(5000)-L20.18.051

12-scene, 300-iteration finalists

LossMean JODΔ vs tuned L2Mean PSNRMean SSIM
Amplitude L28.205+0.00023.840.5532
sRGB Charbonnier7.913-0.29223.700.5487
sRGB L28.268+0.06323.970.5532
PU21-L28.237+0.03223.590.5430

Hard reconstruction gallery

Interpretation

The gain is real but not the requested +0.5 JOD. The spatial bottleneck reveals a useful distinction: dense monotonic HDR encodings retain the pixel accountability needed by Gumbel binary optimization, while pure similarity metrics can improve their own score by sacrificing absolute brightness. PU21/PQ help low-luminance Toys most but can hurt bright Water; sRGB-L2 is more consistent. This points to brightness-conditioned encoding as the next hypothesis, but the present report does not count per-scene oracle selection as a result.

Sources

Raw output: 2.79 GiB under /home/fy277/ph_opt_runs (10 GiB cap).