4-frame capacity-limited HDR CGH · hard binary evaluation

Gradient-budgeted perceptual loss screening

Eight non-LR ideas were tested on Toys, Rushmore, Water, and Castle. The screen used 100 iterations; the top three perceptual candidates and independently selected global-LR L2/Charbonnier baselines were rerun for 300 iterations. All results use the same native ROI, no padding, Gumbel/binary hologram pipeline, seed 0, and deterministic hard reconstruction metrics.

Outcome. The best finalist is L2-GB · PU21-MS-SSIM at mean JOD 9.448: +0.094 versus independently tuned amplitude L2 and +0.335 versus independently tuned sRGB Charbonnier. It beats L2 in 4/4 scenes. This four-scene, four-frame result does not clear the +0.5 JOD target relative to L2.

What is new in the optimization

The first eight-way screen used display-referred sRGB Charbonnier as the anchor. It revealed that the independently tuned amplitude L2 baseline was substantially stronger in the 4-frame regime, so the four most promising auxiliaries were rechecked with amplitude L2 as the anchor. Each candidate adds one perceptual direction with a fixed gradient budget, not a raw scalar weight:

L = Lanchor + r(t) · stopgrad(clamp(RMS(∂Lanchor/∂A) / RMS(∂Laux/∂A), 0.01, 100)) · Laux
r(t): 20-iteration anchor-only warm-up → smooth 60-iteration ramp → 0.25

A is the native reconstructed optical amplitude. The detached ratio equalizes auxiliary gradient RMS across metrics, so LPIPS/DISTS/MS-SSIM numerical scale cannot silently act as a different learning rate. Finalists labeled L2-GB use amplitude L2 as Lanchor; the original GB candidates use sRGB Charbonnier.

The eight ideas

PU21-MS-SSIM

Global multiscale structure after absolute-nit PU21 encoding.

Laux = 1 − MS-SSIM(PU(Lrec), PU(Ltgt))

Screen mean JOD 9.090

PU21-NLPD

Laplacian bands divided by frozen target local contrast; emphasizes visible structure without losing the pixel anchor.

Laux = meanₖ ρ((Bk(rec)−Bk(tgt))/(σk(tgt)+0.02))

Screen mean JOD 9.046

PU21 gradient

First-order edge agreement in perceptually uniform absolute luminance.

Laux = ½[ρ(∂xΔPU)+ρ(∂yΔPU)]

Screen mean JOD 8.918

PU21 multiscale

Robust absolute-PU residual at four progressively low-pass/downsampled scales.

Laux = ¼ Σₖ ρ(Gk(PUrec)−Gk(PUtgt))

Screen mean JOD 8.912

LPIPS-Alex

Learned deep feature similarity on display-referred sRGB, resized only for the feature trunk.

Laux = LPIPS-Alex(sRGBrec, sRGBtgt)

Screen mean JOD 8.977

DISTS

Learned texture/structure similarity on display-referred sRGB.

Laux = DISTS(sRGBrec, sRGBtgt)

Screen mean JOD 8.887

HaarPSI

Wavelet similarity with local perceptual pooling.

Laux = HaarPSI-loss(sRGBrec, sRGBtgt)

Screen mean JOD 9.101

MDSI

Gradient/chromaticity similarity with deviation pooling.

Laux = MDSI-loss(sRGBrec, sRGBtgt)

Screen mean JOD 8.664

100-iteration screen

RankLossGlobal LRMean JODMean PSNRMean SSIM
1L2-GB · HaarPSI0.29.30628.970.6527
2L2-GB · PU21-MS-SSIM0.29.27629.240.6693
3L2-GB · PU21-NLPD0.29.26629.270.6738
4Amplitude L20.29.17829.250.6623
5GB · HaarPSI0.19.10128.550.6467
6GB · PU21-MS-SSIM0.19.09028.620.6544
7GB · PU21-NLPD0.19.04628.620.6555
8L2-GB · LPIPS-Alex0.29.02328.250.6222
9GB · LPIPS-Alex0.18.97728.540.6456
10sRGB Charbonnier0.18.93028.550.6468
11GB · PU21 gradient0.18.91828.590.6556
12GB · PU21 multiscale0.18.91228.590.6414
13GB · DISTS0.18.88728.480.6432
14GB · MDSI0.18.66426.070.5374

300-iteration finalists

RankLossGlobal LRMean JODΔ vs L2Δ vs CharbMean PSNRMean SSIM
1L2-GB · PU21-MS-SSIM0.29.448+0.094+0.33530.980.7391
2L2-GB · PU21-NLPD0.29.438+0.084+0.32531.000.7441
3L2-GB · HaarPSI0.29.437+0.082+0.32330.540.7149
4Amplitude L20.29.355+0.000+0.24130.980.7330
5sRGB Charbonnier0.19.113-0.241+0.00030.260.7120

Visual comparison

Toys

Target
L2-GB · PU21-MS-SSIM
JOD 9.629 · PSNR 37.58 · SSIM 0.8770
L2-GB · PU21-NLPD
JOD 9.656 · PSNR 37.70 · SSIM 0.8852
L2-GB · HaarPSI
JOD 9.755 · PSNR 37.00 · SSIM 0.8654
Amplitude L2
JOD 9.599 · PSNR 37.44 · SSIM 0.8726
sRGB Charbonnier
JOD 9.326 · PSNR 36.51 · SSIM 0.8636

Rushmore

Target
L2-GB · PU21-MS-SSIM
JOD 9.407 · PSNR 29.67 · SSIM 0.6905
L2-GB · PU21-NLPD
JOD 9.399 · PSNR 29.64 · SSIM 0.6979
L2-GB · HaarPSI
JOD 9.201 · PSNR 29.10 · SSIM 0.6572
Amplitude L2
JOD 9.258 · PSNR 29.70 · SSIM 0.6833
sRGB Charbonnier
JOD 8.861 · PSNR 28.60 · SSIM 0.6467

Water

Target
L2-GB · PU21-MS-SSIM
JOD 9.422 · PSNR 28.39 · SSIM 0.7288
L2-GB · PU21-NLPD
JOD 9.401 · PSNR 28.39 · SSIM 0.7303
L2-GB · HaarPSI
JOD 9.446 · PSNR 28.19 · SSIM 0.7061
Amplitude L2
JOD 9.342 · PSNR 28.45 · SSIM 0.7232
sRGB Charbonnier
JOD 9.200 · PSNR 27.96 · SSIM 0.7024

Castle

Target
L2-GB · PU21-MS-SSIM
JOD 9.335 · PSNR 28.26 · SSIM 0.6601
L2-GB · PU21-NLPD
JOD 9.297 · PSNR 28.26 · SSIM 0.6631
L2-GB · HaarPSI
JOD 9.344 · PSNR 27.88 · SSIM 0.6310
Amplitude L2
JOD 9.220 · PSNR 28.33 · SSIM 0.6530
sRGB Charbonnier
JOD 9.066 · PSNR 27.99 · SSIM 0.6355

Interpretation and next decision

The best finalist is L2-GB · PU21-MS-SSIM at mean JOD 9.448: +0.094 versus independently tuned amplitude L2 and +0.335 versus independently tuned sRGB Charbonnier. It beats L2 in 4/4 scenes.

This four-scene, four-frame result does not clear the +0.5 JOD target relative to L2.

The decisive comparison is against the independently LR-selected baselines at the same 4-frame capacity and 300-iteration budget. A positive mean alone is not enough: check whether the winner improves most scenes, preserves PSNR/SSIM, and shows a stable advantage after the ramp rather than only at iteration 100.

If the winner is consistent, the next confirmation should be the full 12-scene 4-frame set and then an 8/4/2-frame capacity curve. If it is inconsistent or below +0.5 JOD, the honest conclusion is that gradient-budgeted perceptual guidance is a controlled negative/weak result, not yet a publishable perceptual improvement.

Machine-readable selection: selection.json. Final logs and manifest are included beside this report.