PU21-MS-SSIM
Global multiscale structure after absolute-nit PU21 encoding.
Laux = 1 − MS-SSIM(PU(Lrec), PU(Ltgt))Screen mean JOD 9.090
Eight non-LR ideas were tested on Toys, Rushmore, Water, and Castle. The screen used 100 iterations; the top three perceptual candidates and independently selected global-LR L2/Charbonnier baselines were rerun for 300 iterations. All results use the same native ROI, no padding, Gumbel/binary hologram pipeline, seed 0, and deterministic hard reconstruction metrics.
The first eight-way screen used display-referred sRGB Charbonnier as the anchor. It revealed that the independently tuned amplitude L2 baseline was substantially stronger in the 4-frame regime, so the four most promising auxiliaries were rechecked with amplitude L2 as the anchor. Each candidate adds one perceptual direction with a fixed gradient budget, not a raw scalar weight:
A is the native reconstructed optical amplitude. The detached ratio equalizes auxiliary gradient RMS across metrics, so LPIPS/DISTS/MS-SSIM numerical scale cannot silently act as a different learning rate. Finalists labeled L2-GB use amplitude L2 as Lanchor; the original GB candidates use sRGB Charbonnier.
Global multiscale structure after absolute-nit PU21 encoding.
Laux = 1 − MS-SSIM(PU(Lrec), PU(Ltgt))Screen mean JOD 9.090
Laplacian bands divided by frozen target local contrast; emphasizes visible structure without losing the pixel anchor.
Laux = meanₖ ρ((Bk(rec)−Bk(tgt))/(σk(tgt)+0.02))Screen mean JOD 9.046
First-order edge agreement in perceptually uniform absolute luminance.
Laux = ½[ρ(∂xΔPU)+ρ(∂yΔPU)]Screen mean JOD 8.918
Robust absolute-PU residual at four progressively low-pass/downsampled scales.
Laux = ¼ Σₖ ρ(Gk(PUrec)−Gk(PUtgt))Screen mean JOD 8.912
Learned deep feature similarity on display-referred sRGB, resized only for the feature trunk.
Laux = LPIPS-Alex(sRGBrec, sRGBtgt)Screen mean JOD 8.977
Learned texture/structure similarity on display-referred sRGB.
Laux = DISTS(sRGBrec, sRGBtgt)Screen mean JOD 8.887
Wavelet similarity with local perceptual pooling.
Laux = HaarPSI-loss(sRGBrec, sRGBtgt)Screen mean JOD 9.101
Gradient/chromaticity similarity with deviation pooling.
Laux = MDSI-loss(sRGBrec, sRGBtgt)Screen mean JOD 8.664

| Rank | Loss | Global LR | Mean JOD | Mean PSNR | Mean SSIM |
|---|---|---|---|---|---|
| 1 | L2-GB · HaarPSI | 0.2 | 9.306 | 28.97 | 0.6527 |
| 2 | L2-GB · PU21-MS-SSIM | 0.2 | 9.276 | 29.24 | 0.6693 |
| 3 | L2-GB · PU21-NLPD | 0.2 | 9.266 | 29.27 | 0.6738 |
| 4 | Amplitude L2 | 0.2 | 9.178 | 29.25 | 0.6623 |
| 5 | GB · HaarPSI | 0.1 | 9.101 | 28.55 | 0.6467 |
| 6 | GB · PU21-MS-SSIM | 0.1 | 9.090 | 28.62 | 0.6544 |
| 7 | GB · PU21-NLPD | 0.1 | 9.046 | 28.62 | 0.6555 |
| 8 | L2-GB · LPIPS-Alex | 0.2 | 9.023 | 28.25 | 0.6222 |
| 9 | GB · LPIPS-Alex | 0.1 | 8.977 | 28.54 | 0.6456 |
| 10 | sRGB Charbonnier | 0.1 | 8.930 | 28.55 | 0.6468 |
| 11 | GB · PU21 gradient | 0.1 | 8.918 | 28.59 | 0.6556 |
| 12 | GB · PU21 multiscale | 0.1 | 8.912 | 28.59 | 0.6414 |
| 13 | GB · DISTS | 0.1 | 8.887 | 28.48 | 0.6432 |
| 14 | GB · MDSI | 0.1 | 8.664 | 26.07 | 0.5374 |



| Rank | Loss | Global LR | Mean JOD | Δ vs L2 | Δ vs Charb | Mean PSNR | Mean SSIM |
|---|---|---|---|---|---|---|---|
| 1 | L2-GB · PU21-MS-SSIM | 0.2 | 9.448 | +0.094 | +0.335 | 30.98 | 0.7391 |
| 2 | L2-GB · PU21-NLPD | 0.2 | 9.438 | +0.084 | +0.325 | 31.00 | 0.7441 |
| 3 | L2-GB · HaarPSI | 0.2 | 9.437 | +0.082 | +0.323 | 30.54 | 0.7149 |
| 4 | Amplitude L2 | 0.2 | 9.355 | +0.000 | +0.241 | 30.98 | 0.7330 |
| 5 | sRGB Charbonnier | 0.1 | 9.113 | -0.241 | +0.000 | 30.26 | 0.7120 |
























The best finalist is L2-GB · PU21-MS-SSIM at mean JOD 9.448: +0.094 versus independently tuned amplitude L2 and +0.335 versus independently tuned sRGB Charbonnier. It beats L2 in 4/4 scenes.
This four-scene, four-frame result does not clear the +0.5 JOD target relative to L2.
The decisive comparison is against the independently LR-selected baselines at the same 4-frame capacity and 300-iteration budget. A positive mean alone is not enough: check whether the winner improves most scenes, preserves PSNR/SSIM, and shows a stable advantage after the ramp rather than only at iteration 100.
If the winner is consistent, the next confirmation should be the full 12-scene 4-frame set and then an 8/4/2-frame capacity curve. If it is inconsistent or below +0.5 JOD, the honest conclusion is that gradient-budgeted perceptual guidance is a controlled negative/weak result, not yet a publishable perceptual improvement.
Machine-readable selection: selection.json. Final logs and manifest are included beside this report.