Doc: Record 4K fast-mode accuracy limits and error distribution

Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README.

- Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2.
- Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile.
- Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
This commit is contained in:
wyj committed 2026-09-26 01:02:41 -04:00
1 parent 229f50cd86
commit e4d093a9fe
5 files changed
+400 -8

No files matched your search

+41 -4
View File
@@ -278,11 +278,12 @@ supersampled scale, so the requested FWHM and beta are preserved by the
downsample itself.
`--fast-deposit nearest` (default) deposits into the single nearest
supersampled pixel. It keeps the PSF shape exactly but quantizes the image
position to `1/(2N)` output pixels. `--fast-deposit bilinear` deposits into 4
supersampled pixel. It keeps the PSF shape exactly on that grid. The grid
spacing is `1/N` output pixel per axis, while the instantaneous rounding error
is bounded by `1/(2N)` per axis. `--fast-deposit bilinear` deposits into 4
adjacent pixels; it preserves the continuous centroid but broadens the profile
(measured +4.5%/+1.9%/+1.0% FWHM at N=2/3/4). The deposition experiment and its
raw output are in
(measured +4.5%/+1.9%/+1.0% FWHM at N=2/3/4). The deposition experiment and
its raw output are in
[benchmarks/fast_mode_deposit_2026-09-18.md](benchmarks/fast_mode_deposit_2026-09-18.md).
Fast mode is available only in the CPU PSF backend. It ignores
@@ -291,6 +292,42 @@ bright event whose requested support exceeds that radius is wing-clipped and
counted in the report. `--psf-min-y` is still applied per event. Fast mode is a
preview approximation, not the physically exact per-event PSF path.
Its accuracy is governed by the deposit mode and `N`, not by the output
resolution. A nearest deposit shifts each image by at most `1/(2N)` output
pixels per axis. The separate two-dimensional centroid-distance experiment
measured worst cases of ~`0.30 px` at `N=2` and ~`0.12 px` at `N=4` for a FWHM
`2.7`, beta `4.5` PSF. These offsets remain fixed fractions of one output
pixel: with `--psf-fwhm-pixels` held fixed, a larger image does not reduce
them, and only a larger `N` (or the centroid-preserving bilinear deposit, at
the cost of a broadened profile) does.
That per-star number is not how one dense still image behaves. In the measured
production-density 4K Galactic-center frame, fast-versus-standard HDR relative
L2 is `10.44%` at `N=2` and `4.37%` at `N=4`, with a `+0.018%` total
linear-RGB-sum offset. The difference is concentrated in a small set of scalar
RGB channel samples: the top `0.01%` ranked by `|fast-base|` account for
`95.7%` (`N=2`) or `94.5%` (`N=4`) of squared-error energy, while containing
about `9%` of the summed baseline linear-RGB values. The disjoint base-value
partition puts the flux-bearing faint ranges (`[1e-3,0.06)`, `[0.06,0.1)`,
`[0.1,0.23)`, together ~`47%` of the linear-RGB sum) at ~`2-5%` local relative
RMS with signed bias at or below `0.19%`. This pattern is consistent with
partial cancellation among many faint PSFs, while resolved bright cores retain
bounded sub-pixel displacement. It does not make fast mode pixelwise or
visually lossless.
These measurements characterize independent still frames only. With nearest
deposition, a smoothly moving image jumps by one complete grid spacing when it
crosses a cell boundary: `1/N` output pixel per axis, or `0.5 px` at `N=2` and
`0.25 px` at `N=4`. Such discontinuities can appear as visible flicker even at
4K. Neither nearest mode is therefore qualified as a temporally coherent movie
output path. Bilinear deposition removes this particular centroid jump but has
the profile-broadening tradeoff above and has not been qualified as a final
movie path either. For single-frame dense-star previews, `N=2` remains the
throughput default and `N=4` the higher-fidelity point; the exact per-event
reference PSF path remains the quantitative and movie reference. The
single-frame error decomposition is reproducible with
[`scripts/fast_mode_error_decomposition.py`](scripts/fast_mode_error_decomposition.py).
In CPU builds the global convolution is a zero-padded FFTW linear convolution
(`fftw3`/`fftw3_omp`, double precision). The immutable kernel spectrum, FFTW
plans, and scratch buffers are built once when the accumulator is initialized