Files
GR-raytracing/benchmarks/fast_mode_fftw_2026-09-25.md
T
wyj e4d093a9fe Doc: Record 4K fast-mode accuracy limits and error distribution
Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README.

- Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2.
- Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile.
- Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
2026-09-26 01:02:41 -04:00

510 lines
26 KiB
Markdown

# Fast-mode FFTW convolution benchmark (2026-09-25)
Replaces the spatial fast-mode global convolution with a reusable FFTW linear
convolution. This record separates one-time plan/setup from steady-state
execution and repeats the bounded Galactic-center comparison. Raw resolver and
end-to-end output is pasted verbatim from the runs below.
The later production-density 4K comparison is recorded separately against the
first committed FFTW implementation, revision
`229f50cd864d9768324042c10edbb56bfc120f08`.
## Environment
- Git revision: `3deebfb` plus the uncommitted FFTW work.
- Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB.
- FFTW: `pkg-config --modversion fftw3_omp` = `3.3.10`;
`pkg-config --libs fftw3_omp` = `-lfftw3_omp -lfftw3`.
- Build:
`make -j8 BUILD_TYPE=Release PSF_BACKEND=cpu SPACETIME=minkowski backend`
(`-O2 -DNDEBUG -fopenmp -march=native`, libpng, `-DFAST_PSF_FFTW`).
- Binary SHA-256 (this run):
- `build/Release/minkowski_sky`:
`ead2a543b8166abe6f82bc1cd41b2e985a6db517aa45a6416d08d188138b8c47`
- `build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw`:
`8da8bbd532104a0c4f524aabf40b66a50c8c715e44479c7db467669ab5a21215`
- Timing wrapper: `/usr/bin/time` is not installed on this host. The
reproducible wrapper used for the end-to-end runs is zsh's built-in `time`
with `TIMEFMT='real=%*E user=%U sys=%S cpu=%P'`, i.e. each command was run as
`TIMEFMT='real=%*E user=%U sys=%S cpu=%P'; time <command>`. The `real=...`
lines below are the verbatim wrapper output.
## Correctness
```sh
make -j8 BUILD_TYPE=Debug PSF_BACKEND=cpu SPACETIME=minkowski test
```
All C tests pass, including the dedicated FFTW-vs-spatial test (258 comparisons
over sizes `17x13`/`16x16`/`23x31`/`33x17`, `N = 1..4`, nearest and bilinear,
eight scenes; worst `max_abs = 1.8e-14`), the `test_frame` fast-mode
regressions, the non-fast FITS references, and both camera CLI checks. HIP and
dummy renderers compile and link without FFTW; `make PSF_BACKEND=dummy test`
now fails fast with `make test requires PSF_BACKEND=cpu` instead of reusing a
stale CPU test binary.
## Resolver-only benchmark
One deterministic sparse impulse buffer, default PSF (FWHM 2.7, beta 4.5,
relative tail 1e-8), `OMP_NUM_THREADS=16`. `--spatial` additionally runs the
reference resolver on the identical buffer. Binary:
`build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw`.
### Default 320x240 N=2, FFTW_ESTIMATE
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 320 --height 240 --supersample 2 --repeats 7 --spatial
```
```text
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.039444 s fftw_plan=0.008085 s kernel_fft=0.005237 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.008085 s, kernel_fft=0.005237 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.005602 s median=0.006455 s max=0.008144 s
fftw samples: 0.006697 0.005784 0.005602 0.005898 0.006455 0.007675 0.008144
fftw_hdr_hash=9ad60d92992c9bbd
spatial: 0.702470 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=7.49e-16 peak=0.359 speedup=108.831x
rss_max_kb=52908
```
### Default 320x240 N=4, FFTW_ESTIMATE
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 320 --height 240 --supersample 4 --repeats 7 --spatial
```
```text
impulse: 320x240 supersample=4 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.114719 s fftw_plan=0.007169 s kernel_fft=0.011942 s
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007169 s, kernel_fft=0.011942 s, scratch=120.7 MiB
impulse_hash=2639864ea9d3a26a
fftw: min=0.034756 s median=0.039591 s max=0.042474 s
fftw samples: 0.041352 0.039591 0.034756 0.040866 0.042474 0.035364 0.036241
fftw_hdr_hash=1ff070fdaac85f4a
spatial: 11.054385 s
spatial_hdr_hash=7682ee8550f88efe
difference: max_abs=5.38e-15 peak=0.533 speedup=279.214x
rss_max_kb=161032
```
### Medium 960x540 N=2, FFTW_ESTIMATE (FFTW only)
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 960 --height 540 --supersample 2 --repeats 7
```
```text
impulse: 960x540 supersample=2 deposit=nearest repeats=7 spatial=0
plan_mode=estimate
setup: init=0.042270 s fftw_plan=0.006317 s kernel_fft=0.010302 s
Fast FFTW: linear min=1266x2106, fft=1280x2160, workers=16, plan=estimate, plan=0.006317 s, kernel_fft=0.010302 s, scratch=147.7 MiB
impulse_hash=6e2de17f6859cd33
fftw: min=0.038361 s median=0.039872 s max=0.055834 s
fftw samples: 0.055834 0.038745 0.039872 0.038823 0.038361 0.048436 0.040141
fftw_hdr_hash=c493dbb793d04706
rss_max_kb=216272
```
### 4K 3840x2160 N=2, FFTW_ESTIMATE (FFTW only)
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 3840 --height 2160 --supersample 2 --repeats 5
```
```text
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.524840 s fftw_plan=0.008154 s kernel_fft=0.482906 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=estimate, plan=0.008154 s, kernel_fft=0.482906 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=2.368634 s median=2.481927 s max=2.553895 s
fftw samples: 2.552212 2.553895 2.481927 2.389895 2.368634
fftw_hdr_hash=1c5720f72a040e6d
rss_max_kb=2879596
```
### 4K 3840x2160 N=4, FFTW_ESTIMATE (FFTW only)
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 3840 --height 2160 --supersample 4 --repeats 5
```
```text
impulse: 3840x2160 supersample=4 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.738771 s fftw_plan=0.003636 s kernel_fft=0.599235 s
Fast FFTW: linear min=9010x15730, fft=9072x15750, workers=16, plan=estimate, plan=0.003636 s, kernel_fft=0.599235 s, scratch=7631.4 MiB
impulse_hash=e6a528db0fb16c33
fftw: min=3.276609 s median=3.339285 s max=3.369411 s
fftw samples: 3.309935 3.369411 3.276609 3.345614 3.339285
fftw_hdr_hash=e55bed24ef3f5506
rss_max_kb=10919208
```
### FFTW_ESTIMATE versus FFTW_MEASURE
`FFTW_MEASURE` cuts steady-state time but costs a large one-time plan. Because
the plan is created once per process and reused across frames, it only pays off
for long movie renders; production therefore defaults to `FFTW_ESTIMATE`.
Wisdom import/export would make `FFTW_MEASURE` the better default and remains a
follow-up. Measured FFTW_MEASURE plan output is not bit-reproducible (plan
measurement is timing-dependent), so its `fftw_hdr_hash` may differ run to run;
the spatial comparison stays at roundoff.
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 320 --height 240 --supersample 2 --repeats 7 --measure --spatial
```
```text
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=measure
setup: init=1.584388 s fftw_plan=1.091687 s kernel_fft=0.467405 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=measure, plan=1.091687 s, kernel_fft=0.467405 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.003995 s median=0.004702 s max=0.005804 s
fftw samples: 0.003995 0.004702 0.004330 0.005804 0.004939 0.004986 0.004619
fftw_hdr_hash=626e386cba92ffa5
spatial: 0.706244 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=8.05e-16 peak=0.359 speedup=150.198x
rss_max_kb=51860
```
```sh
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 3840 --height 2160 --supersample 2 --repeats 5 --measure
```
```text
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=measure
setup: init=41.408673 s fftw_plan=41.239987 s kernel_fft=0.134272 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=measure, plan=41.239987 s, kernel_fft=0.134272 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=0.518707 s median=0.522181 s max=0.539465 s
fftw samples: 0.523380 0.539465 0.518707 0.520357 0.522181
fftw_hdr_hash=ab90c4b076c775e1
rss_max_kb=2884132
```
## End-to-end bounded comparison
Same catalog, geometry, exposure, supersample factors, and worker environment as
`benchmarks/fast_mode_cpu_2026-09-18.md`. Raw terminal output including the
timing wrapper:
```sh
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--output /tmp/opencode/fastmode_fftw_bench/normal.png
```
```text
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000046 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.776 s
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.729 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/normal.png (ok)
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
real=17.526 user=225.46s sys=2.44s cpu=1300%
```
```sh
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 2 \
--output /tmp/opencode/fastmode_fftw_bench/fast_n2.png
```
```text
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000049 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.007770 s, kernel_fft=0.005020 s, scratch=30.2 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.735 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=0.995 user=5.30s sys=0.18s cpu=550%
```
```sh
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 4 \
--output /tmp/opencode/fastmode_fftw_bench/fast_n4.png
```
```text
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000044 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007294 s, kernel_fft=0.012307 s, scratch=120.7 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.722 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n4.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=1.033 user=5.56s sys=0.22s cpu=559%
```
Image counts are identical across normal and both fast runs (5,912,840), so the
FFTW resolver did not change catalog/deposit behavior. Output SHA-256:
```text
normal.png e9f0a98ae7a3c3db546a3f276f400be98c96446fab7bfbc632f388702d4722d1
fast_n2.png 1f144aefb4a5520d9d76a13ab69026a72580663e5e45f53128a87a0801176ccf
fast_n4.png 679f14a67a085f76c144ecbe36af51b3b2534c152e8ef59708a7884a3f7c8464
```
## Production-density 4K Galactic-center comparison
This is the full 3840x2160, 50-degree-field run over the local processed 2MASS
all-sky catalog, not the bounded 320x240 comparison above.
### Provenance
- Git revision: `229f50cd864d9768324042c10edbb56bfc120f08`
(`Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve`).
- The tracked worktree matched that revision when this record was written; the
unrelated untracked `nmesh_spacetime_output_format.md` was not an input.
- Executable: `build/Release/minkowski_sky`, built at
`2026-09-25 23:16:45 -0400`, before the three runs below; size 146,856 bytes;
SHA-256
`7f6486285b2ff498adb2d4f1acaf62dc96b11125c9cec114f572b9379648962c`.
A Make dry-run after the commit requested only the Makefile's unconditional
`FORCE` relink and no object recompilation, so the recorded binary used the
object set corresponding to the committed tracked sources.
- Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB. The commands
did not explicitly set `OMP_NUM_THREADS`; fast-mode initialization reported
`workers=16`.
- Input: `assets/2mass/processed/all_sky`, observed when recording as 64,801
regular files and 16,207,861,727 bytes. The renderer selected 1,742 tiles and
loaded 54,509,992 stars in each run. No content manifest hash was recorded for
this external processed catalog.
- Timing: zsh built-in `time`; its final timing line is included verbatim for
each command.
### Standard cached spatial PSF
```sh
time ./build/Release/minkowski_sky \
--psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 3840 --height 2160 \
--look-ra-deg 262.5 --look-dec-deg -30 \
--fov-deg 50 \
--coarse-cell-pixels 16 \
--exposure 1e13 \
--psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
--catalog-load-workers 4 \
--hdr-output \
--output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png
```
```text
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000130 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 85 px, relative tail 1e-06, tail abs 1e-06, boundary 1e-07, build 2.176 s
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.886 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png (ok)
PSF splats: cached 52477581, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 66954
Warning: --psf-min-y discarded one or more PSF events.
./build/Release/minkowski_sky --psf-relative-tail 1e-6 --psf-min-y 1e-8 1e8 938.60s user 15.80s system 1391% cpu 1:08.61 total
```
### Fast mode, N=2
```sh
time ./build/Release/minkowski_sky --fast-mode --fast-supersample 2 \
--psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 3840 --height 2160 \
--look-ra-deg 262.5 --look-dec-deg -30 \
--fov-deg 50 \
--coarse-cell-pixels 16 \
--exposure 1e13 \
--psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
--catalog-load-workers 4 \
--hdr-output \
--output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png
```
```text
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000101 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 169 ss px (84.50 final px), retained flux 0.999999, cached 4-point quadrature
Fast FFTW: linear min=4658x8018, fft=4704x8064, workers=16, plan=estimate, plan=0.005275 s, kernel_fft=0.141445 s, scratch=2026.1 MiB
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.873 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png (ok)
Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954
./build/Release/minkowski_sky --fast-mode --fast-supersample 2 1e-6 1e-8 115.74s user 2.08s system 763% cpu 15.439 total
```
### Fast mode, N=4
```sh
time ./build/Release/minkowski_sky --fast-mode --fast-supersample 4 \
--psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 3840 --height 2160 \
--look-ra-deg 262.5 --look-dec-deg -30 \
--fov-deg 50 \
--coarse-cell-pixels 16 \
--exposure 1e13 \
--psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
--catalog-load-workers 4 \
--hdr-output \
--output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png
```
```text
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000043 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 336 ss px (84.00 final px), retained flux 0.999999, cached 4-point quadrature
Fast FFTW: linear min=9312x16032, fft=9375x16128, workers=16, plan=estimate, plan=0.005384 s, kernel_fft=0.743673 s, scratch=8075.5 MiB
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.796 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png (ok)
Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954
./build/Release/minkowski_sky --fast-mode --fast-supersample 4 1e-6 1e-8 135.62s user 5.04s system 722% cpu 19.469 total
```
### Artifacts and comparison
```text
05517bde1383327043691ac4b713f638b03ebf6dc72207fe8710a8dce747a03e output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png
86e9220feac277d77eba87aa8d0febadca5f4eea35cb56164a4eb1578b20fccd output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png
5dfef3b4bd39d9809463594a1065c32c65ad397396514664c8aa22a5da6d780b output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png
1ea508d1ee787c2e5f35ccdf37ca5f4a58e9ff5e9984b50928d66ecf88c44c3d output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits
71de3f1e11cb9f811f5d08ea45b51baf5968e8523873873d6f7c414c1849ba53 output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits
554848e39fdfb3b4815fb8f2a591b8d356cbdeaaea235a73f19b46017a18e42a output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits
```
All three runs produced 52,544,535 images and applied the same event filtering:
52,477,581 cached/deposited events, no wing clipping or direct fallback, and
66,954 events (0.12742334% of rendered images) discarded by `--psf-min-y`.
Using the standard HDR buffer as the reference, with
`relative_L2 = ||fast - base||_2 / ||base||_2`, the existing FITS artifacts give:
| Mode | Wall time | Wall speedup | Approx. speedup after subtracting catalog prefetch | User-CPU speedup | Scratch | HDR relative L2 | HDR total-flux ratio | Tone-mapped PNG normalized RMSE |
|---|---:|---:|---:|---:|---:|---:|---:|---:|
| Standard | 68.610 s | 1.000x | 1.000x | 1.000x | n/a | 0 | 1 | 0 |
| Fast N=2 | 15.439 s | 4.444x | 6.557x | 8.110x | 2026.1 MiB | 10.4384% | 1.000176588 | 0.00678916 |
| Fast N=4 | 19.469 s | 3.524x | 4.587x | 6.921x | 8075.5 MiB | 4.37249% | 1.000178510 | 0.00365798 |
The post-prefetch figures subtract only the renderer-reported catalog prefetch
duration from wall time; they are useful context, not isolated PSF timings. N=4
costs 26.1% more wall time and 3.986x the scratch of N=2, while reducing the HDR
relative-L2 difference by about 58% and the tone-mapped PNG RMSE by about 46%.
The nearly identical `+0.0177%` total-flux offset for N=2 and N=4 shows that the
larger N primarily improves spatial placement rather than global photometry.
For the intended rapid single-frame exposure/PSF preview, N=2 is therefore the
default throughput point; N=4 is the higher-fidelity preview option.
### Error distribution across the image
The unweighted relative L2 above is dominated by a small number of very bright
PSF cores, so it understates how well the dense faint-star texture is
reproduced in this still frame. `scripts/fast_mode_error_decomposition.py`
reports the error over the flattened scalar RGB channel samples:
```sh
python3 scripts/fast_mode_error_decomposition.py \
output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits \
output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits \
output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits
```
```text
===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits ===== relative_L2=10.4384%
top |d| RGB samples: sq-error energy base energy linear-RGB sum sample share
top 0.0010%: 69.136% 48.948% 2.822% 0.0010%
top 0.0100%: 95.732% 87.994% 9.141% 0.0100%
top 0.1000%: 99.365% 97.633% 16.396% 0.1000%
top 1.0000%: 99.893% 99.224% 25.984% 1.0000%
top 10.0000%: 99.988% 99.679% 46.006% 10.0000%
base-value range (disjoint): rel_RMS signed bias base energy % linear-RGB sum % samples
[1e-06,0.001): 4.7788% +0.7665% 0.000% 0.001% 67026
[0.001,0.0599): 4.5658% +0.1902% 0.013% 11.063% 12374574
[0.0599,0.1): 4.6294% +0.1297% 0.027% 10.881% 4486724
[0.1,0.229): 4.9079% +0.0866% 0.124% 25.365% 5466556
[0.229,0.769): 6.9340% +0.0100% 0.301% 24.638% 2239488
[0.769,4.62): 11.4421% -0.0990% 0.615% 10.315% 223948
[4.62,inf): 10.4475% -0.1789% 98.921% 17.738% 24884
===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits ===== relative_L2=4.3725%
top |d| RGB samples: sq-error energy base energy linear-RGB sum sample share
top 0.0010%: 65.911% 35.557% 2.418% 0.0010%
top 0.0100%: 94.540% 86.397% 8.879% 0.0100%
top 0.1000%: 99.108% 96.708% 16.179% 0.1000%
top 1.0000%: 99.848% 99.208% 25.909% 1.0000%
top 10.0000%: 99.982% 99.694% 45.966% 10.0000%
base-value range (disjoint): rel_RMS signed bias base energy % linear-RGB sum % samples
[1e-06,0.001): 2.4187% +0.5689% 0.000% 0.001% 67026
[0.001,0.0599): 2.2772% +0.0729% 0.013% 11.063% 12374574
[0.0599,0.1): 2.3129% +0.0550% 0.027% 10.881% 4486724
[0.1,0.229): 2.4542% +0.0393% 0.124% 25.365% 5466556
[0.229,0.769): 3.4700% +0.0166% 0.301% 24.638% 2239488
[0.769,4.62): 5.6752% -0.0189% 0.615% 10.315% 223948
[4.62,inf): 4.3682% -0.0469% 98.921% 17.738% 24884
```
The FITS RGB channels are flattened into scalar samples: `sq-error energy` is
the fraction of `||fast-base||^2`, `linear-RGB sum` is the sum of baseline
linear-RGB sample values (proportional to total flux, not a per-source
photometry), and the base-value ranges partition all samples and sum to 100%.
The top 0.01% of RGB channel samples account for 95.7% (`N=2`) or 94.5%
(`N=4`) of squared-error energy and about 9% of the baseline linear-RGB sum.
The flux-bearing faint ranges (`[1e-3,0.0599)`, `[0.0599,0.1)`, `[0.1,0.229)`,
together ~47% of the linear-RGB sum) have 2.3-4.9% local relative RMS and
signed bias of `+0.04%` to `+0.19%`; the many faint PSFs partially cancel. This
is consistent with the concentrated bright-core difference being predominantly a
bounded sub-pixel redistribution rather than a large net linear-RGB loss, and
explains why this tone-mapped still PNG remains smooth despite the global
relative L2.
This comparison contains one independent frame only. It does not measure
temporal coherence: nearest deposition jumps by `1/N` output pixel per axis
when a moving image crosses a supersampled-cell boundary (`0.5 px` at `N=2`,
`0.25 px` at `N=4`). Both can produce visible movie flicker and are not
qualified here as visually faithful movie-output paths.
## Reading
- Steady-state FFTW is ~109x (N=2) and ~279x (N=4) faster than the spatial
reference at 320x240, and stays under ~0.06 s per frame at 960x540.
- 4K N=2 resolves in ~2.5 s (ESTIMATE) or ~0.52 s (MEASURE) on this host; 4K N=4
needs ~7.6 GiB scratch and ~3.3 s (ESTIMATE).
- FFTW-versus-spatial differences are floating-point roundoff (`max_abs` at the
`1e-15` level against peaks near `0.5`), not crop, wrap, channel, or
normalization errors.
- At the committed revision, the production-density 4K Galactic-center run
completes in 15.439 s at N=2 versus 68.610 s for the standard path (4.444x
end-to-end), with identical catalog image/deposit/filter counts. N=4 trades
26.1% more wall time for materially lower spatial error.
- The 4K still-frame error is concentrated: the top 0.01% of RGB channel
samples hold ~95% of the squared-error energy but only ~9% of the baseline
linear-RGB sum, while the flux-bearing faint-value ranges stay at ~2-5% local
relative RMS with signed bias at or below ~0.2%. This does not establish
temporal fidelity.