Files
GR-raytracing/benchmarks/fast_mode_fftw_2026-09-25.md
wyj e4d093a9fe Doc: Record 4K fast-mode accuracy limits and error distribution
Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README.

- Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2.
- Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile.
- Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
2026-09-26 01:02:41 -04:00

26 KiB

Fast-mode FFTW convolution benchmark (2026-09-25)

Replaces the spatial fast-mode global convolution with a reusable FFTW linear convolution. This record separates one-time plan/setup from steady-state execution and repeats the bounded Galactic-center comparison. Raw resolver and end-to-end output is pasted verbatim from the runs below.

The later production-density 4K comparison is recorded separately against the first committed FFTW implementation, revision 229f50cd864d9768324042c10edbb56bfc120f08.

Environment

  • Git revision: 3deebfb plus the uncommitted FFTW work.
  • Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB.
  • FFTW: pkg-config --modversion fftw3_omp = 3.3.10; pkg-config --libs fftw3_omp = -lfftw3_omp -lfftw3.
  • Build: make -j8 BUILD_TYPE=Release PSF_BACKEND=cpu SPACETIME=minkowski backend (-O2 -DNDEBUG -fopenmp -march=native, libpng, -DFAST_PSF_FFTW).
  • Binary SHA-256 (this run):
    • build/Release/minkowski_sky: ead2a543b8166abe6f82bc1cd41b2e985a6db517aa45a6416d08d188138b8c47
    • build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw: 8da8bbd532104a0c4f524aabf40b66a50c8c715e44479c7db467669ab5a21215
  • Timing wrapper: /usr/bin/time is not installed on this host. The reproducible wrapper used for the end-to-end runs is zsh's built-in time with TIMEFMT='real=%*E user=%U sys=%S cpu=%P', i.e. each command was run as TIMEFMT='real=%*E user=%U sys=%S cpu=%P'; time <command>. The real=... lines below are the verbatim wrapper output.

Correctness

make -j8 BUILD_TYPE=Debug PSF_BACKEND=cpu SPACETIME=minkowski test

All C tests pass, including the dedicated FFTW-vs-spatial test (258 comparisons over sizes 17x13/16x16/23x31/33x17, N = 1..4, nearest and bilinear, eight scenes; worst max_abs = 1.8e-14), the test_frame fast-mode regressions, the non-fast FITS references, and both camera CLI checks. HIP and dummy renderers compile and link without FFTW; make PSF_BACKEND=dummy test now fails fast with make test requires PSF_BACKEND=cpu instead of reusing a stale CPU test binary.

Resolver-only benchmark

One deterministic sparse impulse buffer, default PSF (FWHM 2.7, beta 4.5, relative tail 1e-8), OMP_NUM_THREADS=16. --spatial additionally runs the reference resolver on the identical buffer. Binary: build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw.

Default 320x240 N=2, FFTW_ESTIMATE

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 320 --height 240 --supersample 2 --repeats 7 --spatial
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.039444 s fftw_plan=0.008085 s kernel_fft=0.005237 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.008085 s, kernel_fft=0.005237 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.005602 s median=0.006455 s max=0.008144 s
fftw samples: 0.006697 0.005784 0.005602 0.005898 0.006455 0.007675 0.008144
fftw_hdr_hash=9ad60d92992c9bbd
spatial: 0.702470 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=7.49e-16 peak=0.359 speedup=108.831x
rss_max_kb=52908

Default 320x240 N=4, FFTW_ESTIMATE

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 320 --height 240 --supersample 4 --repeats 7 --spatial
impulse: 320x240 supersample=4 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.114719 s fftw_plan=0.007169 s kernel_fft=0.011942 s
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007169 s, kernel_fft=0.011942 s, scratch=120.7 MiB
impulse_hash=2639864ea9d3a26a
fftw: min=0.034756 s median=0.039591 s max=0.042474 s
fftw samples: 0.041352 0.039591 0.034756 0.040866 0.042474 0.035364 0.036241
fftw_hdr_hash=1ff070fdaac85f4a
spatial: 11.054385 s
spatial_hdr_hash=7682ee8550f88efe
difference: max_abs=5.38e-15 peak=0.533 speedup=279.214x
rss_max_kb=161032

Medium 960x540 N=2, FFTW_ESTIMATE (FFTW only)

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 960 --height 540 --supersample 2 --repeats 7
impulse: 960x540 supersample=2 deposit=nearest repeats=7 spatial=0
plan_mode=estimate
setup: init=0.042270 s fftw_plan=0.006317 s kernel_fft=0.010302 s
Fast FFTW: linear min=1266x2106, fft=1280x2160, workers=16, plan=estimate, plan=0.006317 s, kernel_fft=0.010302 s, scratch=147.7 MiB
impulse_hash=6e2de17f6859cd33
fftw: min=0.038361 s median=0.039872 s max=0.055834 s
fftw samples: 0.055834 0.038745 0.039872 0.038823 0.038361 0.048436 0.040141
fftw_hdr_hash=c493dbb793d04706
rss_max_kb=216272

4K 3840x2160 N=2, FFTW_ESTIMATE (FFTW only)

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 3840 --height 2160 --supersample 2 --repeats 5
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.524840 s fftw_plan=0.008154 s kernel_fft=0.482906 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=estimate, plan=0.008154 s, kernel_fft=0.482906 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=2.368634 s median=2.481927 s max=2.553895 s
fftw samples: 2.552212 2.553895 2.481927 2.389895 2.368634
fftw_hdr_hash=1c5720f72a040e6d
rss_max_kb=2879596

4K 3840x2160 N=4, FFTW_ESTIMATE (FFTW only)

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 3840 --height 2160 --supersample 4 --repeats 5
impulse: 3840x2160 supersample=4 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.738771 s fftw_plan=0.003636 s kernel_fft=0.599235 s
Fast FFTW: linear min=9010x15730, fft=9072x15750, workers=16, plan=estimate, plan=0.003636 s, kernel_fft=0.599235 s, scratch=7631.4 MiB
impulse_hash=e6a528db0fb16c33
fftw: min=3.276609 s median=3.339285 s max=3.369411 s
fftw samples: 3.309935 3.369411 3.276609 3.345614 3.339285
fftw_hdr_hash=e55bed24ef3f5506
rss_max_kb=10919208

FFTW_ESTIMATE versus FFTW_MEASURE

FFTW_MEASURE cuts steady-state time but costs a large one-time plan. Because the plan is created once per process and reused across frames, it only pays off for long movie renders; production therefore defaults to FFTW_ESTIMATE. Wisdom import/export would make FFTW_MEASURE the better default and remains a follow-up. Measured FFTW_MEASURE plan output is not bit-reproducible (plan measurement is timing-dependent), so its fftw_hdr_hash may differ run to run; the spatial comparison stays at roundoff.

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 320 --height 240 --supersample 2 --repeats 7 --measure --spatial
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=measure
setup: init=1.584388 s fftw_plan=1.091687 s kernel_fft=0.467405 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=measure, plan=1.091687 s, kernel_fft=0.467405 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.003995 s median=0.004702 s max=0.005804 s
fftw samples: 0.003995 0.004702 0.004330 0.005804 0.004939 0.004986 0.004619
fftw_hdr_hash=626e386cba92ffa5
spatial: 0.706244 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=8.05e-16 peak=0.359 speedup=150.198x
rss_max_kb=51860
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 3840 --height 2160 --supersample 2 --repeats 5 --measure
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=measure
setup: init=41.408673 s fftw_plan=41.239987 s kernel_fft=0.134272 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=measure, plan=41.239987 s, kernel_fft=0.134272 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=0.518707 s median=0.522181 s max=0.539465 s
fftw samples: 0.523380 0.539465 0.518707 0.520357 0.522181
fftw_hdr_hash=ab90c4b076c775e1
rss_max_kb=2884132

End-to-end bounded comparison

Same catalog, geometry, exposure, supersample factors, and worker environment as benchmarks/fast_mode_cpu_2026-09-18.md. Raw terminal output including the timing wrapper:

OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
  time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --output /tmp/opencode/fastmode_fftw_bench/normal.png
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000046 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.776 s
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.729 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/normal.png (ok)
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
real=17.526 user=225.46s sys=2.44s cpu=1300%
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
  time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 2 \
  --output /tmp/opencode/fastmode_fftw_bench/fast_n2.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000049 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.007770 s, kernel_fft=0.005020 s, scratch=30.2 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.735 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=0.995 user=5.30s sys=0.18s cpu=550%
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
  time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 4 \
  --output /tmp/opencode/fastmode_fftw_bench/fast_n4.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000044 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007294 s, kernel_fft=0.012307 s, scratch=120.7 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.722 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n4.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=1.033 user=5.56s sys=0.22s cpu=559%

Image counts are identical across normal and both fast runs (5,912,840), so the FFTW resolver did not change catalog/deposit behavior. Output SHA-256:

normal.png   e9f0a98ae7a3c3db546a3f276f400be98c96446fab7bfbc632f388702d4722d1
fast_n2.png  1f144aefb4a5520d9d76a13ab69026a72580663e5e45f53128a87a0801176ccf
fast_n4.png  679f14a67a085f76c144ecbe36af51b3b2534c152e8ef59708a7884a3f7c8464

Production-density 4K Galactic-center comparison

This is the full 3840x2160, 50-degree-field run over the local processed 2MASS all-sky catalog, not the bounded 320x240 comparison above.

Provenance

  • Git revision: 229f50cd864d9768324042c10edbb56bfc120f08 (Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve).
  • The tracked worktree matched that revision when this record was written; the unrelated untracked nmesh_spacetime_output_format.md was not an input.
  • Executable: build/Release/minkowski_sky, built at 2026-09-25 23:16:45 -0400, before the three runs below; size 146,856 bytes; SHA-256 7f6486285b2ff498adb2d4f1acaf62dc96b11125c9cec114f572b9379648962c. A Make dry-run after the commit requested only the Makefile's unconditional FORCE relink and no object recompilation, so the recorded binary used the object set corresponding to the committed tracked sources.
  • Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB. The commands did not explicitly set OMP_NUM_THREADS; fast-mode initialization reported workers=16.
  • Input: assets/2mass/processed/all_sky, observed when recording as 64,801 regular files and 16,207,861,727 bytes. The renderer selected 1,742 tiles and loaded 54,509,992 stars in each run. No content manifest hash was recorded for this external processed catalog.
  • Timing: zsh built-in time; its final timing line is included verbatim for each command.

Standard cached spatial PSF

time ./build/Release/minkowski_sky \
  --psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
  --all-sky-catalog assets/2mass/processed/all_sky \
  --width 3840 --height 2160 \
  --look-ra-deg 262.5 --look-dec-deg -30 \
  --fov-deg 50 \
  --coarse-cell-pixels 16 \
  --exposure 1e13 \
  --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
  --catalog-load-workers 4 \
  --hdr-output \
  --output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000130 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 85 px, relative tail 1e-06, tail abs 1e-06, boundary 1e-07, build 2.176 s
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.886 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png (ok)
PSF splats: cached 52477581, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 66954
Warning: --psf-min-y discarded one or more PSF events.
./build/Release/minkowski_sky --psf-relative-tail 1e-6 --psf-min-y 1e-8  1e8   938.60s user 15.80s system 1391% cpu 1:08.61 total

Fast mode, N=2

time ./build/Release/minkowski_sky --fast-mode --fast-supersample 2 \
  --psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
  --all-sky-catalog assets/2mass/processed/all_sky \
  --width 3840 --height 2160 \
  --look-ra-deg 262.5 --look-dec-deg -30 \
  --fov-deg 50 \
  --coarse-cell-pixels 16 \
  --exposure 1e13 \
  --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
  --catalog-load-workers 4 \
  --hdr-output \
  --output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000101 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 169 ss px (84.50 final px), retained flux 0.999999, cached 4-point quadrature
Fast FFTW: linear min=4658x8018, fft=4704x8064, workers=16, plan=estimate, plan=0.005275 s, kernel_fft=0.141445 s, scratch=2026.1 MiB
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.873 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png (ok)
Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954
./build/Release/minkowski_sky --fast-mode --fast-supersample 2  1e-6  1e-8     115.74s user 2.08s system 763% cpu 15.439 total

Fast mode, N=4

time ./build/Release/minkowski_sky --fast-mode --fast-supersample 4 \
  --psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
  --all-sky-catalog assets/2mass/processed/all_sky \
  --width 3840 --height 2160 \
  --look-ra-deg 262.5 --look-dec-deg -30 \
  --fov-deg 50 \
  --coarse-cell-pixels 16 \
  --exposure 1e13 \
  --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
  --catalog-load-workers 4 \
  --hdr-output \
  --output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000043 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 336 ss px (84.00 final px), retained flux 0.999999, cached 4-point quadrature
Fast FFTW: linear min=9312x16032, fft=9375x16128, workers=16, plan=estimate, plan=0.005384 s, kernel_fft=0.743673 s, scratch=8075.5 MiB
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.796 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png (ok)
Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954
./build/Release/minkowski_sky --fast-mode --fast-supersample 4  1e-6  1e-8     135.62s user 5.04s system 722% cpu 19.469 total

Artifacts and comparison

05517bde1383327043691ac4b713f638b03ebf6dc72207fe8710a8dce747a03e  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png
86e9220feac277d77eba87aa8d0febadca5f4eea35cb56164a4eb1578b20fccd  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png
5dfef3b4bd39d9809463594a1065c32c65ad397396514664c8aa22a5da6d780b  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png
1ea508d1ee787c2e5f35ccdf37ca5f4a58e9ff5e9984b50928d66ecf88c44c3d  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits
71de3f1e11cb9f811f5d08ea45b51baf5968e8523873873d6f7c414c1849ba53  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits
554848e39fdfb3b4815fb8f2a591b8d356cbdeaaea235a73f19b46017a18e42a  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits

All three runs produced 52,544,535 images and applied the same event filtering: 52,477,581 cached/deposited events, no wing clipping or direct fallback, and 66,954 events (0.12742334% of rendered images) discarded by --psf-min-y.

Using the standard HDR buffer as the reference, with relative_L2 = ||fast - base||_2 / ||base||_2, the existing FITS artifacts give:

Mode Wall time Wall speedup Approx. speedup after subtracting catalog prefetch User-CPU speedup Scratch HDR relative L2 HDR total-flux ratio Tone-mapped PNG normalized RMSE
Standard 68.610 s 1.000x 1.000x 1.000x n/a 0 1 0
Fast N=2 15.439 s 4.444x 6.557x 8.110x 2026.1 MiB 10.4384% 1.000176588 0.00678916
Fast N=4 19.469 s 3.524x 4.587x 6.921x 8075.5 MiB 4.37249% 1.000178510 0.00365798

The post-prefetch figures subtract only the renderer-reported catalog prefetch duration from wall time; they are useful context, not isolated PSF timings. N=4 costs 26.1% more wall time and 3.986x the scratch of N=2, while reducing the HDR relative-L2 difference by about 58% and the tone-mapped PNG RMSE by about 46%. The nearly identical +0.0177% total-flux offset for N=2 and N=4 shows that the larger N primarily improves spatial placement rather than global photometry. For the intended rapid single-frame exposure/PSF preview, N=2 is therefore the default throughput point; N=4 is the higher-fidelity preview option.

Error distribution across the image

The unweighted relative L2 above is dominated by a small number of very bright PSF cores, so it understates how well the dense faint-star texture is reproduced in this still frame. scripts/fast_mode_error_decomposition.py reports the error over the flattened scalar RGB channel samples:

python3 scripts/fast_mode_error_decomposition.py \
  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits \
  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits \
  output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits
===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits =====  relative_L2=10.4384%
top |d| RGB samples:  sq-error energy   base energy   linear-RGB sum   sample share
  top   0.0010%:  69.136%     48.948%     2.822%     0.0010%
  top   0.0100%:  95.732%     87.994%     9.141%     0.0100%
  top   0.1000%:  99.365%     97.633%    16.396%     0.1000%
  top   1.0000%:  99.893%     99.224%    25.984%     1.0000%
  top  10.0000%:  99.988%     99.679%    46.006%    10.0000%
base-value range (disjoint):  rel_RMS   signed bias   base energy %   linear-RGB sum %   samples
  [1e-06,0.001):    4.7788%     +0.7665%     0.000%    0.001%  67026
  [0.001,0.0599):    4.5658%     +0.1902%     0.013%   11.063%  12374574
  [0.0599,0.1):    4.6294%     +0.1297%     0.027%   10.881%  4486724
  [0.1,0.229):    4.9079%     +0.0866%     0.124%   25.365%  5466556
  [0.229,0.769):    6.9340%     +0.0100%     0.301%   24.638%  2239488
  [0.769,4.62):   11.4421%     -0.0990%     0.615%   10.315%  223948
  [4.62,inf):   10.4475%     -0.1789%    98.921%   17.738%  24884

===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits =====  relative_L2=4.3725%
top |d| RGB samples:  sq-error energy   base energy   linear-RGB sum   sample share
  top   0.0010%:  65.911%     35.557%     2.418%     0.0010%
  top   0.0100%:  94.540%     86.397%     8.879%     0.0100%
  top   0.1000%:  99.108%     96.708%    16.179%     0.1000%
  top   1.0000%:  99.848%     99.208%    25.909%     1.0000%
  top  10.0000%:  99.982%     99.694%    45.966%    10.0000%
base-value range (disjoint):  rel_RMS   signed bias   base energy %   linear-RGB sum %   samples
  [1e-06,0.001):    2.4187%     +0.5689%     0.000%    0.001%  67026
  [0.001,0.0599):    2.2772%     +0.0729%     0.013%   11.063%  12374574
  [0.0599,0.1):    2.3129%     +0.0550%     0.027%   10.881%  4486724
  [0.1,0.229):    2.4542%     +0.0393%     0.124%   25.365%  5466556
  [0.229,0.769):    3.4700%     +0.0166%     0.301%   24.638%  2239488
  [0.769,4.62):    5.6752%     -0.0189%     0.615%   10.315%  223948
  [4.62,inf):    4.3682%     -0.0469%    98.921%   17.738%  24884

The FITS RGB channels are flattened into scalar samples: sq-error energy is the fraction of ||fast-base||^2, linear-RGB sum is the sum of baseline linear-RGB sample values (proportional to total flux, not a per-source photometry), and the base-value ranges partition all samples and sum to 100%.

The top 0.01% of RGB channel samples account for 95.7% (N=2) or 94.5% (N=4) of squared-error energy and about 9% of the baseline linear-RGB sum. The flux-bearing faint ranges ([1e-3,0.0599), [0.0599,0.1), [0.1,0.229), together ~47% of the linear-RGB sum) have 2.3-4.9% local relative RMS and signed bias of +0.04% to +0.19%; the many faint PSFs partially cancel. This is consistent with the concentrated bright-core difference being predominantly a bounded sub-pixel redistribution rather than a large net linear-RGB loss, and explains why this tone-mapped still PNG remains smooth despite the global relative L2.

This comparison contains one independent frame only. It does not measure temporal coherence: nearest deposition jumps by 1/N output pixel per axis when a moving image crosses a supersampled-cell boundary (0.5 px at N=2, 0.25 px at N=4). Both can produce visible movie flicker and are not qualified here as visually faithful movie-output paths.

Reading

  • Steady-state FFTW is ~109x (N=2) and ~279x (N=4) faster than the spatial reference at 320x240, and stays under ~0.06 s per frame at 960x540.
  • 4K N=2 resolves in ~2.5 s (ESTIMATE) or ~0.52 s (MEASURE) on this host; 4K N=4 needs ~7.6 GiB scratch and ~3.3 s (ESTIMATE).
  • FFTW-versus-spatial differences are floating-point roundoff (max_abs at the 1e-15 level against peaks near 0.5), not crop, wrap, channel, or normalization errors.
  • At the committed revision, the production-density 4K Galactic-center run completes in 15.439 s at N=2 versus 68.610 s for the standard path (4.444x end-to-end), with identical catalog image/deposit/filter counts. N=4 trades 26.1% more wall time for materially lower spatial error.
  • The 4K still-frame error is concentrated: the top 0.01% of RGB channel samples hold ~95% of the squared-error energy but only ~9% of the baseline linear-RGB sum, while the flux-bearing faint-value ranges stay at ~2-5% local relative RMS with signed bias at or below ~0.2%. This does not establish temporal fidelity.