Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README. - Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2. - Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile. - Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
26 KiB
Fast-mode FFTW convolution benchmark (2026-09-25)
Replaces the spatial fast-mode global convolution with a reusable FFTW linear convolution. This record separates one-time plan/setup from steady-state execution and repeats the bounded Galactic-center comparison. Raw resolver and end-to-end output is pasted verbatim from the runs below.
The later production-density 4K comparison is recorded separately against the
first committed FFTW implementation, revision
229f50cd864d9768324042c10edbb56bfc120f08.
Environment
- Git revision:
3deebfbplus the uncommitted FFTW work. - Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB.
- FFTW:
pkg-config --modversion fftw3_omp=3.3.10;pkg-config --libs fftw3_omp=-lfftw3_omp -lfftw3. - Build:
make -j8 BUILD_TYPE=Release PSF_BACKEND=cpu SPACETIME=minkowski backend(-O2 -DNDEBUG -fopenmp -march=native, libpng,-DFAST_PSF_FFTW). - Binary SHA-256 (this run):
build/Release/minkowski_sky:ead2a543b8166abe6f82bc1cd41b2e985a6db517aa45a6416d08d188138b8c47build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw:8da8bbd532104a0c4f524aabf40b66a50c8c715e44479c7db467669ab5a21215
- Timing wrapper:
/usr/bin/timeis not installed on this host. The reproducible wrapper used for the end-to-end runs is zsh's built-intimewithTIMEFMT='real=%*E user=%U sys=%S cpu=%P', i.e. each command was run asTIMEFMT='real=%*E user=%U sys=%S cpu=%P'; time <command>. Thereal=...lines below are the verbatim wrapper output.
Correctness
make -j8 BUILD_TYPE=Debug PSF_BACKEND=cpu SPACETIME=minkowski test
All C tests pass, including the dedicated FFTW-vs-spatial test (258 comparisons
over sizes 17x13/16x16/23x31/33x17, N = 1..4, nearest and bilinear,
eight scenes; worst max_abs = 1.8e-14), the test_frame fast-mode
regressions, the non-fast FITS references, and both camera CLI checks. HIP and
dummy renderers compile and link without FFTW; make PSF_BACKEND=dummy test
now fails fast with make test requires PSF_BACKEND=cpu instead of reusing a
stale CPU test binary.
Resolver-only benchmark
One deterministic sparse impulse buffer, default PSF (FWHM 2.7, beta 4.5,
relative tail 1e-8), OMP_NUM_THREADS=16. --spatial additionally runs the
reference resolver on the identical buffer. Binary:
build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw.
Default 320x240 N=2, FFTW_ESTIMATE
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 320 --height 240 --supersample 2 --repeats 7 --spatial
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.039444 s fftw_plan=0.008085 s kernel_fft=0.005237 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.008085 s, kernel_fft=0.005237 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.005602 s median=0.006455 s max=0.008144 s
fftw samples: 0.006697 0.005784 0.005602 0.005898 0.006455 0.007675 0.008144
fftw_hdr_hash=9ad60d92992c9bbd
spatial: 0.702470 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=7.49e-16 peak=0.359 speedup=108.831x
rss_max_kb=52908
Default 320x240 N=4, FFTW_ESTIMATE
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 320 --height 240 --supersample 4 --repeats 7 --spatial
impulse: 320x240 supersample=4 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.114719 s fftw_plan=0.007169 s kernel_fft=0.011942 s
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007169 s, kernel_fft=0.011942 s, scratch=120.7 MiB
impulse_hash=2639864ea9d3a26a
fftw: min=0.034756 s median=0.039591 s max=0.042474 s
fftw samples: 0.041352 0.039591 0.034756 0.040866 0.042474 0.035364 0.036241
fftw_hdr_hash=1ff070fdaac85f4a
spatial: 11.054385 s
spatial_hdr_hash=7682ee8550f88efe
difference: max_abs=5.38e-15 peak=0.533 speedup=279.214x
rss_max_kb=161032
Medium 960x540 N=2, FFTW_ESTIMATE (FFTW only)
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 960 --height 540 --supersample 2 --repeats 7
impulse: 960x540 supersample=2 deposit=nearest repeats=7 spatial=0
plan_mode=estimate
setup: init=0.042270 s fftw_plan=0.006317 s kernel_fft=0.010302 s
Fast FFTW: linear min=1266x2106, fft=1280x2160, workers=16, plan=estimate, plan=0.006317 s, kernel_fft=0.010302 s, scratch=147.7 MiB
impulse_hash=6e2de17f6859cd33
fftw: min=0.038361 s median=0.039872 s max=0.055834 s
fftw samples: 0.055834 0.038745 0.039872 0.038823 0.038361 0.048436 0.040141
fftw_hdr_hash=c493dbb793d04706
rss_max_kb=216272
4K 3840x2160 N=2, FFTW_ESTIMATE (FFTW only)
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 3840 --height 2160 --supersample 2 --repeats 5
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.524840 s fftw_plan=0.008154 s kernel_fft=0.482906 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=estimate, plan=0.008154 s, kernel_fft=0.482906 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=2.368634 s median=2.481927 s max=2.553895 s
fftw samples: 2.552212 2.553895 2.481927 2.389895 2.368634
fftw_hdr_hash=1c5720f72a040e6d
rss_max_kb=2879596
4K 3840x2160 N=4, FFTW_ESTIMATE (FFTW only)
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 3840 --height 2160 --supersample 4 --repeats 5
impulse: 3840x2160 supersample=4 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.738771 s fftw_plan=0.003636 s kernel_fft=0.599235 s
Fast FFTW: linear min=9010x15730, fft=9072x15750, workers=16, plan=estimate, plan=0.003636 s, kernel_fft=0.599235 s, scratch=7631.4 MiB
impulse_hash=e6a528db0fb16c33
fftw: min=3.276609 s median=3.339285 s max=3.369411 s
fftw samples: 3.309935 3.369411 3.276609 3.345614 3.339285
fftw_hdr_hash=e55bed24ef3f5506
rss_max_kb=10919208
FFTW_ESTIMATE versus FFTW_MEASURE
FFTW_MEASURE cuts steady-state time but costs a large one-time plan. Because
the plan is created once per process and reused across frames, it only pays off
for long movie renders; production therefore defaults to FFTW_ESTIMATE.
Wisdom import/export would make FFTW_MEASURE the better default and remains a
follow-up. Measured FFTW_MEASURE plan output is not bit-reproducible (plan
measurement is timing-dependent), so its fftw_hdr_hash may differ run to run;
the spatial comparison stays at roundoff.
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 320 --height 240 --supersample 2 --repeats 7 --measure --spatial
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=measure
setup: init=1.584388 s fftw_plan=1.091687 s kernel_fft=0.467405 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=measure, plan=1.091687 s, kernel_fft=0.467405 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.003995 s median=0.004702 s max=0.005804 s
fftw samples: 0.003995 0.004702 0.004330 0.005804 0.004939 0.004986 0.004619
fftw_hdr_hash=626e386cba92ffa5
spatial: 0.706244 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=8.05e-16 peak=0.359 speedup=150.198x
rss_max_kb=51860
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
--width 3840 --height 2160 --supersample 2 --repeats 5 --measure
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=measure
setup: init=41.408673 s fftw_plan=41.239987 s kernel_fft=0.134272 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=measure, plan=41.239987 s, kernel_fft=0.134272 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=0.518707 s median=0.522181 s max=0.539465 s
fftw samples: 0.523380 0.539465 0.518707 0.520357 0.522181
fftw_hdr_hash=ab90c4b076c775e1
rss_max_kb=2884132
End-to-end bounded comparison
Same catalog, geometry, exposure, supersample factors, and worker environment as
benchmarks/fast_mode_cpu_2026-09-18.md. Raw terminal output including the
timing wrapper:
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--output /tmp/opencode/fastmode_fftw_bench/normal.png
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000046 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.776 s
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.729 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/normal.png (ok)
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
real=17.526 user=225.46s sys=2.44s cpu=1300%
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 2 \
--output /tmp/opencode/fastmode_fftw_bench/fast_n2.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000049 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.007770 s, kernel_fft=0.005020 s, scratch=30.2 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.735 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=0.995 user=5.30s sys=0.18s cpu=550%
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 4 \
--output /tmp/opencode/fastmode_fftw_bench/fast_n4.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000044 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007294 s, kernel_fft=0.012307 s, scratch=120.7 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.722 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n4.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=1.033 user=5.56s sys=0.22s cpu=559%
Image counts are identical across normal and both fast runs (5,912,840), so the FFTW resolver did not change catalog/deposit behavior. Output SHA-256:
normal.png e9f0a98ae7a3c3db546a3f276f400be98c96446fab7bfbc632f388702d4722d1
fast_n2.png 1f144aefb4a5520d9d76a13ab69026a72580663e5e45f53128a87a0801176ccf
fast_n4.png 679f14a67a085f76c144ecbe36af51b3b2534c152e8ef59708a7884a3f7c8464
Production-density 4K Galactic-center comparison
This is the full 3840x2160, 50-degree-field run over the local processed 2MASS all-sky catalog, not the bounded 320x240 comparison above.
Provenance
- Git revision:
229f50cd864d9768324042c10edbb56bfc120f08(Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve). - The tracked worktree matched that revision when this record was written; the
unrelated untracked
nmesh_spacetime_output_format.mdwas not an input. - Executable:
build/Release/minkowski_sky, built at2026-09-25 23:16:45 -0400, before the three runs below; size 146,856 bytes; SHA-2567f6486285b2ff498adb2d4f1acaf62dc96b11125c9cec114f572b9379648962c. A Make dry-run after the commit requested only the Makefile's unconditionalFORCErelink and no object recompilation, so the recorded binary used the object set corresponding to the committed tracked sources. - Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB. The commands
did not explicitly set
OMP_NUM_THREADS; fast-mode initialization reportedworkers=16. - Input:
assets/2mass/processed/all_sky, observed when recording as 64,801 regular files and 16,207,861,727 bytes. The renderer selected 1,742 tiles and loaded 54,509,992 stars in each run. No content manifest hash was recorded for this external processed catalog. - Timing: zsh built-in
time; its final timing line is included verbatim for each command.
Standard cached spatial PSF
time ./build/Release/minkowski_sky \
--psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 3840 --height 2160 \
--look-ra-deg 262.5 --look-dec-deg -30 \
--fov-deg 50 \
--coarse-cell-pixels 16 \
--exposure 1e13 \
--psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
--catalog-load-workers 4 \
--hdr-output \
--output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000130 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 85 px, relative tail 1e-06, tail abs 1e-06, boundary 1e-07, build 2.176 s
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.886 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png (ok)
PSF splats: cached 52477581, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 66954
Warning: --psf-min-y discarded one or more PSF events.
./build/Release/minkowski_sky --psf-relative-tail 1e-6 --psf-min-y 1e-8 1e8 938.60s user 15.80s system 1391% cpu 1:08.61 total
Fast mode, N=2
time ./build/Release/minkowski_sky --fast-mode --fast-supersample 2 \
--psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 3840 --height 2160 \
--look-ra-deg 262.5 --look-dec-deg -30 \
--fov-deg 50 \
--coarse-cell-pixels 16 \
--exposure 1e13 \
--psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
--catalog-load-workers 4 \
--hdr-output \
--output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000101 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 169 ss px (84.50 final px), retained flux 0.999999, cached 4-point quadrature
Fast FFTW: linear min=4658x8018, fft=4704x8064, workers=16, plan=estimate, plan=0.005275 s, kernel_fft=0.141445 s, scratch=2026.1 MiB
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.873 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png (ok)
Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954
./build/Release/minkowski_sky --fast-mode --fast-supersample 2 1e-6 1e-8 115.74s user 2.08s system 763% cpu 15.439 total
Fast mode, N=4
time ./build/Release/minkowski_sky --fast-mode --fast-supersample 4 \
--psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 3840 --height 2160 \
--look-ra-deg 262.5 --look-dec-deg -30 \
--fov-deg 50 \
--coarse-cell-pixels 16 \
--exposure 1e13 \
--psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
--catalog-load-workers 4 \
--hdr-output \
--output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000043 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 336 ss px (84.00 final px), retained flux 0.999999, cached 4-point quadrature
Fast FFTW: linear min=9312x16032, fft=9375x16128, workers=16, plan=estimate, plan=0.005384 s, kernel_fft=0.743673 s, scratch=8075.5 MiB
Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.796 s; 4 loader workers
Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png (ok)
Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954
./build/Release/minkowski_sky --fast-mode --fast-supersample 4 1e-6 1e-8 135.62s user 5.04s system 722% cpu 19.469 total
Artifacts and comparison
05517bde1383327043691ac4b713f638b03ebf6dc72207fe8710a8dce747a03e output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png
86e9220feac277d77eba87aa8d0febadca5f4eea35cb56164a4eb1578b20fccd output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png
5dfef3b4bd39d9809463594a1065c32c65ad397396514664c8aa22a5da6d780b output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png
1ea508d1ee787c2e5f35ccdf37ca5f4a58e9ff5e9984b50928d66ecf88c44c3d output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits
71de3f1e11cb9f811f5d08ea45b51baf5968e8523873873d6f7c414c1849ba53 output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits
554848e39fdfb3b4815fb8f2a591b8d356cbdeaaea235a73f19b46017a18e42a output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits
All three runs produced 52,544,535 images and applied the same event filtering:
52,477,581 cached/deposited events, no wing clipping or direct fallback, and
66,954 events (0.12742334% of rendered images) discarded by --psf-min-y.
Using the standard HDR buffer as the reference, with
relative_L2 = ||fast - base||_2 / ||base||_2, the existing FITS artifacts give:
| Mode | Wall time | Wall speedup | Approx. speedup after subtracting catalog prefetch | User-CPU speedup | Scratch | HDR relative L2 | HDR total-flux ratio | Tone-mapped PNG normalized RMSE |
|---|---|---|---|---|---|---|---|---|
| Standard | 68.610 s | 1.000x | 1.000x | 1.000x | n/a | 0 | 1 | 0 |
| Fast N=2 | 15.439 s | 4.444x | 6.557x | 8.110x | 2026.1 MiB | 10.4384% | 1.000176588 | 0.00678916 |
| Fast N=4 | 19.469 s | 3.524x | 4.587x | 6.921x | 8075.5 MiB | 4.37249% | 1.000178510 | 0.00365798 |
The post-prefetch figures subtract only the renderer-reported catalog prefetch
duration from wall time; they are useful context, not isolated PSF timings. N=4
costs 26.1% more wall time and 3.986x the scratch of N=2, while reducing the HDR
relative-L2 difference by about 58% and the tone-mapped PNG RMSE by about 46%.
The nearly identical +0.0177% total-flux offset for N=2 and N=4 shows that the
larger N primarily improves spatial placement rather than global photometry.
For the intended rapid single-frame exposure/PSF preview, N=2 is therefore the
default throughput point; N=4 is the higher-fidelity preview option.
Error distribution across the image
The unweighted relative L2 above is dominated by a small number of very bright
PSF cores, so it understates how well the dense faint-star texture is
reproduced in this still frame. scripts/fast_mode_error_decomposition.py
reports the error over the flattened scalar RGB channel samples:
python3 scripts/fast_mode_error_decomposition.py \
output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits \
output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits \
output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits
===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits ===== relative_L2=10.4384%
top |d| RGB samples: sq-error energy base energy linear-RGB sum sample share
top 0.0010%: 69.136% 48.948% 2.822% 0.0010%
top 0.0100%: 95.732% 87.994% 9.141% 0.0100%
top 0.1000%: 99.365% 97.633% 16.396% 0.1000%
top 1.0000%: 99.893% 99.224% 25.984% 1.0000%
top 10.0000%: 99.988% 99.679% 46.006% 10.0000%
base-value range (disjoint): rel_RMS signed bias base energy % linear-RGB sum % samples
[1e-06,0.001): 4.7788% +0.7665% 0.000% 0.001% 67026
[0.001,0.0599): 4.5658% +0.1902% 0.013% 11.063% 12374574
[0.0599,0.1): 4.6294% +0.1297% 0.027% 10.881% 4486724
[0.1,0.229): 4.9079% +0.0866% 0.124% 25.365% 5466556
[0.229,0.769): 6.9340% +0.0100% 0.301% 24.638% 2239488
[0.769,4.62): 11.4421% -0.0990% 0.615% 10.315% 223948
[4.62,inf): 10.4475% -0.1789% 98.921% 17.738% 24884
===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits ===== relative_L2=4.3725%
top |d| RGB samples: sq-error energy base energy linear-RGB sum sample share
top 0.0010%: 65.911% 35.557% 2.418% 0.0010%
top 0.0100%: 94.540% 86.397% 8.879% 0.0100%
top 0.1000%: 99.108% 96.708% 16.179% 0.1000%
top 1.0000%: 99.848% 99.208% 25.909% 1.0000%
top 10.0000%: 99.982% 99.694% 45.966% 10.0000%
base-value range (disjoint): rel_RMS signed bias base energy % linear-RGB sum % samples
[1e-06,0.001): 2.4187% +0.5689% 0.000% 0.001% 67026
[0.001,0.0599): 2.2772% +0.0729% 0.013% 11.063% 12374574
[0.0599,0.1): 2.3129% +0.0550% 0.027% 10.881% 4486724
[0.1,0.229): 2.4542% +0.0393% 0.124% 25.365% 5466556
[0.229,0.769): 3.4700% +0.0166% 0.301% 24.638% 2239488
[0.769,4.62): 5.6752% -0.0189% 0.615% 10.315% 223948
[4.62,inf): 4.3682% -0.0469% 98.921% 17.738% 24884
The FITS RGB channels are flattened into scalar samples: sq-error energy is
the fraction of ||fast-base||^2, linear-RGB sum is the sum of baseline
linear-RGB sample values (proportional to total flux, not a per-source
photometry), and the base-value ranges partition all samples and sum to 100%.
The top 0.01% of RGB channel samples account for 95.7% (N=2) or 94.5%
(N=4) of squared-error energy and about 9% of the baseline linear-RGB sum.
The flux-bearing faint ranges ([1e-3,0.0599), [0.0599,0.1), [0.1,0.229),
together ~47% of the linear-RGB sum) have 2.3-4.9% local relative RMS and
signed bias of +0.04% to +0.19%; the many faint PSFs partially cancel. This
is consistent with the concentrated bright-core difference being predominantly a
bounded sub-pixel redistribution rather than a large net linear-RGB loss, and
explains why this tone-mapped still PNG remains smooth despite the global
relative L2.
This comparison contains one independent frame only. It does not measure
temporal coherence: nearest deposition jumps by 1/N output pixel per axis
when a moving image crosses a supersampled-cell boundary (0.5 px at N=2,
0.25 px at N=4). Both can produce visible movie flicker and are not
qualified here as visually faithful movie-output paths.
Reading
- Steady-state FFTW is ~109x (N=2) and ~279x (N=4) faster than the spatial reference at 320x240, and stays under ~0.06 s per frame at 960x540.
- 4K N=2 resolves in ~2.5 s (ESTIMATE) or ~0.52 s (MEASURE) on this host; 4K N=4 needs ~7.6 GiB scratch and ~3.3 s (ESTIMATE).
- FFTW-versus-spatial differences are floating-point roundoff (
max_absat the1e-15level against peaks near0.5), not crop, wrap, channel, or normalization errors. - At the committed revision, the production-density 4K Galactic-center run completes in 15.439 s at N=2 versus 68.610 s for the standard path (4.444x end-to-end), with identical catalog image/deposit/filter counts. N=4 trades 26.1% more wall time for materially lower spatial error.
- The 4K still-frame error is concentrated: the top 0.01% of RGB channel samples hold ~95% of the squared-error energy but only ~9% of the baseline linear-RGB sum, while the flux-bearing faint-value ranges stay at ~2-5% local relative RMS with signed bias at or below ~0.2%. This does not establish temporal fidelity.