# Fast-mode FFTW convolution benchmark (2026-09-25) Replaces the spatial fast-mode global convolution with a reusable FFTW linear convolution. This record separates one-time plan/setup from steady-state execution and repeats the bounded Galactic-center comparison. Raw resolver and end-to-end output is pasted verbatim from the runs below. The later production-density 4K comparison is recorded separately against the first committed FFTW implementation, revision `229f50cd864d9768324042c10edbb56bfc120f08`. ## Environment - Git revision: `3deebfb` plus the uncommitted FFTW work. - Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB. - FFTW: `pkg-config --modversion fftw3_omp` = `3.3.10`; `pkg-config --libs fftw3_omp` = `-lfftw3_omp -lfftw3`. - Build: `make -j8 BUILD_TYPE=Release PSF_BACKEND=cpu SPACETIME=minkowski backend` (`-O2 -DNDEBUG -fopenmp -march=native`, libpng, `-DFAST_PSF_FFTW`). - Binary SHA-256 (this run): - `build/Release/minkowski_sky`: `ead2a543b8166abe6f82bc1cd41b2e985a6db517aa45a6416d08d188138b8c47` - `build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw`: `8da8bbd532104a0c4f524aabf40b66a50c8c715e44479c7db467669ab5a21215` - Timing wrapper: `/usr/bin/time` is not installed on this host. The reproducible wrapper used for the end-to-end runs is zsh's built-in `time` with `TIMEFMT='real=%*E user=%U sys=%S cpu=%P'`, i.e. each command was run as `TIMEFMT='real=%*E user=%U sys=%S cpu=%P'; time `. The `real=...` lines below are the verbatim wrapper output. ## Correctness ```sh make -j8 BUILD_TYPE=Debug PSF_BACKEND=cpu SPACETIME=minkowski test ``` All C tests pass, including the dedicated FFTW-vs-spatial test (258 comparisons over sizes `17x13`/`16x16`/`23x31`/`33x17`, `N = 1..4`, nearest and bilinear, eight scenes; worst `max_abs = 1.8e-14`), the `test_frame` fast-mode regressions, the non-fast FITS references, and both camera CLI checks. HIP and dummy renderers compile and link without FFTW; `make PSF_BACKEND=dummy test` now fails fast with `make test requires PSF_BACKEND=cpu` instead of reusing a stale CPU test binary. ## Resolver-only benchmark One deterministic sparse impulse buffer, default PSF (FWHM 2.7, beta 4.5, relative tail 1e-8), `OMP_NUM_THREADS=16`. `--spatial` additionally runs the reference resolver on the identical buffer. Binary: `build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw`. ### Default 320x240 N=2, FFTW_ESTIMATE ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 320 --height 240 --supersample 2 --repeats 7 --spatial ``` ```text impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1 plan_mode=estimate setup: init=0.039444 s fftw_plan=0.008085 s kernel_fft=0.005237 s Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.008085 s, kernel_fft=0.005237 s, scratch=30.2 MiB impulse_hash=5334fa90cfbb1089 fftw: min=0.005602 s median=0.006455 s max=0.008144 s fftw samples: 0.006697 0.005784 0.005602 0.005898 0.006455 0.007675 0.008144 fftw_hdr_hash=9ad60d92992c9bbd spatial: 0.702470 s spatial_hdr_hash=9327ac3e8e10d670 difference: max_abs=7.49e-16 peak=0.359 speedup=108.831x rss_max_kb=52908 ``` ### Default 320x240 N=4, FFTW_ESTIMATE ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 320 --height 240 --supersample 4 --repeats 7 --spatial ``` ```text impulse: 320x240 supersample=4 deposit=nearest repeats=7 spatial=1 plan_mode=estimate setup: init=0.114719 s fftw_plan=0.007169 s kernel_fft=0.011942 s Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007169 s, kernel_fft=0.011942 s, scratch=120.7 MiB impulse_hash=2639864ea9d3a26a fftw: min=0.034756 s median=0.039591 s max=0.042474 s fftw samples: 0.041352 0.039591 0.034756 0.040866 0.042474 0.035364 0.036241 fftw_hdr_hash=1ff070fdaac85f4a spatial: 11.054385 s spatial_hdr_hash=7682ee8550f88efe difference: max_abs=5.38e-15 peak=0.533 speedup=279.214x rss_max_kb=161032 ``` ### Medium 960x540 N=2, FFTW_ESTIMATE (FFTW only) ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 960 --height 540 --supersample 2 --repeats 7 ``` ```text impulse: 960x540 supersample=2 deposit=nearest repeats=7 spatial=0 plan_mode=estimate setup: init=0.042270 s fftw_plan=0.006317 s kernel_fft=0.010302 s Fast FFTW: linear min=1266x2106, fft=1280x2160, workers=16, plan=estimate, plan=0.006317 s, kernel_fft=0.010302 s, scratch=147.7 MiB impulse_hash=6e2de17f6859cd33 fftw: min=0.038361 s median=0.039872 s max=0.055834 s fftw samples: 0.055834 0.038745 0.039872 0.038823 0.038361 0.048436 0.040141 fftw_hdr_hash=c493dbb793d04706 rss_max_kb=216272 ``` ### 4K 3840x2160 N=2, FFTW_ESTIMATE (FFTW only) ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 3840 --height 2160 --supersample 2 --repeats 5 ``` ```text impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0 plan_mode=estimate setup: init=0.524840 s fftw_plan=0.008154 s kernel_fft=0.482906 s Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=estimate, plan=0.008154 s, kernel_fft=0.482906 s, scratch=1907.8 MiB impulse_hash=e3b6d7f3d621630d fftw: min=2.368634 s median=2.481927 s max=2.553895 s fftw samples: 2.552212 2.553895 2.481927 2.389895 2.368634 fftw_hdr_hash=1c5720f72a040e6d rss_max_kb=2879596 ``` ### 4K 3840x2160 N=4, FFTW_ESTIMATE (FFTW only) ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 3840 --height 2160 --supersample 4 --repeats 5 ``` ```text impulse: 3840x2160 supersample=4 deposit=nearest repeats=5 spatial=0 plan_mode=estimate setup: init=0.738771 s fftw_plan=0.003636 s kernel_fft=0.599235 s Fast FFTW: linear min=9010x15730, fft=9072x15750, workers=16, plan=estimate, plan=0.003636 s, kernel_fft=0.599235 s, scratch=7631.4 MiB impulse_hash=e6a528db0fb16c33 fftw: min=3.276609 s median=3.339285 s max=3.369411 s fftw samples: 3.309935 3.369411 3.276609 3.345614 3.339285 fftw_hdr_hash=e55bed24ef3f5506 rss_max_kb=10919208 ``` ### FFTW_ESTIMATE versus FFTW_MEASURE `FFTW_MEASURE` cuts steady-state time but costs a large one-time plan. Because the plan is created once per process and reused across frames, it only pays off for long movie renders; production therefore defaults to `FFTW_ESTIMATE`. Wisdom import/export would make `FFTW_MEASURE` the better default and remains a follow-up. Measured FFTW_MEASURE plan output is not bit-reproducible (plan measurement is timing-dependent), so its `fftw_hdr_hash` may differ run to run; the spatial comparison stays at roundoff. ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 320 --height 240 --supersample 2 --repeats 7 --measure --spatial ``` ```text impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1 plan_mode=measure setup: init=1.584388 s fftw_plan=1.091687 s kernel_fft=0.467405 s Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=measure, plan=1.091687 s, kernel_fft=0.467405 s, scratch=30.2 MiB impulse_hash=5334fa90cfbb1089 fftw: min=0.003995 s median=0.004702 s max=0.005804 s fftw samples: 0.003995 0.004702 0.004330 0.005804 0.004939 0.004986 0.004619 fftw_hdr_hash=626e386cba92ffa5 spatial: 0.706244 s spatial_hdr_hash=9327ac3e8e10d670 difference: max_abs=8.05e-16 peak=0.359 speedup=150.198x rss_max_kb=51860 ``` ```sh OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \ --width 3840 --height 2160 --supersample 2 --repeats 5 --measure ``` ```text impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0 plan_mode=measure setup: init=41.408673 s fftw_plan=41.239987 s kernel_fft=0.134272 s Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=measure, plan=41.239987 s, kernel_fft=0.134272 s, scratch=1907.8 MiB impulse_hash=e3b6d7f3d621630d fftw: min=0.518707 s median=0.522181 s max=0.539465 s fftw samples: 0.523380 0.539465 0.518707 0.520357 0.522181 fftw_hdr_hash=ab90c4b076c775e1 rss_max_kb=2884132 ``` ## End-to-end bounded comparison Same catalog, geometry, exposure, supersample factors, and worker environment as `benchmarks/fast_mode_cpu_2026-09-18.md`. Raw terminal output including the timing wrapper: ```sh OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \ time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \ --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \ --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \ --output /tmp/opencode/fastmode_fftw_bench/normal.png ``` ```text Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000046 s; payload fnv1a64=0e39867b0e70a809 Blackbody backend: lut PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.776 s Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.729 s; 4 loader workers Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/normal.png (ok) PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0 real=17.526 user=225.46s sys=2.44s cpu=1300% ``` ```sh OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \ time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \ --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \ --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \ --fast-mode --fast-supersample 2 \ --output /tmp/opencode/fastmode_fftw_bench/fast_n2.png ``` ```text Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event. Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000049 s; payload fnv1a64=0e39867b0e70a809 Blackbody backend: lut Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.007770 s, kernel_fft=0.005020 s, scratch=30.2 MiB Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.735 s; 4 loader workers Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n2.png (ok) Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0 real=0.995 user=5.30s sys=0.18s cpu=550% ``` ```sh OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \ time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \ --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \ --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \ --fast-mode --fast-supersample 4 \ --output /tmp/opencode/fastmode_fftw_bench/fast_n4.png ``` ```text Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event. Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000044 s; payload fnv1a64=0e39867b0e70a809 Blackbody backend: lut Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007294 s, kernel_fft=0.012307 s, scratch=120.7 MiB Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.722 s; 4 loader workers Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n4.png (ok) Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0 real=1.033 user=5.56s sys=0.22s cpu=559% ``` Image counts are identical across normal and both fast runs (5,912,840), so the FFTW resolver did not change catalog/deposit behavior. Output SHA-256: ```text normal.png e9f0a98ae7a3c3db546a3f276f400be98c96446fab7bfbc632f388702d4722d1 fast_n2.png 1f144aefb4a5520d9d76a13ab69026a72580663e5e45f53128a87a0801176ccf fast_n4.png 679f14a67a085f76c144ecbe36af51b3b2534c152e8ef59708a7884a3f7c8464 ``` ## Production-density 4K Galactic-center comparison This is the full 3840x2160, 50-degree-field run over the local processed 2MASS all-sky catalog, not the bounded 320x240 comparison above. ### Provenance - Git revision: `229f50cd864d9768324042c10edbb56bfc120f08` (`Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve`). - The tracked worktree matched that revision when this record was written; the unrelated untracked `nmesh_spacetime_output_format.md` was not an input. - Executable: `build/Release/minkowski_sky`, built at `2026-09-25 23:16:45 -0400`, before the three runs below; size 146,856 bytes; SHA-256 `7f6486285b2ff498adb2d4f1acaf62dc96b11125c9cec114f572b9379648962c`. A Make dry-run after the commit requested only the Makefile's unconditional `FORCE` relink and no object recompilation, so the recorded binary used the object set corresponding to the committed tracked sources. - Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB. The commands did not explicitly set `OMP_NUM_THREADS`; fast-mode initialization reported `workers=16`. - Input: `assets/2mass/processed/all_sky`, observed when recording as 64,801 regular files and 16,207,861,727 bytes. The renderer selected 1,742 tiles and loaded 54,509,992 stars in each run. No content manifest hash was recorded for this external processed catalog. - Timing: zsh built-in `time`; its final timing line is included verbatim for each command. ### Standard cached spatial PSF ```sh time ./build/Release/minkowski_sky \ --psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \ --all-sky-catalog assets/2mass/processed/all_sky \ --width 3840 --height 2160 \ --look-ra-deg 262.5 --look-dec-deg -30 \ --fov-deg 50 \ --coarse-cell-pixels 16 \ --exposure 1e13 \ --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \ --catalog-load-workers 4 \ --hdr-output \ --output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png ``` ```text Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000130 s; payload fnv1a64=0e39867b0e70a809 Blackbody backend: lut PSF cache ready: 64x64 phases, radius 85 px, relative tail 1e-06, tail abs 1e-06, boundary 1e-07, build 2.176 s Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.886 s; 4 loader workers Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png (ok) PSF splats: cached 52477581, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 66954 Warning: --psf-min-y discarded one or more PSF events. ./build/Release/minkowski_sky --psf-relative-tail 1e-6 --psf-min-y 1e-8 1e8 938.60s user 15.80s system 1391% cpu 1:08.61 total ``` ### Fast mode, N=2 ```sh time ./build/Release/minkowski_sky --fast-mode --fast-supersample 2 \ --psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \ --all-sky-catalog assets/2mass/processed/all_sky \ --width 3840 --height 2160 \ --look-ra-deg 262.5 --look-dec-deg -30 \ --fov-deg 50 \ --coarse-cell-pixels 16 \ --exposure 1e13 \ --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \ --catalog-load-workers 4 \ --hdr-output \ --output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png ``` ```text Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event. Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000101 s; payload fnv1a64=0e39867b0e70a809 Blackbody backend: lut Fast PSF: supersample 2x, deposit nearest, kernel radius 169 ss px (84.50 final px), retained flux 0.999999, cached 4-point quadrature Fast FFTW: linear min=4658x8018, fft=4704x8064, workers=16, plan=estimate, plan=0.005275 s, kernel_fft=0.141445 s, scratch=2026.1 MiB Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.873 s; 4 loader workers Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png (ok) Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954 ./build/Release/minkowski_sky --fast-mode --fast-supersample 2 1e-6 1e-8 115.74s user 2.08s system 763% cpu 15.439 total ``` ### Fast mode, N=4 ```sh time ./build/Release/minkowski_sky --fast-mode --fast-supersample 4 \ --psf-relative-tail 1e-6 --psf-min-y 1e-8 --max-cache-psf-flux 1e8 \ --all-sky-catalog assets/2mass/processed/all_sky \ --width 3840 --height 2160 \ --look-ra-deg 262.5 --look-dec-deg -30 \ --fov-deg 50 \ --coarse-cell-pixels 16 \ --exposure 1e13 \ --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \ --catalog-load-workers 4 \ --hdr-output \ --output output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png ``` ```text Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event. Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000043 s; payload fnv1a64=0e39867b0e70a809 Blackbody backend: lut Fast PSF: supersample 4x, deposit nearest, kernel radius 336 ss px (84.00 final px), retained flux 0.999999, cached 4-point quadrature Fast FFTW: linear min=9312x16032, fft=9375x16128, workers=16, plan=estimate, plan=0.005384 s, kernel_fft=0.743673 s, scratch=8075.5 MiB Catalog prefetch: 1742 requested, 1742 newly loaded (54509992 stars), 0 unavailable in 5.796 s; 4 loader workers Rendered 52544535 images from 54509992 catalog stars to output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png (ok) Fast PSF splats: deposited 52477581, wing-clipped 0, discarded below min-Y 66954 ./build/Release/minkowski_sky --fast-mode --fast-supersample 4 1e-6 1e-8 135.62s user 5.04s system 722% cpu 19.469 total ``` ### Artifacts and comparison ```text 05517bde1383327043691ac4b713f638b03ebf6dc72207fe8710a8dce747a03e output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base.png 86e9220feac277d77eba87aa8d0febadca5f4eea35cb56164a4eb1578b20fccd output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2.png 5dfef3b4bd39d9809463594a1065c32c65ad397396514664c8aa22a5da6d780b output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4.png 1ea508d1ee787c2e5f35ccdf37ca5f4a58e9ff5e9984b50928d66ecf88c44c3d output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits 71de3f1e11cb9f811f5d08ea45b51baf5968e8523873873d6f7c414c1849ba53 output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits 554848e39fdfb3b4815fb8f2a591b8d356cbdeaaea235a73f19b46017a18e42a output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits ``` All three runs produced 52,544,535 images and applied the same event filtering: 52,477,581 cached/deposited events, no wing clipping or direct fallback, and 66,954 events (0.12742334% of rendered images) discarded by `--psf-min-y`. Using the standard HDR buffer as the reference, with `relative_L2 = ||fast - base||_2 / ||base||_2`, the existing FITS artifacts give: | Mode | Wall time | Wall speedup | Approx. speedup after subtracting catalog prefetch | User-CPU speedup | Scratch | HDR relative L2 | HDR total-flux ratio | Tone-mapped PNG normalized RMSE | |---|---:|---:|---:|---:|---:|---:|---:|---:| | Standard | 68.610 s | 1.000x | 1.000x | 1.000x | n/a | 0 | 1 | 0 | | Fast N=2 | 15.439 s | 4.444x | 6.557x | 8.110x | 2026.1 MiB | 10.4384% | 1.000176588 | 0.00678916 | | Fast N=4 | 19.469 s | 3.524x | 4.587x | 6.921x | 8075.5 MiB | 4.37249% | 1.000178510 | 0.00365798 | The post-prefetch figures subtract only the renderer-reported catalog prefetch duration from wall time; they are useful context, not isolated PSF timings. N=4 costs 26.1% more wall time and 3.986x the scratch of N=2, while reducing the HDR relative-L2 difference by about 58% and the tone-mapped PNG RMSE by about 46%. The nearly identical `+0.0177%` total-flux offset for N=2 and N=4 shows that the larger N primarily improves spatial placement rather than global photometry. For the intended rapid single-frame exposure/PSF preview, N=2 is therefore the default throughput point; N=4 is the higher-fidelity preview option. ### Error distribution across the image The unweighted relative L2 above is dominated by a small number of very bright PSF cores, so it understates how well the dense faint-star texture is reproduced in this still frame. `scripts/fast_mode_error_decomposition.py` reports the error over the flattened scalar RGB channel samples: ```sh python3 scripts/fast_mode_error_decomposition.py \ output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-base_HDR.fits \ output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits \ output/imgs/2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits ``` ```text ===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-2_HDR.fits ===== relative_L2=10.4384% top |d| RGB samples: sq-error energy base energy linear-RGB sum sample share top 0.0010%: 69.136% 48.948% 2.822% 0.0010% top 0.0100%: 95.732% 87.994% 9.141% 0.0100% top 0.1000%: 99.365% 97.633% 16.396% 0.1000% top 1.0000%: 99.893% 99.224% 25.984% 1.0000% top 10.0000%: 99.988% 99.679% 46.006% 10.0000% base-value range (disjoint): rel_RMS signed bias base energy % linear-RGB sum % samples [1e-06,0.001): 4.7788% +0.7665% 0.000% 0.001% 67026 [0.001,0.0599): 4.5658% +0.1902% 0.013% 11.063% 12374574 [0.0599,0.1): 4.6294% +0.1297% 0.027% 10.881% 4486724 [0.1,0.229): 4.9079% +0.0866% 0.124% 25.365% 5466556 [0.229,0.769): 6.9340% +0.0100% 0.301% 24.638% 2239488 [0.769,4.62): 11.4421% -0.0990% 0.615% 10.315% 223948 [4.62,inf): 10.4475% -0.1789% 98.921% 17.738% 24884 ===== 2mass_galactic_center_flat_4k_50deg_1e13_beta3-fast-4_HDR.fits ===== relative_L2=4.3725% top |d| RGB samples: sq-error energy base energy linear-RGB sum sample share top 0.0010%: 65.911% 35.557% 2.418% 0.0010% top 0.0100%: 94.540% 86.397% 8.879% 0.0100% top 0.1000%: 99.108% 96.708% 16.179% 0.1000% top 1.0000%: 99.848% 99.208% 25.909% 1.0000% top 10.0000%: 99.982% 99.694% 45.966% 10.0000% base-value range (disjoint): rel_RMS signed bias base energy % linear-RGB sum % samples [1e-06,0.001): 2.4187% +0.5689% 0.000% 0.001% 67026 [0.001,0.0599): 2.2772% +0.0729% 0.013% 11.063% 12374574 [0.0599,0.1): 2.3129% +0.0550% 0.027% 10.881% 4486724 [0.1,0.229): 2.4542% +0.0393% 0.124% 25.365% 5466556 [0.229,0.769): 3.4700% +0.0166% 0.301% 24.638% 2239488 [0.769,4.62): 5.6752% -0.0189% 0.615% 10.315% 223948 [4.62,inf): 4.3682% -0.0469% 98.921% 17.738% 24884 ``` The FITS RGB channels are flattened into scalar samples: `sq-error energy` is the fraction of `||fast-base||^2`, `linear-RGB sum` is the sum of baseline linear-RGB sample values (proportional to total flux, not a per-source photometry), and the base-value ranges partition all samples and sum to 100%. The top 0.01% of RGB channel samples account for 95.7% (`N=2`) or 94.5% (`N=4`) of squared-error energy and about 9% of the baseline linear-RGB sum. The flux-bearing faint ranges (`[1e-3,0.0599)`, `[0.0599,0.1)`, `[0.1,0.229)`, together ~47% of the linear-RGB sum) have 2.3-4.9% local relative RMS and signed bias of `+0.04%` to `+0.19%`; the many faint PSFs partially cancel. This is consistent with the concentrated bright-core difference being predominantly a bounded sub-pixel redistribution rather than a large net linear-RGB loss, and explains why this tone-mapped still PNG remains smooth despite the global relative L2. This comparison contains one independent frame only. It does not measure temporal coherence: nearest deposition jumps by `1/N` output pixel per axis when a moving image crosses a supersampled-cell boundary (`0.5 px` at `N=2`, `0.25 px` at `N=4`). Both can produce visible movie flicker and are not qualified here as visually faithful movie-output paths. ## Reading - Steady-state FFTW is ~109x (N=2) and ~279x (N=4) faster than the spatial reference at 320x240, and stays under ~0.06 s per frame at 960x540. - 4K N=2 resolves in ~2.5 s (ESTIMATE) or ~0.52 s (MEASURE) on this host; 4K N=4 needs ~7.6 GiB scratch and ~3.3 s (ESTIMATE). - FFTW-versus-spatial differences are floating-point roundoff (`max_abs` at the `1e-15` level against peaks near `0.5`), not crop, wrap, channel, or normalization errors. - At the committed revision, the production-density 4K Galactic-center run completes in 15.439 s at N=2 versus 68.610 s for the standard path (4.444x end-to-end), with identical catalog image/deposit/filter counts. N=4 trades 26.1% more wall time for materially lower spatial error. - The 4K still-frame error is concentrated: the top 0.01% of RGB channel samples hold ~95% of the squared-error energy but only ~9% of the baseline linear-RGB sum, while the flux-bearing faint-value ranges stay at ~2-5% local relative RMS with signed bias at or below ~0.2%. This does not establish temporal fidelity.