Files
GR-raytracing/benchmarks/fast_mode_fftw_2026-09-25.md
T
wyj 229f50cd86 Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.

- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
2026-09-25 23:35:45 -04:00

13 KiB

Fast-mode FFTW convolution benchmark (2026-09-25)

Replaces the spatial fast-mode global convolution with a reusable FFTW linear convolution. This record separates one-time plan/setup from steady-state execution and repeats the bounded Galactic-center comparison. Raw resolver and end-to-end output is pasted verbatim from the runs below.

Environment

  • Git revision: 3deebfb plus the uncommitted FFTW work.
  • Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB.
  • FFTW: pkg-config --modversion fftw3_omp = 3.3.10; pkg-config --libs fftw3_omp = -lfftw3_omp -lfftw3.
  • Build: make -j8 BUILD_TYPE=Release PSF_BACKEND=cpu SPACETIME=minkowski backend (-O2 -DNDEBUG -fopenmp -march=native, libpng, -DFAST_PSF_FFTW).
  • Binary SHA-256 (this run):
    • build/Release/minkowski_sky: ead2a543b8166abe6f82bc1cd41b2e985a6db517aa45a6416d08d188138b8c47
    • build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw: 8da8bbd532104a0c4f524aabf40b66a50c8c715e44479c7db467669ab5a21215
  • Timing wrapper: /usr/bin/time is not installed on this host. The reproducible wrapper used for the end-to-end runs is zsh's built-in time with TIMEFMT='real=%*E user=%U sys=%S cpu=%P', i.e. each command was run as TIMEFMT='real=%*E user=%U sys=%S cpu=%P'; time <command>. The real=... lines below are the verbatim wrapper output.

Correctness

make -j8 BUILD_TYPE=Debug PSF_BACKEND=cpu SPACETIME=minkowski test

All C tests pass, including the dedicated FFTW-vs-spatial test (258 comparisons over sizes 17x13/16x16/23x31/33x17, N = 1..4, nearest and bilinear, eight scenes; worst max_abs = 1.8e-14), the test_frame fast-mode regressions, the non-fast FITS references, and both camera CLI checks. HIP and dummy renderers compile and link without FFTW; make PSF_BACKEND=dummy test now fails fast with make test requires PSF_BACKEND=cpu instead of reusing a stale CPU test binary.

Resolver-only benchmark

One deterministic sparse impulse buffer, default PSF (FWHM 2.7, beta 4.5, relative tail 1e-8), OMP_NUM_THREADS=16. --spatial additionally runs the reference resolver on the identical buffer. Binary: build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw.

Default 320x240 N=2, FFTW_ESTIMATE

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 320 --height 240 --supersample 2 --repeats 7 --spatial
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.039444 s fftw_plan=0.008085 s kernel_fft=0.005237 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.008085 s, kernel_fft=0.005237 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.005602 s median=0.006455 s max=0.008144 s
fftw samples: 0.006697 0.005784 0.005602 0.005898 0.006455 0.007675 0.008144
fftw_hdr_hash=9ad60d92992c9bbd
spatial: 0.702470 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=7.49e-16 peak=0.359 speedup=108.831x
rss_max_kb=52908

Default 320x240 N=4, FFTW_ESTIMATE

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 320 --height 240 --supersample 4 --repeats 7 --spatial
impulse: 320x240 supersample=4 deposit=nearest repeats=7 spatial=1
plan_mode=estimate
setup: init=0.114719 s fftw_plan=0.007169 s kernel_fft=0.011942 s
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007169 s, kernel_fft=0.011942 s, scratch=120.7 MiB
impulse_hash=2639864ea9d3a26a
fftw: min=0.034756 s median=0.039591 s max=0.042474 s
fftw samples: 0.041352 0.039591 0.034756 0.040866 0.042474 0.035364 0.036241
fftw_hdr_hash=1ff070fdaac85f4a
spatial: 11.054385 s
spatial_hdr_hash=7682ee8550f88efe
difference: max_abs=5.38e-15 peak=0.533 speedup=279.214x
rss_max_kb=161032

Medium 960x540 N=2, FFTW_ESTIMATE (FFTW only)

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 960 --height 540 --supersample 2 --repeats 7
impulse: 960x540 supersample=2 deposit=nearest repeats=7 spatial=0
plan_mode=estimate
setup: init=0.042270 s fftw_plan=0.006317 s kernel_fft=0.010302 s
Fast FFTW: linear min=1266x2106, fft=1280x2160, workers=16, plan=estimate, plan=0.006317 s, kernel_fft=0.010302 s, scratch=147.7 MiB
impulse_hash=6e2de17f6859cd33
fftw: min=0.038361 s median=0.039872 s max=0.055834 s
fftw samples: 0.055834 0.038745 0.039872 0.038823 0.038361 0.048436 0.040141
fftw_hdr_hash=c493dbb793d04706
rss_max_kb=216272

4K 3840x2160 N=2, FFTW_ESTIMATE (FFTW only)

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 3840 --height 2160 --supersample 2 --repeats 5
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.524840 s fftw_plan=0.008154 s kernel_fft=0.482906 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=estimate, plan=0.008154 s, kernel_fft=0.482906 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=2.368634 s median=2.481927 s max=2.553895 s
fftw samples: 2.552212 2.553895 2.481927 2.389895 2.368634
fftw_hdr_hash=1c5720f72a040e6d
rss_max_kb=2879596

4K 3840x2160 N=4, FFTW_ESTIMATE (FFTW only)

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 3840 --height 2160 --supersample 4 --repeats 5
impulse: 3840x2160 supersample=4 deposit=nearest repeats=5 spatial=0
plan_mode=estimate
setup: init=0.738771 s fftw_plan=0.003636 s kernel_fft=0.599235 s
Fast FFTW: linear min=9010x15730, fft=9072x15750, workers=16, plan=estimate, plan=0.003636 s, kernel_fft=0.599235 s, scratch=7631.4 MiB
impulse_hash=e6a528db0fb16c33
fftw: min=3.276609 s median=3.339285 s max=3.369411 s
fftw samples: 3.309935 3.369411 3.276609 3.345614 3.339285
fftw_hdr_hash=e55bed24ef3f5506
rss_max_kb=10919208

FFTW_ESTIMATE versus FFTW_MEASURE

FFTW_MEASURE cuts steady-state time but costs a large one-time plan. Because the plan is created once per process and reused across frames, it only pays off for long movie renders; production therefore defaults to FFTW_ESTIMATE. Wisdom import/export would make FFTW_MEASURE the better default and remains a follow-up. Measured FFTW_MEASURE plan output is not bit-reproducible (plan measurement is timing-dependent), so its fftw_hdr_hash may differ run to run; the spatial comparison stays at roundoff.

OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 320 --height 240 --supersample 2 --repeats 7 --measure --spatial
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
plan_mode=measure
setup: init=1.584388 s fftw_plan=1.091687 s kernel_fft=0.467405 s
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=measure, plan=1.091687 s, kernel_fft=0.467405 s, scratch=30.2 MiB
impulse_hash=5334fa90cfbb1089
fftw: min=0.003995 s median=0.004702 s max=0.005804 s
fftw samples: 0.003995 0.004702 0.004330 0.005804 0.004939 0.004986 0.004619
fftw_hdr_hash=626e386cba92ffa5
spatial: 0.706244 s
spatial_hdr_hash=9327ac3e8e10d670
difference: max_abs=8.05e-16 peak=0.359 speedup=150.198x
rss_max_kb=51860
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
  --width 3840 --height 2160 --supersample 2 --repeats 5 --measure
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
plan_mode=measure
setup: init=41.408673 s fftw_plan=41.239987 s kernel_fft=0.134272 s
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=measure, plan=41.239987 s, kernel_fft=0.134272 s, scratch=1907.8 MiB
impulse_hash=e3b6d7f3d621630d
fftw: min=0.518707 s median=0.522181 s max=0.539465 s
fftw samples: 0.523380 0.539465 0.518707 0.520357 0.522181
fftw_hdr_hash=ab90c4b076c775e1
rss_max_kb=2884132

End-to-end bounded comparison

Same catalog, geometry, exposure, supersample factors, and worker environment as benchmarks/fast_mode_cpu_2026-09-18.md. Raw terminal output including the timing wrapper:

OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
  time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --output /tmp/opencode/fastmode_fftw_bench/normal.png
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000046 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.776 s
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.729 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/normal.png (ok)
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
real=17.526 user=225.46s sys=2.44s cpu=1300%
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
  time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 2 \
  --output /tmp/opencode/fastmode_fftw_bench/fast_n2.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000049 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.007770 s, kernel_fft=0.005020 s, scratch=30.2 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.735 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=0.995 user=5.30s sys=0.18s cpu=550%
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
  time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 4 \
  --output /tmp/opencode/fastmode_fftw_bench/fast_n4.png
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000044 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007294 s, kernel_fft=0.012307 s, scratch=120.7 MiB
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.722 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n4.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
real=1.033 user=5.56s sys=0.22s cpu=559%

Image counts are identical across normal and both fast runs (5,912,840), so the FFTW resolver did not change catalog/deposit behavior. Output SHA-256:

normal.png   e9f0a98ae7a3c3db546a3f276f400be98c96446fab7bfbc632f388702d4722d1
fast_n2.png  1f144aefb4a5520d9d76a13ab69026a72580663e5e45f53128a87a0801176ccf
fast_n4.png  679f14a67a085f76c144ecbe36af51b3b2534c152e8ef59708a7884a3f7c8464

Reading

  • Steady-state FFTW is ~109x (N=2) and ~279x (N=4) faster than the spatial reference at 320x240, and stays under ~0.06 s per frame at 960x540.
  • 4K N=2 resolves in ~2.5 s (ESTIMATE) or ~0.52 s (MEASURE) on this host; 4K N=4 needs ~7.6 GiB scratch and ~3.3 s (ESTIMATE).
  • FFTW-versus-spatial differences are floating-point roundoff (max_abs at the 1e-15 level against peaks near 0.5), not crop, wrap, channel, or normalization errors.