Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.
- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
281 lines
13 KiB
Markdown
281 lines
13 KiB
Markdown
# Fast-mode FFTW convolution benchmark (2026-09-25)
|
|
|
|
Replaces the spatial fast-mode global convolution with a reusable FFTW linear
|
|
convolution. This record separates one-time plan/setup from steady-state
|
|
execution and repeats the bounded Galactic-center comparison. Raw resolver and
|
|
end-to-end output is pasted verbatim from the runs below.
|
|
|
|
## Environment
|
|
|
|
- Git revision: `3deebfb` plus the uncommitted FFTW work.
|
|
- Host: 12th Gen Intel Core i7-12700K, 16 hardware threads, 128 GB.
|
|
- FFTW: `pkg-config --modversion fftw3_omp` = `3.3.10`;
|
|
`pkg-config --libs fftw3_omp` = `-lfftw3_omp -lfftw3`.
|
|
- Build:
|
|
`make -j8 BUILD_TYPE=Release PSF_BACKEND=cpu SPACETIME=minkowski backend`
|
|
(`-O2 -DNDEBUG -fopenmp -march=native`, libpng, `-DFAST_PSF_FFTW`).
|
|
- Binary SHA-256 (this run):
|
|
- `build/Release/minkowski_sky`:
|
|
`ead2a543b8166abe6f82bc1cd41b2e985a6db517aa45a6416d08d188138b8c47`
|
|
- `build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw`:
|
|
`8da8bbd532104a0c4f524aabf40b66a50c8c715e44479c7db467669ab5a21215`
|
|
- Timing wrapper: `/usr/bin/time` is not installed on this host. The
|
|
reproducible wrapper used for the end-to-end runs is zsh's built-in `time`
|
|
with `TIMEFMT='real=%*E user=%U sys=%S cpu=%P'`, i.e. each command was run as
|
|
`TIMEFMT='real=%*E user=%U sys=%S cpu=%P'; time <command>`. The `real=...`
|
|
lines below are the verbatim wrapper output.
|
|
|
|
## Correctness
|
|
|
|
```sh
|
|
make -j8 BUILD_TYPE=Debug PSF_BACKEND=cpu SPACETIME=minkowski test
|
|
```
|
|
|
|
All C tests pass, including the dedicated FFTW-vs-spatial test (258 comparisons
|
|
over sizes `17x13`/`16x16`/`23x31`/`33x17`, `N = 1..4`, nearest and bilinear,
|
|
eight scenes; worst `max_abs = 1.8e-14`), the `test_frame` fast-mode
|
|
regressions, the non-fast FITS references, and both camera CLI checks. HIP and
|
|
dummy renderers compile and link without FFTW; `make PSF_BACKEND=dummy test`
|
|
now fails fast with `make test requires PSF_BACKEND=cpu` instead of reusing a
|
|
stale CPU test binary.
|
|
|
|
## Resolver-only benchmark
|
|
|
|
One deterministic sparse impulse buffer, default PSF (FWHM 2.7, beta 4.5,
|
|
relative tail 1e-8), `OMP_NUM_THREADS=16`. `--spatial` additionally runs the
|
|
reference resolver on the identical buffer. Binary:
|
|
`build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw`.
|
|
|
|
### Default 320x240 N=2, FFTW_ESTIMATE
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 320 --height 240 --supersample 2 --repeats 7 --spatial
|
|
```
|
|
|
|
```text
|
|
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
|
|
plan_mode=estimate
|
|
setup: init=0.039444 s fftw_plan=0.008085 s kernel_fft=0.005237 s
|
|
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.008085 s, kernel_fft=0.005237 s, scratch=30.2 MiB
|
|
impulse_hash=5334fa90cfbb1089
|
|
fftw: min=0.005602 s median=0.006455 s max=0.008144 s
|
|
fftw samples: 0.006697 0.005784 0.005602 0.005898 0.006455 0.007675 0.008144
|
|
fftw_hdr_hash=9ad60d92992c9bbd
|
|
spatial: 0.702470 s
|
|
spatial_hdr_hash=9327ac3e8e10d670
|
|
difference: max_abs=7.49e-16 peak=0.359 speedup=108.831x
|
|
rss_max_kb=52908
|
|
```
|
|
|
|
### Default 320x240 N=4, FFTW_ESTIMATE
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 320 --height 240 --supersample 4 --repeats 7 --spatial
|
|
```
|
|
|
|
```text
|
|
impulse: 320x240 supersample=4 deposit=nearest repeats=7 spatial=1
|
|
plan_mode=estimate
|
|
setup: init=0.114719 s fftw_plan=0.007169 s kernel_fft=0.011942 s
|
|
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007169 s, kernel_fft=0.011942 s, scratch=120.7 MiB
|
|
impulse_hash=2639864ea9d3a26a
|
|
fftw: min=0.034756 s median=0.039591 s max=0.042474 s
|
|
fftw samples: 0.041352 0.039591 0.034756 0.040866 0.042474 0.035364 0.036241
|
|
fftw_hdr_hash=1ff070fdaac85f4a
|
|
spatial: 11.054385 s
|
|
spatial_hdr_hash=7682ee8550f88efe
|
|
difference: max_abs=5.38e-15 peak=0.533 speedup=279.214x
|
|
rss_max_kb=161032
|
|
```
|
|
|
|
### Medium 960x540 N=2, FFTW_ESTIMATE (FFTW only)
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 960 --height 540 --supersample 2 --repeats 7
|
|
```
|
|
|
|
```text
|
|
impulse: 960x540 supersample=2 deposit=nearest repeats=7 spatial=0
|
|
plan_mode=estimate
|
|
setup: init=0.042270 s fftw_plan=0.006317 s kernel_fft=0.010302 s
|
|
Fast FFTW: linear min=1266x2106, fft=1280x2160, workers=16, plan=estimate, plan=0.006317 s, kernel_fft=0.010302 s, scratch=147.7 MiB
|
|
impulse_hash=6e2de17f6859cd33
|
|
fftw: min=0.038361 s median=0.039872 s max=0.055834 s
|
|
fftw samples: 0.055834 0.038745 0.039872 0.038823 0.038361 0.048436 0.040141
|
|
fftw_hdr_hash=c493dbb793d04706
|
|
rss_max_kb=216272
|
|
```
|
|
|
|
### 4K 3840x2160 N=2, FFTW_ESTIMATE (FFTW only)
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 3840 --height 2160 --supersample 2 --repeats 5
|
|
```
|
|
|
|
```text
|
|
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
|
|
plan_mode=estimate
|
|
setup: init=0.524840 s fftw_plan=0.008154 s kernel_fft=0.482906 s
|
|
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=estimate, plan=0.008154 s, kernel_fft=0.482906 s, scratch=1907.8 MiB
|
|
impulse_hash=e3b6d7f3d621630d
|
|
fftw: min=2.368634 s median=2.481927 s max=2.553895 s
|
|
fftw samples: 2.552212 2.553895 2.481927 2.389895 2.368634
|
|
fftw_hdr_hash=1c5720f72a040e6d
|
|
rss_max_kb=2879596
|
|
```
|
|
|
|
### 4K 3840x2160 N=4, FFTW_ESTIMATE (FFTW only)
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 3840 --height 2160 --supersample 4 --repeats 5
|
|
```
|
|
|
|
```text
|
|
impulse: 3840x2160 supersample=4 deposit=nearest repeats=5 spatial=0
|
|
plan_mode=estimate
|
|
setup: init=0.738771 s fftw_plan=0.003636 s kernel_fft=0.599235 s
|
|
Fast FFTW: linear min=9010x15730, fft=9072x15750, workers=16, plan=estimate, plan=0.003636 s, kernel_fft=0.599235 s, scratch=7631.4 MiB
|
|
impulse_hash=e6a528db0fb16c33
|
|
fftw: min=3.276609 s median=3.339285 s max=3.369411 s
|
|
fftw samples: 3.309935 3.369411 3.276609 3.345614 3.339285
|
|
fftw_hdr_hash=e55bed24ef3f5506
|
|
rss_max_kb=10919208
|
|
```
|
|
|
|
### FFTW_ESTIMATE versus FFTW_MEASURE
|
|
|
|
`FFTW_MEASURE` cuts steady-state time but costs a large one-time plan. Because
|
|
the plan is created once per process and reused across frames, it only pays off
|
|
for long movie renders; production therefore defaults to `FFTW_ESTIMATE`.
|
|
Wisdom import/export would make `FFTW_MEASURE` the better default and remains a
|
|
follow-up. Measured FFTW_MEASURE plan output is not bit-reproducible (plan
|
|
measurement is timing-dependent), so its `fftw_hdr_hash` may differ run to run;
|
|
the spatial comparison stays at roundoff.
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 320 --height 240 --supersample 2 --repeats 7 --measure --spatial
|
|
```
|
|
|
|
```text
|
|
impulse: 320x240 supersample=2 deposit=nearest repeats=7 spatial=1
|
|
plan_mode=measure
|
|
setup: init=1.584388 s fftw_plan=1.091687 s kernel_fft=0.467405 s
|
|
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=measure, plan=1.091687 s, kernel_fft=0.467405 s, scratch=30.2 MiB
|
|
impulse_hash=5334fa90cfbb1089
|
|
fftw: min=0.003995 s median=0.004702 s max=0.005804 s
|
|
fftw samples: 0.003995 0.004702 0.004330 0.005804 0.004939 0.004986 0.004619
|
|
fftw_hdr_hash=626e386cba92ffa5
|
|
spatial: 0.706244 s
|
|
spatial_hdr_hash=9327ac3e8e10d670
|
|
difference: max_abs=8.05e-16 peak=0.359 speedup=150.198x
|
|
rss_max_kb=51860
|
|
```
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 ./build/Release/obj/minkowski/standard_sink1_cpu/benchmark_fast_psf_fftw \
|
|
--width 3840 --height 2160 --supersample 2 --repeats 5 --measure
|
|
```
|
|
|
|
```text
|
|
impulse: 3840x2160 supersample=2 deposit=nearest repeats=5 spatial=0
|
|
plan_mode=measure
|
|
setup: init=41.408673 s fftw_plan=41.239987 s kernel_fft=0.134272 s
|
|
Fast FFTW: linear min=4506x7866, fft=4536x7875, workers=16, plan=measure, plan=41.239987 s, kernel_fft=0.134272 s, scratch=1907.8 MiB
|
|
impulse_hash=e3b6d7f3d621630d
|
|
fftw: min=0.518707 s median=0.522181 s max=0.539465 s
|
|
fftw samples: 0.523380 0.539465 0.518707 0.520357 0.522181
|
|
fftw_hdr_hash=ab90c4b076c775e1
|
|
rss_max_kb=2884132
|
|
```
|
|
|
|
## End-to-end bounded comparison
|
|
|
|
Same catalog, geometry, exposure, supersample factors, and worker environment as
|
|
`benchmarks/fast_mode_cpu_2026-09-18.md`. Raw terminal output including the
|
|
timing wrapper:
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
|
|
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
|
|
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
|
|
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
|
|
--output /tmp/opencode/fastmode_fftw_bench/normal.png
|
|
```
|
|
|
|
```text
|
|
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000046 s; payload fnv1a64=0e39867b0e70a809
|
|
Blackbody backend: lut
|
|
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.776 s
|
|
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.729 s; 4 loader workers
|
|
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/normal.png (ok)
|
|
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
|
|
real=17.526 user=225.46s sys=2.44s cpu=1300%
|
|
```
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
|
|
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
|
|
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
|
|
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
|
|
--fast-mode --fast-supersample 2 \
|
|
--output /tmp/opencode/fastmode_fftw_bench/fast_n2.png
|
|
```
|
|
|
|
```text
|
|
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
|
|
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000049 s; payload fnv1a64=0e39867b0e70a809
|
|
Blackbody backend: lut
|
|
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
|
|
Fast FFTW: linear min=666x826, fft=672x840, workers=16, plan=estimate, plan=0.007770 s, kernel_fft=0.005020 s, scratch=30.2 MiB
|
|
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.735 s; 4 loader workers
|
|
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n2.png (ok)
|
|
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
|
|
real=0.995 user=5.30s sys=0.18s cpu=550%
|
|
```
|
|
|
|
```sh
|
|
OMP_NUM_THREADS=16 TIMEFMT='real=%*E user=%U sys=%S cpu=%P' \
|
|
time ./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
|
|
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
|
|
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
|
|
--fast-mode --fast-supersample 4 \
|
|
--output /tmp/opencode/fastmode_fftw_bench/fast_n4.png
|
|
```
|
|
|
|
```text
|
|
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
|
|
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000044 s; payload fnv1a64=0e39867b0e70a809
|
|
Blackbody backend: lut
|
|
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
|
|
Fast FFTW: linear min=1330x1650, fft=1344x1680, workers=16, plan=estimate, plan=0.007294 s, kernel_fft=0.012307 s, scratch=120.7 MiB
|
|
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.722 s; 4 loader workers
|
|
Rendered 5912840 images from 6618807 catalog stars to /tmp/opencode/fastmode_fftw_bench/fast_n4.png (ok)
|
|
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
|
|
real=1.033 user=5.56s sys=0.22s cpu=559%
|
|
```
|
|
|
|
Image counts are identical across normal and both fast runs (5,912,840), so the
|
|
FFTW resolver did not change catalog/deposit behavior. Output SHA-256:
|
|
|
|
```text
|
|
normal.png e9f0a98ae7a3c3db546a3f276f400be98c96446fab7bfbc632f388702d4722d1
|
|
fast_n2.png 1f144aefb4a5520d9d76a13ab69026a72580663e5e45f53128a87a0801176ccf
|
|
fast_n4.png 679f14a67a085f76c144ecbe36af51b3b2534c152e8ef59708a7884a3f7c8464
|
|
```
|
|
|
|
## Reading
|
|
|
|
- Steady-state FFTW is ~109x (N=2) and ~279x (N=4) faster than the spatial
|
|
reference at 320x240, and stays under ~0.06 s per frame at 960x540.
|
|
- 4K N=2 resolves in ~2.5 s (ESTIMATE) or ~0.52 s (MEASURE) on this host; 4K N=4
|
|
needs ~7.6 GiB scratch and ~3.3 s (ESTIMATE).
|
|
- FFTW-versus-spatial differences are floating-point roundoff (`max_abs` at the
|
|
`1e-15` level against peaks near `0.5`), not crop, wrap, channel, or
|
|
normalization errors.
|