Files
GR-raytracing/benchmarks/fast_mode_cpu_2026-09-18.md
wyj 3deebfb2fa Feat: Add --fast-mode supersampled point-source accumulation
Add an optional CPU preview path that deposits each point-source image as a
supersampled delta and resolves the whole frame with one global Moffat
convolution plus an N x N box average, instead of splatting a per-event PSF.

- optics: FastPsfAccumulator builds the pixel-area-integral kernel of the
  target Moffat at the supersampled scale (width N*alpha, same beta).  The
  1/N^2 box average then reproduces the final pixel-area integral, so the
  requested FWHM and beta are preserved without renormalisation.  Deposits are
  per-cell atomic adds; resolve accumulates into the caller's HDR buffer.
- frame: fast branch in frame_splat_catalog with one shared supersampled
  buffer and a single resolve per frame; the accumulator is reused across
  movie frames and built from the map dimensions on lens-map import.
- main: --fast-mode, --fast-supersample N (1..8, default 2) and
  --fast-deposit nearest|bilinear (default nearest).  CPU-only and rejected in
  the HIP/dummy backends; --psf-min-y still applies per event while
  --max-cache-psf-flux does not.
- The deposition scheme was chosen by scripts/fast_mode_deposit_error.py:
  nearest keeps the PSF shape exactly with <= 0.5/N px position quantization;
  bilinear keeps the exact centroid but broadens FWHM and beta.  Recorded in
  benchmarks/fast_mode_deposit_2026-09-18.md.
- tests/test_frame.c covers fast nearest vs the direct evaluator at the snapped
  centre, flux conservation, bilinear centroid, min-Y discard, frame plumbing,
  and HDR accumulation onto a non-zero background.
- benchmarks/fast_mode_cpu_2026-09-18.md records a ~10x speedup on the 2MASS
  galactic-centre field with small tone-mapped differences.
2026-09-24 01:36:25 -04:00

5.8 KiB

Fast-mode CPU benchmark (2026-09-18)

Compares the default per-event phase-cache PSF splat with --fast-mode on a 7.4-degree 2MASS galactic-center field. Commands, raw terminal output, and wall-clock totals are preserved below.

Environment

  • Git revision: 7ddc585 plus the uncommitted fast-mode work.
  • Build: make -j4 PSF_BACKEND=cpu SPACETIME=minkowski backend (-O2 -DNDEBUG -fopenmp, libpng).
  • Host: 12th Gen Intel Core i7-12700K, 16 hardware threads.

Commands

All runs use the same catalog, geometry, and exposure; only the accumulation mode changes. --all-sky-catalog assets/2mass/processed/all_sky, --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0.

./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --output /tmp/fastmode/bench/normal.png

./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 2 \
  --output /tmp/fastmode/bench/fast_n2.png

./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 4 \
  --output /tmp/fastmode/bench/fast_n4.png

./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
  --width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
  --exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
  --fast-mode --fast-supersample 2 --fast-deposit bilinear \
  --output /tmp/fastmode/bench/fast_bil_n2.png

Wall-clock totals

Measured with date +%s.%N around each command.

run wall time images
normal (phase cache) 16.645 s 5,912,840
fast nearest N=2 1.663 s 5,912,840
fast bilinear N=2 1.648 s 5,912,840
fast nearest N=4 12.229 s 5,912,840

Fast nearest N=2 is about 10x faster than the phase-cache path on this PSF-dominated field. N=4 increases the supersampled buffer and kernel area by 4x each and becomes convolution-bound.

Raw terminal output

===== normal.log =====
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000056 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.725 s
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.696 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/normal.png (ok)
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
===== fast_n2.log =====
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000045 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.708 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/fast_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
===== fast_n4.log =====
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000058 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.709 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/fast_n4.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
===== fast_bil_n2.log =====
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000047 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit bilinear, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.712 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/fast_bil_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0

Tone-mapped difference vs the normal render

fast_n2     mean=1.1213 rms=1.7255 max=14 px>16=0/76800
fast_n4     mean=0.5516 rms=0.9146 max=8 px>16=0/76800
fast_bil_n2 mean=0.2091 rms=0.4793 max=16 px>16=0/76800

mean/rms are per-channel 8-bit differences; px>16 counts pixels with any channel more than 16 levels from the normal render. No pixel exceeds 16 in any run. These are display-scale differences from the measured deposition tradeoff; they are not a bit-exact regression baseline.