Files
GR-raytracing/benchmarks/fast_mode_cpu_2026-09-18.md
wyj 3deebfb2fa Feat: Add --fast-mode supersampled point-source accumulation
Add an optional CPU preview path that deposits each point-source image as a
supersampled delta and resolves the whole frame with one global Moffat
convolution plus an N x N box average, instead of splatting a per-event PSF.

- optics: FastPsfAccumulator builds the pixel-area-integral kernel of the
  target Moffat at the supersampled scale (width N*alpha, same beta).  The
  1/N^2 box average then reproduces the final pixel-area integral, so the
  requested FWHM and beta are preserved without renormalisation.  Deposits are
  per-cell atomic adds; resolve accumulates into the caller's HDR buffer.
- frame: fast branch in frame_splat_catalog with one shared supersampled
  buffer and a single resolve per frame; the accumulator is reused across
  movie frames and built from the map dimensions on lens-map import.
- main: --fast-mode, --fast-supersample N (1..8, default 2) and
  --fast-deposit nearest|bilinear (default nearest).  CPU-only and rejected in
  the HIP/dummy backends; --psf-min-y still applies per event while
  --max-cache-psf-flux does not.
- The deposition scheme was chosen by scripts/fast_mode_deposit_error.py:
  nearest keeps the PSF shape exactly with <= 0.5/N px position quantization;
  bilinear keeps the exact centroid but broadens FWHM and beta.  Recorded in
  benchmarks/fast_mode_deposit_2026-09-18.md.
- tests/test_frame.c covers fast nearest vs the direct evaluator at the snapped
  centre, flux conservation, bilinear centroid, min-Y discard, frame plumbing,
  and HDR accumulation onto a non-zero background.
- benchmarks/fast_mode_cpu_2026-09-18.md records a ~10x speedup on the 2MASS
  galactic-centre field with small tone-mapped differences.
2026-09-24 01:36:25 -04:00

109 lines
5.8 KiB
Markdown

# Fast-mode CPU benchmark (2026-09-18)
Compares the default per-event phase-cache PSF splat with `--fast-mode` on a
7.4-degree 2MASS galactic-center field. Commands, raw terminal output, and
wall-clock totals are preserved below.
## Environment
- Git revision: `7ddc585` plus the uncommitted fast-mode work.
- Build: `make -j4 PSF_BACKEND=cpu SPACETIME=minkowski backend`
(`-O2 -DNDEBUG -fopenmp`, libpng).
- Host: 12th Gen Intel Core i7-12700K, 16 hardware threads.
## Commands
All runs use the same catalog, geometry, and exposure; only the accumulation
mode changes. `--all-sky-catalog assets/2mass/processed/all_sky`,
`--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0`.
```sh
./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--output /tmp/fastmode/bench/normal.png
./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 2 \
--output /tmp/fastmode/bench/fast_n2.png
./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 4 \
--output /tmp/fastmode/bench/fast_n4.png
./build/Release/minkowski_sky --all-sky-catalog assets/2mass/processed/all_sky \
--width 320 --height 240 --fov-deg 10 --look-ra-deg 266 --look-dec-deg -29 \
--exposure 1e13 --coarse-cell-pixels 16 --refine-max-level 0 \
--fast-mode --fast-supersample 2 --fast-deposit bilinear \
--output /tmp/fastmode/bench/fast_bil_n2.png
```
## Wall-clock totals
Measured with `date +%s.%N` around each command.
| run | wall time | images |
| --- | --- | --- |
| normal (phase cache) | 16.645 s | 5,912,840 |
| fast nearest N=2 | 1.663 s | 5,912,840 |
| fast bilinear N=2 | 1.648 s | 5,912,840 |
| fast nearest N=4 | 12.229 s | 5,912,840 |
Fast nearest N=2 is about 10x faster than the phase-cache path on this
PSF-dominated field. N=4 increases the supersampled buffer and kernel area by
4x each and becomes convolution-bound.
## Raw terminal output
```text
===== normal.log =====
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000056 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.725 s
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.696 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/normal.png (ok)
PSF splats: cached 5912840, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
===== fast_n2.log =====
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000045 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit nearest, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.708 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/fast_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
===== fast_n4.log =====
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000058 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 4x, deposit nearest, kernel radius 185 ss px (46.25 final px), retained flux 1.000000, cached 4-point quadrature
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.709 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/fast_n4.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
===== fast_bil_n2.log =====
Fast mode is a preview approximation; --max-cache-psf-flux is ignored and one global kernel is used for every event.
Blackbody LUT: 1024 CIE 1931 2-deg 1 nm-linear XYZ nodes, T=[670.146556, 101408.88] K, loaded assets/blackbody/cie1931_2deg_xyz_1024.grbblut in 0.000047 s; payload fnv1a64=0e39867b0e70a809
Blackbody backend: lut
Fast PSF: supersample 2x, deposit bilinear, kernel radius 93 ss px (46.50 final px), retained flux 1.000000, cached 4-point quadrature
Catalog prefetch: 96 requested, 96 newly loaded (6618807 stars), 0 unavailable in 0.712 s; 4 loader workers
Rendered 5912840 images from 6618807 catalog stars to /tmp/fastmode/bench/fast_bil_n2.png (ok)
Fast PSF splats: deposited 5912840, wing-clipped 0, discarded below min-Y 0
```
## Tone-mapped difference vs the normal render
```text
fast_n2 mean=1.1213 rms=1.7255 max=14 px>16=0/76800
fast_n4 mean=0.5516 rms=0.9146 max=8 px>16=0/76800
fast_bil_n2 mean=0.2091 rms=0.4793 max=16 px>16=0/76800
```
`mean`/`rms` are per-channel 8-bit differences; `px>16` counts pixels with any
channel more than 16 levels from the normal render. No pixel exceeds 16 in any
run. These are display-scale differences from the measured deposition tradeoff;
they are not a bit-exact regression baseline.