Commit Graph
17 Commits
Author SHA1 Message Date
wyj 229f50cd86 Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.

- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
2026-09-25 23:35:45 -04:00
wyj d02ecb2181 Feat: Add dummy PSF chunk diagnostics
Reuse production catalog mapping and per-worker chunk boundaries to report triangle and spatial distributions without initializing HIP or writing image outputs. Include a protected full-catalog runner and build documentation.
2026-09-11 23:27:42 -04:00
wyj 7f5ec1a488 Optics: add bounded HIP PSF replay and tile reduction experiments
Capture prepared events through a test-only producer consumer and measure serial and indexed parallel generation independently. Compare the production atomic kernel with bounded pixel-owned tile reductions while preserving double HDR, cache support, and direct fallback ordering.

Validate captured inputs, edge and multichunk fixtures, CPU/HIP renderer regressions, and 45 isolated final replays. Keep the production HIP algorithm unchanged because gains depend on event distribution.
2026-09-11 16:43:53 -04:00
wyj e1ec480669 HIP: accelerate PSF accumulation and restore parallel producers
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.

Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
2026-09-06 21:38:53 -04:00
wyj 6ae223a642 Feat: support general single-frame cameras and add a near-horizon example 2026-09-06 04:00:21 -04:00
wyj 7a80086f63 Feat: add initial HIP PSF backend 2026-09-05 04:23:57 -04:00
wyj 2a03cbf724 Frame: buffer cached PSF events per worker 2026-09-05 03:06:46 -04:00
wyj 4953775cdd Test: add fixed PSF image references 2026-09-04 20:28:54 -04:00
wyj 19f3a91b7f Makefile: build all backends with optional HDR output 2026-08-30 02:29:41 -04:00
wyj 6f8ec50c64 Makefile: add Debug build diagnostics 2026-08-29 19:23:50 -04:00
wyj d5d21c7659 Output: write HDR FITS with camera metadata 2026-08-28 19:56:07 -04:00
wyj 80b891d2c8 Optics: cache pixel-area Moffat kernels 2026-08-28 13:53:06 -04:00
wyj 5d80bb1b51 Makefile: tune default CFLAGS 2026-08-27 13:04:36 -04:00
wyj b0fdf1faca Output: default to PNG with PPM fallback 2026-08-27 04:25:45 -04:00
wyj 457c0728b4 Catalog: parallelize all-sky tile prefetch 2026-08-27 03:53:45 -04:00
wyj c8b321a62a Add observer track PNG movie mode 2026-08-25 22:56:10 -04:00
wyj d30f9ac56c init 2026-08-25 22:45:43 -04:00