- Gather the union of every movie frame's all-sky tiles once and read them in bounded batches, replacing the per-frame prefetch scan and log.
- Fuse the fast supersampled-buffer clear into the FFTW pack pass so a resolved frame starts clean with no serial memset.
- Split parallel HDR->RGB8 tone mapping from PNG encoding.
- Add a bounded single-producer/single-writer movie output queue used by both observer movies and multi-frame imported lens maps, with writer timing and error propagation.
- Add --png-compression-level and --movie-output-workers shared|reserve-one.
- Add --fast-fftw-plan estimate|measure|wisdom|wisdom-update with strict wisdom identity sidecars.
- Add staged movie timing, regression tests, and docs.
Add a third analytic spacetime provider for the moving Alcubierre warp
bubble, x_s(t) = v_s*t with x_s(0) = 0. The lab slices stay flat, so
alpha = 1, gamma_ij = delta_ij, beta^x = -v_s f(r_s), and K_ij follows
from the flat spatial metric; the time dependence enters through the
moving shape argument. The exotic matter is treated as transparent, so
there is no capture: rays are only ACTIVE or ESCAPED, with a bubble-
centered escape radius R + 20/sigma.
Expose --alcubierre-vs, --alcubierre-radius, and --alcubierre-sigma
(|v_s| < 1). Scale the per-ray step budget with the escape radius and
1/(1-|v_s|) so near-luminal grazing rays still escape, and reject
parameter combinations whose worst-case budget exceeds the cap. Use a
cancellation-free shape formula for small sigma*R and reject derived
escape radii that overflow.
The regression test covers metric reconstruction, d_beta/K finite
differences, the translation isometry, small-sigma stability, the flat
limit, reflection symmetry, step convergence, and a near-luminal slow
ray. build.md, usage.md, and README.md document the backend.
Add an opt-in, post-processing limited-response model applied to the
finished linear HDR before tone mapping. Each RGB channel is processed
independently and isotropically: overflow above E spreads to the eight
neighbours with a fixed 9-point stencil, while the rest is absorbed or
lost at the image boundary. The synchronous ping-pong update uses a
monotonic bounding box and a row-parallel, deterministic reduction; the
conservative round bound reserves fp guard rounds inside a 4096 hard
limit and fails before touching HDR when exceeded.
Expose --sensor-bloom-limit E and --sensor-bloom-transfer e (both
required together, default disabled), validate them before expensive
initialization, and route every output path through the same hook in
write_frame_outputs: raw FITS first, bloom, tone-mapped PNG/PPM, then the
mesh overlay. The raw --hdr-output FITS therefore stays pre-bloom.
Add a standalone unit test (stencil, boundary loss, cascade reference,
symmetry, thread determinism, convergence limits, validation, allocation
failure), CLI integration and regression coverage, an isolated
sensor-bloom-bench target, and document the model in the design, usage,
README, and build docs.
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.
- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
Reuse production catalog mapping and per-worker chunk boundaries to report triangle and spatial distributions without initializing HIP or writing image outputs. Include a protected full-catalog runner and build documentation.
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.
Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.