- Gather the union of every movie frame's all-sky tiles once and read them in bounded batches, replacing the per-frame prefetch scan and log.
- Fuse the fast supersampled-buffer clear into the FFTW pack pass so a resolved frame starts clean with no serial memset.
- Split parallel HDR->RGB8 tone mapping from PNG encoding.
- Add a bounded single-producer/single-writer movie output queue used by both observer movies and multi-frame imported lens maps, with writer timing and error propagation.
- Add --png-compression-level and --movie-output-workers shared|reserve-one.
- Add --fast-fftw-plan estimate|measure|wisdom|wisdom-update with strict wisdom identity sidecars.
- Add staged movie timing, regression tests, and docs.
Add a third analytic spacetime provider for the moving Alcubierre warp
bubble, x_s(t) = v_s*t with x_s(0) = 0. The lab slices stay flat, so
alpha = 1, gamma_ij = delta_ij, beta^x = -v_s f(r_s), and K_ij follows
from the flat spatial metric; the time dependence enters through the
moving shape argument. The exotic matter is treated as transparent, so
there is no capture: rays are only ACTIVE or ESCAPED, with a bubble-
centered escape radius R + 20/sigma.
Expose --alcubierre-vs, --alcubierre-radius, and --alcubierre-sigma
(|v_s| < 1). Scale the per-ray step budget with the escape radius and
1/(1-|v_s|) so near-luminal grazing rays still escape, and reject
parameter combinations whose worst-case budget exceeds the cap. Use a
cancellation-free shape formula for small sigma*R and reject derived
escape radii that overflow.
The regression test covers metric reconstruction, d_beta/K finite
differences, the translation isometry, small-sigma stability, the flat
limit, reflection symmetry, step convergence, and a near-luminal slow
ray. build.md, usage.md, and README.md document the backend.
Add an opt-in, post-processing limited-response model applied to the
finished linear HDR before tone mapping. Each RGB channel is processed
independently and isotropically: overflow above E spreads to the eight
neighbours with a fixed 9-point stencil, while the rest is absorbed or
lost at the image boundary. The synchronous ping-pong update uses a
monotonic bounding box and a row-parallel, deterministic reduction; the
conservative round bound reserves fp guard rounds inside a 4096 hard
limit and fails before touching HDR when exceeded.
Expose --sensor-bloom-limit E and --sensor-bloom-transfer e (both
required together, default disabled), validate them before expensive
initialization, and route every output path through the same hook in
write_frame_outputs: raw FITS first, bloom, tone-mapped PNG/PPM, then the
mesh overlay. The raw --hdr-output FITS therefore stays pre-bloom.
Add a standalone unit test (stencil, boundary loss, cascade reference,
symmetry, thread determinism, convergence limits, validation, allocation
failure), CLI integration and regression coverage, an isolated
sensor-bloom-bench target, and document the model in the design, usage,
README, and build docs.
Replace the per-channel Reinhard display transform with a parameterized
soft clip T_p(x) = tanh(x^p)^(1/p), default softclip p=2, exposed through
--tone-map and --tone-map-p. Keep --tone-map reinhard bit-compatible with
the previous x/(1+x) curve for existing images and reject combining it
with an explicit --tone-map-p.
Route the primary image and the mesh overlay through the same
ToneMapSettings; the linear HDR FITS writer stays pre-tone-map. Add a
focused optics-linked tone-map test target, CLI success/error coverage,
and document the display operator versus the Moffat effective PSF.
--draw-mesh no longer modifies the primary render. The clean HDR FITS and
tone-mapped image are written first, then the final image-plane mesh is
overlaid in place on the already-consumed HDR buffer and saved as a
<stem>_mesh sibling. This applies uniformly to single-frame, movie-frame,
and imported lens-map renders, with derived paths validated before catalog,
spacetime, and render initialization.
Dummy PSF backends still write nothing and report ignoring --draw-mesh. The
reference-image target now uses the configured image extension so the
ENABLE_PNG=0 test path passes, and the CLI tests cover clean-main identity,
sibling naming, PATH_MAX overflow, and zero-tolerance HDR comparison.
README, README.zh-CN, and usage.md describe the new naming; the mesh reference
asset is renamed to match.
Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README.
- Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2.
- Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile.
- Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.
- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
Add an optional CPU preview path that deposits each point-source image as a
supersampled delta and resolves the whole frame with one global Moffat
convolution plus an N x N box average, instead of splatting a per-event PSF.
- optics: FastPsfAccumulator builds the pixel-area-integral kernel of the
target Moffat at the supersampled scale (width N*alpha, same beta). The
1/N^2 box average then reproduces the final pixel-area integral, so the
requested FWHM and beta are preserved without renormalisation. Deposits are
per-cell atomic adds; resolve accumulates into the caller's HDR buffer.
- frame: fast branch in frame_splat_catalog with one shared supersampled
buffer and a single resolve per frame; the accumulator is reused across
movie frames and built from the map dimensions on lens-map import.
- main: --fast-mode, --fast-supersample N (1..8, default 2) and
--fast-deposit nearest|bilinear (default nearest). CPU-only and rejected in
the HIP/dummy backends; --psf-min-y still applies per event while
--max-cache-psf-flux does not.
- The deposition scheme was chosen by scripts/fast_mode_deposit_error.py:
nearest keeps the PSF shape exactly with <= 0.5/N px position quantization;
bilinear keeps the exact centroid but broadens FWHM and beta. Recorded in
benchmarks/fast_mode_deposit_2026-09-18.md.
- tests/test_frame.c covers fast nearest vs the direct evaluator at the snapped
centre, flux conservation, bilinear centroid, min-Y discard, frame plumbing,
and HDR accumulation onto a non-zero background.
- benchmarks/fast_mode_cpu_2026-09-18.md records a ~10x speedup on the 2MASS
galactic-centre field with small tone-mapped differences.