--draw-mesh no longer modifies the primary render. The clean HDR FITS and
tone-mapped image are written first, then the final image-plane mesh is
overlaid in place on the already-consumed HDR buffer and saved as a
<stem>_mesh sibling. This applies uniformly to single-frame, movie-frame,
and imported lens-map renders, with derived paths validated before catalog,
spacetime, and render initialization.
Dummy PSF backends still write nothing and report ignoring --draw-mesh. The
reference-image target now uses the configured image extension so the
ENABLE_PNG=0 test path passes, and the CLI tests cover clean-main identity,
sibling naming, PATH_MAX overflow, and zero-tolerance HDR comparison.
README, README.zh-CN, and usage.md describe the new naming; the mesh reference
asset is renamed to match.
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.
- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations.
Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
Allow frames to finish refinement independently without aborting the movie. Report genuine refinement errors with generation and frame identifiers, and clarify the sample batch contract.
Add a two-frame Schwarzschild regression covering unequal refinement progress. Validated Release and Debug make test suites and 640x360 rendering of the final eight free-fall frames in both builds.