Feat: Overlap movie output and cut per-frame fast-mode work

- Gather the union of every movie frame's all-sky tiles once and read them in bounded batches, replacing the per-frame prefetch scan and log.
- Fuse the fast supersampled-buffer clear into the FFTW pack pass so a resolved frame starts clean with no serial memset.
- Split parallel HDR->RGB8 tone mapping from PNG encoding.
- Add a bounded single-producer/single-writer movie output queue used by both observer movies and multi-frame imported lens maps, with writer timing and error propagation.
- Add --png-compression-level and --movie-output-workers shared|reserve-one.
- Add --fast-fftw-plan estimate|measure|wisdom|wisdom-update with strict wisdom identity sidecars.
- Add staged movie timing, regression tests, and docs.
This commit is contained in:
wyj committed 2026-10-03 19:32:35 -04:00
1 parent a510fec23e
commit 04611e3e5a
20 files changed
+2282 -212

No files matched your search

+53 -1
View File
@@ -168,6 +168,28 @@ shifts change the brightness. The trajectory generator is a flat-spacetime
example; other camera motions are supplied through observer-track CSVs.
Output is an image sequence; video encoding is a separate step.
An all-sky movie prefetches the union of every finalized frame's requested
one-degree tiles exactly once (one `Movie catalog prefetch: ...` line) and then
renders every frame from the immutable cache, instead of re-running the
per-frame prefetch scan and log. A single-frame render keeps its per-frame
prefetch. The prefetch reads at most 512 unseen tiles per parallel batch, so an
unbounded all-sky union never stages an unbounded temporary star copy.
Movie PNG encoding and writing run on one bounded writer thread and overlap
the next frame's ray/catalog/FFTW work. The default bounds two queued frames
plus the one the writer is currently processing (at most three jobs in memory).
The same queue also serves multi-frame imported `.grlens` maps. `--png-compression-level N`
sets the zlib/libpng level in `0..9`; the default is libpng's own default.
`--fast-mode` never changes it silently, so a preview movie can opt into
`--png-compression-level 1` explicitly. `--movie-output-workers shared`
(default) lets the writer share the machine with the render pool.
`--movie-output-workers reserve-one` shrinks the render OpenMP pool by one core
for the whole movie render (ray tracing, catalog work, and the fast-mode FFTW
plans all see the reduced count); it is applied before the fast accumulator is
built so the FFTW plans actually use the reduced worker count, not only while
the writer is active. A writer error is recorded and fails the render rather
than silently dropping frames.
Movie rays from all frames share a newest-to-oldest coordinate-time sweep.
`--slab-duration` sets its time-slab width (default `64`). The current analytic
backends use logical slabs without metric I/O; [Nmesh](https://github.com/nmeshsource/nmesh) metric loading remains
@@ -462,6 +484,19 @@ substantially but costs a much longer one-time plan for large frames, as
recorded in
[benchmarks/fast_mode_fftw_2026-09-25.md](benchmarks/fast_mode_fftw_2026-09-25.md).
Planning can be selected explicitly. `--fast-fftw-plan measure` uses
`FFTW_MEASURE`. `--fast-fftw-plan wisdom-update --fast-fftw-wisdom FILE`
measures once and writes reusable wisdom plus a `FILE.meta` sidecar recording
FFTW version, precision, supersample, FFT dimensions, kernel radius, and worker
count. Each file is written to a unique temporary and renamed into place
individually; the pair is not a transactional update, and the sidecar is
committed last, so a wisdom without a matching sidecar is treated as a miss.
`--fast-fftw-plan wisdom --fast-fftw-wisdom FILE` imports
that wisdom and creates the plans with `FFTW_WISDOM_ONLY`; if the file, size,
supersample, worker count, version, or precision does not match, it fails with a
diagnostic instead of silently replanning. Wisdom writes go to a temporary file
and are renamed into place, so a failed write never leaves a half file.
## Progress and diagnostics
The PSF-cache completion line is printed before tracing and catalog splatting
@@ -473,7 +508,24 @@ Verbose ray-trace output reports the initial mesh trace and refinement stages
for single frames; movie mode additionally reports each generation's sample
count and each time slab's activation and terminal-ray summary.
Movie renders always print one summary per time slab; `--verbose` also prints
the ray counts before each slab is loaded.
the ray counts before each slab is loaded. In movie mode, `--verbose` adds one
`Movie frame N timing:` line per frame and a `Movie timing total/avg/max`
summary. The stages are `catalog_mark`, `catalog_load`, `fast_clear`, `splat`,
the FFTW breakdown (`fftw_zero_pack`, `fftw_forward`, `fftw_multiply`,
`fftw_inverse`, `fftw_crop`, `fftw_total`), `tone_map`, `output`, and
`wait` (writer-queue backpressure), plus the frame total. `output` covers only
producer-side writes, so it is zero in the async movie path; the writer's own
time is reported separately by `Movie writer summary: jobs/total/avg/max/drain`.
`Movie output pipeline wall:` starts after track load, ray tracing, lens-map
write, and the catalog union prefetch, and covers per-frame render plus async
output. `Movie end-to-end wall:` starts at `render_movie()` entry, so it
additionally includes track load, all ray tracing, the union prefetch, and the
final writer drain; it does not include the catalog, blackbody, and fast
accumulator initialization done earlier in `main()`. `Movie timing total` is
the sum of producer frame times only (it excludes tracing, prefetch, and the
final queue drain). The all-sky `Movie catalog prefetch:` line reports mark,
load+commit, and total tile time. Timing uses one clock read per bulk phase,
never inside the per-star or per-pixel hot loops.
Pass `--draw-mesh` to also write the final image-plane triangle mesh as a
`<output-stem>_mesh.png` sibling (`.ppm` in non-PNG builds). The main