Feat: Overlap movie output and cut per-frame fast-mode work
- Gather the union of every movie frame's all-sky tiles once and read them in bounded batches, replacing the per-frame prefetch scan and log. - Fuse the fast supersampled-buffer clear into the FFTW pack pass so a resolved frame starts clean with no serial memset. - Split parallel HDR->RGB8 tone mapping from PNG encoding. - Add a bounded single-producer/single-writer movie output queue used by both observer movies and multi-frame imported lens maps, with writer timing and error propagation. - Add --png-compression-level and --movie-output-workers shared|reserve-one. - Add --fast-fftw-plan estimate|measure|wisdom|wisdom-update with strict wisdom identity sidecars. - Add staged movie timing, regression tests, and docs.
This commit is contained in:
1 parent
a510fec23e
commit
04611e3e5a
20 files changed
+2282
-212
No files matched your search
@@ -168,6 +168,28 @@ shifts change the brightness. The trajectory generator is a flat-spacetime
|
||||
example; other camera motions are supplied through observer-track CSVs.
|
||||
Output is an image sequence; video encoding is a separate step.
|
||||
|
||||
An all-sky movie prefetches the union of every finalized frame's requested
|
||||
one-degree tiles exactly once (one `Movie catalog prefetch: ...` line) and then
|
||||
renders every frame from the immutable cache, instead of re-running the
|
||||
per-frame prefetch scan and log. A single-frame render keeps its per-frame
|
||||
prefetch. The prefetch reads at most 512 unseen tiles per parallel batch, so an
|
||||
unbounded all-sky union never stages an unbounded temporary star copy.
|
||||
|
||||
Movie PNG encoding and writing run on one bounded writer thread and overlap
|
||||
the next frame's ray/catalog/FFTW work. The default bounds two queued frames
|
||||
plus the one the writer is currently processing (at most three jobs in memory).
|
||||
The same queue also serves multi-frame imported `.grlens` maps. `--png-compression-level N`
|
||||
sets the zlib/libpng level in `0..9`; the default is libpng's own default.
|
||||
`--fast-mode` never changes it silently, so a preview movie can opt into
|
||||
`--png-compression-level 1` explicitly. `--movie-output-workers shared`
|
||||
(default) lets the writer share the machine with the render pool.
|
||||
`--movie-output-workers reserve-one` shrinks the render OpenMP pool by one core
|
||||
for the whole movie render (ray tracing, catalog work, and the fast-mode FFTW
|
||||
plans all see the reduced count); it is applied before the fast accumulator is
|
||||
built so the FFTW plans actually use the reduced worker count, not only while
|
||||
the writer is active. A writer error is recorded and fails the render rather
|
||||
than silently dropping frames.
|
||||
|
||||
Movie rays from all frames share a newest-to-oldest coordinate-time sweep.
|
||||
`--slab-duration` sets its time-slab width (default `64`). The current analytic
|
||||
backends use logical slabs without metric I/O; [Nmesh](https://github.com/nmeshsource/nmesh) metric loading remains
|
||||
@@ -462,6 +484,19 @@ substantially but costs a much longer one-time plan for large frames, as
|
||||
recorded in
|
||||
[benchmarks/fast_mode_fftw_2026-09-25.md](benchmarks/fast_mode_fftw_2026-09-25.md).
|
||||
|
||||
Planning can be selected explicitly. `--fast-fftw-plan measure` uses
|
||||
`FFTW_MEASURE`. `--fast-fftw-plan wisdom-update --fast-fftw-wisdom FILE`
|
||||
measures once and writes reusable wisdom plus a `FILE.meta` sidecar recording
|
||||
FFTW version, precision, supersample, FFT dimensions, kernel radius, and worker
|
||||
count. Each file is written to a unique temporary and renamed into place
|
||||
individually; the pair is not a transactional update, and the sidecar is
|
||||
committed last, so a wisdom without a matching sidecar is treated as a miss.
|
||||
`--fast-fftw-plan wisdom --fast-fftw-wisdom FILE` imports
|
||||
that wisdom and creates the plans with `FFTW_WISDOM_ONLY`; if the file, size,
|
||||
supersample, worker count, version, or precision does not match, it fails with a
|
||||
diagnostic instead of silently replanning. Wisdom writes go to a temporary file
|
||||
and are renamed into place, so a failed write never leaves a half file.
|
||||
|
||||
## Progress and diagnostics
|
||||
|
||||
The PSF-cache completion line is printed before tracing and catalog splatting
|
||||
@@ -473,7 +508,24 @@ Verbose ray-trace output reports the initial mesh trace and refinement stages
|
||||
for single frames; movie mode additionally reports each generation's sample
|
||||
count and each time slab's activation and terminal-ray summary.
|
||||
Movie renders always print one summary per time slab; `--verbose` also prints
|
||||
the ray counts before each slab is loaded.
|
||||
the ray counts before each slab is loaded. In movie mode, `--verbose` adds one
|
||||
`Movie frame N timing:` line per frame and a `Movie timing total/avg/max`
|
||||
summary. The stages are `catalog_mark`, `catalog_load`, `fast_clear`, `splat`,
|
||||
the FFTW breakdown (`fftw_zero_pack`, `fftw_forward`, `fftw_multiply`,
|
||||
`fftw_inverse`, `fftw_crop`, `fftw_total`), `tone_map`, `output`, and
|
||||
`wait` (writer-queue backpressure), plus the frame total. `output` covers only
|
||||
producer-side writes, so it is zero in the async movie path; the writer's own
|
||||
time is reported separately by `Movie writer summary: jobs/total/avg/max/drain`.
|
||||
`Movie output pipeline wall:` starts after track load, ray tracing, lens-map
|
||||
write, and the catalog union prefetch, and covers per-frame render plus async
|
||||
output. `Movie end-to-end wall:` starts at `render_movie()` entry, so it
|
||||
additionally includes track load, all ray tracing, the union prefetch, and the
|
||||
final writer drain; it does not include the catalog, blackbody, and fast
|
||||
accumulator initialization done earlier in `main()`. `Movie timing total` is
|
||||
the sum of producer frame times only (it excludes tracing, prefetch, and the
|
||||
final queue drain). The all-sky `Movie catalog prefetch:` line reports mark,
|
||||
load+commit, and total tile time. Timing uses one clock read per bulk phase,
|
||||
never inside the per-star or per-pixel hot loops.
|
||||
|
||||
Pass `--draw-mesh` to also write the final image-plane triangle mesh as a
|
||||
`<output-stem>_mesh.png` sibling (`.ppm` in non-PNG builds). The main
|
||||
|
||||
Reference in new issue
Block a user