# Movie ray-pool scheduling benchmark (2026-09-06) ## Change and scope Movie samples are appended in frame order. `schedule(static)` without a chunk size assigns each worker one contiguous range of the entire pool, including pending and terminated rays. A slab can therefore leave most workers with no active rays even though there is substantial tracing work. The production change is `schedule(dynamic, 32)` in `ray_pool_advance_active()`. Each worker advances a chunk of up to 32 pool slots and then takes another chunk. Pool slots and frame/sample IDs remain stable; there are no new allocations, caches, tasks, or changes to the geodesic solver. The parallel-region barrier still completes the slab before results are read. Activation, status scans, and mesh analysis remain serial. The full pool is still scanned; active-index compaction is not part of this change. The 32-slot chunk is an empirical choice for this workload, not a universal optimum. Cyclic static chunks perform similarly here. Dynamic chunks retain load balancing when strong-field integration costs differ, while avoiding per-ray scheduling. ## Dataset and build - Baseline: `8cf6106127fc19732cee123e9890d401a6838a51`, including the fix for frames with empty refinement generations. - CPU: Intel Core i7-12700K; 16 online logical CPUs, 16 OpenMP workers. - C compiler and full CPU details are saved in each `*-environment.json`. - CPU PSF backend, event sink enabled, PNG output; C11, `-march=native -O2`, `-DNDEBUG`, OpenMP. Both versions are built from baseline sources, with the candidate substituting only the current `src/ray.c`. - Input: all 393 samples of the reported inward free-fall trajectory, preserved verbatim in [camera.csv](ray_pool_scheduling_2026-09-06/camera.csv). - `assets/sky_grid_5deg.csv`: 5654 catalog stars. Input hashes are recorded. - 64x36, 60-degree horizontal FOV, coarse cell 16 pixels, refinement level 3, exposure 0.2. The temporal span and original 16:9 aspect ratio are retained. - Every run writes all 393 PNGs and a complete movie lens map. This is a reduced-resolution benchmark, not the original 640x360 run. Do not extrapolate its speedup to the full-resolution movie or numerical backends. - Runs were sequential; build/test jobs were not run concurrently with timing. `OMP_NUM_THREADS=16`, `OMP_DYNAMIC=FALSE`. PSF cache build and image output are included in end-to-end wall time. No cache files are reused by these commands. Filesystem caches were not flushed; this is not a cold-I/O benchmark. ## Schedule selection A temporary baseline source copy uses `schedule(runtime)` so `OMP_SCHEDULE` selects the experiment. Each worker records active-ray visits, integrator steps, and loop elapsed time without the final barrier; batch wall time includes the barrier. Counters are collected in worker-owned slots outside solver state. These counters and the runtime schedule override are **not in production**. | Schedule | First run wall (s) | Reverse-order repeat wall (s) | First run ray-batch time (s) | | --- | ---: | ---: | ---: | | contiguous static | 61.040 | 61.093 | 55.688 | | cyclic static,32 | 39.540 | 40.119 | 34.101 | | dynamic,32 | 39.292 | 38.679 | 33.808 | | dynamic,128 | 39.792 | not repeated | 34.316 | Every instrumented run executes 47 slab advances, 402534 active-ray visits and 206670193 integrator steps. The first slab has 1856 active rays: contiguous static assigns them to only 3 workers, while all chunked schedules use 16. The small difference between static,32 and dynamic,32 is not evidence that one is universally faster. The major improvement is removing large contiguous frame ranges from each worker's assignment. User/system CPU seconds are preserved in `*-results.json`. These include OpenMP waiting and should not be interpreted as useful solver work alone. Exact per-worker counts and times remain in the raw logs, and aggregate counters are in [instrumentation-summary.json](ray_pool_scheduling_2026-09-06/instrumentation-summary.json). ## Reproduction and original output Run from the repository root: ```sh python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \ --tag repeat --schedules dynamic,32 static,32 static python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \ --production --tag production-before --schedules static python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \ --production --tag production-after --schedules dynamic,32 ``` [run.py](ray_pool_scheduling_2026-09-06/run.py) builds temporary executables under `/tmp/gr-ray-schedules`, and stores original output and metadata beside the script. Use a fresh `--tag` and `--work-dir` to preserve an existing run. The production comparisons compile literal scheduling directives and have no instrumentation; their logged `OMP_SCHEDULE` value does not override those compiled directives. Production-after requires the candidate `src/ray.c`. Every run has a `*-command.txt` containing the complete, shell-copyable renderer command, `*-results.json` with exit status, wall/user/system time and artifact hash, and an unfiltered `*.log.gz` preserving all stdout/stderr bytes, including PSF build/cache, worker, processed-image and fallback statistics. Build command files and build stdout/stderr logs are also preserved; successful silent compiler output can produce an empty build log. For example: ```sh gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-0-static.log.gz gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-2-dynamic-32.log.gz ``` Output PNGs and lens maps remain under the run's work directory; they are not stored in Git. The reference lens-map SHA-256 is `708a218eca26b9eac42cb4954b0dbd4dfe316cc55d4315d650cbc8ab7e616080`. All schedule-selection runs match it byte-for-byte, including the geometry, endpoint statuses, frequency shifts, frame metadata and mesh topology. ## Uninstrumented production-source comparison | Source | Wall (s) | User CPU (s) | System CPU (s) | | --- | ---: | ---: | ---: | | Baseline contiguous static | 62.126 | 411.940 | 1.964 | | Candidate dynamic,32 | 38.686 | 554.787 | 2.748 | The end-to-end speedup is 1.606x (37.7% less wall time). The CPU-seconds / wall ratio increases from 6.66 to 14.41, including runtime waiting. This comparison has no per-worker diagnostic counters and no runtime-selected schedule. It is one pair of runs, supported by the reverse-order instrumented checks above; there is no statistical confidence interval or claim about all workloads. Both runs exit 0 and write 393 PNGs. Every PNG is byte-identical before/after; all candidate PNGs passed signature, dimension, chunk CRC and IDAT decompression checks. All nine measured runs have the same complete lens-map hash. See [artifact-validation.json](ray_pool_scheduling_2026-09-06/artifact-validation.json), [baseline results](ray_pool_scheduling_2026-09-06/production-before-results.json), [candidate results](ray_pool_scheduling_2026-09-06/production-after-results.json), and their corresponding unfiltered compressed logs in the same directory. ## Regression verification `make -j8 BUILD_TYPE=Release all` and `make -j8 BUILD_TYPE=Debug all` rebuilt the actual production binaries. Both corresponding `make test` suites passed, including a Schwarzschild movie regression that compares complete lens maps with 1, 4 and 16 requested OpenMP threads. `git diff --check` passed. Exact commands and compressed original build/test output are preserved in `verification-commands.txt`, `release-build.log.gz`, `release-tests.log.gz`, `debug-build.log.gz` and `debug-tests.log.gz` beside the benchmark data.