Files
GR-raytracing/benchmarks/ray_pool_scheduling_2026-09-06.md
T
wyj 8741597d9d Ray: Balance movie tracing with dynamic chunks
Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations.

Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
2026-09-06 15:56:37 -04:00

141 lines
7.5 KiB
Markdown

# Movie ray-pool scheduling benchmark (2026-09-06)
## Change and scope
Movie samples are appended in frame order. `schedule(static)` without a chunk
size assigns each worker one contiguous range of the entire pool, including
pending and terminated rays. A slab can therefore leave most workers with no
active rays even though there is substantial tracing work.
The production change is `schedule(dynamic, 32)` in
`ray_pool_advance_active()`. Each worker advances a chunk of up to 32 pool slots
and then takes another chunk. Pool slots and frame/sample IDs remain stable;
there are no new allocations, caches, tasks, or changes to the geodesic solver.
The parallel-region barrier still completes the slab before results are read.
Activation, status scans, and mesh analysis remain serial. The full pool is
still scanned; active-index compaction is not part of this change.
The 32-slot chunk is an empirical choice for this workload, not a universal
optimum. Cyclic static chunks perform similarly here. Dynamic chunks retain
load balancing when strong-field integration costs differ, while avoiding
per-ray scheduling.
## Dataset and build
- Baseline: `8cf6106127fc19732cee123e9890d401a6838a51`, including the fix for
frames with empty refinement generations.
- CPU: Intel Core i7-12700K; 16 online logical CPUs, 16 OpenMP workers.
- C compiler and full CPU details are saved in each `*-environment.json`.
- CPU PSF backend, event sink enabled, PNG output; C11, `-march=native -O2`,
`-DNDEBUG`, OpenMP. Both versions are built from baseline sources, with the
candidate substituting only the current `src/ray.c`.
- Input: all 393 samples of the reported inward free-fall trajectory, preserved
verbatim in [camera.csv](ray_pool_scheduling_2026-09-06/camera.csv).
- `assets/sky_grid_5deg.csv`: 5654 catalog stars. Input hashes are recorded.
- 64x36, 60-degree horizontal FOV, coarse cell 16 pixels, refinement level 3,
exposure 0.2. The temporal span and original 16:9 aspect ratio are retained.
- Every run writes all 393 PNGs and a complete movie lens map. This is a
reduced-resolution benchmark, not the original 640x360 run. Do not extrapolate
its speedup to the full-resolution movie or numerical backends.
- Runs were sequential; build/test jobs were not run concurrently with timing.
`OMP_NUM_THREADS=16`, `OMP_DYNAMIC=FALSE`. PSF cache build and image output are
included in end-to-end wall time. No cache files are reused by these commands.
Filesystem caches were not flushed; this is not a cold-I/O benchmark.
## Schedule selection
A temporary baseline source copy uses `schedule(runtime)` so `OMP_SCHEDULE`
selects the experiment. Each worker records active-ray visits, integrator
steps, and loop elapsed time without the final barrier; batch wall time includes
the barrier. Counters are collected in worker-owned slots outside solver state.
These counters and the runtime schedule override are **not in production**.
| Schedule | First run wall (s) | Reverse-order repeat wall (s) | First run ray-batch time (s) |
| --- | ---: | ---: | ---: |
| contiguous static | 61.040 | 61.093 | 55.688 |
| cyclic static,32 | 39.540 | 40.119 | 34.101 |
| dynamic,32 | 39.292 | 38.679 | 33.808 |
| dynamic,128 | 39.792 | not repeated | 34.316 |
Every instrumented run executes 47 slab advances, 402534 active-ray visits and
206670193 integrator steps. The first slab has 1856 active rays: contiguous
static assigns them to only 3 workers, while all chunked schedules use 16.
The small difference between static,32 and dynamic,32 is not evidence that one
is universally faster. The major improvement is removing large contiguous
frame ranges from each worker's assignment.
User/system CPU seconds are preserved in `*-results.json`. These include OpenMP
waiting and should not be interpreted as useful solver work alone. Exact
per-worker counts and times remain in the raw logs, and aggregate counters are
in [instrumentation-summary.json](ray_pool_scheduling_2026-09-06/instrumentation-summary.json).
## Reproduction and original output
Run from the repository root:
```sh
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
--tag repeat --schedules dynamic,32 static,32 static
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
--production --tag production-before --schedules static
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
--production --tag production-after --schedules dynamic,32
```
[run.py](ray_pool_scheduling_2026-09-06/run.py) builds temporary executables under
`/tmp/gr-ray-schedules`, and stores original output and metadata beside the
script. Use a fresh `--tag` and `--work-dir` to preserve an existing run. The
production comparisons compile literal scheduling directives and have no
instrumentation; their logged `OMP_SCHEDULE` value does not override those
compiled directives. Production-after requires the candidate `src/ray.c`.
Every run has a `*-command.txt` containing the complete, shell-copyable renderer
command, `*-results.json` with exit status, wall/user/system time and artifact
hash, and an unfiltered `*.log.gz` preserving all stdout/stderr bytes, including
PSF build/cache, worker, processed-image and fallback statistics. Build command
files and build stdout/stderr logs are also preserved; successful silent
compiler output can produce an empty build log. For example:
```sh
gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-0-static.log.gz
gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-2-dynamic-32.log.gz
```
Output PNGs and lens maps remain under the run's work directory; they are not
stored in Git. The reference lens-map SHA-256 is
`708a218eca26b9eac42cb4954b0dbd4dfe316cc55d4315d650cbc8ab7e616080`.
All schedule-selection runs match it byte-for-byte, including the geometry,
endpoint statuses, frequency shifts, frame metadata and mesh topology.
## Uninstrumented production-source comparison
| Source | Wall (s) | User CPU (s) | System CPU (s) |
| --- | ---: | ---: | ---: |
| Baseline contiguous static | 62.126 | 411.940 | 1.964 |
| Candidate dynamic,32 | 38.686 | 554.787 | 2.748 |
The end-to-end speedup is 1.606x (37.7% less wall time). The CPU-seconds / wall
ratio increases from 6.66 to 14.41, including runtime waiting. This comparison
has no per-worker diagnostic counters and no runtime-selected schedule. It is
one pair of runs, supported by the reverse-order instrumented checks above;
there is no statistical confidence interval or claim about all workloads.
Both runs exit 0 and write 393 PNGs. Every PNG is byte-identical before/after;
all candidate PNGs passed signature, dimension, chunk CRC and IDAT decompression
checks. All nine measured runs have the same complete lens-map hash.
See [artifact-validation.json](ray_pool_scheduling_2026-09-06/artifact-validation.json),
[baseline results](ray_pool_scheduling_2026-09-06/production-before-results.json),
[candidate results](ray_pool_scheduling_2026-09-06/production-after-results.json),
and their corresponding unfiltered compressed logs in the same directory.
## Regression verification
`make -j8 BUILD_TYPE=Release all` and `make -j8 BUILD_TYPE=Debug all` rebuilt the
actual production binaries. Both corresponding `make test` suites passed,
including a Schwarzschild movie regression that compares complete lens maps
with 1, 4 and 16 requested OpenMP threads. `git diff --check` passed. Exact
commands and compressed original build/test output are preserved in
`verification-commands.txt`, `release-build.log.gz`, `release-tests.log.gz`,
`debug-build.log.gz` and `debug-tests.log.gz` beside the benchmark data.