Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations. Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
141 lines
7.5 KiB
Markdown
141 lines
7.5 KiB
Markdown
# Movie ray-pool scheduling benchmark (2026-09-06)
|
|
|
|
## Change and scope
|
|
|
|
Movie samples are appended in frame order. `schedule(static)` without a chunk
|
|
size assigns each worker one contiguous range of the entire pool, including
|
|
pending and terminated rays. A slab can therefore leave most workers with no
|
|
active rays even though there is substantial tracing work.
|
|
|
|
The production change is `schedule(dynamic, 32)` in
|
|
`ray_pool_advance_active()`. Each worker advances a chunk of up to 32 pool slots
|
|
and then takes another chunk. Pool slots and frame/sample IDs remain stable;
|
|
there are no new allocations, caches, tasks, or changes to the geodesic solver.
|
|
The parallel-region barrier still completes the slab before results are read.
|
|
Activation, status scans, and mesh analysis remain serial. The full pool is
|
|
still scanned; active-index compaction is not part of this change.
|
|
|
|
The 32-slot chunk is an empirical choice for this workload, not a universal
|
|
optimum. Cyclic static chunks perform similarly here. Dynamic chunks retain
|
|
load balancing when strong-field integration costs differ, while avoiding
|
|
per-ray scheduling.
|
|
|
|
## Dataset and build
|
|
|
|
- Baseline: `8cf6106127fc19732cee123e9890d401a6838a51`, including the fix for
|
|
frames with empty refinement generations.
|
|
- CPU: Intel Core i7-12700K; 16 online logical CPUs, 16 OpenMP workers.
|
|
- C compiler and full CPU details are saved in each `*-environment.json`.
|
|
- CPU PSF backend, event sink enabled, PNG output; C11, `-march=native -O2`,
|
|
`-DNDEBUG`, OpenMP. Both versions are built from baseline sources, with the
|
|
candidate substituting only the current `src/ray.c`.
|
|
- Input: all 393 samples of the reported inward free-fall trajectory, preserved
|
|
verbatim in [camera.csv](ray_pool_scheduling_2026-09-06/camera.csv).
|
|
- `assets/sky_grid_5deg.csv`: 5654 catalog stars. Input hashes are recorded.
|
|
- 64x36, 60-degree horizontal FOV, coarse cell 16 pixels, refinement level 3,
|
|
exposure 0.2. The temporal span and original 16:9 aspect ratio are retained.
|
|
- Every run writes all 393 PNGs and a complete movie lens map. This is a
|
|
reduced-resolution benchmark, not the original 640x360 run. Do not extrapolate
|
|
its speedup to the full-resolution movie or numerical backends.
|
|
- Runs were sequential; build/test jobs were not run concurrently with timing.
|
|
`OMP_NUM_THREADS=16`, `OMP_DYNAMIC=FALSE`. PSF cache build and image output are
|
|
included in end-to-end wall time. No cache files are reused by these commands.
|
|
Filesystem caches were not flushed; this is not a cold-I/O benchmark.
|
|
|
|
## Schedule selection
|
|
|
|
A temporary baseline source copy uses `schedule(runtime)` so `OMP_SCHEDULE`
|
|
selects the experiment. Each worker records active-ray visits, integrator
|
|
steps, and loop elapsed time without the final barrier; batch wall time includes
|
|
the barrier. Counters are collected in worker-owned slots outside solver state.
|
|
These counters and the runtime schedule override are **not in production**.
|
|
|
|
| Schedule | First run wall (s) | Reverse-order repeat wall (s) | First run ray-batch time (s) |
|
|
| --- | ---: | ---: | ---: |
|
|
| contiguous static | 61.040 | 61.093 | 55.688 |
|
|
| cyclic static,32 | 39.540 | 40.119 | 34.101 |
|
|
| dynamic,32 | 39.292 | 38.679 | 33.808 |
|
|
| dynamic,128 | 39.792 | not repeated | 34.316 |
|
|
|
|
Every instrumented run executes 47 slab advances, 402534 active-ray visits and
|
|
206670193 integrator steps. The first slab has 1856 active rays: contiguous
|
|
static assigns them to only 3 workers, while all chunked schedules use 16.
|
|
The small difference between static,32 and dynamic,32 is not evidence that one
|
|
is universally faster. The major improvement is removing large contiguous
|
|
frame ranges from each worker's assignment.
|
|
|
|
User/system CPU seconds are preserved in `*-results.json`. These include OpenMP
|
|
waiting and should not be interpreted as useful solver work alone. Exact
|
|
per-worker counts and times remain in the raw logs, and aggregate counters are
|
|
in [instrumentation-summary.json](ray_pool_scheduling_2026-09-06/instrumentation-summary.json).
|
|
|
|
## Reproduction and original output
|
|
|
|
Run from the repository root:
|
|
|
|
```sh
|
|
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py
|
|
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
|
|
--tag repeat --schedules dynamic,32 static,32 static
|
|
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
|
|
--production --tag production-before --schedules static
|
|
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
|
|
--production --tag production-after --schedules dynamic,32
|
|
```
|
|
|
|
[run.py](ray_pool_scheduling_2026-09-06/run.py) builds temporary executables under
|
|
`/tmp/gr-ray-schedules`, and stores original output and metadata beside the
|
|
script. Use a fresh `--tag` and `--work-dir` to preserve an existing run. The
|
|
production comparisons compile literal scheduling directives and have no
|
|
instrumentation; their logged `OMP_SCHEDULE` value does not override those
|
|
compiled directives. Production-after requires the candidate `src/ray.c`.
|
|
|
|
Every run has a `*-command.txt` containing the complete, shell-copyable renderer
|
|
command, `*-results.json` with exit status, wall/user/system time and artifact
|
|
hash, and an unfiltered `*.log.gz` preserving all stdout/stderr bytes, including
|
|
PSF build/cache, worker, processed-image and fallback statistics. Build command
|
|
files and build stdout/stderr logs are also preserved; successful silent
|
|
compiler output can produce an empty build log. For example:
|
|
|
|
```sh
|
|
gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-0-static.log.gz
|
|
gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-2-dynamic-32.log.gz
|
|
```
|
|
|
|
Output PNGs and lens maps remain under the run's work directory; they are not
|
|
stored in Git. The reference lens-map SHA-256 is
|
|
`708a218eca26b9eac42cb4954b0dbd4dfe316cc55d4315d650cbc8ab7e616080`.
|
|
All schedule-selection runs match it byte-for-byte, including the geometry,
|
|
endpoint statuses, frequency shifts, frame metadata and mesh topology.
|
|
|
|
## Uninstrumented production-source comparison
|
|
|
|
| Source | Wall (s) | User CPU (s) | System CPU (s) |
|
|
| --- | ---: | ---: | ---: |
|
|
| Baseline contiguous static | 62.126 | 411.940 | 1.964 |
|
|
| Candidate dynamic,32 | 38.686 | 554.787 | 2.748 |
|
|
|
|
The end-to-end speedup is 1.606x (37.7% less wall time). The CPU-seconds / wall
|
|
ratio increases from 6.66 to 14.41, including runtime waiting. This comparison
|
|
has no per-worker diagnostic counters and no runtime-selected schedule. It is
|
|
one pair of runs, supported by the reverse-order instrumented checks above;
|
|
there is no statistical confidence interval or claim about all workloads.
|
|
|
|
Both runs exit 0 and write 393 PNGs. Every PNG is byte-identical before/after;
|
|
all candidate PNGs passed signature, dimension, chunk CRC and IDAT decompression
|
|
checks. All nine measured runs have the same complete lens-map hash.
|
|
See [artifact-validation.json](ray_pool_scheduling_2026-09-06/artifact-validation.json),
|
|
[baseline results](ray_pool_scheduling_2026-09-06/production-before-results.json),
|
|
[candidate results](ray_pool_scheduling_2026-09-06/production-after-results.json),
|
|
and their corresponding unfiltered compressed logs in the same directory.
|
|
|
|
## Regression verification
|
|
|
|
`make -j8 BUILD_TYPE=Release all` and `make -j8 BUILD_TYPE=Debug all` rebuilt the
|
|
actual production binaries. Both corresponding `make test` suites passed,
|
|
including a Schwarzschild movie regression that compares complete lens maps
|
|
with 1, 4 and 16 requested OpenMP threads. `git diff --check` passed. Exact
|
|
commands and compressed original build/test output are preserved in
|
|
`verification-commands.txt`, `release-build.log.gz`, `release-tests.log.gz`,
|
|
`debug-build.log.gz` and `debug-tests.log.gz` beside the benchmark data.
|