Ray: Balance movie tracing with dynamic chunks
Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations. Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
This commit is contained in:
1 parent
8cf6106127
commit
8741597d9d
48 files changed
+895
-4
No files matched your search
@@ -0,0 +1,140 @@
|
||||
# Movie ray-pool scheduling benchmark (2026-09-06)
|
||||
|
||||
## Change and scope
|
||||
|
||||
Movie samples are appended in frame order. `schedule(static)` without a chunk
|
||||
size assigns each worker one contiguous range of the entire pool, including
|
||||
pending and terminated rays. A slab can therefore leave most workers with no
|
||||
active rays even though there is substantial tracing work.
|
||||
|
||||
The production change is `schedule(dynamic, 32)` in
|
||||
`ray_pool_advance_active()`. Each worker advances a chunk of up to 32 pool slots
|
||||
and then takes another chunk. Pool slots and frame/sample IDs remain stable;
|
||||
there are no new allocations, caches, tasks, or changes to the geodesic solver.
|
||||
The parallel-region barrier still completes the slab before results are read.
|
||||
Activation, status scans, and mesh analysis remain serial. The full pool is
|
||||
still scanned; active-index compaction is not part of this change.
|
||||
|
||||
The 32-slot chunk is an empirical choice for this workload, not a universal
|
||||
optimum. Cyclic static chunks perform similarly here. Dynamic chunks retain
|
||||
load balancing when strong-field integration costs differ, while avoiding
|
||||
per-ray scheduling.
|
||||
|
||||
## Dataset and build
|
||||
|
||||
- Baseline: `8cf6106127fc19732cee123e9890d401a6838a51`, including the fix for
|
||||
frames with empty refinement generations.
|
||||
- CPU: Intel Core i7-12700K; 16 online logical CPUs, 16 OpenMP workers.
|
||||
- C compiler and full CPU details are saved in each `*-environment.json`.
|
||||
- CPU PSF backend, event sink enabled, PNG output; C11, `-march=native -O2`,
|
||||
`-DNDEBUG`, OpenMP. Both versions are built from baseline sources, with the
|
||||
candidate substituting only the current `src/ray.c`.
|
||||
- Input: all 393 samples of the reported inward free-fall trajectory, preserved
|
||||
verbatim in [camera.csv](ray_pool_scheduling_2026-09-06/camera.csv).
|
||||
- `assets/sky_grid_5deg.csv`: 5654 catalog stars. Input hashes are recorded.
|
||||
- 64x36, 60-degree horizontal FOV, coarse cell 16 pixels, refinement level 3,
|
||||
exposure 0.2. The temporal span and original 16:9 aspect ratio are retained.
|
||||
- Every run writes all 393 PNGs and a complete movie lens map. This is a
|
||||
reduced-resolution benchmark, not the original 640x360 run. Do not extrapolate
|
||||
its speedup to the full-resolution movie or numerical backends.
|
||||
- Runs were sequential; build/test jobs were not run concurrently with timing.
|
||||
`OMP_NUM_THREADS=16`, `OMP_DYNAMIC=FALSE`. PSF cache build and image output are
|
||||
included in end-to-end wall time. No cache files are reused by these commands.
|
||||
Filesystem caches were not flushed; this is not a cold-I/O benchmark.
|
||||
|
||||
## Schedule selection
|
||||
|
||||
A temporary baseline source copy uses `schedule(runtime)` so `OMP_SCHEDULE`
|
||||
selects the experiment. Each worker records active-ray visits, integrator
|
||||
steps, and loop elapsed time without the final barrier; batch wall time includes
|
||||
the barrier. Counters are collected in worker-owned slots outside solver state.
|
||||
These counters and the runtime schedule override are **not in production**.
|
||||
|
||||
| Schedule | First run wall (s) | Reverse-order repeat wall (s) | First run ray-batch time (s) |
|
||||
| --- | ---: | ---: | ---: |
|
||||
| contiguous static | 61.040 | 61.093 | 55.688 |
|
||||
| cyclic static,32 | 39.540 | 40.119 | 34.101 |
|
||||
| dynamic,32 | 39.292 | 38.679 | 33.808 |
|
||||
| dynamic,128 | 39.792 | not repeated | 34.316 |
|
||||
|
||||
Every instrumented run executes 47 slab advances, 402534 active-ray visits and
|
||||
206670193 integrator steps. The first slab has 1856 active rays: contiguous
|
||||
static assigns them to only 3 workers, while all chunked schedules use 16.
|
||||
The small difference between static,32 and dynamic,32 is not evidence that one
|
||||
is universally faster. The major improvement is removing large contiguous
|
||||
frame ranges from each worker's assignment.
|
||||
|
||||
User/system CPU seconds are preserved in `*-results.json`. These include OpenMP
|
||||
waiting and should not be interpreted as useful solver work alone. Exact
|
||||
per-worker counts and times remain in the raw logs, and aggregate counters are
|
||||
in [instrumentation-summary.json](ray_pool_scheduling_2026-09-06/instrumentation-summary.json).
|
||||
|
||||
## Reproduction and original output
|
||||
|
||||
Run from the repository root:
|
||||
|
||||
```sh
|
||||
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py
|
||||
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
|
||||
--tag repeat --schedules dynamic,32 static,32 static
|
||||
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
|
||||
--production --tag production-before --schedules static
|
||||
python3 benchmarks/ray_pool_scheduling_2026-09-06/run.py \
|
||||
--production --tag production-after --schedules dynamic,32
|
||||
```
|
||||
|
||||
[run.py](ray_pool_scheduling_2026-09-06/run.py) builds temporary executables under
|
||||
`/tmp/gr-ray-schedules`, and stores original output and metadata beside the
|
||||
script. Use a fresh `--tag` and `--work-dir` to preserve an existing run. The
|
||||
production comparisons compile literal scheduling directives and have no
|
||||
instrumentation; their logged `OMP_SCHEDULE` value does not override those
|
||||
compiled directives. Production-after requires the candidate `src/ray.c`.
|
||||
|
||||
Every run has a `*-command.txt` containing the complete, shell-copyable renderer
|
||||
command, `*-results.json` with exit status, wall/user/system time and artifact
|
||||
hash, and an unfiltered `*.log.gz` preserving all stdout/stderr bytes, including
|
||||
PSF build/cache, worker, processed-image and fallback statistics. Build command
|
||||
files and build stdout/stderr logs are also preserved; successful silent
|
||||
compiler output can produce an empty build log. For example:
|
||||
|
||||
```sh
|
||||
gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-0-static.log.gz
|
||||
gzip -dc benchmarks/ray_pool_scheduling_2026-09-06/sweep-2-dynamic-32.log.gz
|
||||
```
|
||||
|
||||
Output PNGs and lens maps remain under the run's work directory; they are not
|
||||
stored in Git. The reference lens-map SHA-256 is
|
||||
`708a218eca26b9eac42cb4954b0dbd4dfe316cc55d4315d650cbc8ab7e616080`.
|
||||
All schedule-selection runs match it byte-for-byte, including the geometry,
|
||||
endpoint statuses, frequency shifts, frame metadata and mesh topology.
|
||||
|
||||
## Uninstrumented production-source comparison
|
||||
|
||||
| Source | Wall (s) | User CPU (s) | System CPU (s) |
|
||||
| --- | ---: | ---: | ---: |
|
||||
| Baseline contiguous static | 62.126 | 411.940 | 1.964 |
|
||||
| Candidate dynamic,32 | 38.686 | 554.787 | 2.748 |
|
||||
|
||||
The end-to-end speedup is 1.606x (37.7% less wall time). The CPU-seconds / wall
|
||||
ratio increases from 6.66 to 14.41, including runtime waiting. This comparison
|
||||
has no per-worker diagnostic counters and no runtime-selected schedule. It is
|
||||
one pair of runs, supported by the reverse-order instrumented checks above;
|
||||
there is no statistical confidence interval or claim about all workloads.
|
||||
|
||||
Both runs exit 0 and write 393 PNGs. Every PNG is byte-identical before/after;
|
||||
all candidate PNGs passed signature, dimension, chunk CRC and IDAT decompression
|
||||
checks. All nine measured runs have the same complete lens-map hash.
|
||||
See [artifact-validation.json](ray_pool_scheduling_2026-09-06/artifact-validation.json),
|
||||
[baseline results](ray_pool_scheduling_2026-09-06/production-before-results.json),
|
||||
[candidate results](ray_pool_scheduling_2026-09-06/production-after-results.json),
|
||||
and their corresponding unfiltered compressed logs in the same directory.
|
||||
|
||||
## Regression verification
|
||||
|
||||
`make -j8 BUILD_TYPE=Release all` and `make -j8 BUILD_TYPE=Debug all` rebuilt the
|
||||
actual production binaries. Both corresponding `make test` suites passed,
|
||||
including a Schwarzschild movie regression that compares complete lens maps
|
||||
with 1, 4 and 16 requested OpenMP threads. `git diff --check` passed. Exact
|
||||
commands and compressed original build/test output are preserved in
|
||||
`verification-commands.txt`, `release-build.log.gz`, `release-tests.log.gz`,
|
||||
`debug-build.log.gz` and `debug-tests.log.gz` beside the benchmark data.
|
||||
Reference in new issue
Block a user