Document the producer-private adaptive path, the dummy support workload diagnostics, the negative result of the old serial binning and the 2026-09-13 full-day comparison (splat 282.944 s vs atomic 481.417 s). Add the corresponding raw logs, hotspot fixtures and runner evidence, and lay out the staged plan toward a sub-240 s full-day render without changing the default atomic path.
11 lines
1.4 KiB
Plaintext
11 lines
1.4 KiB
Plaintext
COMMAND OMP_NUM_THREADS=16 build/Release/replay_psf benchmarks/hip_psf_replay_2026-09-07/center_65536.events 65536 atomic 256
|
|
TIMEOUT 45s
|
|
GIT 2d01c2c101ccea8193a64bb46ceafca3009f0ef5
|
|
BINARY_SHA256 4b2e99b04076bce7856d0ca2437ec69809dc9846cc1f28eb31e9441da7fb5554
|
|
PSF cache ready: 64x64 phases, radius 47 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 0.671 s
|
|
device=AMD Radeon RX 9070 arch=gfx1201 events=65536 rounds=256 frame=3840x2160 mode=atomic distribution=captured contributions_per_round=432818979 CPU_reference=0.811035872 create=0.039489031
|
|
HIP PSF progress: 5799936 submitted / 5767168 completed events; 352 / 354 batches timed; kernel 4.944 s
|
|
HIP PSF progress: 11698176 submitted / 11665408 completed events; 712 / 714 batches timed; kernel 9.936 s
|
|
complete_replay=14.364993095 replay_wall=14.325504065 select=0.000000000 bin=0.000000000 upload=0.047864901 prepare=0.000000000 accumulation=14.264586411 merge=0.000000000 download=0.004812979 cleanup=0.000000000 events_s=1171143.153 contributions_s=7734573116.915 center_tiles=0 tile_chunks=0 atomic_chunks=1024 refs=0 tasks=0 scratch_used_peak=0 scratch_device_reserved=0 RSS_KiB=706160 max_abs=0.015890534383196337 max_rel=255.00000000039253 flux_R=254.99999999972784 flux_G=254.99999999963809 flux_B=254.9999999993957 flux_Y=254.99999999966304 validation=skipped-for-sustained-timing
|
|
PROCESS_WALL 16.201255127 EXIT 0
|