Preserve user-provided complete CPU and HIP logs and associate the optimized implementation with e1ec480669. Record matching event counts, user-confirmed bit-identical PNG output, and the 2.77x workflow speedup with tracing/import differences noted.
Mark the pre-artifact-fix CPU benchmark as historical and link the new paired result.
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.
Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.