Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries. Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
2 lines
88 B
Plaintext
2 lines
88 B
Plaintext
HIP PSF cache comparison: max_abs=2.8421709430404007e-14 max_rel=7.5319184269262976e-16
|