Move adaptive tile16 selection and support-expanded CPU binning into a
per-producer HipPsfPreparedChunk computed before acquiring the shared
sink lock. The sink keeps sole ownership of the GPU HDR, stream, device
scratch and upload staging; submit_prepared runs under the lock and
preserves the existing single-stream ordering and direct fallback path.
The legacy hip_psf_sink_submit entry point prepares an internal chunk so
single-threaded replay stays compatible, and default accumulation remains
atomic.
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.
Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.