HIP: accelerate PSF accumulation and restore parallel producers

Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.

Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
This commit is contained in:
wyj committed 2026-09-06 21:38:53 -04:00
1 parent 265d7b95d5
commit e1ec480669
45 files changed
+10251 -95

No files matched your search

+3
View File
@@ -33,6 +33,9 @@ adaptive lens meshes, and reusable lens-map files. It is written primarily in
C with OpenMP CPU parallelism; an optional HIP backend accelerates PSF
accumulation.
HIP retains parallel CPU catalog mapping and uses bounded, completion-protected
event uploads. See [HIP configuration and bounded performance checks](build.md#optional-hip-psf-acceleration).
The [Nmesh](https://github.com/nmeshsource/nmesh) numerical-spacetime backend and BBH rendering are still planned.
The current scope is black-hole capture and distant stellar backgrounds;
local matter emission, accretion disks, and plasma are outside this stage.