Add an opt-in DUMMY_PSF_SUPPORT_STATS mode to the diagnostic-only dummy
consumer. It records per-chunk cache-limited 16px support references,
touched tiles, 256-reference tasks, maximum tile list and first-pass
counting time, then reports percentiles. No HDR is produced and no
references are materialized or uploaded.
Move adaptive tile16 selection and support-expanded CPU binning into a
per-producer HipPsfPreparedChunk computed before acquiring the shared
sink lock. The sink keeps sole ownership of the GPU HDR, stream, device
scratch and upload staging; submit_prepared runs under the lock and
preserves the existing single-stream ordering and direct fallback path.
The legacy hip_psf_sink_submit entry point prepares an internal chunk so
single-threaded replay stays compatible, and default accumulation remains
atomic.
Replace the in-tree Wyman analytic CIE fit with a repository CIE 1931
2-degree LUT derived from the 360-830 nm, 1 nm-linear CSV reference. The
GRBBLUT3 table stores 1024 uniform log(T) XYZ nodes over
[670.146556, 101408.88] K with four-point cubic Lagrange interpolation, and
the same CSV reference supplies a three-term Rayleigh-Jeans form above the
table and an inverse-temperature/log-XYZ crossover with channel-specific
endpoint forms below it. The loader validates the header and payload
checksum and rejects legacy formats and runtime generation.
Initialize the backend in the renderer, frame test, and PSF capture tool so
the new default is active wherever colors are produced.
Preserve the protected full-catalog dummy run, its initial harness failure, input and output hashes, distribution quantiles, and the resulting sparse-tile scheduling conclusions.
Reuse production catalog mapping and per-worker chunk boundaries to report triangle and spatial distributions without initializing HIP or writing image outputs. Include a protected full-catalog runner and build documentation.
Add a bounded per-chunk center-tile selector to the real-event replay prototype and retain the paired timing, mixed-boundary, CPU, HIP, and sandbox-failure evidence.
Archive fixed catalog subsets and captured events, manifests, reproducible commands, raw timing and regression logs, and device resource evidence. Document distribution-dependent tile reduction gains and independent indexed producer throughput.
Update the GPU plan and renderer design while retaining the production atomic path. Preserve historical measurement metadata and exclude generated Python bytecode.
Capture prepared events through a test-only producer consumer and measure serial and indexed parallel generation independently. Compare the production atomic kernel with bounded pixel-owned tile reductions while preserving double HDR, cache support, and direct fallback ordering.
Validate captured inputs, edge and multichunk fixtures, CPU/HIP renderer regressions, and 45 isolated final replays. Keep the production HIP algorithm unchanged because gains depend on event distribution.
Replace the artifact-affected Schwarzschild image with the corrected 4K render. Link both READMEs to the paired CPU/HIP benchmark and clarify that this run had no PSF wing clipping or direct fallbacks.
Preserve user-provided complete CPU and HIP logs and associate the optimized implementation with e1ec480669. Record matching event counts, user-confirmed bit-identical PNG output, and the 2.77x workflow speedup with tracing/import differences noted.
Mark the pre-artifact-fix CPU benchmark as historical and link the new paired result.
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.
Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
Check spherical area weight sums against a fixed 1e-8 tolerance and normalize accepted weights before interpolation. Skip invalid images in both catalog paths.
Add a thin-triangle regression for both parities, document the measured tolerance margin, and refresh HDR fixtures with strict comparisons restored.
Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations.
Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
Allow frames to finish refinement independently without aborting the movie. Report genuine refinement errors with generation and frame identifiers, and clarify the sample batch contract.
Add a two-frame Schwarzschild regression covering unequal refinement progress. Validated Release and Debug make test suites and 640x360 rendering of the final eight free-fall frames in both builds.
Integrate timelike geodesics and Fermi-Walker tetrads in ingoing Kerr-Schild coordinates, sampled at a configurable proper-time cadence.
Add movie-track-samples to preserve CSV events as frames. Document usage and singularity guards, and cover analytic orbits, transport convergence, sampling, and CSV rendering.