Files
GR-raytracing/benchmarks/dummy_psf_chunks_2026-09-11
wyj 2fc2b43e8d Doc: Record production PSF chunk geometry
Preserve the protected full-catalog dummy run, its initial harness failure, input and output hashes, distribution quantiles, and the resulting sparse-tile scheduling conclusions.
2026-09-11 23:28:02 -04:00
..
2026-09-11 23:27:42 -04:00

Production chunk geometry diagnostic (2026-09-11)

Scope

This run measures the event stream that the production HIP sink would receive for the supplied Schwarzschild Galactic-center render. PSF_BACKEND=dummy keeps the real imported lens map, all-sky prefetch, OpenMP dynamic triangle scheduling, inverse mapping, redshift/color/flux work, PSF cache eligibility, and 16,384-event per-worker chunk boundaries. It does not initialize HIP, allocate a full HDR framebuffer, splat pixels, or write PNG/FITS output.

The run used repository revision 11e703e33dffb1ff168ba11dd922e68b448116b9 plus the uncommitted dummy-backend work. The complete stdout/stderr, executable and input hashes, process resource use, and protected-output hashes are in full_catalog.log. An initial harness-only failure before the renderer started is retained as full_catalog_missing_time.log.

Build and run

make -j1 PSF_BACKEND=dummy SPACETIME=schwarzschild ENABLE_HDR=1 backend
python3 benchmarks/dummy_psf_chunks_2026-09-11/run.py

The runner fixes OMP_NUM_THREADS=16 and OMP_DYNAMIC=FALSE, uses a 600-second timeout and a process lock, refuses to overwrite its log, and hashes both the named PNG and companion FITS before and after the run. Its renderer arguments are exactly the user-supplied command except that the _hip executable is replaced by _dummy; the exact expanded command is the first line of the log.

Inputs and protected outputs:

  • lens map SHA-256: 905f960aea2d276f9586c79fba545a72cf742311e9e7e18faea293835c68bf28
  • existing PNG SHA-256 before/after: 3e01feb911e60acce545a6a86870ca05f96250022e874eb8097bec88a8879bfd
  • existing FITS SHA-256 before/after: 2408e62dcc22f147da284d87af9a2b0fdd0050eda2dabc23b6ec2554b2f1d38e

Both protected files retained the same size and nanosecond mtime as well as the same hash. No HIP backend was run.

Result

The catalog prefetch loaded 343,363,141 stars for 130,882 lens triangles. The producer classified 598,264,924 cache events, with zero direct fallbacks, wing-clipped events, or min-Y discards. Sixteen workers produced 36,524 chunks: 36,508 full chunks and one final partial chunk per worker.

full-chunk metric min p10 p25 p50 p75 p90 p99 max
event-producing triangles 1 1 1 2 3 10 31 59
occupied 32px center tiles 1 1 1 1 2 5 30 45
event bbox diagonal (px) 0.199 0.742 3.250 12.651 57.018 429.306 3759.261 3834.557
triangle-centroid bbox diagonal (px) 0 0 0 9.618 49.477 424.149 3744.061 3824.075
maximum successive centroid jump (px) 0 0 0 8.724 25.157 144.000 3744.034 3792.034

All 36,508 full chunks satisfy the replay prototype's current selector: events >= 8192 && events / occupied_32px_tiles >= 32.

Catalog prefetch took 55.469 s. Producer classification and chunk accounting took 232.340 s; total process wall time was 288.872 s. These are diagnostic CPU costs and do not estimate HIP tile-kernel speed.

Interpretation and limits

The measured production event order strongly supports the hypothesis that a full worker chunk is usually filled by very few image-plane triangles: the median is two and 75% contain at most three. Center-tile occupancy is even more concentrated. This is favorable for a tile16 path because the same output pixels receive many contributions that can be reduced before the final double HDR write.

The rare tail matters: p99 chunks can contain a few clusters separated by nearly the full 4K image width. “Few triangles” therefore does not imply one small contiguous bounding rectangle. A production implementation should use a sparse occupied-tile list, retain the atomic fallback, and measure the complete streaming sink including binning and synchronization.

This is one Schwarzschild Galactic-center frame with one 16-worker dynamic schedule. Worker-to-triangle assignment is nondeterministic, so exact quantiles may vary between runs. The result establishes event geometry for this input; it does not by itself establish single-BH/BBH universality or a HIP speedup.