Files
GR-raytracing/benchmarks/hip_psf_2026-09-07.md
T
wyj ab7d879f6c Doc: record full-sky CPU and HIP benchmark comparison
Preserve user-provided complete CPU and HIP logs and associate the optimized implementation with e1ec480669. Record matching event counts, user-confirmed bit-identical PNG output, and the 2.77x workflow speedup with tracing/import differences noted.

Mark the pre-artifact-fix CPU benchmark as historical and link the new paired result.
2026-09-06 21:40:51 -04:00

205 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# HIP PSF 有界排查与优化(2026-09-07)
后续用户完成了全天空 CPU/HIP 配对运行,并确认 PNG 比特级一致;见
[完整 4K 记录](2mass_galactic_center_cpu_hip_2026-09-07.md),其中记录了本次优化提交的实际 hash。
本页保留提交前有界排查的原始上下文。
## 范围与结论
本次没有运行用户的完整 343,363,141 星 / 598,308,919 星像任务。
所有 renderer/benchmark 都单进程顺序执行,前一个退出后才启动下一个;单次微基准
45 秒超时,缩小 renderer 60 秒超时。没有并发性能测试,没有删除原 render 输出。
源码基线为 `265d7b95d5b916ec4768d174d73ad963a43f0088`,优化版是此基线上的未提交改动。
最初保留了用户现存 HIP binary;另从该 Git commit 的干净源码重建旧版,以复核对照。
主机 HIP 7.2.53210-9999,AMD Radeon RX 9070 / gfx1201;性能测试设置
`OMP_NUM_THREADS=16`,catalog loaders 为 4。Release / `-O2`;HIP 显式
`--offload-arch=gfx1201`。完整元数据见 [metadata.log](hip_psf_2026-09-07/metadata.log)。
已确认的实现问题与处理:
- 旧 HIP 分支将 catalog 查询、逆映射、频移与黑体积分串行化。现在按 triangle 使用
OpenMP producer,每 worker 一个 16,384-event 私有块,GPU cache/HDR 只分配一次。
- 旧 kernel 一线程串行遍历一个 PSF,double atomic 编译为 64-bit CAS 重试循环。
现在每个 event 用 32 线程访问相邻像素;保持四点 phase cache、圆形 support 和
double RGB 累积。分布实验支持这一改变,不把全部 kernel 时间归因为 atomic 竞争。
- 旧 host 缓冲是可立即复用的普通 malloc 数组;现在 sink 同步复制进自己拥有的两个
pinned slot,重用 host/device slot 前等待上一 kernel 完成,生命周期不依赖 pageable
`hipMemcpyAsync` 的运行时暂存行为。GPU copy/kernel 仍在同一 stream 内顺序执行。
- 旧计时只覆盖前 64 批。现在两个 slot 的计时 event 循环使用,统计全部批次,并约每
5 秒在提交时打印进度。完成数是已回收计时的完成数,最多有两个未回收 slot。
- chunk 提交与整个 CPU direct-fallback 的 download/evaluate/upload 持有同一共享锁。
没有 per-star 锁,没有共享可变 CPU scratch,没有改变 backend 错误时立即失败的行为。
未使用分桶规约、float HDR、近似黑体积分或改变 min-Y/tail 参数。
## 受控 kernel 实验
使用 `tests/benchmark_hip_psf.hip` 生成固定 LCG seed 12345 的相同事件。frame 3840×2160,
FWHM 2.7,beta 4.5,tail 1e-8,cache radius 47;每个事件 56 B。
每个进程先计算串行 CPU double HDR reference,再执行一个 HIP replay。CPU reference
耗时不是 production 16-thread CPU render 的性能基线。
| 事件数 | 中心分布范围 | 旧 kernel | 32-thread kernel 原型 | 比值 |
| ---: | ---: | ---: | ---: | ---: |
| 4,096 | 16 px | 22.490 ms | 5.702 ms | 3.94× |
| 16,384 | 1 px | 137.888 ms | 23.976 ms | 5.75× |
| 16,384 | 3840 px(y 限于2160) | 55.630 ms | 21.492 ms | 2.59× |
原型阶段仅改变 kernel 的线程分工,保留旧上传和计时实现,便于隔离效果。
原始输出:[旧密集](hip_psf_2026-09-07/baseline_dense.log)、
[新密集](hip_psf_2026-09-07/wave_dense.log)、
[旧热点](hip_psf_2026-09-07/baseline_hotspot.log)、
[新热点](hip_psf_2026-09-07/wave_hotspot.log)、
[旧分散](hip_psf_2026-09-07/baseline_spread.log)、
[新分散](hip_psf_2026-09-07/wave_spread.log)。
整合双缓冲后的 65,536-event/16px replay:4/4 批均计时,kernel 78.268 ms,replay wall
83.762 ms;CPU reference 842.402 ms。double HDR max_abs `1.7053025658242404e-13`,
max_rel `3.3472443502765663e-15`,总 RGB flux 相对差 `-1.3316774058238102e-17`。
[原始输出](hip_psf_2026-09-07/final_multibatch.log)。
## 原 4K lens map 的真实 catalog 子集
保留原文件 `output/lens/schwarzschild_galactic_center_R100_45deg_16_4_4k_.grlens`:
3840×2160、45°、65,817 vertices、130,882 triangles。未重新 tracing/refine。
使用 tile 索引的主对照输入只有三个 tile:RA252/Dec048、RA262/Dec060、RA082/Dec090。
每个 tile 均匀抽取最多 8192 行,共 21,087 颗星,产生 156,837 个星像。
准备脚本和源/子集 SHA256:[prepare_inputs.py](hip_psf_2026-09-07/prepare_inputs.py)、
[input_manifest.log](hip_psf_2026-09-07/input_manifest.log)。
其余 tile 故意缺失,因此 prefetch 的 64,437 unavailable 是缩小输入条件,不是生产数据错误。
| 实现 | catalog/splat(日志四舍五入) | GPU kernel 合计 | 完整进程 |
| --- | ---: | ---: | ---: |
| 原有 HIP binary | 1.9 s | 1.469529 s | 2.971 s |
| 干净旧源码重建 HIP | 1.9 s | 1.496493 s | 3.016 s |
| 优化 HIP | 0.4 s | 0.168641 s | 1.487 s |
| CPU 16 workers | 0.6 s | — | 1.674 s |
所有实现均 cached=156837、wing-clipped=0、direct=0、discarded=0。
优化版 GPU 16 批全部计时;producer wall=0.326 s,worker generation 总和=2.053 s,
submission/fallback 总和=1.157 s。worker wall 总和包含并发时间,不能加到 GPU 时间上。
原始日志:[旧 HIP](hip_psf_2026-09-07/baseline_multi.log)、
[优化 HIP](hip_psf_2026-09-07/final_multi.log)、[CPU](hip_psf_2026-09-07/cpu_multi.log)。
干净源码旧版复核另见 [clean_baseline_multi.log](hip_psf_2026-09-07/clean_baseline_multi.log)。
旧 HIP 与优化 HIP 的 kernel 时间相差约 8.7 倍、完整进程约 2 倍;这些是小规模单次测量,
不是完整全天空的加速比。相较 CPU,本子集的完整进程收益较小,cache 初始化和 4K 输出
占比较高。没有依此预测用户完整任务的完成时间。
辅助诊断保留了同一 lens map 的单 CSV 模式:256 星及 8192 星。
8192 星时旧/新 splat 为 4.4/0.6 秒,完整进程 5.489/1.706 秒。
但单 CSV 会放大逐 triangle 扫描成本,不能用它代表生产 tile 索引的 CPU 占比。
同 8192 星改用 tile 索引后,旧 kernel=0.171626 秒,新 kernel=0.028882 秒。
日志:[256](hip_psf_2026-09-07/baseline_lens256.log)、
[CSV旧](hip_psf_2026-09-07/baseline_lens8192.log)、
[CSV新](hip_psf_2026-09-07/final_lens8192.log)、
[索引旧](hip_psf_2026-09-07/baseline_indexed.log)、
[索引新](hip_psf_2026-09-07/final_indexed.log)。
## 验证
- `make -j1 test`:通过;固定 HDR、camera、geodesic、frame、Schwarzschild、observer track、
catalog prefetch、CLI 和多线程 movie/lens map 回归均通过。
[完整日志](hip_psf_2026-09-07/cpu_regression.log)。
- `test_hip_psf`:213 事件、71 次 submit,全部计时;覆盖 host 数组立即覆写、slot 复用、
画面边缘、wing clipping 和中途 CPU HDR 修改后继续 GPU。
[初次日志](hip_psf_2026-09-07/streaming_test.log)、
[最终重建后日志](hip_psf_2026-09-07/streaming_test_final.log)。
- `tests/test_hip_renderer.py`:96×64、4 workers,先 CPU 后 HIP;验证 cached=2/direct=1
的混合场和 discarded=3 的 min-Y 场。统计一致、FITS float32 HDR payload 完全一致。
[完整日志](hip_psf_2026-09-07/renderer_regression.log)。
- 4K 子集的旧/新 HIP 以及 CPU/新 HIP:24,883,200 个 FITS float32 样本完全一致;PNG
也一致。double 精度由上述 sink/replay tests 单独验证,未将 FITS 当作 double 文件。
[比较日志](hip_psf_2026-09-07/hdr_comparison.log)。
- 在单个子进程设置 `ROCR_VISIBLE_DEVICES=''`,32×24 的显式 HIP render 报设备不可用、
非零退出且没有写图。[原始输出](hip_psf_2026-09-07/fail_fast.log)。
- 初次 mixed fixture 使用 exposure=2 未越过 support 阈值,测试断言正确拒绝了无 fallback
的输入;依据源码改为 exposure=1000 后才通过。这是测试输入校正,非 production 修复。
[初次输出](hip_psf_2026-09-07/fixture_calibration_failed.log)。
## 可复制命令
以下使用 zsh 默认 `time`,不设置 GNU 风格 `TIMEFMT`。命令必须顺序执行;任何失败或
超时后先确认前一个进程已退出。原始输出保存在上述链接,未只保留汇总值。
```zsh
# 从本仓库根目录执行;准备固定样本,输出源与子集哈希。
python3 benchmarks/hip_psf_2026-09-07/prepare_inputs.py
# 独立重建旧源码,不改变当前 checkout。
mkdir -p /tmp/gr-hip-baseline-src
git archive 265d7b95d5b916ec4768d174d73ad963a43f0088 Makefile src mk | tar -x -C /tmp/gr-hip-baseline-src
make -j4 -C /tmp/gr-hip-baseline-src PSF_BACKEND=hip SPACETIME=schwarzschild ENABLE_HDR=1 \
HIP_CXXFLAGS='--offload-arch=gfx1201 -std=c++17 -O2 -Wall -Wextra -Wpedantic' backend
# 当前实现与 CPU reference。
make -j4 PSF_BACKEND=hip SPACETIME=schwarzschild ENABLE_HDR=1 \
HIP_CXXFLAGS='--offload-arch=gfx1201 -std=c++17 -O2 -Wall -Wextra -Wpedantic' backend hip-psf-test hip-psf-bench
make -j4 PSF_BACKEND=cpu SPACETIME=schwarzschild ENABLE_HDR=1 backend
render_args=(
--all-sky-catalog /tmp/gr-hip-catalog-multi --catalog-load-workers 4
--psf-relative-tail 1e-8 --psf-min-y 0 --max-cache-psf-flux 1e8
--exposure 1e13 --psf-fwhm-pixels 2.7 --psf-moffat-beta 4.5 --hdr-output
--lens-map-input output/lens/schwarzschild_galactic_center_R100_45deg_16_4_4k_.grlens
--verbose
)
time OMP_NUM_THREADS=16 timeout --kill-after=5s 60s /tmp/gr-hip-baseline-src/build/Release/schwarzschild_sky_hip \
"${render_args[@]}" --output /tmp/gr-hip-clean-baseline-multi.png
time OMP_NUM_THREADS=16 timeout --kill-after=5s 60s ./build/Release/schwarzschild_sky_hip \
"${render_args[@]}" --output /tmp/gr-hip-final-multi.png
time OMP_NUM_THREADS=16 timeout --kill-after=5s 60s ./build/Release/schwarzschild_sky \
"${render_args[@]}" --output /tmp/gr-cpu-multi.png
# 原有 binary 的初始记录使用 /tmp/gr-schwarzschild-hip-baseline,其他参数完全相同。
# 单 tile 索引测试仅将 root 换为 /tmp/gr-hip-catalog-subset。
# 单 CSV 诊断改用 --catalog benchmarks/hip_psf_2026-09-07/catalog_8192.csv
# (或 catalog_256.csv),不用 --all-sky-catalog,输出到单独的 /tmp PNG。
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s ./build/Release/benchmark_hip_psf 65536 16
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s ./build/Release/test_hip_psf
make -j1 test
# 小尺寸 renderer 回归;先分别构建带 HDR 的 Minkowski CPU/HIP。
make PSF_BACKEND=cpu SPACETIME=minkowski ENABLE_HDR=1 backend
make PSF_BACKEND=hip SPACETIME=minkowski ENABLE_HDR=1 \
HIP_CXXFLAGS='--offload-arch=gfx1201 -std=c++17 -O2 -Wall -Wextra -Wpedantic' backend
python3 tests/test_hip_renderer.py build/Release/minkowski_sky build/Release/minkowski_sky_hip
python3 scripts/fits_floatdiff.py /tmp/gr-cpu-multi_HDR.fits /tmp/gr-hip-final-multi_HDR.fits
```
旧/新原型微基准可使用同一 `tests/benchmark_hip_psf.hip` 与基线 sink 源码编译:
```zsh
cp /tmp/gr-hip-baseline-src/src/hip_psf.hip /tmp/gr-hip-psf-baseline.hip
hipcc --offload-arch=gfx1201 -std=c++17 -O2 -fopenmp -Isrc -x hip \
tests/benchmark_hip_psf.hip /tmp/gr-hip-psf-baseline.hip -x none \
build/Release/obj/schwarzschild/hdr_sink1_hip/src/optics.o \
-lm -lpng -lcfitsio -o /tmp/hip_psf_bench_baseline
python3 - <<'PY'
from pathlib import Path
s = Path('/tmp/gr-hip-psf-baseline.hip').read_text()
s = s.replace('const size_t index = (size_t)blockIdx.x * blockDim.x + threadIdx.x;',
'const size_t thread = (size_t)blockIdx.x * blockDim.x + threadIdx.x;\n'
' const size_t index = thread / 32;\n const int lane = thread % 32;')
s = s.replace('int offset_x = first; offset_x <= last; ++offset_x',
'int offset_x = first + lane; offset_x <= last; offset_x += 32')
s = s.replace('(event_count + threads - 1) / threads',
'(event_count * 32 + threads - 1) / threads')
Path('/tmp/gr-hip-psf-wave.hip').write_text(s)
PY
hipcc --offload-arch=gfx1201 -std=c++17 -O2 -fopenmp -Isrc -x hip \
tests/benchmark_hip_psf.hip /tmp/gr-hip-psf-wave.hip -x none \
build/Release/obj/schwarzschild/hdr_sink1_hip/src/optics.o \
-lm -lpng -lcfitsio -o /tmp/hip_psf_bench_wave
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s /tmp/hip_psf_bench_baseline 4096 16
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s /tmp/hip_psf_bench_wave 4096 16
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s /tmp/hip_psf_bench_baseline 16384 1
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s /tmp/hip_psf_bench_wave 16384 1
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s /tmp/hip_psf_bench_baseline 16384 3840
OMP_NUM_THREADS=16 timeout --kill-after=5s 45s /tmp/hip_psf_bench_wave 16384 3840
```