Files
GR-raytracing/benchmarks/2mass_galactic_center_cpu_hip_2026-09-07.md
T
wyj ab7d879f6c Doc: record full-sky CPU and HIP benchmark comparison
Preserve user-provided complete CPU and HIP logs and associate the optimized implementation with e1ec480669. Record matching event counts, user-confirmed bit-identical PNG output, and the 2.77x workflow speedup with tracing/import differences noted.

Mark the pre-artifact-fix CPU benchmark as historical and link the new paired result.
2026-09-06 21:40:51 -04:00

90 lines
5.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 2MASS 银河中心完整 4K:CPU / 优化 HIP(2026-09-07 归档)
## 代码版本与证据来源
HIP 优化提交:**`e1ec480669b0d24a0266ee135edf8a64a0d80ec9`**
(`HIP: accelerate PSF accumulation and restore parallel producers`)。
该提交包含 32-thread/event kernel、并行 CPU producer、有界双缓冲、全批次计时及回归测试。
HIP 完整运行发生在提交之前,对应此次提交的实现;运行日志本身没有嵌入 Git 或 binary hash。
用户指定本页 CPU 运行作为可比较基准,并说明
`265d7b95d5b916ec4768d174d73ad963a43f0088` 修复了此前的一些伪像。
因此,旧 `dd8cf29` 记录的 598,308,919 星像和 1,365 wing-clipped 不用于本次性能/正确性配对。
CPU 日志同样未嵌入 binary hash,不将代码版本说明冒充运行时二进制校验。
两份完整命令与原始终端输出均由用户提供:
- [CPU 原始日志](2mass_galactic_center_cpu_hip_2026-09-07/cpu.log):生成并导出 lens map。
- [HIP 原始日志](2mass_galactic_center_cpu_hip_2026-09-07/hip.log):导入同一 lens map;保留全部
worker 进度、每五秒 batch 进度、最终计数及 zsh `time` 输出,未删减成汇总。
用户确认:**两个 PNG 比特级完全一致。** 本次归档没有重跑渲染,也没有重新执行文件比较;
原始比较命令和两个 PNG 的 SHA256 未提供。不将 PNG 一致性扩大为完整任务的 linear HDR
逐样本一致性声明。此前有界样本的 HDR 验证另见
[有界排查记录](hip_psf_2026-09-07.md)。
## 工作负载与环境
- 主机沿用有界排查的 AMD Radeon RX 9070(gfx1201)与 HIP 7.2 环境;用户说明测试时段无其他
持续负载。完整日志没有独立输出驱动/编译器版本,前期环境证据见
[metadata.log](hip_psf_2026-09-07/metadata.log)。
- Release CPU/HIP 可执行文件,启用 HDR 输出;完整运行的实际编译命令未附在用户日志中。
- 3840×2160,45° FOV,RA 262.5° / Dec −30°,observer radius 100 M。
- CPU coarse cell 16,refine max level 4,Jacobian min 0.2;导出的 lens map 包含
65,817 vertices 和 130,882 triangles,HIP 直接导入它。
- 两次均请求并加载 64,440 tiles、343,363,141 stars,无 unavailable tile;catalog loaders=4。
- HIP 明确报告 16 producer workers。CPU 日志报告 1521% CPU,没有单独打印 worker 数。
- PSF FWHM 2.7 px、Moffat beta 4.5、relative tail 1e-8、min-Y=0、max-cache-psf-flux=1e8、
exposure=1e13;cache 为 64×64 phases、radius 47 px。
- 用户原始 CPU/HIP 命令使用相同 PNG 输出路径,且 CPU 命令的文件名也带 `full-gpu`;
原样保留此命名,不根据文件名判断 backend。未来复现应分别指定输出文件名,避免覆盖对照。
## 完整运行结果
| 项目 | CPU | 优化 HIP |
| --- | ---: | ---: |
| 总墙钟时间 | 1474.32 s(24:34.32) | 533.06 s(8:53.06) |
| user | 22234.08 s | 7733.61 s |
| system | 200.30 s | 61.18 s |
| CPU 百分比 | 1521% | 1462% |
| PSF cache build | 0.676 s | 0.691 s |
| catalog prefetch | 49.253 s | 48.201 s |
| 星像 / cached | 598264924 | 598264924 |
| wing-clipped | 0 | 0 |
| direct fallback | 0 | 0 |
| discarded below min-Y | 0 | 0 |
总墙钟时间比 `1474.32 / 533.06 = 2.76577`,耗时减少约 **63.84%**,节省 **941.26 秒**
(15 分 41.26 秒)。这是两次工作流的端到端时间比:CPU 包含 tracing 与 lens-map 写出,
HIP 导入文件;没有独立的 CPU PSF 阶段计时,不能把该比值称作严格隔离的 PSF backend 加速比。
HIP 全程诊断:
| 项目 | 结果 |
| --- | ---: |
| producer/splat wall | 481.155 s(frame 日志四舍五入为481.2 s) |
| GPU kernel | 476.906037 s |
| event upload | 1.504570 s |
| HDR download | 0.021966 s |
| batch / timed batch | 36526 / 36526 |
| summed worker generation | 3302.459 s |
| summed worker submission/fallback | 4395.356 s |
GPU kernel 时间约为 producer/splat wall 的 **99.12%**,kernel 吞吐约 **125.45 万事件/秒**。
上传占比很小,CPU producer 在中段能够持续供给 GPU;下一步性能工作应优先研究 kernel,
而不是仅增加 CPU worker 或上传带宽。这不单独证明 atomic contention 是 kernel 的主导成本。
多数进度行的 submitted−completed=32,768,等于两个 16,384-event staging slot,符合
有界队列与计时回收的设计。最后全部 36,526 批计时,没有沿用旧版前 64 批采样限制。
两个 worker 时间是线程墙钟之和,其中 submission/fallback 包含锁与设备等待:
`3302.459 + 4395.356 = 7697.815 ≈ 16 × 481.155`。不能将它们当作额外串行时间加到
GPU 或总墙钟时间上,1462% CPU 也不能全部解释为有效计算。
## 比较结论
本次完整任务确认了优化的实际收益,星像及 PSF 分类计数全部匹配,PNG 比特级一致由用户确认。
旧 CPU 文档中的计数差异属于不同修复版本,已不构成本次 CPU/HIP 配对的疑点。
原有完整日志、硬件/构建证据的限制与 tracing/import 差异均保留,不用小规模 kernel 加速比
替代这里的完整运行结果。