HIP: record direct atomic PSF timings
This commit is contained in:
1 parent
7a80086f63
commit
a18ebfbf1a
9 files changed
+279
-10
No files matched your search
@@ -41,7 +41,10 @@ fallbacks remain CPU reference evaluations, with an ordered HDR transfer before
|
|||||||
and after each fallback. A HIP initialization, upload, kernel, or download
|
and after each fallback. A HIP initialization, upload, kernel, or download
|
||||||
error terminates the render; it never switches to the CPU backend silently.
|
error terminates the render; it never switches to the CPU backend silently.
|
||||||
Use `make PSF_BACKEND=hip SPACETIME=minkowski hip-psf-test` for the small
|
Use `make PSF_BACKEND=hip SPACETIME=minkowski hip-psf-test` for the small
|
||||||
GPU-vs-CPU cache-HDR regression.
|
GPU-vs-CPU cache-HDR regression. HIP renders also report event, batch, H2D,
|
||||||
|
kernel, and HDR-download timing after each frame. To bound diagnostic-event
|
||||||
|
storage, the report times at most the first 64 batches and labels the timed
|
||||||
|
batch count explicitly; the event and total-batch counts are not sampled.
|
||||||
|
|
||||||
## Movie PNG sequence (Phase A)
|
## Movie PNG sequence (Phase A)
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,139 @@
|
|||||||
|
# HIP direct-atomic 高密度 PSF endpoint(2026-09-05)
|
||||||
|
|
||||||
|
## 目的和边界
|
||||||
|
|
||||||
|
测量 production HIP `PsfEventSink` 在一个高表面密度 2MASS tile 上的端到端表现,并将 GPU
|
||||||
|
传输、kernel 与 HDR 下载分开记录。此记录验证 direct-global-atomic 的实际成本边界;它不是
|
||||||
|
tile-local reduction 的性能声明,也不把 kernel 时间直接归因为 atomic contention。
|
||||||
|
|
||||||
|
## 控制条件
|
||||||
|
|
||||||
|
- Source baseline: `7a80086f63e8b1e5f27d9524ba0455f2d6e49be5`
|
||||||
|
(`Feat: add initial HIP PSF backend`);HIP timing counters 是该基线之上的未提交诊断改动,
|
||||||
|
不改变 event、HDR 或输出路径。
|
||||||
|
- Host: i7-12700K;`OMP_NUM_THREADS=16`。
|
||||||
|
- GPU: AMD Radeon RX 9070 (`gfx1201`), 56 CUs, 15.9 GiB global memory;HIP runtime 是主机
|
||||||
|
实际 probe 所报告的设备。
|
||||||
|
- Build: Release;C11/`-march=native -O2 -pipe`、OpenMP、PNG;HIP 后端由
|
||||||
|
`PSF_BACKEND=hip` 选择。
|
||||||
|
- Input: `assets/2mass/processed/all_sky/tile_ra252_dec048.csv`,54,160 stars。
|
||||||
|
- Camera: Minkowski;RA 252.5 deg、Dec -41.5 deg、FOV 0.6 deg、1920 x 1920 px、
|
||||||
|
`--coarse-cell-pixels 240`(128 triangles / 81 initial rays)。
|
||||||
|
- PSF: FWHM 2.7 px、Moffat beta 3、exposure `1e15`、
|
||||||
|
`--max-cache-psf-flux 1e300`。cache radius 266 px;本次没有 min-Y discard 或 direct fallback。
|
||||||
|
|
||||||
|
这里刻意以较大的 coarse cell 减少 triangle 枚举工作,使固定 catalog 与相同镜头映射下的
|
||||||
|
splat 阶段成为可比较对象;它不是全天空或生产 4K 吞吐量估计。
|
||||||
|
|
||||||
|
## 命令与原始输出
|
||||||
|
|
||||||
|
CPU reference(同一输入与参数;在 HIP 命令确认退出后顺序运行):
|
||||||
|
|
||||||
|
```console
|
||||||
|
export OMP_NUM_THREADS=16
|
||||||
|
time build/Release/minkowski_sky \
|
||||||
|
--catalog assets/2mass/processed/all_sky/tile_ra252_dec048.csv \
|
||||||
|
--look-ra-deg 252.5 --look-dec-deg -41.5 --fov-deg 0.6 \
|
||||||
|
--width 1920 --height 1920 --coarse-cell-pixels 240 \
|
||||||
|
--exposure 1e15 --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
|
||||||
|
--max-cache-psf-flux 1e300 --output /tmp/cpu_psf_dense.png --verbose
|
||||||
|
PSF cache ready: 64x64 phases, radius 266 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 20.738 s
|
||||||
|
Frame 0: tracing 81 initial rays from 128 mesh triangles...
|
||||||
|
Frame 0: initial ray trace finished in 0.009 s.
|
||||||
|
Frame 0: traced 81 lens vertices; starting catalog render.
|
||||||
|
Frame 0: finding and prefetching catalog tiles...
|
||||||
|
Frame 0: catalog prefetch finished (0 candidate tiles).
|
||||||
|
Frame 0: splatting 128 lens triangles...
|
||||||
|
Frame 0: splat worker 14/16 started.
|
||||||
|
Frame 0: splat worker 10/16 started.
|
||||||
|
Frame 0: splat worker 16/16 started.
|
||||||
|
Frame 0: splat worker 1/16 started.
|
||||||
|
Frame 0: splat worker 12/16 started.
|
||||||
|
Frame 0: splat worker 13/16 started.
|
||||||
|
Frame 0: splat worker 5/16 started.
|
||||||
|
Frame 0: splat worker 11/16 started.
|
||||||
|
Frame 0: splat worker 7/16 started.
|
||||||
|
Frame 0: splat worker 6/16 started.
|
||||||
|
Frame 0: splat worker 3/16 started.
|
||||||
|
Frame 0: splat worker 4/16 started.
|
||||||
|
Frame 0: splat worker 2/16 started.
|
||||||
|
Frame 0: splat worker 9/16 started.
|
||||||
|
Frame 0: splat worker 8/16 started.
|
||||||
|
Frame 0: splat worker 15/16 started.
|
||||||
|
Frame 0: splat worker 14/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 2/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 10/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 7/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 1/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 11/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 3/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 5/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 4/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 12/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 16/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 13/16 reached 8 local triangles.
|
||||||
|
Frame 0: splat worker 9/16 finished after 4 local triangles.
|
||||||
|
Frame 0: splat worker 6/16 finished after 5 local triangles.
|
||||||
|
Frame 0: splat worker 15/16 finished after 5 local triangles.
|
||||||
|
Frame 0: splat worker 8/16 finished after 5 local triangles.
|
||||||
|
Frame 0: splat worker 11/16 finished after 9 local triangles.
|
||||||
|
Frame 0: splat worker 1/16 finished after 9 local triangles.
|
||||||
|
Frame 0: splat worker 5/16 finished after 9 local triangles.
|
||||||
|
Frame 0: splat worker 7/16 finished after 9 local triangles.
|
||||||
|
Frame 0: splat worker 10/16 finished after 9 local triangles.
|
||||||
|
Frame 0: splat worker 16/16 finished after 8 local triangles.
|
||||||
|
Frame 0: splat worker 13/16 finished after 8 local triangles.
|
||||||
|
Frame 0: splat worker 4/16 finished after 8 local triangles.
|
||||||
|
Frame 0: splat worker 3/16 finished after 9 local triangles.
|
||||||
|
Frame 0: splat worker 12/16 finished after 8 local triangles.
|
||||||
|
Frame 0: splat worker 2/16 finished after 10 local triangles.
|
||||||
|
Frame 0: splat worker 14/16 finished after 13 local triangles.
|
||||||
|
Frame 0: catalog splatting finished in 7.1 s; writing image...
|
||||||
|
Rendered 26172 images from 54160 catalog stars to /tmp/cpu_psf_dense.png (ok)
|
||||||
|
PSF splats: cached 26172, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
|
||||||
|
build/Release/minkowski_sky --catalog --look-ra-deg 252.5 --look-dec-deg 418.79s user 3.19s system 1472% cpu 28.654 total
|
||||||
|
```
|
||||||
|
|
||||||
|
HIP measurement(加入 timing counters 后;其余参数相同):
|
||||||
|
|
||||||
|
```console
|
||||||
|
make -B PSF_BACKEND=hip SPACETIME=minkowski backend
|
||||||
|
export OMP_NUM_THREADS=16
|
||||||
|
time build/Release/minkowski_sky_hip \
|
||||||
|
--catalog assets/2mass/processed/all_sky/tile_ra252_dec048.csv \
|
||||||
|
--look-ra-deg 252.5 --look-dec-deg -41.5 --fov-deg 0.6 \
|
||||||
|
--width 1920 --height 1920 --coarse-cell-pixels 240 \
|
||||||
|
--exposure 1e15 --psf-fwhm-pixels 2.7 --psf-moffat-beta 3 \
|
||||||
|
--max-cache-psf-flux 1e300 --output /tmp/hip_psf_dense_timing.png --verbose
|
||||||
|
PSF cache ready: 64x64 phases, radius 266 px, relative tail 1e-08, tail abs 1e-06, boundary 1e-07, build 20.915 s
|
||||||
|
Frame 0: tracing 81 initial rays from 128 mesh triangles...
|
||||||
|
Frame 0: initial ray trace finished in 0.008 s.
|
||||||
|
Frame 0: traced 81 lens vertices; starting catalog render.
|
||||||
|
Frame 0: finding and prefetching catalog tiles...
|
||||||
|
Frame 0: catalog prefetch finished (0 candidate tiles).
|
||||||
|
Frame 0: splatting 128 lens triangles...
|
||||||
|
Frame 0: catalog splatting finished in 3.2 s; writing image...
|
||||||
|
Rendered 26172 images from 54160 catalog stars to /tmp/hip_psf_dense_timing.png (ok)
|
||||||
|
PSF splats: cached 26172, cached wing-clipped 0, direct fallbacks 0, discarded below min-Y 0
|
||||||
|
HIP PSF: 26172 events in 2 batches (2 timed); upload 0.000073 s, kernel 2.950377 s, download 0.011256 s
|
||||||
|
build/Release/minkowski_sky_hip --catalog --look-ra-deg 252.5 --look-dec-deg 326.85s user 1.53s system 1314% cpu 24.985 total
|
||||||
|
```
|
||||||
|
|
||||||
|
两次命令都没有设置 `TIMEFMT`;末行是 zsh 默认 `%J %U user %S system %P cpu %*E total`
|
||||||
|
格式的原始输出。它们按顺序运行:先确认没有 renderer benchmark,再执行 HIP,确认其退出后才
|
||||||
|
执行 CPU,期间没有第二个 renderer benchmark。
|
||||||
|
|
||||||
|
`cmp` 比较 `/tmp/cpu_psf_dense.png` 与 `/tmp/hip_psf_dense_timing.png`:`PNG byte comparison:
|
||||||
|
identical`。
|
||||||
|
|
||||||
|
## 结果和决策
|
||||||
|
|
||||||
|
CPU catalog splatting 为 7.1 s,HIP 为 3.2 s(约 2.22x);完整进程的 zsh elapsed time 为
|
||||||
|
28.654 s(CPU)与 24.985 s(HIP)。在 HIP 的 3.2 s 中,两个 batch 的 kernel event 合计
|
||||||
|
2.950377 s;上传为 0.000073 s、最终 HDR 下载为 0.011256 s。因此这个 endpoint 的 GPU 端
|
||||||
|
时间由 kernel 主导,而不是 PCIe。
|
||||||
|
|
||||||
|
但 kernel 还包括每个 event 的圆形 PSF 支持域遍历、cache 权重读取和 HDR atomic 写入。这里只有
|
||||||
|
一个密度/支持域组合,也没有关闭或替代 atomic 的对照,所以不能以此声称 global atomics 的热点
|
||||||
|
竞争是主导瓶颈。结论是保留 direct atomic,并先测密度与支持域扫描;暂不引入 tile-local reduction
|
||||||
|
或其分桶/重排复杂度。
|
||||||
@@ -174,10 +174,16 @@ HIP direct-atomic 的当前实现状态:
|
|||||||
3. RX 9070 上的 640×360 固定 Minkowski catalog HDR 对照逐样本一致;含 2 个 cache 事件与
|
3. RX 9070 上的 640×360 固定 Minkowski catalog HDR 对照逐样本一致;含 2 个 cache 事件与
|
||||||
1 个 direct fallback 的 `tests/data/hip_psf_mixed_catalog.csv` 场同样逐样本一致。GPU failure
|
1 个 direct fallback 的 `tests/data/hip_psf_mixed_catalog.csv` 场同样逐样本一致。GPU failure
|
||||||
会打印阶段与 HIP 错误并令该帧返回失败,不会切换到 CPU backend。
|
会打印阶段与 HIP 错误并令该帧返回失败,不会切换到 CPU backend。
|
||||||
4. 待记录端到端基准的完整命令和原始输出:Git hash、CPU threads、GPU/runtime、块大小、上传、
|
4. 已在 RX 9070 上记录一个受控的高密度 endpoint:54,160 星的 1920×1920、0.6 deg
|
||||||
kernel、download、分辨率、catalog、PSF 及所有 fallback 统计。
|
2MASS tile 产生 26,172 个 cache event、零 direct fallback。HIP 的 2 个 batch 中 H2D
|
||||||
5. 只在 direct atomic profile 显示热点竞争是主要瓶颈时,比较 tile-local reduction;ATRI 的
|
为 0.073 ms、kernel 为 2.950 s、D2H 为 11.256 ms;CPU/HIP PNG 字节一致。完整命令、
|
||||||
整帧收集/重排同样必须以端到端收益证明。
|
原始输出、Git 基线、线程、设备、分辨率和 PSF 参数见
|
||||||
|
`benchmarks/hip_psf_direct_atomic_dense_tile_2026-09-05.md`。
|
||||||
|
5. 该 endpoint 说明当前时间主要位于 device kernel,而非 PCIe;它**不能单独证明** kernel
|
||||||
|
时间主要来自热点 atomic contention(仍混有每 event 的 PSF 遍历与 cache 读取)。因此暂不
|
||||||
|
实现 tile-local reduction。下一步应在资源隔离条件下,以不同屏幕密度和 PSF 支持域分离竞争
|
||||||
|
效应;只有该证据显示竞争为主导且 tile-local 的整帧端到端测量获益时,才保留分桶/重排实现。
|
||||||
|
ATRI 的整帧收集/重排同样必须以端到端收益证明。
|
||||||
6. HIP 路径稳定后再评估 Vulkan 的可移植后端及其他机器的实际设备矩阵。
|
6. HIP 路径稳定后再评估 Vulkan 的可移植后端及其他机器的实际设备矩阵。
|
||||||
|
|
||||||
每一步独立提交,并保持 CPU reference、GPU runtime 层和 shader 资源的提交边界清晰。
|
每一步独立提交,并保持 CPU reference、GPU runtime 层和 shader 资源的提交边界清晰。
|
||||||
+25
-2
@@ -134,6 +134,7 @@ typedef struct {
|
|||||||
#ifdef PSF_BACKEND_HIP
|
#ifdef PSF_BACKEND_HIP
|
||||||
HipPsfSink *hip;
|
HipPsfSink *hip;
|
||||||
char hip_message[256];
|
char hip_message[256];
|
||||||
|
HipPsfTiming hip_timing;
|
||||||
#endif
|
#endif
|
||||||
int failed;
|
int failed;
|
||||||
} PsfEventSink;
|
} PsfEventSink;
|
||||||
@@ -227,6 +228,10 @@ static int psf_event_sink_resume_gpu(PsfEventSink *sink) {
|
|||||||
static int psf_event_sink_destroy(PsfEventSink *sink) {
|
static int psf_event_sink_destroy(PsfEventSink *sink) {
|
||||||
int result = psf_event_sink_finish_for_cpu(sink);
|
int result = psf_event_sink_finish_for_cpu(sink);
|
||||||
#ifdef PSF_BACKEND_HIP
|
#ifdef PSF_BACKEND_HIP
|
||||||
|
if (sink->hip != NULL && hip_psf_sink_get_timing(sink->hip, &sink->hip_timing)) {
|
||||||
|
psf_event_sink_mark_failed(sink, "timing collection");
|
||||||
|
result = -1;
|
||||||
|
}
|
||||||
hip_psf_sink_destroy(sink->hip);
|
hip_psf_sink_destroy(sink->hip);
|
||||||
#endif
|
#endif
|
||||||
free(sink->events);
|
free(sink->events);
|
||||||
@@ -998,6 +1003,8 @@ typedef struct {
|
|||||||
size_t direct_fallbacks;
|
size_t direct_fallbacks;
|
||||||
size_t cached_wing_clipped;
|
size_t cached_wing_clipped;
|
||||||
size_t discarded_below_min_y;
|
size_t discarded_below_min_y;
|
||||||
|
size_t gpu_event_count, gpu_batch_count, gpu_timed_batch_count;
|
||||||
|
double gpu_upload_seconds, gpu_kernel_seconds, gpu_download_seconds;
|
||||||
int failed;
|
int failed;
|
||||||
#ifdef GR_DEBUG
|
#ifdef GR_DEBUG
|
||||||
double max_raw_magnification;
|
double max_raw_magnification;
|
||||||
@@ -1016,6 +1023,12 @@ static void copy_psf_splat_stats(PsfSplatStats *destination,
|
|||||||
.cached_wing_clipped = source.cached_wing_clipped,
|
.cached_wing_clipped = source.cached_wing_clipped,
|
||||||
.direct_fallbacks = source.direct_fallbacks,
|
.direct_fallbacks = source.direct_fallbacks,
|
||||||
.discarded_below_min_y = source.discarded_below_min_y,
|
.discarded_below_min_y = source.discarded_below_min_y,
|
||||||
|
.gpu_event_count = source.gpu_event_count,
|
||||||
|
.gpu_batch_count = source.gpu_batch_count,
|
||||||
|
.gpu_timed_batch_count = source.gpu_timed_batch_count,
|
||||||
|
.gpu_upload_seconds = source.gpu_upload_seconds,
|
||||||
|
.gpu_kernel_seconds = source.gpu_kernel_seconds,
|
||||||
|
.gpu_download_seconds = source.gpu_download_seconds,
|
||||||
#ifdef GR_DEBUG
|
#ifdef GR_DEBUG
|
||||||
.max_raw_magnification = source.max_raw_magnification,
|
.max_raw_magnification = source.max_raw_magnification,
|
||||||
.magnification_clamped_triangles =
|
.magnification_clamped_triangles =
|
||||||
@@ -1162,8 +1175,18 @@ static CatalogSplatStats splat_catalog_triangles(
|
|||||||
stats.cached_wing_clipped += context.cached_wing_clipped;
|
stats.cached_wing_clipped += context.cached_wing_clipped;
|
||||||
stats.discarded_below_min_y += context.discarded_below_min_y;
|
stats.discarded_below_min_y += context.discarded_below_min_y;
|
||||||
}
|
}
|
||||||
if (owns_sink && psf_event_sink_destroy(&owned_sink))
|
if (owns_sink) {
|
||||||
stats.failed = 1;
|
if (psf_event_sink_destroy(&owned_sink))
|
||||||
|
stats.failed = 1;
|
||||||
|
#ifdef PSF_BACKEND_HIP
|
||||||
|
stats.gpu_event_count = owned_sink.hip_timing.event_count;
|
||||||
|
stats.gpu_batch_count = owned_sink.hip_timing.batch_count;
|
||||||
|
stats.gpu_timed_batch_count = owned_sink.hip_timing.timed_batch_count;
|
||||||
|
stats.gpu_upload_seconds = owned_sink.hip_timing.upload_seconds;
|
||||||
|
stats.gpu_kernel_seconds = owned_sink.hip_timing.kernel_seconds;
|
||||||
|
stats.gpu_download_seconds = owned_sink.hip_timing.download_seconds;
|
||||||
|
#endif
|
||||||
|
}
|
||||||
return stats;
|
return stats;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -13,6 +13,11 @@ extern "C" {
|
|||||||
* for PsfCachedEvent chunks, immutable cache weights, and double RGB HDR. */
|
* for PsfCachedEvent chunks, immutable cache weights, and double RGB HDR. */
|
||||||
typedef struct HipPsfSink HipPsfSink;
|
typedef struct HipPsfSink HipPsfSink;
|
||||||
|
|
||||||
|
typedef struct {
|
||||||
|
size_t event_count, batch_count, timed_batch_count;
|
||||||
|
double upload_seconds, kernel_seconds, download_seconds;
|
||||||
|
} HipPsfTiming;
|
||||||
|
|
||||||
int hip_psf_available(char *message, size_t message_size);
|
int hip_psf_available(char *message, size_t message_size);
|
||||||
|
|
||||||
/* Creates a zeroed GPU HDR framebuffer. event_capacity is the reusable upload
|
/* Creates a zeroed GPU HDR framebuffer. event_capacity is the reusable upload
|
||||||
@@ -35,6 +40,9 @@ int hip_psf_sink_submit(HipPsfSink *sink, const PsfCachedEvent *events,
|
|||||||
/* Completes all work and overwrites hdr with double linear RGB device HDR. */
|
/* Completes all work and overwrites hdr with double linear RGB device HDR. */
|
||||||
int hip_psf_sink_finish(HipPsfSink *sink, double *hdr, char *message,
|
int hip_psf_sink_finish(HipPsfSink *sink, double *hdr, char *message,
|
||||||
size_t message_size);
|
size_t message_size);
|
||||||
|
/* Valid after hip_psf_sink_finish(); timings are accumulated without a
|
||||||
|
* per-batch host synchronization. */
|
||||||
|
int hip_psf_sink_get_timing(const HipPsfSink *sink, HipPsfTiming *timing);
|
||||||
void hip_psf_sink_destroy(HipPsfSink *sink);
|
void hip_psf_sink_destroy(HipPsfSink *sink);
|
||||||
|
|
||||||
#ifdef __cplusplus
|
#ifdef __cplusplus
|
||||||
|
|||||||
+74
-3
@@ -5,6 +5,14 @@
|
|||||||
#include <cmath>
|
#include <cmath>
|
||||||
#include <cstdio>
|
#include <cstdio>
|
||||||
#include <limits>
|
#include <limits>
|
||||||
|
#include <vector>
|
||||||
|
|
||||||
|
struct HipPsfBatchTiming {
|
||||||
|
hipEvent_t upload_start = nullptr, upload_end = nullptr;
|
||||||
|
hipEvent_t kernel_start = nullptr, kernel_end = nullptr;
|
||||||
|
};
|
||||||
|
|
||||||
|
enum { HIP_PSF_MAX_TIMED_BATCHES = 64 };
|
||||||
|
|
||||||
struct HipPsfSink {
|
struct HipPsfSink {
|
||||||
PsfCachedEvent *events = nullptr;
|
PsfCachedEvent *events = nullptr;
|
||||||
@@ -14,6 +22,10 @@ struct HipPsfSink {
|
|||||||
size_t event_capacity = 0, hdr_values = 0;
|
size_t event_capacity = 0, hdr_values = 0;
|
||||||
int width = 0, height = 0, phase_resolution = 0, radius_pixels = 0;
|
int width = 0, height = 0, phase_resolution = 0, radius_pixels = 0;
|
||||||
double max_radius_pixels = 0.0;
|
double max_radius_pixels = 0.0;
|
||||||
|
hipEvent_t download_start = nullptr, download_end = nullptr;
|
||||||
|
std::vector<HipPsfBatchTiming> batches;
|
||||||
|
size_t measured_batches = 0;
|
||||||
|
HipPsfTiming timing = {};
|
||||||
};
|
};
|
||||||
|
|
||||||
static int report(hipError_t status, char *message, size_t message_size) {
|
static int report(hipError_t status, char *message, size_t message_size) {
|
||||||
@@ -25,6 +37,15 @@ static void ok(char *message, size_t message_size) {
|
|||||||
if (message && message_size) std::snprintf(message, message_size, "ok");
|
if (message && message_size) std::snprintf(message, message_size, "ok");
|
||||||
}
|
}
|
||||||
|
|
||||||
|
static void destroy_batch_timing(HipPsfBatchTiming *timing) {
|
||||||
|
if (!timing) return;
|
||||||
|
(void)hipEventDestroy(timing->upload_start);
|
||||||
|
(void)hipEventDestroy(timing->upload_end);
|
||||||
|
(void)hipEventDestroy(timing->kernel_start);
|
||||||
|
(void)hipEventDestroy(timing->kernel_end);
|
||||||
|
*timing = {};
|
||||||
|
}
|
||||||
|
|
||||||
__device__ static size_t weight_index(int phase_resolution, int radius_pixels,
|
__device__ static size_t weight_index(int phase_resolution, int radius_pixels,
|
||||||
int phase_x, int phase_y, int offset_x, int offset_y) {
|
int phase_x, int phase_y, int offset_x, int offset_y) {
|
||||||
const size_t nodes = (size_t)phase_resolution + 1;
|
const size_t nodes = (size_t)phase_resolution + 1;
|
||||||
@@ -118,6 +139,8 @@ extern "C" int hip_psf_sink_create(HipPsfSink **out, int width, int height,
|
|||||||
sink->radius_pixels = cache->radius_pixels; sink->max_radius_pixels = cache->max_radius_pixels;
|
sink->radius_pixels = cache->radius_pixels; sink->max_radius_pixels = cache->max_radius_pixels;
|
||||||
const size_t weight_count = nodes * nodes * side * side;
|
const size_t weight_count = nodes * nodes * side * side;
|
||||||
if (report(hipStreamCreate(&sink->stream), message, message_size) ||
|
if (report(hipStreamCreate(&sink->stream), message, message_size) ||
|
||||||
|
report(hipEventCreate(&sink->download_start), message, message_size) ||
|
||||||
|
report(hipEventCreate(&sink->download_end), message, message_size) ||
|
||||||
report(hipMalloc(&sink->events, event_capacity * sizeof *sink->events), message, message_size) ||
|
report(hipMalloc(&sink->events, event_capacity * sizeof *sink->events), message, message_size) ||
|
||||||
report(hipMalloc(&sink->weights, weight_count * sizeof *sink->weights), message, message_size) ||
|
report(hipMalloc(&sink->weights, weight_count * sizeof *sink->weights), message, message_size) ||
|
||||||
report(hipMalloc(&sink->hdr, sink->hdr_values * sizeof *sink->hdr), message, message_size) ||
|
report(hipMalloc(&sink->hdr, sink->hdr_values * sizeof *sink->hdr), message, message_size) ||
|
||||||
@@ -137,13 +160,35 @@ extern "C" int hip_psf_sink_submit(HipPsfSink *sink, const PsfCachedEvent *event
|
|||||||
if (!event_count) { ok(message, message_size); return 0; }
|
if (!event_count) { ok(message, message_size); return 0; }
|
||||||
const unsigned int threads = 128;
|
const unsigned int threads = 128;
|
||||||
const size_t blocks = (event_count + threads - 1) / threads;
|
const size_t blocks = (event_count + threads - 1) / threads;
|
||||||
|
HipPsfBatchTiming timing;
|
||||||
|
const int timed = sink->batches.size() < HIP_PSF_MAX_TIMED_BATCHES;
|
||||||
if (blocks > std::numeric_limits<unsigned int>::max() ||
|
if (blocks > std::numeric_limits<unsigned int>::max() ||
|
||||||
report(hipMemcpyAsync(sink->events, events, event_count * sizeof *events, hipMemcpyHostToDevice, sink->stream), message, message_size))
|
(timed && (report(hipEventCreate(&timing.upload_start), message, message_size) ||
|
||||||
|
report(hipEventCreate(&timing.upload_end), message, message_size) ||
|
||||||
|
report(hipEventCreate(&timing.kernel_start), message, message_size) ||
|
||||||
|
report(hipEventCreate(&timing.kernel_end), message, message_size) ||
|
||||||
|
report(hipEventRecord(timing.upload_start, sink->stream), message, message_size))) ||
|
||||||
|
report(hipMemcpyAsync(sink->events, events, event_count * sizeof *events,
|
||||||
|
hipMemcpyHostToDevice, sink->stream), message, message_size) ||
|
||||||
|
(timed && (report(hipEventRecord(timing.upload_end, sink->stream), message, message_size) ||
|
||||||
|
report(hipEventRecord(timing.kernel_start, sink->stream), message, message_size)))) {
|
||||||
|
destroy_batch_timing(&timing);
|
||||||
return -1;
|
return -1;
|
||||||
|
}
|
||||||
hipLaunchKernelGGL(splat_kernel, dim3((unsigned int)blocks), dim3(threads), 0, sink->stream,
|
hipLaunchKernelGGL(splat_kernel, dim3((unsigned int)blocks), dim3(threads), 0, sink->stream,
|
||||||
sink->events, event_count, sink->width, sink->height, sink->weights,
|
sink->events, event_count, sink->width, sink->height, sink->weights,
|
||||||
sink->phase_resolution, sink->radius_pixels, sink->max_radius_pixels, sink->hdr);
|
sink->phase_resolution, sink->radius_pixels, sink->max_radius_pixels, sink->hdr);
|
||||||
if (report(hipGetLastError(), message, message_size)) return -1;
|
if (report(hipGetLastError(), message, message_size) ||
|
||||||
|
(timed && report(hipEventRecord(timing.kernel_end, sink->stream), message, message_size))) {
|
||||||
|
destroy_batch_timing(&timing);
|
||||||
|
return -1;
|
||||||
|
}
|
||||||
|
if (timed) {
|
||||||
|
sink->batches.push_back(timing);
|
||||||
|
++sink->timing.timed_batch_count;
|
||||||
|
}
|
||||||
|
++sink->timing.batch_count;
|
||||||
|
sink->timing.event_count += event_count;
|
||||||
ok(message, message_size); return 0;
|
ok(message, message_size); return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -161,13 +206,39 @@ extern "C" int hip_psf_sink_load_hdr(HipPsfSink *sink, const double *hdr,
|
|||||||
|
|
||||||
extern "C" int hip_psf_sink_finish(HipPsfSink *sink, double *hdr, char *message, size_t message_size) {
|
extern "C" int hip_psf_sink_finish(HipPsfSink *sink, double *hdr, char *message, size_t message_size) {
|
||||||
if (!sink || !hdr) { if (message && message_size) std::snprintf(message, message_size, "invalid HIP PSF HDR download"); return -1; }
|
if (!sink || !hdr) { if (message && message_size) std::snprintf(message, message_size, "invalid HIP PSF HDR download"); return -1; }
|
||||||
if (report(hipMemcpyAsync(hdr, sink->hdr, sink->hdr_values * sizeof *hdr, hipMemcpyDeviceToHost, sink->stream), message, message_size) ||
|
if (report(hipEventRecord(sink->download_start, sink->stream), message, message_size) ||
|
||||||
|
report(hipMemcpyAsync(hdr, sink->hdr, sink->hdr_values * sizeof *hdr, hipMemcpyDeviceToHost, sink->stream), message, message_size) ||
|
||||||
|
report(hipEventRecord(sink->download_end, sink->stream), message, message_size) ||
|
||||||
report(hipStreamSynchronize(sink->stream), message, message_size)) return -1;
|
report(hipStreamSynchronize(sink->stream), message, message_size)) return -1;
|
||||||
|
float milliseconds = 0.0f;
|
||||||
|
for (size_t i = sink->measured_batches; i < sink->batches.size(); ++i) {
|
||||||
|
const HipPsfBatchTiming &batch = sink->batches[i];
|
||||||
|
if (report(hipEventElapsedTime(&milliseconds, batch.upload_start, batch.upload_end),
|
||||||
|
message, message_size)) return -1;
|
||||||
|
sink->timing.upload_seconds += milliseconds * 1e-3;
|
||||||
|
if (report(hipEventElapsedTime(&milliseconds, batch.kernel_start, batch.kernel_end),
|
||||||
|
message, message_size)) return -1;
|
||||||
|
sink->timing.kernel_seconds += milliseconds * 1e-3;
|
||||||
|
}
|
||||||
|
sink->measured_batches = sink->batches.size();
|
||||||
|
if (report(hipEventElapsedTime(&milliseconds, sink->download_start, sink->download_end),
|
||||||
|
message, message_size)) return -1;
|
||||||
|
sink->timing.download_seconds += milliseconds * 1e-3;
|
||||||
ok(message, message_size); return 0;
|
ok(message, message_size); return 0;
|
||||||
}
|
}
|
||||||
|
|
||||||
|
extern "C" int hip_psf_sink_get_timing(const HipPsfSink *sink, HipPsfTiming *timing) {
|
||||||
|
if (!sink || !timing) return -1;
|
||||||
|
*timing = sink->timing;
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
extern "C" void hip_psf_sink_destroy(HipPsfSink *sink) {
|
extern "C" void hip_psf_sink_destroy(HipPsfSink *sink) {
|
||||||
if (!sink) return;
|
if (!sink) return;
|
||||||
|
for (HipPsfBatchTiming &timing : sink->batches)
|
||||||
|
destroy_batch_timing(&timing);
|
||||||
|
(void)hipEventDestroy(sink->download_start);
|
||||||
|
(void)hipEventDestroy(sink->download_end);
|
||||||
(void)hipFree(sink->events); (void)hipFree(sink->weights); (void)hipFree(sink->hdr);
|
(void)hipFree(sink->events); (void)hipFree(sink->weights); (void)hipFree(sink->hdr);
|
||||||
(void)hipStreamDestroy(sink->stream); delete sink;
|
(void)hipStreamDestroy(sink->stream); delete sink;
|
||||||
}
|
}
|
||||||
@@ -301,6 +301,13 @@ void psf_kernel_cache_report(const PsfKernelCache *cache,
|
|||||||
stats == NULL ? 0u : stats->cached_wing_clipped,
|
stats == NULL ? 0u : stats->cached_wing_clipped,
|
||||||
stats == NULL ? 0u : stats->direct_fallbacks,
|
stats == NULL ? 0u : stats->direct_fallbacks,
|
||||||
stats == NULL ? 0u : stats->discarded_below_min_y);
|
stats == NULL ? 0u : stats->discarded_below_min_y);
|
||||||
|
if (stats != NULL && stats->gpu_batch_count != 0)
|
||||||
|
fprintf(stream, "HIP PSF: %zu events in %zu batches (%zu timed); upload %.6f s, "
|
||||||
|
"kernel %.6f s, download %.6f s\n",
|
||||||
|
stats->gpu_event_count, stats->gpu_batch_count,
|
||||||
|
stats->gpu_timed_batch_count,
|
||||||
|
stats->gpu_upload_seconds, stats->gpu_kernel_seconds,
|
||||||
|
stats->gpu_download_seconds);
|
||||||
if (stats != NULL && stats->discarded_below_min_y != 0)
|
if (stats != NULL && stats->discarded_below_min_y != 0)
|
||||||
fputs("Warning: --psf-min-y discarded one or more PSF events.\n", stream);
|
fputs("Warning: --psf-min-y discarded one or more PSF events.\n", stream);
|
||||||
#ifdef GR_DEBUG
|
#ifdef GR_DEBUG
|
||||||
|
|||||||
@@ -26,6 +26,8 @@ typedef struct {
|
|||||||
size_t cached_wing_clipped;
|
size_t cached_wing_clipped;
|
||||||
size_t direct_fallbacks;
|
size_t direct_fallbacks;
|
||||||
size_t discarded_below_min_y;
|
size_t discarded_below_min_y;
|
||||||
|
size_t gpu_event_count, gpu_batch_count, gpu_timed_batch_count;
|
||||||
|
double gpu_upload_seconds, gpu_kernel_seconds, gpu_download_seconds;
|
||||||
#ifdef GR_DEBUG
|
#ifdef GR_DEBUG
|
||||||
double max_raw_magnification;
|
double max_raw_magnification;
|
||||||
size_t magnification_clamped_triangles;
|
size_t magnification_clamped_triangles;
|
||||||
|
|||||||
@@ -46,6 +46,16 @@ int main() {
|
|||||||
std::free(cpu_hdr); std::free(gpu_hdr); psf_kernel_cache_destroy(&cache);
|
std::free(cpu_hdr); std::free(gpu_hdr); psf_kernel_cache_destroy(&cache);
|
||||||
return 1;
|
return 1;
|
||||||
}
|
}
|
||||||
|
HipPsfTiming timing = {};
|
||||||
|
if (hip_psf_sink_get_timing(sink, &timing) || timing.event_count != 3 ||
|
||||||
|
timing.batch_count != 1 || timing.timed_batch_count != 1 ||
|
||||||
|
timing.upload_seconds < 0.0 || timing.kernel_seconds < 0.0 ||
|
||||||
|
timing.download_seconds < 0.0) {
|
||||||
|
std::fputs("HIP PSF timing accounting failed\n", stderr);
|
||||||
|
hip_psf_sink_destroy(sink);
|
||||||
|
std::free(cpu_hdr); std::free(gpu_hdr); psf_kernel_cache_destroy(&cache);
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
double max_abs = 0.0, max_rel = 0.0;
|
double max_abs = 0.0, max_rel = 0.0;
|
||||||
for (size_t i = 0; i < values; ++i) {
|
for (size_t i = 0; i < values; ++i) {
|
||||||
const double absolute = std::fabs(cpu_hdr[i] - gpu_hdr[i]);
|
const double absolute = std::fabs(cpu_hdr[i] - gpu_hdr[i]);
|
||||||
|
|||||||
Reference in new issue
Block a user