6.4 KiB
CPU PSF lookup-table optimization plan
Goal
Speed up the current CPU point-source rendering path when the physically correct lens map produces very many star images. This plan changes only how a known image's Moffat PSF is evaluated and accumulated. It does not change the catalog representation, inverse lens map, number of images, frequency shift, magnification, adaptive-mesh work, or use of the GPU.
The target is to replace the per-covered-pixel Moffat pow() evaluation with
a lookup of the same continuously translated PSF. It must preserve continuous
sub-pixel image positions and have an explicit, testable bound on PSF-tail
error, including for bright sources.
Interface and ownership
Add a PsfKernelCache owned by the rendering settings for the lifetime of one
program invocation. Exactly one cache is built after command-line parsing,
using the process-wide values of:
--psf-fwhm-pixels;--psf-moffat-beta;- a named PSF-tail/display-error configuration;
- a fixed sub-pixel phase resolution.
It is immutable after construction, shared read-only by splat workers, and is destroyed after rendering. It is not serialized to disk and no multi-parameter cache or LRU is required: the current program has one PSF parameter pair for the entire image or movie.
splat_moffat() gains a cache-aware path (or is replaced by an equivalently
named cache-aware function). A direct-evaluation implementation remains
available as a reference and as a fallback if cache construction fails.
Kernel construction
For each sub-pixel phase (fx, fy), where
fx = x - floor(x) and fy = y - floor(y), build a scalar kernel containing
the pixel-area integral of the normalized circular Moffat profile over each
integer pixel footprint. Kernel weights are independent of color and flux.
Use a fixed, documented numerical quadrature whose convergence is verified
against a higher-accuracy reference.
At runtime, retain the image event's full-precision (x, y). Select the four
neighboring phase tables and bilinearly interpolate their weights. Do not
round to a nearest phase table. Normalize every interior phase kernel so its
weight sum is one; an image clipped by the framebuffer boundary retains the
current behavior of losing the off-frame contribution rather than being
renormalized.
Start with a 64 by 64 phase grid. This is an initial implementation value, not an accuracy claim: increase it if the phase-continuity and reference-image tests below do not meet their thresholds.
Tail support and bright sources
A Moffat profile has infinite support, so a finite renderer needs an explicit finite-support approximation. Do not rely solely on the current fixed relative omitted-flux fraction for every source. For each image event, choose the support radius from its final post-lens, post-frequency-shift, post-exposure flux so that both of these hold:
- omitted integrated flux is below a configured absolute frame error budget;
- the PSF value at the support boundary is below a configured per-pixel HDR error budget.
The cache contains weights through the largest radius required by the selected configuration. A dim image uses a cropped subset; an unusually bright image uses a larger subset. If an event requires a radius beyond the cache maximum, use the direct Moffat evaluator for that event and record it in diagnostics. This prevents a visible hard cutoff in bright-star wings without silently discarding their flux.
The selected tail budgets, maximum cache radius, and fallback count must be reported with the render diagnostics. Their final numerical values are to be chosen from the validation experiments, not inferred from a performance goal.
Accumulation and parallel behavior
For every image event, compute color and final flux exactly as in the existing path. Then multiply the scalar cached PSF weights by that color and flux and add them to the worker's existing private HDR buffer. Preserve the current private-HDR ownership, memory cap, reduction ordering, and serial fallback; this plan does not introduce atomics, locks, or a different parallel partitioning.
Validation and benchmark
Add tests and a reproducible benchmark before making the lookup path default:
- Compare cached and high-accuracy direct pixel-area integration for a grid of FWHM, beta, phase, flux, and support-radius cases. Check kernel sum, per-pixel error, and omitted-flux bound.
- Translate a star in sub-pixel increments across pixel boundaries. Verify continuous image values and absence of phase quantization jumps.
- Test bright stars whose required support is larger than the ordinary cache crop, including the direct-fallback case. Verify the configured tail-error bound and no visible-radius discontinuity relative to the reference.
- Compare serial and OpenMP outputs under the existing private-HDR contract. Preserve the current expected identity/tolerance policy explicitly.
- Benchmark the Schwarzschild 1080p, 60-degree, FWHM 1.0 all-sky case and separately report cache-build time, number of image events, cached-splat time, direct-fallback count, and total render time. Compare output against the direct evaluator with the same error budgets.
The lookup path becomes the default only after it meets the agreed visual and numerical tolerances and demonstrates a material speedup in the PSF-dominated case. The direct evaluator remains a selectable regression reference.
Implementation record (2026-08-28)
Implemented as one immutable PsfKernelCache per invocation, built after CLI
parsing and passed read-only through the existing frame splat workers. The
cache uses 64 by 64 phase intervals (65 nodes on each axis), 4-point
Gauss-Legendre pixel-area quadrature, and bilinear phase interpolation. The
--psf-direct regression mode uses an independent 8-point quadrature
reference; events whose requested support exceeds the cache use that path
automatically.
The selected HDR budgets are omitted flux 1e-6 and boundary contribution
1e-7; the former 1e-8 relative tail remains part of the support selection.
The cache reports build time, maximum radius, cached splat count, and direct
fallback count. tests/test_frame.c covers cache construction, phases near
pixel boundaries, comparison with the direct reference, and bright-event
fallback. The float-HDR default-PSF acceptance result and the 1080p full-sky
performance result are recorded in
benchmarks/2mass_all_sky_psf_lookup_2026-08-28.md.