Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.
- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
18 KiB
Rendering guide
See README.md for the project overview and a first 2MASS render,
and build.md for compilation. Commands below run from the repository
root and use the default Release binaries. Run a binary with --help for its
complete option list and defaults.
Catalog input
Use --all-sky-catalog assets/2mass/processed/all_sky for the tiled 2MASS
catalog, or --catalog PATH for a single processed CSV. Acquisition and
processing are documented in assets/2mass/README.md.
2MASS blackbody amplitudes need a much larger display exposure than the
synthetic test catalog; --exposure 1e15 is a starting point, to be adjusted
for the field and camera.
Single-frame camera
Both backends accept the same instantaneous camera parameters. Position and
velocity use the backend's coordinates; velocity means dx/dt, dy/dt, dz/dt,
not a local physical speed. The analytic single-frame event is at t=0.
| Option | Meaning / default |
|---|---|
--observer-position X Y Z |
Coordinate position; if look is omitted, point toward the origin |
--look-ra-deg RA, --look-dec-deg DEC |
Coordinate look direction; missing angle defaults to RA=90°, Dec=-90° |
--observer-radius R |
Positive radius used only to infer position, default 30; conflicts with explicit position |
--observer-velocity VX VY VZ |
Coordinate velocity, default 0 0 0; must be future-timelike in the local metric |
--camera-roll-deg ANGLE |
Rotate the camera up axis toward its right axis, default 0° |
When position and look are both supplied, they are independent. Look alone
implies position -R * look_direction, in either backend. Radius alone uses
the default look. With no position/look/radius options, Minkowski keeps the
origin camera and Schwarzschild keeps its radius-30 camera on +Z; both look
along -Z. Velocity and roll alone do not alter these defaults.
An explicit origin position requires a look angle. Position-derived directions
on the Z axis use RA=0° to fix the pole orientation deterministically.
The axes are right-handed ICRS: X=RA 0°/Dec 0°, Y=RA 90°/Dec 0°,
Z=Dec +90°. The coordinate look vector (cos(dec) cos(ra), cos(dec) sin(ra), sin(dec)) is projected into the moving camera's rest space. The coordinate
north and west seeds are then orthogonalized there to form up and right.
Consequently the look angles specify a projected coordinate orientation;
they do not lock a lensed sky object to the image center. Positive roll uses
up' = cos(roll) up + sin(roll) right,
right' = -sin(roll) up + cos(roll) right.
No static observer is needed. The metric determines whether (1,VX,VY,VZ)
is timelike; there is no Euclidean |V| < 1 check in curved coordinates.
Invalid velocities are rejected before loading catalogs or building the PSF
cache, with the velocity and its metric norm in the diagnostic.
--observer-inward-speed has been removed.
Single-frame camera options cannot be combined with --observer-track,
--frames-dir, or --lens-map-input.
Schwarzschild uses Cartesian ingoing Kerr–Schild coordinates with M=1.
Cameras at and inside the horizon r=2 are allowed with a valid timelike
coordinate velocity. The current backend excludes camera positions at or
inside its capture cutoff r=1.5; its finite escape radius is 256.
These remain analytic demonstration settings, not criteria for future NR data.
Zero coordinate velocity at or inside the horizon is not timelike and is rejected.
The following complete examples use the bundled synthetic catalog:
make -j PSF_BACKEND=cpu SPACETIME=minkowski backend
mkdir -p output/imgs
./build/Release/minkowski_sky --catalog assets/sky_grid_5deg.csv \
--observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
--look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
--output output/imgs/minkowski_moving.png
make -j PSF_BACKEND=cpu SPACETIME=schwarzschild backend
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
--observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
--look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
--output output/imgs/schwarzschild_off_axis.png
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
--observer-position 1.75 0 0 --observer-velocity -0.5 0 0 \
--look-ra-deg 0 --look-dec-deg 0 \
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
--output output/imgs/schwarzschild_inside_horizon.png
The last example points outward from a camera moving inward inside the horizon. It needs no CSV trajectory or movie wrapper. These small images are camera checks; increase resolution and refinement for production renders.
Movie image sequences
Movie mode reads an observer-track CSV containing coordinate time, proper time, Cartesian position, and a full four-by-four tetrad (21 columns). The track must cover the requested frame times. The following generates a sample accelerating Minkowski trajectory and renders it against the 2MASS sky:
mkdir -p output/imgs
./build/Release/minkowski_sky --write-minkowski-accel-track output/minkowski_accel_2s.csv \
--duration 2 --fps 30 --proper-acceleration 1.52
./build/Release/minkowski_sky --observer-track output/minkowski_accel_2s.csv \
--all-sky-catalog assets/2mass/processed/all_sky \
--frames-dir output/imgs --frames-prefix minkowski_accel \
--start-time 0 --duration 2 --fps 30 --exposure 1e15
The inclusive time interval produces frames minkowski_accel_000000.png
through minkowski_accel_000060.png. Exposure may need adjustment as Doppler
shifts change the brightness. The trajectory generator is a flat-spacetime
example; other camera motions are supplied through observer-track CSVs.
Output is an image sequence; video encoding is a separate step.
Movie rays from all frames share a newest-to-oldest coordinate-time sweep.
--slab-duration sets its time-slab width (default 64). The current analytic
backends use logical slabs without metric I/O; Nmesh metric loading remains
future work.
Reuse a completed lens map
--lens-map-output FILE saves the finalized inverse-lens mesh after tracing
and refinement, while also rendering the image normally. It stores the
geometry, frequency shifts, terminal statuses, triangle topology, and frame
metadata in a versioned .grlens file. This allows later catalog, PSF, and
exposure changes without repeating ray tracing.
mkdir -p output/imgs output/maps
./build/Release/schwarzschild_sky \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 640 --height 360 --coarse-cell-pixels 8 --fov-deg 60 \
--refine-max-level 2 --exposure 1e15 \
--lens-map-output output/maps/schwarzschild.grlens \
--output output/imgs/schwarzschild_trace.png
./build/Release/schwarzschild_sky \
--lens-map-input output/maps/schwarzschild.grlens \
--all-sky-catalog assets/2mass/processed/all_sky \
--psf-fwhm-pixels 6 --psf-moffat-beta 3 --exposure 1e15 \
--output output/imgs/schwarzschild_restyled.png
Import and export are mutually exclusive. Imported maps retain their original pixel dimensions, horizontal FOV, and mesh resolution; they cannot be refined on import. Explicit conflicting dimensions or FOV are rejected. The reader checks the format version, finite values, unit directions, triangle indices, and per-frame CRCs.
A movie export stores all final frame meshes in one file. To render it again,
pass --lens-map-input FILE, the catalog, --frames-dir DIR, and
--frames-prefix NAME; no observer track is needed on import.
Adaptive image mesh refinement
Adaptive refinement is disabled by default (--refine-max-level 0), so the
existing coarse-mesh renders remain unchanged. When enabled, its defaults are
an absolute direction error of 1e-3 degrees, relative error 0.1, minimum
long edge 0.5 pixels, minimum area 0.25 pixel-squared, and a provisional
minimum discrete-Jacobian magnitude of 1e-3. Each value can be overridden
independently:
--refine-max-level N
--refine-angle-abs-deg D
--refine-angle-rel R
--refine-jacobian-min J
--refine-min-edge-pixels P
--refine-min-area-pixels2 A
N caps the triangle refinement level. Let e be the angle between the
traced longest-edge midpoint direction and the normalized endpoint
interpolation, and let s be the angle between those two endpoint camera
directions. Both are evaluated internally in radians; the absolute CLI
threshold D is specified in degrees and converted before comparison. s
is the angular geometric size of the image triangle's test edge, not a
source-sky/lens-map length. A locally escaped triangle is split only when
both e > D_rad (the converted --refine-angle-abs-deg D) and
e / max(s, 1e-15) > --refine-angle-rel. P and A
prevent selecting a leaf already at or below the requested image-plane
long-edge and area scales.
Triangles whose three vertices disagree between capture and escape are split
independently of the direction-error thresholds, allowing the mesh to follow a
shadow boundary.
Independently of the midpoint geometry test, an all-escaped triangle also
computes the discrete lens Jacobian
J = Omega_source / Omega_image. Both signed solid angles use
2 atan2(dot(a, cross(b,c)), 1 + dot(a,b) + dot(b,c) + dot(c,a)), with the
ordered camera directions for Omega_image and their traced infinity
directions for Omega_source. A J-driven split requires both a shared
image edge whose incident triangles have opposite nonzero signs of J and
min(abs(J_left), abs(J_right)) < --refine-jacobian-min. It then requests
that shared edge on both leaves. Thus |J| bounds the fold selection instead
of widening it as a standalone critical-curve band. A negative sign is
physical parity and is retained. The 1e-3 default is deliberately
provisional and should be tuned with the small Schwarzschild refinement
diagnostic before being treated as a production threshold.
For a small mesh diagnostic using the synthetic catalog:
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
--width 48 --height 48 --coarse-cell-pixels 24 --fov-deg 40 \
--refine-max-level 1 --refine-angle-abs-deg 0.001 \
--refine-angle-rel 0.001 --refine-jacobian-min 0.001 \
--refine-min-edge-pixels 1 \
--refine-min-area-pixels2 1 --draw-mesh \
--output output/imgs/schwarzschild_refinement.png
For movies, each refinement generation completes the full newest-to-oldest
time-slab sweep before any probe becomes a mesh vertex. Newly added vertices
are therefore traced only by the next generation; the renderer never returns
to a slab that has already been released. At the start of every generation,
newly inserted vertices and geometry-only longest-edge probes for its new
leaves are collected together, so both ray sets use the same parallel
RayPool pass.
Camera directions and optics
Catalog directions and --look-ra-deg/--look-dec-deg use standard
right-handed ICRS Cartesian axes: +X is RA 0 degrees/Dec 0 degrees, +Y is
RA 90 degrees/Dec 0 degrees, and +Z is the north celestial pole. The local
camera axes are forward, celestial north, and celestial west, so an image with
north up has decreasing RA to the right. The defaults preserve the original
-Z view. --exposure converts a catalog's physical flux
normalization to the prototype HDR scale. The current synthetic catalog is
calibrated for default exposure 1e-3; a 2MASS blackbody normalization in
steradians requires a much larger display exposure such as the example above.
The optics path
integrates each fitted Planck spectrum through CIE 1931 color-matching functions
and converts the resulting radiance to linear sRGB; it does not use an empirical
color-temperature RGB approximation.
Point sources use a flux-normalized circular Moffat PSF by default
(--psf-fwhm-pixels 2.7 --psf-moffat-beta 4.5). The FWHM matches the former
1.15-pixel Gaussian core while the Moffat wings remain continuous; both values
are display/optics calibration parameters.
The default renderer builds one immutable, process-wide 64-by-64 sub-pixel
Moffat lookup kernel. Its weights are pixel-area integrals and are bilinearly
interpolated between phase tables. The renderer reports its build time and
cached/direct-fallback image counts. Pass --psf-direct to use the slower
8-point quadrature reference evaluator for regression comparisons; an image
whose required HDR-tail support exceeds the cache radius selects that reference
path automatically.
--psf-relative-tail R controls the maximum omitted PSF tail fraction per
image; it defaults to 1e-8. Larger values intentionally shorten the Moffat
support and rebuild the immutable cache at the corresponding radius, which is
useful when a faster, lower-fidelity render is acceptable.
--psf-min-y Y defaults to 0 (disabled). A positive value is a linear-HDR
luminance cutoff: a PSF stops where its continuous Moffat Y profile falls below
Y, and an image whose central value is already below Y is omitted. The
final PSF report counts such omitted events and emits a warning when any occur.
For bounded preview renders, --max-cache-psf-flux F (default 1) allows
images with 1 < flux <= F to use the existing cache instead of the direct
evaluator. This does not rebuild or enlarge the cache: the image retains its
true core flux, color, and sub-pixel position, while its Moffat wing is clipped
at the existing cache radius. Images above F retain the direct fallback.
--max-magnification M (default unlimited) caps the per-triangle rendering
magnification before flux is formed; it is an explicit preview approximation.
Fast preview mode
--fast-mode replaces the per-event PSF splat with a two-stage approximation:
each point-source image is deposited as a delta into an N×N supersampled HDR
buffer (--fast-supersample N, default 2), then the whole frame is convolved
once with a single global Moffat kernel and averaged down by the N×N block.
The kernel is the pixel-area integral of the requested Moffat at the
supersampled scale, so the requested FWHM and beta are preserved by the
downsample itself.
--fast-deposit nearest (default) deposits into the single nearest
supersampled pixel. It keeps the PSF shape exactly but quantizes the image
position to 1/(2N) output pixels. --fast-deposit bilinear deposits into 4
adjacent pixels; it preserves the continuous centroid but broadens the profile
(measured +4.5%/+1.9%/+1.0% FWHM at N=2/3/4). The deposition experiment and its
raw output are in
benchmarks/fast_mode_deposit_2026-09-18.md.
Fast mode is available only in the CPU PSF backend. It ignores
--max-cache-psf-flux and uses one global kernel radius for every event, so a
bright event whose requested support exceeds that radius is wing-clipped and
counted in the report. --psf-min-y is still applied per event. Fast mode is a
preview approximation, not the physically exact per-event PSF path.
In CPU builds the global convolution is a zero-padded FFTW linear convolution
(fftw3/fftw3_omp, double precision). The immutable kernel spectrum, FFTW
plans, and scratch buffers are built once when the accumulator is initialized
and reused for every frame; the disposable spatial reference remains only for
tests and benchmarks. Startup prints one Fast FFTW: line with the minimum
linear-convolution extent, the selected smooth FFT dimensions, worker count,
plan mode, one-time plan and kernel-transform time, and scratch bytes.
--verbose additionally prints per-frame stage timings (zero/pack, forward FFT,
frequency multiply, inverse FFT, crop/downsample, total). The default FFTW
planning mode is FFTW_ESTIMATE; FFTW_MEASURE can cut steady-state execution
substantially but costs a much longer one-time plan for large frames, as
recorded in
benchmarks/fast_mode_fftw_2026-09-25.md.
Progress and diagnostics
The PSF-cache completion line is printed before tracing and catalog splatting
begin. For long renders, pass --verbose to print catalog-prefetch state,
splat-worker local heartbeats (8, 16, 32, ... completed triangles per worker),
and image-write boundaries. The worker heartbeats use neither global progress
accounting nor cross-worker synchronization.
Verbose ray-trace output reports the initial mesh trace and refinement stages
for single frames; movie mode additionally reports each generation's sample
count and each time slab's activation and terminal-ray summary.
Movie renders always print one summary per time slab; --verbose also prints
the ray counts before each slab is loaded.
Pass --draw-mesh to alpha-composite image-plane triangle edges as
one-pixel-wide 0.5 linear-gray diagnostic lines at 0.5 opacity. The line
rasterizer uses coverage-based antialiasing.
HDR output
To preserve a single-frame render for later exposure and tone-mapping work,
build with ENABLE_HDR=1 as described in build.md. This produces build/Release/schwarzschild_sky,
which accepts --hdr-output. This switch writes the HDR file next to the
ordinary output, replacing its extension with _HDR.fits; for example,
--output output/imgs/ring.png --hdr-output writes
output/imgs/ring_HDR.fits. It writes the pre-tone-mapping RGB
framebuffer as a three-plane, 32-bit float FITS image. Values remain linear HDR
at the renderer's arbitrary scale; no tone mapping or per-frame normalization
is applied. The ordinary
binaries do not contain this option or writer.
The FITS header describes a synthetic 8640-by-5760, 36-by-24 mm full-frame
sensor with 4.1667 um pixels. Each render records its active centered crop and
derives FOCALLEN from that render's width and horizontal --fov-deg; these
camera fields support plate-solving workflows but do not calibrate flux.