Files
GR-raytracing/usage.md
T

35 KiB
Raw Blame History

Rendering guide

See README.md for the project overview and a first 2MASS render, and build.md for compilation. Commands below run from the repository root and use the default Release binaries. Run a binary with --help for its complete option list and defaults.

Catalog input

Use --all-sky-catalog assets/2mass/processed/all_sky for the tiled 2MASS catalog, or --catalog PATH for a single processed CSV. Acquisition and processing are documented in assets/2mass/README.md. 2MASS blackbody amplitudes need a much larger display exposure than the synthetic test catalog; --exposure 1e15 is a starting point, to be adjusted for the field and camera.

Single-frame camera

Both backends accept the same instantaneous camera parameters. Position and velocity use the backend's coordinates; velocity means dx/dt, dy/dt, dz/dt, not a local physical speed. The single-frame event defaults to t=0; use --observer-time T to select another coordinate time (any finite value, including negative times). Metric evaluation and past-directed ray tracing start at that event, and saved lens maps retain its coordinate time.

Option Meaning / default
--observer-time T Single-frame camera coordinate time, default 0; cannot be combined with movie/track or lens-map input
--observer-position X Y Z Coordinate position; if look is omitted, point toward the origin
--look-ra-deg RA, --look-dec-deg DEC Coordinate look direction; missing angle defaults to RA=90°, Dec=-90°
--observer-radius R Positive radius used only to infer position, default 30; conflicts with explicit position
--observer-velocity VX VY VZ Coordinate velocity, default 0 0 0; must be future-timelike in the local metric
--camera-roll-deg ANGLE Rotate the camera up axis toward its right axis, default 0°

When position and look are both supplied, they are independent. Look alone implies position -R * look_direction, in either backend. Radius alone uses the default look. With no position/look/radius options, Minkowski keeps the origin camera and Schwarzschild keeps its radius-30 camera on +Z; both look along -Z. Velocity and roll alone do not alter these defaults. An explicit origin position requires a look angle. Position-derived directions on the Z axis use RA=0° to fix the pole orientation deterministically.

The axes are right-handed ICRS: X=RA 0°/Dec 0°, Y=RA 90°/Dec 0°, Z=Dec +90°. The coordinate look vector (cos(dec) cos(ra), cos(dec) sin(ra), sin(dec)) is projected into the moving camera's rest space. The coordinate north and west seeds are then orthogonalized there to form up and right. Consequently the look angles specify a projected coordinate orientation; they do not lock a lensed sky object to the image center. Positive roll uses up' = cos(roll) up + sin(roll) right, right' = -sin(roll) up + cos(roll) right.

No static observer is needed. The metric determines whether (1,VX,VY,VZ) is timelike; there is no Euclidean |V| < 1 check in curved coordinates. Invalid velocities are rejected before loading catalogs or building the PSF cache, with the velocity and its metric norm in the diagnostic. --observer-inward-speed has been removed. Single-frame camera options cannot be combined with --observer-track, --frames-dir, or --lens-map-input.

Schwarzschild uses Cartesian ingoing Kerr–Schild coordinates with M=1. Cameras at and inside the horizon r=2 (and inside the old r=1.5 guard) are allowed with a valid timelike coordinate velocity; position never decides a ray endpoint. Its finite escape radius is 256. These remain analytic demonstration settings, not criteria for future NR data. Zero coordinate velocity at or inside the horizon is not timelike and is rejected.

The normal dark terminal, for every backend, is the camera-relative local energy growth L - L0 >= T (default T = 8, overridable with --dark-threshold), where L = ln(alpha p^0) and L0 is the photon's L at the camera event (kept distinct from the escape-worldtube entry energy for an external camera). A constant camera boost cancels, so a large initial L alone does not produce a dark ray. Neither the photon energy nor the frequency ratio is reset; budget-exhausted and data/integration failures are separate unresolved/incomplete outcomes. Failed and unresolved midpoint probes are retained as diagnostic samples, not discarded after refinement. Lens-map replay uses its saved geometric policy and the same incomplete-output check as live tracing.

The following complete examples use the bundled synthetic catalog:

make -j PSF_BACKEND=cpu SPACETIME=minkowski backend
mkdir -p output/imgs
./build/Release/minkowski_sky --catalog assets/sky_grid_5deg.csv \
  --observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
  --look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
  --width 64 --height 48 --fov-deg 80 --exposure 0.1 \
  --coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
  --output output/imgs/minkowski_moving.png

make -j PSF_BACKEND=cpu SPACETIME=schwarzschild backend
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
  --observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
  --look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
  --width 64 --height 48 --fov-deg 80 --exposure 0.1 \
  --coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
  --output output/imgs/schwarzschild_off_axis.png

./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
  --observer-position 1.75 0 0 --observer-velocity -0.5 0 0 \
  --look-ra-deg 0 --look-dec-deg 0 \
  --width 64 --height 48 --fov-deg 80 --exposure 0.1 \
  --coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
  --output output/imgs/schwarzschild_inside_horizon.png

The last example points outward from a camera moving inward inside the horizon. It needs no CSV trajectory or movie wrapper. These small images are camera checks; increase resolution and refinement for production renders.

Alcubierre warp-bubble spacetime

alcubierre_sky renders the moving Alcubierre line element

ds^2 = -dt^2 + \left[dx - v_s f(r_s)\,dt\right]^2 + dy^2 + dz^2,

with the bubble center following the constant-velocity worldline x_s(t) = v_s t (a lab/non-comoving slicing fixed by x_s(0) = 0) and

$$f(r) = \frac{\tanh(\sigma(r+R)) - \tanh(\sigma(r-R))}{2\tanh(\sigma R)},\qquad r_s = \sqrt{(x-x_s)^2 + y^2 + z^2}.$$

The renderer evaluates the moving bubble's time-dependent metric along each ray. The exotic matter sourcing the bubble is treated as optically transparent. The shared dark policy terminates rays when L - L0 >= T (T = 8 by default, set with --dark-threshold); this finite threshold can also be reached at sub-luminal bubble velocities.

Option Meaning / default
--alcubierre-vs V Constant bubble velocity v_s (any finite value, default 0.5)
--alcubierre-radius R Bubble radius R > 0 (default 5)
--alcubierre-sigma S Wall sharpness S > 0 (default 1)

A camera must be timelike with an orthonormal tetrad. Its coordinate velocity V = dx/dt must satisfy |V - v_s f e_x| < 1. A static camera requires |v_s f| < 1; at the bubble center, the comoving velocity (v_s, 0, 0) is timelike even for super-luminal bubbles. Cameras may lie inside or outside the escape sphere; exterior rays are routed to their first entry or to infinity.

The escape sphere follows the bubble with radius R + 20/sigma; escaping rays continue through a Minkowski exterior. The single-frame camera defaults to (0,0,15) at t = 0. The finite-radius truncation leaves a residual shift of order |v_s| e^{-40}; accuracy at extremely large velocities is not guaranteed.

Per-ray coordinate-time coverage is a resource allowance

$$B = \frac{5,R_\text{escape}}{\max!\big(|1-|v_s||,; e^{-T}\big)},\qquad R_\text{escape} = R + \frac{20}{\sigma},$$

with T = --dark-threshold, clamped to [DBL_MIN, DBL_MAX/4]. This is not a completion guarantee: quota exhaustion returns UNRESOLVED/BUDGET_EXHAUSTED. --trace-lookback-time overrides B independently of --trace-max-steps. RK4 rejects a derived default step estimate above its internal cap; an explicit --trace-max-steps bypasses that check.

Lensing and frequency shifts come from the bubble wall. The configuration is invariant under the isometry (t, x) -> (t + T, x + v_s T), so observers related by it see identical escaping directions and frequency ratios. For example:

make -j PSF_BACKEND=cpu SPACETIME=alcubierre backend
./build/Release/alcubierre_sky --catalog assets/sky_grid_5deg.csv \
  --observer-radius 15 --look-ra-deg 90 --look-dec-deg -90 \
  --alcubierre-vs 1.5 --alcubierre-radius 1 --alcubierre-sigma 1 \
  --width 640 --height 360 --fov-deg 60 --exposure 1 \
  --coarse-cell-pixels 16 --refine-max-level 2 --psf-direct \
  --output output/imgs/alcubierre_wall.png

Movie image sequences

Movie mode reads an observer-track CSV containing coordinate time, proper time, Cartesian position, and a full four-by-four tetrad (21 columns). The track must cover the requested frame times. The following generates a sample accelerating Minkowski trajectory and renders it against the 2MASS sky:

mkdir -p output/imgs
./build/Release/minkowski_sky --write-minkowski-accel-track output/minkowski_accel_2s.csv \
  --duration 2 --fps 30 --proper-acceleration 1.52
./build/Release/minkowski_sky --observer-track output/minkowski_accel_2s.csv \
  --all-sky-catalog assets/2mass/processed/all_sky \
  --frames-dir output/imgs --frames-prefix minkowski_accel \
  --start-time 0 --duration 2 --fps 30 --exposure 1e15

The inclusive time interval produces frames minkowski_accel_000000.png through minkowski_accel_000060.png. Exposure may need adjustment as Doppler shifts change the brightness. The trajectory generator is a flat-spacetime example; other camera motions are supplied through observer-track CSVs. Output is an image sequence; video encoding is a separate step.

An all-sky movie prefetches the union of every finalized frame's requested one-degree tiles exactly once (one Movie catalog prefetch: ... line) and then renders every frame from the immutable cache, instead of re-running the per-frame prefetch scan and log. A single-frame render keeps its per-frame prefetch. The prefetch reads at most 512 unseen tiles per parallel batch, so an unbounded all-sky union never stages an unbounded temporary star copy.

Movie PNG encoding and writing run on one bounded writer thread and overlap the next frame's ray/catalog/FFTW work. The default bounds two queued frames plus the one the writer is currently processing (at most three jobs in memory). The same queue also serves multi-frame imported .grlens maps. --png-compression-level N sets the zlib/libpng level in 0..9; the default is libpng's own default. --fast-mode never changes it silently, so a preview movie can opt into --png-compression-level 1 explicitly. --movie-output-workers shared (default) lets the writer share the machine with the render pool. --movie-output-workers reserve-one shrinks the render OpenMP pool by one core for the whole movie render (ray tracing, catalog work, and the fast-mode FFTW plans all see the reduced count); it is applied before the fast accumulator is built so the FFTW plans actually use the reduced worker count, not only while the writer is active. A writer error is recorded and fails the render rather than silently dropping frames.

Movie rays from all frames share a newest-to-oldest coordinate-time sweep. --slab-duration sets its time-slab width (default 64). The current analytic backends use logical slabs without metric I/O; Nmesh metric loading remains future work.

Geodesic integration and tracing budgets

The default --integrator dp54 selects adaptive Dormand–Prince 5(4). Position, photon direction/momentum, and log-energy errors have separate absolute tolerances:

--ode-rtol R
--ode-atol-x X
--ode-atol-pi P
--ode-atol-l L
--ode-initial-step H
--ode-min-step HMIN
--ode-max-step HMAX
--ode-max-rejections N

All tolerance and step values must be finite and positive; the initial step must lie between the step bounds. The position absolute tolerance has the backend's coordinate-length units. Smaller tolerances control local ODE error, not a guaranteed bound on final sky-direction or image error near critical rays.

Default minimum steps are 1e-12 in Minkowski and Schwarzschild (M=1), and 1e-12 * R in Alcubierre. This is a conservative numerical guard, not a measured physical minimum. Default maximum steps are 16, 8M, and eight times min(0.1, 0.05/sigma), respectively. Schwarzschild starts at 0.1M; increasing the upper bound does not force large steps through strong-field regions. Use tolerance convergence for near-critical rays rather than interpreting the upper bound as a global accuracy guarantee. See the bounded step-bound experiments.

Tracing has independent accepted-step and coordinate-time budgets:

--trace-max-steps N
--trace-lookback-time T
--retry-step-increment N
--max-total-steps N
--retry-lookback-increment T
--max-total-lookback-time T

Changing the accepted-step allowance does not change the time allowance. A trustworthy ray that exhausts either allowance is unresolved and can be resumed by refinement with additional resources; it is not a physical dark endpoint. Retries keep the last accepted state and camera energy reference, without relaxing numerical tolerances. A hard limit that prevents a required retry blocks normal output; --allow-incomplete is a diagnostic override.

--integrator rk4 retains the fixed-step comparison path. Adaptive-only tolerances and time-budget overrides are not applicable to that path. Use an explicit stepper when reproducing a fixed-step reference rather than relying on the executable's default.

Reuse a completed lens map

--lens-map-output FILE saves the finalized inverse-lens mesh after tracing and refinement, while also rendering the image normally. It stores the geometry, frequency shifts, terminal statuses, triangle topology, and frame metadata in a versioned .grlens file. This allows later catalog, PSF, and exposure changes without repeating ray tracing.

mkdir -p output/imgs output/maps
./build/Release/schwarzschild_sky \
  --all-sky-catalog assets/2mass/processed/all_sky \
  --width 640 --height 360 --coarse-cell-pixels 8 --fov-deg 60 \
  --refine-max-level 2 --exposure 1e15 \
  --lens-map-output output/maps/schwarzschild.grlens \
  --output output/imgs/schwarzschild_trace.png
./build/Release/schwarzschild_sky \
  --lens-map-input output/maps/schwarzschild.grlens \
  --all-sky-catalog assets/2mass/processed/all_sky \
  --psf-fwhm-pixels 6 --psf-moffat-beta 3 --exposure 1e15 \
  --output output/imgs/schwarzschild_restyled.png

Import and export are mutually exclusive. Imported maps retain their original pixel dimensions, horizontal FOV, and mesh resolution; they cannot be refined on import. Explicit conflicting dimensions or FOV are rejected. The reader checks the format version, finite values, unit directions, triangle indices, and per-frame CRCs.

The map also records the stepper, numerical tolerances and step bounds, tracing and retry allowances, and accepted/rejected/RHS costs. Replay uses this saved provenance; tracing-option overrides are rejected because replay does not integrate rays. Version 3 preserves adaptive provenance. Version 2 imports as fixed RK4 with unavailable adaptive fields and cost diagnostics; version 1 is rejected because its capture semantics cannot be reconstructed reliably.

A movie export stores all final frame meshes in one file. To render it again, pass --lens-map-input FILE, the catalog, --frames-dir DIR, and --frames-prefix NAME; no observer track is needed on import.

Adaptive image mesh refinement

Adaptive refinement is disabled by default (--refine-max-level 0), so the existing coarse-mesh renders remain unchanged. When enabled, its defaults are an absolute direction error of 1e-3 degrees, relative error 0.1, minimum long edge 0.5 pixels, minimum area 0.25 pixel-squared, and a provisional minimum discrete-Jacobian magnitude of 1e-3. Each value can be overridden independently:

--refine-max-level N
--refine-angle-abs-deg D
--refine-angle-rel R
--refine-jacobian-min J
--refine-min-edge-pixels P
--refine-min-area-pixels2 A

N caps the triangle refinement level. Let e be the angle between the traced longest-edge midpoint direction and the normalized endpoint interpolation, and let s be the angle between those two endpoint camera directions. Both are evaluated internally in radians; the absolute CLI threshold D is specified in degrees and converted before comparison. s is the angular geometric size of the image triangle's test edge, not a source-sky/lens-map length. A locally escaped triangle is split only when both e > D_rad (the converted --refine-angle-abs-deg D) and e / max(s, 1e-15) > --refine-angle-rel. P and A prevent selecting a leaf already at or below the requested image-plane long-edge and area scales. Triangles whose three vertices straddle a dark/escape or unresolved/dark boundary are split independently of the direction-error thresholds, allowing the mesh to follow a shadow boundary. Unresolved vertices with an escape vertex (or three unresolved vertices) are retried with more step budget before any split; at the configured total cap the render is reported incomplete unless --allow-incomplete is given. A UUD/UDD boundary triangle at the geometric stop scale is approximately blackened and recorded with its image-plane area.

Independently of the midpoint geometry test, an all-escaped triangle also computes the discrete lens Jacobian J = Omega_source / Omega_image. Both signed solid angles use 2 atan2(dot(a, cross(b,c)), 1 + dot(a,b) + dot(b,c) + dot(c,a)), with the ordered camera directions for Omega_image and their traced infinity directions for Omega_source. A J-driven split requires both a shared image edge whose incident triangles have opposite nonzero signs of J and min(abs(J_left), abs(J_right)) < --refine-jacobian-min. It then requests that shared edge on both leaves. Thus |J| bounds the fold selection instead of widening it as a standalone critical-curve band. A negative sign is physical parity and is retained. The 1e-3 default is deliberately provisional and should be tuned with the small Schwarzschild refinement diagnostic before being treated as a production threshold.

For a small mesh diagnostic using the synthetic catalog:

./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
  --width 48 --height 48 --coarse-cell-pixels 24 --fov-deg 40 \
  --refine-max-level 1 --refine-angle-abs-deg 0.001 \
  --refine-angle-rel 0.001 --refine-jacobian-min 0.001 \
  --refine-min-edge-pixels 1 \
  --refine-min-area-pixels2 1 --draw-mesh \
  --output output/imgs/schwarzschild_refinement.png

For movies, each refinement generation completes the full newest-to-oldest time-slab sweep before any probe becomes a mesh vertex. Newly added vertices are therefore traced only by the next generation; the renderer never returns to a slab that has already been released. At the start of every generation, newly inserted vertices and geometry-only longest-edge probes for its new leaves are collected together, so both ray sets use the same parallel RayPool pass.

Camera directions and optics

Catalog directions and --look-ra-deg/--look-dec-deg use standard right-handed ICRS Cartesian axes: +X is RA 0 degrees/Dec 0 degrees, +Y is RA 90 degrees/Dec 0 degrees, and +Z is the north celestial pole. The local camera axes are forward, celestial north, and celestial west, so an image with north up has decreasing RA to the right. The defaults preserve the original -Z view. --exposure converts a catalog's physical flux normalization to the prototype HDR scale. The current synthetic catalog is calibrated for default exposure 1e-3; a 2MASS blackbody normalization in steradians requires a much larger display exposure such as the example above. The optics path integrates each fitted Planck spectrum through CIE 1931 color-matching functions and converts the resulting radiance to linear sRGB; it does not use an empirical color-temperature RGB approximation.

Point sources use a flux-normalized circular Moffat PSF by default (--psf-fwhm-pixels 2.7 --psf-moffat-beta 4.5). The FWHM matches the former 1.15-pixel Gaussian core while the Moffat wings remain continuous; both values are display/optics calibration parameters.

The default renderer builds one immutable, process-wide 64-by-64 sub-pixel Moffat lookup kernel. Its weights are pixel-area integrals and are bilinearly interpolated between phase tables. The renderer reports its build time and cached/direct-fallback image counts. Pass --psf-direct to use the slower 8-point quadrature reference evaluator for regression comparisons; an image whose required HDR-tail support exceeds the cache radius selects that reference path automatically.

--psf-relative-tail R controls the maximum omitted PSF tail fraction per image; it defaults to 1e-8. Larger values intentionally shorten the Moffat support and rebuild the immutable cache at the corresponding radius, which is useful when a faster, lower-fidelity render is acceptable.

--psf-min-y Y defaults to 0 (disabled). A positive value is a linear-HDR luminance cutoff: a PSF stops where its continuous Moffat Y profile falls below Y, and an image whose central value is already below Y is omitted. The final PSF report counts such omitted events and emits a warning when any occur.

For bounded preview renders, --max-cache-psf-flux F (default 1) allows images with 1 < flux <= F to use the existing cache instead of the direct evaluator. This does not rebuild or enlarge the cache: the image retains its true core flux, color, and sub-pixel position, while its Moffat wing is clipped at the existing cache radius. Images above F retain the direct fallback. --max-magnification M (default unlimited) caps the per-triangle rendering magnification before flux is formed; it is an explicit preview approximation.

Tone mapping and display

--exposure is a linear multiplier applied while the HDR framebuffer and its PSF splats are accumulated. The Moffat profile is the empirical effective optical PSF, absorbing seeing, jitter, and residual aberrations; it is not changed by any display choice.

The tone map is a later camera/display response applied to the finished linear HDR image, so it does not alter the linear HDR PSF. The default operator is a parameterized, per-channel soft clip

[ T_p(x) = \left[\tanh(x^p)\right]^{1/p}, \qquad p \ge 1, ]

selected with --tone-map softclip --tone-map-p 2, i.e.

[ T_2(x) = \sqrt{\tanh(x^2)}. ]

It is close to the identity in the dark, then rolls smoothly into a saturated shoulder above (x \sim 1). Because it is applied independently to each linear RGB channel, a bright star transitions from its colored Moffat wings into a white saturated core. Larger p approaches (\min(x,1)). This per-channel response is what gives the current simplified preview pipeline the crisp white cores seen in real astrophotography, rather than the gray, soft cores a Reinhard curve leaves on bright stars.

--tone-map reinhard selects the historical (x/(1+x)) display transform and is retained only to reproduce output rendered before the soft-clip operator was introduced. FITS output always remains the pre-tone-map linear RGB framebuffer, independent of operator and p; use it to measure physical flux, peak values, and FWHM. Tone-mapped PNG/PPM data are display-referred and must not be used for those measurements.

Commands written before this change that omit --tone-map actually used Reinhard; add --tone-map reinhard explicitly to reproduce their PNG/PPM output. Historical FITS/HDR benchmarks are unaffected.

Sensor bloom (optional)

--sensor-bloom-limit E --sensor-bloom-transfer e (both required together, default disabled) enable an optional post-PSF sensor saturation model applied to the finished linear HDR framebuffer before the display operator. Each RGB channel is processed independently and isotropically. A channel value above the finite response limit (E) contributes overflow (D=\max(H-E,0)). A fraction (e\in[0,1)) of that overflow is spread to the eight neighbours with a fixed 9-point stencil (four axial weights (4/20), four diagonal weights (1/20)), while the remaining ((1-e)D) is absorbed. Overflow directed outside the image is lost at the boundary and the stencil is not renormalized there. e=0 clamps every over-limit channel to (E) without spreading to neighbours.

(E) is expressed in the post-exposure linear HDR renderer scale, so it scales with --exposure. The model is a phenomenological limited-response approximation, not a specific CCD/CMOS/Bayer structure. It is non-conservative: signal is lost both to the drain ((1-e)D) and to the image boundary, and (e) controls both the per-round retained fraction and the effective propagation distance. The mesh overlay is drawn after the model, so the diagnostic lines neither bloom nor feed back into overflow propagation. Visual multi-scale bloom is not implemented.

When enabled, each frame prints one Sensor bloom: line with the initial saturated and final clamped channel counts, the actual and conservative round counts, the peak and maximum initial overflow, absorbed and boundary loss, residual clamp loss, and elapsed time.

Fast preview mode

--fast-mode replaces the per-event PSF splat with a two-stage approximation: each point-source image is deposited as a delta into an N×N supersampled HDR buffer (--fast-supersample N, default 2), then the whole frame is convolved once with a single global Moffat kernel and averaged down by the N×N block. The kernel is the pixel-area integral of the requested Moffat at the supersampled scale, so the requested FWHM and beta are preserved by the downsample itself.

--fast-deposit nearest (default) deposits into the single nearest supersampled pixel. It keeps the PSF shape exactly on that grid. The grid spacing is 1/N output pixel per axis, while the instantaneous rounding error is bounded by 1/(2N) per axis. --fast-deposit bilinear deposits into 4 adjacent pixels; it preserves the continuous centroid but broadens the profile (measured +4.5%/+1.9%/+1.0% FWHM at N=2/3/4). The deposition experiment and its raw output are in benchmarks/fast_mode_deposit_2026-09-18.md.

Fast mode is available only in the CPU PSF backend. It ignores --max-cache-psf-flux and uses one global kernel radius for every event, so a bright event whose requested support exceeds that radius is wing-clipped and counted in the report. --psf-min-y is still applied per event. Fast mode is a preview approximation, not the physically exact per-event PSF path.

Its accuracy is governed by the deposit mode and N, not by the output resolution. A nearest deposit shifts each image by at most 1/(2N) output pixels per axis. The separate two-dimensional centroid-distance experiment measured worst cases of ~0.30 px at N=2 and ~0.12 px at N=4 for a FWHM 2.7, beta 4.5 PSF. These offsets remain fixed fractions of one output pixel: with --psf-fwhm-pixels held fixed, a larger image does not reduce them, and only a larger N (or the centroid-preserving bilinear deposit, at the cost of a broadened profile) does.

That per-star number is not how one dense still image behaves. In the measured production-density 4K Galactic-center frame, fast-versus-standard HDR relative L2 is 10.44% at N=2 and 4.37% at N=4, with a +0.018% total linear-RGB-sum offset. The difference is concentrated in a small set of scalar RGB channel samples: the top 0.01% ranked by |fast-base| account for 95.7% (N=2) or 94.5% (N=4) of squared-error energy, while containing about 9% of the summed baseline linear-RGB values. The disjoint base-value partition puts the flux-bearing faint ranges ([1e-3,0.06), [0.06,0.1), [0.1,0.23), together ~47% of the linear-RGB sum) at ~2-5% local relative RMS with signed bias at or below 0.19%. This pattern is consistent with partial cancellation among many faint PSFs, while resolved bright cores retain bounded sub-pixel displacement. It does not make fast mode pixelwise or visually lossless.

These measurements characterize independent still frames only. With nearest deposition, a smoothly moving image jumps by one complete grid spacing when it crosses a cell boundary: 1/N output pixel per axis, or 0.5 px at N=2 and 0.25 px at N=4. Such discontinuities can appear as visible flicker even at 4K. Neither nearest mode is therefore qualified as a temporally coherent movie output path. Bilinear deposition removes this particular centroid jump but has the profile-broadening tradeoff above and has not been qualified as a final movie path either. For single-frame dense-star previews, N=2 remains the throughput default and N=4 the higher-fidelity point; the exact per-event reference PSF path remains the quantitative and movie reference. The single-frame error decomposition is reproducible with scripts/fast_mode_error_decomposition.py.

In CPU builds the global convolution is a zero-padded FFTW linear convolution (fftw3/fftw3_omp, double precision). The immutable kernel spectrum, FFTW plans, and scratch buffers are built once when the accumulator is initialized and reused for every frame; the disposable spatial reference remains only for tests and benchmarks. Startup prints one Fast FFTW: line with the minimum linear-convolution extent, the selected smooth FFT dimensions, worker count, plan mode, one-time plan and kernel-transform time, and scratch bytes. --verbose additionally prints per-frame stage timings (zero/pack, forward FFT, frequency multiply, inverse FFT, crop/downsample, total). The default FFTW planning mode is FFTW_ESTIMATE; FFTW_MEASURE can cut steady-state execution substantially but costs a much longer one-time plan for large frames, as recorded in benchmarks/fast_mode_fftw_2026-09-25.md.

Planning can be selected explicitly. --fast-fftw-plan measure uses FFTW_MEASURE. --fast-fftw-plan wisdom-update --fast-fftw-wisdom FILE measures once and writes reusable wisdom plus a FILE.meta sidecar recording FFTW version, precision, supersample, FFT dimensions, kernel radius, and worker count. Each file is written to a unique temporary and renamed into place individually; the pair is not a transactional update, and the sidecar is committed last, so a wisdom without a matching sidecar is treated as a miss. --fast-fftw-plan wisdom --fast-fftw-wisdom FILE imports that wisdom and creates the plans with FFTW_WISDOM_ONLY; if the file, size, supersample, worker count, version, or precision does not match, it fails with a diagnostic instead of silently replanning. Wisdom writes go to a temporary file and are renamed into place, so a failed write never leaves a half file.

Progress and diagnostics

The PSF-cache completion line is printed before tracing and catalog splatting begin. For long renders, pass --verbose to print catalog-prefetch state, splat-worker local heartbeats (8, 16, 32, ... completed triangles per worker), and image-write boundaries. The worker heartbeats use neither global progress accounting nor cross-worker synchronization. Verbose ray-trace output reports the initial mesh trace and refinement stages for single frames; movie mode additionally reports each generation's sample count and each time slab's activation and terminal-ray summary. Movie renders always print one summary per time slab; --verbose also prints the ray counts before each slab is loaded. In movie mode, --verbose adds one Movie frame N timing: line per frame and a Movie timing total/avg/max summary. The stages are catalog_mark, catalog_load, fast_clear, splat, the FFTW breakdown (fftw_zero_pack, fftw_forward, fftw_multiply, fftw_inverse, fftw_crop, fftw_total), tone_map, output, and wait (writer-queue backpressure), plus the frame total. output covers only producer-side writes, so it is zero in the async movie path; the writer's own time is reported separately by Movie writer summary: jobs/total/avg/max/drain. Movie output pipeline wall: starts after track load, ray tracing, lens-map write, and the catalog union prefetch, and covers per-frame render plus async output. Movie end-to-end wall: starts at render_movie() entry, so it additionally includes track load, all ray tracing, the union prefetch, and the final writer drain; it does not include the catalog, blackbody, and fast accumulator initialization done earlier in main(). Movie timing total is the sum of producer frame times only (it excludes tracing, prefetch, and the final queue drain). The all-sky Movie catalog prefetch: line reports mark, load+commit, and total tile time. Timing uses one clock read per bulk phase, never inside the per-star or per-pixel hot loops.

Pass --draw-mesh to also write the final image-plane triangle mesh as a <output-stem>_mesh.png sibling (.ppm in non-PNG builds). The main tone-mapped image and any --hdr-output FITS file remain mesh-free. The overlay alpha-composites image-plane triangle edges as one-pixel-wide 0.5 linear-gray diagnostic lines at 0.5 opacity. The line rasterizer uses coverage-based antialiasing.

Normal progress and summaries go to stdout; warnings, errors, and Debug diagnostics go to stderr. Successful runs exit 0 even if warnings are emitted.

HDR output

To preserve a single-frame render for later exposure and tone-mapping work, build with ENABLE_HDR=1 as described in build.md. This produces build/Release/schwarzschild_sky, which accepts --hdr-output. This switch writes the HDR file next to the ordinary output, replacing its extension with _HDR.fits; for example, --output output/imgs/ring.png --hdr-output writes output/imgs/ring_HDR.fits. It writes the pre-tone-mapping RGB framebuffer as a three-plane, 32-bit float FITS image. Values remain linear HDR at the renderer's arbitrary scale; no tone mapping or per-frame normalization is applied, and the --tone-map operator and --tone-map-p value never affect this file. The --sensor-bloom-* model runs only after this file is written, so enabling bloom never changes the stored FITS. The ordinary binaries do not contain this option or writer.

The FITS header describes a synthetic 8640-by-5760, 36-by-24 mm full-frame sensor with 4.1667 um pixels. Each render records its active centered crop and derives FOCALLEN from that render's width and horizontal --fov-deg; these camera fields support plate-solving workflows but do not calibrate flux.