Replace the position capture cutoff with a camera-relative dark threshold shared by every backend, and carry explicit outcome/reason provenance through the ray, RayPool, adaptive mesh, lens-map and replay paths. - eval/eval_slab return SpacetimePointStatus; remove SPACETIME_RAY_CAPTURED and the Schwarzschild capture radius; decouple observer construction from ray position. - RayEndpoint stores RayOutcome/RayReason plus the last trusted state; budget exhaustion is retryable UNRESOLVED, data/integration failures are INCOMPLETE. - Normal dark terminal is L - L0 >= --dark-threshold (default 8), with L0 taken at the camera event and kept distinct from the worldtube entry energy; photon energy and frequency ratio are never reset. - Implement E/D/U triangle decisions with merged budget retries, persistent probe witnesses promoted in place by vertex identity, conformity settling, and approximate-black boundary provenance with achieved-scale statistics. - Add RayPool continuation state and per-ray step budgets. - Bump lens-map to v2 with explicit end/outcome/reason, approx_black, threshold/retry/geometry provenance and per-frame retry counts; reject v1. - Gate production output on incomplete/error results, overridable with --allow-incomplete. - Update AGENTS.md, the design document and usage docs; add the termination oracle and regression coverage. make -B -j4 BUILD_TYPE=Debug test passes with bit-identical reference HDRs.
31 KiB
Rendering guide
See README.md for the project overview and a first 2MASS render,
and build.md for compilation. Commands below run from the repository
root and use the default Release binaries. Run a binary with --help for its
complete option list and defaults.
Catalog input
Use --all-sky-catalog assets/2mass/processed/all_sky for the tiled 2MASS
catalog, or --catalog PATH for a single processed CSV. Acquisition and
processing are documented in assets/2mass/README.md.
2MASS blackbody amplitudes need a much larger display exposure than the
synthetic test catalog; --exposure 1e15 is a starting point, to be adjusted
for the field and camera.
Single-frame camera
Both backends accept the same instantaneous camera parameters. Position and
velocity use the backend's coordinates; velocity means dx/dt, dy/dt, dz/dt,
not a local physical speed. The analytic single-frame event is at t=0.
| Option | Meaning / default |
|---|---|
--observer-position X Y Z |
Coordinate position; if look is omitted, point toward the origin |
--look-ra-deg RA, --look-dec-deg DEC |
Coordinate look direction; missing angle defaults to RA=90°, Dec=-90° |
--observer-radius R |
Positive radius used only to infer position, default 30; conflicts with explicit position |
--observer-velocity VX VY VZ |
Coordinate velocity, default 0 0 0; must be future-timelike in the local metric |
--camera-roll-deg ANGLE |
Rotate the camera up axis toward its right axis, default 0° |
When position and look are both supplied, they are independent. Look alone
implies position -R * look_direction, in either backend. Radius alone uses
the default look. With no position/look/radius options, Minkowski keeps the
origin camera and Schwarzschild keeps its radius-30 camera on +Z; both look
along -Z. Velocity and roll alone do not alter these defaults.
An explicit origin position requires a look angle. Position-derived directions
on the Z axis use RA=0° to fix the pole orientation deterministically.
The axes are right-handed ICRS: X=RA 0°/Dec 0°, Y=RA 90°/Dec 0°,
Z=Dec +90°. The coordinate look vector (cos(dec) cos(ra), cos(dec) sin(ra), sin(dec)) is projected into the moving camera's rest space. The coordinate
north and west seeds are then orthogonalized there to form up and right.
Consequently the look angles specify a projected coordinate orientation;
they do not lock a lensed sky object to the image center. Positive roll uses
up' = cos(roll) up + sin(roll) right,
right' = -sin(roll) up + cos(roll) right.
No static observer is needed. The metric determines whether (1,VX,VY,VZ)
is timelike; there is no Euclidean |V| < 1 check in curved coordinates.
Invalid velocities are rejected before loading catalogs or building the PSF
cache, with the velocity and its metric norm in the diagnostic.
--observer-inward-speed has been removed.
Single-frame camera options cannot be combined with --observer-track,
--frames-dir, or --lens-map-input.
Schwarzschild uses Cartesian ingoing Kerr–Schild coordinates with M=1.
Cameras at and inside the horizon r=2 (and inside the old r=1.5 guard) are
allowed with a valid timelike coordinate velocity; position never decides a ray
endpoint. Its finite escape radius is 256. These remain analytic demonstration
settings, not criteria for future NR data.
Zero coordinate velocity at or inside the horizon is not timelike and is rejected.
The normal dark terminal, for every backend, is the camera-relative local energy
growth L - L0 >= T (default T = 8, overridable with --dark-threshold),
where L = ln(alpha p^0) and L0 is the photon's L at the camera event
(kept distinct from the escape-worldtube entry energy for an external camera).
A constant camera boost cancels, so a large initial L alone does not produce a
dark ray. Neither the photon energy nor the frequency ratio is reset;
budget-exhausted and data/integration failures are separate
unresolved/incomplete outcomes.
Failed and unresolved midpoint probes are retained as diagnostic samples, not
discarded after refinement. Lens-map replay uses its saved geometric policy and
the same incomplete-output check as live tracing.
The following complete examples use the bundled synthetic catalog:
make -j PSF_BACKEND=cpu SPACETIME=minkowski backend
mkdir -p output/imgs
./build/Release/minkowski_sky --catalog assets/sky_grid_5deg.csv \
--observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
--look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
--output output/imgs/minkowski_moving.png
make -j PSF_BACKEND=cpu SPACETIME=schwarzschild backend
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
--observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
--look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
--output output/imgs/schwarzschild_off_axis.png
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
--observer-position 1.75 0 0 --observer-velocity -0.5 0 0 \
--look-ra-deg 0 --look-dec-deg 0 \
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
--output output/imgs/schwarzschild_inside_horizon.png
The last example points outward from a camera moving inward inside the horizon. It needs no CSV trajectory or movie wrapper. These small images are camera checks; increase resolution and refinement for production renders.
Alcubierre warp-bubble spacetime
alcubierre_sky renders the moving Alcubierre line element
ds^2 = -dt^2 + \left[dx - v_s f(r_s)\,dt\right]^2 + dy^2 + dz^2,
with the bubble center following the constant-velocity worldline
x_s(t) = v_s t (a lab/non-comoving slicing fixed by x_s(0) = 0) and
$$f(r) = \frac{\tanh(\sigma(r+R)) - \tanh(\sigma(r-R))}{2\tanh(\sigma R)},\qquad r_s = \sqrt{(x-x_s)^2 + y^2 + z^2}.$$
The bubble therefore propagates through the coordinates, and the metric is
time-dependent: the renderer evaluates f(r_s) and its spatial derivatives at
each coordinate time, while the extrinsic curvature supplies the required
d_t gamma information to the 3+1 null-ray equations. The exotic matter that
would source the bubble is treated as optically transparent and there is no
horizon, so in practice rays are active or escaped: the bubble has no causal
boundary at which L - L0 can diverge, and the shared dark policy is not
expected to trigger. This is why the backend requires a sub-luminal
|v_s| < 1; at or above 1 the metric develops an ergoregion/event horizon and
static observers cease to exist, which is outside the current scope.
| Option | Meaning / default |
|---|---|
--alcubierre-vs V |
Constant shift parameter, ` |
--alcubierre-radius R |
Bubble radius R > 0 (default 5) |
--alcubierre-sigma S |
Wall sharpness S > 0 (default 1) |
f decays to zero past r_s = R over a transition width ~1/sigma, so the
finite escape sphere is bubble-centered with radius R + 20/sigma and needs no
CLI option; it follows the moving bubble, so rays terminate only once the local
metric is flat to below double precision. The single-frame camera default is
(0,0,15) at t = 0, when the bubble is still at the origin; it must lie
inside the escape sphere, or the observer build fails with an explicit error.
The per-ray step budget scales with the escape radius and 1/(1-|v_s|), so
near-luminal v_s still lets grazing rays escape; combinations whose
worst-case budget would exceed the internal cap are rejected at startup.
Lensing and frequency shifts come from the bubble wall. The configuration is
invariant under the isometry (t, x) -> (t + T, x + v_s T), so observers
related by it see identical escaping directions and frequency ratios. For
example:
make -j PSF_BACKEND=cpu SPACETIME=alcubierre backend
./build/Release/alcubierre_sky --catalog assets/sky_grid_5deg.csv \
--observer-radius 15 --look-ra-deg 90 --look-dec-deg -90 \
--alcubierre-vs 0.5 --alcubierre-radius 5 --alcubierre-sigma 1 \
--width 640 --height 360 --fov-deg 60 --exposure 1 \
--coarse-cell-pixels 16 --refine-max-level 2 --psf-direct \
--output output/imgs/alcubierre_wall.png
Movie image sequences
Movie mode reads an observer-track CSV containing coordinate time, proper time, Cartesian position, and a full four-by-four tetrad (21 columns). The track must cover the requested frame times. The following generates a sample accelerating Minkowski trajectory and renders it against the 2MASS sky:
mkdir -p output/imgs
./build/Release/minkowski_sky --write-minkowski-accel-track output/minkowski_accel_2s.csv \
--duration 2 --fps 30 --proper-acceleration 1.52
./build/Release/minkowski_sky --observer-track output/minkowski_accel_2s.csv \
--all-sky-catalog assets/2mass/processed/all_sky \
--frames-dir output/imgs --frames-prefix minkowski_accel \
--start-time 0 --duration 2 --fps 30 --exposure 1e15
The inclusive time interval produces frames minkowski_accel_000000.png
through minkowski_accel_000060.png. Exposure may need adjustment as Doppler
shifts change the brightness. The trajectory generator is a flat-spacetime
example; other camera motions are supplied through observer-track CSVs.
Output is an image sequence; video encoding is a separate step.
An all-sky movie prefetches the union of every finalized frame's requested
one-degree tiles exactly once (one Movie catalog prefetch: ... line) and then
renders every frame from the immutable cache, instead of re-running the
per-frame prefetch scan and log. A single-frame render keeps its per-frame
prefetch. The prefetch reads at most 512 unseen tiles per parallel batch, so an
unbounded all-sky union never stages an unbounded temporary star copy.
Movie PNG encoding and writing run on one bounded writer thread and overlap
the next frame's ray/catalog/FFTW work. The default bounds two queued frames
plus the one the writer is currently processing (at most three jobs in memory).
The same queue also serves multi-frame imported .grlens maps. --png-compression-level N
sets the zlib/libpng level in 0..9; the default is libpng's own default.
--fast-mode never changes it silently, so a preview movie can opt into
--png-compression-level 1 explicitly. --movie-output-workers shared
(default) lets the writer share the machine with the render pool.
--movie-output-workers reserve-one shrinks the render OpenMP pool by one core
for the whole movie render (ray tracing, catalog work, and the fast-mode FFTW
plans all see the reduced count); it is applied before the fast accumulator is
built so the FFTW plans actually use the reduced worker count, not only while
the writer is active. A writer error is recorded and fails the render rather
than silently dropping frames.
Movie rays from all frames share a newest-to-oldest coordinate-time sweep.
--slab-duration sets its time-slab width (default 64). The current analytic
backends use logical slabs without metric I/O; Nmesh metric loading remains
future work.
Reuse a completed lens map
--lens-map-output FILE saves the finalized inverse-lens mesh after tracing
and refinement, while also rendering the image normally. It stores the
geometry, frequency shifts, terminal statuses, triangle topology, and frame
metadata in a versioned .grlens file. This allows later catalog, PSF, and
exposure changes without repeating ray tracing.
mkdir -p output/imgs output/maps
./build/Release/schwarzschild_sky \
--all-sky-catalog assets/2mass/processed/all_sky \
--width 640 --height 360 --coarse-cell-pixels 8 --fov-deg 60 \
--refine-max-level 2 --exposure 1e15 \
--lens-map-output output/maps/schwarzschild.grlens \
--output output/imgs/schwarzschild_trace.png
./build/Release/schwarzschild_sky \
--lens-map-input output/maps/schwarzschild.grlens \
--all-sky-catalog assets/2mass/processed/all_sky \
--psf-fwhm-pixels 6 --psf-moffat-beta 3 --exposure 1e15 \
--output output/imgs/schwarzschild_restyled.png
Import and export are mutually exclusive. Imported maps retain their original pixel dimensions, horizontal FOV, and mesh resolution; they cannot be refined on import. Explicit conflicting dimensions or FOV are rejected. The reader checks the format version, finite values, unit directions, triangle indices, and per-frame CRCs.
A movie export stores all final frame meshes in one file. To render it again,
pass --lens-map-input FILE, the catalog, --frames-dir DIR, and
--frames-prefix NAME; no observer track is needed on import.
Adaptive image mesh refinement
Adaptive refinement is disabled by default (--refine-max-level 0), so the
existing coarse-mesh renders remain unchanged. When enabled, its defaults are
an absolute direction error of 1e-3 degrees, relative error 0.1, minimum
long edge 0.5 pixels, minimum area 0.25 pixel-squared, and a provisional
minimum discrete-Jacobian magnitude of 1e-3. Each value can be overridden
independently:
--refine-max-level N
--refine-angle-abs-deg D
--refine-angle-rel R
--refine-jacobian-min J
--refine-min-edge-pixels P
--refine-min-area-pixels2 A
N caps the triangle refinement level. Let e be the angle between the
traced longest-edge midpoint direction and the normalized endpoint
interpolation, and let s be the angle between those two endpoint camera
directions. Both are evaluated internally in radians; the absolute CLI
threshold D is specified in degrees and converted before comparison. s
is the angular geometric size of the image triangle's test edge, not a
source-sky/lens-map length. A locally escaped triangle is split only when
both e > D_rad (the converted --refine-angle-abs-deg D) and
e / max(s, 1e-15) > --refine-angle-rel. P and A
prevent selecting a leaf already at or below the requested image-plane
long-edge and area scales.
Triangles whose three vertices straddle a dark/escape or unresolved/dark
boundary are split independently of the direction-error thresholds, allowing the
mesh to follow a shadow boundary. Unresolved vertices with an escape vertex (or
three unresolved vertices) are retried with more step budget before any split;
at the configured total cap the render is reported incomplete unless
--allow-incomplete is given. A UUD/UDD boundary triangle at the geometric stop
scale is approximately blackened and recorded with its image-plane area.
Independently of the midpoint geometry test, an all-escaped triangle also
computes the discrete lens Jacobian
J = Omega_source / Omega_image. Both signed solid angles use
2 atan2(dot(a, cross(b,c)), 1 + dot(a,b) + dot(b,c) + dot(c,a)), with the
ordered camera directions for Omega_image and their traced infinity
directions for Omega_source. A J-driven split requires both a shared
image edge whose incident triangles have opposite nonzero signs of J and
min(abs(J_left), abs(J_right)) < --refine-jacobian-min. It then requests
that shared edge on both leaves. Thus |J| bounds the fold selection instead
of widening it as a standalone critical-curve band. A negative sign is
physical parity and is retained. The 1e-3 default is deliberately
provisional and should be tuned with the small Schwarzschild refinement
diagnostic before being treated as a production threshold.
For a small mesh diagnostic using the synthetic catalog:
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
--width 48 --height 48 --coarse-cell-pixels 24 --fov-deg 40 \
--refine-max-level 1 --refine-angle-abs-deg 0.001 \
--refine-angle-rel 0.001 --refine-jacobian-min 0.001 \
--refine-min-edge-pixels 1 \
--refine-min-area-pixels2 1 --draw-mesh \
--output output/imgs/schwarzschild_refinement.png
For movies, each refinement generation completes the full newest-to-oldest
time-slab sweep before any probe becomes a mesh vertex. Newly added vertices
are therefore traced only by the next generation; the renderer never returns
to a slab that has already been released. At the start of every generation,
newly inserted vertices and geometry-only longest-edge probes for its new
leaves are collected together, so both ray sets use the same parallel
RayPool pass.
Camera directions and optics
Catalog directions and --look-ra-deg/--look-dec-deg use standard
right-handed ICRS Cartesian axes: +X is RA 0 degrees/Dec 0 degrees, +Y is
RA 90 degrees/Dec 0 degrees, and +Z is the north celestial pole. The local
camera axes are forward, celestial north, and celestial west, so an image with
north up has decreasing RA to the right. The defaults preserve the original
-Z view. --exposure converts a catalog's physical flux
normalization to the prototype HDR scale. The current synthetic catalog is
calibrated for default exposure 1e-3; a 2MASS blackbody normalization in
steradians requires a much larger display exposure such as the example above.
The optics path
integrates each fitted Planck spectrum through CIE 1931 color-matching functions
and converts the resulting radiance to linear sRGB; it does not use an empirical
color-temperature RGB approximation.
Point sources use a flux-normalized circular Moffat PSF by default
(--psf-fwhm-pixels 2.7 --psf-moffat-beta 4.5). The FWHM matches the former
1.15-pixel Gaussian core while the Moffat wings remain continuous; both values
are display/optics calibration parameters.
The default renderer builds one immutable, process-wide 64-by-64 sub-pixel
Moffat lookup kernel. Its weights are pixel-area integrals and are bilinearly
interpolated between phase tables. The renderer reports its build time and
cached/direct-fallback image counts. Pass --psf-direct to use the slower
8-point quadrature reference evaluator for regression comparisons; an image
whose required HDR-tail support exceeds the cache radius selects that reference
path automatically.
--psf-relative-tail R controls the maximum omitted PSF tail fraction per
image; it defaults to 1e-8. Larger values intentionally shorten the Moffat
support and rebuild the immutable cache at the corresponding radius, which is
useful when a faster, lower-fidelity render is acceptable.
--psf-min-y Y defaults to 0 (disabled). A positive value is a linear-HDR
luminance cutoff: a PSF stops where its continuous Moffat Y profile falls below
Y, and an image whose central value is already below Y is omitted. The
final PSF report counts such omitted events and emits a warning when any occur.
For bounded preview renders, --max-cache-psf-flux F (default 1) allows
images with 1 < flux <= F to use the existing cache instead of the direct
evaluator. This does not rebuild or enlarge the cache: the image retains its
true core flux, color, and sub-pixel position, while its Moffat wing is clipped
at the existing cache radius. Images above F retain the direct fallback.
--max-magnification M (default unlimited) caps the per-triangle rendering
magnification before flux is formed; it is an explicit preview approximation.
Tone mapping and display
--exposure is a linear multiplier applied while the HDR framebuffer and its
PSF splats are accumulated. The Moffat profile is the empirical effective
optical PSF, absorbing seeing, jitter, and residual aberrations; it is not
changed by any display choice.
The tone map is a later camera/display response applied to the finished linear HDR image, so it does not alter the linear HDR PSF. The default operator is a parameterized, per-channel soft clip
[ T_p(x) = \left[\tanh(x^p)\right]^{1/p}, \qquad p \ge 1, ]
selected with --tone-map softclip --tone-map-p 2, i.e.
[ T_2(x) = \sqrt{\tanh(x^2)}. ]
It is close to the identity in the dark, then rolls smoothly into a saturated
shoulder above (x \sim 1). Because it is applied independently to each linear
RGB channel, a bright star transitions from its colored Moffat wings into a
white saturated core. Larger p approaches (\min(x,1)). This per-channel
response is what gives the current simplified preview pipeline the crisp white
cores seen in real astrophotography, rather than the gray, soft cores a
Reinhard curve leaves on bright stars.
--tone-map reinhard selects the historical (x/(1+x)) display transform and
is retained only to reproduce output rendered before the soft-clip operator was
introduced. FITS output always remains the pre-tone-map linear RGB framebuffer,
independent of operator and p; use it to measure physical flux, peak values,
and FWHM. Tone-mapped PNG/PPM data are display-referred and must not be used for
those measurements.
Commands written before this change that omit --tone-map actually used
Reinhard; add --tone-map reinhard explicitly to reproduce their PNG/PPM
output. Historical FITS/HDR benchmarks are unaffected.
Sensor bloom (optional)
--sensor-bloom-limit E --sensor-bloom-transfer e (both required together,
default disabled) enable an optional post-PSF sensor saturation model applied to
the finished linear HDR framebuffer before the display operator. Each RGB
channel is processed independently and isotropically. A channel value above the
finite response limit (E) contributes overflow (D=\max(H-E,0)). A fraction
(e\in[0,1)) of that overflow is spread to the eight neighbours with a fixed
9-point stencil (four axial weights (4/20), four diagonal weights (1/20)),
while the remaining ((1-e)D) is absorbed. Overflow directed outside the image
is lost at the boundary and the stencil is not renormalized there. e=0 clamps
every over-limit channel to (E) without spreading to neighbours.
(E) is expressed in the post-exposure linear HDR renderer scale, so it scales
with --exposure. The model is a phenomenological limited-response
approximation, not a specific CCD/CMOS/Bayer structure. It is non-conservative:
signal is lost both to the drain ((1-e)D) and to the image boundary, and (e)
controls both the per-round retained fraction and the effective propagation
distance. The mesh overlay is drawn after the model, so the diagnostic lines
neither bloom nor feed back into overflow propagation. Visual multi-scale bloom
is not implemented.
When enabled, each frame prints one Sensor bloom: line with the initial
saturated and final clamped channel counts, the actual and conservative round
counts, the peak and maximum initial overflow, absorbed and boundary loss,
residual clamp loss, and elapsed time.
Fast preview mode
--fast-mode replaces the per-event PSF splat with a two-stage approximation:
each point-source image is deposited as a delta into an N×N supersampled HDR
buffer (--fast-supersample N, default 2), then the whole frame is convolved
once with a single global Moffat kernel and averaged down by the N×N block.
The kernel is the pixel-area integral of the requested Moffat at the
supersampled scale, so the requested FWHM and beta are preserved by the
downsample itself.
--fast-deposit nearest (default) deposits into the single nearest
supersampled pixel. It keeps the PSF shape exactly on that grid. The grid
spacing is 1/N output pixel per axis, while the instantaneous rounding error
is bounded by 1/(2N) per axis. --fast-deposit bilinear deposits into 4
adjacent pixels; it preserves the continuous centroid but broadens the profile
(measured +4.5%/+1.9%/+1.0% FWHM at N=2/3/4). The deposition experiment and
its raw output are in
benchmarks/fast_mode_deposit_2026-09-18.md.
Fast mode is available only in the CPU PSF backend. It ignores
--max-cache-psf-flux and uses one global kernel radius for every event, so a
bright event whose requested support exceeds that radius is wing-clipped and
counted in the report. --psf-min-y is still applied per event. Fast mode is a
preview approximation, not the physically exact per-event PSF path.
Its accuracy is governed by the deposit mode and N, not by the output
resolution. A nearest deposit shifts each image by at most 1/(2N) output
pixels per axis. The separate two-dimensional centroid-distance experiment
measured worst cases of ~0.30 px at N=2 and ~0.12 px at N=4 for a FWHM
2.7, beta 4.5 PSF. These offsets remain fixed fractions of one output
pixel: with --psf-fwhm-pixels held fixed, a larger image does not reduce
them, and only a larger N (or the centroid-preserving bilinear deposit, at
the cost of a broadened profile) does.
That per-star number is not how one dense still image behaves. In the measured
production-density 4K Galactic-center frame, fast-versus-standard HDR relative
L2 is 10.44% at N=2 and 4.37% at N=4, with a +0.018% total
linear-RGB-sum offset. The difference is concentrated in a small set of scalar
RGB channel samples: the top 0.01% ranked by |fast-base| account for
95.7% (N=2) or 94.5% (N=4) of squared-error energy, while containing
about 9% of the summed baseline linear-RGB values. The disjoint base-value
partition puts the flux-bearing faint ranges ([1e-3,0.06), [0.06,0.1),
[0.1,0.23), together ~47% of the linear-RGB sum) at ~2-5% local relative
RMS with signed bias at or below 0.19%. This pattern is consistent with
partial cancellation among many faint PSFs, while resolved bright cores retain
bounded sub-pixel displacement. It does not make fast mode pixelwise or
visually lossless.
These measurements characterize independent still frames only. With nearest
deposition, a smoothly moving image jumps by one complete grid spacing when it
crosses a cell boundary: 1/N output pixel per axis, or 0.5 px at N=2 and
0.25 px at N=4. Such discontinuities can appear as visible flicker even at
4K. Neither nearest mode is therefore qualified as a temporally coherent movie
output path. Bilinear deposition removes this particular centroid jump but has
the profile-broadening tradeoff above and has not been qualified as a final
movie path either. For single-frame dense-star previews, N=2 remains the
throughput default and N=4 the higher-fidelity point; the exact per-event
reference PSF path remains the quantitative and movie reference. The
single-frame error decomposition is reproducible with
scripts/fast_mode_error_decomposition.py.
In CPU builds the global convolution is a zero-padded FFTW linear convolution
(fftw3/fftw3_omp, double precision). The immutable kernel spectrum, FFTW
plans, and scratch buffers are built once when the accumulator is initialized
and reused for every frame; the disposable spatial reference remains only for
tests and benchmarks. Startup prints one Fast FFTW: line with the minimum
linear-convolution extent, the selected smooth FFT dimensions, worker count,
plan mode, one-time plan and kernel-transform time, and scratch bytes.
--verbose additionally prints per-frame stage timings (zero/pack, forward FFT,
frequency multiply, inverse FFT, crop/downsample, total). The default FFTW
planning mode is FFTW_ESTIMATE; FFTW_MEASURE can cut steady-state execution
substantially but costs a much longer one-time plan for large frames, as
recorded in
benchmarks/fast_mode_fftw_2026-09-25.md.
Planning can be selected explicitly. --fast-fftw-plan measure uses
FFTW_MEASURE. --fast-fftw-plan wisdom-update --fast-fftw-wisdom FILE
measures once and writes reusable wisdom plus a FILE.meta sidecar recording
FFTW version, precision, supersample, FFT dimensions, kernel radius, and worker
count. Each file is written to a unique temporary and renamed into place
individually; the pair is not a transactional update, and the sidecar is
committed last, so a wisdom without a matching sidecar is treated as a miss.
--fast-fftw-plan wisdom --fast-fftw-wisdom FILE imports
that wisdom and creates the plans with FFTW_WISDOM_ONLY; if the file, size,
supersample, worker count, version, or precision does not match, it fails with a
diagnostic instead of silently replanning. Wisdom writes go to a temporary file
and are renamed into place, so a failed write never leaves a half file.
Progress and diagnostics
The PSF-cache completion line is printed before tracing and catalog splatting
begin. For long renders, pass --verbose to print catalog-prefetch state,
splat-worker local heartbeats (8, 16, 32, ... completed triangles per worker),
and image-write boundaries. The worker heartbeats use neither global progress
accounting nor cross-worker synchronization.
Verbose ray-trace output reports the initial mesh trace and refinement stages
for single frames; movie mode additionally reports each generation's sample
count and each time slab's activation and terminal-ray summary.
Movie renders always print one summary per time slab; --verbose also prints
the ray counts before each slab is loaded. In movie mode, --verbose adds one
Movie frame N timing: line per frame and a Movie timing total/avg/max
summary. The stages are catalog_mark, catalog_load, fast_clear, splat,
the FFTW breakdown (fftw_zero_pack, fftw_forward, fftw_multiply,
fftw_inverse, fftw_crop, fftw_total), tone_map, output, and
wait (writer-queue backpressure), plus the frame total. output covers only
producer-side writes, so it is zero in the async movie path; the writer's own
time is reported separately by Movie writer summary: jobs/total/avg/max/drain.
Movie output pipeline wall: starts after track load, ray tracing, lens-map
write, and the catalog union prefetch, and covers per-frame render plus async
output. Movie end-to-end wall: starts at render_movie() entry, so it
additionally includes track load, all ray tracing, the union prefetch, and the
final writer drain; it does not include the catalog, blackbody, and fast
accumulator initialization done earlier in main(). Movie timing total is
the sum of producer frame times only (it excludes tracing, prefetch, and the
final queue drain). The all-sky Movie catalog prefetch: line reports mark,
load+commit, and total tile time. Timing uses one clock read per bulk phase,
never inside the per-star or per-pixel hot loops.
Pass --draw-mesh to also write the final image-plane triangle mesh as a
<output-stem>_mesh.png sibling (.ppm in non-PNG builds). The main
tone-mapped image and any --hdr-output FITS file remain mesh-free. The
overlay alpha-composites image-plane triangle edges as one-pixel-wide 0.5
linear-gray diagnostic lines at 0.5 opacity. The line rasterizer uses
coverage-based antialiasing.
HDR output
To preserve a single-frame render for later exposure and tone-mapping work,
build with ENABLE_HDR=1 as described in build.md. This produces build/Release/schwarzschild_sky,
which accepts --hdr-output. This switch writes the HDR file next to the
ordinary output, replacing its extension with _HDR.fits; for example,
--output output/imgs/ring.png --hdr-output writes
output/imgs/ring_HDR.fits. It writes the pre-tone-mapping RGB
framebuffer as a three-plane, 32-bit float FITS image. Values remain linear HDR
at the renderer's arbitrary scale; no tone mapping or per-frame normalization
is applied, and the --tone-map operator and --tone-map-p value never affect
this file. The --sensor-bloom-* model runs only after this file is written, so
enabling bloom never changes the stored FITS. The ordinary
binaries do not contain this option or writer.
The FITS header describes a synthetic 8640-by-5760, 36-by-24 mm full-frame
sensor with 4.1667 um pixels. Each render records its active centered crop and
derives FOCALLEN from that render's width and horizontal --fov-deg; these
camera fields support plate-solving workflows but do not calibrate flux.