Color antialiased half-edges by ray outcome with configurable Catppuccin colors and default opacity 0.5. Rasterize premultiplied RGBA8 overlays on the producer and composite in place after writing the clean image. Keep single-frame, movie, and replay output consistent. Add overlay, CLI, and queue ownership regressions and document the final output architecture.
683 lines
37 KiB
Markdown
683 lines
37 KiB
Markdown
# Rendering guide
|
||
|
||
See [README.md](README.md) for the project overview and a first 2MASS render,
|
||
and [build.md](build.md) for compilation. Commands below run from the repository
|
||
root and use the default Release binaries. Run a binary with `--help` for its
|
||
complete option list and defaults.
|
||
|
||
## Catalog input
|
||
|
||
Use `--all-sky-catalog assets/2mass/processed/all_sky` for the tiled 2MASS
|
||
catalog, or `--catalog PATH` for a single processed CSV. Acquisition and
|
||
processing are documented in [assets/2mass/README.md](assets/2mass/README.md).
|
||
2MASS blackbody amplitudes need a much larger display exposure than the
|
||
synthetic test catalog; `--exposure 1e15` is a starting point, to be adjusted
|
||
for the field and camera.
|
||
|
||
## Single-frame camera
|
||
|
||
Both backends accept the same instantaneous camera parameters. Position and
|
||
velocity use the backend's coordinates; velocity means `dx/dt, dy/dt, dz/dt`,
|
||
not a local physical speed. The single-frame event defaults to `t=0`; use
|
||
`--observer-time T` to select another coordinate time (any finite value,
|
||
including negative times). Metric evaluation and past-directed ray tracing
|
||
start at that event, and saved lens maps retain its coordinate time.
|
||
|
||
| Option | Meaning / default |
|
||
| --- | --- |
|
||
| `--observer-time T` | Single-frame camera coordinate time, default `0`; cannot be combined with movie/track or lens-map input |
|
||
| `--observer-position X Y Z` | Coordinate position; if look is omitted, point toward the origin |
|
||
| `--look-ra-deg RA`, `--look-dec-deg DEC` | Coordinate look direction; missing angle defaults to RA=90°, Dec=-90° |
|
||
| `--observer-radius R` | Positive radius used only to infer position, default 30; conflicts with explicit position |
|
||
| `--observer-velocity VX VY VZ` | Coordinate velocity, default `0 0 0`; must be future-timelike in the local metric |
|
||
| `--camera-roll-deg ANGLE` | Rotate the camera up axis toward its right axis, default 0° |
|
||
|
||
When position and look are both supplied, they are independent. Look alone
|
||
implies position `-R * look_direction`, in either backend. Radius alone uses
|
||
the default look. With no position/look/radius options, Minkowski keeps the
|
||
origin camera and Schwarzschild keeps its radius-30 camera on +Z; both look
|
||
along -Z. Velocity and roll alone do not alter these defaults.
|
||
An explicit origin position requires a look angle. Position-derived directions
|
||
on the Z axis use RA=0° to fix the pole orientation deterministically.
|
||
|
||
The axes are right-handed ICRS: X=RA 0°/Dec 0°, Y=RA 90°/Dec 0°,
|
||
Z=Dec +90°. The coordinate look vector `(cos(dec) cos(ra), cos(dec) sin(ra),
|
||
sin(dec))` is projected into the moving camera's rest space. The coordinate
|
||
north and west seeds are then orthogonalized there to form up and right.
|
||
Consequently the look angles specify a projected coordinate orientation;
|
||
they do not lock a lensed sky object to the image center. Positive roll uses
|
||
`up' = cos(roll) up + sin(roll) right`,
|
||
`right' = -sin(roll) up + cos(roll) right`.
|
||
|
||
No static observer is needed. The metric determines whether `(1,VX,VY,VZ)`
|
||
is timelike; there is no Euclidean `|V| < 1` check in curved coordinates.
|
||
Invalid velocities are rejected before loading catalogs or building the PSF
|
||
cache, with the velocity and its metric norm in the diagnostic.
|
||
`--observer-inward-speed` has been removed.
|
||
Single-frame camera options cannot be combined with `--observer-track`,
|
||
`--frames-dir`, or `--lens-map-input`.
|
||
|
||
Schwarzschild uses Cartesian ingoing Kerr–Schild coordinates with `M=1`.
|
||
Cameras at and inside the horizon `r=2` (and inside the old `r=1.5` guard) are
|
||
allowed with a valid timelike coordinate velocity; position never decides a ray
|
||
endpoint. Its finite escape radius is `256`. These remain analytic demonstration
|
||
settings, not criteria for future NR data.
|
||
Zero coordinate velocity at or inside the horizon is not timelike and is rejected.
|
||
|
||
The normal dark terminal, for every backend, is the camera-relative local energy
|
||
growth `L - L0 >= T` (default `T = 8`, overridable with `--dark-threshold`),
|
||
where `L = ln(alpha p^0)` and `L0` is the photon's `L` at the **camera event**
|
||
(kept distinct from the escape-worldtube entry energy for an external camera).
|
||
A constant camera boost cancels, so a large initial `L` alone does not produce a
|
||
dark ray. Neither the photon energy nor the frequency ratio is reset;
|
||
budget-exhausted and data/integration failures are separate
|
||
unresolved/incomplete outcomes.
|
||
Failed and unresolved midpoint probes are retained as diagnostic samples, not
|
||
discarded after refinement. Lens-map replay uses its saved geometric policy and
|
||
the same incomplete-output check as live tracing.
|
||
|
||
The following complete examples use the bundled synthetic catalog:
|
||
|
||
```sh
|
||
make -j PSF_BACKEND=cpu SPACETIME=minkowski backend
|
||
mkdir -p output/imgs
|
||
./build/Release/minkowski_sky --catalog assets/sky_grid_5deg.csv \
|
||
--observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
|
||
--look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
|
||
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
|
||
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
|
||
--output output/imgs/minkowski_moving.png
|
||
|
||
make -j PSF_BACKEND=cpu SPACETIME=schwarzschild backend
|
||
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
|
||
--observer-position 3 -4 5 --observer-velocity 0.2 -0.1 0.3 \
|
||
--look-ra-deg 37 --look-dec-deg -23 --camera-roll-deg 19 \
|
||
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
|
||
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
|
||
--output output/imgs/schwarzschild_off_axis.png
|
||
|
||
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
|
||
--observer-position 1.75 0 0 --observer-velocity -0.5 0 0 \
|
||
--look-ra-deg 0 --look-dec-deg 0 \
|
||
--width 64 --height 48 --fov-deg 80 --exposure 0.1 \
|
||
--coarse-cell-pixels 8 --refine-max-level 0 --psf-direct \
|
||
--output output/imgs/schwarzschild_inside_horizon.png
|
||
```
|
||
|
||
The last example points outward from a camera moving inward inside the horizon.
|
||
It needs no CSV trajectory or movie wrapper. These small images are camera
|
||
checks; increase resolution and refinement for production renders.
|
||
|
||
## Alcubierre warp-bubble spacetime
|
||
|
||
`alcubierre_sky` renders the moving Alcubierre line element
|
||
|
||
$$ds^2 = -dt^2 + \left[dx - v_s f(r_s)\,dt\right]^2 + dy^2 + dz^2,$$
|
||
|
||
with the bubble center following the constant-velocity worldline
|
||
`x_s(t) = v_s t` (a lab/non-comoving slicing fixed by `x_s(0) = 0`) and
|
||
|
||
$$f(r) = \frac{\tanh(\sigma(r+R)) - \tanh(\sigma(r-R))}{2\tanh(\sigma R)},\qquad
|
||
r_s = \sqrt{(x-x_s)^2 + y^2 + z^2}.$$
|
||
|
||
The renderer evaluates the moving bubble's time-dependent metric along each
|
||
ray. The exotic matter sourcing the bubble is treated as optically transparent.
|
||
The shared dark policy terminates rays when `L - L0 >= T` (`T = 8` by default,
|
||
set with `--dark-threshold`); this finite threshold can also be reached at
|
||
sub-luminal bubble velocities.
|
||
|
||
| Option | Meaning / default |
|
||
| --- | --- |
|
||
| `--alcubierre-vs V` | Constant bubble velocity `v_s` (any finite value, default 0.5) |
|
||
| `--alcubierre-radius R` | Bubble radius `R > 0` (default 5) |
|
||
| `--alcubierre-sigma S` | Wall sharpness `S > 0` (default 1) |
|
||
|
||
A camera must be timelike with an orthonormal tetrad. Its coordinate velocity
|
||
`V = dx/dt` must satisfy `|V - v_s f e_x| < 1`. A static camera requires
|
||
`|v_s f| < 1`; at the bubble center, the comoving velocity `(v_s, 0, 0)` is
|
||
timelike even for super-luminal bubbles. Cameras may lie inside or outside the
|
||
escape sphere; exterior rays are routed to their first entry or to infinity.
|
||
|
||
The escape sphere follows the bubble with radius `R + 20/sigma`; escaping rays
|
||
continue through a Minkowski exterior. The single-frame camera defaults to
|
||
`(0,0,15)` at `t = 0`. The finite-radius truncation leaves a residual shift of
|
||
order `|v_s| e^{-40}`; accuracy at extremely large velocities is not guaranteed.
|
||
|
||
Per-ray coordinate-time coverage is a **resource allowance**
|
||
|
||
$$B = \frac{5\,R_\text{escape}}{\max\!\big(|1-|v_s||,\; e^{-T}\big)},\qquad
|
||
R_\text{escape} = R + \frac{20}{\sigma},$$
|
||
|
||
with `T = --dark-threshold`, clamped to `[DBL_MIN, DBL_MAX/4]`. This is not a
|
||
completion guarantee: quota exhaustion returns `UNRESOLVED/BUDGET_EXHAUSTED`.
|
||
`--trace-lookback-time` overrides `B` independently of `--trace-max-steps`.
|
||
RK4 rejects a derived default step estimate above its internal cap; an explicit
|
||
`--trace-max-steps` bypasses that check.
|
||
|
||
Lensing and frequency shifts come from the bubble wall. The configuration is
|
||
invariant under the isometry `(t, x) -> (t + T, x + v_s T)`, so observers
|
||
related by it see identical escaping directions and frequency ratios. For
|
||
example:
|
||
|
||
```sh
|
||
make -j PSF_BACKEND=cpu SPACETIME=alcubierre backend
|
||
./build/Release/alcubierre_sky --catalog assets/sky_grid_5deg.csv \
|
||
--observer-radius 15 --look-ra-deg 90 --look-dec-deg -90 \
|
||
--alcubierre-vs 1.5 --alcubierre-radius 1 --alcubierre-sigma 1 \
|
||
--width 640 --height 360 --fov-deg 60 --exposure 1 \
|
||
--coarse-cell-pixels 16 --refine-max-level 2 --psf-direct \
|
||
--output output/imgs/alcubierre_wall.png
|
||
```
|
||
|
||
## Movie image sequences
|
||
|
||
Movie mode reads an observer-track CSV containing coordinate time, proper
|
||
time, Cartesian position, and a full four-by-four tetrad (21 columns). The
|
||
track must cover the requested frame times. The following generates a sample
|
||
accelerating Minkowski trajectory and renders it against the 2MASS sky:
|
||
|
||
```sh
|
||
mkdir -p output/imgs
|
||
./build/Release/minkowski_sky --write-minkowski-accel-track output/minkowski_accel_2s.csv \
|
||
--duration 2 --fps 30 --proper-acceleration 1.52
|
||
./build/Release/minkowski_sky --observer-track output/minkowski_accel_2s.csv \
|
||
--all-sky-catalog assets/2mass/processed/all_sky \
|
||
--frames-dir output/imgs --frames-prefix minkowski_accel \
|
||
--start-time 0 --duration 2 --fps 30 --exposure 1e15
|
||
```
|
||
|
||
The inclusive time interval produces frames `minkowski_accel_000000.png`
|
||
through `minkowski_accel_000060.png`. Exposure may need adjustment as Doppler
|
||
shifts change the brightness. The trajectory generator is a flat-spacetime
|
||
example; other camera motions are supplied through observer-track CSVs.
|
||
Output is an image sequence; video encoding is a separate step.
|
||
|
||
An all-sky movie prefetches the union of every finalized frame's requested
|
||
one-degree tiles exactly once (one `Movie catalog prefetch: ...` line) and then
|
||
renders every frame from the immutable cache, instead of re-running the
|
||
per-frame prefetch scan and log. A single-frame render keeps its per-frame
|
||
prefetch. The prefetch reads at most 512 unseen tiles per parallel batch, so an
|
||
unbounded all-sky union never stages an unbounded temporary star copy.
|
||
|
||
Movie PNG encoding and writing run on one bounded writer thread and overlap
|
||
the next frame's ray/catalog/FFTW work. The default bounds two queued frames
|
||
plus the one the writer is currently processing (at most three jobs in memory).
|
||
The same queue also serves multi-frame imported `.grlens` maps. `--png-compression-level N`
|
||
sets the zlib/libpng level in `0..9`; the default is libpng's own default.
|
||
`--fast-mode` never changes it silently, so a preview movie can opt into
|
||
`--png-compression-level 1` explicitly. `--movie-output-workers shared`
|
||
(default) lets the writer share the machine with the render pool.
|
||
`--movie-output-workers reserve-one` shrinks the render OpenMP pool by one core
|
||
for the whole movie render (ray tracing, catalog work, and the fast-mode FFTW
|
||
plans all see the reduced count); it is applied before the fast accumulator is
|
||
built so the FFTW plans actually use the reduced worker count, not only while
|
||
the writer is active. A writer error is recorded and fails the render rather
|
||
than silently dropping frames.
|
||
|
||
Movie rays from all frames share a newest-to-oldest coordinate-time sweep.
|
||
`--slab-duration` sets its time-slab width (default `64`). The current analytic
|
||
backends use logical slabs without metric I/O; [Nmesh](https://github.com/nmeshsource/nmesh) metric loading remains
|
||
future work.
|
||
|
||
## Geodesic integration and tracing budgets
|
||
|
||
The default `--integrator dp54` selects adaptive Dormand–Prince 5(4). Position, photon
|
||
direction/momentum, and log-energy errors have separate absolute tolerances:
|
||
|
||
```text
|
||
--ode-rtol R
|
||
--ode-atol-x X
|
||
--ode-atol-pi P
|
||
--ode-atol-l L
|
||
--ode-initial-step H
|
||
--ode-min-step HMIN
|
||
--ode-max-step HMAX
|
||
--ode-max-rejections N
|
||
```
|
||
|
||
All tolerance and step values must be finite and positive; the initial step
|
||
must lie between the step bounds. The position absolute tolerance has the
|
||
backend's coordinate-length units. Smaller tolerances control local ODE error,
|
||
not a guaranteed bound on final sky-direction or image error near critical rays.
|
||
|
||
Default minimum steps are `1e-12` in Minkowski and Schwarzschild (`M=1`), and
|
||
`1e-12 * R` in Alcubierre. This is a conservative numerical guard, not a measured
|
||
physical minimum. Default maximum steps are `16`, `8M`, and eight times
|
||
`min(0.1, 0.05/sigma)`, respectively. Schwarzschild starts at `0.1M`; increasing
|
||
the upper bound does not force large steps through strong-field regions.
|
||
Use tolerance convergence for near-critical rays rather than interpreting the
|
||
upper bound as a global accuracy guarantee. See the bounded
|
||
[step-bound experiments](benchmarks/adaptive_step_bounds_2026-10-05/README.md).
|
||
|
||
Tracing has independent accepted-step and coordinate-time budgets:
|
||
|
||
```text
|
||
--trace-max-steps N
|
||
--trace-lookback-time T
|
||
--retry-step-increment N
|
||
--max-total-steps N
|
||
--retry-lookback-increment T
|
||
--max-total-lookback-time T
|
||
```
|
||
|
||
Changing the accepted-step allowance does not change the time allowance. A
|
||
trustworthy ray that exhausts either allowance is unresolved and can be resumed
|
||
by refinement with additional resources; it is not a physical dark endpoint.
|
||
Retries keep the last accepted state and camera energy reference, without
|
||
relaxing numerical tolerances. A hard limit that prevents a required retry
|
||
blocks normal output; `--allow-incomplete` is a diagnostic override.
|
||
|
||
`--integrator rk4` retains the fixed-step comparison path. Adaptive-only
|
||
tolerances and time-budget overrides are not applicable to that path. Use an
|
||
explicit stepper when reproducing a fixed-step reference rather than relying
|
||
on the executable's default.
|
||
|
||
## Reuse a completed lens map
|
||
|
||
`--lens-map-output FILE` saves the finalized inverse-lens mesh after tracing
|
||
and refinement, while also rendering the image normally. It stores the
|
||
geometry, frequency shifts, terminal statuses, triangle topology, and frame
|
||
metadata in a versioned `.grlens` file. This allows later catalog, PSF, and
|
||
exposure changes without repeating ray tracing.
|
||
|
||
```sh
|
||
mkdir -p output/imgs output/maps
|
||
./build/Release/schwarzschild_sky \
|
||
--all-sky-catalog assets/2mass/processed/all_sky \
|
||
--width 640 --height 360 --coarse-cell-pixels 8 --fov-deg 60 \
|
||
--refine-max-level 2 --exposure 1e15 \
|
||
--lens-map-output output/maps/schwarzschild.grlens \
|
||
--output output/imgs/schwarzschild_trace.png
|
||
./build/Release/schwarzschild_sky \
|
||
--lens-map-input output/maps/schwarzschild.grlens \
|
||
--all-sky-catalog assets/2mass/processed/all_sky \
|
||
--psf-fwhm-pixels 6 --psf-moffat-beta 3 --exposure 1e15 \
|
||
--output output/imgs/schwarzschild_restyled.png
|
||
```
|
||
|
||
Import and export are mutually exclusive. Imported maps retain their original
|
||
pixel dimensions, horizontal FOV, and mesh resolution; they cannot be refined
|
||
on import. Explicit conflicting dimensions or FOV are rejected. The reader
|
||
checks the format version, finite values, unit directions, triangle indices,
|
||
and per-frame CRCs.
|
||
|
||
The map also records the stepper, numerical tolerances and step bounds, tracing
|
||
and retry allowances, and accepted/rejected/RHS costs. Replay uses this saved
|
||
provenance; tracing-option overrides are rejected because replay does not
|
||
integrate rays. Version 3 preserves adaptive provenance. Version 2 imports as
|
||
fixed RK4 with unavailable adaptive fields and cost diagnostics; version 1 is
|
||
rejected because its capture semantics cannot be reconstructed reliably.
|
||
|
||
A movie export stores all final frame meshes in one file. To render it again,
|
||
pass `--lens-map-input FILE`, the catalog, `--frames-dir DIR`, and
|
||
`--frames-prefix NAME`; no observer track is needed on import.
|
||
|
||
## Adaptive image mesh refinement
|
||
|
||
Adaptive refinement is disabled by default (`--refine-max-level 0`), so the
|
||
existing coarse-mesh renders remain unchanged. When enabled, its defaults are
|
||
an absolute direction error of `1e-3` degrees, relative error `0.1`, minimum
|
||
long edge `0.5` pixels, minimum area `0.25` pixel-squared, and a provisional
|
||
minimum discrete-Jacobian magnitude of `1e-3`. Each value can be overridden
|
||
independently:
|
||
|
||
```text
|
||
--refine-max-level N
|
||
--refine-angle-abs-deg D
|
||
--refine-angle-rel R
|
||
--refine-jacobian-min J
|
||
--refine-min-edge-pixels P
|
||
--refine-min-area-pixels2 A
|
||
```
|
||
|
||
`N` caps the triangle refinement level. Let `e` be the angle between the
|
||
traced longest-edge midpoint direction and the normalized endpoint
|
||
interpolation, and let `s` be the angle between those two endpoint **camera
|
||
directions**. Both are evaluated internally in radians; the absolute CLI
|
||
threshold `D` is specified in degrees and converted before comparison. `s`
|
||
is the angular geometric size of the image triangle's test edge, not a
|
||
source-sky/lens-map length. A locally escaped triangle is split only when
|
||
**both** `e > D_rad` (the converted `--refine-angle-abs-deg D`) and
|
||
`e / max(s, 1e-15) > --refine-angle-rel`. `P` and `A`
|
||
prevent selecting a leaf already at or below the requested image-plane
|
||
long-edge and area scales.
|
||
Triangles whose three vertices straddle a dark/escape or unresolved/dark
|
||
boundary are split independently of the direction-error thresholds, allowing the
|
||
mesh to follow a shadow boundary. Unresolved vertices with an escape vertex (or
|
||
three unresolved vertices) are retried with more step budget before any split;
|
||
at the configured total cap the render is reported incomplete unless
|
||
`--allow-incomplete` is given. A UUD/UDD boundary triangle at the geometric stop
|
||
scale is approximately blackened and recorded with its image-plane area.
|
||
|
||
Independently of the midpoint geometry test, an all-escaped triangle also
|
||
computes the discrete lens Jacobian
|
||
`J = Omega_source / Omega_image`. Both signed solid angles use
|
||
`2 atan2(dot(a, cross(b,c)), 1 + dot(a,b) + dot(b,c) + dot(c,a))`, with the
|
||
ordered camera directions for `Omega_image` and their traced infinity
|
||
directions for `Omega_source`. A J-driven split requires **both** a shared
|
||
image edge whose incident triangles have opposite nonzero signs of `J` and
|
||
`min(abs(J_left), abs(J_right)) < --refine-jacobian-min`. It then requests
|
||
that shared edge on both leaves. Thus `|J|` bounds the fold selection instead
|
||
of widening it as a standalone critical-curve band. A negative sign is
|
||
physical parity and is retained. The `1e-3` default is deliberately
|
||
provisional and should be tuned with the small Schwarzschild refinement
|
||
diagnostic before being treated as a production threshold.
|
||
|
||
For a small mesh diagnostic using the synthetic catalog:
|
||
|
||
```sh
|
||
./build/Release/schwarzschild_sky --catalog assets/sky_grid_5deg.csv \
|
||
--width 48 --height 48 --coarse-cell-pixels 24 --fov-deg 40 \
|
||
--refine-max-level 1 --refine-angle-abs-deg 0.001 \
|
||
--refine-angle-rel 0.001 --refine-jacobian-min 0.001 \
|
||
--refine-min-edge-pixels 1 \
|
||
--refine-min-area-pixels2 1 --draw-mesh \
|
||
--output output/imgs/schwarzschild_refinement.png
|
||
```
|
||
|
||
For movies, each refinement generation completes the full newest-to-oldest
|
||
time-slab sweep before any probe becomes a mesh vertex. Newly added vertices
|
||
are therefore traced only by the next generation; the renderer never returns
|
||
to a slab that has already been released. At the start of every generation,
|
||
newly inserted vertices and geometry-only longest-edge probes for its new
|
||
leaves are collected together, so both ray sets use the same parallel
|
||
`RayPool` pass.
|
||
|
||
## Camera directions and optics
|
||
|
||
Catalog directions and `--look-ra-deg`/`--look-dec-deg` use standard
|
||
right-handed ICRS Cartesian axes: `+X` is RA 0 degrees/Dec 0 degrees, `+Y` is
|
||
RA 90 degrees/Dec 0 degrees, and `+Z` is the north celestial pole. The local
|
||
camera axes are forward, celestial north, and celestial west, so an image with
|
||
north up has decreasing RA to the right. The defaults preserve the original
|
||
`-Z` view. `--exposure` converts a catalog's physical flux
|
||
normalization to the prototype HDR scale. The current synthetic catalog is
|
||
calibrated for default exposure `1e-3`; a 2MASS blackbody normalization in
|
||
steradians requires a much larger display exposure such as the example above.
|
||
The optics path
|
||
integrates each fitted Planck spectrum through CIE 1931 color-matching functions
|
||
and converts the resulting radiance to linear sRGB; it does not use an empirical
|
||
color-temperature RGB approximation.
|
||
|
||
Point sources use a flux-normalized circular Moffat PSF by default
|
||
(`--psf-fwhm-pixels 2.7 --psf-moffat-beta 4.5`). The FWHM matches the former
|
||
1.15-pixel Gaussian core while the Moffat wings remain continuous; both values
|
||
are display/optics calibration parameters.
|
||
|
||
The default renderer builds one immutable, process-wide 64-by-64 sub-pixel
|
||
Moffat lookup kernel. Its weights are pixel-area integrals and are bilinearly
|
||
interpolated between phase tables. The renderer reports its build time and
|
||
cached/direct-fallback image counts. Pass `--psf-direct` to use the slower
|
||
8-point quadrature reference evaluator for regression comparisons; an image
|
||
whose required HDR-tail support exceeds the cache radius selects that reference
|
||
path automatically.
|
||
|
||
`--psf-relative-tail R` controls the maximum omitted PSF tail fraction per
|
||
image; it defaults to `1e-8`. Larger values intentionally shorten the Moffat
|
||
support and rebuild the immutable cache at the corresponding radius, which is
|
||
useful when a faster, lower-fidelity render is acceptable.
|
||
|
||
`--psf-min-y Y` defaults to `0` (disabled). A positive value is a linear-HDR
|
||
luminance cutoff: a PSF stops where its continuous Moffat Y profile falls below
|
||
`Y`, and an image whose central value is already below `Y` is omitted. The
|
||
final PSF report counts such omitted events and emits a warning when any occur.
|
||
|
||
For bounded preview renders, `--max-cache-psf-flux F` (default `1`) allows
|
||
images with `1 < flux <= F` to use the existing cache instead of the direct
|
||
evaluator. This does not rebuild or enlarge the cache: the image retains its
|
||
true core flux, color, and sub-pixel position, while its Moffat wing is clipped
|
||
at the existing cache radius. Images above `F` retain the direct fallback.
|
||
`--max-magnification M` (default unlimited) caps the per-triangle rendering
|
||
magnification before flux is formed; it is an explicit preview approximation.
|
||
|
||
### Tone mapping and display
|
||
|
||
`--exposure` is a linear multiplier applied while the HDR framebuffer and its
|
||
PSF splats are accumulated. The Moffat profile is the empirical effective
|
||
optical PSF, absorbing seeing, jitter, and residual aberrations; it is not
|
||
changed by any display choice.
|
||
|
||
The tone map is a later camera/display response applied to the finished linear
|
||
HDR image, so it does not alter the linear HDR PSF. The default operator is a
|
||
parameterized, per-channel soft clip
|
||
|
||
\[
|
||
T_p(x) = \left[\tanh(x^p)\right]^{1/p}, \qquad p \ge 1,
|
||
\]
|
||
|
||
selected with `--tone-map softclip --tone-map-p 2`, i.e.
|
||
|
||
\[
|
||
T_2(x) = \sqrt{\tanh(x^2)}.
|
||
\]
|
||
|
||
It is close to the identity in the dark, then rolls smoothly into a saturated
|
||
shoulder above \(x \sim 1\). Because it is applied independently to each linear
|
||
RGB channel, a bright star transitions from its colored Moffat wings into a
|
||
white saturated core. Larger `p` approaches \(\min(x,1)\). This per-channel
|
||
response is what gives the current simplified preview pipeline the crisp white
|
||
cores seen in real astrophotography, rather than the gray, soft cores a
|
||
Reinhard curve leaves on bright stars.
|
||
|
||
`--tone-map reinhard` selects the historical \(x/(1+x)\) display transform and
|
||
is retained only to reproduce output rendered before the soft-clip operator was
|
||
introduced. FITS output always remains the pre-tone-map linear RGB framebuffer,
|
||
independent of operator and `p`; use it to measure physical flux, peak values,
|
||
and FWHM. Tone-mapped PNG/PPM data are display-referred and must not be used for
|
||
those measurements.
|
||
|
||
Commands written before this change that omit `--tone-map` actually used
|
||
Reinhard; add `--tone-map reinhard` explicitly to reproduce their PNG/PPM
|
||
output. Historical FITS/HDR benchmarks are unaffected.
|
||
|
||
### Sensor bloom (optional)
|
||
|
||
`--sensor-bloom-limit E --sensor-bloom-transfer e` (both required together,
|
||
default disabled) enable an optional post-PSF sensor saturation model applied to
|
||
the finished linear HDR framebuffer before the display operator. Each RGB
|
||
channel is processed independently and isotropically. A channel value above the
|
||
finite response limit \(E\) contributes overflow \(D=\max(H-E,0)\). A fraction
|
||
\(e\in[0,1)\) of that overflow is spread to the eight neighbours with a fixed
|
||
9-point stencil (four axial weights \(4/20\), four diagonal weights \(1/20\)),
|
||
while the remaining \((1-e)D\) is absorbed. Overflow directed outside the image
|
||
is lost at the boundary and the stencil is not renormalized there. `e=0` clamps
|
||
every over-limit channel to \(E\) without spreading to neighbours.
|
||
|
||
\(E\) is expressed in the post-exposure linear HDR renderer scale, so it scales
|
||
with `--exposure`. The model is a phenomenological limited-response
|
||
approximation, not a specific CCD/CMOS/Bayer structure. It is non-conservative:
|
||
signal is lost both to the drain \((1-e)D\) and to the image boundary, and \(e\)
|
||
controls both the per-round retained fraction and the effective propagation
|
||
distance. The mesh overlay is drawn after the model, so the diagnostic lines
|
||
neither bloom nor feed back into overflow propagation. Visual multi-scale bloom
|
||
is not implemented.
|
||
|
||
When enabled, each frame prints one `Sensor bloom:` line with the initial
|
||
saturated and final clamped channel counts, the actual and conservative round
|
||
counts, the peak and maximum initial overflow, absorbed and boundary loss,
|
||
residual clamp loss, and elapsed time.
|
||
|
||
### Fast preview mode
|
||
|
||
`--fast-mode` replaces the per-event PSF splat with a two-stage approximation:
|
||
each point-source image is deposited as a delta into an `N×N` supersampled HDR
|
||
buffer (`--fast-supersample N`, default 2), then the whole frame is convolved
|
||
once with a single global Moffat kernel and averaged down by the `N×N` block.
|
||
The kernel is the pixel-area integral of the requested Moffat at the
|
||
supersampled scale, so the requested FWHM and beta are preserved by the
|
||
downsample itself.
|
||
|
||
`--fast-deposit nearest` (default) deposits into the single nearest
|
||
supersampled pixel. It keeps the PSF shape exactly on that grid. The grid
|
||
spacing is `1/N` output pixel per axis, while the instantaneous rounding error
|
||
is bounded by `1/(2N)` per axis. `--fast-deposit bilinear` deposits into 4
|
||
adjacent pixels; it preserves the continuous centroid but broadens the profile
|
||
(measured +4.5%/+1.9%/+1.0% FWHM at N=2/3/4). The deposition experiment and
|
||
its raw output are in
|
||
[benchmarks/fast_mode_deposit_2026-09-18.md](benchmarks/fast_mode_deposit_2026-09-18.md).
|
||
|
||
Fast mode is available only in the CPU PSF backend. It ignores
|
||
`--max-cache-psf-flux` and uses one global kernel radius for every event, so a
|
||
bright event whose requested support exceeds that radius is wing-clipped and
|
||
counted in the report. `--psf-min-y` is still applied per event. Fast mode is a
|
||
preview approximation, not the physically exact per-event PSF path.
|
||
|
||
Its accuracy is governed by the deposit mode and `N`, not by the output
|
||
resolution. A nearest deposit shifts each image by at most `1/(2N)` output
|
||
pixels per axis. The separate two-dimensional centroid-distance experiment
|
||
measured worst cases of ~`0.30 px` at `N=2` and ~`0.12 px` at `N=4` for a FWHM
|
||
`2.7`, beta `4.5` PSF. These offsets remain fixed fractions of one output
|
||
pixel: with `--psf-fwhm-pixels` held fixed, a larger image does not reduce
|
||
them, and only a larger `N` (or the centroid-preserving bilinear deposit, at
|
||
the cost of a broadened profile) does.
|
||
|
||
That per-star number is not how one dense still image behaves. In the measured
|
||
production-density 4K Galactic-center frame, fast-versus-standard HDR relative
|
||
L2 is `10.44%` at `N=2` and `4.37%` at `N=4`, with a `+0.018%` total
|
||
linear-RGB-sum offset. The difference is concentrated in a small set of scalar
|
||
RGB channel samples: the top `0.01%` ranked by `|fast-base|` account for
|
||
`95.7%` (`N=2`) or `94.5%` (`N=4`) of squared-error energy, while containing
|
||
about `9%` of the summed baseline linear-RGB values. The disjoint base-value
|
||
partition puts the flux-bearing faint ranges (`[1e-3,0.06)`, `[0.06,0.1)`,
|
||
`[0.1,0.23)`, together ~`47%` of the linear-RGB sum) at ~`2-5%` local relative
|
||
RMS with signed bias at or below `0.19%`. This pattern is consistent with
|
||
partial cancellation among many faint PSFs, while resolved bright cores retain
|
||
bounded sub-pixel displacement. It does not make fast mode pixelwise or
|
||
visually lossless.
|
||
|
||
These measurements characterize independent still frames only. With nearest
|
||
deposition, a smoothly moving image jumps by one complete grid spacing when it
|
||
crosses a cell boundary: `1/N` output pixel per axis, or `0.5 px` at `N=2` and
|
||
`0.25 px` at `N=4`. Such discontinuities can appear as visible flicker even at
|
||
4K. Neither nearest mode is therefore qualified as a temporally coherent movie
|
||
output path. Bilinear deposition removes this particular centroid jump but has
|
||
the profile-broadening tradeoff above and has not been qualified as a final
|
||
movie path either. For single-frame dense-star previews, `N=2` remains the
|
||
throughput default and `N=4` the higher-fidelity point; the exact per-event
|
||
reference PSF path remains the quantitative and movie reference. The
|
||
single-frame error decomposition is reproducible with
|
||
[`scripts/fast_mode_error_decomposition.py`](scripts/fast_mode_error_decomposition.py).
|
||
|
||
In CPU builds the global convolution is a zero-padded FFTW linear convolution
|
||
(`fftw3`/`fftw3_omp`, double precision). The immutable kernel spectrum, FFTW
|
||
plans, and scratch buffers are built once when the accumulator is initialized
|
||
and reused for every frame; the disposable spatial reference remains only for
|
||
tests and benchmarks. Startup prints one `Fast FFTW:` line with the minimum
|
||
linear-convolution extent, the selected smooth FFT dimensions, worker count,
|
||
plan mode, one-time plan and kernel-transform time, and scratch bytes.
|
||
`--verbose` additionally prints per-frame stage timings (zero/pack, forward FFT,
|
||
frequency multiply, inverse FFT, crop/downsample, total). The default FFTW
|
||
planning mode is `FFTW_ESTIMATE`; `FFTW_MEASURE` can cut steady-state execution
|
||
substantially but costs a much longer one-time plan for large frames, as
|
||
recorded in
|
||
[benchmarks/fast_mode_fftw_2026-09-25.md](benchmarks/fast_mode_fftw_2026-09-25.md).
|
||
|
||
Planning can be selected explicitly. `--fast-fftw-plan measure` uses
|
||
`FFTW_MEASURE`. `--fast-fftw-plan wisdom-update --fast-fftw-wisdom FILE`
|
||
measures once and writes reusable wisdom plus a `FILE.meta` sidecar recording
|
||
FFTW version, precision, supersample, FFT dimensions, kernel radius, and worker
|
||
count. Each file is written to a unique temporary and renamed into place
|
||
individually; the pair is not a transactional update, and the sidecar is
|
||
committed last, so a wisdom without a matching sidecar is treated as a miss.
|
||
`--fast-fftw-plan wisdom --fast-fftw-wisdom FILE` imports
|
||
that wisdom and creates the plans with `FFTW_WISDOM_ONLY`; if the file, size,
|
||
supersample, worker count, version, or precision does not match, it fails with a
|
||
diagnostic instead of silently replanning. Wisdom writes go to a temporary file
|
||
and are renamed into place, so a failed write never leaves a half file.
|
||
|
||
## Progress and diagnostics
|
||
|
||
The PSF-cache completion line is printed before tracing and catalog splatting
|
||
begin. For long renders, pass `--verbose` to print catalog-prefetch state,
|
||
splat-worker local heartbeats (8, 16, 32, ... completed triangles per worker),
|
||
and image-write boundaries. The worker heartbeats use neither global progress
|
||
accounting nor cross-worker synchronization.
|
||
Verbose ray-trace output reports the initial mesh trace and refinement stages
|
||
for single frames; movie mode additionally reports each generation's sample
|
||
count and each time slab's activation and terminal-ray summary.
|
||
Movie renders always print one summary per time slab; `--verbose` also prints
|
||
the ray counts before each slab is loaded. In movie mode, `--verbose` adds one
|
||
`Movie frame N timing:` line per frame and a `Movie timing total/avg/max`
|
||
summary. The stages are `catalog_mark`, `catalog_load`, `fast_clear`, `splat`,
|
||
the FFTW breakdown (`fftw_zero_pack`, `fftw_forward`, `fftw_multiply`,
|
||
`fftw_inverse`, `fftw_crop`, `fftw_total`), `tone_map`, `output`, and
|
||
`wait` (writer-queue backpressure), plus the frame total. `output` covers only
|
||
producer-side writes, so it is zero in the async movie path; the writer's own
|
||
time is reported separately by `Movie writer summary: jobs/total/avg/max/drain`.
|
||
`Movie output pipeline wall:` starts after track load, ray tracing, lens-map
|
||
write, and the catalog union prefetch, and covers per-frame render plus async
|
||
output. `Movie end-to-end wall:` starts at `render_movie()` entry, so it
|
||
additionally includes track load, all ray tracing, the union prefetch, and the
|
||
final writer drain; it does not include the catalog, blackbody, and fast
|
||
accumulator initialization done earlier in `main()`. `Movie timing total` is
|
||
the sum of producer frame times only (it excludes tracing, prefetch, and the
|
||
final queue drain). The all-sky `Movie catalog prefetch:` line reports mark,
|
||
load+commit, and total tile time. Timing uses one clock read per bulk phase,
|
||
never inside the per-star or per-pixel hot loops.
|
||
With `--draw-mesh`, tone mapping runs only once per frame. The async writer
|
||
writes the clean RGB8 image first, then composites the producer-rasterized
|
||
premultiplied RGBA8 layer in place and writes the mesh sibling. Mesh preparation
|
||
and rasterization are included in the producer's frame total; composition and
|
||
image output are included in the writer summary.
|
||
|
||
Pass `--draw-mesh` to also write the final image-plane triangle mesh as a
|
||
`<output-stem>_mesh.png` sibling (`.ppm` in non-PNG builds). The main
|
||
tone-mapped image and any `--hdr-output` FITS file remain mesh-free. The
|
||
overlay alpha-composites one-pixel-wide, coverage-antialiased triangle edges
|
||
onto the final sRGB8 image **after** sensor bloom, tone mapping, and the sRGB
|
||
transfer, preserving mesh contrast on saturated highlights. Each vertex colors
|
||
its incident half-edges; differently classified endpoints switch color at the
|
||
edge midpoint. Shared edges are drawn once. The premultiplied sRGB RGBA8 layer
|
||
uses source-over accumulation and composition, with the same rules for
|
||
single-frame, movie, and replay output.
|
||
|
||
The default palette is Catppuccin Mocha, with opacity `0.5`:
|
||
|
||
| Vertex category | Default color | CLI override |
|
||
| --- | --- | --- |
|
||
| `ESCAPED` | Overlay1 `#7F849C` (gray) | `--mesh-color-escape` |
|
||
| `DARK` | Mauve `#CBA6F7` (purple) | `--mesh-color-dark` |
|
||
| `UNRESOLVED` | Yellow `#F9E2AF` | `--mesh-color-unresolved` |
|
||
| `INCOMPLETE` | Red `#F38BA8` | `--mesh-color-incomplete` |
|
||
| Untraced | Blue `#89B4FA` | `--mesh-color-untraced` |
|
||
|
||
Color arguments are strict sRGB `#RRGGBB` values; quote them in the shell.
|
||
`--mesh-opacity` accepts a finite number in `[0,1]`. These settings do not
|
||
implicitly enable `--draw-mesh`. For example:
|
||
|
||
```sh
|
||
--draw-mesh --mesh-color-dark '#CBA6F7' --mesh-color-unresolved '#F9E2AF' --mesh-opacity 0.8
|
||
```
|
||
|
||
`UNRESOLVED` denotes trustworthy trajectories with exhausted compute budgets
|
||
(not just accepted-step limits), while `INCOMPLETE` denotes actual history,
|
||
domain, metric, integration, I/O, or protocol failures. Different dark reasons
|
||
share one color. Coloring is a read-only visualization of the finalized mesh.
|
||
|
||
Normal progress and summaries go to stdout; warnings, errors, and Debug
|
||
diagnostics go to stderr. Successful runs exit `0` even if warnings are emitted.
|
||
|
||
Incomplete ray failures and budget-exhausted frames report their reasons and
|
||
affected frames on stderr even without `--verbose`; `--verbose` and Debug builds
|
||
add bounded per-sample localization (film position and integration cost).
|
||
|
||
## HDR output
|
||
|
||
To preserve a single-frame render for later exposure and tone-mapping work,
|
||
build with `ENABLE_HDR=1` as described in [build.md](build.md). This produces `build/Release/schwarzschild_sky`,
|
||
which accepts `--hdr-output`. This switch writes the HDR file next to the
|
||
ordinary output, replacing its extension with `_HDR.fits`; for example,
|
||
`--output output/imgs/ring.png --hdr-output` writes
|
||
`output/imgs/ring_HDR.fits`. It writes the pre-tone-mapping RGB
|
||
framebuffer as a three-plane, 32-bit float FITS image. Values remain linear HDR
|
||
at the renderer's arbitrary scale; no tone mapping or per-frame normalization
|
||
is applied, and the `--tone-map` operator and `--tone-map-p` value never affect
|
||
this file. The `--sensor-bloom-*` model runs only after this file is written, so
|
||
enabling bloom never changes the stored FITS. The ordinary
|
||
binaries do not contain this option or writer.
|
||
|
||
The FITS header describes a synthetic 8640-by-5760, 36-by-24 mm full-frame
|
||
sensor with 4.1667 um pixels. Each render records its active centered crop and
|
||
derives `FOCALLEN` from that render's width and horizontal `--fov-deg`; these
|
||
camera fields support plate-solving workflows but do not calibrate flux.
|