Replace the position capture cutoff with a camera-relative dark threshold
shared by every backend, and carry explicit outcome/reason provenance
through the ray, RayPool, adaptive mesh, lens-map and replay paths.
- eval/eval_slab return SpacetimePointStatus; remove SPACETIME_RAY_CAPTURED
and the Schwarzschild capture radius; decouple observer construction from
ray position.
- RayEndpoint stores RayOutcome/RayReason plus the last trusted state;
budget exhaustion is retryable UNRESOLVED, data/integration failures are
INCOMPLETE.
- Normal dark terminal is L - L0 >= --dark-threshold (default 8), with L0
taken at the camera event and kept distinct from the worldtube entry
energy; photon energy and frequency ratio are never reset.
- Implement E/D/U triangle decisions with merged budget retries, persistent
probe witnesses promoted in place by vertex identity, conformity settling,
and approximate-black boundary provenance with achieved-scale statistics.
- Add RayPool continuation state and per-ray step budgets.
- Bump lens-map to v2 with explicit end/outcome/reason, approx_black,
threshold/retry/geometry provenance and per-frame retry counts; reject v1.
- Gate production output on incomplete/error results, overridable with
--allow-incomplete.
- Update AGENTS.md, the design document and usage docs; add the termination
oracle and regression coverage.
make -B -j4 BUILD_TYPE=Debug test passes with bit-identical reference HDRs.
Replace radius-only escape termination with a common asymptotic exterior protocol: declared ends, moving escape worldtubes, directed inside->outside crossings, and a PENDING_ENTRY lifecycle shared by single-frame and movie tracing.
Add an analytic Carlson-integral Schwarzschild monopole exterior (angle primitive, bracketed turning radius, ingoing Kerr-Schild coordinate-time transfer, conserved-energy frequency) so a camera outside the escape sphere is traced through an entry event.
Make the lifecycle tri-state (no ends / ready / protocol error), carry end_id through the endpoint and lens mesh, validate sources in constructors via spacetime_source_finalize(), and refresh the Schwarzschild reference images for the corrected finish.
- Gather the union of every movie frame's all-sky tiles once and read them in bounded batches, replacing the per-frame prefetch scan and log.
- Fuse the fast supersampled-buffer clear into the FFTW pack pass so a resolved frame starts clean with no serial memset.
- Split parallel HDR->RGB8 tone mapping from PNG encoding.
- Add a bounded single-producer/single-writer movie output queue used by both observer movies and multi-frame imported lens maps, with writer timing and error propagation.
- Add --png-compression-level and --movie-output-workers shared|reserve-one.
- Add --fast-fftw-plan estimate|measure|wisdom|wisdom-update with strict wisdom identity sidecars.
- Add staged movie timing, regression tests, and docs.
Add an optional CPU preview path that deposits each point-source image as a
supersampled delta and resolves the whole frame with one global Moffat
convolution plus an N x N box average, instead of splatting a per-event PSF.
- optics: FastPsfAccumulator builds the pixel-area-integral kernel of the
target Moffat at the supersampled scale (width N*alpha, same beta). The
1/N^2 box average then reproduces the final pixel-area integral, so the
requested FWHM and beta are preserved without renormalisation. Deposits are
per-cell atomic adds; resolve accumulates into the caller's HDR buffer.
- frame: fast branch in frame_splat_catalog with one shared supersampled
buffer and a single resolve per frame; the accumulator is reused across
movie frames and built from the map dimensions on lens-map import.
- main: --fast-mode, --fast-supersample N (1..8, default 2) and
--fast-deposit nearest|bilinear (default nearest). CPU-only and rejected in
the HIP/dummy backends; --psf-min-y still applies per event while
--max-cache-psf-flux does not.
- The deposition scheme was chosen by scripts/fast_mode_deposit_error.py:
nearest keeps the PSF shape exactly with <= 0.5/N px position quantization;
bilinear keeps the exact centroid but broadens FWHM and beta. Recorded in
benchmarks/fast_mode_deposit_2026-09-18.md.
- tests/test_frame.c covers fast nearest vs the direct evaluator at the snapped
centre, flux conservation, bilinear centroid, min-Y discard, frame plumbing,
and HDR accumulation onto a non-zero background.
- benchmarks/fast_mode_cpu_2026-09-18.md records a ~10x speedup on the 2MASS
galactic-centre field with small tone-mapped differences.
Move adaptive tile16 selection and support-expanded CPU binning into a
per-producer HipPsfPreparedChunk computed before acquiring the shared
sink lock. The sink keeps sole ownership of the GPU HDR, stream, device
scratch and upload staging; submit_prepared runs under the lock and
preserves the existing single-stream ordering and direct fallback path.
The legacy hip_psf_sink_submit entry point prepares an internal chunk so
single-threaded replay stays compatible, and default accumulation remains
atomic.
Reuse production catalog mapping and per-worker chunk boundaries to report triangle and spatial distributions without initializing HIP or writing image outputs. Include a protected full-catalog runner and build documentation.
Capture prepared events through a test-only producer consumer and measure serial and indexed parallel generation independently. Compare the production atomic kernel with bounded pixel-owned tile reductions while preserving double HDR, cache support, and direct fallback ordering.
Validate captured inputs, edge and multichunk fixtures, CPU/HIP renderer regressions, and 45 isolated final replays. Keep the production HIP algorithm unchanged because gains depend on event distribution.
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.
Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
Check spherical area weight sums against a fixed 1e-8 tolerance and normalize accepted weights before interpolation. Skip invalid images in both catalog paths.
Add a thin-triangle regression for both parities, document the measured tolerance margin, and refresh HDR fixtures with strict comparisons restored.