Files
GR-raytracing/benchmarks/adaptive_step_bounds_2026-10-05/README.md
T
wyj 0a46a7095b Feat: Complete adaptive geodesic tracing with DP54
Add error-controlled DP5(4) integration and trusted first-crossing localization, including non-monotonic energy thresholds and representable-time stepping.

Preserve adaptive state and independent step/time retry grants across RayPool, refinement and movie scheduling. Expose numerical controls, record actual persistent-sample costs, and add v3 lens-map provenance with legacy v2 RK4 import.

Use DP54 by default and select an 8M Schwarzschild maximum step from bounded scans and a two-run 4K comparison. Retain the conservative minimum-step guard and document critical-ray and backend capability limits. Archive self-contained benchmark inputs and raw output; keep fixed RK4 HDR references explicit.

Validation: make -B -j4 BUILD_TYPE=Debug test passed; explicit RK4 HDR references have zero differences. Bounded convergence checks, benchmark reproduction, Release build and focused reviews passed. No numerical-relativity backend is added.
2026-10-05 20:27:42 -04:00

12 KiB
Raw Blame History

adaptive_step_bounds_2026-10-05

Long-term benchmark for calibrating the DP54 adaptive-integrator step bounds (min_step / max_step) of the past-directed null-geodesic tracer, and the public record of the one authorised 4K Schwarzschild max_step 2-vs-8 pair.

The benchmark runs the public endpoint geodesic_trace_past(), the production DP core, and bounded adaptive-mesh renders against the analytic Minkowski / Schwarzschild / Alcubierre backends. It does not modify production sources.

Build (required before renderer scripts)

The Python mesh/4K drivers start from an existing renderer binary. The shell endpoint/core drivers compile their own small harnesses against production code.

make -j4 BUILD_TYPE=Release SPACETIME=schwarzschild backend
# binary: build/Release/schwarzschild_sky

Files

file role
a_public_endpoints.c Experiment A: step-bound scans through the public endpoint, tol=1e-12 reference, RK4 h-halving on near-critical rays.
b_actual_h.c Experiment B: textually includes src/geodesic.c, drives the production dp_advance_one on finite trusted intervals to measure observed accepted h.
critical_ref_check.c Critical-classification probe (r30/r100 inward, r2.1 outward; theta_crit and theta_crit±1e-7; tol/max_step ladder).
mesh_maps.py Pure GRLENS v3 load/compare; no renderer, no untracked import.
run_mesh.py 8 bounded 320x180 candidates (R100 and R2.1, max_step .5/2/8/32).
run_mesh_reference.py 2 tight-tolerance references and comparisons against the 8 maps.
summarize_4k.py Postprocess existing maps only; never invokes a renderer.
run_4k_pair.py One-task, at-most-two-attempt 4K pair; refuses to run without --authorize-two-4k.
record_environment.py Provenance (git/hashes/compiler/cpu/build/run commands); no render.
failure_floor.c Failure-path floor diagnostic; includes ../../tests/test_geodesic_adaptive.c.
run_failure_floor.sh Builds/runs failure_floor.c with -DGEODESIC_EVENT_TESTING.
run_limited.sh Builds/runs Experiment A + B; all output under OUT_DIR.
run_critical_ref.sh Builds/runs critical_ref_check.c; all output under OUT_DIR.
summarize.py OUT_DIR CLI arg (default local/adaptive_step_bounds_2026-10-05); writes OUT_DIR/logs/summary.txt.
results/ Tracked long-term evidence (see below).

Reproduction

AB_OUT="$PWD/local/adaptive_step_bounds_2026-10-05"
OUT="$PWD/local/adaptive_bounds_mesh"

bash  benchmarks/adaptive_step_bounds_2026-10-05/run_limited.sh "$AB_OUT"
bash  benchmarks/adaptive_step_bounds_2026-10-05/run_critical_ref.sh "$AB_OUT"
python3 benchmarks/adaptive_step_bounds_2026-10-05/summarize.py "$AB_OUT"

python3 benchmarks/adaptive_step_bounds_2026-10-05/run_mesh.py "$OUT"
python3 benchmarks/adaptive_step_bounds_2026-10-05/run_mesh_reference.py "$OUT"
bash  benchmarks/adaptive_step_bounds_2026-10-05/run_failure_floor.sh "$OUT"

# 4K only after explicit fresh user authorization for that task:
python3 benchmarks/adaptive_step_bounds_2026-10-05/run_4k_pair.py \
        --authorize-two-4k --out local/adaptive_bounds_4k

Defaults: A/B OUT_DIR is local/adaptive_step_bounds_2026-10-05; mesh OUT_DIR is local/adaptive_bounds_mesh; 4K --out default is local/adaptive_bounds_4k. All logs/CSV go to OUT_DIR; the tracked benchmark directory holds only sources and results/. Scratch binaries go to /tmp/opencode/step_bounds/.

The run_4k_pair.py driver refuses to run without --authorize-two-4k, keeps an attempt ledger, and has no automatic retry. Authorisation does not extend to later tasks or routine tests.

All catalog inputs are generated inline by the scripts (column header longitude_deg,latitude_deg,temperature_K,amplitude, one 1e-30 star); no external survey CSV is read.

Results

1. Public-endpoint step-bound scan (Experiment A/B)

Environment and full table inputs: results/source_environment.txt, results/summary_tables.txt; full stdout: results/expA.log, results/expB.log. Per-ray raw CSV (not tracked) is reproduced by run_limited.sh.

backend worst near-critical sky error vs tol=1e-12 ref notes
Schwarzschild r30 2.82e-4 rad (max_step≥1); 8.67e-5 rad at 0.25 r100 8.0e-5 / 2.6e-5; r2.1 2.87e-4
Minkowski Δn = 0, Δg/g = 0 (analytic) crossing x/t ≤ 7.7e-11 at max_step≤256
Alcubierre ≤1.8e-10 rad, no class change 8x vs 1x initial RHS cost 4.3x–5.7x
  • Dark terminal L−L0 margin agrees with the reference to ≤2.1e-13 in every config; near-critical dark stop times differ up to ~41 time units between tol=1e-9 and tol=1e-12 (orbit-count sensitivity, reported not hidden).
  • Schwarzschild floor plateau: min_step 1e-14 … 1e-2 give bit-identical results and RHS; only 0.1 breaks one r1.5 ray (INCOMPLETE/INTEGRATION_ERROR). No evidence to raise the default floor, and no evidence that 1e-12 is uniquely optimal (it is never active).
  • Observed accepted h (Experiment B, finite window T=min(20, ref span)): Schwarzschild only — r1.5 radial 0.0921, all other sampled Schwarzschild rays ≥0.1; Alcubierre minimum = its initial step (0.05 / 0.005 / 0.0005); Minkowski 1.0. This is a finite-window diagnostic, not a full-trajectory minimum.
  • RK4 h-halving worst (r2.1, theta_crit−1e-7): 1.04e-5 rad.

2. Bounded 320x180 adaptive mesh

Numeric inputs: results/mesh_summary.json, results/mesh_reference_summary.json. Scene: R100/FOV45/jacobian0.2 and R2.1/outward/FOV90; DP54, hmin 1e-12, initial 0.1, tol 1e-9 (tight reference 1e-12), 8 threads, one-point dim catalog.

scene max_step persistent vertices RHS wall (s)
R100 0.5 3158 16813349 2.071
R100 2 3158 6030185 0.979
R100 8 3158 4230835 0.823
R100 32 3158 4081609 0.830
R2.1 0.5 5496 20136508 1.974
R2.1 2 5496 9548434 1.178
R2.1 8 5496 7389760 0.970
R2.1 32 5496 7197806 0.972
  • 2→8 RHS reduction 29.84% (R100) / 22.61% (R2.1); 8→32 only 3.53% / 2.60%.
  • Identical terminal classifications and mesh/sample counts for all four caps.
  • Error vs tight reference: R100 max sky 9.216e-8 (hmax2) / 8.744e-8 (hmax8), mutual 2-vs-8 difference 7.063e-9 rad; R2.1 max sky 1.4772e-5 rad at both, mutual difference 8.370e-11 rad. The larger R2.1 reference error is tolerance/conditioning sensitive and is not cured by a smaller max_step.
  • Max log-frequency differences ≤8.51e-10 (R100) / ≤2.10e-9 (R2.1).
  • Short single-run wall times include setup/cache/output; not a timing study.

3. One authorised 4K pair (R100, 3840x2160, FOV45)

Raw terminal output (byte-for-byte): results/4k_hmax2.log, results/4k_hmax8.log; provenance: results/4k_metadata.json, results/4k_attempts.json; postprocessed: results/4k_final_summary.json, results/4k_raw_summary.log.

metric max_step=2 max_step=8
persistent vertices / triangles 66045 / 131338 66045 / 131338
outcomes 62419 ESC, 3626 DARK 62419 ESC, 3626 DARK
unresolved / error 0 0
persistent accepted 15,768,677 10,183,020
persistent rejected 97,176 194,781
persistent RHS 130,409,139 91,991,144
trace+refine time (s) 10.774 7.805
process wall (s) 11.3893 8.38925
  • Triangle payload is bit-identical; shared film identities 66045, terminal mismatches 0; max mutual endpoint sky difference 7.9544e-7 rad, max logg difference 3.4711e-11.
  • Persistent-vertex RHS reduction −29.46%; accepted −35.42%; wall −26.34%.
  • Cost scope: the v3 map stores final persistent vertices only. The requested-sample count including discarded refinement probes is 116308 (> 66045), so discarded probes contribute no stored cost and these sums must not be read as full-render executed RHS totals.
  • Timing: one sequential pair only; no statistical timing confidence.
  • PSF splat cached/direct = 0 and --psf-min-y discarded 1965 (hmax2) / 1954 (hmax8) events; this does not measure image error. The PNGs are dark diagnostic backgrounds; no 2MASS catalog was used.

4. Critical-classification reference check

Full output: results/critical_ref.log, results/critical_ref.csv (36 endpoint traces). The exact theta_crit direction is a classification separatrix; it is reported, never asserted.

  • r2.1 theta_crit: ref tol1e-12/h0.25 vs tol1e-13/h0.125 sky difference 2.33638 rad — the exactly-critical escape direction is not independently converged. EXACT_CRITICAL_REFERENCE_NOT_CONVERGED.
  • r30 theta_crit (DARK) stop-time difference 13.056; r100 theta_crit (DARK) 12.700. At tol1e-11, h2-vs-h8 at r2.1 theta_crit is 5.09e-5 rad.
  • Neighbours ±1e-7: both reference tiers agree (same outcome/reason/end, clean, positive steps) at every neighbour. Escaped-neighbour sky differences between the layers are ~1e-7 rad (9.47e-8 r30, 2.74e-8 r100, 1.42e-7 r2.1), consistent with the loose 1e-4 magnitude trend; DARK neighbours report stop-time differences 1.5e-7 … 7.5e-7.
  • CRITICAL_REFERENCE_CHECK status=0 (no neighbour reference disagreement or error endpoint).

Decision summary

  • Schwarzschild max_step = 8: supported by the far-field cost plateau (2→8 saves ~23–30% persistent RHS; 8→32 only ~3%) together with the tight reference and 4K 2-vs-8 comparisons. This is not claimed as globally optimal and gives no guarantee for arbitrary near-critical/shadow directions.
  • min_step = 1e-12: retained as a non-binding numerical guard; not a measured optimum.
  • Minkowski max_step = 16 and Alcubierre max_step = 8 × initial: retained as conservative caps. Their bounded accuracy checks support keeping these values, not claiming cost optimality; larger Alcubierre caps still reduce RHS, but have larger frequency errors in some cases.
  • The R2.1 exactly-critical difference is tolerance-sensitive; critical_ref_check.c confirms the exact direction is not independently converged and the report does not treat it as certified.

Caveats

  • Analytic backends only; no time-dependent numerical-relativity metric was exercised. Bounds and the unit = 1 calibration are not transferable to NR.
  • Experiment B samples T = min(20, reference span); no full-trajectory minimum-step claim is made.
  • The 4K numbers are one sequential pair; no timing-stability confidence and no image-error bound (inverse-lens-map Jacobian amplification is separate).
  • Sky-angle differences divided by pixel angle are not inverse-lens-map image error bounds.
  • failure_floor.c reuses the regression invalid-metric temporal-domain fixture; it is not a physical shadow or floor calibration.

Verification of the selected defaults and reproduction tools

make -B -j4 BUILD_TYPE=Debug test passed after selecting the Schwarzschild cap of 8 and adding CLI provenance assertions for the default bounds. Full terminal output is retained locally at local/step_bounds_mesh/final_debug_tests.log (STEP_BOUNDS_FULL_DEBUG_EXIT=0). The explicit-RK4 HDR reference tests remain bit-identical. The Release Schwarzschild binary was rebuilt with the selected default.

The self-contained limited driver was also executed with the relative output path local/step_bounds_reproduction: 1500 endpoint traces and 2427 observed finite-window steps completed, including summary generation. Small angular differences use atan2(|cross|,dot) rather than rounded acos(dot). Running the 4K driver without fresh explicit authorization refuses with exit code 2; this check launches no renderer. No further 4K attempt was made.