# adaptive_step_bounds_2026-10-05 Long-term benchmark for calibrating the DP54 adaptive-integrator step bounds (`min_step` / `max_step`) of the past-directed null-geodesic tracer, and the public record of the one authorised 4K Schwarzschild `max_step` 2-vs-8 pair. The benchmark runs the public endpoint `geodesic_trace_past()`, the production DP core, and bounded adaptive-mesh renders against the analytic Minkowski / Schwarzschild / Alcubierre backends. It does not modify production sources. ## Build (required before renderer scripts) The Python mesh/4K drivers start from an existing renderer binary. The shell endpoint/core drivers compile their own small harnesses against production code. ``` make -j4 BUILD_TYPE=Release SPACETIME=schwarzschild backend # binary: build/Release/schwarzschild_sky ``` ## Files | file | role | |---|---| | `a_public_endpoints.c` | Experiment A: step-bound scans through the public endpoint, tol=1e-12 reference, RK4 h-halving on near-critical rays. | | `b_actual_h.c` | Experiment B: textually includes `src/geodesic.c`, drives the production `dp_advance_one` on finite trusted intervals to measure observed accepted `h`. | | `critical_ref_check.c` | Critical-classification probe (r30/r100 inward, r2.1 outward; `theta_crit` and `theta_crit±1e-7`; tol/max_step ladder). | | `mesh_maps.py` | Pure GRLENS v3 `load`/`compare`; no renderer, no untracked import. | | `run_mesh.py` | 8 bounded 320x180 candidates (R100 and R2.1, `max_step` .5/2/8/32). | | `run_mesh_reference.py` | 2 tight-tolerance references and comparisons against the 8 maps. | | `summarize_4k.py` | Postprocess existing maps only; never invokes a renderer. | | `run_4k_pair.py` | One-task, at-most-two-attempt 4K pair; refuses to run without `--authorize-two-4k`. | | `record_environment.py` | Provenance (git/hashes/compiler/cpu/build/run commands); no render. | | `failure_floor.c` | Failure-path floor diagnostic; includes `../../tests/test_geodesic_adaptive.c`. | | `run_failure_floor.sh` | Builds/runs `failure_floor.c` with `-DGEODESIC_EVENT_TESTING`. | | `run_limited.sh` | Builds/runs Experiment A + B; all output under `OUT_DIR`. | | `run_critical_ref.sh` | Builds/runs `critical_ref_check.c`; all output under `OUT_DIR`. | | `summarize.py` | `OUT_DIR` CLI arg (default `local/adaptive_step_bounds_2026-10-05`); writes `OUT_DIR/logs/summary.txt`. | | `results/` | Tracked long-term evidence (see below). | ## Reproduction ``` AB_OUT="$PWD/local/adaptive_step_bounds_2026-10-05" OUT="$PWD/local/adaptive_bounds_mesh" bash benchmarks/adaptive_step_bounds_2026-10-05/run_limited.sh "$AB_OUT" bash benchmarks/adaptive_step_bounds_2026-10-05/run_critical_ref.sh "$AB_OUT" python3 benchmarks/adaptive_step_bounds_2026-10-05/summarize.py "$AB_OUT" python3 benchmarks/adaptive_step_bounds_2026-10-05/run_mesh.py "$OUT" python3 benchmarks/adaptive_step_bounds_2026-10-05/run_mesh_reference.py "$OUT" bash benchmarks/adaptive_step_bounds_2026-10-05/run_failure_floor.sh "$OUT" # 4K only after explicit fresh user authorization for that task: python3 benchmarks/adaptive_step_bounds_2026-10-05/run_4k_pair.py \ --authorize-two-4k --out local/adaptive_bounds_4k ``` Defaults: A/B `OUT_DIR` is `local/adaptive_step_bounds_2026-10-05`; mesh `OUT_DIR` is `local/adaptive_bounds_mesh`; 4K `--out` default is `local/adaptive_bounds_4k`. All logs/CSV go to `OUT_DIR`; the tracked benchmark directory holds only sources and `results/`. Scratch binaries go to `/tmp/opencode/step_bounds/`. The `run_4k_pair.py` driver refuses to run without `--authorize-two-4k`, keeps an attempt ledger, and has no automatic retry. Authorisation does not extend to later tasks or routine tests. All catalog inputs are generated inline by the scripts (column header `longitude_deg,latitude_deg,temperature_K,amplitude`, one 1e-30 star); no external survey CSV is read. ## Results ### 1. Public-endpoint step-bound scan (Experiment A/B) Environment and full table inputs: `results/source_environment.txt`, `results/summary_tables.txt`; full stdout: `results/expA.log`, `results/expB.log`. Per-ray raw CSV (not tracked) is reproduced by `run_limited.sh`. | backend | worst near-critical sky error vs tol=1e-12 ref | notes | |---|---|---| | Schwarzschild r30 | 2.82e-4 rad (`max_step`≥1); 8.67e-5 rad at 0.25 | r100 8.0e-5 / 2.6e-5; r2.1 2.87e-4 | | Minkowski | Δn = 0, Δg/g = 0 (analytic) | crossing x/t ≤ 7.7e-11 at `max_step`≤256 | | Alcubierre | ≤1.8e-10 rad, no class change | 8x vs 1x `initial` RHS cost 4.3x–5.7x | - Dark terminal `L−L0` margin agrees with the reference to ≤2.1e-13 in every config; near-critical dark stop times differ up to ~41 time units between tol=1e-9 and tol=1e-12 (orbit-count sensitivity, reported not hidden). - Schwarzschild floor plateau: `min_step` `1e-14 … 1e-2` give bit-identical results and RHS; only `0.1` breaks one r1.5 ray (`INCOMPLETE/INTEGRATION_ERROR`). No evidence to raise the default floor, and no evidence that `1e-12` is uniquely optimal (it is never active). - Observed accepted `h` (Experiment B, finite window `T=min(20, ref span)`): **Schwarzschild only** — r1.5 radial 0.0921, all other sampled Schwarzschild rays ≥0.1; Alcubierre minimum = its initial step (0.05 / 0.005 / 0.0005); Minkowski 1.0. This is a finite-window diagnostic, **not** a full-trajectory minimum. - RK4 h-halving worst (r2.1, `theta_crit−1e-7`): 1.04e-5 rad. ### 2. Bounded 320x180 adaptive mesh Numeric inputs: `results/mesh_summary.json`, `results/mesh_reference_summary.json`. Scene: R100/FOV45/jacobian0.2 and R2.1/outward/FOV90; DP54, hmin 1e-12, initial 0.1, tol 1e-9 (tight reference 1e-12), 8 threads, one-point dim catalog. | scene | `max_step` | persistent vertices | RHS | wall (s) | |---|---:|---:|---:|---:| | R100 | 0.5 | 3158 | 16813349 | 2.071 | | R100 | 2 | 3158 | 6030185 | 0.979 | | R100 | 8 | 3158 | 4230835 | 0.823 | | R100 | 32 | 3158 | 4081609 | 0.830 | | R2.1 | 0.5 | 5496 | 20136508 | 1.974 | | R2.1 | 2 | 5496 | 9548434 | 1.178 | | R2.1 | 8 | 5496 | 7389760 | 0.970 | | R2.1 | 32 | 5496 | 7197806 | 0.972 | - 2→8 RHS reduction 29.84% (R100) / 22.61% (R2.1); 8→32 only 3.53% / 2.60%. - Identical terminal classifications and mesh/sample counts for all four caps. - Error vs tight reference: R100 max sky 9.216e-8 (hmax2) / 8.744e-8 (hmax8), mutual 2-vs-8 difference 7.063e-9 rad; R2.1 max sky 1.4772e-5 rad at both, mutual difference 8.370e-11 rad. The larger R2.1 reference error is tolerance/conditioning sensitive and is not cured by a smaller `max_step`. - Max log-frequency differences ≤8.51e-10 (R100) / ≤2.10e-9 (R2.1). - Short single-run wall times include setup/cache/output; not a timing study. ### 3. One authorised 4K pair (R100, 3840x2160, FOV45) Raw terminal output (byte-for-byte): `results/4k_hmax2.log`, `results/4k_hmax8.log`; provenance: `results/4k_metadata.json`, `results/4k_attempts.json`; postprocessed: `results/4k_final_summary.json`, `results/4k_raw_summary.log`. | metric | `max_step`=2 | `max_step`=8 | |---|---:|---:| | persistent vertices / triangles | 66045 / 131338 | 66045 / 131338 | | outcomes | 62419 ESC, 3626 DARK | 62419 ESC, 3626 DARK | | unresolved / error | 0 | 0 | | persistent accepted | 15,768,677 | 10,183,020 | | persistent rejected | 97,176 | 194,781 | | persistent RHS | 130,409,139 | 91,991,144 | | trace+refine time (s) | 10.774 | 7.805 | | process wall (s) | 11.3893 | 8.38925 | - Triangle payload is bit-identical; shared film identities 66045, terminal mismatches 0; max mutual endpoint sky difference 7.9544e-7 rad, max logg difference 3.4711e-11. - Persistent-vertex RHS reduction −29.46%; accepted −35.42%; wall −26.34%. - **Cost scope:** the v3 map stores final persistent vertices only. The requested-sample count including discarded refinement probes is 116308 (> 66045), so discarded probes contribute no stored cost and these sums must not be read as full-render executed RHS totals. - **Timing:** one sequential pair only; no statistical timing confidence. - PSF splat cached/direct = 0 and `--psf-min-y` discarded 1965 (hmax2) / 1954 (hmax8) events; this does not measure image error. The PNGs are dark diagnostic backgrounds; no 2MASS catalog was used. ### 4. Critical-classification reference check Full output: `results/critical_ref.log`, `results/critical_ref.csv` (36 endpoint traces). The exact `theta_crit` direction is a classification separatrix; it is reported, never asserted. - r2.1 `theta_crit`: ref tol1e-12/h0.25 vs tol1e-13/h0.125 sky difference **2.33638 rad** — the exactly-critical escape direction is not independently converged. `EXACT_CRITICAL_REFERENCE_NOT_CONVERGED`. - r30 `theta_crit` (DARK) stop-time difference 13.056; r100 `theta_crit` (DARK) 12.700. At tol1e-11, h2-vs-h8 at r2.1 `theta_crit` is 5.09e-5 rad. - Neighbours `±1e-7`: both reference tiers agree (same outcome/reason/end, clean, positive steps) at every neighbour. Escaped-neighbour sky differences between the layers are ~1e-7 rad (9.47e-8 r30, 2.74e-8 r100, 1.42e-7 r2.1), consistent with the loose 1e-4 magnitude trend; DARK neighbours report stop-time differences 1.5e-7 … 7.5e-7. - `CRITICAL_REFERENCE_CHECK status=0` (no neighbour reference disagreement or error endpoint). ## Decision summary - **Schwarzschild `max_step = 8`**: supported by the far-field cost plateau (2→8 saves ~23–30% persistent RHS; 8→32 only ~3%) together with the tight reference and 4K 2-vs-8 comparisons. This is not claimed as globally optimal and gives no guarantee for arbitrary near-critical/shadow directions. - **`min_step = 1e-12`**: retained as a non-binding numerical guard; not a measured optimum. - **Minkowski `max_step = 16`** and **Alcubierre `max_step = 8 × initial`**: retained as conservative caps. Their bounded accuracy checks support keeping these values, not claiming cost optimality; larger Alcubierre caps still reduce RHS, but have larger frequency errors in some cases. - The R2.1 exactly-critical difference is tolerance-sensitive; `critical_ref_check.c` confirms the exact direction is not independently converged and the report does not treat it as certified. ## Caveats - Analytic backends only; no time-dependent numerical-relativity metric was exercised. Bounds and the `unit = 1` calibration are not transferable to NR. - Experiment B samples `T = min(20, reference span)`; no full-trajectory minimum-step claim is made. - The 4K numbers are one sequential pair; no timing-stability confidence and no image-error bound (inverse-lens-map Jacobian amplification is separate). - Sky-angle differences divided by pixel angle are not inverse-lens-map image error bounds. - `failure_floor.c` reuses the regression invalid-metric temporal-domain fixture; it is not a physical shadow or floor calibration. ## Verification of the selected defaults and reproduction tools `make -B -j4 BUILD_TYPE=Debug test` passed after selecting the Schwarzschild cap of 8 and adding CLI provenance assertions for the default bounds. Full terminal output is retained locally at `local/step_bounds_mesh/final_debug_tests.log` (`STEP_BOUNDS_FULL_DEBUG_EXIT=0`). The explicit-RK4 HDR reference tests remain bit-identical. The Release Schwarzschild binary was rebuilt with the selected default. The self-contained limited driver was also executed with the relative output path `local/step_bounds_reproduction`: 1500 endpoint traces and 2427 observed finite-window steps completed, including summary generation. Small angular differences use `atan2(|cross|,dot)` rather than rounded `acos(dot)`. Running the 4K driver without fresh explicit authorization refuses with exit code 2; this check launches no renderer. No further 4K attempt was made.