Files
GR-raytracing/benchmarks/quadratic_precision
wyj 7c980c35aa Benchmark: Compare quadratic entry precision and arithmetic cost
Generate plain long-double and experimental FMA variants from the production entry kernel. Compare exact-rational references, geometric validation, fallback counts and repeated timings with self-contained fixtures.

Validate cached build dependencies and reject stale timing output. Count all unconfirmed entry outcomes independently of reference classification.
2026-10-09 00:15:02 -04:00
..

quadratic_precision — stable entry quadratic: precision vs. cost

Self-contained, reproducible benchmark comparing four arithmetic realisations of the same asymptotic-entry algebra. Nothing here modifies production code, the Makefile, or git state.

Question

The production camera pre-route solves the relative-distance quadratic

F(s) = |d + q s|^2 - (R0 - rr s)^2 = a s^2 + b s + c
d = x_cur - c_frame,  q = w_frame + v_frame

in long double, with a stable root formula, an ordinary discriminant b*b - 4*a*c and entry slope, a common power-of-two scaling, an uncertainty band, geometric validation and a numerical fallback. Which matters for accuracy and which for speed?

Four variants share the entire rest of the module (scaling, uncertainty band, validation, fallback, dispatch) and are generated from the current src/asymptotic.c by rewriting exactly one lexical region:

variant coefficients type discriminant / slope
ld_plain as production long double plain b*b - 4*a*c, plain slope
ld_fma as production long double compensated fmal (experimental arm)
double_fma ordinary += double compensated fma
double_fma_coeff fma dots, fma products double compensated fma

double_fma is the control that isolates type precision from FMA; the two double variants isolate coefficient accumulation.

Generation and build

build.py locates the kernel region between two lexical anchors in the current working-tree src/asymptotic.c:

  • start: the comment /* Long-double coefficients of the relative-distance quadratic
  • end: the forward declaration static void minkowski_route_entry(...)

It asserts each anchor and the presence of entry_quadratic_coeffs, entry_solve, entry_discriminant exactly once, then emits four full-module copies under <output>/generated/. Every other line (dispatch, fallback, asymptotic_route_camera, Schwarzschild path, …) is copied verbatim. If a future edit moves an anchor or changes the region, generation fails rather than silently measuring the wrong code.

ld_plain retains the production kernel verbatim; ld_fma rewrites only the discriminant body and the entry slope. Double variants replace long double→double, fmal/fabsl/fmaxl/frexpl/ scalbnl/sqrtl/copysignl→their double forms, LDBL_*→DBL_*, and strip L suffixes from literals. A guard rejects any remaining long double promotion, L literal or *l/fmal call in the double kernel section, so no double expression can be re-promoted to long double. double_fma_coeff additionally changes the coefficient accumulation to fma.

The probe #includes the generated module, so the private static entry_quadratic_coeffs/entry_solve and the public asymptotic_route_camera that are exercised are the actual generated code, not an independent copy.

Compile flags, identical for all four variants:

-std=c11 -march=native -O2 -DNDEBUG -ffp-contract=off -fopenmp

-ffp-contract=off prevents implicit FMA contraction from silently changing the ld_plain arm; explicit fma/fmal still lower as written.

The probe links a minimal production subset (geodesic.c, asymptotic_entry.c, asymptotic_schwarzschild.c, spacetime_common.c, spacetime_minkowski.c) plus libm. No FFTW, PNG, catalog or observer-track data are needed.

Inputs

All fixtures are frozen as IEEE-754 hex literals (cases.json and the generated C header are bit-identical), including the two production grazing rows reproduced from the Alcubierre pre-route (x=(0,-24,0), centre=2t, v=2, R=5, canonical w hex values). Families:

  • curated / adversarial: head-on hit/miss, clear miss, grazing y=nextafter(R,0), exact tangent, near-boundary, 1e10 + 0.5, R=1 cancellation, large-t centre=0.1t, exact linear a==0 (inward and initially outward), translated-origin cancellation.
  • fixed: axes and three oblique frames, radii 1e-100…1e100, D/R ∈ {1+ulp,2,10,100,512,1024,1e4,1e8,1e10}, impact ratios {0,.5,.99,1-1e-6,nextafter(1,0),1,nextafter(1,∞),1+1e-6,1.1}. The 512/1024 ratios sample the crossover where long-double coefficient rounding meets the 128 ε_D R² geometry tolerance (≈512) and the double crossover (≈11), separating fallback rates from the all-UNCERTAIN 1e10 regime.
  • moving: v ∈ {0,.1,2,10} along a transverse direction.
  • growing: rr = -1 and nextafter(-1,∓∞), inward and initially-outward photons (w = ∓tow), axis and oblique.
  • shrinking: rr>0, labelled/domain-checked algebra (radius would go negative on the open past; roots outside R0 - rr·s > 0 are non-physical).
  • route: rr = 0 plus rr<0 growing inward/outward families (positive radius on the whole past, valid_t_min=-1e300) across all three direction frames, driven through the public asymptotic_route_camera. The oblique-2 growing-outward family reproduces the double kernel false-MISS cases, where a kernel MISS bypasses the fallback and the route escapes directly; the route reference uses the actual normalised canonical w, so normalisation effects are explicit.

Exact counts are emitted by the run (and in summary.json); they are not hard-coded here.

Reference oracle

Inputs are exact, so they are represented as fractions.Fraction (Python stdlib). d, q, a, b, c, the discriminant and polynomial residuals are exact rationals; only sqrt uses decimal (--precision, default 160). A custom hex parser handles both %a (double) and %La (long double) without float.fromhex, which would silently drop a 64-bit long-double mantissa to 53 bits. Reference classification mirrors production semantics: c<0 is INSIDE (not an entry failure), c==0 uses the boundary slope, a==0 is the exact linear branch, a real double root/tangent is MISS.

Route references are built from the actual canonical state the probe printed (canonx*, canonw*) joined with the callback samples at t0, so the 1-ulp normalisation of w and both callback models (centre = c0 + v(t-t0) and the production centre = v t) are handled explicitly.

Metrics

Kernel:

  • status counts (ENTRY/MISS/UNCERTAIN) vs the reference, false MISS (reference ENTER, kernel MISS) and false candidate (reference MISS, kernel ENTRY);
  • root error in ULP of the returned double root and relative error, for the common reference-ENTER set and the intersection where all variants returned ENTRY (so more UNCERTAIN cannot look better by selection);
  • polynomial residual at the candidate root using the actual printed coefficients, the exact ideal polynomial, and the nearest-double-rounded reference root as a baseline;
  • a exact-zero vs kernel-zero, a sign mismatches, a collapse/spurious non-zero, and conditioning labels (near_linear, late, ill_conditioned). Outward/growing and near-linear cases are ill-conditioned; a rounded-zero a can turn a very late true entry into a false MISS, so these are reported separately and are not treated as a universal precision claim.

Route:

  • kind/status vs reference, false MISS, false candidate, camera-INSIDE mismatches, total INVALID/ENTRY_UNCONFIRMED outcomes independent of reference classification (a conservative refusal to confirm is not a false escape);
  • the public kernel classification at the actual canonical inputs (kern_status in the route CSV), so a reference ENTER / kernel MISS / route ESCAPED (kernel_false_miss_escaped) is visible and is not confused with canonical normalisation;
  • fast-path vs fallback (entry_fallback_evaluations), F/tol for accepted entries, and Pi/L_camera preservation.

The reference enforces R0 - rr·s > 0 for quadratic roots; a positive root outside the physical radius domain is reported as domain_clipped and is not counted as a physical ENTER/false-MISS. A 160-vs-240-digit consistency check is run for every curated and near-linear/late/domain-clipped id and must report the same status, discriminant sign and nearest-double root.

Timing:

  • kernel: coefficient assembly + entry_solve, serial, noinline + volatile sink, and an inline-asm opaque loop index so GCC cannot hoist the pure kernel out of the repetition loop (identical for all variants). A linearity check runs the kernel at 1× and 2× calls and reports the ratio (expect ≈2) to prove per-call execution.
  • route: public pre-route, OpenMP static schedule + reductions, with route-kind and fallback counts captured so a timing gain cannot come from silently escaping more rays. --threads defaults to 4.

Codegen evidence is collected with objdump/nm: for the double variants entry_solve contains hardware vfmadd and zero x87; for ld_fma it contains x87 plus fmal PLT relocations (extended-precision FMA is a libm software routine on x86-64).

Run

python3 benchmarks/quadratic_precision/run.py \
    --output-dir /tmp/opencode/quadratic-comparison

Useful overrides: --rounds N, --threads T, --precision P, --kernel-target-calls N, --route-target-calls N, --skip-build (reuse the last build), --repo PATH, --cc CC.

Artifacts (all under --output-dir):

generated/{ld_plain,ld_fma,double_fma,double_fma_coeff}.c
cases/quadratic_cases.h, cases/cases.json
environment.json, build_manifest.json
raw/kernel_<v>.csv, raw/route_<v>.csv
raw/kernel_metrics_<v>.csv, raw/route_metrics_<v>.csv
raw/curated_kernel.csv, raw/curated_route.csv
raw/microbench_<v>_r*.json, raw/routebench_<v>_r*.json
raw/fma_verify.json, raw/{linearity,curated}*
summary.json
logs/{build_commands,accuracy_*,microbench_*,routebench_*}.log

summary.json is the machine-readable aggregate; the script also prints the coverage counts, per-variant ULP distributions, the curated case table, codegen evidence and median ns/call.

Interpretation caveats

  • FMA improves the rounding of one product only. It cannot restore the information already lost in the double inputs and in the coefficient sums (d = x - c, q = w + v, and the qq/dq/dd accumulations). The data should be read stage by stage.
  • A kernel UNCERTAIN is not a failure: the shared route fallback may confirm the entry. A kernel MISS bypasses the fallback entirely, so growing/outward false-MISS families are exercised through dedicated rr<0 route fixtures and reported both as the raw kernel decision (kern_status) and the route outcome. This benchmark does not claim that the fallback rescues a raw kernel MISS on the same case; it only observes that more fallback is used to confirm inaccurate candidates.
  • The double uncertainty band differs from the long-double one by the precision term: production uses 64·(DBL_EPSILON + LDBL_EPSILON)·scale while the double variants use 64·(DBL_EPSILON + DBL_EPSILON)·scale; since LDBL_EPSILON ≪ DBL_EPSILON this is roughly a factor of two in the band width, which only matters exactly at the UNCERTAIN/MISS border.
  • Number/scale results are machine- and compiler-specific; environment.json records the source hash, compiler, flags and platform.
  • This benchmark does not establish a universal guarantee, only the sampled behaviour on the frozen fixture set.