Generate plain long-double and experimental FMA variants from the production entry kernel. Compare exact-rational references, geometric validation, fallback counts and repeated timings with self-contained fixtures. Validate cached build dependencies and reject stale timing output. Count all unconfirmed entry outcomes independently of reference classification.
233 lines
11 KiB
Markdown
233 lines
11 KiB
Markdown
# quadratic_precision — stable entry quadratic: precision vs. cost
|
||
|
||
Self-contained, reproducible benchmark comparing four arithmetic realisations of
|
||
the **same** asymptotic-entry algebra. Nothing here modifies production code,
|
||
the `Makefile`, or git state.
|
||
|
||
## Question
|
||
|
||
The production camera pre-route solves the relative-distance quadratic
|
||
|
||
```
|
||
F(s) = |d + q s|^2 - (R0 - rr s)^2 = a s^2 + b s + c
|
||
d = x_cur - c_frame, q = w_frame + v_frame
|
||
```
|
||
|
||
in `long double`, with a stable root formula, an ordinary discriminant
|
||
`b*b - 4*a*c` and entry slope, a common power-of-two scaling, an
|
||
uncertainty band, geometric validation and a numerical fallback. Which matters
|
||
for accuracy and which for speed?
|
||
|
||
Four variants share the *entire* rest of the module (scaling, uncertainty band,
|
||
validation, fallback, dispatch) and are generated from the current
|
||
`src/asymptotic.c` by rewriting exactly one lexical region:
|
||
|
||
| variant | coefficients | type | discriminant / slope |
|
||
|---------------------|--------------|-------------|----------------------|
|
||
| `ld_plain` | as production| `long double` | plain `b*b - 4*a*c`, plain slope |
|
||
| `ld_fma` | as production| `long double` | compensated `fmal` (experimental arm) |
|
||
| `double_fma` | ordinary `+=`| `double` | compensated `fma` |
|
||
| `double_fma_coeff` | `fma` dots, `fma` products | `double` | compensated `fma` |
|
||
|
||
`double_fma` is the control that isolates *type* precision from *FMA*; the two
|
||
double variants isolate *coefficient accumulation*.
|
||
|
||
## Generation and build
|
||
|
||
`build.py` locates the kernel region between two lexical anchors in the current
|
||
working-tree `src/asymptotic.c`:
|
||
|
||
* start: the comment `/* Long-double coefficients of the relative-distance
|
||
quadratic`
|
||
* end: the forward declaration `static void minkowski_route_entry(...)`
|
||
|
||
It asserts each anchor and the presence of `entry_quadratic_coeffs`,
|
||
`entry_solve`, `entry_discriminant` exactly once, then emits four full-module
|
||
copies under `<output>/generated/`. Every other line (dispatch, fallback,
|
||
`asymptotic_route_camera`, Schwarzschild path, …) is copied verbatim. If a
|
||
future edit moves an anchor or changes the region, generation **fails** rather
|
||
than silently measuring the wrong code.
|
||
|
||
`ld_plain` retains the production kernel verbatim; `ld_fma` rewrites only the
|
||
discriminant body and the entry slope. Double
|
||
variants replace `long double`→`double`, `fmal`/`fabsl`/`fmaxl`/`frexpl`/
|
||
`scalbnl`/`sqrtl`/`copysignl`→their double forms, `LDBL_*`→`DBL_*`, and strip
|
||
`L` suffixes from literals. A guard rejects any remaining `long double`
|
||
promotion, `L` literal or `*l`/`fmal` call in the double kernel section, so no
|
||
double expression can be re-promoted to `long double`. `double_fma_coeff`
|
||
additionally changes the coefficient accumulation to `fma`.
|
||
|
||
The probe `#include`s the generated module, so the private static
|
||
`entry_quadratic_coeffs`/`entry_solve` and the public `asymptotic_route_camera`
|
||
that are exercised are the **actual generated code**, not an independent copy.
|
||
|
||
Compile flags, identical for all four variants:
|
||
|
||
```
|
||
-std=c11 -march=native -O2 -DNDEBUG -ffp-contract=off -fopenmp
|
||
```
|
||
|
||
`-ffp-contract=off` prevents implicit FMA contraction from silently changing
|
||
the `ld_plain` arm; explicit `fma`/`fmal` still lower as written.
|
||
|
||
The probe links a minimal production subset (`geodesic.c`,
|
||
`asymptotic_entry.c`, `asymptotic_schwarzschild.c`, `spacetime_common.c`,
|
||
`spacetime_minkowski.c`) plus libm. No FFTW, PNG, catalog or observer-track
|
||
data are needed.
|
||
|
||
## Inputs
|
||
|
||
All fixtures are frozen as IEEE-754 hex literals (`cases.json` and the
|
||
generated C header are bit-identical), including the two production grazing
|
||
rows reproduced from the Alcubierre pre-route (`x=(0,-24,0)`, `centre=2t`,
|
||
`v=2`, `R=5`, canonical `w` hex values). Families:
|
||
|
||
* **curated / adversarial**: head-on hit/miss, clear miss, grazing
|
||
`y=nextafter(R,0)`, exact tangent, near-boundary, `1e10 + 0.5, R=1`
|
||
cancellation, large-`t` `centre=0.1t`, exact linear `a==0` (inward and
|
||
initially outward), translated-origin cancellation.
|
||
* **fixed**: axes and three oblique frames, radii
|
||
`1e-100…1e100`, `D/R ∈ {1+ulp,2,10,100,512,1024,1e4,1e8,1e10}`, impact
|
||
ratios `{0,.5,.99,1-1e-6,nextafter(1,0),1,nextafter(1,∞),1+1e-6,1.1}`.
|
||
The `512`/`1024` ratios sample the crossover where long-double coefficient
|
||
rounding meets the `128 ε_D R²` geometry tolerance (`≈512`) and the double
|
||
crossover (`≈11`), separating fallback rates from the all-UNCERTAIN `1e10`
|
||
regime.
|
||
* **moving**: `v ∈ {0,.1,2,10}` along a transverse direction.
|
||
* **growing**: `rr = -1` and `nextafter(-1,∓∞)`, **inward and
|
||
initially-outward** photons (`w = ∓tow`), axis and oblique.
|
||
* **shrinking**: `rr>0`, labelled/domain-checked algebra (radius would go
|
||
negative on the open past; roots outside `R0 - rr·s > 0` are non-physical).
|
||
* **route**: `rr = 0` plus **`rr<0` growing inward/outward** families
|
||
(positive radius on the whole past, `valid_t_min=-1e300`) across all three
|
||
direction frames, driven through the public `asymptotic_route_camera`. The
|
||
oblique-2 growing-outward family reproduces the double kernel false-MISS
|
||
cases, where a kernel MISS bypasses the fallback and the route escapes
|
||
directly; the route reference uses the actual normalised canonical `w`, so
|
||
normalisation effects are explicit.
|
||
|
||
Exact counts are emitted by the run (and in `summary.json`); they are not
|
||
hard-coded here.
|
||
|
||
## Reference oracle
|
||
|
||
Inputs are exact, so they are represented as `fractions.Fraction` (Python
|
||
stdlib). `d`, `q`, `a`, `b`, `c`, the discriminant and polynomial residuals
|
||
are **exact rationals**; only `sqrt` uses `decimal` (`--precision`, default
|
||
160). A custom hex parser handles both `%a` (double) and `%La` (long double)
|
||
without `float.fromhex`, which would silently drop a 64-bit long-double mantissa
|
||
to 53 bits. Reference classification mirrors production semantics: `c<0` is
|
||
INSIDE (not an entry failure), `c==0` uses the boundary slope, `a==0` is the
|
||
exact linear branch, a real double root/tangent is MISS.
|
||
|
||
Route references are built from the **actual canonical state the probe printed**
|
||
(`canonx*`, `canonw*`) joined with the callback samples at `t0`, so the 1-ulp
|
||
normalisation of `w` and both callback models (`centre = c0 + v(t-t0)` and the
|
||
production `centre = v t`) are handled explicitly.
|
||
|
||
## Metrics
|
||
|
||
Kernel:
|
||
|
||
* status counts (ENTRY/MISS/UNCERTAIN) vs the reference, **false MISS**
|
||
(reference ENTER, kernel MISS) and **false candidate** (reference MISS,
|
||
kernel ENTRY);
|
||
* root error in ULP of the returned double root and relative error, for the
|
||
common reference-ENTER set and the intersection where *all* variants returned
|
||
ENTRY (so more UNCERTAIN cannot look better by selection);
|
||
* polynomial residual at the candidate root using the actual printed
|
||
coefficients, the exact ideal polynomial, and the nearest-double-rounded
|
||
reference root as a baseline;
|
||
* `a` exact-zero vs kernel-zero, `a` sign mismatches, `a` collapse/spurious
|
||
non-zero, and conditioning labels (`near_linear`, `late`, `ill_conditioned`).
|
||
Outward/growing and near-linear cases are ill-conditioned; a rounded-zero `a`
|
||
can turn a very late true entry into a false MISS, so these are reported
|
||
separately and are not treated as a universal precision claim.
|
||
|
||
Route:
|
||
|
||
* kind/status vs reference, false MISS, false candidate, camera-INSIDE
|
||
mismatches, total `INVALID/ENTRY_UNCONFIRMED` outcomes independent of reference
|
||
classification (a conservative refusal to confirm is not a false escape);
|
||
* the **public kernel classification at the actual canonical inputs**
|
||
(`kern_status` in the route CSV), so a reference ENTER / kernel MISS / route
|
||
ESCAPED (`kernel_false_miss_escaped`) is visible and is not confused with
|
||
canonical normalisation;
|
||
* fast-path vs fallback (`entry_fallback_evaluations`), `F/tol` for accepted
|
||
entries, and Pi/`L_camera` preservation.
|
||
|
||
The reference enforces `R0 - rr·s > 0` for quadratic roots; a positive root
|
||
outside the physical radius domain is reported as `domain_clipped` and is not
|
||
counted as a physical ENTER/false-MISS. A 160-vs-240-digit consistency check is
|
||
run for every curated and near-linear/late/domain-clipped id and must report the
|
||
same status, discriminant sign and nearest-double root.
|
||
|
||
Timing:
|
||
|
||
* kernel: coefficient assembly + `entry_solve`, serial, `noinline` +
|
||
`volatile` sink, and an inline-asm opaque loop index so GCC cannot hoist the
|
||
pure kernel out of the repetition loop (identical for all variants). A
|
||
linearity check runs the kernel at 1× and 2× calls and reports the ratio
|
||
(expect ≈2) to prove per-call execution.
|
||
* route: public pre-route, OpenMP `static` schedule + reductions, with
|
||
route-kind and fallback counts captured so a timing gain cannot come from
|
||
silently escaping more rays. `--threads` defaults to 4.
|
||
|
||
Codegen evidence is collected with `objdump`/`nm`: for the double variants
|
||
`entry_solve` contains hardware `vfmadd` and zero x87; for `ld_fma` it contains
|
||
x87 plus `fmal` PLT relocations (extended-precision FMA is a libm software
|
||
routine on x86-64).
|
||
|
||
## Run
|
||
|
||
```sh
|
||
python3 benchmarks/quadratic_precision/run.py \
|
||
--output-dir /tmp/opencode/quadratic-comparison
|
||
```
|
||
|
||
Useful overrides: `--rounds N`, `--threads T`, `--precision P`,
|
||
`--kernel-target-calls N`, `--route-target-calls N`, `--skip-build` (reuse the
|
||
last build), `--repo PATH`, `--cc CC`.
|
||
|
||
Artifacts (all under `--output-dir`):
|
||
|
||
```
|
||
generated/{ld_plain,ld_fma,double_fma,double_fma_coeff}.c
|
||
cases/quadratic_cases.h, cases/cases.json
|
||
environment.json, build_manifest.json
|
||
raw/kernel_<v>.csv, raw/route_<v>.csv
|
||
raw/kernel_metrics_<v>.csv, raw/route_metrics_<v>.csv
|
||
raw/curated_kernel.csv, raw/curated_route.csv
|
||
raw/microbench_<v>_r*.json, raw/routebench_<v>_r*.json
|
||
raw/fma_verify.json, raw/{linearity,curated}*
|
||
summary.json
|
||
logs/{build_commands,accuracy_*,microbench_*,routebench_*}.log
|
||
```
|
||
|
||
`summary.json` is the machine-readable aggregate; the script also prints the
|
||
coverage counts, per-variant ULP distributions, the curated case table, codegen
|
||
evidence and median ns/call.
|
||
|
||
## Interpretation caveats
|
||
|
||
* FMA improves the rounding of one product only. It **cannot** restore the
|
||
information already lost in the double inputs and in the coefficient sums
|
||
(`d = x - c`, `q = w + v`, and the `qq`/`dq`/`dd` accumulations). The data
|
||
should be read stage by stage.
|
||
* A kernel UNCERTAIN is not a failure: the shared route fallback may confirm the
|
||
entry. A kernel **MISS bypasses the fallback entirely**, so growing/outward
|
||
false-MISS families are exercised through dedicated `rr<0` route fixtures and
|
||
reported both as the raw kernel decision (`kern_status`) and the route
|
||
outcome. This benchmark does **not** claim that the fallback rescues a raw
|
||
kernel MISS on the same case; it only observes that more fallback is used to
|
||
confirm inaccurate candidates.
|
||
* The double uncertainty band differs from the long-double one by the precision
|
||
term: production uses `64·(DBL_EPSILON + LDBL_EPSILON)·scale` while the
|
||
double variants use `64·(DBL_EPSILON + DBL_EPSILON)·scale`; since
|
||
`LDBL_EPSILON ≪ DBL_EPSILON` this is roughly a factor of two in the band
|
||
width, which only matters exactly at the UNCERTAIN/MISS border.
|
||
* Number/scale results are machine- and compiler-specific; `environment.json`
|
||
records the source hash, compiler, flags and platform.
|
||
* This benchmark does not establish a universal guarantee, only the sampled
|
||
behaviour on the frozen fixture set.
|