Commit Graph
21 Commits
Author SHA1 Message Date
wyj 7c980c35aa Benchmark: Compare quadratic entry precision and arithmetic cost
Generate plain long-double and experimental FMA variants from the production entry kernel. Compare exact-rational references, geometric validation, fallback counts and repeated timings with self-contained fixtures.

Validate cached build dependencies and reject stale timing output. Count all unconfirmed entry outcomes independently of reference classification.
2026-10-09 00:15:02 -04:00
wyj 0a46a7095b Feat: Complete adaptive geodesic tracing with DP54
Add error-controlled DP5(4) integration and trusted first-crossing localization, including non-monotonic energy thresholds and representable-time stepping.

Preserve adaptive state and independent step/time retry grants across RayPool, refinement and movie scheduling. Expose numerical controls, record actual persistent-sample costs, and add v3 lens-map provenance with legacy v2 RK4 import.

Use DP54 by default and select an 8M Schwarzschild maximum step from bounded scans and a two-run 4K comparison. Retain the conservative minimum-step guard and document critical-ray and backend capability limits. Archive self-contained benchmark inputs and raw output; keep fixed RK4 HDR references explicit.

Validation: make -B -j4 BUILD_TYPE=Debug test passed; explicit RK4 HDR references have zero differences. Bounded convergence checks, benchmark reproduction, Release build and focused reviews passed. No numerical-relativity backend is added.
2026-10-05 20:27:42 -04:00
wyj 4e34780fa9 Doc: Record the isolated sensor-bloom benchmark
Store the bounded 512x288 synthetic-frame benchmark for the optional
sensor bloom model: the tested commit and file hashes, the exact build
and run commands, and the raw OMP_NUM_THREADS=1/4/16 outputs. It does not
claim any full-resolution or production performance.
2026-09-27 04:39:57 -04:00
wyj e4d093a9fe Doc: Record 4K fast-mode accuracy limits and error distribution
Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README.

- Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2.
- Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile.
- Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
2026-09-26 01:02:41 -04:00
wyj 229f50cd86 Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.

- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
2026-09-25 23:35:45 -04:00
wyj 3deebfb2fa Feat: Add --fast-mode supersampled point-source accumulation
Add an optional CPU preview path that deposits each point-source image as a
supersampled delta and resolves the whole frame with one global Moffat
convolution plus an N x N box average, instead of splatting a per-event PSF.

- optics: FastPsfAccumulator builds the pixel-area-integral kernel of the
  target Moffat at the supersampled scale (width N*alpha, same beta).  The
  1/N^2 box average then reproduces the final pixel-area integral, so the
  requested FWHM and beta are preserved without renormalisation.  Deposits are
  per-cell atomic adds; resolve accumulates into the caller's HDR buffer.
- frame: fast branch in frame_splat_catalog with one shared supersampled
  buffer and a single resolve per frame; the accumulator is reused across
  movie frames and built from the map dimensions on lens-map import.
- main: --fast-mode, --fast-supersample N (1..8, default 2) and
  --fast-deposit nearest|bilinear (default nearest).  CPU-only and rejected in
  the HIP/dummy backends; --psf-min-y still applies per event while
  --max-cache-psf-flux does not.
- The deposition scheme was chosen by scripts/fast_mode_deposit_error.py:
  nearest keeps the PSF shape exactly with <= 0.5/N px position quantization;
  bilinear keeps the exact centroid but broadens FWHM and beta.  Recorded in
  benchmarks/fast_mode_deposit_2026-09-18.md.
- tests/test_frame.c covers fast nearest vs the direct evaluator at the snapped
  centre, flux conservation, bilinear centroid, min-Y discard, frame plumbing,
  and HDR accumulation onto a non-zero background.
- benchmarks/fast_mode_cpu_2026-09-18.md records a ~10x speedup on the 2MASS
  galactic-centre field with small tone-mapped differences.
2026-09-24 01:36:25 -04:00
wyj 7ddc58515a Doc: Record full-day parallel prebinning results and next steps
Document the producer-private adaptive path, the dummy support workload
diagnostics, the negative result of the old serial binning and the
2026-09-13 full-day comparison (splat 282.944 s vs atomic 481.417 s).
Add the corresponding raw logs, hotspot fixtures and runner evidence,
and lay out the staged plan toward a sub-240 s full-day render without
changing the default atomic path.
2026-09-14 19:16:00 -04:00
wyj a88ec56097 Fix: Honor OMP_DYNAMIC and env binary hash in replay runner
Pin OMP_DYNAMIC=FALSE so thread counts stay deterministic, record it in
the command line, and locate the binary hash correctly when the command
is wrapped in env with inline assignments.
2026-09-14 19:15:56 -04:00
wyj 2fc2b43e8d Doc: Record production PSF chunk geometry
Preserve the protected full-catalog dummy run, its initial harness failure, input and output hashes, distribution quantiles, and the resulting sparse-tile scheduling conclusions.
2026-09-11 23:28:02 -04:00
wyj d02ecb2181 Feat: Add dummy PSF chunk diagnostics
Reuse production catalog mapping and per-worker chunk boundaries to report triangle and spatial distributions without initializing HIP or writing image outputs. Include a protected full-catalog runner and build documentation.
2026-09-11 23:27:42 -04:00
wyj 776f9b7247 Test: Add adaptive HIP tile replay selection
Add a bounded per-chunk center-tile selector to the real-event replay prototype and retain the paired timing, mixed-boundary, CPU, HIP, and sandbox-failure evidence.
2026-09-11 23:27:24 -04:00
wyj 11e703e33d Doc: record bounded HIP PSF replay results and provenance
Archive fixed catalog subsets and captured events, manifests, reproducible commands, raw timing and regression logs, and device resource evidence. Document distribution-dependent tile reduction gains and independent indexed producer throughput.

Update the GPU plan and renderer design while retaining the production atomic path. Preserve historical measurement metadata and exclude generated Python bytecode.
2026-09-11 16:45:48 -04:00
wyj ab7d879f6c Doc: record full-sky CPU and HIP benchmark comparison
Preserve user-provided complete CPU and HIP logs and associate the optimized implementation with e1ec480669. Record matching event counts, user-confirmed bit-identical PNG output, and the 2.77x workflow speedup with tracing/import differences noted.

Mark the pre-artifact-fix CPU benchmark as historical and link the new paired result.
2026-09-06 21:40:51 -04:00
wyj e1ec480669 HIP: accelerate PSF accumulation and restore parallel producers
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.

Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
2026-09-06 21:38:53 -04:00
wyj 8741597d9d Ray: Balance movie tracing with dynamic chunks
Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations.

Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
2026-09-06 15:56:37 -04:00
wyj 94149d75e4 Doc: reorganize bilingual READMEs and build guide with reference renders 2026-09-06 02:03:50 -04:00
wyj a18ebfbf1a HIP: record direct atomic PSF timings 2026-09-05 22:13:14 -04:00
wyj 92d8d294b5 Bench: record PSF event sink overhead 2026-09-05 03:20:40 -04:00
wyj 4953775cdd Test: add fixed PSF image references 2026-09-04 20:28:54 -04:00
wyj 80b891d2c8 Optics: cache pixel-area Moffat kernels 2026-08-28 13:53:06 -04:00
wyj 172738d932 Doc: record dense 2MASS tile benchmark 2026-08-27 19:29:55 -04:00