Commit Graph
83 Commits
Author SHA1 Message Date
wyj 9cd933d1f8 Feat: Add optional three-channel sensor bloom model
Add an opt-in, post-processing limited-response model applied to the
finished linear HDR before tone mapping. Each RGB channel is processed
independently and isotropically: overflow above E spreads to the eight
neighbours with a fixed 9-point stencil, while the rest is absorbed or
lost at the image boundary. The synchronous ping-pong update uses a
monotonic bounding box and a row-parallel, deterministic reduction; the
conservative round bound reserves fp guard rounds inside a 4096 hard
limit and fails before touching HDR when exceeded.

Expose --sensor-bloom-limit E and --sensor-bloom-transfer e (both
required together, default disabled), validate them before expensive
initialization, and route every output path through the same hook in
write_frame_outputs: raw FITS first, bloom, tone-mapped PNG/PPM, then the
mesh overlay. The raw --hdr-output FITS therefore stays pre-bloom.

Add a standalone unit test (stencil, boundary loss, cascade reference,
symmetry, thread determinism, convergence limits, validation, allocation
failure), CLI integration and regression coverage, an isolated
sensor-bloom-bench target, and document the model in the design, usage,
README, and build docs.
2026-09-27 04:39:38 -04:00
wyj 85fce1bdeb Misc: add a test catalog for testing PSF 2026-09-27 02:48:52 -04:00
wyj e31e1ed27e Feat: Add soft-clip tone mapping with legacy Reinhard
Replace the per-channel Reinhard display transform with a parameterized
soft clip T_p(x) = tanh(x^p)^(1/p), default softclip p=2, exposed through
--tone-map and --tone-map-p. Keep --tone-map reinhard bit-compatible with
the previous x/(1+x) curve for existing images and reject combining it
with an explicit --tone-map-p.

Route the primary image and the mesh overlay through the same
ToneMapSettings; the linear HDR FITS writer stays pre-tone-map. Add a
focused optics-linked tone-map test target, CLI success/error coverage,
and document the display operator versus the Moffat effective PSF.
2026-09-27 00:44:19 -04:00
wyj 7399eb6904 Tests: Use image-neutral wording in camera CLI logs
The camera CLI test now runs for both PNG and PPM builds, so the single/movie
comparison messages should say "image" rather than "PNG".
2026-09-26 17:02:31 -04:00
wyj 2a0197229a Feat: Write mesh overlay as a separate <stem>_mesh image
--draw-mesh no longer modifies the primary render.  The clean HDR FITS and
tone-mapped image are written first, then the final image-plane mesh is
overlaid in place on the already-consumed HDR buffer and saved as a
<stem>_mesh sibling.  This applies uniformly to single-frame, movie-frame,
and imported lens-map renders, with derived paths validated before catalog,
spacetime, and render initialization.

Dummy PSF backends still write nothing and report ignoring --draw-mesh.  The
reference-image target now uses the configured image extension so the
ENABLE_PNG=0 test path passes, and the CLI tests cover clean-main identity,
sibling naming, PATH_MAX overflow, and zero-tolerance HDR comparison.

README, README.zh-CN, and usage.md describe the new naming; the mesh reference
asset is renamed to match.
2026-09-26 17:01:25 -04:00
wyj e4d093a9fe Doc: Record 4K fast-mode accuracy limits and error distribution
Add the production-density 4K Galactic-center comparison to the 2026-09-25 benchmark record, including the per-image error decomposition, and refine the fast-mode positioning in usage.md, the design document, and README.

- Document that nearest/bilinear deposit accuracy is set by the deposit mode and N rather than output resolution, and that the still-frame error is dominated by a small number of resolved bright cores while the flux-dominant faint texture is much better than the global relative L2.
- Add the temporal-coherence caveat: nearest deposition jumps by 1/N output pixel per axis across supersampled-cell boundaries, so fast mode is not qualified as a temporally coherent movie path; bilinear removes the centroid jump but broadens the profile.
- Add scripts/fast_mode_error_decomposition.py, which ranks the largest |fast-base| RGB samples and partitions the error by a disjoint base-value range.
2026-09-26 01:02:41 -04:00
wyj 229f50cd86 Feat: Use FFTW linear convolution for CPU fast-mode PSF resolve
Replace the nested spatial global convolution in fast_psf_accumulator_resolve with a reusable double-precision FFTW linear convolution on the CPU PSF backend.

- Add the private src/fast_psf_fftw.{c,h} module: zero-padded R2C/C2R plans, cached kernel spectrum and planar scratch, exact 1/(Pwidth*Pheight) and 1/N^2 normalization, (R,R) crop, and additive HDR output.
- Keep the previous nested loops as fast_psf_accumulator_resolve_spatial_reference for tests/benchmarks only; it is not a runtime fallback.
- Cache the circular row spans on FastPsfAccumulator and report one-time plan, kernel transform, scratch, and per-frame stage timings.
- Require fftw3_omp for CPU builds; HIP and dummy builds do not link FFTW.
- Namespace test/helper binaries by spacetime and build tag, and reject make test / psf-capture for non-CPU backends.
- Add tests/test_fast_psf_fftw.c (FFTW versus spatial), tests/benchmark_fast_psf_fftw.c, an FFTW CLI smoke check, and the 2026-09-25 benchmark record.
2026-09-25 23:35:45 -04:00
wyj 3deebfb2fa Feat: Add --fast-mode supersampled point-source accumulation
Add an optional CPU preview path that deposits each point-source image as a
supersampled delta and resolves the whole frame with one global Moffat
convolution plus an N x N box average, instead of splatting a per-event PSF.

- optics: FastPsfAccumulator builds the pixel-area-integral kernel of the
  target Moffat at the supersampled scale (width N*alpha, same beta).  The
  1/N^2 box average then reproduces the final pixel-area integral, so the
  requested FWHM and beta are preserved without renormalisation.  Deposits are
  per-cell atomic adds; resolve accumulates into the caller's HDR buffer.
- frame: fast branch in frame_splat_catalog with one shared supersampled
  buffer and a single resolve per frame; the accumulator is reused across
  movie frames and built from the map dimensions on lens-map import.
- main: --fast-mode, --fast-supersample N (1..8, default 2) and
  --fast-deposit nearest|bilinear (default nearest).  CPU-only and rejected in
  the HIP/dummy backends; --psf-min-y still applies per event while
  --max-cache-psf-flux does not.
- The deposition scheme was chosen by scripts/fast_mode_deposit_error.py:
  nearest keeps the PSF shape exactly with <= 0.5/N px position quantization;
  bilinear keeps the exact centroid but broadens FWHM and beta.  Recorded in
  benchmarks/fast_mode_deposit_2026-09-18.md.
- tests/test_frame.c covers fast nearest vs the direct evaluator at the snapped
  centre, flux conservation, bilinear centroid, min-Y discard, frame plumbing,
  and HDR accumulation onto a non-zero background.
- benchmarks/fast_mode_cpu_2026-09-18.md records a ~10x speedup on the 2MASS
  galactic-centre field with small tone-mapped differences.
2026-09-24 01:36:25 -04:00
wyj 7ddc58515a Doc: Record full-day parallel prebinning results and next steps
Document the producer-private adaptive path, the dummy support workload
diagnostics, the negative result of the old serial binning and the
2026-09-13 full-day comparison (splat 282.944 s vs atomic 481.417 s).
Add the corresponding raw logs, hotspot fixtures and runner evidence,
and lay out the staged plan toward a sub-240 s full-day render without
changing the default atomic path.
2026-09-14 19:16:00 -04:00
wyj a88ec56097 Fix: Honor OMP_DYNAMIC and env binary hash in replay runner
Pin OMP_DYNAMIC=FALSE so thread counts stay deterministic, record it in
the command line, and locate the binary hash correctly when the command
is wrapped in env with inline assignments.
2026-09-14 19:15:56 -04:00
wyj ce691b11ad Test: Extend HIP PSF replay for production and sustained rounds
Add production, production-prepared, production-parallel and
production-busy modes plus an optional 1..512 round count. Repeated
rounds validate against a round-scaled CPU double reference and report
per-round throughput; production-parallel exercises concurrent
prepared-chunk submission and production-busy adds host spin workers to
probe contention.
2026-09-14 19:15:55 -04:00
wyj 4bf2de1b40 Feat: Add dummy PSF support workload diagnostics
Add an opt-in DUMMY_PSF_SUPPORT_STATS mode to the diagnostic-only dummy
consumer. It records per-chunk cache-limited 16px support references,
touched tiles, 256-reference tasks, maximum tile list and first-pass
counting time, then reports percentiles. No HDR is produced and no
references are materialized or uploaded.
2026-09-14 19:15:53 -04:00
wyj eec5dd215a Feat: Add producer-private adaptive HIP PSF prebinning
Move adaptive tile16 selection and support-expanded CPU binning into a
per-producer HipPsfPreparedChunk computed before acquiring the shared
sink lock. The sink keeps sole ownership of the GPU HDR, stream, device
scratch and upload staging; submit_prepared runs under the lock and
preserves the existing single-stream ordering and direct fallback path.
The legacy hip_psf_sink_submit entry point prepares an internal chunk so
single-threaded replay stays compatible, and default accumulation remains
atomic.
2026-09-14 19:15:51 -04:00
wyj 2d01c2c101 Doc: Adopt CSV CIE 1931 blackbody reference 2026-09-12 15:29:36 -04:00
wyj ce9c4bc16f Test: Refresh reference images for CSV blackbody reference 2026-09-12 15:29:36 -04:00
wyj d2a901cab1 Feat: Add CSV CIE 1931 blackbody LUT backend
Replace the in-tree Wyman analytic CIE fit with a repository CIE 1931
2-degree LUT derived from the 360-830 nm, 1 nm-linear CSV reference. The
GRBBLUT3 table stores 1024 uniform log(T) XYZ nodes over
[670.146556, 101408.88] K with four-point cubic Lagrange interpolation, and
the same CSV reference supplies a three-term Rayleigh-Jeans form above the
table and an inverse-temperature/log-XYZ crossover with channel-specific
endpoint forms below it. The loader validates the header and payload
checksum and rejects legacy formats and runtime generation.

Initialize the backend in the renderer, frame test, and PSF capture tool so
the new default is active wherever colors are produced.
2026-09-12 15:29:33 -04:00
wyj 2fc2b43e8d Doc: Record production PSF chunk geometry
Preserve the protected full-catalog dummy run, its initial harness failure, input and output hashes, distribution quantiles, and the resulting sparse-tile scheduling conclusions.
2026-09-11 23:28:02 -04:00
wyj d02ecb2181 Feat: Add dummy PSF chunk diagnostics
Reuse production catalog mapping and per-worker chunk boundaries to report triangle and spatial distributions without initializing HIP or writing image outputs. Include a protected full-catalog runner and build documentation.
2026-09-11 23:27:42 -04:00
wyj 776f9b7247 Test: Add adaptive HIP tile replay selection
Add a bounded per-chunk center-tile selector to the real-event replay prototype and retain the paired timing, mixed-boundary, CPU, HIP, and sandbox-failure evidence.
2026-09-11 23:27:24 -04:00
wyj 11e703e33d Doc: record bounded HIP PSF replay results and provenance
Archive fixed catalog subsets and captured events, manifests, reproducible commands, raw timing and regression logs, and device resource evidence. Document distribution-dependent tile reduction gains and independent indexed producer throughput.

Update the GPU plan and renderer design while retaining the production atomic path. Preserve historical measurement metadata and exclude generated Python bytecode.
2026-09-11 16:45:48 -04:00
wyj 7f5ec1a488 Optics: add bounded HIP PSF replay and tile reduction experiments
Capture prepared events through a test-only producer consumer and measure serial and indexed parallel generation independently. Compare the production atomic kernel with bounded pixel-owned tile reductions while preserving double HDR, cache support, and direct fallback ordering.

Validate captured inputs, edge and multichunk fixtures, CPU/HIP renderer regressions, and 45 isolated final replays. Keep the production HIP algorithm unchanged because gains depend on event distribution.
2026-09-11 16:43:53 -04:00
wyj 5eccc31c99 Doc: refresh Galactic-center reference image and benchmark links
Replace the artifact-affected Schwarzschild image with the corrected 4K render. Link both READMEs to the paired CPU/HIP benchmark and clarify that this run had no PSF wing clipping or direct fallbacks.
2026-09-06 22:02:39 -04:00
wyj ab7d879f6c Doc: record full-sky CPU and HIP benchmark comparison
Preserve user-provided complete CPU and HIP logs and associate the optimized implementation with e1ec480669. Record matching event counts, user-confirmed bit-identical PNG output, and the 2.77x workflow speedup with tracing/import differences noted.

Mark the pre-artifact-fix CPU benchmark as historical and link the new paired result.
2026-09-06 21:40:51 -04:00
wyj e1ec480669 HIP: accelerate PSF accumulation and restore parallel producers
Cooperate across 32 lanes per PSF and use two completion-protected staging slots with complete batch timing. Restore coarse OpenMP event production while serializing shared GPU submissions and direct fallback boundaries.

Add bounded benchmarks, streaming and renderer regressions, and preserve validation evidence and ownership documentation.
2026-09-06 21:38:53 -04:00
wyj 265d7b95d5 Fix: reject invalid inverse lens-map weights
Check spherical area weight sums against a fixed 1e-8 tolerance and normalize accepted weights before interpolation. Skip invalid images in both catalog paths.

Add a thin-triangle regression for both parities, document the measured tolerance margin, and refresh HDR fixtures with strict comparisons restored.
2026-09-06 17:47:09 -04:00
wyj 8741597d9d Ray: Balance movie tracing with dynamic chunks
Schedule ray-pool advancement in dynamic chunks of 32 slots to distribute frame-grouped active rays across workers while preserving stable endpoint destinations.

Add 1/4/16-thread lens-map equivalence coverage and document schedule comparisons with complete commands, input data and original logs. The 393-frame 64x36 production-source benchmark improves from 62.13 to 38.69 seconds with byte-identical PNGs and lens maps. Release and Debug test suites pass.
2026-09-06 15:56:37 -04:00
wyj 8cf6106127 Fix: Skip empty movie refinement generations
Allow frames to finish refinement independently without aborting the movie. Report genuine refinement errors with generation and frame identifiers, and clarify the sample batch contract.

Add a two-frame Schwarzschild regression covering unequal refinement progress. Validated Release and Debug make test suites and 640x360 rendering of the final eight free-fall frames in both builds.
2026-09-06 08:10:29 -04:00
wyj b69c9cfd45 Feat: generate freely falling Schwarzschild camera tracks
Integrate timelike geodesics and Fermi-Walker tetrads in ingoing Kerr-Schild coordinates, sampled at a configurable proper-time cadence.

Add movie-track-samples to preserve CSV events as frames. Document usage and singularity guards, and cover analytic orbits, transport convergence, sampling, and CSV rendering.
2026-09-06 05:56:51 -04:00
wyj 8d011b7ec5 Doc: require English commit messages 2026-09-06 04:01:38 -04:00
wyj 6ae223a642 Feat: support general single-frame cameras and add a near-horizon example 2026-09-06 04:00:21 -04:00
wyj 94149d75e4 Doc: reorganize bilingual READMEs and build guide with reference renders 2026-09-06 02:03:50 -04:00
wyj a18ebfbf1a HIP: record direct atomic PSF timings 2026-09-05 22:13:14 -04:00
wyj 7a80086f63 Feat: add initial HIP PSF backend 2026-09-05 04:23:57 -04:00
wyj e48c490337 Doc: update GPU PSF acceleration status 2026-09-05 03:47:42 -04:00
wyj e8bc553128 DOC: preserve complete benchmark output 2026-09-05 03:22:15 -04:00
wyj 92d8d294b5 Bench: record PSF event sink overhead 2026-09-05 03:20:40 -04:00
wyj 2a03cbf724 Frame: buffer cached PSF events per worker 2026-09-05 03:06:46 -04:00
wyj 4953775cdd Test: add fixed PSF image references 2026-09-04 20:28:54 -04:00
wyj ab56e23fc9 Doc: document reusable lens maps 2026-09-03 20:58:20 -04:00
wyj d43ad262fc CLI: document lens map options 2026-09-03 20:55:08 -04:00
wyj 1ee3a8cc26 Feat: add reusable lens map files 2026-09-03 19:46:05 -04:00
wyj 1a54e59eb4 Output: add CLI help text 2026-09-03 19:24:32 -04:00
wyj dd8cf29dac Doc: add GPU PSF acceleration plan 2026-08-31 15:26:47 -04:00
wyj b2ac642de7 Fix: correct ICRS camera orientation 2026-08-31 00:13:37 -04:00
wyj e8deb1ff19 Optics: add PSF luminance cutoff 2026-08-30 15:10:14 -04:00
wyj 9e23c3fcdb Optics: iterate circular PSF rows directly 2026-08-30 15:06:53 -04:00
wyj 8662064044 Optics: parameterize PSF relative tail 2026-08-30 15:05:40 -04:00
wyj ff2ba43333 Output: simplify HDR output naming 2026-08-30 02:35:09 -04:00
wyj 19f3a91b7f Makefile: build all backends with optional HDR output 2026-08-30 02:29:41 -04:00
wyj 09923b9b18 Frame: localize shadow boundary refinement 2026-08-29 22:58:04 -04:00
wyj e47005ea6b Optics: bound cache-only bright PSF splats 2026-08-29 20:02:45 -04:00
wyj 6f8ec50c64 Makefile: add Debug build diagnostics 2026-08-29 19:23:50 -04:00
wyj 9e22d59213 Frame: optimize midpoint lookup 2026-08-29 19:05:43 -04:00
wyj 6cb4184b54 Output: expand verbose render progress 2026-08-29 16:26:44 -04:00
wyj d1320fffc4 Revert: remove triangle quality refinement 2026-08-29 03:26:17 -04:00
wyj 451b700b32 Frame: guard sliver-producing refinement splits 2026-08-29 02:06:06 -04:00
wyj b95c6579bd Frame: gate fold refinement by Jacobian 2026-08-29 01:17:47 -04:00
wyj 4c575fa8b5 Frame: refine critical lens regions 2026-08-29 00:44:20 -04:00
wyj 48dcf4e707 Frame: add adaptive mesh refinement 2026-08-28 23:06:53 -04:00
wyj 4aa5f6666a Output: validate image format before rendering 2026-08-28 20:38:39 -04:00
wyj d5d21c7659 Output: write HDR FITS with camera metadata 2026-08-28 19:56:07 -04:00
wyj 5fd8e3819d Spacetime: free analytic render worker count 2026-08-28 17:20:00 -04:00
wyj 9b23f3717e Output: Add render progress logging 2026-08-28 16:19:01 -04:00
wyj 69819bb24f Optics: parallelize PSF cache construction 2026-08-28 15:32:06 -04:00
wyj 7af3a52cb9 Optics: scale PSF cache radius by parameters 2026-08-28 14:43:21 -04:00
wyj 80b891d2c8 Optics: cache pixel-area Moffat kernels 2026-08-28 13:53:06 -04:00
wyj a71180706f DOC: update video_rendering_plan.md 2026-08-28 12:54:58 -04:00
wyj 172738d932 Doc: record dense 2MASS tile benchmark 2026-08-27 19:29:55 -04:00
wyj 5d80bb1b51 Makefile: tune default CFLAGS 2026-08-27 13:04:36 -04:00
wyj b0fdf1faca Output: default to PNG with PPM fallback 2026-08-27 04:25:45 -04:00
wyj 83dc05fd52 Observer: parameterize Schwarzschild camera view 2026-08-27 04:12:17 -04:00
wyj 457c0728b4 Catalog: parallelize all-sky tile prefetch 2026-08-27 03:53:45 -04:00
wyj d9f2c83b9f Frame: Dynamically balance catalog splats 2026-08-27 00:36:38 -04:00
wyj 7b034f42e8 Doc: define commit message categories 2026-08-26 23:31:40 -04:00
wyj b0ca7c6df1 Add lazy tiled 2MASS catalog support 2026-08-26 23:19:07 -04:00
wyj 747f8eb695 Add resumable all-sky 2MASS acquisition 2026-08-26 20:04:08 -04:00
wyj 388055c01e Decouple movie tracing through metric slabs 2026-08-26 01:35:57 -04:00
wyj d3522886f0 Aggregate movie rays through time slabs 2026-08-26 00:14:26 -04:00
wyj 2a00d7a2d1 Recalibrate default test catalog exposure 2026-08-26 00:00:52 -04:00
wyj 5e8286675d Use lower default exposure for movie tests 2026-08-25 23:58:20 -04:00
wyj c8b321a62a Add observer track PNG movie mode 2026-08-25 22:56:10 -04:00
wyj d3dc2e5836 Add video rendering implementation plan 2026-08-25 22:50:06 -04:00
wyj d30f9ac56c init 2026-08-25 22:45:43 -04:00