Files
sras-viewer/docs/design.md
T
Thomas Ales 8154c57066 Move design essays to docs/design.md, leave pointers
The three block essays in sras_compute.py — memory budget / row
chunking, angle-alignment coordinate frames, and the sidecar placement
+ schema history — move to docs/design.md (joined by a new section on
the zoom FFT peak search), each replaced by a 2-4 line pointer. The
save_manual_alignment docstring no longer restates the JSON schema its
own eight lines of code construct.

Load-bearing trap notes (normalization=None, subpixel-score mixing,
memmap lazy reads) stay in place. Zero code changes — golden-hash diff
verified empty.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-06 11:01:49 -05:00

7.6 KiB
Raw Blame History

sras-viewer design notes

Rationale that outgrew code comments. Each section is referenced by a short pointer comment at the relevant definition, so the code stays scannable and the reasoning stays findable.

Memory budget and row chunking (sras_compute.py)

DC images are computed over row chunks so the float32 working buffers for one chunk stay under a memory budget. A fixed row count (the original design) works fine for small legacy scans but is catastrophic for a v6 scan with a large per-angle frame/sample count — e.g. a 7500-frame × 2500-sample angle needs ~2.4 GB for a single 32-row chunk.

With chunks running concurrently the budget has to cover all live chunks at once. On a large scan chunk_rows is already clamped to its floor of one row (one row alone is ~75 MB of float32 at 7507×2500), so shrinking the per-chunk size cannot buy more concurrency — the worker count must be derived from the budget instead: _plan_chunks picks the worker count first and sizes the chunk to it. Sizing the chunk first is the trap: a single chunk would always consume the whole budget and leave room for exactly one worker, precisely on the large scans that need concurrency most.

The 1024 MB default (SRAS_MEM_BUDGET_MB) is the measured knee on a 16-core machine against a 7507-frame × 2500-sample angle: 512 MB left ~20% of the speedup on the table, and 1536+ MB cost ~0.4 GB more resident memory for no further gain.

A caller that itself runs several computations concurrently (angle-level parallelism, plan_angle_level) must pass both max_workers=1 and its share of the budget. Capping the workers alone is not enough: the chunk would still be sized against the whole budget, and N concurrent callers would each allocate all of it.

FFT peak search: block-parallel zoom refinement (sras_compute.py)

The displayed RF value per pixel is the argmax of the zero-padded power spectrum of that pixel's CH1 waveform. At the pad factor of 40 needed for mapping resolution, materialising padded spectra is hopeless: ~9 GB per scan row, which is what used to collapse the old row-chunk planner to one worker and make synthesis single-threaded.

_peak_bins_zoom never materialises the padded spectrum:

  1. a coarse rfft at next_fast_len(2*spf) — 2× oversampled, so the padded power spectrum (a trig polynomial of degree spf−1) cannot hide its global max between coarse samples;
  2. every coarse bin within _ZOOM_CAND_RATIO (0.7) of its row's coarse max becomes a refinement candidate. Quarter-natural-bin scalloping at the 2× grid can understate a peak's power by at most ~19%, so 0.7 keeps a wide margin. The DC-adjacent window is always refined too: the coarse DC bin is zeroed for suppression, which would otherwise blind the scan to fine bins closer to DC than the first coarse sample (where the leakage skirt of an un-subtracted offset peaks);
  3. each candidate window (±_ZOOM_HALFWIDTH = 0.75 coarse spacings; every fine bin lies within 0.5 spacings of its nearest coarse bin) is evaluated on the exact n_fft grid by one small complex gemm, with np.argmax's lowest-bin tie-break preserved across windows.

The selected bin is bit-identical to the full padded argmax — enforced by tests/test_compute.py::test_zoom_identity, a fuzz test over adversarial spectra, and the golden-hash harness (tools/check_equivalence.py), whose baseline was captured on the old full-padded path.

Work fans out over a persistent thread pool in _FFT_BLOCK = 512-waveform tasks: smaller blocks serialise on GIL-held numpy dispatch, larger ones lose cache residency and task granularity (measured on a 16-core machine, where this path runs ~35× faster than the old serial padded transform at pad 40). pyFFTW runs through per-thread builders plans (FFTW_MEASURE, wisdom persisted under ~/.cache/sras-viewer/), and threadpoolctl clamps BLAS to one thread under the pool so the refinement gemm cannot oversubscribe. compute_rf_image(exact=True) (or SRAS_FFT_EXACT=1) keeps the reference full-padded path for audits.

Angle alignment coordinate frames (sras_compute.py)

Alignment puts every angle's images onto one shared, zero-padded pixel grid using a rigid transform only — rotation + translation, never scale.

Angle 0 (the reference) is the sole coordinate authority: it is the only angle whose stage XY (x_start_mm / y_positions_mm) is ever read, and the shared canvas is literally an extension of angle 0's own pixel grid, so the aligned view carries angle 0's real X/Y axes. Every other angle is placed purely by content — its rotation and translation come from cross-correlating its CH4 image against angle 0's (register_angle_to_reference) — and its own stage XY is deliberately never consulted. That is not an oversight: the rotation stage moves the sample relative to the scan window, so where a window sat in stage coordinates says nothing about where the sample is, and an earlier design that pivoted each angle on a signal-weighted centroid of its own window put every angle on a ~20 mm circle around the optical center instead of stacking them into one shape.

Only two coordinate frames exist:

  • local mm — one angle's own physical frame: origin at the center of its own pixel array, x along +column, y along +row, scaled by that angle's own pitches. Carries no stage position whatsoever.
  • ref mm — the reference angle's local mm. A registration result (rotation_deg, shift_mm) is exactly the rigid map from an angle's local mm to ref mm: q = R(rotation_deg) @ l + shift_mm. Stage coordinates re-enter once, at the very end, when the canvas origin is converted to angle 0's stage mm (AlignmentResult.canvas_origin_mm).

Rotation is done in mm, never on raw pixel indices: the x pitch (SrasFile.pixel_x_mm, 5 µm on a real scan) and the y/row pitch (50 µm) differ by 10×, so rotating the raw index grid would shear the image — an unwanted anisotropic scale. Registration runs on a resampled isotropic grid for the same reason, and every affine maps shared-grid index → mm → undo rotation/shift → that angle's own local mm → that angle's own raw index, matching the output→input convention scipy.ndimage.affine_transform wants.

Manual-alignment sidecar (sras_compute.py)

<name>.sras.align.json lives next to the scan file. The code lives in sras_compute, not sras_format: sras_format is scoped to the versioned binary .sras spec itself (see scan_format.md), while a manual alignment is a viewer-computed derived artifact, analogous in kind to AlignmentResult — so it belongs with the alignment math it serialises. json + pathlib are stdlib, so this adds no dependency to a module whose load-bearing constraint is staying free of Qt/matplotlib for cheap multiprocessing-child imports.

Schema history

The stored rotation_deg/shift_mm are meaningless without the frame they were measured in, so _SIDECAR_SCHEMA_VERSION is bumped whenever that frame changes. Each bump makes older files describe a different (and, for the bugs each bump fixed, actively wrong) transform than the same numbers would today; loading one unchanged would silently reproduce the very "scans show up everywhere" symptom the bump fixed — so older sidecars are treated as absent rather than migrated.

  • 1 → 2 — pivot moved from the scan-window bbox center to a content-derived centroid, and the rotation sign convention was corrected.
  • 2 → 3 — the content centroid was abandoned entirely: rotation is now about each angle's own array center, mapped onto the reference's array center, with shift_mm in the reference's local mm frame. No angle but the reference contributes stage coordinates any more.