Add row-averaged FFT feature with configurable window size (row_avg_n parameter) for improved signal-to-noise ratio on noisy scans. Includes: - Gaussian-weighted same-row neighbor averaging (never crosses rows) - Masked/renormalized convolution handling for edge cases and masked samples - Cache format v2 with row_avg_n tracking to prevent silent cache mismatches - GUI dialog option for row-average window configuration - Comprehensive tests validating kernel properties, background subtraction invariance, and cache dispatch Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
11 KiB
sras-viewer design notes
Rationale that outgrew code comments. Each section is referenced by a short pointer comment at the relevant definition, so the code stays scannable and the reasoning stays findable.
Memory budget and row chunking (sras_compute.py)
DC images are computed over row chunks so the float32 working buffers for one chunk stay under a memory budget. A fixed row count (the original design) works fine for small legacy scans but is catastrophic for a v6 scan with a large per-angle frame/sample count — e.g. a 7500-frame × 2500-sample angle needs ~2.4 GB for a single 32-row chunk.
With chunks running concurrently the budget has to cover all live chunks at
once. On a large scan chunk_rows is already clamped to its floor of one row
(one row alone is ~75 MB of float32 at 7507×2500), so shrinking the per-chunk
size cannot buy more concurrency — the worker count must be derived from the
budget instead: _plan_chunks picks the worker count first and sizes the
chunk to it. Sizing the chunk first is the trap: a single chunk would always
consume the whole budget and leave room for exactly one worker, precisely on
the large scans that need concurrency most.
The 1024 MB default (SRAS_MEM_BUDGET_MB) is the measured knee on a 16-core
machine against a 7507-frame × 2500-sample angle: 512 MB left ~20% of the
speedup on the table, and 1536+ MB cost ~0.4 GB more resident memory for no
further gain.
A caller that itself runs several computations concurrently (angle-level
parallelism, plan_angle_level) must pass both max_workers=1 and its
share of the budget. Capping the workers alone is not enough: the chunk would
still be sized against the whole budget, and N concurrent callers would each
allocate all of it.
FFT peak search: block-parallel zoom refinement (sras_compute.py)
The displayed RF value per pixel is the argmax of the zero-padded power spectrum of that pixel's CH1 waveform. At the pad factor of 40 needed for mapping resolution, materialising padded spectra is hopeless: ~9 GB per scan row, which is what used to collapse the old row-chunk planner to one worker and make synthesis single-threaded.
_peak_bins_zoom never materialises the padded spectrum:
- a coarse rfft at
next_fast_len(2*spf)— 2× oversampled, so the padded power spectrum (a trig polynomial of degree spf−1) cannot hide its global max between coarse samples; - every coarse bin within
_ZOOM_CAND_RATIO(0.7) of its row's coarse max becomes a refinement candidate. Quarter-natural-bin scalloping at the 2× grid can understate a peak's power by at most ~19%, so 0.7 keeps a wide margin. The DC-adjacent window is always refined too: the coarse DC bin is zeroed for suppression, which would otherwise blind the scan to fine bins closer to DC than the first coarse sample (where the leakage skirt of an un-subtracted offset peaks); - each candidate window (±
_ZOOM_HALFWIDTH= 0.75 coarse spacings; every fine bin lies within 0.5 spacings of its nearest coarse bin) is evaluated on the exactn_fftgrid by one small complex gemm, with np.argmax's lowest-bin tie-break preserved across windows.
The selected bin is bit-identical to the full padded argmax — enforced by
tests/test_compute.py::test_zoom_identity, a fuzz test over adversarial
spectra, and the golden-hash harness (tools/check_equivalence.py), whose
baseline was captured on the old full-padded path.
Work fans out over a persistent thread pool in _FFT_BLOCK = 512-waveform
tasks: smaller blocks serialise on GIL-held numpy dispatch, larger ones lose
cache residency and task granularity (measured on a 16-core machine, where
this path runs ~35× faster than the old serial padded transform at pad 40).
pyFFTW runs through per-thread builders plans (FFTW_MEASURE, wisdom
persisted under ~/.cache/sras-viewer/), and threadpoolctl clamps BLAS to
one thread under the pool so the refinement gemm cannot oversubscribe.
compute_rf_image(exact=True) (or SRAS_FFT_EXACT=1) keeps the reference
full-padded path for audits.
Row-averaged FFT: same-row, distance-weighted SNR cleanup (sras_compute.py)
compute_rf_image's row_avg_n parameter averages each pixel's CH1
waveform with its up-to-n same-row neighbors before the FFT peak search, to
improve SNR on noisy scans. Never crosses rows: pixel pitch is strongly
anisotropic and varies by scan (5 µm × 50 µm on a typical scan, but as
stretched as 5 µm × 1 mm on others), so a physically meaningful "neighbor"
set can't be a fixed-shape 2-D window — but the X pitch within one row is
a single file-wide constant (SrasFile.pixel_x_mm), so restricting to the
row axis sidesteps the anisotropy question entirely rather than solving it
with an elliptical or physically-scaled 2-D kernel.
_row_average_weights is a Gaussian in pixel-index distance, not physical
mm distance — deliberately: within one row those are the same function up
to a fixed scale factor (pixel_x_mm is constant along a row), so the
kernel itself needs no pitch at all. pixel_x_mm is used for real exactly
once, in the GUI's options dialog, to show the window's physical width —
not in the kernel math, where it would only ever cancel out.
_row_average_waveforms is a masked/renormalized convolution (two
correlate1d calls, numerator and denominator, divided) rather than a
single fixed-normalized convolution, because a masked neighbor must
contribute zero weight, not a zero-amplitude sample at full weight — the
latter would bias every average near a masked run or a row's own edge
toward zero. The same two-correlation trick handles row-edge truncation for
free: mode="constant", cval=0.0 zero-pads both the numerator and the
denominator beyond a row's own ends, so the output renormalizes by whatever
weight sum actually landed inside the row, no separate edge case.
Background subtraction stays exactly where it already was (subtracted once
from the fully-assembled waves buffer) rather than being threaded into the
per-neighbor gather. This is exact, not an approximation: because
_row_average_waveforms's denominator is always the actual sum of
included, valid weights (never a fixed total), Σwᵢ·(rawᵢ−bg) / Σwᵢ
distributes to avg − bg·(Σwᵢ/Σwᵢ) = avg − bg regardless of which or how
many neighbors were included — subtracting background from the averaged
waveform is identical to subtracting it from every neighbor first, for any
window, at any row edge, with any number of masked-out neighbors.
No cross-row halo is needed: compute_rf_image's chunk loop already splits
on rows only, and read_row already reads one row's complete
(n_frames, spf) slice at a time — averaging happens entirely inside that
one row's own frame axis, so a chunk boundary (which falls between rows)
can never truncate a window. Only a row's own start/end can, and that's the
same edge case the masked convolution already handles.
The averaging step doubles the live per-row scratch memory (a full-width
(n_frames, spf) buffer on top of the existing compacted waves buffer),
so compute_rf_image halves its byte budget when row_avg_n > 0 before
_plan_fft_rows/the exact-path sizing runs — see "Memory budget and row
chunking" above. On the largest real scans _plan_chunks is already
clamped to its floor of one row regardless, so this costs no concurrency
where it matters most; it mainly protects moderate-sized scans from an
unexpected regression.
Persistence: cached_rf_image (the extracted fast-path check) requires
sras.precomputed_row_avg_n == row_avg_n exactly, so a raw request can
never be silently served a row-averaged cache or vice versa, and a request
at one window size can never be served a cache at another — see
scan_format.md's Cache Tail / CACH tail version history sections for the
on-disk row_avg_n field this depends on.
Angle alignment coordinate frames (sras_compute.py)
Alignment puts every angle's images onto one shared, zero-padded pixel grid using a rigid transform only — rotation + translation, never scale.
Angle 0 (the reference) is the sole coordinate authority: it is the only
angle whose stage XY (x_start_mm / y_positions_mm) is ever read, and the
shared canvas is literally an extension of angle 0's own pixel grid, so the
aligned view carries angle 0's real X/Y axes. Every other angle is placed
purely by content — its rotation and translation come from cross-correlating
its CH4 image against angle 0's (register_angle_to_reference) — and its own
stage XY is deliberately never consulted. That is not an oversight: the
rotation stage moves the sample relative to the scan window, so where a
window sat in stage coordinates says nothing about where the sample is, and
an earlier design that pivoted each angle on a signal-weighted centroid of
its own window put every angle on a ~20 mm circle around the optical center
instead of stacking them into one shape.
Only two coordinate frames exist:
- local mm — one angle's own physical frame: origin at the center of its own pixel array, x along +column, y along +row, scaled by that angle's own pitches. Carries no stage position whatsoever.
- ref mm — the reference angle's local mm. A registration result
(rotation_deg, shift_mm)is exactly the rigid map from an angle's local mm to ref mm:q = R(rotation_deg) @ l + shift_mm. Stage coordinates re-enter once, at the very end, when the canvas origin is converted to angle 0's stage mm (AlignmentResult.canvas_origin_mm).
Rotation is done in mm, never on raw pixel indices: the x pitch
(SrasFile.pixel_x_mm, 5 µm on a real scan) and the y/row pitch (50 µm)
differ by 10×, so rotating the raw index grid would shear the image — an
unwanted anisotropic scale. Registration runs on a resampled isotropic grid
for the same reason, and every affine maps shared-grid index → mm → undo
rotation/shift → that angle's own local mm → that angle's own raw index,
matching the output→input convention scipy.ndimage.affine_transform wants.
Manual-alignment sidecar (sras_compute.py)
<name>.sras.align.json lives next to the scan file. The code lives in
sras_compute, not sras_format: sras_format is scoped to the versioned
binary .sras spec itself (see scan_format.md), while a manual alignment is
a viewer-computed derived artifact, analogous in kind to AlignmentResult
— so it belongs with the alignment math it serialises. json + pathlib are
stdlib, so this adds no dependency to a module whose load-bearing constraint
is staying free of Qt/matplotlib for cheap multiprocessing-child imports.
Schema history
The stored rotation_deg/shift_mm are meaningless without the frame they
were measured in, so _SIDECAR_SCHEMA_VERSION is bumped whenever that frame
changes. Each bump makes older files describe a different (and, for the bugs
each bump fixed, actively wrong) transform than the same numbers would today;
loading one unchanged would silently reproduce the very "scans show up
everywhere" symptom the bump fixed — so older sidecars are treated as absent
rather than migrated.
- 1 → 2 — pivot moved from the scan-window bbox center to a content-derived centroid, and the rotation sign convention was corrected.
- 2 → 3 — the content centroid was abandoned entirely: rotation is now
about each angle's own array center, mapped onto the reference's array
center, with
shift_mmin the reference's local mm frame. No angle but the reference contributes stage coordinates any more.