3989c2a1b8
Add row-averaged FFT feature with configurable window size (row_avg_n parameter) for improved signal-to-noise ratio on noisy scans. Includes: - Gaussian-weighted same-row neighbor averaging (never crosses rows) - Masked/renormalized convolution handling for edge cases and masked samples - Cache format v2 with row_avg_n tracking to prevent silent cache mismatches - GUI dialog option for row-average window configuration - Comprehensive tests validating kernel properties, background subtraction invariance, and cache dispatch Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
200 lines
11 KiB
Markdown
200 lines
11 KiB
Markdown
# sras-viewer design notes
|
||
|
||
Rationale that outgrew code comments. Each section is referenced by a short
|
||
pointer comment at the relevant definition, so the code stays scannable and
|
||
the reasoning stays findable.
|
||
|
||
## Memory budget and row chunking (`sras_compute.py`)
|
||
|
||
DC images are computed over row chunks so the float32 working buffers for one
|
||
chunk stay under a memory budget. A fixed row count (the original design)
|
||
works fine for small legacy scans but is catastrophic for a v6 scan with a
|
||
large per-angle frame/sample count — e.g. a 7500-frame × 2500-sample angle
|
||
needs ~2.4 GB for a single 32-row chunk.
|
||
|
||
With chunks running concurrently the budget has to cover *all* live chunks at
|
||
once. On a large scan `chunk_rows` is already clamped to its floor of one row
|
||
(one row alone is ~75 MB of float32 at 7507×2500), so shrinking the per-chunk
|
||
size cannot buy more concurrency — the worker count must be derived from the
|
||
budget instead: `_plan_chunks` picks the worker count *first* and sizes the
|
||
chunk to it. Sizing the chunk first is the trap: a single chunk would always
|
||
consume the whole budget and leave room for exactly one worker, precisely on
|
||
the large scans that need concurrency most.
|
||
|
||
The 1024 MB default (`SRAS_MEM_BUDGET_MB`) is the measured knee on a 16-core
|
||
machine against a 7507-frame × 2500-sample angle: 512 MB left ~20% of the
|
||
speedup on the table, and 1536+ MB cost ~0.4 GB more resident memory for no
|
||
further gain.
|
||
|
||
A caller that itself runs several computations concurrently (angle-level
|
||
parallelism, `plan_angle_level`) must pass *both* `max_workers=1` and its
|
||
share of the budget. Capping the workers alone is not enough: the chunk would
|
||
still be sized against the whole budget, and N concurrent callers would each
|
||
allocate all of it.
|
||
|
||
## FFT peak search: block-parallel zoom refinement (`sras_compute.py`)
|
||
|
||
The displayed RF value per pixel is the argmax of the zero-padded power
|
||
spectrum of that pixel's CH1 waveform. At the pad factor of 40 needed for
|
||
mapping resolution, materialising padded spectra is hopeless: ~9 GB per scan
|
||
row, which is what used to collapse the old row-chunk planner to one worker
|
||
and make synthesis single-threaded.
|
||
|
||
`_peak_bins_zoom` never materialises the padded spectrum:
|
||
|
||
1. a coarse rfft at `next_fast_len(2*spf)` — 2× oversampled, so the padded
|
||
power spectrum (a trig polynomial of degree spf−1) cannot hide its global
|
||
max between coarse samples;
|
||
2. every coarse bin within `_ZOOM_CAND_RATIO` (0.7) of its row's coarse max
|
||
becomes a refinement candidate. Quarter-natural-bin scalloping at the 2×
|
||
grid can understate a peak's power by at most ~19%, so 0.7 keeps a wide
|
||
margin. The DC-adjacent window is always refined too: the coarse DC bin
|
||
is zeroed for suppression, which would otherwise blind the scan to fine
|
||
bins closer to DC than the first coarse sample (where the leakage skirt
|
||
of an un-subtracted offset peaks);
|
||
3. each candidate window (±`_ZOOM_HALFWIDTH` = 0.75 coarse spacings; every
|
||
fine bin lies within 0.5 spacings of its nearest coarse bin) is evaluated
|
||
on the exact `n_fft` grid by one small complex gemm, with np.argmax's
|
||
lowest-bin tie-break preserved across windows.
|
||
|
||
The selected bin is bit-identical to the full padded argmax — enforced by
|
||
`tests/test_compute.py::test_zoom_identity`, a fuzz test over adversarial
|
||
spectra, and the golden-hash harness (`tools/check_equivalence.py`), whose
|
||
baseline was captured on the old full-padded path.
|
||
|
||
Work fans out over a persistent thread pool in `_FFT_BLOCK` = 512-waveform
|
||
tasks: smaller blocks serialise on GIL-held numpy dispatch, larger ones lose
|
||
cache residency and task granularity (measured on a 16-core machine, where
|
||
this path runs ~35× faster than the old serial padded transform at pad 40).
|
||
pyFFTW runs through per-thread `builders` plans (FFTW_MEASURE, wisdom
|
||
persisted under `~/.cache/sras-viewer/`), and `threadpoolctl` clamps BLAS to
|
||
one thread under the pool so the refinement gemm cannot oversubscribe.
|
||
`compute_rf_image(exact=True)` (or `SRAS_FFT_EXACT=1`) keeps the reference
|
||
full-padded path for audits.
|
||
|
||
## Row-averaged FFT: same-row, distance-weighted SNR cleanup (`sras_compute.py`)
|
||
|
||
`compute_rf_image`'s `row_avg_n` parameter averages each pixel's CH1
|
||
waveform with its up-to-n same-row neighbors before the FFT peak search, to
|
||
improve SNR on noisy scans. Never crosses rows: pixel pitch is strongly
|
||
anisotropic and varies by scan (5 µm × 50 µm on a typical scan, but as
|
||
stretched as 5 µm × 1 mm on others), so a physically meaningful "neighbor"
|
||
set can't be a fixed-shape 2-D window — but the X pitch *within one row* is
|
||
a single file-wide constant (`SrasFile.pixel_x_mm`), so restricting to the
|
||
row axis sidesteps the anisotropy question entirely rather than solving it
|
||
with an elliptical or physically-scaled 2-D kernel.
|
||
|
||
`_row_average_weights` is a Gaussian in pixel-index distance, not physical
|
||
mm distance — deliberately: within one row those are the same function up
|
||
to a fixed scale factor (`pixel_x_mm` is constant along a row), so the
|
||
kernel itself needs no pitch at all. `pixel_x_mm` is used for real exactly
|
||
once, in the GUI's options dialog, to show the window's physical width —
|
||
not in the kernel math, where it would only ever cancel out.
|
||
|
||
`_row_average_waveforms` is a masked/renormalized convolution (two
|
||
`correlate1d` calls, numerator and denominator, divided) rather than a
|
||
single fixed-normalized convolution, because a masked neighbor must
|
||
contribute *zero weight*, not a zero-amplitude sample at full weight — the
|
||
latter would bias every average near a masked run or a row's own edge
|
||
toward zero. The same two-correlation trick handles row-edge truncation for
|
||
free: `mode="constant", cval=0.0` zero-pads both the numerator and the
|
||
denominator beyond a row's own ends, so the output renormalizes by whatever
|
||
weight sum actually landed inside the row, no separate edge case.
|
||
|
||
Background subtraction stays exactly where it already was (subtracted once
|
||
from the fully-assembled `waves` buffer) rather than being threaded into the
|
||
per-neighbor gather. This is exact, not an approximation: because
|
||
`_row_average_waveforms`'s denominator is always the *actual* sum of
|
||
included, valid weights (never a fixed total), `Σwᵢ·(rawᵢ−bg) / Σwᵢ`
|
||
distributes to `avg − bg·(Σwᵢ/Σwᵢ) = avg − bg` regardless of which or how
|
||
many neighbors were included — subtracting background from the averaged
|
||
waveform is identical to subtracting it from every neighbor first, for any
|
||
window, at any row edge, with any number of masked-out neighbors.
|
||
|
||
No cross-row halo is needed: `compute_rf_image`'s chunk loop already splits
|
||
on rows only, and `read_row` already reads one row's complete
|
||
`(n_frames, spf)` slice at a time — averaging happens entirely inside that
|
||
one row's own frame axis, so a chunk boundary (which falls between rows)
|
||
can never truncate a window. Only a row's own start/end can, and that's the
|
||
same edge case the masked convolution already handles.
|
||
|
||
The averaging step doubles the live per-row scratch memory (a full-width
|
||
`(n_frames, spf)` buffer on top of the existing compacted `waves` buffer),
|
||
so `compute_rf_image` halves its byte budget when `row_avg_n > 0` before
|
||
`_plan_fft_rows`/the exact-path sizing runs — see "Memory budget and row
|
||
chunking" above. On the largest real scans `_plan_chunks` is already
|
||
clamped to its floor of one row regardless, so this costs no concurrency
|
||
where it matters most; it mainly protects moderate-sized scans from an
|
||
unexpected regression.
|
||
|
||
Persistence: `cached_rf_image` (the extracted fast-path check) requires
|
||
`sras.precomputed_row_avg_n == row_avg_n` exactly, so a raw request can
|
||
never be silently served a row-averaged cache or vice versa, and a request
|
||
at one window size can never be served a cache at another — see
|
||
`scan_format.md`'s Cache Tail / CACH tail version history sections for the
|
||
on-disk `row_avg_n` field this depends on.
|
||
|
||
## Angle alignment coordinate frames (`sras_compute.py`)
|
||
|
||
Alignment puts every angle's images onto one shared, zero-padded pixel grid
|
||
using a rigid transform only — rotation + translation, never scale.
|
||
|
||
Angle 0 (the reference) is the sole coordinate authority: it is the only
|
||
angle whose stage XY (`x_start_mm` / `y_positions_mm`) is ever read, and the
|
||
shared canvas is literally an extension of angle 0's own pixel grid, so the
|
||
aligned view carries angle 0's real X/Y axes. Every *other* angle is placed
|
||
purely by content — its rotation and translation come from cross-correlating
|
||
its CH4 image against angle 0's (`register_angle_to_reference`) — and its own
|
||
stage XY is deliberately never consulted. That is not an oversight: the
|
||
rotation stage moves the sample relative to the scan window, so where a
|
||
window sat in stage coordinates says nothing about where the sample is, and
|
||
an earlier design that pivoted each angle on a signal-weighted centroid of
|
||
its own window put every angle on a ~20 mm circle around the optical center
|
||
instead of stacking them into one shape.
|
||
|
||
Only two coordinate frames exist:
|
||
|
||
* **local mm** — one angle's own physical frame: origin at the *center of its
|
||
own pixel array*, x along +column, y along +row, scaled by that angle's own
|
||
pitches. Carries no stage position whatsoever.
|
||
* **ref mm** — the reference angle's local mm. A registration result
|
||
`(rotation_deg, shift_mm)` is exactly the rigid map from an angle's local
|
||
mm to ref mm: `q = R(rotation_deg) @ l + shift_mm`. Stage coordinates
|
||
re-enter once, at the very end, when the canvas origin is converted to
|
||
angle 0's stage mm (`AlignmentResult.canvas_origin_mm`).
|
||
|
||
Rotation is done in mm, never on raw pixel indices: the x pitch
|
||
(`SrasFile.pixel_x_mm`, 5 µm on a real scan) and the y/row pitch (50 µm)
|
||
differ by 10×, so rotating the raw index grid would shear the image — an
|
||
unwanted anisotropic scale. Registration runs on a resampled *isotropic* grid
|
||
for the same reason, and every affine maps shared-grid index → mm → undo
|
||
rotation/shift → that angle's own local mm → that angle's own raw index,
|
||
matching the output→input convention `scipy.ndimage.affine_transform` wants.
|
||
|
||
## Manual-alignment sidecar (`sras_compute.py`)
|
||
|
||
`<name>.sras.align.json` lives next to the scan file. The code lives in
|
||
`sras_compute`, not `sras_format`: `sras_format` is scoped to the versioned
|
||
binary .sras spec itself (see `scan_format.md`), while a manual alignment is
|
||
a viewer-computed *derived* artifact, analogous in kind to `AlignmentResult`
|
||
— so it belongs with the alignment math it serialises. json + pathlib are
|
||
stdlib, so this adds no dependency to a module whose load-bearing constraint
|
||
is staying free of Qt/matplotlib for cheap multiprocessing-child imports.
|
||
|
||
### Schema history
|
||
|
||
The stored `rotation_deg`/`shift_mm` are meaningless without the frame they
|
||
were measured in, so `_SIDECAR_SCHEMA_VERSION` is bumped whenever that frame
|
||
changes. Each bump makes older files describe a different (and, for the bugs
|
||
each bump fixed, actively wrong) transform than the same numbers would today;
|
||
loading one unchanged would silently reproduce the very "scans show up
|
||
everywhere" symptom the bump fixed — so older sidecars are treated as absent
|
||
rather than migrated.
|
||
|
||
* **1 → 2** — pivot moved from the scan-window bbox center to a
|
||
content-derived centroid, and the rotation sign convention was corrected.
|
||
* **2 → 3** — the content centroid was abandoned entirely: rotation is now
|
||
about each angle's own array center, mapped onto the reference's array
|
||
center, with `shift_mm` in the reference's local mm frame. No angle but the
|
||
reference contributes stage coordinates any more.
|