Move design essays to docs/design.md, leave pointers
The three block essays in sras_compute.py — memory budget / row chunking, angle-alignment coordinate frames, and the sidecar placement + schema history — move to docs/design.md (joined by a new section on the zoom FFT peak search), each replaced by a 2-4 line pointer. The save_manual_alignment docstring no longer restates the JSON schema its own eight lines of code construct. Load-bearing trap notes (normalization=None, subpixel-score mixing, memmap lazy reads) stay in place. Zero code changes — golden-hash diff verified empty. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
+137
@@ -0,0 +1,137 @@
|
||||
# sras-viewer design notes
|
||||
|
||||
Rationale that outgrew code comments. Each section is referenced by a short
|
||||
pointer comment at the relevant definition, so the code stays scannable and
|
||||
the reasoning stays findable.
|
||||
|
||||
## Memory budget and row chunking (`sras_compute.py`)
|
||||
|
||||
DC images are computed over row chunks so the float32 working buffers for one
|
||||
chunk stay under a memory budget. A fixed row count (the original design)
|
||||
works fine for small legacy scans but is catastrophic for a v6 scan with a
|
||||
large per-angle frame/sample count — e.g. a 7500-frame × 2500-sample angle
|
||||
needs ~2.4 GB for a single 32-row chunk.
|
||||
|
||||
With chunks running concurrently the budget has to cover *all* live chunks at
|
||||
once. On a large scan `chunk_rows` is already clamped to its floor of one row
|
||||
(one row alone is ~75 MB of float32 at 7507×2500), so shrinking the per-chunk
|
||||
size cannot buy more concurrency — the worker count must be derived from the
|
||||
budget instead: `_plan_chunks` picks the worker count *first* and sizes the
|
||||
chunk to it. Sizing the chunk first is the trap: a single chunk would always
|
||||
consume the whole budget and leave room for exactly one worker, precisely on
|
||||
the large scans that need concurrency most.
|
||||
|
||||
The 1024 MB default (`SRAS_MEM_BUDGET_MB`) is the measured knee on a 16-core
|
||||
machine against a 7507-frame × 2500-sample angle: 512 MB left ~20% of the
|
||||
speedup on the table, and 1536+ MB cost ~0.4 GB more resident memory for no
|
||||
further gain.
|
||||
|
||||
A caller that itself runs several computations concurrently (angle-level
|
||||
parallelism, `plan_angle_level`) must pass *both* `max_workers=1` and its
|
||||
share of the budget. Capping the workers alone is not enough: the chunk would
|
||||
still be sized against the whole budget, and N concurrent callers would each
|
||||
allocate all of it.
|
||||
|
||||
## FFT peak search: block-parallel zoom refinement (`sras_compute.py`)
|
||||
|
||||
The displayed RF value per pixel is the argmax of the zero-padded power
|
||||
spectrum of that pixel's CH1 waveform. At the pad factor of 40 needed for
|
||||
mapping resolution, materialising padded spectra is hopeless: ~9 GB per scan
|
||||
row, which is what used to collapse the old row-chunk planner to one worker
|
||||
and make synthesis single-threaded.
|
||||
|
||||
`_peak_bins_zoom` never materialises the padded spectrum:
|
||||
|
||||
1. a coarse rfft at `next_fast_len(2*spf)` — 2× oversampled, so the padded
|
||||
power spectrum (a trig polynomial of degree spf−1) cannot hide its global
|
||||
max between coarse samples;
|
||||
2. every coarse bin within `_ZOOM_CAND_RATIO` (0.7) of its row's coarse max
|
||||
becomes a refinement candidate. Quarter-natural-bin scalloping at the 2×
|
||||
grid can understate a peak's power by at most ~19%, so 0.7 keeps a wide
|
||||
margin. The DC-adjacent window is always refined too: the coarse DC bin
|
||||
is zeroed for suppression, which would otherwise blind the scan to fine
|
||||
bins closer to DC than the first coarse sample (where the leakage skirt
|
||||
of an un-subtracted offset peaks);
|
||||
3. each candidate window (±`_ZOOM_HALFWIDTH` = 0.75 coarse spacings; every
|
||||
fine bin lies within 0.5 spacings of its nearest coarse bin) is evaluated
|
||||
on the exact `n_fft` grid by one small complex gemm, with np.argmax's
|
||||
lowest-bin tie-break preserved across windows.
|
||||
|
||||
The selected bin is bit-identical to the full padded argmax — enforced by
|
||||
`tests/test_compute.py::test_zoom_identity`, a fuzz test over adversarial
|
||||
spectra, and the golden-hash harness (`tools/check_equivalence.py`), whose
|
||||
baseline was captured on the old full-padded path.
|
||||
|
||||
Work fans out over a persistent thread pool in `_FFT_BLOCK` = 512-waveform
|
||||
tasks: smaller blocks serialise on GIL-held numpy dispatch, larger ones lose
|
||||
cache residency and task granularity (measured on a 16-core machine, where
|
||||
this path runs ~35× faster than the old serial padded transform at pad 40).
|
||||
pyFFTW runs through per-thread `builders` plans (FFTW_MEASURE, wisdom
|
||||
persisted under `~/.cache/sras-viewer/`), and `threadpoolctl` clamps BLAS to
|
||||
one thread under the pool so the refinement gemm cannot oversubscribe.
|
||||
`compute_rf_image(exact=True)` (or `SRAS_FFT_EXACT=1`) keeps the reference
|
||||
full-padded path for audits.
|
||||
|
||||
## Angle alignment coordinate frames (`sras_compute.py`)
|
||||
|
||||
Alignment puts every angle's images onto one shared, zero-padded pixel grid
|
||||
using a rigid transform only — rotation + translation, never scale.
|
||||
|
||||
Angle 0 (the reference) is the sole coordinate authority: it is the only
|
||||
angle whose stage XY (`x_start_mm` / `y_positions_mm`) is ever read, and the
|
||||
shared canvas is literally an extension of angle 0's own pixel grid, so the
|
||||
aligned view carries angle 0's real X/Y axes. Every *other* angle is placed
|
||||
purely by content — its rotation and translation come from cross-correlating
|
||||
its CH4 image against angle 0's (`register_angle_to_reference`) — and its own
|
||||
stage XY is deliberately never consulted. That is not an oversight: the
|
||||
rotation stage moves the sample relative to the scan window, so where a
|
||||
window sat in stage coordinates says nothing about where the sample is, and
|
||||
an earlier design that pivoted each angle on a signal-weighted centroid of
|
||||
its own window put every angle on a ~20 mm circle around the optical center
|
||||
instead of stacking them into one shape.
|
||||
|
||||
Only two coordinate frames exist:
|
||||
|
||||
* **local mm** — one angle's own physical frame: origin at the *center of its
|
||||
own pixel array*, x along +column, y along +row, scaled by that angle's own
|
||||
pitches. Carries no stage position whatsoever.
|
||||
* **ref mm** — the reference angle's local mm. A registration result
|
||||
`(rotation_deg, shift_mm)` is exactly the rigid map from an angle's local
|
||||
mm to ref mm: `q = R(rotation_deg) @ l + shift_mm`. Stage coordinates
|
||||
re-enter once, at the very end, when the canvas origin is converted to
|
||||
angle 0's stage mm (`AlignmentResult.canvas_origin_mm`).
|
||||
|
||||
Rotation is done in mm, never on raw pixel indices: the x pitch
|
||||
(`SrasFile.pixel_x_mm`, 5 µm on a real scan) and the y/row pitch (50 µm)
|
||||
differ by 10×, so rotating the raw index grid would shear the image — an
|
||||
unwanted anisotropic scale. Registration runs on a resampled *isotropic* grid
|
||||
for the same reason, and every affine maps shared-grid index → mm → undo
|
||||
rotation/shift → that angle's own local mm → that angle's own raw index,
|
||||
matching the output→input convention `scipy.ndimage.affine_transform` wants.
|
||||
|
||||
## Manual-alignment sidecar (`sras_compute.py`)
|
||||
|
||||
`<name>.sras.align.json` lives next to the scan file. The code lives in
|
||||
`sras_compute`, not `sras_format`: `sras_format` is scoped to the versioned
|
||||
binary .sras spec itself (see `scan_format.md`), while a manual alignment is
|
||||
a viewer-computed *derived* artifact, analogous in kind to `AlignmentResult`
|
||||
— so it belongs with the alignment math it serialises. json + pathlib are
|
||||
stdlib, so this adds no dependency to a module whose load-bearing constraint
|
||||
is staying free of Qt/matplotlib for cheap multiprocessing-child imports.
|
||||
|
||||
### Schema history
|
||||
|
||||
The stored `rotation_deg`/`shift_mm` are meaningless without the frame they
|
||||
were measured in, so `_SIDECAR_SCHEMA_VERSION` is bumped whenever that frame
|
||||
changes. Each bump makes older files describe a different (and, for the bugs
|
||||
each bump fixed, actively wrong) transform than the same numbers would today;
|
||||
loading one unchanged would silently reproduce the very "scans show up
|
||||
everywhere" symptom the bump fixed — so older sidecars are treated as absent
|
||||
rather than migrated.
|
||||
|
||||
* **1 → 2** — pivot moved from the scan-window bbox center to a
|
||||
content-derived centroid, and the rotation sign convention was corrected.
|
||||
* **2 → 3** — the content centroid was abandoned entirely: rotation is now
|
||||
about each angle's own array center, mapped onto the reference's array
|
||||
center, with `shift_mm` in the reference's local mm frame. No angle but the
|
||||
reference contributes stage coordinates any more.
|
||||
Reference in New Issue
Block a user