Move design essays to docs/design.md, leave pointers

The three block essays in sras_compute.py — memory budget / row
chunking, angle-alignment coordinate frames, and the sidecar placement
+ schema history — move to docs/design.md (joined by a new section on
the zoom FFT peak search), each replaced by a 2-4 line pointer. The
save_manual_alignment docstring no longer restates the JSON schema its
own eight lines of code construct.

Load-bearing trap notes (normalization=None, subpixel-score mixing,
memmap lazy reads) stay in place. Zero code changes — golden-hash diff
verified empty.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Thomas Ales
2026-08-06 11:01:25 -05:00
parent f0f622b9ab
commit 8154c57066
3 changed files with 155 additions and 86 deletions
+17 -86
View File
@@ -276,22 +276,10 @@ def _peak_bins_direct(waves: np.ndarray, n_len: int) -> np.ndarray:
# Chunking / parallel budget
# ---------------------------------------------------------------------------
#
# Rows are batched so the float32 working buffers for one chunk stay under a
# memory budget. A fixed row count (the original design) works fine for small
# legacy scans but is catastrophic for a v6 scan with a large per-angle
# frame/sample count — e.g. a 7500-frame x 2500-sample angle needs ~2.4 GB for
# a single 32-row chunk.
#
# With chunks running concurrently the budget has to cover *all* live chunks at
# once. Note that on a large scan chunk_rows is already clamped to its floor of
# 1 row (one row alone is ~75 MB of float32 at 7507x2500), so shrinking the
# per-chunk size cannot buy more concurrency — the worker count must be derived
# from the budget instead. See _plan_chunks.
# 1024 MB is the measured knee on a 16-core machine against a 7507-frame x
# 2500-sample angle: 512 MB leaves ~20% of the FFT speedup on the table, and
# 1536+ MB costs ~0.4 GB more resident for no further gain. Override with
# SRAS_MEM_BUDGET_MB on a smaller machine.
# Row chunks are budgeted so all concurrently-live working buffers fit in
# memory; the worker count is derived from the budget, not vice versa.
# Rationale and the measured 1024 MB default: docs/design.md ("Memory budget
# and row chunking").
_TOTAL_BYTES_BUDGET = int(os.environ.get("SRAS_MEM_BUDGET_MB", 1024)) * 1024 * 1024
_CHUNK_ROWS_MAX = 32 # cap for small scans (original behavior)
_MAX_WORKERS = int(os.environ.get("SRAS_MAX_WORKERS", 0)) or (os.cpu_count() or 4)
@@ -659,41 +647,12 @@ def cache_file(path: str, mode: str, apply_bg_sub: bool,
# ---------------------------------------------------------------------------
# Angle alignment (Fusion menu)
#
# Puts every angle's images onto one shared, zero-padded pixel grid using a
# rigid transform only — rotation + translation, never scale.
#
# Angle 0 (the reference) is the sole coordinate authority: it is the only
# angle whose stage XY (x_start_mm / y_positions_mm) is ever read, and the
# shared canvas is literally an extension of angle 0's own pixel grid, so the
# aligned view carries angle 0's real X/Y axes. Every *other* angle is placed
# purely by content — its rotation and translation come from cross-correlating
# its CH4 image against angle 0's (register_angle_to_reference) — and its own
# stage XY is deliberately never consulted. That is not an oversight: the
# rotation stage moves the sample relative to the scan window, so where a
# window sat in stage coordinates says nothing about where the sample is, and
# an earlier design that pivoted each angle on a signal-weighted centroid of
# its own window put every angle on a ~20 mm circle around the optical center
# instead of stacking them into one shape.
#
# Only two coordinate frames exist here:
#
# local mm — one angle's own physical frame: origin at the *center of its own
# pixel array*, x along +column, y along +row, scaled by that angle's own
# pitches. Carries no stage position whatsoever.
#
# ref mm — the reference angle's local mm. A registration result
# (rotation_deg, shift_mm) is exactly the rigid map from an angle's local
# mm to ref mm: q = R(rotation_deg) @ l + shift_mm. Stage coordinates
# re-enter once, at the very end, when the canvas origin is converted to
# angle 0's stage mm (AlignmentResult.canvas_origin_mm).
#
# Rotation is done in mm, never on raw pixel indices: the x pitch
# (SrasFile.pixel_x_mm, 5 µm on a real scan) and the y/row pitch (50 µm) differ
# by 10x, so rotating the raw index grid would shear the image — an unwanted
# anisotropic scale. Registration runs on a resampled *isotropic* grid for the
# same reason, and every affine here maps shared-grid index -> mm -> undo
# rotation/shift -> that angle's own local mm -> that angle's own raw index,
# matching the output->input convention scipy.ndimage.affine_transform wants.
# Rigid transforms only (rotation + translation, never scale), computed in mm
# on two frames: each angle's "local mm" (origin at its own array center) and
# the reference angle's local mm. Angle 0 is the sole coordinate authority;
# every other angle is placed purely by image content. Why, and the full
# frame/affine conventions: docs/design.md ("Angle alignment coordinate
# frames").
# ---------------------------------------------------------------------------
@dataclass
@@ -1555,14 +1514,8 @@ def build_manual_alignment(sras: SrasFile, ref_angle_idx: int,
# ---- Sidecar persistence (<name>.sras.align.json) -------------------------
#
# Lives here, not sras_format.py: sras_format.py is scoped to the versioned
# binary .sras spec itself (see scan_format.md); a manual alignment is a
# viewer-computed *derived* artifact, analogous in kind to AlignmentResult —
# so it belongs with the alignment math it serialises, which already lives
# in this module. json + pathlib are both stdlib, so this doesn't add a new
# dependency to a module whose only load-bearing constraint is staying free
# of Qt/matplotlib for cheap multiprocessing-child imports.
# A viewer-computed derived artifact, so it lives with the alignment math
# rather than in sras_format (see docs/design.md, "Manual-alignment sidecar").
@dataclass
class ManualAlignmentSidecar:
@@ -1579,18 +1532,9 @@ def sidecar_path(sras_path) -> Path:
return p.with_name(p.name + ".align.json")
# The stored rotation_deg/shift_mm are meaningless without the frame they were
# measured in, so this is bumped whenever that frame changes. Each bump makes
# older files describe a different (and, for the bugs each bump fixed, actively
# wrong) transform than the same numbers would today, and loading one unchanged
# would silently reproduce the very "scans show up everywhere" symptom the bump
# fixed — so older sidecars are treated as absent rather than migrated.
# 1 -> 2 pivot moved from the scan-window bbox center to a content-derived
# centroid, and the rotation sign convention was corrected.
# 2 -> 3 the content centroid was abandoned entirely: rotation is now about
# each angle's own array center, mapped onto the reference's array
# center, with shift_mm in the reference's local mm frame. No angle
# but the reference contributes stage coordinates any more.
# Bumped whenever the frame the stored numbers are measured in changes; older
# sidecars are treated as absent, never migrated. Bump history:
# docs/design.md ("Schema history").
_SIDECAR_SCHEMA_VERSION = 3
@@ -1598,21 +1542,8 @@ def save_manual_alignment(sras: SrasFile, ref_angle_idx: int,
dc_threshold_mv: float,
per_angle: dict[int, ManualAngleParams]) -> Path:
"""Write the sidecar JSON for sras.path (overwriting any existing one)
and return the path written.
Schema (schema_version 3):
{
"schema_version": 3,
"ref_angle_idx": <int>,
"dc_threshold_mv": <float>,
"per_angle": {
"<angle_idx>": {"rotation_deg": <float>, "shift_mm": [<dx_mm>, <dy_mm>]},
...
}
}
Angle indices are JSON object keys, so they round-trip as strings —
load_manual_alignment converts them back to int.
"""
and return the path written. Angle indices become JSON object keys, so
they round-trip as strings — load_manual_alignment converts them back."""
path = sidecar_path(sras.path)
payload = {
"schema_version": _SIDECAR_SCHEMA_VERSION,