Files
scanengine-3/core/sras_format.py
T
Thomas Ales 844fcd0297 SAW quality check: one middle row per angle, and a viewer that overlays them
A full multi-angle scan takes hours, and a rig whose angles disagree produces
all of them before anyone finds out. This adds a test mode that acquires one
row per angle — the row-wise middle of the ROI — and a viewer that puts every
angle's SAW frequency on one graph. The default 80×50 mm ROI at 5 angles goes
from 1461 rows to 5.

Why the middle row answers an alignment question at all: build_plan centres
every angle's rotated bounding box on the same nominal ROI centre, so each
angle's middle row crosses that one point on the sample. All the angles
measure the same material, so a spread in their frequencies belongs to the rig
rather than to where each row happened to land. test_every_angles_middle_row_
crosses_the_roi_centre pins that premise, since the whole comparison rests on
it and nothing else in the geometry code would notice it breaking.

core/saw_check.py — both halves of the mode, kept together because neither is
much use alone. middle_row_plan() reduces a ScanPlan to one row per angle
(n_rows // 2, the upper of two centre rows when even); frequency_traces() and
alignment_summary() turn the resulting file back into per-angle frequency
traces and the scalars an operator is actually asking about — the spread of
the per-angle medians, the worst drift along a row, the sparsest row. The
verdict thresholds are labelled as rules of thumb, not physics: an anisotropic
sample genuinely varies with angle, so a wide spread is a prompt to look at
the curves rather than a verdict.

Format v10: byte-identical to v6, one row per angle. The version byte earns
its keep because the two are otherwise indistinguishable — a v6 scan aborted
after its first row is not a check, and a reader guessing from the row count
would read a failed scan as a deliberate measurement. create_scan_file()
enforces the one-row rule at write time, since nothing downstream can recover
from a v10 file that breaks it. ScanEngine gains file_version and is otherwise
untouched: the acquisition, the abort/pause path and the background capture
are the scan's, unchanged.

sras_scan_manager.py now carries the source file's version through an export
instead of stamping v6 on everything, which the wider reader would otherwise
have made a lie.

saw_check_viewer.py — frequency along the row, one curve per angle, over a
common offset axis so the curves lie on the same piece of sample; a summary of
each angle's median ±1σ against angle; and the per-angle numbers in a table.
Analysis parameters (DC threshold, background, time gate) recompute on a
worker thread; display ones (smoothing, axis, MHz↔m/s) only redraw. A full v6
scan opens too — the same middle row is pulled out of it — so a finished scan
can be re-examined with the check's own read-out.

In the app, a check finishes by handing the operator the file and an "Open
Viewer" button rather than shutting the rig down the way a completed scan
does. Burst mode is not offered: one row per angle means every burst would be
a single row, so it buys nothing and still pays for the gate preflight.

137 tests passing, ruff clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-04 08:13:54 -05:00

369 lines
15 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""SRAS binary scan-file format (v6 and v10) — the single implementation.
Full byte-level spec: scan_format.md. Summary:
header >4sBHfffffffIdBB magic ver n_angles xs_nom ys_nom xd_nom yd_nom
row_spacing velocity laser_freq spf sample_rate
bytes_per_sample n_channels
angle table n_angles × >f
geometry table n_angles × >ffIH (x_start x_delta n_frames n_rows)
row tables (ragged) per angle: n_rows × >f (y positions)
preambles n_channels × (>H length + utf-8 WFMOutpre string)
background block >I length + raw int8 CH1 average
waveform data angle-major, row-minor, channel-inner:
for each angle, for each row, for each channel,
n_frames × samples_per_frame × bytes_per_sample
Incomplete files are valid: the data block is one contiguous append-only
stream, so the readable prefix defines a single frontier past which nothing
has been written yet (see ``SrasFile.angle_status``).
Version 10 is the SAW quality check (core.saw_check): byte layout identical
to v6, but every angle declares exactly one row — the row-wise middle of the
ROI. The version byte is the whole difference, and it exists so a reader can
tell a one-row-per-angle check from a full scan that was aborted after its
first row. ``create_scan_file`` enforces the one-row rule at write time.
"""
from __future__ import annotations
import mmap
import struct
from dataclasses import dataclass, field
from pathlib import Path
from typing import BinaryIO
import numpy as np
from core.scan_geometry import AngleGeometry, ScanPlan
MAGIC = b"SRAS"
VERSION = 6
# One row per angle, taken from the middle of the ROI — see core.saw_check.
VERSION_SAW_CHECK = 10
SUPPORTED_VERSIONS = (VERSION, VERSION_SAW_CHECK)
HDR_FMT = ">4sBHfffffffIdBB"
HDR_SIZE = struct.calcsize(HDR_FMT) # 49 bytes
GEOM_FMT = ">ffIH"
GEOM_SIZE = struct.calcsize(GEOM_FMT) # 14 bytes
# Oscilloscope channels recorded, in on-disk order.
SCAN_CHANNELS = [1, 3, 4]
STATUS_OK = "OK"
STATUS_TRUNCATED = "TRUNCATED"
STATUS_MISSING = "MISSING"
@dataclass
class ScanHeader:
"""The fixed v6 global header (everything but magic/version)."""
n_angles: int
x_start_nominal: float
y_start_nominal: float
x_delta_nominal: float
y_delta_nominal: float
row_spacing: float
velocity: float
laser_freq: float
samples_per_frame: int
sample_rate: float
bytes_per_sample: int
n_channels: int
@dataclass
class AngleStatus:
"""How much of one angle's declared data is actually on disk."""
index: int
angle_deg: float
n_rows: int # declared
row_bytes: int
data_offset: int
n_rows_available: int
status: str # STATUS_OK / STATUS_TRUNCATED / STATUS_MISSING
@property
def complete(self) -> bool:
return self.status == STATUS_OK
def create_scan_file(path: Path, plan: ScanPlan, samples_per_frame: int,
sample_rate: float, preambles: list[str],
background_waveform: bytes,
version: int = VERSION) -> BinaryIO:
"""Create a new .sras file and write the header + tables.
``version`` selects which kind of file this is — VERSION for a full scan,
VERSION_SAW_CHECK for a middle-row quality check. The layout is the same
either way; the one-row-per-angle rule that gives v10 its meaning is
checked here, since nothing downstream can recover from a v10 file that
breaks it.
Returns an open binary file positioned at the start of the data block;
the caller appends waveform rows and must close it (try/finally).
"""
if version not in SUPPORTED_VERSIONS:
raise ValueError(
f"Cannot write SRAS format version {version} "
f"(supported: {', '.join(str(v) for v in SUPPORTED_VERSIONS)})"
)
if version == VERSION_SAW_CHECK:
bad = [f"{pa.angle_deg:.1f}° has {pa.n_rows}"
for pa in plan.per_angle if pa.n_rows != 1]
if bad:
raise ValueError(
"A v10 SAW-check file holds exactly one row per angle, but "
+ ", ".join(bad) + " — build the plan with "
"core.saw_check.middle_row_plan()."
)
path.parent.mkdir(parents=True, exist_ok=True)
f = open(path, "wb")
f.write(struct.pack(
HDR_FMT, MAGIC, version,
plan.n_angles,
plan.x_start_nominal, plan.y_start_nominal,
plan.x_delta_nominal, plan.y_delta_nominal,
plan.row_spacing,
plan.velocity_mm_s, plan.laser_freq_hz,
samples_per_frame,
sample_rate,
1, # bytes_per_sample: int8 from scope default
len(SCAN_CHANNELS),
))
f.write(struct.pack(f">{plan.n_angles}f", *plan.angles))
for pa in plan.per_angle:
f.write(struct.pack(GEOM_FMT, pa.x_start, pa.x_delta, pa.n_frames, pa.n_rows))
for pa in plan.per_angle:
f.write(struct.pack(f">{pa.n_rows}f", *pa.y_positions))
for p in preambles:
enc = p.encode("utf-8")
f.write(struct.pack(">H", len(enc)))
f.write(enc)
f.write(struct.pack(">I", len(background_waveform)))
f.write(background_waveform)
return f
@dataclass
class SrasFile:
"""Parsed .sras file (v6 or v10): header, tables, and lazy (memmap) access.
Parsing reads only the header/tables — never the waveform block — so
opening a multi-GB file is cheap. ``load_angle``/``load_row`` return
read-only numpy views backed by a shared mmap; no data is copied until
the caller computes on it.
"""
path: Path
version: int = field(init=False)
header: ScanHeader = field(init=False)
per_angle: list[AngleGeometry] = field(init=False)
preambles: list[str] = field(init=False)
preambles_raw: list[bytes] = field(init=False)
background: bytes = field(init=False)
data_start_offset: int = field(init=False)
file_size: int = field(init=False)
def __post_init__(self):
self.path = Path(self.path)
self._mmap: mmap.mmap | None = None
self._parse()
def _parse(self):
self.file_size = self.path.stat().st_size
with open(self.path, "rb") as f:
raw = f.read(HDR_SIZE)
if len(raw) < HDR_SIZE:
raise ValueError(f"{self.path.name}: file too short to contain a valid header")
(magic, version, n_angles, x_start_nominal, y_start_nominal,
x_delta_nominal, y_delta_nominal, row_spacing, velocity, laser_freq,
samples_per_frame, sample_rate, bytes_per_sample,
n_channels) = struct.unpack(HDR_FMT, raw)
if magic != MAGIC:
raise ValueError(f"{self.path.name}: not a valid SRAS file (bad magic)")
if version not in SUPPORTED_VERSIONS:
raise ValueError(
f"{self.path.name}: unsupported SRAS format version {version} "
f"(supported: {', '.join(str(v) for v in SUPPORTED_VERSIONS)})"
)
self.version = version
self.header = ScanHeader(
n_angles=n_angles,
x_start_nominal=x_start_nominal, y_start_nominal=y_start_nominal,
x_delta_nominal=x_delta_nominal, y_delta_nominal=y_delta_nominal,
row_spacing=row_spacing, velocity=velocity, laser_freq=laser_freq,
samples_per_frame=samples_per_frame, sample_rate=sample_rate,
bytes_per_sample=bytes_per_sample, n_channels=n_channels,
)
angles = struct.unpack(f">{n_angles}f", f.read(4 * n_angles))
self.per_angle = []
for a in angles:
x_start, x_delta, n_frames, n_rows = struct.unpack(GEOM_FMT, f.read(GEOM_SIZE))
self.per_angle.append(AngleGeometry(
angle_deg=a, x_start=x_start, x_delta=x_delta,
n_frames=n_frames, n_rows=n_rows,
))
for pa in self.per_angle:
pa.y_positions = list(struct.unpack(f">{pa.n_rows}f", f.read(4 * pa.n_rows)))
self.preambles_raw = []
for _ in range(n_channels):
(plen,) = struct.unpack(">H", f.read(2))
self.preambles_raw.append(f.read(plen))
self.preambles = [p.decode("utf-8", errors="replace") for p in self.preambles_raw]
(n_bg,) = struct.unpack(">I", f.read(4))
self.background = f.read(n_bg)
self.data_start_offset = f.tell()
@property
def is_saw_check(self) -> bool:
"""True for a v10 middle-row SAW quality check rather than a scan."""
return self.version == VERSION_SAW_CHECK
# ── Frontier / truncation analysis ───────────────────────────────────────
def row_bytes(self, angle_idx: int) -> int:
pa = self.per_angle[angle_idx]
return (self.header.n_channels * pa.n_frames
* self.header.samples_per_frame * self.header.bytes_per_sample)
def angle_status(self) -> list[AngleStatus]:
"""Walk declared per-row byte counts against the actual file size.
Because the data is one contiguous append-only stream, once an angle
is found short every later angle is necessarily absent too — there is
a single frontier past which nothing has been written yet.
"""
statuses = []
cursor = self.data_start_offset
frontier_seen = False
for ai, pa in enumerate(self.per_angle):
row_bytes = self.row_bytes(ai)
data_offset = cursor
if frontier_seen:
n_rows_available = 0
status = STATUS_MISSING
else:
declared_bytes = row_bytes * pa.n_rows
if row_bytes > 0 and cursor + declared_bytes <= self.file_size:
n_rows_available = pa.n_rows
status = STATUS_OK
cursor += declared_bytes
else:
remaining = max(0, self.file_size - cursor)
n_rows_available = remaining // row_bytes if row_bytes > 0 else 0
status = STATUS_MISSING if n_rows_available == 0 else STATUS_TRUNCATED
frontier_seen = True
statuses.append(AngleStatus(
index=ai, angle_deg=pa.angle_deg, n_rows=pa.n_rows,
row_bytes=row_bytes, data_offset=data_offset,
n_rows_available=n_rows_available, status=status,
))
return statuses
def angle_data_offset(self, angle_idx: int) -> int:
offset = self.data_start_offset
for ai in range(angle_idx):
offset += self.row_bytes(ai) * self.per_angle[ai].n_rows
return offset
# ── Lazy data access ─────────────────────────────────────────────────────
def _ensure_mmap(self) -> mmap.mmap:
if self._mmap is None:
# The mapping stays valid after the file object is closed, so
# don't hold the descriptor open for the (long) life of a viewer
# session.
with open(self.path, "rb") as f:
self._mmap = mmap.mmap(f.fileno(), 0, access=mmap.ACCESS_READ)
return self._mmap
def _dtype(self) -> np.dtype:
return np.dtype(np.int16 if self.header.bytes_per_sample == 2 else np.int8)
def load_angle(self, angle_idx: int, n_rows: int | None = None) -> np.ndarray:
"""Read-only view of one angle's data block, shape
(n_rows, n_channels, n_frames, samples_per_frame).
``n_rows`` limits the view to the rows actually on disk (pass
``AngleStatus.n_rows_available`` for truncated files); default is the
declared row count.
"""
pa = self.per_angle[angle_idx]
h = self.header
if n_rows is None:
n_rows = pa.n_rows
start = self.angle_data_offset(angle_idx)
count = n_rows * h.n_channels * pa.n_frames * h.samples_per_frame
arr = np.frombuffer(self._ensure_mmap(), dtype=self._dtype(),
count=count, offset=start)
arr = arr.reshape(n_rows, h.n_channels, pa.n_frames, h.samples_per_frame)
arr.flags.writeable = False
return arr
def load_row(self, angle_idx: int, row: int, channel_idx: int) -> np.ndarray:
"""Read-only view of one row/channel, shape (n_frames, samples_per_frame)."""
pa = self.per_angle[angle_idx]
h = self.header
ch_bytes = pa.n_frames * h.samples_per_frame * h.bytes_per_sample
start = (self.angle_data_offset(angle_idx) + row * self.row_bytes(angle_idx)
+ channel_idx * ch_bytes)
arr = np.frombuffer(self._ensure_mmap(), dtype=self._dtype(),
count=pa.n_frames * h.samples_per_frame, offset=start)
arr = arr.reshape(pa.n_frames, h.samples_per_frame)
arr.flags.writeable = False
return arr
def close(self):
"""Release this file's hold on the mapping.
Views handed out earlier stay valid — they keep the mapping alive
until they are garbage-collected, at which point the OS frees it.
"""
if self._mmap is not None:
try:
self._mmap.close()
except BufferError:
pass # live numpy views still reference the buffer
self._mmap = None
def __enter__(self):
return self
def __exit__(self, exc_type, exc_val, exc_tb):
self.close()
# ── Axes helpers (viewer conveniences, derived from header fields) ───────
def pixel_pitch_x_mm(self) -> float:
"""Distance between adjacent frames along X."""
return self.header.velocity / self.header.laser_freq
def x_axis_mm(self, angle_idx: int) -> np.ndarray:
pa = self.per_angle[angle_idx]
return pa.x_start + np.arange(pa.n_frames) * self.pixel_pitch_x_mm()
def time_axis_ns(self) -> np.ndarray:
h = self.header
return np.arange(h.samples_per_frame) / h.sample_rate * 1e9
def freq_axis_mhz(self, nfft: int) -> np.ndarray:
return np.fft.rfftfreq(nfft, d=1.0 / self.header.sample_rate) / 1e6
def plan_from_header(sras: SrasFile) -> ScanPlan:
"""Reconstruct the ScanPlan a file was written with (for resume)."""
h = sras.header
return ScanPlan(
x_start_nominal=h.x_start_nominal, y_start_nominal=h.y_start_nominal,
x_delta_nominal=h.x_delta_nominal, y_delta_nominal=h.y_delta_nominal,
row_spacing=h.row_spacing,
velocity_mm_s=h.velocity, laser_freq_hz=h.laser_freq,
per_angle=list(sras.per_angle),
)