56 lines
7.0 KiB
Markdown
56 lines
7.0 KiB
Markdown
# Proposed architecture
|
||
|
||
All components below are a design proposal, not an implemented or measured system. The shared architecture is intended for multiple machine-vision use cases; the camera geometry, coverage calculations and material experiments here describe the initial Carbofol use case. Keep material-specific settings, labels and model versions in explicit configurations rather than hard-coding them into the application.
|
||
|
||
## Cameras and illumination
|
||
|
||
Compare two arrangements before fixing the rig:
|
||
|
||
1. Two views of the same surface: approximately perpendicular with diffuse illumination, plus an oblique view with grazing light. This was the latest discussion preference for complementary evidence.
|
||
2. One camera per membrane surface if both surfaces require inspection. Confirm whether “each side” means opposite surfaces or lateral regions. Each arrangement needs its own coverage and calibration plan.
|
||
|
||
Experiment with diffuse light, grazing light from either direction, polarization and shielding from ambient light. Record the lighting recipe with every capture. Projected lines are a later option if ordinary illumination cannot reveal the target geometry; depth measurement requires calibration and is not implied by seeing a distorted line.
|
||
|
||
Use representative material, normal texture and known defects. Fix working distance, field of view, focus, exposure and gain. Machine-vision cameras with manual controls, suitable lenses and hardware triggering are candidates; existing webcams/action cameras can support initial visibility experiments. Check Linux/ARM driver support, sustained dual-camera transfer and power requirements before selecting interfaces.
|
||
|
||
## Coverage and triggering
|
||
|
||
4 m/min equals about 66.7 mm/s. At 1/500 s exposure, calculated travel is about 0.133 mm; whether that blur is acceptable depends on the smallest relevant defect. These are geometric calculations, not measurements.
|
||
|
||
The specified 30–50 cm is cross-web width. Along-web field of view is still unknown. For speed `v`, usable along-web view length `L` and overlap fraction `o`, a gap-free nominal schedule needs `interval <= L * (1-o) / v`; include margins for jitter and unusable image edges. The earlier example of one capture per 250 mm gives 3.75 s at maximum speed, but only works if usable coverage and overlap permit it. A proposed 1–2 fps per camera is an experiment, not a coverage guarantee.
|
||
|
||
Start with timestamped acquisition; later prefer encoder-triggered groups for repeatable production positions. Store actual per-camera timestamps, trigger identifiers and incomplete groups. Apply physical camera offsets when relating images to the same material. Do not assume that simultaneous frames show the same location. Without an encoder, mark estimated positions explicitly or leave them unknown.
|
||
|
||
## Processing and edge AI
|
||
|
||
- Capture adapters assign stable image/group IDs and save original evidence with camera settings.
|
||
- Preprocessing applies versioned regions of interest, calibration and normalization; retain the mapping back to original pixels.
|
||
- A vision worker starts with classical contrast/texture baselines and an anomaly-detection candidate such as Anomalib. Training uses reviewed normal material. Compare later supervised detection/segmentation once sufficient labels exist.
|
||
- Thresholded scores produce suspected anomaly regions, boxes and/or masks. An anomaly score is not a calibrated probability or a confirmed defect class.
|
||
- Persist the inspection outcome and enqueue its UI event. Review and alarm operate on that result immediately.
|
||
- An optional bounded VLM job receives selected crops and verified metadata. A small quantized 2–4B candidate was discussed, but fit, image support, quality and latency must be tested on the actual runtime. Preserve prompt/model versions and distinguish generated suggestions from human labels.
|
||
|
||
Use a bounded inference queue and expose queue age. Never silently discard required inspection frames: overload, missing cameras or failed inference produce an explicit incomplete/error condition. Preview frames may be replaced by newer ones without implying inspection coverage. Prioritize detection over VLM work; disable or offload descriptions to an available LAN server if resources are insufficient. No server availability is assumed.
|
||
|
||
Separate training from online inspection. Record the model artifact hash, preprocessing configuration and threshold for reproducible replay. Verify runtime/export compatibility on the selected device instead of assuming every model supports TensorRT.
|
||
|
||
## Backend and frontend
|
||
|
||
Proposed stack: Python with FastAPI; server-rendered HTML/HTMX initially, with polling or SSE for updates. Choose the simplest working transport during implementation. Serve two latest images with timestamps, capture IDs, processing state and per-image overlays. Also show the latest completed inspection, pending work, camera health, storage health and alarm state.
|
||
|
||
Never apply an older result's overlay to a newer image. “No anomaly detected”, “not inspected yet” and “inspection failed” must remain distinct. An alarm should remain visible until acknowledged; acknowledgement is separate from accepting or rejecting the detection. An initial alarm is visual in the UI; audible/physical outputs and escalation rules are open decisions.
|
||
|
||
The catalogue supports filtering by run, time, camera, review status and class, with crops, originals and correction controls. Generated text must be visibly marked as unreviewed. Persist events before publishing updates; after reconnection, the client reloads authoritative state from the backend.
|
||
|
||
## SQLite and filesystem
|
||
|
||
Use SQLite for relationships and metadata, with images and masks stored as files; see [data model](data-model.md). Keep transactions short and serialize writes initially. Evaluate WAL mode on local storage when concurrent reads are introduced. A separate database server is unnecessary for the proposed single-device PoC, but revisit this if multiple hosts write concurrently.
|
||
|
||
Use atomic file finalization and record failure states because filesystem writes and database commits are not one transaction. Monitor free space, define retention and test coherent backups of both images and metadata. On storage failure, report degraded inspection and preserve the distinction between detection and successful evidence storage.
|
||
|
||
## Timing and validation
|
||
|
||
Proposed interpretation of the <60 s requirement: capture timestamp through completed evaluation, durable result and visible alarm when applicable, including queueing. Confirm this definition with the operator. Measure optional description completion separately; its deadline is undecided. Earlier <1 s detection and 10–20 s description figures were aspirations, not acceptance commitments or results.
|
||
|
||
Measure full-pipeline latency, backlog, dropped/incomplete captures, memory, storage growth and false alarms under sustained two-camera load. Low frame rate alone does not establish that the target hardware is sufficient. At maximum speed the membrane travels 4 m in 60 s; tracking and operator response must account for this displacement.
|