Files
machine-vision-poc/docs/roadmap.md
T

40 lines
4.2 KiB
Markdown

# PoC roadmap
No phase has produced measured results yet. The repository establishes the proposed scope and design.
| Phase | Work | Exit evidence |
|---|---|---|
| 0 — Define the inspection task | Confirm material variants, target defects, minimum size, observed surfaces, field of view and operator response | Agreed defect examples and acceptance protocol, including tolerated missed defects/false alarms and latency definition |
| 1 — Establish visibility | Compare diffuse/grazing light, polarization, geometry and exposure on static and moving material | Versioned image set with settings; domain expert can identify target defects reproducibly |
| 2 — Capture and catalogue | Implement recorded-image replay, then two-camera capture; SQLite/filesystem persistence and basic review UI | Traceable images, explicit missing-camera states, repeatable capture coverage, successful backup/restore |
| 3 — Detection baseline | Compare classical CV with anomaly detection; calibrate thresholds and map regions to originals | Held-out results per defect class/size and production run; false alarms, missed defects and localization documented |
| 4 — Integrated edge trial | Deploy chosen vision worker and live UI; benchmark sustained two-camera load and failures | Measured end-to-end latency against <60 s, queue stability, coverage, memory and storage; alarms and recovery demonstrated |
| 5 — Optional descriptions | Evaluate a small VLM locally or on a separate available server | Reviewed description quality, measured resource use and proof that failure/overload cannot block detection |
| 6 — Next-stage decision | Compare visibility, detection quality, usability and complete costs | Evidence-based decision to stop, adjust optics/data, upgrade compute or plan production integration |
## First software slice
Implement replay of recorded images through a placeholder inspection interface, persistent metadata and a page showing two timestamped images. Use explicit “not evaluated” states until a detector exists. Add human review, then a measured baseline detector and camera adapters. Add dependency locking and installation instructions with that first executable implementation.
Tests should cover real failure modes: wrong-image overlays, incomplete camera groups, inference/storage failure, restart after interrupted writes, duplicate retry events and review history. Use a small synthetic fixture set for software checks, clearly separated from real detection evaluation.
## Open decisions
- Which Carbofol variants and defects matter, and what is the smallest relevant defect?
- Do the two cameras observe one surface from two angles, opposite surfaces, or lateral areas?
- What are working distance, usable along-web view, required overlap and position accuracy?
- Is encoder/trigger access available, and what offsets separate observation points?
- Does <60 s include durable storage, UI alarm and descriptions? Proposed baseline excludes optional descriptions.
- What false-negative and false-positive rates are acceptable, using which evaluation units (region, frame, metre or physical defect)?
- Which camera/interface/lens/lighting combination actually reveals the target defects?
- Is a suitable existing computer available? Does a Nano meet measured resource needs?
- What should alarms do, how are they acknowledged, and what happens on system failure?
- How long are originals/normal samples retained; who reviews data and owns backups?
- Which model/runtime versions and licences fit the intended use? Which project licence and remote hosting should the owner choose?
## Evidence discipline
For each experiment record date, material/lot, rig configuration, dataset manifest, code/model versions, threshold, hardware/runtime settings and observed results. Report sample counts and limitations alongside metrics. Measure latency from actual timestamps including queueing, with percentiles and worst observed value under a stated workload. Neither a low average nor advertised TOPS proves the <60 s condition.
Keep the evaluation set separate from threshold tuning and training. Retain negative samples and review missed defects, not only alarms. Promotions of human labels into a training dataset must be explicit and versioned.