# PoC roadmap No phase has produced measured results yet. The repository establishes the proposed scope and design. | Phase | Work | Exit evidence | |---|---|---| | 0 — Define the inspection task | Confirm material variants, target defects, minimum size, observed surfaces, field of view and operator response | Agreed defect examples and acceptance protocol, including tolerated missed defects/false alarms and latency definition | | 1 — Establish visibility | Compare diffuse/grazing light, polarization, geometry and exposure on static and moving material | Versioned image set with settings; domain expert can identify target defects reproducibly | | 2 — Capture and catalogue | Implement recorded-image replay, then two-camera capture; SQLite/filesystem persistence and basic review UI | Traceable images, explicit missing-camera states, repeatable capture coverage, successful backup/restore | | 3 — Detection baseline | Compare classical CV with anomaly detection; calibrate thresholds and map regions to originals | Held-out results per defect class/size and production run; false alarms, missed defects and localization documented | | 4 — Integrated edge trial | Deploy chosen vision worker and live UI; benchmark sustained two-camera load and failures | Measured end-to-end latency against <60 s, queue stability, coverage, memory and storage; alarms and recovery demonstrated | | 5 — Optional descriptions | Evaluate a small VLM locally or on a separate available server | Reviewed description quality, measured resource use and proof that failure/overload cannot block detection | | 6 — Next-stage decision | Compare visibility, detection quality, usability and complete costs | Evidence-based decision to stop, adjust optics/data, upgrade compute or plan production integration | ## First software slice Implement replay of recorded images through a placeholder inspection interface, persistent metadata and a page showing two timestamped images. Use explicit “not evaluated” states until a detector exists. Add human review, then a measured baseline detector and camera adapters. Add dependency locking and installation instructions with that first executable implementation. Tests should cover real failure modes: wrong-image overlays, incomplete camera groups, inference/storage failure, restart after interrupted writes, duplicate retry events and review history. Use a small synthetic fixture set for software checks, clearly separated from real detection evaluation. ## Open decisions - Which Carbofol variants and defects matter, and what is the smallest relevant defect? - Do the two cameras observe one surface from two angles, opposite surfaces, or lateral areas? - What are working distance, usable along-web view, required overlap and position accuracy? - Is encoder/trigger access available, and what offsets separate observation points? - Does <60 s include durable storage, UI alarm and descriptions? Proposed baseline excludes optional descriptions. - What false-negative and false-positive rates are acceptable, using which evaluation units (region, frame, metre or physical defect)? - Which camera/interface/lens/lighting combination actually reveals the target defects? - Is a suitable existing computer available? Does a Nano meet measured resource needs? - What should alarms do, how are they acknowledged, and what happens on system failure? - How long are originals/normal samples retained; who reviews data and owns backups? - Which model/runtime versions and licences fit the intended use? Which project licence and remote hosting should the owner choose? ## Evidence discipline For each experiment record date, material/lot, rig configuration, dataset manifest, code/model versions, threshold, hardware/runtime settings and observed results. Report sample counts and limitations alongside metrics. Measure latency from actual timestamps including queueing, with percentiles and worst observed value under a stated workload. Neither a low average nor advertised TOPS proves the <60 s condition. Keep the evaluation set separate from threshold tuning and training. Retain negative samples and review missed defects, not only alarms. Promotions of human labels into a training dataset must be explicit and versioned.