Files
machine-vision-poc/docs/data-model.md
T

62 lines
6.2 KiB
Markdown

# Proposed SQLite and filesystem data model
This is a logical design, not an installed schema or migration. Use stable UUID-style text IDs, UTC timestamps and explicit units. Enable SQLite foreign keys on every connection. Store original images once; reference them from inspections and reviews.
## Entities
| Entity | Key fields and relationships |
|---|---|
| `production_runs` | `id`, product/surface variant, lot, start/end timestamps, notes |
| `capture_configs` | `id`, immutable camera/lighting settings, calibration version/hash, trigger mode, preprocessing configuration |
| `capture_groups` | `id`, `production_run_id`, trigger ID, expected camera IDs, group status, optional `position_m`, position source (`encoder`, `estimated`, `unknown`) |
| `captures` | `id`, `capture_group_id`, camera ID, captured timestamp, `capture_config_id`, original asset ID, width/height in pixels, acquisition status/error |
| `model_versions` | `id`, task, model name, artifact hash, runtime and preprocessing versions, configuration snapshot |
| `inspections` | `id`, `capture_id`, `model_version_id`, start/completion timestamps, threshold, nullable score and anomaly decision, status (`pending`, `complete`, `error`), error text |
| `defects` | `id`, `inspection_id`, bounding box, region score, optional mask asset ID, predicted class and nullable classifier confidence; one row per suspected region |
| `descriptions` | `id`, `defect_id`, model/prompt version, generated text, status, created/completed timestamps, error; machine suggestions only |
| `reviews` | `id`, `defect_id`, reviewer, timestamp, verdict (`confirmed`, `false_positive`, `uncertain`), human class ID, comment, optional superseded review ID |
| `defect_classes` | `id`, stable code, display label, description, active flag |
| `assets` | `id`, relative path, kind (`original`, `crop`, `overlay`, `mask`), checksum, byte size, dimensions, storage status (`staging`, `ready`, `missing`, `deleted`) |
| `defect_assets` | `defect_id`, `asset_id`, role; links crops/overlays/masks without duplicating original images |
| `alarms` | `id`, `inspection_id`, created timestamp, state, acknowledged timestamp and user; acknowledgement does not imply review |
Relationships: one run has many groups; a group has one capture per expected camera, including failed acquisition records. One capture can have many inspection attempts/model versions. One inspection can produce zero or more suspected regions. A defect can have many descriptions and append-only reviews.
A missing camera or failed inspection never becomes a normal result. Keep score and anomaly decision null until successful evaluation. Enforce uniqueness of group/camera pairs and asset paths. Use indexes on group/run, capture/group, inspection/capture, defect/inspection and review/defect/time; add catalogue query indexes when queries exist.
## Coordinates and evidence
Use bounding boxes `(x, y, width, height)` in original-image pixels, top-left origin, positive size and bounds inside the associated image. Map model output back through preprocessing. Store mask dimensions and transforms explicitly. Physical size is nullable and requires recorded calibration; do not infer millimetres or surface height from generated text.
A defect row is a **camera observation**, not automatically a distinct physical defect. Overlapping frames and complementary views may duplicate the same event. If useful later, add a reviewed `physical_events` table plus observation links, recording association method and confidence. Do not fuse the two surfaces or deduplicate by timestamp alone.
## File layout and consistency
```text
<data-root>/
inspection.sqlite3
images/<run-id>/<group-id>/<camera-id>/<capture-id>.png
derived/<inspection-id>/<defect-id>/crop.png
derived/<inspection-id>/<defect-id>/overlay.png
derived/<inspection-id>/<defect-id>/mask.png
exports/<export-id>/manifest.json
```
Paths in SQLite are relative to the configured data root. PNG is an illustrative lossless format; choose format and bit depth against camera data and storage tests. Avoid moving originals when a review changes the verdict. Derived evidence should remain associated with the model that generated it.
Write assets to temporary paths on the same filesystem, finalize with atomic rename, then commit ready asset references and associated results in a short database transaction. Failures between steps may leave orphan files; reconcile these explicitly after restart. Never publish success with a broken asset reference. Use stable job IDs for retries to avoid duplicate alarms/records.
Define retention separately for defect originals, reviewed normal samples, unreviewed data and regenerable previews. Keep enough normal material to evaluate false negatives and drift. Record deletions rather than leaving unexplained broken paths. Back up SQLite consistently along with its referenced files; account for WAL state and test restoration. Do not copy only the live main database file and assume a complete backup.
## Human-in-the-loop catalogue
1. Persist a suspected region and any machine label; new observations are unreviewed.
2. Present original context, crop, overlay, score and generated description separately.
3. The reviewer confirms, rejects as false positive, or marks uncertain; they can correct the class and add comments.
4. Append a review with identity and time. Derive current review status from the latest non-superseded review; preserve previous decisions and machine output.
5. Export a versioned dataset manifest listing evidence hashes, model versions and approved labels. Exclude uncertain observations from ground truth unless explicitly resolved.
Candidate classes from the discussion: material accumulation, foreign body, fold, hole, surface-structure anomaly and other/unknown. These require domain review. “No defect” is a review verdict, not a physical defect category. An unreviewed model label or generated description is not ground truth.
Allow reviewers to inspect normal captures and add missed regions through a future manual annotation workflow. Detection-only review cannot measure missed defects. Split training and evaluation by production run/lot and linked physical event to prevent overlapping images of the same material leaking across splits. Define representative held-out ground truth before reporting accuracy.