Files
production-analytics/PROJECT_KNOWLEDGE.md
T

161 lines
8.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Project knowledge
## Purpose
This service is a calculation and persistence layer between ENLYZE and
Grafana. Its primary use case is to continuously integrate material consumption
for the currently active production order and expose its cumulative value in
Grafana while production is running. K7 fiber is the first validated
application, not the scope of the product. It must not become a second
historian: raw process data stays in ENLYZE and is retrieved again when
historical calculations need reproduction.
## Core architectural decisions
- Python with a `src/` package layout.
- FastAPI is the preferred future service layer; no HTTP API is needed now.
- PostgreSQL with TimescaleDB is the target derived-data store.
- Grafana reads derived data from that database as an additional datasource.
- Secrets are environment variables only; no credentials or site-specific
configuration are committed.
- Calculations are modular, testable Python implementations. Configuration
declares instances; it is not a generic low-code language.
- Historical production-order analysis/backfill is secondary. It uses the same
generic integration engine as live processing: one integration engine with
live incremental and historical replay/backfill operating modes.
## Material-consumption integration
The core calculation is generic rate integration: cumulative material
consumption increases by `material_rate * elapsed_time` when the configured
machine/process-specific validity or gating conditions are satisfied. It is
reusable across machines, materials, rate variables, units, and gates. The
project deliberately does not define a broader configuration schema here.
## Peak-cycle detector
The first generic calculation is a timestamp-based `PeakCycleDetector` under
`calculations`. It is transport-independent and works for roll length, roll
weight, and similar cyclic signals. It maintains the maximum in an open cycle;
after a value falls below `drop_ratio * current_peak`, it emits that maximum
only after actual subsequent below-threshold observations support a continuous
`hold_seconds` interval. `min_peak` is a generic detector setting, along with
`drop_ratio` (0 < ratio < 1), `hold_seconds` (>= 0), and an explicit positive
`max_sample_gap_seconds`. They are selected per signal, rather than being
globally fixed process constants. Elapsed time, never a sample count, determines
confirmation. An interval longer than the configured maximum sample gap
restarts a reset candidate, so a timestamp gap is not
continuous-below evidence. Reset values can be negative rather than zero,
equal timestamps add no elapsed time, and backwards timestamps are rejected.
The maximum gap is required rather than inferred or defaulted, making each
source's continuity assumption explicit.
K7 roll length, variable `bf81c547-dccf-4709-aee2-0f78366d1dfc` (m), is the
clean reference validation case. With `min_peak=10`, `drop_ratio=0.75`,
`hold_seconds=60`, and `max_sample_gap_seconds=20`, its continuous 10-second
capture produced exactly four plausible peaks of approximately 49.98 m. Its
shorter-than-requested result was caused by an end time queried in the future,
not observed time-series gaps.
Bento 2 roll weight is the second, more realistic robustness validation case:
machine `6201f3de-b940-45d7-930c-1481c5ae9e1d`, variable
`059c0015-9aff-4270-a4ce-f21b83f99378` (kg). Its observed cadence is 10
seconds, but it has long plateaus, reset values roughly 300–470 kg, and
materially varying peaks. With `min_peak=800`, `drop_ratio=0.75`,
`hold_seconds=60`, and `max_sample_gap_seconds=20`, the extended capture
yielded 13 plausible completed peaks: 936, 919, 935, 930, 930, 920, 909,
1195, 921, 930, 939, 912, and 914 kg (from 00:46:50Z through 03:06:30Z on
2026-09-04). The approximately 1195 kg peak is retained as a real
process/mechanical condition, not special-cased as a defect; it is emitted
only after its following reset has enough below-threshold observations. A
reset is not expected to approach zero, and peak values are not assumed to be
constant. Raw captures remain ignored and deterministic tests use synthetic
fixtures only.
## Central domain context
Production orders connect machine, article/material, source interval, and
derived results. An ENLYZE Production Run is a time segment; one
`production_order` can have one or more Production Runs. Integrate each run
separately, then aggregate derived values on the production-order level. Never
implicitly bridge gaps between runs. Metrics and events must retain calculation
type/version and source-time-range provenance.
## K7 fiber integration
The current K7 candidate fiber mass-flow signal is `Stundenleistung Anlage`
in kg/h. Observed behaviour supports treating it as a process-responsive
band-scale signal, but it can remain frozen at a non-zero value during a
machine stop. It must therefore be integrated only while the validated gate
`Geschwindigkeit Gesamtanlage > 0.5 m/min` is active. The ENLYZE boolean
`Anlage läuft` is not the primary integration gate.
The process-derived result is `fiber_feed_kg` for a Production Run. It is not
automatically a finished-product mass or a per-order material yield. Material
measured at the band scale may still be in the machine at an order transition,
and ERP feedback can arrive asynchronously or periodically. Differences from
reported finished-product quantities must not automatically be labelled scrap;
an exact per-order yield needs WIP/transport-delay accounting.
## Bento 1 bentonite-powder integration
Bento 1 bentonite-powder consumption is the next planned application of the
same generic integration mechanism. Its ENLYZE source variable and its
validity/gating logic have not yet been selected or validated; both must be
determined from Bento 1 process data before implementation. Bento 1 does not
need a separate subsystem or calculation formula.
## Quantity-total interpretation
For this customer setup, `production_run.quantity_total` is populated from
ERP/BI production feedback received by ENLYZE, rather than necessarily being
calculated from machine signals. The tested source CSV field is
`PCO_FEEDB_QUANTITY`. For inspected machines/articles it represents reported
finished-product area in m² (for example, 5.00 m × 50 m = 250.00 m²; 4.75 m ×
50 m = 237.50 m²; 5.80 m × 50 m = 290.00 m²). The unit displayed by ENLYZE is
separately configured and must not be considered authoritative until its
configuration is validated.
## Known versus unknown
Validated K7 historical tests show physically plausible fiber-feed totals when
the band-scale kg/h signal is speed-gated. For production order
`K 7-12026000842`, its run from 2026-08-31T10:57:54Z to
2026-08-31T18:23:11Z had ERP `quantity_total` of 12,420 m², approximately
5.206 active hours, approximately 5,180.1 kg integrated fiber feed,
approximately 2,174.2 m integrated machine length, and approximately
995.1 kg/h average active throughput. Article external ID `213408` has width
6.0 m and static mean basis weight 390.16 g/m²; the corresponding approximate
reported finished-product mass is 4,845.8 kg. This is a plausibility validation
case, not a precise material-yield calculation.
Known from the ENLYZE UI/API exploration: machine and product identity,
Production Runs, and relevant K7 process signals are available. Some API and
time-series semantics remain unverified; do not infer them beyond documented
evidence. The active evidence and open-question list is in
[docs/enlyze-api.md](docs/enlyze-api.md).
## Future refinement
An optional future historical mass-balance refinement is to obtain actual
QA/QS basis-weight measurements from SQL instead of using static article basis
weight. It is not a blocker for the live material-consumption MVP.
## Exploration workflow
A deliberately generic, read-only CLI is available as
`production-analytics enlyze raw PATH`. It only makes GET requests, never
guesses endpoint schemas, and emits sanitized response metadata/body. The
operator must first obtain an authorized base URL, a verified safe path, and
the authentication method. Credentials reside in ignored `secrets/enlyze.env`;
the CLI parses simple assignments without sourcing/executing the file. Public
ENLYZE documentation verifies an `Authorization: Bearer <ENLYZE_API_KEY>`
authentication header; its value is never printed.
Unmodified captures are local in ignored `data/raw/enlyze/`. Fixtures written
to `fixtures/enlyze/` are sanitized, but must still be reviewed before commit.
The exploration deliverable is verified API observations and sanitized
fixtures. The next product deliverable is live incremental material
integration, using a generic material-rate integrator rather than a separate
historical-only calculation path.