# Architecture ## Scope boundary ENLYZE is authoritative for raw process data. This service retrieves raw data only to calculate derived results and persists the results, their context, and the minimum state needed for incremental calculation. It does not mirror raw time series into PostgreSQL/TimescaleDB. ## Product direction The primary product path is live integration of material consumption for the currently active production order, with a continuously updated cumulative value for Grafana. The core mechanism is generic rate integration: cumulative material consumption increases by `material_rate * elapsed_time` subject to machine/process-specific validity or gating conditions. It is reusable for different machines, materials, rate variables, units, and gates. Historical production-order analysis and backfill are secondary. They must reuse this generic integration engine rather than introduce separate formulas: one integration engine, two operating modes (live incremental processing and historical replay/backfill). ## Components | Component | Responsibility | Must not do | | --- | --- | --- | | `enlyze` | Isolate ENLYZE retrieval and API-specific mapping | Leak assumed API schemas into calculations | | `domain` | Small, stable models for orders, provenance, metrics, and events | Encode database or HTTP details | | `calculations` | Versioned, testable operators such as integration and peak detection | Fetch data or write directly to a database | | `persistence` | Store results and calculation state through repository ports | Become a raw-data historian | | `cli` | Operator commands; first use is API exploration | Contain business calculations | | `service` | Future FastAPI composition/API layer | Require endpoints during bootstrap | ## Production Run boundaries and aggregation ENLYZE Production Runs are time segments, not necessarily one-to-one with a production order. A `production_order` may contain multiple runs. The integrator processes every Production Run independently; production-order totals aggregate the resulting derived values. It must not implicitly integrate across a gap between runs. ## Material-consumption integrator Live mode detects the active Production Run/order, ingests new ENLYZE samples, incrementally integrates material consumption while its configured validity conditions are active, and persists calculation state and the cumulative result. Replay mode fetches a bounded completed run and applies that same stateful operator from its initial state. Results retain the run boundary and source interval; aggregation across runs is a distinct production-order-level step. The generic engine does not prescribe a material, source variable, unit, or gate. Those are specific to the supported machine/process use case. ### K7 fiber (first validated implementation) For K7, the current validated calculation integrates `Stundenleistung Anlage` (kg/h) only for intervals where `Geschwindigkeit Gesamtanlage > 0.5 m/min`. The mass-flow signal can remain non-zero while the machine is stopped, so it must not be integrated without this gate. `Anlage läuft` is not the primary gate. `fiber_feed_kg` is an objective process-derived feed quantity, not by itself a finished-product mass or yield. Band-scale material and asynchronously reported ERP production can be shifted by WIP/transport delay, especially around order transitions. Any exact per-order yield calculation therefore needs additional WIP accounting. ### Bento 1 bentonite powder (planned) Bento 1 is the next planned use case for the same material-consumption integrator. Its bentonite-powder source variable and validity/gating conditions are not yet selected or validated and must be determined from ENLYZE process data before implementation. It does not require a separate subsystem. ## Result provenance Persisted metrics/events need machine and production-order context where available, article context, observed/result time or interval, calculation type/version, and the raw ENLYZE source interval used. A calculation run records operational metadata that ties a batch of results to its implementation version and source range. ## Calculation instances Configuration declares a named calculation instance with a calculation type, version, and signal/context references. The engine selects a tested Python operator implementation for that type. This keeps configuration simple and avoids a generic expression language. ## Incremental execution Future calculation state is stored per calculation instance and relevant context partition. For live material-consumption integration it includes the latest processed source position, any interval/carry state needed for correct integration, and the cumulative run value. It enables safe continuation (for example an open peak cycle), while the source interval recorded on results keeps a historical run reproducible by fetching ENLYZE data again. ## Peak-cycle semantics The generic peak detector maintains an open cycle maximum. A reset begins when a sample falls below `drop_ratio * current_peak`; it is confirmed only after actual subsequent below-threshold samples support `hold_seconds` of elapsed time. It emits only maxima at least `min_peak`, starts a fresh cycle after a completed reset, and suppresses duplicates during the continuing low phase. This is calculation-domain logic, with no ENLYZE transport dependency. The detector uses timestamps rather than a sample count. An interval longer than the configured positive `max_sample_gap_seconds` restarts a reset candidate, so a timestamp alone does not establish that a signal was continuously below threshold. These parameters (`min_peak`, `drop_ratio`, `hold_seconds`, and `max_sample_gap_seconds`) are selected per signal rather than treated as global process constants. Equal timestamps are processed in arrival order but contribute no elapsed time, and backwards timestamps are rejected. Reset values are not assumed to approach zero, and peak values may vary materially between cycles.