Files
production-analytics/docs/architecture.md
T

61 lines
3.0 KiB
Markdown

# Architecture
## Scope boundary
ENLYZE is authoritative for raw process data. This service retrieves raw data
only to calculate derived results and persists the results, their context, and
the minimum state needed for incremental calculation. It does not mirror raw
time series into PostgreSQL/TimescaleDB.
## Components
| Component | Responsibility | Must not do |
| --- | --- | --- |
| `enlyze` | Isolate ENLYZE retrieval and API-specific mapping | Leak assumed API schemas into calculations |
| `domain` | Small, stable models for orders, provenance, metrics, and events | Encode database or HTTP details |
| `calculations` | Versioned, testable operators such as integration and peak detection | Fetch data or write directly to a database |
| `persistence` | Store results and calculation state through repository ports | Become a raw-data historian |
| `cli` | Operator commands; first use is API exploration | Contain business calculations |
| `service` | Future FastAPI composition/API layer | Require endpoints during bootstrap |
## Result provenance
Persisted metrics/events need machine and production-order context where
available, article context, observed/result time or interval, calculation
type/version, and the raw ENLYZE source interval used. A calculation run
records operational metadata that ties a batch of results to its implementation
version and source range.
## Calculation instances
Configuration declares a named calculation instance with a calculation type,
version, and signal/context references. The engine selects a tested Python
operator implementation for that type. This keeps configuration simple and
avoids a generic expression language.
## Incremental execution
Future calculation state is stored per calculation instance and relevant
context partition. It enables safe continuation (for example an open peak
cycle), while the source interval recorded on results keeps a historical run
reproducible by fetching ENLYZE data again.
## Peak-cycle semantics
The generic peak detector maintains an open cycle maximum. A reset begins when
a sample falls below `drop_ratio * current_peak`; it is confirmed only after
actual subsequent below-threshold samples support `hold_seconds` of elapsed
time. It emits only maxima at least `min_peak`, starts a fresh cycle after a
completed reset, and suppresses duplicates during the continuing low phase.
This is calculation-domain logic, with no ENLYZE transport dependency.
The detector uses timestamps rather than a sample count. An interval longer
than the configured positive `max_sample_gap_seconds` restarts a reset candidate,
so a timestamp alone does not establish that a signal was continuously below
threshold. These parameters (`min_peak`, `drop_ratio`, `hold_seconds`, and
`max_sample_gap_seconds`) are selected per signal rather than treated as global
process constants. Equal timestamps are processed in arrival order but
contribute no elapsed time, and backwards timestamps are rejected. Reset values
are not assumed to approach zero, and peak values may vary materially between
cycles.