Initial production analytics foundation

This commit is contained in:
2026-09-04 05:34:52 +02:00
commit f017d187eb
30 changed files with 1219 additions and 0 deletions
+41
View File
@@ -0,0 +1,41 @@
# Architecture
## Scope boundary
ENLYZE is authoritative for raw process data. This service retrieves raw data
only to calculate derived results and persists the results, their context, and
the minimum state needed for incremental calculation. It does not mirror raw
time series into PostgreSQL/TimescaleDB.
## Components
| Component | Responsibility | Must not do |
| --- | --- | --- |
| `enlyze` | Isolate ENLYZE retrieval and API-specific mapping | Leak assumed API schemas into calculations |
| `domain` | Small, stable models for orders, provenance, metrics, and events | Encode database or HTTP details |
| `calculations` | Versioned, testable operators such as integration and peak detection | Fetch data or write directly to a database |
| `persistence` | Store results and calculation state through repository ports | Become a raw-data historian |
| `cli` | Operator commands; first use is API exploration | Contain business calculations |
| `service` | Future FastAPI composition/API layer | Require endpoints during bootstrap |
## Result provenance
Persisted metrics/events need machine and production-order context where
available, article context, observed/result time or interval, calculation
type/version, and the raw ENLYZE source interval used. A calculation run
records operational metadata that ties a batch of results to its implementation
version and source range.
## Calculation instances
Configuration declares a named calculation instance with a calculation type,
version, and signal/context references. The engine selects a tested Python
operator implementation for that type. This keeps configuration simple and
avoids a generic expression language.
## Incremental execution
Future calculation state is stored per calculation instance and relevant
context partition. It enables safe continuation (for example an open peak
cycle), while the source interval recorded on results keeps a historical run
reproducible by fetching ENLYZE data again.
+27
View File
@@ -0,0 +1,27 @@
# Data model
The entities below describe intended persistence and domain boundaries; this
bootstrap does not define a database schema or migrations.
| Entity | Intended contents | Purpose |
| --- | --- | --- |
| `ProductionOrder` | source ID, machine ID, article/material reference, known interval | Central result-attribution context |
| `MachineState` | machine ID, state, observed interval | Runtime/downtime and contextual calculations |
| `CalculatedMetric` | name, value/unit, result interval, attribution, provenance | Integrated consumption and future aggregates |
| `ProcessEvent` | event type, time, payload/value, attribution, provenance | Completed peaks and threshold/cycle events |
| `CalculationRun` | type/version, execution time, source interval, status/metadata | Auditing and reproducibility |
| calculation state | calculation-instance/context key and serializable state | Incremental execution only |
## Provenance minimum
Each metric/event must store `calculation_type`, `calculation_version`, and
the ENLYZE source time range. Values should retain their unit. Production order
and article can be absent when ENLYZE cannot provide them for a particular
record; absence is explicit rather than fabricated.
## Database direction
TimescaleDB will store result/event time series and supporting relational
context. Schema design comes after the ENLYZE exploration milestone confirms
source identifiers, timestamp precision/timezone, ordering, correction, and
context-history behavior.
+132
View File
@@ -0,0 +1,132 @@
# ENLYZE API investigation
This document separates public facts, schema facts, and customer-specific observations. Raw captures are local under ignored `data/raw/enlyze/`; reviewed sanitized fixtures may be stored in `fixtures/enlyze/`.
## VERIFIED FROM PUBLIC DOCUMENTATION
- ENLYZE uses Bearer-token authentication with `ENLYZE_API_KEY`.
- Public resources include machines, sites, products, production runs, and downtimes.
## VERIFIED FROM OPENAPI SCHEMA
Source: locally supplied `data/raw/enlyze/openapi.json`, verified by the operator as OpenAPI 3.1.0, `ENLYZE API` v2. The schema server is the relative URL `/api`; resolving it against `https://app.enlyze.com` and combining it with paths such as `/v2/machines` yields `https://app.enlyze.com/api/v2/machines`.
Every operation below declares `Bearer` security: an `apiKey` in the `Authorization` header. The exploration CLI sends `Authorization: Bearer <ENLYZE_API_KEY>` without logging it.
### Common pagination
Paginated list/time-series responses contain `data` and `metadata`. The required `metadata.next_cursor` is a string continuation token or null. Reuse a non-null value as `cursor` with otherwise identical parameters. No general `limit`/page-size parameter is declared. UUID-array filters permit at most 50 items; one time-series request permits at most 100 variables.
### Machines and sites
| Operation | Parameters | Response |
| --- | --- | --- |
| `GET /v2/machines` | Optional `site` UUID array (max 50), `cursor` | `PaginatedMachines` |
| `GET /v2/machines/{uuid}` | Required machine UUID | `Machine` |
| `GET /v2/sites` | Optional `cursor` | `PaginatedSites` |
| `GET /v2/sites/{uuid}` | Required site UUID | `Site` |
`Machine` requires UUID `uuid`, `name`, site UUID `site`, and date-only `genesis_date`. `Site` requires UUID `uuid` and `name`. No current machine-running state is exposed by these schemas.
### Products and production runs
| Operation | Parameters | Response |
| --- | --- | --- |
| `GET /v2/products` | Optional `cursor` | `PaginatedProducts` |
| `GET /v2/products/{uuid}` | Required product UUID | `Product` |
| `GET /v2/production-runs` | `cursor`; optional `machine` UUID array (max 50), `product-external-id`, `product` UUID, `production-order-external-id`, `start`, `end`, `uuid` UUID array (max 50), `q` | `PaginatedProductionRuns` |
| `GET /v2/production-runs/{uuid}` | Required production-run UUID | `ProductionRun` |
`Product` requires UUID `uuid`, ERP/MES identifier `external_id`, and nullable description/name `name`.
`ProductionRun` requires UUID `uuid`, machine UUID `machine`, string `production_order` (the schema calls this the identifier of the production order), product UUID `product`, timezone-aware ISO 8601 `start`, and nullable timezone-aware ISO 8601 `end`. A null end denotes an unfinished run. It also contains quantities (`{value, unit}` when present), OEE components (score 0–1 and seconds of time loss), optional maximum run speed with unit and observation period, data-quality metrics, and arbitrary attributes.
For list filtering, `start` means runs executed starting at/after the supplied time; `end` means runs executed up to the supplied time. This does not prove interval-overlap behavior needed for integration.
### Variables and data sources
| Operation | Parameters | Response |
| --- | --- | --- |
| `GET /v2/variables` | Optional `machine`, `type`, `state`, `data_source`, `data_type`, `uuid`, `depends_on` UUID arrays (each max 50 where applicable), `cursor` | `PaginatedVariables` |
| `GET /v2/data-sources` | Optional `cursor`, `uuid`, `machine`, `spark` UUID arrays (max 50) | `PaginatedDataSources` |
| `GET /v2/data-sources/{uuid}` | Required data-source UUID | `DataSource` |
`VariableResponse` requires variable UUID, associated machine UUID, data type, and discriminated details. It can include `display_name`, nullable string `unit`, nullable `scaling_factor`, comment, and types `FLOAT`, `INTEGER`, `STRING`, `BOOLEAN`, or array variants. Spark details include data-source UUID, structured origin identifier, origin name/comment, and current capture state (`inactive`, `active`, `evaluating`). Derived details list dependencies and description. The deprecated top-level `type` mirrors `details.type` (`spark` or `derived`).
`DataSource` can provide `readout_limit`, `readout_interval`, current `readout_status`, and machine UUID. This is a data-source/readout state, not a documented machine operational/downtime state.
### Historical time series
| Operation | Parameters | Response |
| --- | --- | --- |
| `POST /v2/timeseries` | Required JSON `machine`, timezone-aware ISO 8601 `start`, `end`, and 1–100 `variables`; optional `resampling_interval`, `cursor` | `TimeseriesResponse` |
Each requested variable contains required UUID and optional `resampling_method`: `first`, `last`, `max`, `min`, `count`, `sum`, `avg`, `median`, `std`, `q5`, `q25`, `q75`, `q95`. `resampling_interval` is an optional integer from 10 to 604800; the schema does not define its unit or whether omission means native/raw resolution.
`TimeseriesResponse.data.columns` is ordered with `time` always first, followed by requested variable UUIDs. `data.records` is tabular in that column order. The schema example uses timestamps such as `2023-01-03T14:00:00.0+00:00` and numeric values. It defines no per-sample unit, quality/status field, null/missing-value behavior, boundary inclusion, ordering, or correction/backfill semantics. Units belong to `VariableResponse.unit`, not the time-series envelope.
### Downtimes
| Operation | Parameters | Response |
| --- | --- | --- |
| `GET /v2/downtimes` | Optional UUID/machine UUID arrays (max 50), timezone-aware ISO 8601 `start`, `end`, reason filters, reason state, `cursor` | `PaginatedDowntimes` |
`Downtime` requires UUID, machine UUID, type (`THRESHOLD` or `NO_DATA`), timezone-aware ISO 8601 start, nullable end, and nullable comment/reason/update information. A null end denotes an ongoing downtime. It supports historical downtime retrieval but does not establish a complete machine-state timeline.
## VERIFIED WITH LIVE CUSTOMER API
None. This execution environment has no ENLYZE DNS/network access, so no customer operation has been sent in this milestone.
## STILL UNKNOWN
- Customer response values, permissions, authentication success/failure behavior, default page size, and rate/retention limits.
- Whether production-run filters have the desired interval-overlap behavior; completed-run `start`/`end` are candidate integration boundaries but require live confirmation.
- Which customer variable is a consumption signal and whether its unit/scaling factor makes it usable for integration.
- Whether omitted `resampling_interval` yields native/raw resolution; its unit and exact temporal meaning of resampling methods.
- Sample ordering, boundary inclusion, native sampling behavior, gaps, nulls, quality/bad samples, late arrivals, corrections, and backfills.
- A complete historical machine operational-state resource beyond downtimes.
## Exact bounded manual exploration sequence
Set the schema server URL locally; the key remains only in `secrets/enlyze.env`:
```bash
export ENLYZE_BASE_URL='https://app.enlyze.com/api/'
```
1. First page of machines (the schema has no limit parameter):
```bash
production-analytics enlyze raw /v2/machines --pretty \
--save-raw machines-page-1.raw.json --save-fixture machines-page-1.json
```
2. Select a harmless machine UUID from the ignored raw capture and retrieve a narrow completed-run page. Replace uppercase placeholders with local values:
```bash
production-analytics enlyze raw /v2/production-runs \
--query 'machine=MACHINE_UUID' \
--query 'start=RUN_SEARCH_START_ISO8601' \
--query 'end=RUN_SEARCH_END_ISO8601' --pretty \
--save-raw production-runs.raw.json --save-fixture production-runs.json
```
3. Resolve the selected run’s product and discover variables for its machine:
```bash
production-analytics enlyze raw /v2/products/PRODUCT_UUID --pretty \
--save-raw product.raw.json --save-fixture product.json
production-analytics enlyze raw /v2/variables --query 'machine=MACHINE_UUID' --pretty \
--save-raw variables.raw.json --save-fixture variables.json
```
4. Select a numeric variable with a meaningful unit and query a short completed-run subinterval. Omit resampling initially to test the optional/default behavior:
```bash
production-analytics enlyze timeseries \
--machine MACHINE_UUID --variable VARIABLE_UUID \
--start SAMPLE_START_ISO8601 --end SAMPLE_END_ISO8601 --pretty \
--save-raw timeseries.raw.json --save-fixture timeseries.json
```
`--resampling-interval 10` and `--resampling-method avg` are available only for a subsequent bounded comparison. The CLI reads `secrets/enlyze.env` by default and does not print the token/header. Raw files remain ignored; review every sanitized fixture before committing it.
+39
View File
@@ -0,0 +1,39 @@
# Roadmap
## 0. Bootstrap (this milestone)
Establish package boundaries, architecture documentation, example calculation
configuration, basic domain/protocol models, tests, and an optional local
TimescaleDB Compose design.
## 1. ENLYZE API exploration client/CLI (recommended exact next milestone)
A read-only, GET-only raw inspection CLI and fixture sanitizer are implemented.
The remaining work is live, authorized exploration of machines, signal catalog,
bounded samples, state history, production-order history, and article/material
context. Record sanitized fixtures and update `docs/enlyze-api.md` with
evidence. Do not implement ingestion or a database schema before this is
complete.
## 2. Persistence foundation
Design and migrate a TimescaleDB schema for derived metrics/events, calculation
runs, context, and incremental state. Add Grafana-oriented query examples.
## 3. First vertical slice
Implement one configured integration calculation over a verified signal and
production-order context. Include backfill, rerun/provenance behavior, and
tests against sanitized fixtures.
## 4. Peak detection
Implement the configurable sawtooth peak detector: retain the current maximum;
close a cycle after a configurable below-peak fraction remains true for a
configurable duration; emit one completed peak event; preserve open-cycle
state.
## Later
State/cycle durations, aggregates by order, rolling statistics, threshold
events, derivatives, and specific-consumption calculations.