260 lines
14 KiB
Markdown
260 lines
14 KiB
Markdown
# production-analytics
|
||
|
||
`production-analytics` is a small Python service that derives production
|
||
analytics from ENLYZE data for Grafana. ENLYZE remains the authority for raw
|
||
process data; this project persists only derived metrics, events, relevant
|
||
production-order context, and calculation state.
|
||
|
||
## Status
|
||
|
||
This repository includes an ENLYZE exploration CLI, a production-run/timeseries
|
||
gateway, and a configured continuous material polling runner with atomic JSON
|
||
checkpoints and PostgreSQL cumulative material-consumption snapshots for Grafana.
|
||
The live material-consumption MVP remains in progress.
|
||
There are no HTTP endpoints.
|
||
|
||
## Intended flow
|
||
|
||
```text
|
||
ENLYZE (raw data) -> retrieval boundary -> calculations -> TimescaleDB -> Grafana
|
||
\-> calculation provenance/state
|
||
```
|
||
|
||
Production orders are the primary attribution context for results. Every
|
||
persisted result will ultimately be traceable to its machine, order/article
|
||
context, source time range, and calculation implementation version.
|
||
|
||
## Development
|
||
|
||
Requires Python 3.11 or newer. Install the project and development tools:
|
||
|
||
```bash
|
||
python3 -m pip install -e '.[dev]'
|
||
pytest
|
||
ruff check .
|
||
```
|
||
|
||
Copy `.env.example` to `.env` for public/local configuration documentation.
|
||
Put real ENLYZE credentials only in `secrets/enlyze.env`, which is ignored.
|
||
The CLI reads that file by default without executing it; it accepts only simple
|
||
`KEY=VALUE` lines (quoted values are supported). Environment variables may be
|
||
used instead when appropriate. Never pass keys as command-line arguments.
|
||
|
||
The generic request command only performs `GET` requests and requires the
|
||
operator to provide an API path that is known to be safe and authorized:
|
||
|
||
```bash
|
||
production-analytics enlyze raw /verified/path --pretty
|
||
production-analytics enlyze raw /verified/path --save-fixture response.json
|
||
```
|
||
|
||
Public ENLYZE documentation verifies Bearer-token authentication. Put
|
||
`ENLYZE_API_KEY='...'` in the local secret file; the exploration client sends
|
||
it as an `Authorization: Bearer …` header and never prints that header/value.
|
||
|
||
The second command saves a sanitized, reviewable fixture under
|
||
`fixtures/enlyze/` by default. Raw captures belong in the ignored
|
||
`data/raw/enlyze/` directory and must not be committed.
|
||
|
||
The OpenAPI server URL is `https://app.enlyze.com/api/`; its operation paths
|
||
begin with `/v2/`. Set `ENLYZE_BASE_URL` to that server URL, not to an
|
||
operation path. The documented read-only time-series operation is exposed for
|
||
exploration as `production-analytics enlyze timeseries`.
|
||
|
||
The current Compose file has no application service, so it deliberately does
|
||
not pass ENLYZE credentials to TimescaleDB. A future application service should
|
||
use `env_file: ./secrets/enlyze.env` rather than copying secrets into Compose.
|
||
|
||
## Material-consumption integration
|
||
|
||
`MaterialConsumptionIntegrator` in `calculations` accepts `MaterialSample`
|
||
values (timezone-aware timestamp, material rate in kg/h, numeric gate value).
|
||
Configure a strict `gate_value > gate_threshold` condition and an explicit,
|
||
positive `max_sample_gap_seconds`. Each interval uses the preceding sample's
|
||
rate and gate, converting elapsed seconds to hours to accumulate kg. A longer
|
||
gap contributes neither consumption nor running time and resets the baseline
|
||
to the newer sample. No time before the first or after the last sample is inferred.
|
||
|
||
Call `process(sample)` or `process_many(samples)` on the same instance for live
|
||
input, replay, or successive chunks. Both use the same calculation. The immutable
|
||
`state` snapshot exposes cumulative kg, integrated running seconds, last timestamp
|
||
(UTC), last rate, last gate, and whether the latest gate is active. State restoration
|
||
and JSON persistence are supported. Naive timestamps and non-finite values
|
||
are rejected; backwards timestamps raise without changing state. Duplicate
|
||
timestamps add no consumption but replace the baseline in arrival order.
|
||
Finite negative material rates are currently accepted and decrease cumulative
|
||
consumption during affected integrated intervals. This is intentional generic
|
||
behavior for now; machine-specific validation or clamping may be added later
|
||
at the input/adapter layer if required by process semantics.
|
||
|
||
Run `python scripts/validate_k7_material_consumption.py` from the repository root
|
||
with the ignored `k7-00842-throughput.raw.json` and `k7-00842-speed.raw.json`
|
||
captures in `data/raw/enlyze/`. The utility joins common non-null timestamps
|
||
without filling values and rejects unordered or duplicate capture records.
|
||
K7-specific signal UUIDs are declared in the runner configuration and validation
|
||
utility: `Stundenleistung Anlage` is gated by `Geschwindigkeit Gesamtanlage > 0.5 m/min`. With the explicit
|
||
20-second validation gap limit (`--max-sample-gap-seconds` to override),
|
||
2,672 common samples yield **5.205555556 h** and **5,180.127150811 kg** for run
|
||
00842, matching the previous manual calculation's rounded results. Synthetic
|
||
tests verify equivalence to manual interval integration without local captures
|
||
or live ENLYZE access. Grafana and Bento 1 remain pending.
|
||
|
||
`MaterialPollingService.poll_once(now=...)` starts at the open run's start or
|
||
resumes inclusively at its last processed timestamp. Duplicate boundary samples
|
||
add no time. A new run UUID for the same order retains cumulative kg and running
|
||
seconds but resets the temporal baseline; different orders and machines have
|
||
separate state. Empty responses add no consumption. Naive `now` is rejected;
|
||
no open run or a window beginning after `now` returns `None` without saving.
|
||
|
||
`JsonMaterialStateStore` uses deterministic SHA-256 filenames derived from the
|
||
machine/order pair, preventing path traversal and sanitized-name collisions.
|
||
It flushes and fsyncs a temporary file in the state directory before atomic
|
||
replacement; a failed write leaves the previous primary file intact. Missing
|
||
state returns `None`; malformed JSON or state raises `ValueError` with file
|
||
context. Use an ignored runtime directory and one polling writer per state key;
|
||
atomic replacement does not coordinate concurrent read-modify-write cycles.
|
||
Files from the earlier uncommitted sanitized-filename prototype are not loaded
|
||
under the new names. The gateway rejects ambiguous open runs, naive windows,
|
||
and missing columns or malformed records instead of silently skipping them.
|
||
|
||
## Continuous material polling
|
||
|
||
The committed [K7 configuration](config/k7-material-consumption.yaml) declares
|
||
`k7-fiber-consumption`, a `material_consumption` calculation at version `"1"`.
|
||
It selects `Stundenleistung Anlage` (kg/h) as the rate and `Geschwindigkeit
|
||
Gesamtanlage` (m/min) as the gate, integrating only when the gate is **strictly
|
||
above 0.5**, with a maximum sample gap of 20 seconds. `Anlage läuft` is not the
|
||
primary gate. The output is `material_consumption` in `kg`, grouped by
|
||
`production_order`; production orders remain opaque strings, including leading
|
||
zeros and whitespace.
|
||
|
||
Install the declared dependencies in the repository venv, then run from the
|
||
repository root:
|
||
|
||
```bash
|
||
.venv/bin/python -m pip install -e '.[dev]'
|
||
.venv/bin/python -m production_analytics.cli run material-poll \
|
||
--config config/k7-material-consumption.yaml \
|
||
--calculation-id k7-fiber-consumption \
|
||
--poll-interval-seconds 10 \
|
||
--state-directory data/state/material \
|
||
--secrets-file secrets/enlyze.env
|
||
```
|
||
|
||
The installed `production-analytics run material-poll` command is equivalent.
|
||
The interval, state directory, and secrets path above are defaults. Paths are
|
||
relative to the current working directory. Existing `ENLYZE_BASE_URL`,
|
||
`ENLYZE_API_KEY`, and `ENLYZE_HTTP_TIMEOUT_SECONDS` handling is reused; values
|
||
in the secrets file override environment values. Keys are never CLI arguments.
|
||
|
||
Keep calculation fields in YAML and technical runtime settings in CLI options.
|
||
The loader requires every field shown in the K7 file, rejects unknown fields,
|
||
duplicate keys/ids, unsupported types/versions/outputs, and invalid or non-finite
|
||
numbers. Numeric fields must be YAML numbers, and `version` must be quoted.
|
||
Multiple material calculations may share a file, but then `--calculation-id` is
|
||
required. All entries are validated before selection. The unrelated
|
||
`config/calculations.example.yaml` remains illustrative and is not executable
|
||
by this material-only runner.
|
||
|
||
The foreground runner polls once immediately using UTC, then sleeps for the
|
||
configured positive interval after each completed cycle, including failures.
|
||
A slow poll delays the next cycle; there is no overlap or catch-up scheduling.
|
||
Ctrl-C/SIGINT exits cleanly. Successful cycles print machine, opaque order,
|
||
cumulative kg, integrated running seconds, and cycle-start timestamp. No open
|
||
run (or no eligible window) is normal and visible. Errors on stderr identify
|
||
state loading/saving or ENLYZE/gateway/polling and the exception class; arbitrary
|
||
exception text and response bodies are withheld to avoid leaking credentials.
|
||
Startup/configuration failures return non-zero; cycle failures retry after the
|
||
delay without replacing checkpoints with empty state.
|
||
|
||
Checkpoints live in ignored `data/state/material/` by default, with one hashed
|
||
JSON filename per machine/order pair. Run **one polling writer per state key**;
|
||
there is no multi-process locking. Use a separate state directory when changing
|
||
calculation parameters or running a different calculation for the same machine
|
||
and order: checkpoints are not namespaced by calculation id/version. These files
|
||
are integration checkpoints, not a Grafana metric store.
|
||
|
||
Each successful non-empty poll writes one cumulative snapshot to PostgreSQL before
|
||
printing its result. JSON remains the restart checkpoint; PostgreSQL stores derived
|
||
time-series snapshots only. ENLYZE remains the raw-data source of truth. Database
|
||
write failures emit a concise error and polling continues with the next cycle;
|
||
missed snapshots are not retried or backfilled and JSON state is not rolled back.
|
||
The next successful snapshot includes the continuing cumulative total.
|
||
|
||
A new order's first snapshot is **not guaranteed to be zero**: the service processes
|
||
available samples from the run start before returning, so its first total may
|
||
already be non-zero. A returned zero is stored normally. Totals continue across
|
||
runs for the same order; existing checkpoints for a previously seen order resume.
|
||
|
||
### PostgreSQL setup
|
||
|
||
Set all five runtime settings: `POSTGRES_HOST`, `POSTGRES_PORT`, `POSTGRES_DB`,
|
||
`POSTGRES_USER`, and `POSTGRES_PASSWORD`, using exported environment variables or
|
||
the existing secrets file (file values override the environment). The runner does
|
||
not automatically load `.env`. Missing or invalid settings fail at startup;
|
||
connectivity and schema errors are reported during polling. For the existing
|
||
Compose instance, use host `localhost`, the published port (default `5432`), and
|
||
`production_analytics` for both database and user. Set the same password in
|
||
Compose's `.env` and the runner environment/secrets file.
|
||
|
||
Start the database and apply the repeatable schema from the repository root:
|
||
|
||
```bash
|
||
docker compose up -d timescaledb
|
||
docker compose exec -T timescaledb psql -U production_analytics -d production_analytics \
|
||
-v ON_ERROR_STOP=1 < db/schema.sql
|
||
```
|
||
|
||
Wait until PostgreSQL is ready before applying the schema. Configure Grafana's
|
||
PostgreSQL data source to query `material_consumption_snapshots`, selecting
|
||
`timestamp` as time and `consumption_kg` as the cumulative value, filtered by
|
||
`calculation_id`, `machine_id`, and `production_order`. These are ordinary
|
||
PostgreSQL tables/indexes; no Timescale-specific features or hypertables are used.
|
||
The primary key deduplicates calculation/machine/order/timestamp (first write wins).
|
||
A short transaction opens and closes one synchronous connection per snapshot,
|
||
with 10-second connection and statement timeouts. Snapshot timestamps are poll
|
||
start times, not source-sample timestamps. Schema application is manual.
|
||
Bento 1 bentonite source and gate selection remain intentionally undefined pending
|
||
process validation; the generic integration core is unchanged and reusable.
|
||
|
||
## Peak-cycle detection
|
||
|
||
`PeakCycleDetector` is a pure calculation-domain component for roll length,
|
||
roll weight, and comparable sawtooth/batch signals. It retains the current
|
||
maximum and emits it only after the value is below
|
||
`drop_ratio * current_peak` continuously for `hold_seconds`. Configuration is
|
||
`min_peak`, `drop_ratio` (strictly between 0 and 1), non-negative
|
||
`hold_seconds`, and positive `max_sample_gap_seconds`. These are per-signal
|
||
configuration parameters, not globally fixed process constants. The hold uses elapsed
|
||
timestamps, never a sample count; an interval longer than the configured
|
||
maximum sample gap restarts a reset candidate,
|
||
so a timestamp gap alone is not evidence that a signal remained below threshold.
|
||
`max_sample_gap_seconds` is required rather than defaulted, so the source's
|
||
continuity assumption is explicit for each detector configuration.
|
||
Reset values need not approach zero (and can be negative), and peak magnitude
|
||
can vary materially between cycles. Equal timestamps are accepted in arrival
|
||
order but add no elapsed hold time; backwards timestamps are rejected.
|
||
|
||
Real-world validations use ignored local captures and are not part of the test
|
||
suite. K7 roll length (m), variable `bf81c547-dccf-4709-aee2-0f78366d1dfc`, is
|
||
the clean reference case: its 10-second capture produced exactly four plausible
|
||
peaks of about 49.98 m with `min_peak=10`, `drop_ratio=0.75`,
|
||
`hold_seconds=60`, and `max_sample_gap_seconds=20`. Run
|
||
`python scripts/validate_k7_roll_length.py` when that capture is available.
|
||
|
||
Bento 2 roll weight (kg), machine `6201f3de-b940-45d7-930c-1481c5ae9e1d`,
|
||
variable `059c0015-9aff-4270-a4ce-f21b83f99378`, is the more realistic
|
||
robustness case. Its observed 10-second signal has plateaus, reset values of
|
||
roughly 300–470 kg, and materially varying peak heights. With
|
||
`min_peak=800`, `drop_ratio=0.75`, `hold_seconds=60`, and
|
||
`max_sample_gap_seconds=20`, the extended capture produces 13 plausible
|
||
completed peaks, including a legitimate approximately 1195 kg cycle. Run
|
||
`python scripts/validate_bento2_roll_weight.py` when the ignored capture is
|
||
available. Neither capture should be added to version control or automated
|
||
tests.
|
||
|
||
The normal unit test suite mocks PostgreSQL and requires no live database.
|
||
|
||
See [PROJECT_KNOWLEDGE.md](PROJECT_KNOWLEDGE.md) for durable project context
|
||
and [docs/roadmap.md](docs/roadmap.md) for the implementation sequence.
|