Add continuous material polling runner

This commit is contained in:
2026-09-05 07:55:40 +02:00
parent efc563512f
commit 6ddb76fd03
10 changed files with 672 additions and 10 deletions
+68 -4
View File
@@ -8,8 +8,10 @@ production-order context, and calculation state.
## Status
This repository includes an ENLYZE exploration CLI, a production-run/timeseries
gateway, and single-cycle live material polling with JSON state persistence.
Database migrations, a continuous polling runner, and HTTP endpoints remain pending.
gateway, and a configured continuous material polling runner with atomic JSON
checkpoints. The live material-consumption MVP remains in progress: Grafana-oriented
derived-total persistence/exposure and database migrations are still pending.
There are no HTTP endpoints.
## Intended flow
@@ -89,8 +91,8 @@ Run `python scripts/validate_k7_material_consumption.py` from the repository roo
with the ignored `k7-00842-throughput.raw.json` and `k7-00842-speed.raw.json`
captures in `data/raw/enlyze/`. The utility joins common non-null timestamps
without filling values and rejects unordered or duplicate capture records.
K7-specific signal UUIDs are confined to the utility: `Stundenleistung Anlage`
is gated by `Geschwindigkeit Gesamtanlage > 0.5 m/min`. With the explicit
K7-specific signal UUIDs are declared in the runner configuration and validation
utility: `Stundenleistung Anlage` is gated by `Geschwindigkeit Gesamtanlage > 0.5 m/min`. With the explicit
20-second validation gap limit (`--max-sample-gap-seconds` to override),
2,672 common samples yield **5.205555556 h** and **5,180.127150811 kg** for run
00842, matching the previous manual calculation's rounded results. Synthetic
@@ -115,6 +117,68 @@ Files from the earlier uncommitted sanitized-filename prototype are not loaded
under the new names. The gateway rejects ambiguous open runs, naive windows,
and missing columns or malformed records instead of silently skipping them.
## Continuous material polling
The committed [K7 configuration](config/k7-material-consumption.yaml) declares
`k7-fiber-consumption`, a `material_consumption` calculation at version `"1"`.
It selects `Stundenleistung Anlage` (kg/h) as the rate and `Geschwindigkeit
Gesamtanlage` (m/min) as the gate, integrating only when the gate is **strictly
above 0.5**, with a maximum sample gap of 20 seconds. `Anlage läuft` is not the
primary gate. The output is `material_consumption` in `kg`, grouped by
`production_order`; production orders remain opaque strings, including leading
zeros and whitespace.
Install the declared dependencies in the repository venv, then run from the
repository root:
```bash
.venv/bin/python -m pip install -e '.[dev]'
.venv/bin/python -m production_analytics.cli run material-poll \
--config config/k7-material-consumption.yaml \
--calculation-id k7-fiber-consumption \
--poll-interval-seconds 10 \
--state-directory data/state/material \
--secrets-file secrets/enlyze.env
```
The installed `production-analytics run material-poll` command is equivalent.
The interval, state directory, and secrets path above are defaults. Paths are
relative to the current working directory. Existing `ENLYZE_BASE_URL`,
`ENLYZE_API_KEY`, and `ENLYZE_HTTP_TIMEOUT_SECONDS` handling is reused; values
in the secrets file override environment values. Keys are never CLI arguments.
Keep calculation fields in YAML and technical runtime settings in CLI options.
The loader requires every field shown in the K7 file, rejects unknown fields,
duplicate keys/ids, unsupported types/versions/outputs, and invalid or non-finite
numbers. Numeric fields must be YAML numbers, and `version` must be quoted.
Multiple material calculations may share a file, but then `--calculation-id` is
required. All entries are validated before selection. The unrelated
`config/calculations.example.yaml` remains illustrative and is not executable
by this material-only runner.
The foreground runner polls once immediately using UTC, then sleeps for the
configured positive interval after each completed cycle, including failures.
A slow poll delays the next cycle; there is no overlap or catch-up scheduling.
Ctrl-C/SIGINT exits cleanly. Successful cycles print machine, opaque order,
cumulative kg, integrated running seconds, and cycle-start timestamp. No open
run (or no eligible window) is normal and visible. Errors on stderr identify
state loading/saving or ENLYZE/gateway/polling and the exception class; arbitrary
exception text and response bodies are withheld to avoid leaking credentials.
Startup/configuration failures return non-zero; cycle failures retry after the
delay without replacing checkpoints with empty state.
Checkpoints live in ignored `data/state/material/` by default, with one hashed
JSON filename per machine/order pair. Run **one polling writer per state key**;
there is no multi-process locking. Use a separate state directory when changing
calculation parameters or running a different calculation for the same machine
and order: checkpoints are not namespaced by calculation id/version. These files
are integration checkpoints, not a Grafana metric store.
This runner exposes no Grafana/TimescaleDB metric yet. Grafana-oriented derived-total
persistence/exposure remains the next milestone; the MVP is not complete.
Bento 1 bentonite source and gate selection remain intentionally undefined pending
process validation; the generic integration core is unchanged and reusable.
## Peak-cycle detection
`PeakCycleDetector` is a pure calculation-domain component for roll length,