Files
production-analytics/README.md
T

469 lines
26 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# production-analytics
`production-analytics` is a small Python service that derives production
analytics from ENLYZE data for Grafana. ENLYZE remains the authority for raw
process data; this project persists only derived metrics, events, relevant
production-order context, and calculation state.
## Status
This repository includes an ENLYZE exploration CLI, a production-run/timeseries
gateway, and a configured continuous material polling runner with atomic JSON
checkpoints and PostgreSQL cumulative material-consumption snapshots for Grafana.
The live material-consumption MVP remains in progress.
There are no HTTP endpoints.
## Intended flow
```text
ENLYZE (raw data) -> retrieval boundary -> calculations -> TimescaleDB -> Grafana
\-> calculation provenance/state
```
Production orders are the primary attribution context for results. Every
persisted result will ultimately be traceable to its machine, order/article
context, source time range, and calculation implementation version.
## Development
Requires Python 3.11 or newer. Install the project and development tools:
```bash
python3 -m pip install -e '.[dev]'
pytest
ruff check .
```
Copy `.env.example` to `.env` for public/local configuration documentation.
Put real ENLYZE credentials only in `secrets/enlyze.env`, which is ignored.
The CLI reads that file by default without executing it; it accepts only simple
`KEY=VALUE` lines (quoted values are supported). Environment variables may be
used instead when appropriate. Never pass keys as command-line arguments.
The generic request command only performs `GET` requests and requires the
operator to provide an API path that is known to be safe and authorized:
```bash
production-analytics enlyze raw /verified/path --pretty
production-analytics enlyze raw /verified/path --save-fixture response.json
```
Public ENLYZE documentation verifies Bearer-token authentication. Put
`ENLYZE_API_KEY='...'` in the local secret file; the exploration client sends
it as an `Authorization: Bearer …` header and never prints that header/value.
The second command saves a sanitized, reviewable fixture under
`fixtures/enlyze/` by default. Raw captures belong in the ignored
`data/raw/enlyze/` directory and must not be committed.
The OpenAPI server URL is `https://app.enlyze.com/api/`; its operation paths
begin with `/v2/`. Set `ENLYZE_BASE_URL` to that server URL, not to an
operation path. The documented read-only time-series operation is exposed for
exploration as `production-analytics enlyze timeseries`.
The current Compose file has no application service, so it deliberately does
not pass ENLYZE credentials to TimescaleDB. A future application service should
use `env_file: ./secrets/enlyze.env` rather than copying secrets into Compose.
## Material-consumption integration
`MaterialConsumptionIntegrator` in `calculations` accepts `MaterialSample`
values (timezone-aware timestamp, material rate in kg/h, numeric gate value).
Configure a strict `gate_value > gate_threshold` condition and an explicit,
positive `max_sample_gap_seconds`. Each interval uses the preceding sample's
rate and gate, converting elapsed seconds to hours to accumulate kg. A longer
gap contributes neither consumption nor running time and resets the baseline
to the newer sample. No time before the first or after the last sample is inferred.
Call `process(sample)` or `process_many(samples)` on the same instance for live
input, replay, or successive chunks. Both use the same calculation. The immutable
`state` snapshot exposes cumulative kg, integrated running seconds, last timestamp
(UTC), last rate, last gate, and whether the latest gate is active. State restoration
and JSON persistence are supported. Naive timestamps and non-finite values
are rejected; backwards timestamps raise without changing state. Duplicate
timestamps add no consumption but replace the baseline in arrival order.
Finite negative material rates are currently accepted and decrease cumulative
consumption during affected integrated intervals. This is intentional generic
behavior for now; machine-specific validation or clamping may be added later
at the input/adapter layer if required by process semantics.
Run `python scripts/validate_k7_material_consumption.py` from the repository root
with the ignored `k7-00842-throughput.raw.json` and `k7-00842-speed.raw.json`
captures in `data/raw/enlyze/`. The utility joins common non-null timestamps
without filling values and rejects unordered or duplicate capture records.
K7-specific signal UUIDs are declared in the runner configuration and validation
utility: `Stundenleistung Anlage` is gated by `Geschwindigkeit Gesamtanlage > 0.5 m/min`. With the explicit
20-second validation gap limit (`--max-sample-gap-seconds` to override),
2,672 common samples yield **5.205555556 h** and **5,180.127150811 kg** for run
00842, matching the previous manual calculation's rounded results. Synthetic
tests verify equivalence to manual interval integration without local captures
or live ENLYZE access. Grafana and Bento 1 remain pending.
`MaterialPollingService.poll_once(now=...)` starts at the open run's start or
resumes inclusively at its last processed timestamp. Duplicate boundary samples
add no time. A new run UUID for the same order retains cumulative kg and running
seconds but resets the temporal baseline; different orders and machines have
separate state. Empty responses add no consumption. Naive `now` is rejected;
no open run or a window beginning after `now` returns `None` without saving.
`JsonMaterialStateStore` uses deterministic SHA-256 filenames derived from the
machine/order pair, preventing path traversal and sanitized-name collisions.
It flushes and fsyncs a temporary file in the state directory before atomic
replacement; a failed write leaves the previous primary file intact. Missing
state returns `None`; malformed JSON or state raises `ValueError` with file
context. Use an ignored runtime directory and one polling writer per state key;
atomic replacement does not coordinate concurrent read-modify-write cycles.
Files from the earlier uncommitted sanitized-filename prototype are not loaded
under the new names. The gateway rejects ambiguous open runs, naive windows,
and missing columns or malformed records instead of silently skipping them.
## Continuous material polling
The committed [K7 configuration](config/k7-material-consumption.yaml) declares
`k7-fiber-consumption`, a `material_consumption` calculation at version `"1"`.
It selects `Stundenleistung Anlage` (kg/h) as the rate and `Geschwindigkeit
Gesamtanlage` (m/min) as the gate, integrating only when the gate is **strictly
above 0.5**, with a maximum sample gap of 20 seconds. `Anlage läuft` is not the
primary gate. The output is `material_consumption` in `kg`, grouped by
`production_order`; production orders remain opaque strings, including leading
zeros and whitespace.
Install the declared dependencies in the repository venv, then run from the
repository root:
```bash
.venv/bin/python -m pip install -e '.[dev]'
.venv/bin/python -m production_analytics.cli run material-poll \
--config config/k7-material-consumption.yaml \
--calculation-id k7-fiber-consumption \
--poll-interval-seconds 10 \
--state-directory data/state/material \
--secrets-file secrets/enlyze.env
```
The installed `production-analytics run material-poll` command is equivalent.
The interval, state directory, and secrets path above are defaults. Paths are
relative to the current working directory. Existing `ENLYZE_BASE_URL`,
`ENLYZE_API_KEY`, and `ENLYZE_HTTP_TIMEOUT_SECONDS` handling is reused; values
in the secrets file override environment values. Keys are never CLI arguments.
Keep calculation fields in YAML and technical runtime settings in CLI options.
The loader requires every field shown in the K7 file, rejects unknown fields,
duplicate keys/ids, unsupported types/versions/outputs, and invalid or non-finite
numbers. Numeric fields must be YAML numbers, and `version` must be quoted.
Multiple material calculations may share a file, but then `--calculation-id` is
required. All entries are validated before selection. The unrelated
`config/calculations.example.yaml` remains illustrative and is not executable
by this material-only runner.
The foreground runner polls once immediately using UTC, then sleeps for the
configured positive interval after each completed cycle, including failures.
A slow poll delays the next cycle; there is no overlap or catch-up scheduling.
Ctrl-C/SIGINT exits cleanly. Successful cycles print machine, opaque order,
cumulative kg, integrated running seconds, and cycle-start timestamp. No open
run (or no eligible window) is normal and visible. Errors on stderr identify
state loading/saving or ENLYZE/gateway/polling and the exception class; arbitrary
exception text and response bodies are withheld to avoid leaking credentials.
Startup/configuration failures return non-zero; cycle failures retry after the
delay without replacing checkpoints with empty state.
Checkpoints live in ignored `data/state/material/` by default, with one hashed
JSON filename per machine/order pair. Run **one polling writer per state key**;
there is no multi-process locking. Use a separate state directory when changing
calculation parameters or running a different calculation for the same machine
and order: checkpoints are not namespaced by calculation id/version. These files
are integration checkpoints, not a Grafana metric store.
Each successful non-empty poll writes one cumulative snapshot to PostgreSQL before
printing its result. JSON remains the restart checkpoint; PostgreSQL stores derived
time-series snapshots only. ENLYZE remains the raw-data source of truth. Database
write failures emit a concise error and polling continues with the next cycle;
missed snapshots are not retried or backfilled and JSON state is not rolled back.
The next successful snapshot includes the continuing cumulative total.
A new order's first snapshot is **not guaranteed to be zero**: the service processes
available samples from the run start before returning, so its first total may
already be non-zero. A returned zero is stored normally. Totals continue across
runs for the same order; existing checkpoints for a previously seen order resume.
### PostgreSQL setup
Set all five runtime settings: `POSTGRES_HOST`, `POSTGRES_PORT`, `POSTGRES_DB`,
`POSTGRES_USER`, and `POSTGRES_PASSWORD`, using exported environment variables or
the existing secrets file (file values override the environment). The runner does
not automatically load `.env`. Missing or invalid settings fail at startup;
connectivity and schema errors are reported during polling. For the existing
Compose instance, use host `localhost`, the published port (default `5432`), and
`production_analytics` for both database and user. Set the same password in
Compose's `.env` and the runner environment/secrets file.
Start the database and apply the repeatable schema from the repository root:
```bash
docker compose up -d timescaledb
docker compose exec -T timescaledb psql -U production_analytics -d production_analytics \
-v ON_ERROR_STOP=1 < db/schema.sql
```
Wait until PostgreSQL is ready before applying the schema. Configure Grafana's
PostgreSQL data source to query `material_consumption_snapshots`, selecting
`timestamp` as time and `consumption_kg` as the cumulative value, filtered by
`calculation_id`, `machine_id`, and `production_order`. These are ordinary
PostgreSQL tables/indexes; no Timescale-specific features or hypertables are used.
The primary key deduplicates calculation/machine/order/timestamp (first write wins).
A short transaction opens and closes one synchronous connection per snapshot,
with 10-second connection and statement timeouts. Snapshot timestamps are poll
start times, not source-sample timestamps. Schema application is manual.
Bento 1 bentonite source and gate selection remain intentionally undefined pending
process validation; the generic integration core is unchanged and reusable.
## ERP current workplace status
The read-only ERP adapter reads production-order/article context from MSSQL
`NV_DWH.dbo.GRAFANA_WORKPLACE_STATUS`, for later live material KPIs. Credentials
belong in ignored `secrets/erp.env`, using `ERP_DB_HOST`, `ERP_DB_PORT`,
`ERP_DB_NAME`, `ERP_DB_USER`, and `ERP_DB_PASSWORD`. The verified endpoint is
`192.168.111.24:49601`, database `NV_DWH`, using SQL Server authentication and
a read-only account. All settings are required; ports must be integers 1–65535.
```python
from production_analytics.erp import ErpSettings, ErpWorkplaceStatusGateway
settings = ErpSettings.from_secret_file() # File values override environment values.
status = ErpWorkplaceStatusGateway(settings).get_current_workplace_status("K7")
```
Alternatively, use `ErpSettings.from_environment(mapping)`. The existing simple
dotenv loader is reused; no shell content is executed. Each call opens and closes
one connection, with 10-second login/query timeouts, and performs a parameterized
SELECT filtered by workplace. Zero rows returns `None`, one row returns an
immutable `CurrentWorkplaceStatus`, and multiple rows raise `ErpReadError`.
Driver failures and malformed rows also raise `ErpReadError` with safe messages.
Only trailing whitespace is removed from workplace, production order, article
number, and description; identifiers otherwise remain unchanged. Nullable fields
stay `None`. Decimal/integer quantities are explicitly converted to finite floats
(with normal floating-point precision); non-finite/overflowing values are rejected.
Quantity fields use m² and remaining time uses hours, following the supplied source
semantics. `feedback_timestamp` is preserved exactly, including a naive timezone.
It represents the latest ERP feedback, and the ERP export is delayed relative to
live process data; the adapter makes no freshness inference.
Automated ERP tests use mocks and require no live connectivity.
## ERP context normalization
Two pure helpers in `production_analytics.context` prepare context for later KPIs:
- `build_enlyze_production_order(erp_production_order, format_template) -> str`
uses an explicit template configured by the caller per machine/workplace.
The template requires exactly one literal `{production_order}` placeholder and
no other braces; invalid templates raise `ValueError`.
Outer ERP-order whitespace is stripped and the
remaining order must contain only ASCII digits. Leading zeros are preserved.
Empty/invalid orders raise `ValueError`. There is no default machine rule.
- `extract_nominal_width_m(article_description: str | None) -> float | None`
is generic across machines and conservatively reads a single
`<width> x <length> m` pair, accepting decimal
comma/point and variable whitespace. For example, `Stex R 1501 C (PR) 5,80 x 50 m`
yields `5.8`. Earlier article numbers and trailing descriptive text are ignored.
Missing, malformed, multiple or chained dimension patterns return `None`, as do
non-positive/non-finite dimensions. Signed/scientific notation and other units
are deliberately unsupported; both dimensions must be positive plain numbers.
K7 currently uses the explicit template `K 7-{production_order}`:
```python
k7_order_template = "K 7-{production_order}"
enlyze_order = build_enlyze_production_order("12026000815", k7_order_template)
# "K 7-12026000815"
```
Future machines may supply different verified templates without changing the
generic helper. Template selection belongs to machine/workplace configuration
or adapters; the helper contains no workplace lookup or machine-specific branch.
No other machine rule is introduced here.
ENLYZE production-order identifiers remain opaque everywhere else, with exact
comparison. These helpers never reverse-parse or split ENLYZE orders, including
combined identifiers such as `K 7-12025000074-K 7-12025000075`; passing such an
identifier to the ERP mapping function is rejected.
Nominal finished-product width currently comes from ERP article-description
parsing because neither verified DWH view (`dbo.GRAFANA_WORKPLACE_STATUS` and
`dbo.GRAFANA_PRODUCTION_CONFIRMATION`) has a dedicated width column. ENLYZE
`wLg1MeasuringWidth` (`Produktbreite`) is explicitly not used as K7 nominal
finished-product width:
it belongs to the MAHLO measurement system and observed values differ from nominal
article widths. Structured ERP product master data would be preferred when available.
These helpers are independent of database access. The service below composes them
for material efficiency; no background processing is attached.
## Feedback-aligned material efficiency
`MaterialEfficiencyService` in `service.material_efficiency` combines ERP good area
with cumulative material consumption for one exact production order. Its
`evaluate_current()` reads the existing ERP workplace adapter once; `evaluate(status)`
evaluates an already supplied `CurrentWorkplaceStatus`. Both return an immutable
`MaterialEfficiencySnapshot` or `None` when inputs are unavailable.
`PostgresMaterialSnapshotRepository(settings).latest_at_or_before(...)` in
`service.postgres_material` takes keyword arguments `calculation_id`, `machine_id`,
`production_order`, and an aware `timestamp`. A parameterized query selects the
latest snapshot **at or before ERP feedback time**, matching calculation, machine,
and mapped order exactly. It returns `MaterialConsumptionSnapshot` or `None`.
Run IDs do not restrict this lookup: stored consumption is already cumulative
across the order's runs. The service neither reintegrates nor sums snapshots and
does not bridge run boundaries.
The result includes workplace/machine/calculation identifiers, both order identifiers,
article context, nominal width, the UTC ERP feedback instant and selected material timestamp, good area, aligned
consumption, `material_consumption_kg_per_m2 = consumption_kg / good_quantity_m2`,
and `material_consumption_g_per_m2 = material_consumption_kg_per_m2 * 1000`.
Nominal width comes from generic ERP description parsing and is context only;
ERP good square metres remain authoritative even when width cannot be parsed.
Missing ERP status, missing aligned material, missing/non-positive/non-finite good
area, invalid consumption (negative or non-finite), or non-finite computed ratios
produce `None`. Zero consumption is valid. Configuration errors, invalid timestamp
alignment, and database failures raise rather than masquerading as missing data.
All machine/workplace settings are supplied explicitly. Example using the existing
K7 material calculation (with previously constructed ERP gateway and PG settings):
```python
from zoneinfo import ZoneInfo
from production_analytics.service.material_efficiency import MaterialEfficiencyService
from production_analytics.service.postgres_material import PostgresMaterialSnapshotRepository
service = MaterialEfficiencyService(
erp_gateway, PostgresMaterialSnapshotRepository(postgres_settings),
workplace="K7",
machine_id="c220f95c-a65e-4cb7-99b7-0626d6c7508c",
calculation_id="k7-fiber-consumption",
format_template="K 7-{production_order}",
# Supply only after verifying the ERP timestamp's source timezone:
erp_timezone=ZoneInfo("Europe/Berlin"),
)
result = service.evaluate_current()
```
Aware ERP timestamps are compared as UTC instants. Naive ERP timestamps require
explicit `erp_timezone`; there is no inferred system/database timezone. Ambiguous
or nonexistent local times during DST transitions are rejected. The result carries the
ERP feedback instant normalized to UTC and the selected aware material timestamp.
The original ERP status object remains unchanged.
This alignment avoids knowingly including future material consumption, but does
not eliminate ERP roll-feedback timing uncertainty (feedback may be roughly one
roll ahead or otherwise offset). Live values are plausibility indicators; final
production-order values are more meaningful as relative timing error diminishes.
There is no interpolation, lag correction, smoothing, freshness threshold, or
estimated timestamp. Historical bootstrap snapshots are usable only when their
stored timestamps satisfy the cutoff; a later cumulative total cannot reconstruct
an earlier value. Reliable final evaluation requires retaining final ERP feedback;
the current-workplace adapter alone does not provide historical completed orders.
The service itself returns results in memory without modifying material snapshots.
The standalone persistence runner is described below. Daily per-machine 24h reporting
and aggregation remain future work. Repository SQL tests use an in-memory SQLite
fixture with driver transport adaptation; they require no live PostgreSQL.
## Peak-cycle detection
`PeakCycleDetector` is a pure calculation-domain component for roll length,
roll weight, and comparable sawtooth/batch signals. It retains the current
maximum and emits it only after the value is below
`drop_ratio * current_peak` continuously for `hold_seconds`. Configuration is
`min_peak`, `drop_ratio` (strictly between 0 and 1), non-negative
`hold_seconds`, and positive `max_sample_gap_seconds`. These are per-signal
configuration parameters, not globally fixed process constants. The hold uses elapsed
timestamps, never a sample count; an interval longer than the configured
maximum sample gap restarts a reset candidate,
so a timestamp gap alone is not evidence that a signal remained below threshold.
`max_sample_gap_seconds` is required rather than defaulted, so the source's
continuity assumption is explicit for each detector configuration.
Reset values need not approach zero (and can be negative), and peak magnitude
can vary materially between cycles. Equal timestamps are accepted in arrival
order but add no elapsed hold time; backwards timestamps are rejected.
Real-world validations use ignored local captures and are not part of the test
suite. K7 roll length (m), variable `bf81c547-dccf-4709-aee2-0f78366d1dfc`, is
the clean reference case: its 10-second capture produced exactly four plausible
peaks of about 49.98 m with `min_peak=10`, `drop_ratio=0.75`,
`hold_seconds=60`, and `max_sample_gap_seconds=20`. Run
`python scripts/validate_k7_roll_length.py` when that capture is available.
Bento 2 roll weight (kg), machine `6201f3de-b940-45d7-930c-1481c5ae9e1d`,
variable `059c0015-9aff-4270-a4ce-f21b83f99378`, is the more realistic
robustness case. Its observed 10-second signal has plateaus, reset values of
roughly 300–470 kg, and materially varying peak heights. With
`min_peak=800`, `drop_ratio=0.75`, `hold_seconds=60`, and
`max_sample_gap_seconds=20`, the extended capture produces 13 plausible
completed peaks, including a legitimate approximately 1195 kg cycle. Run
`python scripts/validate_bento2_roll_weight.py` when the ignored capture is
available. Neither capture should be added to version control or automated
tests.
The normal unit test suite mocks PostgreSQL and requires no live database.
See [PROJECT_KNOWLEDGE.md](PROJECT_KNOWLEDGE.md) for durable project context
and [docs/roadmap.md](docs/roadmap.md) for the implementation sequence.
### Feedback-driven material-efficiency persistence
Apply the updated `db/schema.sql` using the PostgreSQL setup command above, then run
this independent foreground process:
```bash
production-analytics run material-efficiency --config config/k7-material-efficiency.yaml
```
`poll_interval_seconds` in YAML defaults to 60 seconds (fixed delay after each
cycle). Workplace, machine, calculation, order format and `erp_timezone` are also
configured in YAML; K7 uses `Europe/Berlin`. PostgreSQL settings use the existing
exported `POSTGRES_*` variables and `--secrets-file` (default `secrets/enlyze.env`);
ERP uses `--erp-secrets-file` (default `secrets/erp.env`). No additional secrets are
needed, and `.env` is not loaded automatically.
`PostgresMaterialEfficiencyWriter.write(snapshot) -> bool` validates aware
timestamps and finite numbers, then inserts all KPI fields in a short transaction.
`material_efficiency_snapshots` retains one immutable point per
`(calculation_id, machine_id, enlyze_production_order, erp_feedback_timestamp)`.
The writer uses `RETURNING 1` to return `True` for a new insert and `False` for a
duplicate. Only new inserts produce a flushed stdout line with feedback time,
workplace, order, good m², material kg and g/m². Unavailable evaluations and
duplicates remain silent.
Repeated evaluations attempt `ON CONFLICT DO NOTHING`; they never update history,
including after a restart. Missing/invalid KPI inputs produce no row and can be
retried on a later cycle. Nullable article context remains SQL NULL.
Persistence frequency follows ERP feedback changes, not ENLYZE sample frequency.
This runner is independent of the 10-second material polling process. Live values
are plausibility indicators; final production-order (FA) values are more meaningful.
Grafana can read PostgreSQL alone:
```sql
SELECT erp_feedback_timestamp AS "time", material_consumption_g_per_m2
FROM material_efficiency_snapshots
WHERE $__timeFilter(erp_feedback_timestamp)
AND calculation_id = 'k7-fiber-consumption'
AND machine_id = 'c220f95c-a65e-4cb7-99b7-0626d6c7508c'
ORDER BY erp_feedback_timestamp;
```
ERP read errors and transient PostgreSQL connection/operational errors are reported
using error classes only and retried after the configured delay. Each database
operation opens a fresh connection. Configuration, schema/programming and invalid
timestamp contract errors terminate with a nonzero CLI exit; ambiguous/nonexistent
DST feedback remains rejected. Ctrl-C stops the foreground runner cleanly.
The current-status ERP source cannot backfill feedback events missed between polls
or during outages, nor guarantee observation of final FA feedback before the order
changes. Daily 24h reports and final order summary tables remain future work.
No dashboard, aggregation or lag correction is added here.
Bento 1: [fresh bentonite consumption configuration and scope](docs/bento1-fresh-bentonite.md).