Files
production-analytics/README.md
T

23 KiB
Raw Blame History

production-analytics

production-analytics is a small Python service that derives production analytics from ENLYZE data for Grafana. ENLYZE remains the authority for raw process data; this project persists only derived metrics, events, relevant production-order context, and calculation state.

Status

This repository includes an ENLYZE exploration CLI, a production-run/timeseries gateway, and a configured continuous material polling runner with atomic JSON checkpoints and PostgreSQL cumulative material-consumption snapshots for Grafana. The live material-consumption MVP remains in progress. There are no HTTP endpoints.

Intended flow

ENLYZE (raw data) -> retrieval boundary -> calculations -> TimescaleDB -> Grafana
                                      \-> calculation provenance/state

Production orders are the primary attribution context for results. Every persisted result will ultimately be traceable to its machine, order/article context, source time range, and calculation implementation version.

Development

Requires Python 3.11 or newer. Install the project and development tools:

python3 -m pip install -e '.[dev]'
pytest
ruff check .

Copy .env.example to .env for public/local configuration documentation. Put real ENLYZE credentials only in secrets/enlyze.env, which is ignored. The CLI reads that file by default without executing it; it accepts only simple KEY=VALUE lines (quoted values are supported). Environment variables may be used instead when appropriate. Never pass keys as command-line arguments.

The generic request command only performs GET requests and requires the operator to provide an API path that is known to be safe and authorized:

production-analytics enlyze raw /verified/path --pretty
production-analytics enlyze raw /verified/path --save-fixture response.json

Public ENLYZE documentation verifies Bearer-token authentication. Put ENLYZE_API_KEY='...' in the local secret file; the exploration client sends it as an Authorization: Bearer … header and never prints that header/value.

The second command saves a sanitized, reviewable fixture under fixtures/enlyze/ by default. Raw captures belong in the ignored data/raw/enlyze/ directory and must not be committed.

The OpenAPI server URL is https://app.enlyze.com/api/; its operation paths begin with /v2/. Set ENLYZE_BASE_URL to that server URL, not to an operation path. The documented read-only time-series operation is exposed for exploration as production-analytics enlyze timeseries.

The current Compose file has no application service, so it deliberately does not pass ENLYZE credentials to TimescaleDB. A future application service should use env_file: ./secrets/enlyze.env rather than copying secrets into Compose.

Material-consumption integration

MaterialConsumptionIntegrator in calculations accepts MaterialSample values (timezone-aware timestamp, material rate in kg/h, numeric gate value). Configure a strict gate_value > gate_threshold condition and an explicit, positive max_sample_gap_seconds. Each interval uses the preceding sample's rate and gate, converting elapsed seconds to hours to accumulate kg. A longer gap contributes neither consumption nor running time and resets the baseline to the newer sample. No time before the first or after the last sample is inferred.

Call process(sample) or process_many(samples) on the same instance for live input, replay, or successive chunks. Both use the same calculation. The immutable state snapshot exposes cumulative kg, integrated running seconds, last timestamp (UTC), last rate, last gate, and whether the latest gate is active. State restoration and JSON persistence are supported. Naive timestamps and non-finite values are rejected; backwards timestamps raise without changing state. Duplicate timestamps add no consumption but replace the baseline in arrival order. Finite negative material rates are currently accepted and decrease cumulative consumption during affected integrated intervals. This is intentional generic behavior for now; machine-specific validation or clamping may be added later at the input/adapter layer if required by process semantics.

Run python scripts/validate_k7_material_consumption.py from the repository root with the ignored k7-00842-throughput.raw.json and k7-00842-speed.raw.json captures in data/raw/enlyze/. The utility joins common non-null timestamps without filling values and rejects unordered or duplicate capture records. K7-specific signal UUIDs are declared in the runner configuration and validation utility: Stundenleistung Anlage is gated by Geschwindigkeit Gesamtanlage > 0.5 m/min. With the explicit 20-second validation gap limit (--max-sample-gap-seconds to override), 2,672 common samples yield 5.205555556 h and 5,180.127150811 kg for run 00842, matching the previous manual calculation's rounded results. Synthetic tests verify equivalence to manual interval integration without local captures or live ENLYZE access. Grafana and Bento 1 remain pending.

MaterialPollingService.poll_once(now=...) starts at the open run's start or resumes inclusively at its last processed timestamp. Duplicate boundary samples add no time. A new run UUID for the same order retains cumulative kg and running seconds but resets the temporal baseline; different orders and machines have separate state. Empty responses add no consumption. Naive now is rejected; no open run or a window beginning after now returns None without saving.

JsonMaterialStateStore uses deterministic SHA-256 filenames derived from the machine/order pair, preventing path traversal and sanitized-name collisions. It flushes and fsyncs a temporary file in the state directory before atomic replacement; a failed write leaves the previous primary file intact. Missing state returns None; malformed JSON or state raises ValueError with file context. Use an ignored runtime directory and one polling writer per state key; atomic replacement does not coordinate concurrent read-modify-write cycles. Files from the earlier uncommitted sanitized-filename prototype are not loaded under the new names. The gateway rejects ambiguous open runs, naive windows, and missing columns or malformed records instead of silently skipping them.

Continuous material polling

The committed K7 configuration declares k7-fiber-consumption, a material_consumption calculation at version "1". It selects Stundenleistung Anlage (kg/h) as the rate and Geschwindigkeit Gesamtanlage (m/min) as the gate, integrating only when the gate is strictly above 0.5, with a maximum sample gap of 20 seconds. Anlage läuft is not the primary gate. The output is material_consumption in kg, grouped by production_order; production orders remain opaque strings, including leading zeros and whitespace.

Install the declared dependencies in the repository venv, then run from the repository root:

.venv/bin/python -m pip install -e '.[dev]'
.venv/bin/python -m production_analytics.cli run material-poll \
  --config config/k7-material-consumption.yaml \
  --calculation-id k7-fiber-consumption \
  --poll-interval-seconds 10 \
  --state-directory data/state/material \
  --secrets-file secrets/enlyze.env

The installed production-analytics run material-poll command is equivalent. The interval, state directory, and secrets path above are defaults. Paths are relative to the current working directory. Existing ENLYZE_BASE_URL, ENLYZE_API_KEY, and ENLYZE_HTTP_TIMEOUT_SECONDS handling is reused; values in the secrets file override environment values. Keys are never CLI arguments.

Keep calculation fields in YAML and technical runtime settings in CLI options. The loader requires every field shown in the K7 file, rejects unknown fields, duplicate keys/ids, unsupported types/versions/outputs, and invalid or non-finite numbers. Numeric fields must be YAML numbers, and version must be quoted. Multiple material calculations may share a file, but then --calculation-id is required. All entries are validated before selection. The unrelated config/calculations.example.yaml remains illustrative and is not executable by this material-only runner.

The foreground runner polls once immediately using UTC, then sleeps for the configured positive interval after each completed cycle, including failures. A slow poll delays the next cycle; there is no overlap or catch-up scheduling. Ctrl-C/SIGINT exits cleanly. Successful cycles print machine, opaque order, cumulative kg, integrated running seconds, and cycle-start timestamp. No open run (or no eligible window) is normal and visible. Errors on stderr identify state loading/saving or ENLYZE/gateway/polling and the exception class; arbitrary exception text and response bodies are withheld to avoid leaking credentials. Startup/configuration failures return non-zero; cycle failures retry after the delay without replacing checkpoints with empty state.

Checkpoints live in ignored data/state/material/ by default, with one hashed JSON filename per machine/order pair. Run one polling writer per state key; there is no multi-process locking. Use a separate state directory when changing calculation parameters or running a different calculation for the same machine and order: checkpoints are not namespaced by calculation id/version. These files are integration checkpoints, not a Grafana metric store.

Each successful non-empty poll writes one cumulative snapshot to PostgreSQL before printing its result. JSON remains the restart checkpoint; PostgreSQL stores derived time-series snapshots only. ENLYZE remains the raw-data source of truth. Database write failures emit a concise error and polling continues with the next cycle; missed snapshots are not retried or backfilled and JSON state is not rolled back. The next successful snapshot includes the continuing cumulative total.

A new order's first snapshot is not guaranteed to be zero: the service processes available samples from the run start before returning, so its first total may already be non-zero. A returned zero is stored normally. Totals continue across runs for the same order; existing checkpoints for a previously seen order resume.

PostgreSQL setup

Set all five runtime settings: POSTGRES_HOST, POSTGRES_PORT, POSTGRES_DB, POSTGRES_USER, and POSTGRES_PASSWORD, using exported environment variables or the existing secrets file (file values override the environment). The runner does not automatically load .env. Missing or invalid settings fail at startup; connectivity and schema errors are reported during polling. For the existing Compose instance, use host localhost, the published port (default 5432), and production_analytics for both database and user. Set the same password in Compose's .env and the runner environment/secrets file.

Start the database and apply the repeatable schema from the repository root:

docker compose up -d timescaledb
docker compose exec -T timescaledb psql -U production_analytics -d production_analytics \
  -v ON_ERROR_STOP=1 < db/schema.sql

Wait until PostgreSQL is ready before applying the schema. Configure Grafana's PostgreSQL data source to query material_consumption_snapshots, selecting timestamp as time and consumption_kg as the cumulative value, filtered by calculation_id, machine_id, and production_order. These are ordinary PostgreSQL tables/indexes; no Timescale-specific features or hypertables are used. The primary key deduplicates calculation/machine/order/timestamp (first write wins). A short transaction opens and closes one synchronous connection per snapshot, with 10-second connection and statement timeouts. Snapshot timestamps are poll start times, not source-sample timestamps. Schema application is manual. Bento 1 bentonite source and gate selection remain intentionally undefined pending process validation; the generic integration core is unchanged and reusable.

ERP current workplace status

The read-only ERP adapter reads production-order/article context from MSSQL NV_DWH.dbo.GRAFANA_WORKPLACE_STATUS, for later live material KPIs. Credentials belong in ignored secrets/erp.env, using ERP_DB_HOST, ERP_DB_PORT, ERP_DB_NAME, ERP_DB_USER, and ERP_DB_PASSWORD. The verified endpoint is 192.168.111.24:49601, database NV_DWH, using SQL Server authentication and a read-only account. All settings are required; ports must be integers 1–65535.

from production_analytics.erp import ErpSettings, ErpWorkplaceStatusGateway

settings = ErpSettings.from_secret_file()  # File values override environment values.
status = ErpWorkplaceStatusGateway(settings).get_current_workplace_status("K7")

Alternatively, use ErpSettings.from_environment(mapping). The existing simple dotenv loader is reused; no shell content is executed. Each call opens and closes one connection, with 10-second login/query timeouts, and performs a parameterized SELECT filtered by workplace. Zero rows returns None, one row returns an immutable CurrentWorkplaceStatus, and multiple rows raise ErpReadError. Driver failures and malformed rows also raise ErpReadError with safe messages.

Only trailing whitespace is removed from workplace, production order, article number, and description; identifiers otherwise remain unchanged. Nullable fields stay None. Decimal/integer quantities are explicitly converted to finite floats (with normal floating-point precision); non-finite/overflowing values are rejected. Quantity fields use m² and remaining time uses hours, following the supplied source semantics. feedback_timestamp is preserved exactly, including a naive timezone. It represents the latest ERP feedback, and the ERP export is delayed relative to live process data; the adapter makes no freshness inference.

Automated ERP tests use mocks and require no live connectivity.

ERP context normalization

Two pure helpers in production_analytics.context prepare context for later KPIs:

  • build_enlyze_production_order(erp_production_order, format_template) -> str uses an explicit template configured by the caller per machine/workplace. The template requires exactly one literal {production_order} placeholder and no other braces; invalid templates raise ValueError. Outer ERP-order whitespace is stripped and the remaining order must contain only ASCII digits. Leading zeros are preserved. Empty/invalid orders raise ValueError. There is no default machine rule.
  • extract_nominal_width_m(article_description: str | None) -> float | None is generic across machines and conservatively reads a single <width> x <length> m pair, accepting decimal comma/point and variable whitespace. For example, Stex R 1501 C (PR) 5,80 x 50 m yields 5.8. Earlier article numbers and trailing descriptive text are ignored. Missing, malformed, multiple or chained dimension patterns return None, as do non-positive/non-finite dimensions. Signed/scientific notation and other units are deliberately unsupported; both dimensions must be positive plain numbers.

K7 currently uses the explicit template K 7-{production_order}:

k7_order_template = "K 7-{production_order}"
enlyze_order = build_enlyze_production_order("12026000815", k7_order_template)
# "K 7-12026000815"

Future machines may supply different verified templates without changing the generic helper. Template selection belongs to machine/workplace configuration or adapters; the helper contains no workplace lookup or machine-specific branch. No other machine rule is introduced here.

ENLYZE production-order identifiers remain opaque everywhere else, with exact comparison. These helpers never reverse-parse or split ENLYZE orders, including combined identifiers such as K 7-12025000074-K 7-12025000075; passing such an identifier to the ERP mapping function is rejected.

Nominal finished-product width currently comes from ERP article-description parsing because neither verified DWH view (dbo.GRAFANA_WORKPLACE_STATUS and dbo.GRAFANA_PRODUCTION_CONFIRMATION) has a dedicated width column. ENLYZE wLg1MeasuringWidth (Produktbreite) is explicitly not used as K7 nominal finished-product width: it belongs to the MAHLO measurement system and observed values differ from nominal article widths. Structured ERP product master data would be preferred when available.

These helpers are independent of database access. The service below composes them for material efficiency; no background processing is attached.

Feedback-aligned material efficiency

MaterialEfficiencyService in service.material_efficiency combines ERP good area with cumulative material consumption for one exact production order. Its evaluate_current() reads the existing ERP workplace adapter once; evaluate(status) evaluates an already supplied CurrentWorkplaceStatus. Both return an immutable MaterialEfficiencySnapshot or None when inputs are unavailable.

PostgresMaterialSnapshotRepository(settings).latest_at_or_before(...) in service.postgres_material takes keyword arguments calculation_id, machine_id, production_order, and an aware timestamp. A parameterized query selects the latest snapshot at or before ERP feedback time, matching calculation, machine, and mapped order exactly. It returns MaterialConsumptionSnapshot or None. Run IDs do not restrict this lookup: stored consumption is already cumulative across the order's runs. The service neither reintegrates nor sums snapshots and does not bridge run boundaries.

The result includes workplace/machine/calculation identifiers, both order identifiers, article context, nominal width, both original source timestamps, good area, aligned consumption, material_consumption_kg_per_m2 = consumption_kg / good_quantity_m2, and material_consumption_g_per_m2 = material_consumption_kg_per_m2 * 1000. Nominal width comes from generic ERP description parsing and is context only; ERP good square metres remain authoritative even when width cannot be parsed.

Missing ERP status, missing aligned material, missing/non-positive/non-finite good area, invalid consumption (negative or non-finite), or non-finite computed ratios produce None. Zero consumption is valid. Configuration errors, invalid timestamp alignment, and database failures raise rather than masquerading as missing data.

All machine/workplace settings are supplied explicitly. Example using the existing K7 material calculation (with previously constructed ERP gateway and PG settings):

from zoneinfo import ZoneInfo

from production_analytics.service.material_efficiency import MaterialEfficiencyService
from production_analytics.service.postgres_material import PostgresMaterialSnapshotRepository

service = MaterialEfficiencyService(
    erp_gateway, PostgresMaterialSnapshotRepository(postgres_settings),
    workplace="K7",
    machine_id="c220f95c-a65e-4cb7-99b7-0626d6c7508c",
    calculation_id="k7-fiber-consumption",
    format_template="K 7-{production_order}",
    # Supply only after verifying the ERP timestamp's source timezone:
    erp_timezone=ZoneInfo("Europe/Berlin"),
)
result = service.evaluate_current()

Aware ERP timestamps are compared as UTC instants. Naive ERP timestamps require explicit erp_timezone; there is no inferred system/database timezone. Ambiguous or nonexistent local times during DST transitions are rejected. The original ERP timestamp (including naivety) and selected material timestamp are preserved in the result; retain the source timezone configuration alongside it for interpretation.

This alignment avoids knowingly including future material consumption, but does not eliminate ERP roll-feedback timing uncertainty (feedback may be roughly one roll ahead or otherwise offset). Live values are plausibility indicators; final production-order values are more meaningful as relative timing error diminishes. There is no interpolation, lag correction, smoothing, freshness threshold, or estimated timestamp. Historical bootstrap snapshots are usable only when their stored timestamps satisfy the cutoff; a later cumulative total cannot reconstruct an earlier value. Reliable final evaluation requires retaining final ERP feedback; the current-workplace adapter alone does not provide historical completed orders.

Call on changed ERP feedback; no new runner or polling loop is provided. Results are returned in memory without modifying earlier results or material snapshots. No schema, Grafana, or timeseries changes are needed. Daily per-machine 24h reporting and aggregation remain future work. Repository SQL selection tests use an in-memory SQLite fixture with driver transport adaptation; they require no live PostgreSQL.

Peak-cycle detection

PeakCycleDetector is a pure calculation-domain component for roll length, roll weight, and comparable sawtooth/batch signals. It retains the current maximum and emits it only after the value is below drop_ratio * current_peak continuously for hold_seconds. Configuration is min_peak, drop_ratio (strictly between 0 and 1), non-negative hold_seconds, and positive max_sample_gap_seconds. These are per-signal configuration parameters, not globally fixed process constants. The hold uses elapsed timestamps, never a sample count; an interval longer than the configured maximum sample gap restarts a reset candidate, so a timestamp gap alone is not evidence that a signal remained below threshold. max_sample_gap_seconds is required rather than defaulted, so the source's continuity assumption is explicit for each detector configuration. Reset values need not approach zero (and can be negative), and peak magnitude can vary materially between cycles. Equal timestamps are accepted in arrival order but add no elapsed hold time; backwards timestamps are rejected.

Real-world validations use ignored local captures and are not part of the test suite. K7 roll length (m), variable bf81c547-dccf-4709-aee2-0f78366d1dfc, is the clean reference case: its 10-second capture produced exactly four plausible peaks of about 49.98 m with min_peak=10, drop_ratio=0.75, hold_seconds=60, and max_sample_gap_seconds=20. Run python scripts/validate_k7_roll_length.py when that capture is available.

Bento 2 roll weight (kg), machine 6201f3de-b940-45d7-930c-1481c5ae9e1d, variable 059c0015-9aff-4270-a4ce-f21b83f99378, is the more realistic robustness case. Its observed 10-second signal has plateaus, reset values of roughly 300–470 kg, and materially varying peak heights. With min_peak=800, drop_ratio=0.75, hold_seconds=60, and max_sample_gap_seconds=20, the extended capture produces 13 plausible completed peaks, including a legitimate approximately 1195 kg cycle. Run python scripts/validate_bento2_roll_weight.py when the ignored capture is available. Neither capture should be added to version control or automated tests.

The normal unit test suite mocks PostgreSQL and requires no live database.

See PROJECT_KNOWLEDGE.md for durable project context and docs/roadmap.md for the implementation sequence.