Files
production-analytics/README.md
T

106 lines
5.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# production-analytics
`production-analytics` is a small Python service that derives production
analytics from ENLYZE data for Grafana. ENLYZE remains the authority for raw
process data; this project persists only derived metrics, events, relevant
production-order context, and calculation state.
## Status
This repository includes a read-only ENLYZE API exploration CLI. It contains
no verified ENLYZE operation wrappers, database migrations, or HTTP endpoints.
## Intended flow
```text
ENLYZE (raw data) -> retrieval boundary -> calculations -> TimescaleDB -> Grafana
\-> calculation provenance/state
```
Production orders are the primary attribution context for results. Every
persisted result will ultimately be traceable to its machine, order/article
context, source time range, and calculation implementation version.
## Development
Requires Python 3.11 or newer. Install the project and development tools:
```bash
python3 -m pip install -e '.[dev]'
pytest
ruff check .
```
Copy `.env.example` to `.env` for public/local configuration documentation.
Put real ENLYZE credentials only in `secrets/enlyze.env`, which is ignored.
The CLI reads that file by default without executing it; it accepts only simple
`KEY=VALUE` lines (quoted values are supported). Environment variables may be
used instead when appropriate. Never pass keys as command-line arguments.
The generic request command only performs `GET` requests and requires the
operator to provide an API path that is known to be safe and authorized:
```bash
production-analytics enlyze raw /verified/path --pretty
production-analytics enlyze raw /verified/path --save-fixture response.json
```
Public ENLYZE documentation verifies Bearer-token authentication. Put
`ENLYZE_API_KEY='...'` in the local secret file; the exploration client sends
it as an `Authorization: Bearer …` header and never prints that header/value.
The second command saves a sanitized, reviewable fixture under
`fixtures/enlyze/` by default. Raw captures belong in the ignored
`data/raw/enlyze/` directory and must not be committed.
The OpenAPI server URL is `https://app.enlyze.com/api/`; its operation paths
begin with `/v2/`. Set `ENLYZE_BASE_URL` to that server URL, not to an
operation path. The documented read-only time-series operation is exposed for
exploration as `production-analytics enlyze timeseries`.
The current Compose file has no application service, so it deliberately does
not pass ENLYZE credentials to TimescaleDB. A future application service should
use `env_file: ./secrets/enlyze.env` rather than copying secrets into Compose.
## Peak-cycle detection
`PeakCycleDetector` is a pure calculation-domain component for roll length,
roll weight, and comparable sawtooth/batch signals. It retains the current
maximum and emits it only after the value is below
`drop_ratio * current_peak` continuously for `hold_seconds`. Configuration is
`min_peak`, `drop_ratio` (strictly between 0 and 1), non-negative
`hold_seconds`, and positive `max_sample_gap_seconds`. These are per-signal
configuration parameters, not globally fixed process constants. The hold uses elapsed
timestamps, never a sample count; an interval longer than the configured
maximum sample gap restarts a reset candidate,
so a timestamp gap alone is not evidence that a signal remained below threshold.
`max_sample_gap_seconds` is required rather than defaulted, so the source's
continuity assumption is explicit for each detector configuration.
Reset values need not approach zero (and can be negative), and peak magnitude
can vary materially between cycles. Equal timestamps are accepted in arrival
order but add no elapsed hold time; backwards timestamps are rejected.
Real-world validations use ignored local captures and are not part of the test
suite. K7 roll length (m), variable `bf81c547-dccf-4709-aee2-0f78366d1dfc`, is
the clean reference case: its 10-second capture produced exactly four plausible
peaks of about 49.98 m with `min_peak=10`, `drop_ratio=0.75`,
`hold_seconds=60`, and `max_sample_gap_seconds=20`. Run
`python scripts/validate_k7_roll_length.py` when that capture is available.
Bento 2 roll weight (kg), machine `6201f3de-b940-45d7-930c-1481c5ae9e1d`,
variable `059c0015-9aff-4270-a4ce-f21b83f99378`, is the more realistic
robustness case. Its observed 10-second signal has plateaus, reset values of
roughly 300–470 kg, and materially varying peak heights. With
`min_peak=800`, `drop_ratio=0.75`, `hold_seconds=60`, and
`max_sample_gap_seconds=20`, the extended capture produces 13 plausible
completed peaks, including a legitimate approximately 1195 kg cycle. Run
`python scripts/validate_bento2_roll_weight.py` when the ignored capture is
available. Neither capture should be added to version control or automated
tests.
`docker compose up -d timescaledb` is an optional local database design for a
future persistence milestone. It is not required for the bootstrap tests.
See [PROJECT_KNOWLEDGE.md](PROJECT_KNOWLEDGE.md) for durable project context
and [docs/roadmap.md](docs/roadmap.md) for the implementation sequence.