Files
production-analytics/PROJECT_KNOWLEDGE.md
T
2026-09-04 08:00:59 +02:00

74 lines
3.7 KiB
Markdown

# Project knowledge
## Purpose
This service is a calculation and persistence layer between ENLYZE and
Grafana. It must not become a second historian: raw process data stays in
ENLYZE and is retrieved again when historical calculations need reproduction.
## Core architectural decisions
- Python with a `src/` package layout.
- FastAPI is the preferred future service layer; no HTTP API is needed now.
- PostgreSQL with TimescaleDB is the target derived-data store.
- Grafana reads derived data from that database as an additional datasource.
- Secrets are environment variables only; no credentials or site-specific
configuration are committed.
- Calculations are modular, testable Python implementations. Configuration
declares instances; it is not a generic low-code language.
## Peak-cycle detector
The first generic calculation is a timestamp-based `PeakCycleDetector` under
`calculations`. It is transport-independent and works for roll length, roll
weight, and similar cyclic signals. It maintains the maximum in an open cycle;
after a value falls below `drop_ratio * current_peak`, it emits that maximum
only after actual subsequent below-threshold observations support a continuous
`hold_seconds` interval. `min_peak` is a generic detector setting, along with
`drop_ratio` (0 < ratio < 1), `hold_seconds` (>= 0), and an explicit positive
`max_sample_gap_seconds`. Elapsed time, never a sample count, determines
confirmation. An interval longer than the configured maximum sample gap
restarts a reset candidate, so a timestamp gap is not
continuous-below evidence. Reset values can be negative rather than zero,
equal timestamps add no elapsed time, and backwards timestamps are rejected.
The maximum gap is required rather than inferred or defaulted, making each
source's continuity assumption explicit.
K7 roll length, variable `bf81c547-dccf-4709-aee2-0f78366d1dfc` (m), is the
first real-world validation case. The tested capture is continuous on a
10-second grid through its latest available sample; its shorter-than-requested
result was caused by an end time queried in the future, not observed time-series
gaps. The raw capture remains ignored and deterministic tests use synthetic
fixtures only.
## Central domain context
Production orders connect machine, article/material, source interval, and
derived results. Metrics and events must retain calculation type/version and
source-time-range provenance.
## Known versus unknown
Known from the ENLYZE UI: machine identity, operational/downtime state,
current production order, article/material number, and process signals exist.
It is **not yet verified** how, or whether, each is exposed by the ENLYZE API.
Do not infer endpoint paths, authentication mechanisms, identifiers, paging,
timestamp semantics, or signal payloads. The active open-question list is in
[docs/enlyze-api.md](docs/enlyze-api.md).
## Exploration workflow
A deliberately generic, read-only CLI is available as
`production-analytics enlyze raw PATH`. It only makes GET requests, never
guesses endpoint schemas, and emits sanitized response metadata/body. The
operator must first obtain an authorized base URL, a verified safe path, and
the authentication method. Credentials reside in ignored `secrets/enlyze.env`;
the CLI parses simple assignments without sourcing/executing the file. Public
ENLYZE documentation verifies an `Authorization: Bearer <ENLYZE_API_KEY>`
authentication header; its value is never printed.
Unmodified captures are local in ignored `data/raw/enlyze/`. Fixtures written
to `fixtures/enlyze/` are sanitized, but must still be reviewed before commit.
The deliverable is verified API observations and sanitized fixtures, not
production calculations or database ingestion.