Compare commits

..
Author SHA1 Message Date
admin caf36a76cb Add ERP enrichment dry run 2026-07-29 14:32:15 +02:00
admin d92dc61277 Add ERP explorer 2026-07-29 14:03:06 +02:00
admin 720d789914 Merge branch 'feature/rollcalc-json-importer' 2026-07-29 13:16:00 +02:00
admin be599d28b3 Add RollCalc JSON importer 2026-07-29 11:31:17 +02:00
16 changed files with 1964 additions and 4 deletions
+3
View File
@@ -2,6 +2,9 @@
## Unreleased ## Unreleased
- ERP Enrichment Dry Run für ausgewählte ERP-Felder und bestehenden RollCalc-Artikelbestand ergänzt.
- ERP Explorer für CSV-Profiling, Dublettenanalyse und optionalen RollCalc-Abgleich ergänzt.
- RollCalc-JSON-Importer mit strikter Struktur- und Typvalidierung ergänzt.
- Initiale Projektstruktur angelegt. - Initiale Projektstruktur angelegt.
- Dokumentation fuer Architektur, Datenmodell, Datenwoerterbuch und Merge-Regeln erstellt. - Dokumentation fuer Architektur, Datenmodell, Datenwoerterbuch und Merge-Regeln erstellt.
- Reduzierte Beispieldaten und Test-Fixtures ergaenzt. - Reduzierte Beispieldaten und Test-Fixtures ergaenzt.
+4 -2
View File
@@ -40,7 +40,7 @@ Produktive ERP-Daten, lokale Quelldaten, generierte Dateien und Reports sind per
## Aktueller Stand ## Aktueller Stand
Phase 0: Struktur, Dokumentation, Fixtures, minimale Python-Bausteine und Tests. Phase 0: Struktur, Dokumentation, Fixtures, minimale Python-Bausteine, Tests, RollCalc-JSON-Importer, ERP Explorer und ERP Enrichment Dry Run.
## Offene fachliche Fragen ## Offene fachliche Fragen
@@ -48,13 +48,15 @@ Siehe `docs/data-model.md` und `docs/merge-rules.md`.
## Naechste Schritte ## Naechste Schritte
- Importer fuer RollCalc-JSON, ERP-CSV und manuelle CSV-Datei implementieren. - ERP-CSV-Importer und Import manueller CSV-Datei implementieren.
- Validierungs- und Merge-Regeln produktiv ausbauen. - Validierungs- und Merge-Regeln produktiv ausbauen.
- Reports und deterministischen Export ergaenzen. - Reports und deterministischen Export ergaenzen.
## Wichtige Dateien ## Wichtige Dateien
- `docs/data-dictionary.md` - `docs/data-dictionary.md`
- `docs/erp-enrichment.md`
- `docs/importers.md`
- `docs/merge-rules.md` - `docs/merge-rules.md`
- `tests/fixtures/` - `tests/fixtures/`
- `src/article_data_manager/` - `src/article_data_manager/`
+39
View File
@@ -86,6 +86,45 @@ article-data-manager validate data/generated/article-data.json
Die CLI-Kommandos sind in Phase 0 nur als Platzhalter vorgesehen. Die CLI-Kommandos sind in Phase 0 nur als Platzhalter vorgesehen.
## Implementierte Importer
Der RollCalc-JSON-Importer ist als Python-API verfügbar:
```python
from pathlib import Path
from article_data_manager.importers.rollcalc_json import load_rollcalc_articles
articles = load_rollcalc_articles(Path("data/source/rollcalc/article-data.json"))
```
Er validiert die bestehende RollCalc-JSON-Struktur strikt, erhält Artikelnummern als Strings und verändert die Quelldatei nicht. Details stehen in `docs/importers.md`.
## Interne Analysewerkzeuge
Der ERP Explorer profiliert einen ERP-CSV-Export, ohne daraus Import- oder Merge-Regeln abzuleiten:
```python
from pathlib import Path
from article_data_manager.tools.erp_explorer import write_erp_profile_report
write_erp_profile_report(
Path("data/source/erp/production-key-data.csv"),
rollcalc_path=Path("data/source/rollcalc/article-data.json"),
)
```
Der Textreport wird standardmäßig unter `data/reports/erp_profile_report.txt` erzeugt. Die Quelldaten werden nicht verändert.
## ERP Enrichment Dry Run
Der ERP Enrichment Dry Run ergänzt vorhandene RollCalc-Artikel um ein separates `erp`-Objekt mit ausgewählten ERP-Produktionsinformationen. Die RollCalc-Datei definiert die relevante Artikelmenge; ERP-Artikel ohne RollCalc-Entsprechung werden ignoriert.
Übernommen werden nur `ROP_PRODUCT_WIDTH`, `ROP_RATE_OF_PRODUCTION`, `SL_MINIMUM_PRODUCTION_QUANTITY` und `WPL_WORKPLACE_TEXT`. QC-Daten wie `area_weight` werden nicht aus dem ERP übernommen. `SL_PRODUCTION_SPEED` wird wegen artikelabhängiger Einheit nicht verwendet.
Details stehen in `docs/erp-enrichment.md`.
## Tests ## Tests
```bash ```bash
+94
View File
@@ -0,0 +1,94 @@
# ERP Enrichment Dry Run
Der ERP Enrichment Dry Run reichert ausschließlich vorhandene RollCalc-Artikel mit ausgewählten ERP-Feldern an. Die RollCalc-Datei definiert die relevante Artikelmenge und Reihenfolge. ERP-Artikel ohne RollCalc-Entsprechung werden ignoriert.
Es handelt sich nicht um einen vollständigen ERP-Import, keinen Merge-Prozess und keinen produktiven RollCalc-Export.
## Öffentliche API
```python
from pathlib import Path
from article_data_manager.enrichment.erp import (
enrich_rollcalc_articles_from_erp,
write_enrichment_outputs,
)
result = enrich_rollcalc_articles_from_erp(
Path("data/source/rollcalc/article-data.json"),
Path("data/source/erp/production-key-data.csv"),
)
write_enrichment_outputs(
result,
Path("data/generated/article-data.erp-enriched-dry-run.json"),
Path("data/reports/erp-enrichment-report.csv"),
)
```
## Übernommene ERP-Felder
| ERP-Feld | Zielfeld | Bedeutung | Einheit |
| --- | --- | --- | --- |
| `ROP_PRODUCT_WIDTH` | `product_width_m` | ERP-Produktbreite | m |
| `ROP_RATE_OF_PRODUCTION` | `line_speed_m_min` | Liniengeschwindigkeit | m/min |
| `SL_MINIMUM_PRODUCTION_QUANTITY` | `minimum_production_quantity` | Mindestproduktionsmenge | ERP-Originaleinheit |
| `WPL_WORKPLACE_TEXT` | `workplace` | geplanter Arbeitsplatz beziehungsweise geplante Anlage | Text |
`ROP_RATE_OF_PRODUCTION` ist die für Rollenwechselbetrachtungen maßgebliche Liniengeschwindigkeit in m/min. `SL_PRODUCTION_SPEED` wird nicht verwendet, weil dessen Einheit artikelabhängig sein kann. `kg_qm` wird nicht übernommen, weil `area_weight` aus QC beziehungsweise aus der bestehenden RollCalc-Datenquelle stammt.
`ROP_PRODUCT_WIDTH` wird als ERP-Produktbreite übernommen, ohne Sonderfälle fachlich zu interpretieren.
## Internes Dry-Run-Format
Die vorhandene RollCalc-Struktur bleibt erhalten und wird um ein separates `erp`-Objekt ergänzt:
```json
{
"nr": "214700",
"name": "Stex H 751 (Tfix 751), 6,00 x 50 m",
"thickness": 6.722,
"area_weight": 0.0,
"core_type": 0.0,
"erp": {
"product_width_m": 6.0,
"line_speed_m_min": 18.5,
"minimum_production_quantity": 5000.0,
"workplace": "Anlage 3"
}
}
```
Das `erp`-Objekt ist auch vorhanden, wenn kein ERP-Treffer existiert. Fehlende, widersprüchliche oder ungültige ERP-Werte erscheinen dort als `null`.
Vor produktiver Verwendung in RollCalc muss dessen Ladeverhalten für das zusätzliche `erp`-Objekt geprüft werden. Langfristig soll ein eigener RollCalc-Export nur die für RollCalc vorgesehenen Felder ausgeben. ERP-Produktionsinformationen müssen nicht zwangsläufig Bestandteil des RollCalc-Exports werden.
## Mehrfachtreffer und Status
ERP-Zeilen werden ausschließlich über `SL_ITEM_NO` der RollCalc-Artikelnummer `nr` zugeordnet. Artikelnummern werden nicht numerisch konvertiert oder umformatiert.
Mehrfachtreffer werden feldweise ausgewertet:
- keine ERP-Zeile: `no_erp_match`
- keine gefüllten Werte: `missing`
- genau ein Wert: `unique`
- mehrere Zeilen mit identischem Wert: `same_value_multiple_rows`
- mehrere unterschiedliche Werte: `conflict`
- nicht interpretierbarer numerischer Wert: `invalid`
Leere Werte zählen nicht als eigener Konfliktwert. Konflikte und ungültige Werte werden berichtet und nicht automatisch aufgelöst.
Artikelstatus:
- `complete`: alle vier Felder haben `unique` oder `same_value_multiple_rows`
- `partial`: mindestens ein Feld ist `missing`, aber kein Feld ist `conflict` oder `invalid`
- `conflict`: mindestens ein Feld ist `conflict` oder `invalid`
- `no_erp_match`: keine ERP-Zeile zur RollCalc-Artikelnummer
## Ausgaben
Standardpfade:
- `data/generated/article-data.erp-enriched-dry-run.json`
- `data/reports/erp-enrichment-report.csv`
Die Quelldateien werden nicht verändert. Produktive Dry-Run-Ergebnisse werden nicht versioniert.
+81
View File
@@ -0,0 +1,81 @@
# Importer
## RollCalc-JSON-Importer
Der RollCalc-JSON-Importer liest die bestehende RollCalc-Datei `article-data.json` ein und überführt gültige Datensätze in typisierte `RollCalcArticle`-Objekte.
Öffentliche API:
```python
from pathlib import Path
from article_data_manager.importers.rollcalc_json import load_rollcalc_articles
articles = load_rollcalc_articles(Path("data/source/rollcalc/article-data.json"))
```
## Eingabeformat
Erwartet wird eine UTF-8-Datei mit einem JSON-Array. Jeder Array-Eintrag muss ein JSON-Objekt sein.
Pflichtfelder:
- `nr`
- `name`
- `thickness`
- `area_weight`
- `core_type`
Datentypen:
- `nr`: String, nicht leer, ohne führende oder nachgestellte Leerzeichen
- `name`: String
- `thickness`: JSON-Zahl, kein Boolean
- `area_weight`: JSON-Zahl, kein Boolean
- `core_type`: JSON-Zahl, kein Boolean
Integer-Zahlen aus JSON werden intern als Python-`float` gespeichert, sofern sie in numerischen Feldern stehen. Strings wie `"6.722"`, `null` und Booleans werden nicht als Zahlen akzeptiert.
## Unbekannte Felder
Zusätzliche unbekannte Felder werden akzeptiert und ignoriert. Sie werden nicht in das aktuelle interne Modell übernommen. Dadurch bleibt der Importer kompatibel mit möglichen RollCalc-Erweiterungen, validiert die bekannten Felder aber weiterhin strikt.
## Dubletten
Doppelte Artikelnummern innerhalb einer Datei sind ein Validierungsfehler. Maßgeblich ist die exakte Stringdarstellung, daher sind `"00001"` und `"1"` unterschiedliche Artikelnummern.
## Fehlerverhalten
Der Importer verwendet eigene Exceptions:
- `RollCalcImportError`
- `RollCalcFileError`
- `RollCalcJsonSyntaxError`
- `RollCalcValidationError`
Fehlermeldungen enthalten Datei, Datensatzindex und Feldname, soweit anwendbar. Vollständige Datensätze werden nicht in Fehlermeldungen ausgegeben.
## Quelldatei
Der Importer liest die Quelldatei nur. Er schreibt, verändert oder repariert die Datei nicht und erzeugt keine Ausgabe auf stdout oder stderr.
## ERP Explorer
Der ERP Explorer unter `article_data_manager.tools.erp_explorer` ist ein internes Analysewerkzeug und kein ERP-Importer. Er liest einen ERP-CSV-Export, profiliert alle Spalten, zählt doppelte Artikelnummern in `SL_ITEM_NO` und kann optional RollCalc-Artikelnummern gegen den ERP-Export abgleichen.
Öffentliche API:
```python
from pathlib import Path
from article_data_manager.tools.erp_explorer import profile_erp_csv, write_erp_profile_report
profile = profile_erp_csv(Path("data/source/erp/production-key-data.csv"))
write_erp_profile_report(Path("data/source/erp/production-key-data.csv"))
```
Der Report wird standardmäßig als `data/reports/erp_profile_report.txt` geschrieben. Der Explorer verändert weder ERP-CSV noch RollCalc-JSON und erzeugt keine `article-data.json`.
## ERP Enrichment Dry Run
Der ERP Enrichment Dry Run ist unter `docs/erp-enrichment.md` dokumentiert. Er ist kein vollständiger ERP-Importer, sondern erzeugt ein internes Dry-Run-Artefakt mit separatem `erp`-Objekt und einen CSV-Report.
@@ -0,0 +1,2 @@
"""Read-only enrichment dry runs."""
+355
View File
@@ -0,0 +1,355 @@
"""Dry-run enrichment of existing RollCalc articles with selected ERP fields."""
from __future__ import annotations
import csv
import json
from collections import Counter, defaultdict
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any, Literal
from article_data_manager.importers.rollcalc_json import load_rollcalc_articles
ERP_ARTICLE_NUMBER_FIELD = "SL_ITEM_NO"
DEFAULT_JSON_OUTPUT_PATH = Path("data/generated/article-data.erp-enriched-dry-run.json")
DEFAULT_REPORT_OUTPUT_PATH = Path("data/reports/erp-enrichment-report.csv")
FieldStatus = Literal[
"unique",
"same_value_multiple_rows",
"missing",
"conflict",
"invalid",
"no_erp_match",
]
ArticleStatus = Literal["complete", "partial", "conflict", "no_erp_match"]
@dataclass(frozen=True, slots=True)
class ErpFieldRule:
"""Mapping rule for one explicitly approved ERP field."""
erp_field: str
output_field: str
numeric: bool
ERP_FIELD_RULES: tuple[ErpFieldRule, ...] = (
ErpFieldRule("ROP_PRODUCT_WIDTH", "product_width_m", True),
ErpFieldRule("ROP_RATE_OF_PRODUCTION", "line_speed_m_min", True),
ErpFieldRule("SL_MINIMUM_PRODUCTION_QUANTITY", "minimum_production_quantity", True),
ErpFieldRule("WPL_WORKPLACE_TEXT", "workplace", False),
)
@dataclass(frozen=True, slots=True)
class EnrichedField:
"""Resolved dry-run value and status for one ERP target field."""
value: float | str | None
status: FieldStatus
@dataclass(frozen=True, slots=True)
class EnrichedArticle:
"""One RollCalc article with separate ERP dry-run data."""
nr: str
name: str
source: dict[str, Any]
erp_match_count: int
fields: dict[str, EnrichedField]
status: ArticleStatus
def to_output_dict(self) -> dict[str, Any]:
"""Return the source article plus the separate `erp` object."""
enriched = dict(self.source)
enriched["erp"] = {
rule.output_field: self.fields[rule.output_field].value for rule in ERP_FIELD_RULES
}
return enriched
@dataclass(frozen=True, slots=True)
class FieldSummary:
"""Status counts for one output field."""
field_name: str
unique: int = 0
same_value_multiple_rows: int = 0
missing: int = 0
conflict: int = 0
invalid: int = 0
no_erp_match: int = 0
@dataclass(frozen=True, slots=True)
class EnrichmentSummary:
"""Compact summary of one enrichment dry run."""
rollcalc_article_count: int
with_erp_match_count: int
without_erp_match_count: int
complete_count: int
partial_count: int
conflict_count: int
field_summaries: dict[str, FieldSummary]
@dataclass(frozen=True, slots=True)
class EnrichmentResult:
"""Complete dry-run enrichment result."""
articles: tuple[EnrichedArticle, ...]
summary: EnrichmentSummary
@dataclass(frozen=True, slots=True)
class _CollectedValues:
valid_values: list[float | str] = field(default_factory=list)
invalid_values: list[str] = field(default_factory=list)
def enrich_rollcalc_articles_from_erp(
rollcalc_path: Path,
erp_csv_path: Path,
) -> EnrichmentResult:
"""Enrich existing RollCalc articles with selected ERP fields as a dry run.
The RollCalc file defines the output article set and order. ERP rows without
matching RollCalc article number are ignored. Source files are read only.
"""
rollcalc_articles = load_rollcalc_articles(Path(rollcalc_path))
raw_rollcalc_articles = _load_raw_rollcalc_articles(Path(rollcalc_path))
erp_rows_by_article_number = _load_erp_rows_by_article_number(Path(erp_csv_path))
enriched_articles: list[EnrichedArticle] = []
for article, source in zip(rollcalc_articles, raw_rollcalc_articles, strict=True):
matching_rows = erp_rows_by_article_number.get(article.nr, [])
fields = {
rule.output_field: _resolve_field(rule, matching_rows)
for rule in ERP_FIELD_RULES
}
enriched_articles.append(
EnrichedArticle(
nr=article.nr,
name=article.name,
source=source,
erp_match_count=len(matching_rows),
fields=fields,
status=_resolve_article_status(len(matching_rows), fields),
)
)
articles_tuple = tuple(enriched_articles)
return EnrichmentResult(
articles=articles_tuple,
summary=_build_summary(articles_tuple),
)
def write_enrichment_outputs(
result: EnrichmentResult,
json_path: Path,
report_path: Path,
) -> None:
"""Write deterministic dry-run JSON and CSV report files."""
json_destination = Path(json_path)
report_destination = Path(report_path)
json_destination.parent.mkdir(parents=True, exist_ok=True)
report_destination.parent.mkdir(parents=True, exist_ok=True)
output_articles = [article.to_output_dict() for article in result.articles]
json_destination.write_text(
json.dumps(output_articles, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
_write_report(result, report_destination)
def run_enrichment_dry_run(
rollcalc_path: Path,
erp_csv_path: Path,
*,
json_path: Path = DEFAULT_JSON_OUTPUT_PATH,
report_path: Path = DEFAULT_REPORT_OUTPUT_PATH,
) -> EnrichmentResult:
"""Run the dry-run enrichment and write the standard output artifacts."""
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_csv_path)
write_enrichment_outputs(result, json_path, report_path)
return result
def _load_raw_rollcalc_articles(path: Path) -> list[dict[str, Any]]:
payload = json.loads(path.read_text(encoding="utf-8"))
return [dict(item) for item in payload]
def _load_erp_rows_by_article_number(path: Path) -> dict[str, list[dict[str, str]]]:
rows_by_article_number: dict[str, list[dict[str, str]]] = defaultdict(list)
with path.open("r", encoding="utf-8-sig", newline="") as csv_file:
sample = csv_file.read(4096)
csv_file.seek(0)
dialect = csv.Sniffer().sniff(sample, delimiters=",;\t")
reader = csv.DictReader(csv_file, dialect=dialect)
for row in reader:
article_number = row.get(ERP_ARTICLE_NUMBER_FIELD, "")
if article_number != "":
rows_by_article_number[article_number].append(row)
return dict(rows_by_article_number)
def _resolve_field(rule: ErpFieldRule, matching_rows: list[dict[str, str]]) -> EnrichedField:
if not matching_rows:
return EnrichedField(value=None, status="no_erp_match")
collected = _collect_values(rule, matching_rows)
if collected.invalid_values:
return EnrichedField(value=None, status="invalid")
if not collected.valid_values:
return EnrichedField(value=None, status="missing")
unique_values = set(collected.valid_values)
if len(unique_values) > 1:
return EnrichedField(value=None, status="conflict")
value = collected.valid_values[0]
status: FieldStatus = (
"same_value_multiple_rows" if len(collected.valid_values) > 1 else "unique"
)
return EnrichedField(value=value, status=status)
def _collect_values(rule: ErpFieldRule, matching_rows: list[dict[str, str]]) -> _CollectedValues:
valid_values: list[float | str] = []
invalid_values: list[str] = []
for row in matching_rows:
raw_value = row.get(rule.erp_field, "")
if raw_value == "":
continue
if rule.numeric:
parsed_value = _parse_erp_number(raw_value)
if parsed_value is None:
invalid_values.append(raw_value)
else:
valid_values.append(parsed_value)
else:
valid_values.append(raw_value)
return _CollectedValues(valid_values=valid_values, invalid_values=invalid_values)
def _parse_erp_number(value: str) -> float | None:
normalized = value.strip()
if normalized == "":
return None
if normalized.count(",") + normalized.count(".") > 1:
return None
normalized = normalized.replace(",", ".")
try:
return float(normalized)
except ValueError:
return None
def _resolve_article_status(
erp_match_count: int,
fields: dict[str, EnrichedField],
) -> ArticleStatus:
if erp_match_count == 0:
return "no_erp_match"
field_statuses = {field.status for field in fields.values()}
if "conflict" in field_statuses or "invalid" in field_statuses:
return "conflict"
if field_statuses <= {"unique", "same_value_multiple_rows"}:
return "complete"
return "partial"
def _build_summary(articles: tuple[EnrichedArticle, ...]) -> EnrichmentSummary:
article_status_counts = Counter(article.status for article in articles)
field_summaries = {
rule.output_field: _build_field_summary(rule.output_field, articles)
for rule in ERP_FIELD_RULES
}
return EnrichmentSummary(
rollcalc_article_count=len(articles),
with_erp_match_count=sum(1 for article in articles if article.erp_match_count > 0),
without_erp_match_count=sum(1 for article in articles if article.erp_match_count == 0),
complete_count=article_status_counts["complete"],
partial_count=article_status_counts["partial"],
conflict_count=article_status_counts["conflict"],
field_summaries=field_summaries,
)
def _build_field_summary(
field_name: str,
articles: tuple[EnrichedArticle, ...],
) -> FieldSummary:
counts = Counter(article.fields[field_name].status for article in articles)
return FieldSummary(
field_name=field_name,
unique=counts["unique"],
same_value_multiple_rows=counts["same_value_multiple_rows"],
missing=counts["missing"],
conflict=counts["conflict"],
invalid=counts["invalid"],
no_erp_match=counts["no_erp_match"],
)
def _write_report(result: EnrichmentResult, report_path: Path) -> None:
with report_path.open("w", encoding="utf-8-sig", newline="") as csv_file:
writer = csv.DictWriter(csv_file, fieldnames=_report_fieldnames(), delimiter=";")
writer.writeheader()
for article in result.articles:
writer.writerow(_report_row(article))
def _report_fieldnames() -> list[str]:
return [
"nr",
"name",
"erp_match_count",
"status",
"product_width_m",
"product_width_status",
"line_speed_m_min",
"line_speed_status",
"minimum_production_quantity",
"minimum_production_quantity_status",
"workplace",
"workplace_status",
]
def _report_row(article: EnrichedArticle) -> dict[str, str | int]:
return {
"nr": article.nr,
"name": article.name,
"erp_match_count": article.erp_match_count,
"status": article.status,
"product_width_m": _format_report_value(article.fields["product_width_m"].value),
"product_width_status": article.fields["product_width_m"].status,
"line_speed_m_min": _format_report_value(article.fields["line_speed_m_min"].value),
"line_speed_status": article.fields["line_speed_m_min"].status,
"minimum_production_quantity": _format_report_value(
article.fields["minimum_production_quantity"].value
),
"minimum_production_quantity_status": article.fields[
"minimum_production_quantity"
].status,
"workplace": _format_report_value(article.fields["workplace"].value),
"workplace_status": article.fields["workplace"].status,
}
def _format_report_value(value: float | str | None) -> str:
if value is None:
return ""
return str(value)
@@ -0,0 +1,161 @@
"""Importer for the existing RollCalc `article-data.json` source file."""
from __future__ import annotations
import json
from json import JSONDecodeError
from pathlib import Path
from typing import Any
from article_data_manager.models import RollCalcArticle
REQUIRED_FIELDS: tuple[str, ...] = ("nr", "name", "thickness", "area_weight", "core_type")
NUMERIC_FIELDS: frozenset[str] = frozenset({"thickness", "area_weight", "core_type"})
class RollCalcImportError(Exception):
"""Base class for RollCalc import failures."""
class RollCalcFileError(RollCalcImportError):
"""Raised when the source file cannot be read as a regular UTF-8 file."""
class RollCalcJsonSyntaxError(RollCalcImportError):
"""Raised when the source file does not contain syntactically valid JSON."""
class RollCalcValidationError(RollCalcImportError):
"""Raised when valid JSON does not match the expected RollCalc structure."""
def load_rollcalc_articles(path: Path) -> list[RollCalcArticle]:
"""Load and validate an existing RollCalc JSON article file.
The source must be a UTF-8 JSON array. Known RollCalc fields are validated
strictly, unknown object fields are accepted and ignored for forward
compatibility. Article numbers are never trimmed, reformatted, or interpreted
numerically; leading or trailing whitespace is treated as invalid input.
"""
source_path = Path(path)
payload = _read_json(source_path)
if not isinstance(payload, list):
raise RollCalcValidationError(
f"{source_path}: RollCalc JSON root must be an array, got {_type_name(payload)}."
)
articles: list[RollCalcArticle] = []
seen_article_numbers: dict[str, int] = {}
for index, item in enumerate(payload):
article = _parse_article(source_path, index, item)
first_index = seen_article_numbers.get(article.nr)
if first_index is not None:
raise RollCalcValidationError(
f"{source_path}: Duplicate article number {article.nr!r} at index {index}; "
f"first occurrence at index {first_index}."
)
seen_article_numbers[article.nr] = index
articles.append(article)
return articles
def _read_json(path: Path) -> Any:
if not path.exists():
raise RollCalcFileError(f"{path}: RollCalc JSON file does not exist.")
if not path.is_file():
raise RollCalcFileError(f"{path}: RollCalc JSON path must be a regular file.")
try:
content = path.read_text(encoding="utf-8")
except UnicodeDecodeError as exc:
raise RollCalcFileError(f"{path}: RollCalc JSON file must be readable as UTF-8.") from exc
try:
return json.loads(content, parse_constant=_reject_non_standard_json_constant)
except JSONDecodeError as exc:
raise RollCalcJsonSyntaxError(
f"{path}: invalid JSON at line {exc.lineno}, column {exc.colno}: {exc.msg}."
) from exc
except ValueError as exc:
raise RollCalcJsonSyntaxError(f"{path}: invalid JSON: {exc}.") from exc
def _reject_non_standard_json_constant(value: str) -> None:
raise ValueError(f"non-standard numeric constant {value!r} is not allowed")
def _parse_article(path: Path, index: int, item: Any) -> RollCalcArticle:
if not isinstance(item, dict):
raise RollCalcValidationError(
f"{path}: Invalid RollCalc article at index {index}: "
f"entry must be an object, got {_type_name(item)}."
)
for field_name in REQUIRED_FIELDS:
if field_name not in item:
raise RollCalcValidationError(
f"{path}: Invalid RollCalc article at index {index}: "
f"missing required field {field_name!r}."
)
nr = _require_string(path, index, "nr", item["nr"])
if nr == "":
raise RollCalcValidationError(
f"{path}: Invalid RollCalc article at index {index}: field 'nr' must not be empty."
)
if nr != nr.strip():
raise RollCalcValidationError(
f"{path}: Invalid RollCalc article at index {index}: "
"field 'nr' must not contain leading or trailing whitespace."
)
return RollCalcArticle.from_validated_values(
nr=nr,
name=_require_string(path, index, "name", item["name"]),
thickness=_require_number(path, index, "thickness", item["thickness"]),
area_weight=_require_number(path, index, "area_weight", item["area_weight"]),
core_type=_require_number(path, index, "core_type", item["core_type"]),
)
def _require_string(path: Path, index: int, field_name: str, value: Any) -> str:
if not isinstance(value, str):
raise RollCalcValidationError(
f"{path}: Invalid RollCalc article at index {index}: field {field_name!r} "
f"must be a string, got {_type_name(value)} ({_format_value(value)})."
)
return value
def _require_number(path: Path, index: int, field_name: str, value: Any) -> int | float:
if not _is_json_number(value):
raise RollCalcValidationError(
f"{path}: Invalid RollCalc article at index {index}: field {field_name!r} "
f"must be a JSON number, got {_type_name(value)} ({_format_value(value)})."
)
return value
def _is_json_number(value: Any) -> bool:
return isinstance(value, int | float) and not isinstance(value, bool)
def _type_name(value: Any) -> str:
if value is None:
return "null"
if isinstance(value, bool):
return "bool"
if isinstance(value, list):
return "array"
if isinstance(value, dict):
return "object"
return type(value).__name__
def _format_value(value: Any, *, max_length: int = 80) -> str:
formatted = repr(value)
if len(formatted) > max_length:
return formatted[: max_length - 3] + "..."
return formatted
+2 -2
View File
@@ -1,5 +1,5 @@
"""Domain models.""" """Domain models."""
from article_data_manager.models.article import Article from article_data_manager.models.article import Article, RollCalcArticle
__all__ = ["Article"] __all__ = ["Article", "RollCalcArticle"]
@@ -1,6 +1,7 @@
"""Canonical article model for Phase 0 assumptions.""" """Canonical article model for Phase 0 assumptions."""
from dataclasses import dataclass from dataclasses import dataclass
from typing import Any
@dataclass(frozen=True, slots=True) @dataclass(frozen=True, slots=True)
@@ -24,3 +25,60 @@ class Article:
raise TypeError("Article number 'nr' must be a string.") raise TypeError("Article number 'nr' must be a string.")
if self.nr == "": if self.nr == "":
raise ValueError("Article number 'nr' must not be empty.") raise ValueError("Article number 'nr' must not be empty.")
@dataclass(frozen=True, slots=True)
class RollCalcArticle:
"""Article shape imported from the existing RollCalc JSON source.
This model preserves the currently observed RollCalc fields. The numeric
`core_type` value is intentionally not translated because its meaning is
still fachlich offen.
"""
nr: str
name: str
thickness: float
area_weight: float
core_type: float
def __post_init__(self) -> None:
if not isinstance(self.nr, str):
raise TypeError("Article number 'nr' must be a string.")
if self.nr == "":
raise ValueError("Article number 'nr' must not be empty.")
@classmethod
def from_validated_values(
cls,
*,
nr: str,
name: str,
thickness: int | float,
area_weight: int | float,
core_type: int | float,
) -> "RollCalcArticle":
"""Create an article after importer validation.
Integer JSON numbers are stored as floats to keep numeric model fields
consistent without changing their value.
"""
return cls(
nr=nr,
name=name,
thickness=float(thickness),
area_weight=float(area_weight),
core_type=float(core_type),
)
def to_dict(self) -> dict[str, Any]:
"""Return the imported RollCalc fields as a plain dictionary."""
return {
"nr": self.nr,
"name": self.name,
"thickness": self.thickness,
"area_weight": self.area_weight,
"core_type": self.core_type,
}
@@ -0,0 +1,2 @@
"""Internal development tools."""
@@ -0,0 +1,277 @@
"""Read-only ERP CSV explorer for data profiling."""
from __future__ import annotations
import csv
from collections import Counter
from dataclasses import dataclass
from pathlib import Path
from article_data_manager.importers.rollcalc_json import load_rollcalc_articles
DEFAULT_ARTICLE_NUMBER_COLUMN = "SL_ITEM_NO"
DEFAULT_REPORT_PATH = Path("data/reports/erp_profile_report.txt")
@dataclass(frozen=True, slots=True)
class ColumnProfile:
"""Profile information for one CSV column."""
name: str
filled_count: int
empty_count: int
distinct_count: int
inferred_type: str
@dataclass(frozen=True, slots=True)
class DuplicateArticleNumber:
"""Repeated ERP article number and its exact occurrence count."""
nr: str
count: int
@dataclass(frozen=True, slots=True)
class ArticleNumberProfile:
"""Profile information for an ERP article number column."""
column_name: str
distinct_count: int
duplicate_count: int
duplicates: tuple[DuplicateArticleNumber, ...]
@dataclass(frozen=True, slots=True)
class RollCalcComparison:
"""Simple ERP/RollCalc article number coverage comparison."""
rollcalc_article_count: int
found_count: int
missing_count: int
missing_article_numbers: tuple[str, ...]
@dataclass(frozen=True, slots=True)
class ErpProfile:
"""Complete read-only profile of one ERP CSV export."""
file_name: str
row_count: int
column_count: int
columns: tuple[ColumnProfile, ...]
article_numbers: ArticleNumberProfile
rollcalc_comparison: RollCalcComparison | None = None
def profile_erp_csv(
csv_path: Path,
*,
article_number_column: str = DEFAULT_ARTICLE_NUMBER_COLUMN,
rollcalc_path: Path | None = None,
) -> ErpProfile:
"""Profile an ERP CSV export without interpreting or changing its source data."""
source_path = Path(csv_path)
rows, fieldnames = _read_csv(source_path)
column_profiles = tuple(
_profile_column(name, [row.get(name, "") for row in rows]) for name in fieldnames
)
article_number_profile = _profile_article_numbers(rows, article_number_column)
erp_article_numbers = {row.get(article_number_column, "") for row in rows}
comparison = (
_compare_rollcalc_articles(rollcalc_path, erp_article_numbers)
if rollcalc_path is not None
else None
)
return ErpProfile(
file_name=source_path.name,
row_count=len(rows),
column_count=len(fieldnames),
columns=column_profiles,
article_numbers=article_number_profile,
rollcalc_comparison=comparison,
)
def write_erp_profile_report(
csv_path: Path,
*,
report_path: Path = DEFAULT_REPORT_PATH,
article_number_column: str = DEFAULT_ARTICLE_NUMBER_COLUMN,
rollcalc_path: Path | None = None,
) -> ErpProfile:
"""Create a text profile report under the requested report path.
The ERP CSV and optional RollCalc JSON are read only. The only write performed
by this function is the requested text report.
"""
profile = profile_erp_csv(
csv_path,
article_number_column=article_number_column,
rollcalc_path=rollcalc_path,
)
destination = Path(report_path)
destination.parent.mkdir(parents=True, exist_ok=True)
destination.write_text(render_erp_profile_report(profile), encoding="utf-8")
return profile
def render_erp_profile_report(profile: ErpProfile) -> str:
"""Render an ERP profile as deterministic plain text."""
lines = [
"ERP Profile Report",
f"File: {profile.file_name}",
f"Rows: {profile.row_count}",
f"Columns: {profile.column_count}",
"",
"Column profiles",
]
for column in profile.columns:
lines.extend(
[
f"Column: {column.name}",
f"Filled: {column.filled_count}",
f"Empty: {column.empty_count}",
f"Distinct: {column.distinct_count}",
f"Type: {column.inferred_type}",
"",
]
)
lines.extend(
[
"Article numbers",
f"Column: {profile.article_numbers.column_name}",
f"Distinct: {profile.article_numbers.distinct_count}",
f"Duplicates: {profile.article_numbers.duplicate_count}",
"",
"Duplicate article numbers",
]
)
for duplicate in profile.article_numbers.duplicates:
lines.append(f"{duplicate.nr}\t{duplicate.count}")
if profile.rollcalc_comparison is not None:
comparison = profile.rollcalc_comparison
lines.extend(
[
"",
"RollCalc comparison",
f"RollCalc articles: {comparison.rollcalc_article_count}",
f"Found: {comparison.found_count}",
f"Missing: {comparison.missing_count}",
"Missing article numbers",
]
)
lines.extend(comparison.missing_article_numbers)
return "\n".join(lines) + "\n"
def _read_csv(path: Path) -> tuple[list[dict[str, str]], list[str]]:
with path.open("r", encoding="utf-8-sig", newline="") as csv_file:
sample = csv_file.read(4096)
csv_file.seek(0)
dialect = csv.Sniffer().sniff(sample, delimiters=",;\t")
reader = csv.DictReader(csv_file, dialect=dialect)
fieldnames = list(reader.fieldnames or [])
return list(reader), fieldnames
def _profile_column(name: str, values: list[str]) -> ColumnProfile:
filled_values = [value for value in values if not _is_empty(value)]
empty_count = len(values) - len(filled_values)
return ColumnProfile(
name=name,
filled_count=len(filled_values),
empty_count=empty_count,
distinct_count=len(set(filled_values)),
inferred_type=infer_value_type(filled_values),
)
def infer_value_type(values: list[str]) -> str:
"""Infer a simple data type for already-filled CSV values."""
if not values:
return "empty"
value_types = {_infer_single_value_type(value) for value in values}
if value_types == {"integer"}:
return "integer"
if value_types <= {"integer", "decimal"}:
return "decimal"
if value_types == {"text"}:
return "text"
return "mixed"
def _infer_single_value_type(value: str) -> str:
normalized = value.strip()
if _is_integer(normalized):
return "integer"
if _is_decimal(normalized):
return "decimal"
return "text"
def _is_integer(value: str) -> bool:
if value.startswith(("+", "-")):
value = value[1:]
return value.isdecimal()
def _is_decimal(value: str) -> bool:
if value.count(",") + value.count(".") != 1:
return False
separator = "," if "," in value else "."
left, right = value.split(separator, 1)
if left.startswith(("+", "-")):
left = left[1:]
return left.isdecimal() and right.isdecimal()
def _profile_article_numbers(
rows: list[dict[str, str]],
article_number_column: str,
) -> ArticleNumberProfile:
values = []
for row in rows:
value = row.get(article_number_column, "")
if not _is_empty(value):
values.append(value)
counts = Counter(values)
duplicates = tuple(
DuplicateArticleNumber(nr=nr, count=count)
for nr, count in sorted(counts.items())
if count > 1
)
return ArticleNumberProfile(
column_name=article_number_column,
distinct_count=len(counts),
duplicate_count=len(duplicates),
duplicates=duplicates,
)
def _compare_rollcalc_articles(
rollcalc_path: Path,
erp_article_numbers: set[str],
) -> RollCalcComparison:
rollcalc_articles = load_rollcalc_articles(Path(rollcalc_path))
rollcalc_numbers = tuple(article.nr for article in rollcalc_articles)
missing_article_numbers = tuple(nr for nr in rollcalc_numbers if nr not in erp_article_numbers)
return RollCalcComparison(
rollcalc_article_count=len(rollcalc_numbers),
found_count=len(rollcalc_numbers) - len(missing_article_numbers),
missing_count=len(missing_article_numbers),
missing_article_numbers=missing_article_numbers,
)
def _is_empty(value: str | None) -> bool:
return value is None or value == ""
@@ -0,0 +1,77 @@
import csv
import json
from pathlib import Path
from article_data_manager.enrichment.erp import (
enrich_rollcalc_articles_from_erp,
write_enrichment_outputs,
)
def test_enrichment_dry_run_writes_json_and_csv_outputs(tmp_path: Path) -> None:
rollcalc_path = tmp_path / "article-data.json"
erp_path = tmp_path / "erp.csv"
json_path = tmp_path / "generated" / "article-data.erp-enriched-dry-run.json"
report_path = tmp_path / "reports" / "erp-enrichment-report.csv"
rollcalc_path.write_text(
json.dumps(
[
{
"nr": "214700",
"name": "RollCalc A",
"thickness": 6.722,
"area_weight": 750.0,
"core_type": 0.0,
},
{
"nr": "999999",
"name": "RollCalc Only",
"thickness": 1.0,
"area_weight": 0.0,
"core_type": 0.0,
},
],
ensure_ascii=False,
indent=2,
)
+ "\n",
encoding="utf-8",
)
erp_path.write_text(
"\n".join(
[
"SL_ITEM_NO,ROP_PRODUCT_WIDTH,ROP_RATE_OF_PRODUCTION,"
"SL_MINIMUM_PRODUCTION_QUANTITY,WPL_WORKPLACE_TEXT,kg_qm,SL_PRODUCTION_SPEED",
"214700,\"6,00\",\"18,5\",5000,Anlage 3,999,999",
"ERPONLY,9.99,1.0,1,Anlage X,1,1",
]
),
encoding="utf-8",
)
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path)
write_enrichment_outputs(result, json_path, report_path)
output_articles = json.loads(json_path.read_text(encoding="utf-8"))
assert [article["nr"] for article in output_articles] == ["214700", "999999"]
assert output_articles[0]["area_weight"] == 750.0
assert output_articles[0]["erp"] == {
"product_width_m": 6.0,
"line_speed_m_min": 18.5,
"minimum_production_quantity": 5000.0,
"workplace": "Anlage 3",
}
assert output_articles[1]["erp"] == {
"product_width_m": None,
"line_speed_m_min": None,
"minimum_production_quantity": None,
"workplace": None,
}
with report_path.open("r", encoding="utf-8-sig", newline="") as report_file:
rows = list(csv.DictReader(report_file, delimiter=";"))
assert [row["nr"] for row in rows] == ["214700", "999999"]
assert rows[0]["status"] == "complete"
assert rows[1]["status"] == "no_erp_match"
@@ -0,0 +1,387 @@
import csv
import json
from pathlib import Path
from article_data_manager.enrichment.erp import (
enrich_rollcalc_articles_from_erp,
write_enrichment_outputs,
)
def write_rollcalc(path: Path, article_numbers: list[str]) -> Path:
payload = [
{
"nr": nr,
"name": f"RollCalc {nr}",
"thickness": float(index + 1),
"area_weight": 100.0 + index,
"core_type": 0.0,
}
for index, nr in enumerate(article_numbers)
]
path.write_text(json.dumps(payload, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
return path
def write_erp(path: Path, rows: list[dict[str, str]]) -> Path:
fieldnames = [
"SL_ITEM_NO",
"ROP_PRODUCT_WIDTH",
"ROP_RATE_OF_PRODUCTION",
"SL_MINIMUM_PRODUCTION_QUANTITY",
"WPL_WORKPLACE_TEXT",
"SL_PRODUCTION_SPEED",
"kg_qm",
]
with path.open("w", encoding="utf-8", newline="") as csv_file:
writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
writer.writeheader()
for row in rows:
writer.writerow({field: row.get(field, "") for field in fieldnames})
return path
def test_unique_erp_match_enriches_selected_fields(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{
"SL_ITEM_NO": "214700",
"ROP_PRODUCT_WIDTH": "6,00",
"ROP_RATE_OF_PRODUCTION": "18.5",
"SL_MINIMUM_PRODUCTION_QUANTITY": "5000",
"WPL_WORKPLACE_TEXT": "Anlage 3",
}
],
)
article = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0]
assert article.status == "complete"
assert article.to_output_dict()["erp"] == {
"product_width_m": 6.0,
"line_speed_m_min": 18.5,
"minimum_production_quantity": 5000.0,
"workplace": "Anlage 3",
}
def test_no_erp_match_keeps_rollcalc_article_with_null_erp_object(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["999999"])
erp_path = write_erp(tmp_path / "erp.csv", [{"SL_ITEM_NO": "214700"}])
article = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0]
assert article.status == "no_erp_match"
assert article.erp_match_count == 0
assert article.to_output_dict()["erp"] == {
"product_width_m": None,
"line_speed_m_min": None,
"minimum_production_quantity": None,
"workplace": None,
}
assert {field.status for field in article.fields.values()} == {"no_erp_match"}
def test_multiple_erp_rows_with_identical_width_are_collapsed(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6,00"},
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6.00"},
],
)
field = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].fields[
"product_width_m"
]
assert field.value == 6.0
assert field.status == "same_value_multiple_rows"
def test_conflicting_width_is_reported_as_null_field(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6,00"},
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6.10"},
],
)
article = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0]
assert article.status == "conflict"
assert article.fields["product_width_m"].value is None
assert article.fields["product_width_m"].status == "conflict"
def test_identical_line_speeds_in_multiple_rows_are_collapsed(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_RATE_OF_PRODUCTION": "18,5"},
{"SL_ITEM_NO": "214700", "ROP_RATE_OF_PRODUCTION": "18.50"},
],
)
field = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].fields[
"line_speed_m_min"
]
assert field.value == 18.5
assert field.status == "same_value_multiple_rows"
def test_conflicting_line_speeds_are_reported(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_RATE_OF_PRODUCTION": "18,5"},
{"SL_ITEM_NO": "214700", "ROP_RATE_OF_PRODUCTION": "19,0"},
],
)
article = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0]
assert article.status == "conflict"
assert article.fields["line_speed_m_min"].status == "conflict"
def test_missing_minimum_production_quantity_is_reported(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(tmp_path / "erp.csv", [{"SL_ITEM_NO": "214700"}])
field = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].fields[
"minimum_production_quantity"
]
assert field.value is None
assert field.status == "missing"
def test_identical_workplaces_in_multiple_rows_are_collapsed(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "WPL_WORKPLACE_TEXT": "Anlage 3"},
{"SL_ITEM_NO": "214700", "WPL_WORKPLACE_TEXT": "Anlage 3"},
],
)
field = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].fields[
"workplace"
]
assert field.value == "Anlage 3"
assert field.status == "same_value_multiple_rows"
def test_different_workplaces_are_reported_as_conflict(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "WPL_WORKPLACE_TEXT": "Anlage 3"},
{"SL_ITEM_NO": "214700", "WPL_WORKPLACE_TEXT": "Anlage 4"},
],
)
field = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].fields[
"workplace"
]
assert field.value is None
assert field.status == "conflict"
def test_decimal_values_with_comma_and_point_are_supported(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700", "000123"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6,00"},
{"SL_ITEM_NO": "000123", "ROP_PRODUCT_WIDTH": "1.20"},
],
)
articles = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles
assert articles[0].fields["product_width_m"].value == 6.0
assert articles[1].fields["product_width_m"].value == 1.2
def test_invalid_numeric_value_sets_field_to_null(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "not-a-number"}],
)
article = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0]
assert article.status == "conflict"
assert article.fields["product_width_m"].value is None
assert article.fields["product_width_m"].status == "invalid"
def test_empty_erp_fields_do_not_create_conflicts(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": ""},
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6.00"},
],
)
field = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].fields[
"product_width_m"
]
assert field.value == 6.0
assert field.status == "unique"
def test_leading_zero_article_number_matches_exactly(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["000123"])
erp_path = write_erp(
tmp_path / "erp.csv",
[{"SL_ITEM_NO": "000123", "ROP_PRODUCT_WIDTH": "1.20"}],
)
article = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0]
assert article.nr == "000123"
assert article.fields["product_width_m"].value == 1.2
def test_erp_article_without_rollcalc_match_is_not_output(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "214700", "ROP_PRODUCT_WIDTH": "6.00"},
{"SL_ITEM_NO": "ERPONLY", "ROP_PRODUCT_WIDTH": "9.99"},
],
)
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path)
assert [article.nr for article in result.articles] == ["214700"]
def test_rollcalc_order_is_preserved(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["300", "100", "200"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{"SL_ITEM_NO": "100"},
{"SL_ITEM_NO": "200"},
{"SL_ITEM_NO": "300"},
],
)
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path)
assert [article.nr for article in result.articles] == ["300", "100", "200"]
def test_existing_area_weight_remains_unchanged_and_erp_weight_fields_are_ignored(
tmp_path: Path,
) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{
"SL_ITEM_NO": "214700",
"kg_qm": "999",
"SL_PRODUCTION_SPEED": "999",
"ROP_PRODUCT_WIDTH": "6.00",
}
],
)
output = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).articles[0].to_output_dict()
assert output["area_weight"] == 100.0
assert "kg_qm" not in output["erp"]
assert "SL_PRODUCTION_SPEED" not in output["erp"]
def test_source_files_remain_unchanged(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(tmp_path / "erp.csv", [{"SL_ITEM_NO": "214700"}])
original_rollcalc = rollcalc_path.read_bytes()
original_erp = erp_path.read_bytes()
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path)
write_enrichment_outputs(result, tmp_path / "generated.json", tmp_path / "report.csv")
assert rollcalc_path.read_bytes() == original_rollcalc
assert erp_path.read_bytes() == original_erp
def test_json_output_is_deterministic(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(tmp_path / "erp.csv", [{"SL_ITEM_NO": "214700"}])
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path)
first_json = tmp_path / "first.json"
second_json = tmp_path / "second.json"
write_enrichment_outputs(result, first_json, tmp_path / "first.csv")
write_enrichment_outputs(result, second_json, tmp_path / "second.csv")
assert first_json.read_text(encoding="utf-8") == second_json.read_text(encoding="utf-8")
def test_csv_report_is_deterministic_and_contains_statuses(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["214700"])
erp_path = write_erp(tmp_path / "erp.csv", [{"SL_ITEM_NO": "214700"}])
result = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path)
first_report = tmp_path / "first.csv"
second_report = tmp_path / "second.csv"
write_enrichment_outputs(result, tmp_path / "first.json", first_report)
write_enrichment_outputs(result, tmp_path / "second.json", second_report)
first_content = first_report.read_text(encoding="utf-8-sig")
assert first_content == second_report.read_text(encoding="utf-8-sig")
assert "nr;name;erp_match_count;status" in first_content
assert "missing" in first_content
def test_summary_counts_articles_and_field_statuses(tmp_path: Path) -> None:
rollcalc_path = write_rollcalc(tmp_path / "article-data.json", ["A", "B", "C"])
erp_path = write_erp(
tmp_path / "erp.csv",
[
{
"SL_ITEM_NO": "A",
"ROP_PRODUCT_WIDTH": "1.0",
"ROP_RATE_OF_PRODUCTION": "2.0",
"SL_MINIMUM_PRODUCTION_QUANTITY": "3",
"WPL_WORKPLACE_TEXT": "Anlage",
},
{"SL_ITEM_NO": "B", "ROP_PRODUCT_WIDTH": "bad"},
],
)
summary = enrich_rollcalc_articles_from_erp(rollcalc_path, erp_path).summary
assert summary.rollcalc_article_count == 3
assert summary.with_erp_match_count == 2
assert summary.without_erp_match_count == 1
assert summary.complete_count == 1
assert summary.partial_count == 0
assert summary.conflict_count == 1
assert summary.field_summaries["product_width_m"].unique == 1
assert summary.field_summaries["product_width_m"].invalid == 1
assert summary.field_summaries["product_width_m"].no_erp_match == 1
+270
View File
@@ -0,0 +1,270 @@
import json
from pathlib import Path
import pytest
from article_data_manager.importers.rollcalc_json import (
RollCalcFileError,
RollCalcJsonSyntaxError,
RollCalcValidationError,
load_rollcalc_articles,
)
def write_json(path: Path, payload: object) -> Path:
path.write_text(json.dumps(payload), encoding="utf-8")
return path
def valid_article(**overrides: object) -> dict[str, object]:
article: dict[str, object] = {
"nr": "214700",
"name": "Example Article",
"thickness": 6.722,
"area_weight": 0.0,
"core_type": 0.0,
}
article.update(overrides)
return article
def test_loads_valid_file_with_one_article(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article()])
articles = load_rollcalc_articles(path)
assert len(articles) == 1
assert articles[0].nr == "214700"
assert articles[0].name == "Example Article"
assert articles[0].thickness == 6.722
assert articles[0].area_weight == 0.0
assert articles[0].core_type == 0.0
def test_loads_valid_file_with_multiple_articles(tmp_path: Path) -> None:
path = write_json(
tmp_path / "article-data.json",
[
valid_article(nr="100001"),
valid_article(nr="100002"),
],
)
articles = load_rollcalc_articles(path)
assert [article.nr for article in articles] == ["100001", "100002"]
def test_preserves_source_order(tmp_path: Path) -> None:
path = write_json(
tmp_path / "article-data.json",
[
valid_article(nr="300003"),
valid_article(nr="100001"),
valid_article(nr="200002"),
],
)
articles = load_rollcalc_articles(path)
assert [article.nr for article in articles] == ["300003", "100001", "200002"]
def test_preserves_leading_zero_in_article_number(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(nr="000123")])
articles = load_rollcalc_articles(path)
assert articles[0].nr == "000123"
def test_numeric_integer_is_stored_as_float(tmp_path: Path) -> None:
path = write_json(
tmp_path / "article-data.json",
[valid_article(thickness=5, area_weight=350, core_type=1)],
)
article = load_rollcalc_articles(path)[0]
assert article.thickness == 5.0
assert article.area_weight == 350.0
assert article.core_type == 1.0
assert isinstance(article.thickness, float)
assert isinstance(article.area_weight, float)
assert isinstance(article.core_type, float)
def test_accepts_and_ignores_unknown_fields(tmp_path: Path) -> None:
path = write_json(
tmp_path / "article-data.json",
[valid_article(width=6.0, unknown_nested={"value": "ignored"})],
)
article = load_rollcalc_articles(path)[0]
assert not hasattr(article, "width")
assert not hasattr(article, "unknown_nested")
assert article.to_dict() == {
"nr": "214700",
"name": "Example Article",
"thickness": 6.722,
"area_weight": 0.0,
"core_type": 0.0,
}
def test_empty_json_array_returns_empty_list(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [])
assert load_rollcalc_articles(path) == []
def test_missing_file_raises_file_error(tmp_path: Path) -> None:
path = tmp_path / "missing.json"
with pytest.raises(RollCalcFileError, match="does not exist"):
load_rollcalc_articles(path)
def test_directory_path_raises_file_error(tmp_path: Path) -> None:
with pytest.raises(RollCalcFileError, match="regular file"):
load_rollcalc_articles(tmp_path)
def test_invalid_utf8_raises_file_error(tmp_path: Path) -> None:
path = tmp_path / "article-data.json"
path.write_bytes(b"\xff\xfe\xfa")
with pytest.raises(RollCalcFileError, match="UTF-8"):
load_rollcalc_articles(path)
def test_empty_file_raises_json_syntax_error(tmp_path: Path) -> None:
path = tmp_path / "article-data.json"
path.write_text("", encoding="utf-8")
with pytest.raises(RollCalcJsonSyntaxError, match="invalid JSON"):
load_rollcalc_articles(path)
def test_syntactically_invalid_json_raises_json_syntax_error(tmp_path: Path) -> None:
path = tmp_path / "article-data.json"
path.write_text("[}", encoding="utf-8")
with pytest.raises(RollCalcJsonSyntaxError, match="invalid JSON"):
load_rollcalc_articles(path)
def test_json_root_must_be_array(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", {"nr": "214700"})
with pytest.raises(RollCalcValidationError, match="root must be an array"):
load_rollcalc_articles(path)
def test_array_entry_must_be_object(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(), "invalid"])
with pytest.raises(RollCalcValidationError, match="index 1: entry must be an object"):
load_rollcalc_articles(path)
def test_missing_required_field(tmp_path: Path) -> None:
article = valid_article()
del article["name"]
path = write_json(tmp_path / "article-data.json", [article])
with pytest.raises(RollCalcValidationError, match="missing required field 'name'"):
load_rollcalc_articles(path)
def test_nr_must_be_string(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(nr=214700)])
with pytest.raises(RollCalcValidationError, match="field 'nr' must be a string, got int"):
load_rollcalc_articles(path)
def test_nr_must_not_be_empty(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(nr="")])
with pytest.raises(RollCalcValidationError, match="field 'nr' must not be empty"):
load_rollcalc_articles(path)
def test_nr_must_not_contain_surrounding_whitespace(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(nr=" 000123")])
with pytest.raises(RollCalcValidationError, match="leading or trailing whitespace"):
load_rollcalc_articles(path)
def test_name_must_be_string(tmp_path: Path) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(name=123)])
with pytest.raises(RollCalcValidationError, match="field 'name' must be a string"):
load_rollcalc_articles(path)
@pytest.mark.parametrize("field_name", ["thickness", "area_weight", "core_type"])
def test_numeric_field_must_not_be_string(tmp_path: Path, field_name: str) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(**{field_name: "6.722"})])
with pytest.raises(
RollCalcValidationError,
match=f"field '{field_name}' must be a JSON number, got str",
):
load_rollcalc_articles(path)
@pytest.mark.parametrize("field_name", ["thickness", "area_weight", "core_type"])
def test_numeric_field_must_not_be_null(tmp_path: Path, field_name: str) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(**{field_name: None})])
with pytest.raises(
RollCalcValidationError,
match=f"field '{field_name}' must be a JSON number, got null",
):
load_rollcalc_articles(path)
@pytest.mark.parametrize("field_name", ["thickness", "area_weight", "core_type"])
def test_numeric_field_must_not_be_boolean(tmp_path: Path, field_name: str) -> None:
path = write_json(tmp_path / "article-data.json", [valid_article(**{field_name: True})])
with pytest.raises(
RollCalcValidationError,
match=f"field '{field_name}' must be a JSON number, got bool",
):
load_rollcalc_articles(path)
def test_duplicate_article_number_raises_validation_error(tmp_path: Path) -> None:
path = write_json(
tmp_path / "article-data.json",
[valid_article(nr="00001"), valid_article(nr="1"), valid_article(nr="00001")],
)
with pytest.raises(
RollCalcValidationError,
match="Duplicate article number '00001' at index 2; first occurrence at index 0",
):
load_rollcalc_articles(path)
def test_multiple_errors_report_first_deterministically(tmp_path: Path) -> None:
path = write_json(
tmp_path / "article-data.json",
[
valid_article(name=123, thickness="6.722"),
valid_article(nr="214700"),
],
)
with pytest.raises(RollCalcValidationError) as exc_info:
load_rollcalc_articles(path)
message = str(exc_info.value)
assert "index 0" in message
assert "field 'name'" in message
assert "thickness" not in message
+152
View File
@@ -0,0 +1,152 @@
import json
from pathlib import Path
from article_data_manager.tools.erp_explorer import (
infer_value_type,
profile_erp_csv,
render_erp_profile_report,
write_erp_profile_report,
)
def write_text(path: Path, content: str) -> Path:
path.write_text(content, encoding="utf-8")
return path
def write_rollcalc_json(path: Path, article_numbers: list[str]) -> Path:
payload = [
{
"nr": nr,
"name": f"Article {nr}",
"thickness": 1.0,
"area_weight": 0.0,
"core_type": 0.0,
}
for nr in article_numbers
]
path.write_text(json.dumps(payload), encoding="utf-8")
return path
def test_profiles_columns(tmp_path: Path) -> None:
csv_path = write_text(
tmp_path / "erp.csv",
"\n".join(
[
"SL_ITEM_NO,ROP_PRODUCT_WIDTH,COMMENT,EMPTY_COLUMN",
"214700,\"6,00\",Alpha,",
"000123,1.20,Beta,",
"777777,,Alpha,",
]
),
)
profile = profile_erp_csv(csv_path)
assert profile.file_name == "erp.csv"
assert profile.row_count == 3
assert profile.column_count == 4
columns = {column.name: column for column in profile.columns}
assert columns["ROP_PRODUCT_WIDTH"].filled_count == 2
assert columns["ROP_PRODUCT_WIDTH"].empty_count == 1
assert columns["ROP_PRODUCT_WIDTH"].distinct_count == 2
assert columns["ROP_PRODUCT_WIDTH"].inferred_type == "decimal"
assert columns["COMMENT"].distinct_count == 2
assert columns["COMMENT"].inferred_type == "text"
assert columns["EMPTY_COLUMN"].inferred_type == "empty"
def test_infers_value_types() -> None:
assert infer_value_type([]) == "empty"
assert infer_value_type(["1", "002", "-3"]) == "integer"
assert infer_value_type(["1", "2.5", "3,75"]) == "decimal"
assert infer_value_type(["Alpha", "Beta"]) == "text"
assert infer_value_type(["1", "Alpha"]) == "mixed"
def test_detects_duplicate_article_numbers(tmp_path: Path) -> None:
csv_path = write_text(
tmp_path / "erp.csv",
"\n".join(
[
"SL_ITEM_NO,NAME",
"00001,A",
"1,B",
"00001,C",
"214700,D",
"214700,E",
]
),
)
article_numbers = profile_erp_csv(csv_path).article_numbers
assert article_numbers.distinct_count == 3
assert article_numbers.duplicate_count == 2
assert [(duplicate.nr, duplicate.count) for duplicate in article_numbers.duplicates] == [
("00001", 2),
("214700", 2),
]
def test_compares_rollcalc_articles_when_path_is_provided(tmp_path: Path) -> None:
csv_path = write_text(
tmp_path / "erp.csv",
"\n".join(
[
"SL_ITEM_NO,NAME",
"214700,A",
"000123,B",
]
),
)
rollcalc_path = write_rollcalc_json(tmp_path / "article-data.json", ["214700", "999999"])
comparison = profile_erp_csv(csv_path, rollcalc_path=rollcalc_path).rollcalc_comparison
assert comparison is not None
assert comparison.rollcalc_article_count == 2
assert comparison.found_count == 1
assert comparison.missing_count == 1
assert comparison.missing_article_numbers == ("999999",)
def test_rollcalc_comparison_is_optional(tmp_path: Path) -> None:
csv_path = write_text(
tmp_path / "erp.csv",
"\n".join(
[
"SL_ITEM_NO,NAME",
"214700,A",
]
),
)
assert profile_erp_csv(csv_path).rollcalc_comparison is None
def test_renders_and_writes_report(tmp_path: Path) -> None:
csv_path = write_text(
tmp_path / "erp.csv",
"\n".join(
[
"SL_ITEM_NO,ROP_PRODUCT_WIDTH",
"214700,\"6,00\"",
"214700,\"6,00\"",
]
),
)
report_path = tmp_path / "reports" / "erp_profile_report.txt"
profile = write_erp_profile_report(csv_path, report_path=report_path)
report = report_path.read_text(encoding="utf-8")
assert report == render_erp_profile_report(profile)
assert "ERP Profile Report" in report
assert "File: erp.csv" in report
assert "Rows: 2" in report
assert "Columns: 2" in report
assert "Column: ROP_PRODUCT_WIDTH" in report
assert "Duplicate article numbers" in report
assert "214700\t2" in report