diff --git a/docs/architecture.md b/docs/architecture.md index ef00cf1..480ad8e 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -265,6 +265,11 @@ Deterministic Canonicalizer: - validates and normalizes extraction objects - assigns stable source references and IDs - normalizes category names and basic field structure +- validates and normalizes action-item responsible fields against Meeting + Context when available: known participant and mentioned-person aliases are + normalized to canonical display names, while dates, locations, projects, + products, technical terms, generic process words and unknown free text are + cleared with a structured validation record - performs only safe deterministic cleanup - may group exact duplicates - preserves all source evidence @@ -277,6 +282,12 @@ Semantic Consolidator: - V0 merges semantically equivalent fact items conservatively - V0 preserves source references and evidence - V0 validates that every source fact appears exactly once +- V0 sizes its Ollama output budget from the actual fact payload instead of + using a fixed response cap for every meeting +- V0 may apply deterministic source-coverage repair after valid model JSON is + parsed: duplicate source IDs are removed after their first occurrence, empty + groups are removed and missing source facts are restored as singleton groups + from canonicalized input before strict validation runs - V0 does not process non-fact categories semantically - later versions should group content by topic, mark contradictions and uncertainty, separate durable information from transient discussion and @@ -295,6 +306,11 @@ It only reformulates the analysis results for a specific audience and purpose. Depending on the output and maturity of the implementation, a renderer may be deterministic, template-based or LLM-assisted. +LLM-assisted renderers preserve raw model output separately and write the final +output artifact only after deterministic contract validation succeeds. Renderer +post-processing may remove non-semantic wrapper text, but must not fabricate +missing semantic sections or relabel an invalid summary as a valid output view. + The planned output products are: - Working Protocol (`working_protocol.md`, Arbeitsprotokoll) diff --git a/docs/output-views.md b/docs/output-views.md index 4db3898..a545811 100644 --- a/docs/output-views.md +++ b/docs/output-views.md @@ -110,6 +110,21 @@ Characteristics: Completeness goal: optimize for recall and traceability. +Current Working Protocol V2 renderer contract: + +- output must begin exactly with `# Working Protocol` +- no explanatory preamble may appear before that heading +- content is organized by topic with `##` topic headings +- each topic may use only the supported `###` sections: `Background`, + `Decisions`, `Action Items`, `Open Questions` +- category-level report framing such as a global `## Decisions` / `## Action + Items` summary is not a valid topic-oriented Working Protocol +- raw model responses are preserved separately +- deterministic cleanup may remove leading prose before an already valid + `# Working Protocol` heading and normalize harmless heading whitespace +- `working_protocol.md` is written only after the cleaned candidate passes the + renderer contract validator + ## Concise Distribution Protocol Suggested filename: `distribution_protocol.md` diff --git a/docs/regression-bugs.md b/docs/regression-bugs.md index 9f5ae22..4b50e0a 100644 --- a/docs/regression-bugs.md +++ b/docs/regression-bugs.md @@ -476,3 +476,366 @@ Notes: This generalizes BUG-002 beyond the specific Jovana assignment case and should be evaluated against the responsibility attribution invariant. + +## BUG-008 + +ID: BUG-008 + +Title: Semantic Consolidator emits duplicate source fact coverage + +Pipeline stage: Semantic Consolidator V0 / Constraint Repair + +Severity: High + +Status: Verified + +Date discovered: 2026-08-04 + +Version first observed: `progeo_meeting_20260804_083849` + +Description: + +The Progeo end-to-end evaluation stopped during Semantic Consolidator V0 +validation because the preserved model grouping output assigned the same source +fact ID to more than one group. + +Expected behaviour: + +Every source fact ID must appear exactly once in the Semantic Consolidator V0 +fact grouping output. The validator must reject duplicate or missing source +coverage before a consolidated document is rendered. + +Actual behaviour: + +The validator correctly rejected the preserved raw consolidator response with: + +```text +Source item IDs appear in multiple groups: ['fact_0012'] +``` + +`fact_0012` appeared once in a merged group with `fact_0002` and again as a +singleton group. The generic validator report recorded this as one +`duplicate_id` violation with both structural occurrences. + +Recovery result: + +The preserved raw consolidator response was repaired with the generic +Constraint Repair V1 interface and deterministic structural operations: + +- remove the repeated `fact_0012` occurrence from the later singleton group +- remove the now-empty singleton group + +No Semantic Consolidator LLM call was rerun. No prompts, Meeting Context, +extraction outputs or canonicalizer output were modified. Second validation +succeeded: every source fact ID appears exactly once, no IDs are missing, no +empty groups remain, and non-fact categories remained unchanged. + +This verifies the recovery path for this structural coverage failure. It does +not fix or change the Semantic Consolidator generation behaviour itself. + +Related files: + +- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/raw_model_response.txt` +- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_before.json` +- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_model_groups.json` +- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_consolidated_extractions.json` +- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_after.json` +- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repair_metadata.json` +- `samples/benchmarks/progeo_meeting_20260804_083849/working_protocol/working_protocol.md` +- `docs/constraint-repair.md` + +Regression test available (yes/no): no + +Current status: + +Verified. The recovery path repaired and revalidated this benchmark artifact, +then allowed the renderer to run once on the repaired consolidated output. + +Notes: + +The renderer output was produced, but it starts with explanatory prose instead +of the required `# Working Protocol` heading. That is a renderer faithfulness +issue, not part of the Semantic Consolidator coverage repair verified here. + +## BUG-009 + +ID: BUG-009 + +Title: Invalid responsible-party values survive canonicalization + +Pipeline stage: Extraction / Deterministic Canonicalizer + +Severity: High + +Status: Verified + +Date discovered: 2026-08-04 + +Version first observed: `progeo_meeting_context_v1_20260804_110913` + +Description: + +The context-aware Progeo benchmark produced action items whose `responsible` +field contained dates or date fragments instead of responsible entities. The +canonicalizer accepted these values as plain strings. + +Expected behaviour: + +When Meeting Context is available, action-item responsible values should be +deterministically resolved to known participants, mentioned people or supported +organization entities. Dates, locations, projects, products, technical terms, +generic process words and unknown free text must not survive as responsible +parties. + +Actual behaviour: + +The previous canonicalized Progeo artifact contained invalid responsible +values such as: + +- `31. August` +- `15. oder 16. September` +- `am 31.` +- `27.8.` + +Root cause: + +The extraction schema allowed `responsible` to be a free-form string or null. +Canonicalizer V1 parsed and preserved that string without checking it against +Meeting Context or obvious non-person/non-organization patterns. + +Fix: + +Canonicalizer V1 now validates responsible fields when Meeting Context is +available either through `--meeting-context` or extraction provenance. It: + +- accepts exact participant and mentioned-person display names +- accepts participant and mentioned-person aliases +- normalizes accepted aliases to canonical display names +- accepts explicit organization/department names only when represented in the + current Meeting Context structure +- rejects dates, relative dates, weekdays, clock times, locations, projects, + products, systems, technical terms, generic process words and unknown free + text +- clears rejected responsible values to null +- records structured `responsibility_validation` data and an aggregate + `responsibility_validations` report + +Verification: + +The existing context-aware Progeo extraction artifacts were re-canonicalized +without rerunning chunking, normalization or extraction: + +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json` + +Action item count remained 28 and semantic content other than responsible +fields remained unchanged. Invalid date-like responsible values were cleared. +`Martin Tazl` normalized to `Martin`, and `Marleen` remained valid. `Marleen +Wever` was rejected because the current Progeo Meeting Context does not list it +as a display name or alias. + +Related files: + +- `src/meeting_lab/consolidation/canonicalize.py` +- `tests/test_canonicalize.py` +- `samples/real_live/progeo_meeting/meeting_context.yaml` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer/canonicalized_extractions.json` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json` + +Regression test available (yes/no): yes + +Current status: + +Verified. Focused tests pass and the Progeo downstream-only regression demo +removes the observed invalid responsible values without changing action-item +count or non-responsible semantic content. + +Notes: + +This does not prove that remaining accepted responsibilities are semantically +supported by transcript evidence. It only prevents invalid entity values from +surviving in the `responsible` field. + +## BUG-010 + +ID: BUG-010 + +Title: Semantic Consolidator JSON truncates on larger real-life meetings + +Pipeline stage: Semantic Consolidator V0 + +Severity: High + +Status: Fixed + +Date discovered: 2026-08-04 + +Version first observed: `progeo_meeting_context_v1_20260804_110913` + +Description: + +The context-aware Progeo benchmark produced substantially more canonical fact +items than the earlier run. Semantic Consolidator V0 used the fixed committed +`num_predict` value of 4096 for its single grouping response and the preserved +raw response stopped in the middle of a JSON group. + +Expected behaviour: + +The consolidator should allocate enough output budget for the expected +fact-group JSON, preserve the raw model response, parse only valid JSON and +validate exact source fact coverage before writing consolidated output. + +Actual behaviour: + +The raw response ended after `fact_0060`/start of the next group, while Ollama +reported `eval_count: 4096`, exactly matching the old response cap. Parsing +failed before source coverage validation: + +```text +Invalid model JSON: Expecting property name enclosed in double quotes +``` + +Root cause: + +Output was truncated at the fixed `num_predict` cap. The prompt asks the model +to emit one JSON group per fact unless duplicates are found, so output size +scales with fact count and fact text size. The fixed cap was adequate for the +previous smaller Progeo run but not for the context-aware run. + +Fix: + +Semantic Consolidator V0 now estimates the response budget from the actual fact +payload and context-window headroom when `--num-predict` is not explicitly set. +Explicit `--num-predict` values are still respected. + +After valid model JSON is parsed, the consolidator also applies deterministic +source-coverage repair before strict validation: + +- duplicate source IDs after the first occurrence are removed +- empty groups created by removal are dropped +- missing source fact IDs are restored as singleton groups from the + canonicalized input + +This repair does not rewrite existing model group text, invent facts or create +new semantic merges. Strict validation still runs after repair and can still +reject the output. + +Verification: + +The Progeo context canonicalized benchmark was rerun through the fixed +Semantic Consolidator into: + +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/` + +The fixed run used `num_predict: 14347`, the model returned valid JSON with +`eval_count: 4568`, deterministic repair restored seven omitted source fact +IDs as singletons and final validation passed with 77 unique source fact IDs, +no duplicates and no missing IDs. + +Related files: + +- `src/meeting_lab/consolidation/consolidate_facts.py` +- `tests/test_consolidate_facts.py` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator/raw_model_response.txt` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/consolidated_extractions.json` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/repair_metadata.json` +- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/report.md` + +Regression test available (yes/no): yes + +Current status: + +Fixed for the observed truncation failure and protected by focused +consolidator tests. This does not improve semantic merge quality; it only +prevents fixed-budget truncation and enforces complete source coverage +deterministically. + +## BUG-011 + +ID: BUG-011 + +Title: Working Protocol Renderer V2 writes contract-invalid output + +Pipeline stage: Working Protocol Renderer V2 + +Severity: High + +Status: Verified + +Date discovered: 2026-08-04 + +Version first observed: `progeo_meeting_rc1_20260804_121502` + +Description: + +The RC1 Progeo renderer completed its LLM call and wrote `working_protocol.md`, +but the generated document did not start with `# Working Protocol`. It started +with explanatory prose and used category-summary/report framing instead of the +committed topic-oriented Working Protocol V2 structure. + +Expected behaviour: + +The renderer must preserve raw model output separately, deterministically clean +only harmless wrapper text, validate the cleaned candidate against the committed +Working Protocol V2 contract and write `working_protocol.md` only when the +candidate is valid. + +Actual behaviour: + +The one-off renderer path used for the benchmark wrote the raw LLM text to +`working_protocol.md` even though metadata recorded `valid=false` and +`readable_markdown=false`. + +Root cause: + +The committed contract existed in `prompts/working_protocol.md`, but there was +no reusable Working Protocol V2 renderer implementation with a strict output +validator. The model ignored explicit prompt instructions, and the pipeline did +not enforce the contract deterministically before preserving the final protocol +file. + +Fix: + +`src/meeting_lab/protocol/render_working_protocol.py` now implements the +Working Protocol V2 renderer path. It: + +- preserves the raw Ollama JSON response and raw response text +- removes leading prose only when a real `# Working Protocol` heading exists +- normalizes harmless heading whitespace +- rejects output without the required heading +- rejects category-summary framing that lacks topic-oriented body structure +- rejects malformed Markdown such as unclosed fenced code blocks +- writes `working_protocol.md` only after validation succeeds + +Verification: + +Focused renderer tests cover valid output, deterministic preamble cleanup, +leading whitespace cleanup, missing heading rejection, categorized-summary +rejection, malformed Markdown rejection, raw response preservation and ensuring +`working_protocol.md` contains only validated content. + +RC1 downstream demo: + +- Existing RC1 raw renderer output contained no embedded valid `# Working + Protocol` body. +- The new renderer path was run once against the existing RC1 consolidated + input with current committed settings. +- The raw response and cleaned candidate were preserved. +- Validation correctly rejected the output and did not write a final + `working_protocol.md`. + +Related files: + +- `src/meeting_lab/protocol/render_working_protocol.py` +- `tests/test_render_working_protocol.py` +- `docs/output-views.md` +- `samples/benchmarks/progeo_meeting_rc1_20260804_121502/working_protocol_bug011/` + +Regression test available (yes/no): yes + +Current status: + +Verified. The renderer contract is now enforced deterministically. This does +not improve semantic quality of the generated prose; it prevents invalid +renderer output from being accepted as a final Working Protocol. diff --git a/samples/real_live/progeo_meeting/meeting_context.yaml b/samples/real_live/progeo_meeting/meeting_context.yaml index e028817..3f2a994 100644 --- a/samples/real_live/progeo_meeting/meeting_context.yaml +++ b/samples/real_live/progeo_meeting/meeting_context.yaml @@ -2,29 +2,124 @@ schema_version: "1" meeting: meeting_id: "progeo-meeting" - title: "Progeo meeting" + title: "Progeo Meeting" language: "de" - date: null - objective: "" - notes: "Meeting Context scaffold for Real-Life Reference Meeting #2. Participant and entity metadata require manual completion." + date: 2026-07-27 + objective: >- + Ergebnisbericht zur abgeschlossenen Versuchsproduktion eines Geogitters + aus recyceltem Polypropylen und Abstimmung weiterer Großversuche. + notes: >- + Real-Life Reference Meeting #2. Beteiligte externe Organisationen: + Fraunhofer, MAS, Universität Leeds, FH Münster, Kiwa und Kockmann. -participants: [] +participants: + - participant_id: "martin" + display_name: "Martin" + aliases: [Martin Tazl, Herr Tazl] + role: "Entwicklungsleiter Naue" + department: "naue" + attendance_status: "present" + notes: null -mentioned_people: [] + - participant_id: "james" + display_name: "James" + aliases: [James Nachtigall, Herr Nachtigall] + role: "Entwicklungsingenieur" + department: "naue" + attendance_status: "present" + notes: null + + - participant_id: "david" + display_name: "David" + aliases: [] + role: "Entwicklungsingenieur" + department: "naue" + attendance_status: "present" + notes: null + + - participant_id: "gotthard" + display_name: "Gotthard" + aliases: [Gothard, Gotthardt, Gothardt, Herr Walter] + role: "Projektleiter" + department: "fh-muenster" + attendance_status: "present" + notes: null + + - participant_id: "tim" + display_name: "Tim" + aliases: [Timm, Tim Schulte-Uebbing, Herr Schulte-Uebbing] + role: "PhD" + department: "fh-muenster" + attendance_status: "present" + notes: null + + - participant_id: "antonius" + display_name: "Antonius" + aliases: [] + role: "PhD" + department: "fh-muenster" + attendance_status: "present" + notes: null + + - participant_id: "marleen" + display_name: "Marleen" + aliases: [Marlene, Frau Wever, Marleen Wever] + role: "Entwicklungsingenieurin" + department: "huesker" + attendance_status: "present" + notes: null + +mentioned_people: + - person_id: "henning" + display_name: "Henning" + aliases: [] + role: null + department: null + attendance_status: "not_present" + notes: null organization: name: null - departments: [] + departments: + - id: "naue" + name: "Naue" + aliases: [Naue GmbH & Co. KG] - abbreviations: {} + - id: "huesker" + name: "Huesker" + aliases: [HÜSKER] + + - id: "fh-muenster" + name: "FH Münster" + aliases: [Fachhochschule Münster] + + abbreviations: + RC: "Recycling" + PP: "Polypropylen" + PET: "Polyethylenterephthalat" + MFI: "Melt Flow Index" + IV: "intrinsische Viskosität" + SSP: "Solid-State Polymerization" known_entities: - projects: [] - products: [] + projects: [Progeo] + products: [Secugrid, Combigrid] systems: [] - locations: [] - technical_terms: [] + locations: [Stettin] + technical_terms: + - Rezyklat + - Virgin-Material + - Glührückstand + - Verstrecken + - Streckwerk + - Bruchspannung + - Zugfestigkeit + - Dehnung + - Trommelsieb + - chemisches Recycling + - mechanisches Recycling + - Nachkondensation context_rules: participant_list_is_authoritative: true @@ -32,4 +127,4 @@ context_rules: do_not_infer_departments: true do_not_infer_responsibilities: true do_not_infer_attendance: true - mentioned_people_are_not_participants: true + mentioned_people_are_not_participants: true \ No newline at end of file diff --git a/src/meeting_lab/consolidation/canonicalize.py b/src/meeting_lab/consolidation/canonicalize.py index 9073d33..cab7036 100644 --- a/src/meeting_lab/consolidation/canonicalize.py +++ b/src/meeting_lab/consolidation/canonicalize.py @@ -8,9 +8,12 @@ import json import re import sys from collections import Counter +from dataclasses import dataclass from pathlib import Path from typing import Any +from src.meeting_lab.models.meeting_context import load_meeting_context + SCHEMA_VERSION = "1" @@ -36,6 +39,51 @@ CANONICAL_CATEGORIES = tuple(CATEGORY_NAMES.values()) CHUNK_EXTRACTION_RE = re.compile(r"^chunk_(\d+)_extraction\.json$") WHITESPACE_RE = re.compile(r"\s+") +RESPONSIBLE_KEY_RE = re.compile(r"\s+") +NUMERIC_DATE_RE = re.compile(r"(?i)\b\d{1,2}\s*[./]\s*(?:\d{1,2}|[a-zäöü]+)?\b") +MONTH_DATE_RE = re.compile( + r"(?i)\b\d{1,2}\.?\s*(?:oder\s+\d{1,2}\.?\s*)?" + r"(januar|februar|märz|maerz|april|mai|juni|juli|august|september|oktober|november|dezember)\b" +) +TIME_RE = re.compile(r"(?i)\b\d{1,2}[:.]\d{2}\s*(?:uhr)?\b|\b\d{1,2}\s*uhr\b") +DATE_FRAGMENT_RE = re.compile(r"(?i)\b(?:am|zum|bis|vor|nach)\s+\d{1,2}\.?\b") + +DATE_WORDS = { + "heute", + "morgen", + "übermorgen", + "uebermorgen", + "gestern", + "vorgestern", +} +WEEKDAY_WORDS = { + "montag", + "dienstag", + "mittwoch", + "donnerstag", + "freitag", + "samstag", + "sonntag", +} +RELATIVE_DATE_PHRASES = { + "nächste woche", + "naechste woche", + "diese woche", + "kommende woche", + "nächsten monat", + "naechsten monat", +} +GENERIC_RESPONSIBLE_WORDS = { + "team", + "projektteam", + "projektleitung", + "logistik", + "alle", + "autor", + "labor", + "partner", + "gruppe", +} def parse_args() -> argparse.Namespace: @@ -59,9 +107,110 @@ def parse_args() -> argparse.Namespace: action="store_true", help="Preserve exact duplicate items instead of merging them.", ) + parser.add_argument( + "--meeting-context", + type=Path, + help=( + "Optional Meeting Context V1 YAML used to validate and normalize " + "action-item responsible fields. If omitted, canonicalizer uses " + "context provenance from extraction JSON when available." + ), + ) return parser.parse_args() +@dataclass(frozen=True) +class ResponsiblePartyNormalizer: + allowed_names: dict[str, str] + invalid_values: dict[str, str] + + @classmethod + def from_meeting_context_data(cls, data: dict[str, Any]) -> "ResponsiblePartyNormalizer": + allowed: dict[str, str] = {} + invalid: dict[str, str] = {} + + for collection, id_key in ( + ("participants", "participant_id"), + ("mentioned_people", "person_id"), + ): + for person in data.get(collection, []) or []: + if not isinstance(person, dict): + continue + display_name = clean_text(person.get("display_name")) + if not display_name: + continue + for value in [display_name, *(person.get("aliases") or [])]: + text = clean_text(value) + if text: + allowed[responsible_key(text)] = display_name + + organization = data.get("organization") + if isinstance(organization, dict): + organization_name = clean_text(organization.get("name")) + if organization_name: + allowed[responsible_key(organization_name)] = organization_name + + for department in organization.get("departments") or []: + if not isinstance(department, dict): + continue + department_name = clean_text(department.get("name")) + if department_name: + allowed[responsible_key(department_name)] = department_name + for alias in department.get("aliases") or []: + text = clean_text(alias) + if text and department_name: + allowed[responsible_key(text)] = department_name + + known_entities = data.get("known_entities") + if isinstance(known_entities, dict): + for collection in ( + "projects", + "products", + "systems", + "locations", + "technical_terms", + ): + for value in known_entities.get(collection) or []: + text = clean_text(value) + if text: + invalid[responsible_key(text)] = collection + + return cls(allowed_names=allowed, invalid_values=invalid) + + def normalize( + self, + value: str | None, + item_id: str, + ) -> tuple[str | None, dict[str, Any] | None]: + original = clean_text(value) + if original is None: + return None, None + + key = responsible_key(original) + canonical = self.allowed_names.get(key) + if canonical is not None: + return canonical, { + "action_item_id": item_id, + "field": "responsible", + "original_value": original, + "normalized_value": canonical, + "status": "accepted", + "reason": "resolved_to_meeting_context_entity", + "cleared_to_null": False, + } + + reason = invalid_responsible_reason(original, self.invalid_values) + return None, { + "action_item_id": item_id, + "field": "responsible", + "original_value": original, + "normalized_value": None, + "status": "rejected", + "reason": reason, + "cleared_to_null": True, + } + + def chunk_sort_key(path: Path) -> tuple[int, str]: match = CHUNK_EXTRACTION_RE.match(path.name) if not match: @@ -109,6 +258,30 @@ def clean_text(value: Any) -> str | None: return text if text else None +def responsible_key(value: str) -> str: + return RESPONSIBLE_KEY_RE.sub(" ", value.casefold()).strip(" .,:;") + + +def invalid_responsible_reason( + value: str, + invalid_values: dict[str, str], +) -> str: + key = responsible_key(value) + if key in invalid_values: + return f"known_{invalid_values[key]}_not_responsible_entity" + if key in GENERIC_RESPONSIBLE_WORDS: + return "generic_process_word" + if key in DATE_WORDS or key in WEEKDAY_WORDS or key in RELATIVE_DATE_PHRASES: + return "date_or_relative_date" + if NUMERIC_DATE_RE.search(value) or MONTH_DATE_RE.search(value): + return "date_or_date_range" + if TIME_RE.search(value): + return "clock_time" + if DATE_FRAGMENT_RE.search(value): + return "date_fragment" + return "unknown_responsible_entity" + + def split_legacy_string(value: str) -> list[str]: return [part.strip() for part in value.split("|")] @@ -324,6 +497,7 @@ def canonicalize_value( source_file: str, source_index: int, counts: Counter[str], + responsible_normalizer: ResponsiblePartyNormalizer | None = None, ) -> dict[str, Any]: parsed = PARSERS[category](value) item: dict[str, Any] = { @@ -337,6 +511,14 @@ def canonicalize_value( } for key, parsed_value in parsed.items(): item[key] = parsed_value + if category == "action_item" and responsible_normalizer is not None: + normalized, validation = responsible_normalizer.normalize( + item.get("responsible"), + item["item_id"], + ) + item["responsible"] = normalized + if validation is not None: + item["responsibility_validation"] = validation item["source_references"] = [ source_reference(source_file, source_index, value, item["evidence"]) ] @@ -367,6 +549,7 @@ def merge_exact_duplicates(items: list[dict[str, Any]]) -> tuple[list[dict[str, def canonicalize_extractions( input_dir: Path, merge_duplicates: bool = True, + meeting_context_path: Path | None = None, ) -> dict[str, Any]: files = find_extraction_files(input_dir) if not files: @@ -375,9 +558,22 @@ def canonicalize_extractions( counts: Counter[str] = Counter() input_counts: Counter[str] = Counter() items: list[dict[str, Any]] = [] + responsible_normalizer: ResponsiblePartyNormalizer | None = None + responsibility_validations: list[dict[str, Any]] = [] + loaded_files: list[tuple[Path, dict[str, Any]]] = [] for path in files: data = load_json_object(path) + loaded_files.append((path, data)) + + context_path = meeting_context_path or infer_meeting_context_path(loaded_files) + if context_path is not None: + meeting_context = load_meeting_context(context_path) + responsible_normalizer = ResponsiblePartyNormalizer.from_meeting_context_data( + meeting_context.data + ) + + for path, data in loaded_files: validate_required_categories(data, path) for raw_category in REQUIRED_CATEGORIES: category = CATEGORY_NAMES[raw_category] @@ -391,8 +587,16 @@ def canonicalize_extractions( source_file=path.name, source_index=source_index, counts=counts, + responsible_normalizer=responsible_normalizer, ) ) + if ( + items[-1].get("category") == "action_item" + and "responsibility_validation" in items[-1] + ): + responsibility_validations.append( + items[-1]["responsibility_validation"] + ) exact_duplicates = 0 if merge_duplicates: @@ -415,11 +619,40 @@ def canonicalize_extractions( "input_item_count_by_category": input_counts_by_category, "output_item_count_by_category": output_counts_by_category, "exact_duplicates_merged": exact_duplicates, + "responsibility_validation_count": len(responsibility_validations), + "responsibility_rejection_count": sum( + 1 + for validation in responsibility_validations + if validation.get("status") == "rejected" + ), + "responsibility_normalization_count": sum( + 1 + for validation in responsibility_validations + if validation.get("status") == "accepted" + and validation.get("original_value") != validation.get("normalized_value") + ), }, + "responsibility_validations": responsibility_validations, "items": items, } +def infer_meeting_context_path( + loaded_files: list[tuple[Path, dict[str, Any]]], +) -> Path | None: + paths: set[str] = set() + for _path, data in loaded_files: + context = data.get("context") + if not isinstance(context, dict): + continue + source_file = clean_text(context.get("source_file")) + if source_file: + paths.add(source_file) + if len(paths) != 1: + return None + return Path(next(iter(paths))) + + def write_canonicalized(output: dict[str, Any], output_path: Path) -> Path: output_path.parent.mkdir(parents=True, exist_ok=True) output_path.write_text( @@ -435,6 +668,7 @@ def main() -> int: output = canonicalize_extractions( args.input_dir, merge_duplicates=not args.no_merge_exact_duplicates, + meeting_context_path=args.meeting_context, ) output_path = write_canonicalized(output, args.output) except (OSError, UnicodeError, ValueError) as exc: diff --git a/src/meeting_lab/consolidation/consolidate_facts.py b/src/meeting_lab/consolidation/consolidate_facts.py index 50948a3..4a55dc2 100644 --- a/src/meeting_lab/consolidation/consolidate_facts.py +++ b/src/meeting_lab/consolidation/consolidate_facts.py @@ -4,6 +4,7 @@ from __future__ import annotations import argparse +import copy import json import sys import time @@ -25,8 +26,13 @@ except ModuleNotFoundError: # pragma: no cover - used by repository-root tests. DEFAULT_MODEL = "qwen3.5:9b" DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate" DEFAULT_NUM_CTX = 32768 -DEFAULT_NUM_PREDICT = 4096 +DEFAULT_MIN_NUM_PREDICT = 4096 +DEFAULT_NUM_PREDICT = DEFAULT_MIN_NUM_PREDICT DEFAULT_PROGRESS_INTERVAL = 30 +OUTPUT_CONTEXT_RESERVE_TOKENS = 1024 +OUTPUT_TOKEN_ESTIMATE_CHARS = 4 +OUTPUT_GROUP_OVERHEAD_CHARS = 320 +OUTPUT_SAFETY_MARGIN = 1.35 PROMPT_NAME = "consolidate_facts.md" @@ -75,10 +81,11 @@ def parse_args() -> argparse.Namespace: parser.add_argument( "--num-predict", type=int, - default=DEFAULT_NUM_PREDICT, + default=None, help=( - "Maximum generated tokens. The default is bounded for the expected " - f"fact-group JSON while leaving truncation headroom (default: {DEFAULT_NUM_PREDICT})." + "Maximum generated tokens. By default this is estimated from the " + "fact payload size and bounded by the context window. Explicit " + "values preserve the previous fixed-budget behavior." ), ) thinking = parser.add_mutually_exclusive_group() @@ -158,6 +165,49 @@ def build_consolidation_prompt(facts: list[dict[str, Any]]) -> str: return f"{task_prompt}\n\nFACT ITEMS:\n{payload}\n" +def estimate_response_tokens(facts: list[dict[str, Any]]) -> int: + """ + Estimate the token budget needed for the model's grouping JSON. + + Semantic Consolidator V0 asks the model to return one group per source fact + unless it finds a conservative duplicate. The response therefore scales with + the number and text size of fact items. The estimate intentionally includes + per-group JSON overhead and a safety margin; strict validation still decides + whether the actual response is usable. + """ + text_chars = 0 + for item in facts: + text_chars += len(str(item.get("text", ""))) + text_chars += len(str(item.get("evidence", ""))) + + estimated_chars = int( + (text_chars + len(facts) * OUTPUT_GROUP_OVERHEAD_CHARS) + * OUTPUT_SAFETY_MARGIN + ) + return max( + DEFAULT_MIN_NUM_PREDICT, + (estimated_chars + OUTPUT_TOKEN_ESTIMATE_CHARS - 1) + // OUTPUT_TOKEN_ESTIMATE_CHARS, + ) + + +def resolve_num_predict( + requested_num_predict: int | None, + facts: list[dict[str, Any]], + prompt_token_estimate: int, + num_ctx: int, +) -> int: + if requested_num_predict is not None: + return requested_num_predict + + estimated = estimate_response_tokens(facts) + max_available = max( + DEFAULT_MIN_NUM_PREDICT, + num_ctx - prompt_token_estimate - OUTPUT_CONTEXT_RESERVE_TOKENS, + ) + return min(estimated, max_available) + + def response_text_from_ollama_data(data: dict[str, Any]) -> str | None: text = data.get("response") if isinstance(text, str) and text.strip(): @@ -370,6 +420,96 @@ def validate_group_shapes(groups: list[dict[str, Any]]) -> None: raise ConsolidationValidationError("Merged groups need at least two IDs.") +def repair_model_group_coverage( + model_output: dict[str, Any], + facts: list[dict[str, Any]], +) -> tuple[dict[str, Any], list[dict[str, Any]]]: + """ + Apply deterministic source-coverage repairs to model grouping JSON. + + The repair is intentionally conservative. Repeated source IDs are removed + after their first occurrence, empty groups created by that removal are + dropped, and missing facts are restored as singleton groups using the + original canonicalized fact text. No existing group text, merge reason or + semantic merge is rewritten. + """ + groups = model_output.get("groups") + if not isinstance(groups, list): + return model_output, [] + + repaired = copy.deepcopy(model_output) + repaired_groups = repaired["groups"] + facts_by_id = {str(item.get("item_id")): item for item in facts} + expected_ids = set(facts_by_id) + seen: set[str] = set() + changes: list[dict[str, Any]] = [] + + for group_index, group in enumerate(repaired_groups): + if not isinstance(group, dict): + continue + source_ids = group.get("source_item_ids") + if not isinstance(source_ids, list): + continue + kept_ids: list[str] = [] + for id_index, item_id in enumerate(source_ids): + if not isinstance(item_id, str) or item_id not in expected_ids: + kept_ids.append(item_id) + continue + if item_id in seen: + changes.append( + { + "operation": "remove_duplicate_source_id", + "id": item_id, + "group_index": group_index, + "id_index": id_index, + } + ) + continue + seen.add(item_id) + kept_ids.append(item_id) + group["source_item_ids"] = kept_ids + + non_empty_groups: list[dict[str, Any]] = [] + for group_index, group in enumerate(repaired_groups): + if ( + isinstance(group, dict) + and isinstance(group.get("source_item_ids"), list) + and len(group["source_item_ids"]) == 0 + ): + changes.append( + { + "operation": "remove_empty_group", + "group_index": group_index, + "canonical_text": group.get("canonical_text"), + } + ) + continue + non_empty_groups.append(group) + repaired["groups"] = non_empty_groups + + missing_ids = sorted(expected_ids - seen) + for item_id in missing_ids: + fact = facts_by_id[item_id] + repaired["groups"].append( + { + "canonical_text": str(fact.get("text", "")).strip(), + "source_item_ids": [item_id], + "merge_reason": ( + "Deterministic coverage repair: source fact was missing " + "from the model grouping and is preserved as a singleton." + ), + } + ) + changes.append( + { + "operation": "restore_missing_source_id_as_singleton", + "id": item_id, + } + ) + + return repaired, changes + + def build_consolidated_fact_item( group: dict[str, Any], fact_by_id: dict[str, dict[str, Any]], @@ -481,9 +621,11 @@ def write_report( runtime: float, prompt_chars: int, prompt_token_estimate: int, + num_predict: int, fact_count: int, groups: list[dict[str, Any]], output_path: Path, + repair_changes: list[dict[str, Any]] | None = None, ) -> None: merged = [group for group in groups if len(group["source_item_ids"]) > 1] singletons = [group for group in groups if len(group["source_item_ids"]) == 1] @@ -497,9 +639,11 @@ def write_report( f"- Fact item count: {fact_count}", f"- Prompt characters: {prompt_chars}", f"- Estimated prompt tokens: {prompt_token_estimate}", + f"- num_predict: {num_predict}", f"- Merged fact groups: {len(merged)}", f"- Source facts involved in merges: {sum(len(group['source_item_ids']) for group in merged)}", f"- Singleton fact groups: {len(singletons)}", + f"- Deterministic repair changes: {len(repair_changes or [])}", f"- Output path: `{output_path}`", "", "## Actual Merges", @@ -518,6 +662,10 @@ def write_report( "", ] ) + if repair_changes: + lines.extend(["", "## Deterministic Coverage Repairs", ""]) + for change in repair_changes: + lines.append(f"- `{change['operation']}`: {json.dumps(change, ensure_ascii=False, sort_keys=True)}") path.write_text("\n".join(lines).rstrip() + "\n", encoding="utf-8") @@ -527,6 +675,7 @@ def main() -> int: raw_response_path = args.output_dir / "raw_model_response.txt" output_path = args.output_dir / "consolidated_extractions.json" report_path = args.output_dir / "report.md" + repair_metadata_path = args.output_dir / "repair_metadata.json" try: canonicalized = load_json_object(args.canonicalized_input) @@ -534,9 +683,16 @@ def main() -> int: prompt = build_consolidation_prompt(facts) prompt_chars = len(prompt) prompt_token_estimate = (prompt_chars + 3) // 4 + num_predict = resolve_num_predict( + requested_num_predict=args.num_predict, + facts=facts, + prompt_token_estimate=prompt_token_estimate, + num_ctx=args.num_ctx, + ) print(f"Fact item count: {len(facts)}") print(f"Estimated prompt size chars: {prompt_chars}") print(f"Estimated prompt tokens: {prompt_token_estimate}") + print(f"Resolved num_predict: {num_predict}") print("Expected LLM call count: 1") print("Expected runtime: 5-10 minutes on current local benchmark basis") @@ -546,12 +702,23 @@ def main() -> int: prompt=prompt, timeout=args.timeout, num_ctx=args.num_ctx, - num_predict=args.num_predict, + num_predict=num_predict, think=args.think, progress_interval=args.progress_interval, ) raw_response_path.write_text(raw_text + "\n", encoding="utf-8") model_output = parse_model_json(raw_text) + model_output, repair_changes = repair_model_group_coverage(model_output, facts) + if repair_changes: + write_json( + repair_metadata_path, + { + "scope": "semantic_consolidator_v0_source_coverage", + "llm_used": False, + "repair_count": len(repair_changes), + "repairs": repair_changes, + }, + ) expected_fact_ids = {item["item_id"] for item in facts} groups = validate_model_groups(model_output, expected_fact_ids) validate_group_shapes(groups) @@ -564,9 +731,11 @@ def main() -> int: runtime=runtime, prompt_chars=prompt_chars, prompt_token_estimate=prompt_token_estimate, + num_predict=num_predict, fact_count=len(facts), groups=groups, output_path=output_path, + repair_changes=repair_changes, ) except requests.ConnectionError as exc: print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr) diff --git a/src/meeting_lab/protocol/render_working_protocol.py b/src/meeting_lab/protocol/render_working_protocol.py new file mode 100644 index 0000000..c18c0c1 --- /dev/null +++ b/src/meeting_lab/protocol/render_working_protocol.py @@ -0,0 +1,519 @@ +#!/usr/bin/env python3 +"""LLM-backed Working Protocol V2 renderer with deterministic contract checks.""" + +from __future__ import annotations + +import argparse +import json +import re +import sys +import time +from dataclasses import dataclass +from datetime import datetime +from pathlib import Path +from typing import Any + +import requests + +try: + from meeting_lab.llm.prompts import load_prompt +except ModuleNotFoundError: # pragma: no cover - used by repository-root tests. + from src.meeting_lab.llm.prompts import load_prompt + + +DEFAULT_MODEL = "qwen3.5:9b" +DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate" +DEFAULT_NUM_CTX = 32768 +DEFAULT_NUM_PREDICT = 4096 +DEFAULT_TIMEOUT = 1800 +PROMPT_NAME = "working_protocol.md" +REQUIRED_TITLE = "# Working Protocol" +ALLOWED_TOPIC_SECTIONS = { + "Background", + "Decisions", + "Action Items", + "Open Questions", +} +GENERIC_SUMMARY_HEADINGS = { + "Entscheidungen", + "Decisions", + "Handlungsaufträge", + "Handlungsauftraege", + "Action Items", + "Offene Fragen", + "Open Questions", + "Technische Details", + "Technical Details", + "Fakten", + "Facts", + "Zusammenfassung", + "Summary", + "Konsolidierter Projektstatusbericht", +} +HEADING_RE = re.compile(r"^(#{1,6})\s+(.+?)\s*$") + + +class WorkingProtocolValidationError(ValueError): + """Raised when rendered Markdown violates the Working Protocol V2 contract.""" + + +@dataclass(frozen=True) +class WorkingProtocolValidationReport: + valid: bool + violations: list[dict[str, Any]] + + def to_dict(self) -> dict[str, Any]: + return {"valid": self.valid, "violations": self.violations} + + +def parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser( + description="Render a consolidated meeting representation as Working Protocol V2." + ) + parser.add_argument("input", type=Path, help="Consolidated meeting JSON input.") + parser.add_argument( + "-o", + "--output-dir", + type=Path, + required=True, + help="Directory for working_protocol.md, raw response, validation and metadata.", + ) + parser.add_argument( + "--model", + default=DEFAULT_MODEL, + help=f"Ollama model name (default: {DEFAULT_MODEL}).", + ) + parser.add_argument( + "--endpoint", + default=DEFAULT_ENDPOINT, + help=f"Ollama generate endpoint (default: {DEFAULT_ENDPOINT}).", + ) + parser.add_argument( + "--timeout", + type=int, + default=DEFAULT_TIMEOUT, + help=f"HTTP timeout in seconds (default: {DEFAULT_TIMEOUT}).", + ) + parser.add_argument( + "--num-ctx", + type=int, + default=DEFAULT_NUM_CTX, + help=f"Context window tokens (default: {DEFAULT_NUM_CTX}).", + ) + parser.add_argument( + "--num-predict", + type=int, + default=DEFAULT_NUM_PREDICT, + help=f"Maximum generated tokens (default: {DEFAULT_NUM_PREDICT}).", + ) + thinking = parser.add_mutually_exclusive_group() + thinking.add_argument("--think", dest="think", action="store_true") + thinking.add_argument("--no-think", dest="think", action="store_false") + parser.set_defaults(think=False) + return parser.parse_args() + + +def build_renderer_prompt(input_text: str) -> str: + return f"{load_prompt(PROMPT_NAME)}\n\nINPUT JSON:\n{input_text}\n" + + +def response_text_from_ollama_data(data: dict[str, Any]) -> str: + text = data.get("response") + if isinstance(text, str): + return text + message = data.get("message") + if isinstance(message, dict) and isinstance(message.get("content"), str): + return message["content"] + return "" + + +def build_ollama_payload( + model: str, + prompt: str, + num_ctx: int, + num_predict: int, + think: bool, +) -> dict[str, Any]: + return { + "model": model, + "prompt": prompt, + "think": think, + "stream": False, + "options": { + "temperature": 0.0, + "num_ctx": num_ctx, + "num_predict": num_predict, + }, + } + + +def call_ollama( + endpoint: str, + payload: dict[str, Any], + timeout: int, +) -> tuple[dict[str, Any], float]: + start = time.perf_counter() + response = requests.post(endpoint, json=payload, timeout=timeout) + runtime = time.perf_counter() - start + response.raise_for_status() + data = response.json() + if not isinstance(data, dict): + raise ValueError("Ollama response must be a JSON object.") + return data, runtime + + +def normalize_heading_whitespace(line: str) -> str: + match = HEADING_RE.match(line.strip()) + if not match: + return line.rstrip() + return f"{match.group(1)} {match.group(2).strip()}" + + +def clean_working_protocol_markdown(text: str) -> str: + """ + Remove only non-semantic wrapper text before an existing protocol heading. + + The cleanup does not create headings or transform a categorized summary into + a Working Protocol. It only makes an already present contract heading the + first byte of the candidate document. + """ + text = text.replace("\r\n", "\n").replace("\r", "\n") + lines = text.split("\n") + start_index: int | None = None + for index, line in enumerate(lines): + if normalize_heading_whitespace(line) == REQUIRED_TITLE: + start_index = index + break + if start_index is None: + return text.lstrip() + + cleaned_lines = lines[start_index:] + if cleaned_lines: + cleaned_lines[0] = REQUIRED_TITLE + return "\n".join(normalize_heading_whitespace(line) for line in cleaned_lines).strip() + "\n" + + +def validate_markdown_shape(markdown: str) -> list[dict[str, Any]]: + violations: list[dict[str, Any]] = [] + if markdown.count("```") % 2 != 0: + violations.append( + { + "type": "malformed_markdown", + "reason": "unclosed_fenced_code_block", + } + ) + + seen_title = False + seen_topic = False + current_topic_has_section = False + topic_count = 0 + section_count = 0 + + for line_number, line in enumerate(markdown.splitlines(), start=1): + match = HEADING_RE.match(line) + if not match: + continue + level = len(match.group(1)) + title = match.group(2).strip().strip("*") + + if level == 1: + if line_number != 1 or title != "Working Protocol": + violations.append( + { + "type": "invalid_heading", + "line": line_number, + "heading": line, + "reason": "only the first line may be '# Working Protocol'", + } + ) + seen_title = True + continue + + if not seen_title: + violations.append( + { + "type": "invalid_heading_order", + "line": line_number, + "heading": line, + "reason": "heading appears before required title", + } + ) + continue + + if level == 2: + topic_count += 1 + seen_topic = True + current_topic_has_section = False + if title in GENERIC_SUMMARY_HEADINGS: + violations.append( + { + "type": "generic_summary_framing", + "line": line_number, + "heading": line, + "reason": "top-level category heading is not a topic", + } + ) + continue + + if level == 3: + if not seen_topic: + violations.append( + { + "type": "missing_topic", + "line": line_number, + "heading": line, + "reason": "section appears before any topic heading", + } + ) + if title not in ALLOWED_TOPIC_SECTIONS: + violations.append( + { + "type": "unknown_topic_section", + "line": line_number, + "heading": line, + "allowed": sorted(ALLOWED_TOPIC_SECTIONS), + } + ) + else: + section_count += 1 + current_topic_has_section = True + continue + + violations.append( + { + "type": "unsupported_heading_level", + "line": line_number, + "heading": line, + "reason": "Working Protocol V2 uses h1 title, h2 topics and h3 sections only", + } + ) + + if seen_topic and not current_topic_has_section: + # This catches the last topic; earlier empty topics are caught below by + # counting consecutive h2 headings. + pass + + h2_without_section = _topic_headings_without_sections(markdown) + violations.extend(h2_without_section) + + if topic_count == 0: + violations.append( + { + "type": "missing_topic", + "reason": "document must contain at least one '## Topic title' section", + } + ) + if section_count == 0: + violations.append( + { + "type": "missing_topic_sections", + "reason": "document must contain at least one allowed h3 topic section", + } + ) + return violations + + +def _topic_headings_without_sections(markdown: str) -> list[dict[str, Any]]: + violations: list[dict[str, Any]] = [] + current_topic: tuple[int, str] | None = None + current_has_section = False + + for line_number, line in enumerate(markdown.splitlines(), start=1): + match = HEADING_RE.match(line) + if not match: + continue + level = len(match.group(1)) + if level == 2: + if current_topic is not None and not current_has_section: + violations.append( + { + "type": "empty_topic", + "line": current_topic[0], + "heading": current_topic[1], + "reason": "topic has no allowed h3 section", + } + ) + current_topic = (line_number, line) + current_has_section = False + elif level == 3 and current_topic is not None: + title = match.group(2).strip().strip("*") + if title in ALLOWED_TOPIC_SECTIONS: + current_has_section = True + + if current_topic is not None and not current_has_section: + violations.append( + { + "type": "empty_topic", + "line": current_topic[0], + "heading": current_topic[1], + "reason": "topic has no allowed h3 section", + } + ) + return violations + + +def validate_working_protocol_markdown(markdown: str) -> WorkingProtocolValidationReport: + violations: list[dict[str, Any]] = [] + if not markdown.startswith(REQUIRED_TITLE): + violations.append( + { + "type": "missing_required_heading", + "expected": REQUIRED_TITLE, + "reason": "document must begin exactly with '# Working Protocol'", + } + ) + elif not markdown.startswith(REQUIRED_TITLE + "\n"): + violations.append( + { + "type": "invalid_required_heading", + "expected": REQUIRED_TITLE, + "reason": "required heading must occupy the complete first line", + } + ) + + if markdown.startswith(REQUIRED_TITLE): + violations.extend(validate_markdown_shape(markdown)) + + return WorkingProtocolValidationReport( + valid=len(violations) == 0, + violations=violations, + ) + + +def write_json(path: Path, data: Any) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(data, ensure_ascii=False, indent=2) + "\n", encoding="utf-8") + + +def render_working_protocol( + input_path: Path, + output_dir: Path, + model: str = DEFAULT_MODEL, + endpoint: str = DEFAULT_ENDPOINT, + timeout: int = DEFAULT_TIMEOUT, + num_ctx: int = DEFAULT_NUM_CTX, + num_predict: int = DEFAULT_NUM_PREDICT, + think: bool = False, +) -> dict[str, Any]: + output_dir.mkdir(parents=True, exist_ok=True) + prompt_path = Path("prompts") / PROMPT_NAME + raw_response_path = output_dir / "raw_model_response.json" + raw_text_path = output_dir / "raw_model_response.txt" + candidate_path = output_dir / "cleaned_candidate.md" + validation_path = output_dir / "validation_report.json" + protocol_path = output_dir / "working_protocol.md" + metadata_path = output_dir / "metadata.json" + report_path = output_dir / "report.md" + + input_text = input_path.read_text(encoding="utf-8-sig") + prompt = build_renderer_prompt(input_text) + payload = build_ollama_payload(model, prompt, num_ctx, num_predict, think) + + data, runtime = call_ollama(endpoint, payload, timeout) + response_text = response_text_from_ollama_data(data) + raw_response_path.write_text( + json.dumps(data, ensure_ascii=False, indent=2) + "\n", + encoding="utf-8", + ) + raw_text_path.write_text(response_text, encoding="utf-8") + + cleaned = clean_working_protocol_markdown(response_text) + candidate_path.write_text(cleaned, encoding="utf-8") + validation = validate_working_protocol_markdown(cleaned) + write_json(validation_path, validation.to_dict()) + + if validation.valid: + protocol_path.write_text(cleaned, encoding="utf-8") + elif protocol_path.exists(): + protocol_path.unlink() + + metadata = { + "model": model, + "endpoint": endpoint, + "think": think, + "stream": False, + "temperature": 0.0, + "num_ctx": num_ctx, + "num_predict": num_predict, + "timeout_seconds": timeout, + "request_count": 1, + "runtime_seconds": runtime, + "prompt_path": str(prompt_path.resolve()), + "input_path": str(input_path.resolve()), + "output_path": str(protocol_path.resolve()) if validation.valid else None, + "raw_response_path": str(raw_response_path.resolve()), + "raw_text_path": str(raw_text_path.resolve()), + "cleaned_candidate_path": str(candidate_path.resolve()), + "validation_report_path": str(validation_path.resolve()), + "http_status_code": 200, + "response_text_length": len(response_text), + "candidate_text_length": len(cleaned), + "valid": validation.valid, + "readable_markdown": validation.valid, + "top_level_json_keys": sorted(data.keys()), + "done": data.get("done"), + "done_reason": data.get("done_reason"), + "total_duration": data.get("total_duration"), + "load_duration": data.get("load_duration"), + "prompt_eval_count": data.get("prompt_eval_count"), + "prompt_eval_duration": data.get("prompt_eval_duration"), + "eval_count": data.get("eval_count"), + "eval_duration": data.get("eval_duration"), + "created_at": datetime.now().isoformat(timespec="seconds"), + } + write_json(metadata_path, metadata) + + lines = [ + "# Working Protocol Renderer V2 Report", + "", + f"- Result: {'valid renderer run' if validation.valid else 'invalid renderer run'}", + f"- Model: `{model}`", + f"- Runtime: {runtime:.3f} seconds", + f"- Request count: 1", + f"- Valid: {validation.valid}", + f"- Violations: {len(validation.violations)}", + f"- Raw response path: `{raw_response_path}`", + f"- Cleaned candidate path: `{candidate_path}`", + f"- Validation report path: `{validation_path}`", + f"- Output path: `{protocol_path if validation.valid else 'not written'}`", + ] + report_path.write_text("\n".join(lines) + "\n", encoding="utf-8") + return metadata + + +def main() -> int: + args = parse_args() + try: + metadata = render_working_protocol( + input_path=args.input, + output_dir=args.output_dir, + model=args.model, + endpoint=args.endpoint, + timeout=args.timeout, + num_ctx=args.num_ctx, + num_predict=args.num_predict, + think=args.think, + ) + except requests.ConnectionError as exc: + print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr) + return 1 + except requests.Timeout as exc: + print(f"Error: Ollama request timed out after {args.timeout} seconds: {exc}", file=sys.stderr) + return 1 + except requests.HTTPError as exc: + print(f"Error: Ollama returned an HTTP error: {exc}", file=sys.stderr) + return 1 + except (OSError, UnicodeError, ValueError, json.JSONDecodeError) as exc: + print(f"Error: {exc}", file=sys.stderr) + return 1 + + print(f"Runtime seconds: {metadata['runtime_seconds']:.3f}") + print(f"Validation result: {'passed' if metadata['valid'] else 'failed'}") + print(f"Output: {metadata['output_path'] or 'not written'}") + print(f"Raw model response: {metadata['raw_response_path']}") + print(f"Validation report: {metadata['validation_report_path']}") + return 0 if metadata["valid"] else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/test_canonicalize.py b/tests/test_canonicalize.py index 0636ee6..f73db8c 100644 --- a/tests/test_canonicalize.py +++ b/tests/test_canonicalize.py @@ -19,6 +19,8 @@ EMPTY_EXTRACTION = { "technical": [], } +PROGEO_CONTEXT = Path("samples/real_live/progeo_meeting/meeting_context.yaml") + def write_extraction(directory: Path, name: str, data: dict) -> Path: path = directory / name @@ -163,6 +165,102 @@ class CanonicalizeTests(unittest.TestCase): self.assertEqual(output["stats"]["input_item_count"], 0) self.assertEqual(output["stats"]["output_item_count"], 0) + def test_responsible_aliases_normalize_with_meeting_context(self) -> None: + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + data = dict(EMPTY_EXTRACTION) + data["todos"] = [ + {"task": "A", "responsible": "Martin", "deadline": None, "evidence": "e"}, + {"task": "B", "responsible": "Herr Tazl", "deadline": None, "evidence": "e"}, + {"task": "C", "responsible": "Marlene", "deadline": None, "evidence": "e"}, + {"task": "D", "responsible": "Gothard", "deadline": None, "evidence": "e"}, + {"task": "E", "responsible": "Henning", "deadline": None, "evidence": "e"}, + ] + write_extraction(root, "chunk_01_extraction.json", data) + + output = canonicalize_extractions(root, meeting_context_path=PROGEO_CONTEXT) + responsibles = [item["responsible"] for item in output["items"]] + + self.assertEqual( + responsibles, + ["Martin", "Martin", "Marleen", "Gotthard", "Henning"], + ) + self.assertEqual(output["stats"]["responsibility_rejection_count"], 0) + self.assertEqual(output["stats"]["responsibility_normalization_count"], 3) + + def test_invalid_responsible_values_are_cleared_with_validation_report(self) -> None: + invalid_values = [ + "31. August", + "15. oder 16. September", + "27.8.", + "am 31.", + "morgen", + "nächste Woche", + "Frankfurt", + "Progeo", + "Secugrid", + "8:30 Uhr", + "Unbekannte Person", + ] + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + data = dict(EMPTY_EXTRACTION) + data["todos"] = [ + { + "task": f"Task {index}", + "responsible": value, + "deadline": None, + "evidence": "e", + } + for index, value in enumerate(invalid_values, start=1) + ] + data["todos"].append( + { + "task": "Null task", + "responsible": None, + "deadline": None, + "evidence": "e", + } + ) + write_extraction(root, "chunk_01_extraction.json", data) + + output = canonicalize_extractions(root, meeting_context_path=PROGEO_CONTEXT) + + self.assertTrue(all(item["responsible"] is None for item in output["items"])) + self.assertEqual( + output["stats"]["responsibility_rejection_count"], + len(invalid_values), + ) + rejected = output["responsibility_validations"] + self.assertEqual([item["original_value"] for item in rejected], invalid_values) + self.assertTrue(all(item["cleared_to_null"] for item in rejected)) + + def test_null_responsible_remains_valid_without_warning(self) -> None: + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + data = dict(EMPTY_EXTRACTION) + data["todos"] = [ + {"task": "Task", "responsible": None, "deadline": None, "evidence": "e"} + ] + write_extraction(root, "chunk_01_extraction.json", data) + + output = canonicalize_extractions(root, meeting_context_path=PROGEO_CONTEXT) + + self.assertIsNone(output["items"][0]["responsible"]) + self.assertEqual(output["responsibility_validations"], []) + + def test_responsible_validation_does_not_run_without_meeting_context(self) -> None: + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + data = dict(EMPTY_EXTRACTION) + data["todos"] = ["Update docs | Mira | Friday | I will update docs"] + write_extraction(root, "chunk_01_extraction.json", data) + + output = canonicalize_extractions(root) + + self.assertEqual(output["items"][0]["responsible"], "Mira") + self.assertEqual(output["responsibility_validations"], []) + if __name__ == "__main__": unittest.main() diff --git a/tests/test_consolidate_facts.py b/tests/test_consolidate_facts.py index 13ad7f9..5fe3c50 100644 --- a/tests/test_consolidate_facts.py +++ b/tests/test_consolidate_facts.py @@ -1,11 +1,18 @@ import unittest +import json +from pathlib import Path from src.meeting_lab.consolidation.consolidate_facts import ( ConsolidationValidationError, + DEFAULT_MIN_NUM_PREDICT, DEFAULT_NUM_PREDICT, build_consolidated_output, build_ollama_payload, + estimate_response_tokens, + fact_items, parse_model_json, + repair_model_group_coverage, + resolve_num_predict, validate_consolidated_output, validate_model_groups, ) @@ -98,6 +105,54 @@ class ConsolidateFactsTests(unittest.TestCase): self.assertEqual(payload["options"]["num_ctx"], 16384) self.assertEqual(payload["options"]["num_predict"], 1024) + def test_explicit_num_predict_is_preserved(self): + facts = [canonicalized_fixture()["items"][0]] + + self.assertEqual( + resolve_num_predict( + requested_num_predict=1234, + facts=facts, + prompt_token_estimate=100, + num_ctx=32768, + ), + 1234, + ) + + def test_adaptive_num_predict_scales_with_fact_payload(self): + fixture = canonicalized_fixture() + base_facts = fact_items(fixture) + larger_facts = [] + for index in range(80): + item = dict(base_facts[index % len(base_facts)]) + item["item_id"] = f"fact_{index + 1:04d}" + item["text"] = item["text"] + " " + ("detail " * 20) + item["evidence"] = item["evidence"] + " " + ("evidence " * 20) + larger_facts.append(item) + + resolved = resolve_num_predict( + requested_num_predict=None, + facts=larger_facts, + prompt_token_estimate=9000, + num_ctx=32768, + ) + + self.assertGreater(resolved, DEFAULT_MIN_NUM_PREDICT) + self.assertLessEqual(resolved, 32768 - 9000 - 1024) + + def test_progeo_context_benchmark_needs_more_than_fixed_default_when_available(self): + path = Path( + "samples/benchmarks/progeo_meeting_context_v1_20260804_110913/" + "canonicalizer/canonicalized_extractions.json" + ) + if not path.exists(): + self.skipTest("Progeo context benchmark artifact is not available.") + + canonicalized = json.loads(path.read_text(encoding="utf-8-sig")) + facts = fact_items(canonicalized) + + self.assertGreater(len(facts), 60) + self.assertGreater(estimate_response_tokens(facts), DEFAULT_NUM_PREDICT) + def test_grouping_validation_accepts_complete_singletons(self): groups = validate_model_groups( { @@ -154,6 +209,61 @@ class ConsolidateFactsTests(unittest.TestCase): {"fact_0001"}, ) + def test_source_coverage_repair_restores_missing_singletons(self): + fixture = canonicalized_fixture() + facts = fact_items(fixture) + repaired, changes = repair_model_group_coverage( + { + "groups": [ + { + "canonical_text": "The lead maintains the list.", + "source_item_ids": ["fact_0001"], + "merge_reason": "Singleton.", + } + ] + }, + facts, + ) + + groups = validate_model_groups(repaired, {"fact_0001", "fact_0002"}) + + self.assertEqual(len(groups), 2) + self.assertEqual(groups[1]["source_item_ids"], ["fact_0002"]) + self.assertEqual( + changes[0]["operation"], + "restore_missing_source_id_as_singleton", + ) + + def test_source_coverage_repair_removes_duplicate_occurrences(self): + fixture = canonicalized_fixture() + facts = fact_items(fixture) + repaired, changes = repair_model_group_coverage( + { + "groups": [ + { + "canonical_text": "Merged.", + "source_item_ids": ["fact_0001", "fact_0002"], + "merge_reason": "Same.", + }, + { + "canonical_text": "Duplicate.", + "source_item_ids": ["fact_0002"], + "merge_reason": "Duplicate.", + }, + ] + }, + facts, + ) + + groups = validate_model_groups(repaired, {"fact_0001", "fact_0002"}) + + self.assertEqual(len(groups), 1) + self.assertEqual(groups[0]["source_item_ids"], ["fact_0001", "fact_0002"]) + self.assertEqual( + [change["operation"] for change in changes], + ["remove_duplicate_source_id", "remove_empty_group"], + ) + def test_merged_group_validation(self): groups = validate_model_groups( { diff --git a/tests/test_render_working_protocol.py b/tests/test_render_working_protocol.py new file mode 100644 index 0000000..b60923f --- /dev/null +++ b/tests/test_render_working_protocol.py @@ -0,0 +1,179 @@ +import json +import tempfile +import unittest +from pathlib import Path +from unittest.mock import patch + +from src.meeting_lab.protocol.render_working_protocol import ( + REQUIRED_TITLE, + clean_working_protocol_markdown, + render_working_protocol, + validate_working_protocol_markdown, +) + + +VALID_PROTOCOL = """# Working Protocol + +## Materialversuche + +### Background + +Die Materialversuche wurden besprochen. + +### Decisions + +- Der nächste Versuch wird vorbereitet. + +### Action Items + +- Verantwortung offen: Materialstatus prüfen. + +### Open Questions + +- Wann liegen die Ergebnisse vor? +""" + + +class WorkingProtocolRendererTests(unittest.TestCase): + def test_valid_protocol_beginning_with_required_heading_passes(self) -> None: + report = validate_working_protocol_markdown(VALID_PROTOCOL) + + self.assertTrue(report.valid) + self.assertEqual(report.violations, []) + + def test_explanatory_prose_before_valid_protocol_is_removed(self) -> None: + cleaned = clean_working_protocol_markdown( + "Hier ist das Protokoll:\n\n" + VALID_PROTOCOL + ) + + self.assertTrue(cleaned.startswith(REQUIRED_TITLE + "\n")) + self.assertNotIn("Hier ist das Protokoll", cleaned) + self.assertTrue(validate_working_protocol_markdown(cleaned).valid) + + def test_leading_whitespace_before_valid_heading_is_removed(self) -> None: + cleaned = clean_working_protocol_markdown("\n \n # Working Protocol\n\n## Thema\n\n### Background\n\nText.\n") + + self.assertTrue(cleaned.startswith(REQUIRED_TITLE + "\n")) + self.assertTrue(validate_working_protocol_markdown(cleaned).valid) + + def test_output_without_required_heading_is_rejected(self) -> None: + report = validate_working_protocol_markdown("## Materialversuche\n\nText.\n") + + self.assertFalse(report.valid) + self.assertEqual(report.violations[0]["type"], "missing_required_heading") + + def test_categorized_summary_without_topic_body_is_rejected(self) -> None: + text = """Hier ist die Zusammenfassung: + +### **Entscheidungen (Decisions)** + +- Entscheidung. + +### **Handlungsaufträge (Action Items)** + +- Aufgabe. +""" + + cleaned = clean_working_protocol_markdown(text) + report = validate_working_protocol_markdown(cleaned) + + self.assertFalse(report.valid) + self.assertEqual(report.violations[0]["type"], "missing_required_heading") + + def test_categorized_summary_after_heading_is_rejected(self) -> None: + text = """# Working Protocol + +## Entscheidungen + +- Entscheidung. + +## Action Items + +- Aufgabe. +""" + + report = validate_working_protocol_markdown(text) + + self.assertFalse(report.valid) + violation_types = {violation["type"] for violation in report.violations} + self.assertIn("generic_summary_framing", violation_types) + self.assertIn("missing_topic_sections", violation_types) + + def test_malformed_markdown_is_rejected(self) -> None: + text = """# Working Protocol + +## Materialversuche + +### Background + +```json +{"unterbrochen": true} +""" + + report = validate_working_protocol_markdown(text) + + self.assertFalse(report.valid) + self.assertIn( + "malformed_markdown", + {violation["type"] for violation in report.violations}, + ) + + def test_raw_response_preserved_and_protocol_contains_only_validated_content(self) -> None: + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + input_path = root / "input.json" + input_path.write_text('{"items": []}\n', encoding="utf-8") + output_dir = root / "working_protocol" + + with patch( + "src.meeting_lab.protocol.render_working_protocol.call_ollama", + return_value=( + { + "response": "Einleitung.\n\n" + VALID_PROTOCOL, + "done": True, + "done_reason": "stop", + }, + 1.25, + ), + ): + metadata = render_working_protocol(input_path, output_dir) + + raw = json.loads((output_dir / "raw_model_response.json").read_text(encoding="utf-8")) + self.assertEqual(raw["response"], "Einleitung.\n\n" + VALID_PROTOCOL) + self.assertTrue((output_dir / "cleaned_candidate.md").exists()) + self.assertTrue((output_dir / "validation_report.json").exists()) + self.assertTrue(metadata["valid"]) + + protocol = (output_dir / "working_protocol.md").read_text(encoding="utf-8") + self.assertEqual(protocol, clean_working_protocol_markdown(raw["response"])) + self.assertNotIn("Einleitung.", protocol) + + def test_invalid_output_preserves_raw_but_does_not_write_protocol(self) -> None: + with tempfile.TemporaryDirectory() as directory: + root = Path(directory) + input_path = root / "input.json" + input_path.write_text('{"items": []}\n', encoding="utf-8") + output_dir = root / "working_protocol" + + with patch( + "src.meeting_lab.protocol.render_working_protocol.call_ollama", + return_value=( + { + "response": "Hier ist die Zusammenfassung:\n\n### Entscheidungen\n\n- X\n", + "done": True, + "done_reason": "stop", + }, + 1.25, + ), + ): + metadata = render_working_protocol(input_path, output_dir) + + self.assertFalse(metadata["valid"]) + self.assertTrue((output_dir / "raw_model_response.json").exists()) + self.assertTrue((output_dir / "cleaned_candidate.md").exists()) + self.assertTrue((output_dir / "validation_report.json").exists()) + self.assertFalse((output_dir / "working_protocol.md").exists()) + + +if __name__ == "__main__": + unittest.main()