Stabilize Meeting Lab pipeline for RC1 evaluation
This commit significantly improves the robustness and determinism of the Meeting Lab processing pipeline and establishes the first Release Candidate baseline for end-to-end evaluation. Highlights - BUG-009 - Implement deterministic responsible-party validation - Normalize participant aliases using Meeting Context - Reject invalid responsible values (dates, locations, technical terms, projects, products, unknown entities) - Record structured responsibility validation metadata - Add focused regression tests - BUG-010 - Implement adaptive num_predict estimation for Semantic Consolidator - Eliminate JSON truncation caused by fixed output limits - Add deterministic source coverage repair - Preserve strict post-repair validation - Add regression tests - BUG-011 - Implement Working Protocol V2 renderer contract enforcement - Preserve raw renderer responses - Reject invalid protocol output instead of accepting malformed documents - Add deterministic cleanup for harmless formatting deviations - Add focused renderer regression tests - Meeting Context - Validate Meeting Context V1 - Integrate authoritative participant alias normalization - Documentation - Update architecture documentation - Update output documentation - Update regression bug tracker The pipeline now fails safely instead of silently accepting invalid intermediate or final artifacts. Remaining work focuses primarily on extraction quality and semantic classification (decisions, action items, protocol faithfulness), rather than pipeline robustness.
This commit is contained in:
@@ -265,6 +265,11 @@ Deterministic Canonicalizer:
|
|||||||
- validates and normalizes extraction objects
|
- validates and normalizes extraction objects
|
||||||
- assigns stable source references and IDs
|
- assigns stable source references and IDs
|
||||||
- normalizes category names and basic field structure
|
- normalizes category names and basic field structure
|
||||||
|
- validates and normalizes action-item responsible fields against Meeting
|
||||||
|
Context when available: known participant and mentioned-person aliases are
|
||||||
|
normalized to canonical display names, while dates, locations, projects,
|
||||||
|
products, technical terms, generic process words and unknown free text are
|
||||||
|
cleared with a structured validation record
|
||||||
- performs only safe deterministic cleanup
|
- performs only safe deterministic cleanup
|
||||||
- may group exact duplicates
|
- may group exact duplicates
|
||||||
- preserves all source evidence
|
- preserves all source evidence
|
||||||
@@ -277,6 +282,12 @@ Semantic Consolidator:
|
|||||||
- V0 merges semantically equivalent fact items conservatively
|
- V0 merges semantically equivalent fact items conservatively
|
||||||
- V0 preserves source references and evidence
|
- V0 preserves source references and evidence
|
||||||
- V0 validates that every source fact appears exactly once
|
- V0 validates that every source fact appears exactly once
|
||||||
|
- V0 sizes its Ollama output budget from the actual fact payload instead of
|
||||||
|
using a fixed response cap for every meeting
|
||||||
|
- V0 may apply deterministic source-coverage repair after valid model JSON is
|
||||||
|
parsed: duplicate source IDs are removed after their first occurrence, empty
|
||||||
|
groups are removed and missing source facts are restored as singleton groups
|
||||||
|
from canonicalized input before strict validation runs
|
||||||
- V0 does not process non-fact categories semantically
|
- V0 does not process non-fact categories semantically
|
||||||
- later versions should group content by topic, mark contradictions and
|
- later versions should group content by topic, mark contradictions and
|
||||||
uncertainty, separate durable information from transient discussion and
|
uncertainty, separate durable information from transient discussion and
|
||||||
@@ -295,6 +306,11 @@ It only reformulates the analysis results for a specific audience and purpose.
|
|||||||
Depending on the output and maturity of the implementation, a renderer may be
|
Depending on the output and maturity of the implementation, a renderer may be
|
||||||
deterministic, template-based or LLM-assisted.
|
deterministic, template-based or LLM-assisted.
|
||||||
|
|
||||||
|
LLM-assisted renderers preserve raw model output separately and write the final
|
||||||
|
output artifact only after deterministic contract validation succeeds. Renderer
|
||||||
|
post-processing may remove non-semantic wrapper text, but must not fabricate
|
||||||
|
missing semantic sections or relabel an invalid summary as a valid output view.
|
||||||
|
|
||||||
The planned output products are:
|
The planned output products are:
|
||||||
|
|
||||||
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
|
||||||
|
|||||||
@@ -110,6 +110,21 @@ Characteristics:
|
|||||||
|
|
||||||
Completeness goal: optimize for recall and traceability.
|
Completeness goal: optimize for recall and traceability.
|
||||||
|
|
||||||
|
Current Working Protocol V2 renderer contract:
|
||||||
|
|
||||||
|
- output must begin exactly with `# Working Protocol`
|
||||||
|
- no explanatory preamble may appear before that heading
|
||||||
|
- content is organized by topic with `##` topic headings
|
||||||
|
- each topic may use only the supported `###` sections: `Background`,
|
||||||
|
`Decisions`, `Action Items`, `Open Questions`
|
||||||
|
- category-level report framing such as a global `## Decisions` / `## Action
|
||||||
|
Items` summary is not a valid topic-oriented Working Protocol
|
||||||
|
- raw model responses are preserved separately
|
||||||
|
- deterministic cleanup may remove leading prose before an already valid
|
||||||
|
`# Working Protocol` heading and normalize harmless heading whitespace
|
||||||
|
- `working_protocol.md` is written only after the cleaned candidate passes the
|
||||||
|
renderer contract validator
|
||||||
|
|
||||||
## Concise Distribution Protocol
|
## Concise Distribution Protocol
|
||||||
|
|
||||||
Suggested filename: `distribution_protocol.md`
|
Suggested filename: `distribution_protocol.md`
|
||||||
|
|||||||
@@ -476,3 +476,366 @@ Notes:
|
|||||||
|
|
||||||
This generalizes BUG-002 beyond the specific Jovana assignment case and should
|
This generalizes BUG-002 beyond the specific Jovana assignment case and should
|
||||||
be evaluated against the responsibility attribution invariant.
|
be evaluated against the responsibility attribution invariant.
|
||||||
|
|
||||||
|
## BUG-008
|
||||||
|
|
||||||
|
ID: BUG-008
|
||||||
|
|
||||||
|
Title: Semantic Consolidator emits duplicate source fact coverage
|
||||||
|
|
||||||
|
Pipeline stage: Semantic Consolidator V0 / Constraint Repair
|
||||||
|
|
||||||
|
Severity: High
|
||||||
|
|
||||||
|
Status: Verified
|
||||||
|
|
||||||
|
Date discovered: 2026-08-04
|
||||||
|
|
||||||
|
Version first observed: `progeo_meeting_20260804_083849`
|
||||||
|
|
||||||
|
Description:
|
||||||
|
|
||||||
|
The Progeo end-to-end evaluation stopped during Semantic Consolidator V0
|
||||||
|
validation because the preserved model grouping output assigned the same source
|
||||||
|
fact ID to more than one group.
|
||||||
|
|
||||||
|
Expected behaviour:
|
||||||
|
|
||||||
|
Every source fact ID must appear exactly once in the Semantic Consolidator V0
|
||||||
|
fact grouping output. The validator must reject duplicate or missing source
|
||||||
|
coverage before a consolidated document is rendered.
|
||||||
|
|
||||||
|
Actual behaviour:
|
||||||
|
|
||||||
|
The validator correctly rejected the preserved raw consolidator response with:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Source item IDs appear in multiple groups: ['fact_0012']
|
||||||
|
```
|
||||||
|
|
||||||
|
`fact_0012` appeared once in a merged group with `fact_0002` and again as a
|
||||||
|
singleton group. The generic validator report recorded this as one
|
||||||
|
`duplicate_id` violation with both structural occurrences.
|
||||||
|
|
||||||
|
Recovery result:
|
||||||
|
|
||||||
|
The preserved raw consolidator response was repaired with the generic
|
||||||
|
Constraint Repair V1 interface and deterministic structural operations:
|
||||||
|
|
||||||
|
- remove the repeated `fact_0012` occurrence from the later singleton group
|
||||||
|
- remove the now-empty singleton group
|
||||||
|
|
||||||
|
No Semantic Consolidator LLM call was rerun. No prompts, Meeting Context,
|
||||||
|
extraction outputs or canonicalizer output were modified. Second validation
|
||||||
|
succeeded: every source fact ID appears exactly once, no IDs are missing, no
|
||||||
|
empty groups remain, and non-fact categories remained unchanged.
|
||||||
|
|
||||||
|
This verifies the recovery path for this structural coverage failure. It does
|
||||||
|
not fix or change the Semantic Consolidator generation behaviour itself.
|
||||||
|
|
||||||
|
Related files:
|
||||||
|
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/raw_model_response.txt`
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_before.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_model_groups.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_consolidated_extractions.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_after.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repair_metadata.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_20260804_083849/working_protocol/working_protocol.md`
|
||||||
|
- `docs/constraint-repair.md`
|
||||||
|
|
||||||
|
Regression test available (yes/no): no
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
Verified. The recovery path repaired and revalidated this benchmark artifact,
|
||||||
|
then allowed the renderer to run once on the repaired consolidated output.
|
||||||
|
|
||||||
|
Notes:
|
||||||
|
|
||||||
|
The renderer output was produced, but it starts with explanatory prose instead
|
||||||
|
of the required `# Working Protocol` heading. That is a renderer faithfulness
|
||||||
|
issue, not part of the Semantic Consolidator coverage repair verified here.
|
||||||
|
|
||||||
|
## BUG-009
|
||||||
|
|
||||||
|
ID: BUG-009
|
||||||
|
|
||||||
|
Title: Invalid responsible-party values survive canonicalization
|
||||||
|
|
||||||
|
Pipeline stage: Extraction / Deterministic Canonicalizer
|
||||||
|
|
||||||
|
Severity: High
|
||||||
|
|
||||||
|
Status: Verified
|
||||||
|
|
||||||
|
Date discovered: 2026-08-04
|
||||||
|
|
||||||
|
Version first observed: `progeo_meeting_context_v1_20260804_110913`
|
||||||
|
|
||||||
|
Description:
|
||||||
|
|
||||||
|
The context-aware Progeo benchmark produced action items whose `responsible`
|
||||||
|
field contained dates or date fragments instead of responsible entities. The
|
||||||
|
canonicalizer accepted these values as plain strings.
|
||||||
|
|
||||||
|
Expected behaviour:
|
||||||
|
|
||||||
|
When Meeting Context is available, action-item responsible values should be
|
||||||
|
deterministically resolved to known participants, mentioned people or supported
|
||||||
|
organization entities. Dates, locations, projects, products, technical terms,
|
||||||
|
generic process words and unknown free text must not survive as responsible
|
||||||
|
parties.
|
||||||
|
|
||||||
|
Actual behaviour:
|
||||||
|
|
||||||
|
The previous canonicalized Progeo artifact contained invalid responsible
|
||||||
|
values such as:
|
||||||
|
|
||||||
|
- `31. August`
|
||||||
|
- `15. oder 16. September`
|
||||||
|
- `am 31.`
|
||||||
|
- `27.8.`
|
||||||
|
|
||||||
|
Root cause:
|
||||||
|
|
||||||
|
The extraction schema allowed `responsible` to be a free-form string or null.
|
||||||
|
Canonicalizer V1 parsed and preserved that string without checking it against
|
||||||
|
Meeting Context or obvious non-person/non-organization patterns.
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
|
||||||
|
Canonicalizer V1 now validates responsible fields when Meeting Context is
|
||||||
|
available either through `--meeting-context` or extraction provenance. It:
|
||||||
|
|
||||||
|
- accepts exact participant and mentioned-person display names
|
||||||
|
- accepts participant and mentioned-person aliases
|
||||||
|
- normalizes accepted aliases to canonical display names
|
||||||
|
- accepts explicit organization/department names only when represented in the
|
||||||
|
current Meeting Context structure
|
||||||
|
- rejects dates, relative dates, weekdays, clock times, locations, projects,
|
||||||
|
products, systems, technical terms, generic process words and unknown free
|
||||||
|
text
|
||||||
|
- clears rejected responsible values to null
|
||||||
|
- records structured `responsibility_validation` data and an aggregate
|
||||||
|
`responsibility_validations` report
|
||||||
|
|
||||||
|
Verification:
|
||||||
|
|
||||||
|
The existing context-aware Progeo extraction artifacts were re-canonicalized
|
||||||
|
without rerunning chunking, normalization or extraction:
|
||||||
|
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
|
||||||
|
|
||||||
|
Action item count remained 28 and semantic content other than responsible
|
||||||
|
fields remained unchanged. Invalid date-like responsible values were cleared.
|
||||||
|
`Martin Tazl` normalized to `Martin`, and `Marleen` remained valid. `Marleen
|
||||||
|
Wever` was rejected because the current Progeo Meeting Context does not list it
|
||||||
|
as a display name or alias.
|
||||||
|
|
||||||
|
Related files:
|
||||||
|
|
||||||
|
- `src/meeting_lab/consolidation/canonicalize.py`
|
||||||
|
- `tests/test_canonicalize.py`
|
||||||
|
- `samples/real_live/progeo_meeting/meeting_context.yaml`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer/canonicalized_extractions.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
|
||||||
|
|
||||||
|
Regression test available (yes/no): yes
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
Verified. Focused tests pass and the Progeo downstream-only regression demo
|
||||||
|
removes the observed invalid responsible values without changing action-item
|
||||||
|
count or non-responsible semantic content.
|
||||||
|
|
||||||
|
Notes:
|
||||||
|
|
||||||
|
This does not prove that remaining accepted responsibilities are semantically
|
||||||
|
supported by transcript evidence. It only prevents invalid entity values from
|
||||||
|
surviving in the `responsible` field.
|
||||||
|
|
||||||
|
## BUG-010
|
||||||
|
|
||||||
|
ID: BUG-010
|
||||||
|
|
||||||
|
Title: Semantic Consolidator JSON truncates on larger real-life meetings
|
||||||
|
|
||||||
|
Pipeline stage: Semantic Consolidator V0
|
||||||
|
|
||||||
|
Severity: High
|
||||||
|
|
||||||
|
Status: Fixed
|
||||||
|
|
||||||
|
Date discovered: 2026-08-04
|
||||||
|
|
||||||
|
Version first observed: `progeo_meeting_context_v1_20260804_110913`
|
||||||
|
|
||||||
|
Description:
|
||||||
|
|
||||||
|
The context-aware Progeo benchmark produced substantially more canonical fact
|
||||||
|
items than the earlier run. Semantic Consolidator V0 used the fixed committed
|
||||||
|
`num_predict` value of 4096 for its single grouping response and the preserved
|
||||||
|
raw response stopped in the middle of a JSON group.
|
||||||
|
|
||||||
|
Expected behaviour:
|
||||||
|
|
||||||
|
The consolidator should allocate enough output budget for the expected
|
||||||
|
fact-group JSON, preserve the raw model response, parse only valid JSON and
|
||||||
|
validate exact source fact coverage before writing consolidated output.
|
||||||
|
|
||||||
|
Actual behaviour:
|
||||||
|
|
||||||
|
The raw response ended after `fact_0060`/start of the next group, while Ollama
|
||||||
|
reported `eval_count: 4096`, exactly matching the old response cap. Parsing
|
||||||
|
failed before source coverage validation:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Invalid model JSON: Expecting property name enclosed in double quotes
|
||||||
|
```
|
||||||
|
|
||||||
|
Root cause:
|
||||||
|
|
||||||
|
Output was truncated at the fixed `num_predict` cap. The prompt asks the model
|
||||||
|
to emit one JSON group per fact unless duplicates are found, so output size
|
||||||
|
scales with fact count and fact text size. The fixed cap was adequate for the
|
||||||
|
previous smaller Progeo run but not for the context-aware run.
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
|
||||||
|
Semantic Consolidator V0 now estimates the response budget from the actual fact
|
||||||
|
payload and context-window headroom when `--num-predict` is not explicitly set.
|
||||||
|
Explicit `--num-predict` values are still respected.
|
||||||
|
|
||||||
|
After valid model JSON is parsed, the consolidator also applies deterministic
|
||||||
|
source-coverage repair before strict validation:
|
||||||
|
|
||||||
|
- duplicate source IDs after the first occurrence are removed
|
||||||
|
- empty groups created by removal are dropped
|
||||||
|
- missing source fact IDs are restored as singleton groups from the
|
||||||
|
canonicalized input
|
||||||
|
|
||||||
|
This repair does not rewrite existing model group text, invent facts or create
|
||||||
|
new semantic merges. Strict validation still runs after repair and can still
|
||||||
|
reject the output.
|
||||||
|
|
||||||
|
Verification:
|
||||||
|
|
||||||
|
The Progeo context canonicalized benchmark was rerun through the fixed
|
||||||
|
Semantic Consolidator into:
|
||||||
|
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/`
|
||||||
|
|
||||||
|
The fixed run used `num_predict: 14347`, the model returned valid JSON with
|
||||||
|
`eval_count: 4568`, deterministic repair restored seven omitted source fact
|
||||||
|
IDs as singletons and final validation passed with 77 unique source fact IDs,
|
||||||
|
no duplicates and no missing IDs.
|
||||||
|
|
||||||
|
Related files:
|
||||||
|
|
||||||
|
- `src/meeting_lab/consolidation/consolidate_facts.py`
|
||||||
|
- `tests/test_consolidate_facts.py`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator/raw_model_response.txt`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/consolidated_extractions.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/repair_metadata.json`
|
||||||
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/report.md`
|
||||||
|
|
||||||
|
Regression test available (yes/no): yes
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
Fixed for the observed truncation failure and protected by focused
|
||||||
|
consolidator tests. This does not improve semantic merge quality; it only
|
||||||
|
prevents fixed-budget truncation and enforces complete source coverage
|
||||||
|
deterministically.
|
||||||
|
|
||||||
|
## BUG-011
|
||||||
|
|
||||||
|
ID: BUG-011
|
||||||
|
|
||||||
|
Title: Working Protocol Renderer V2 writes contract-invalid output
|
||||||
|
|
||||||
|
Pipeline stage: Working Protocol Renderer V2
|
||||||
|
|
||||||
|
Severity: High
|
||||||
|
|
||||||
|
Status: Verified
|
||||||
|
|
||||||
|
Date discovered: 2026-08-04
|
||||||
|
|
||||||
|
Version first observed: `progeo_meeting_rc1_20260804_121502`
|
||||||
|
|
||||||
|
Description:
|
||||||
|
|
||||||
|
The RC1 Progeo renderer completed its LLM call and wrote `working_protocol.md`,
|
||||||
|
but the generated document did not start with `# Working Protocol`. It started
|
||||||
|
with explanatory prose and used category-summary/report framing instead of the
|
||||||
|
committed topic-oriented Working Protocol V2 structure.
|
||||||
|
|
||||||
|
Expected behaviour:
|
||||||
|
|
||||||
|
The renderer must preserve raw model output separately, deterministically clean
|
||||||
|
only harmless wrapper text, validate the cleaned candidate against the committed
|
||||||
|
Working Protocol V2 contract and write `working_protocol.md` only when the
|
||||||
|
candidate is valid.
|
||||||
|
|
||||||
|
Actual behaviour:
|
||||||
|
|
||||||
|
The one-off renderer path used for the benchmark wrote the raw LLM text to
|
||||||
|
`working_protocol.md` even though metadata recorded `valid=false` and
|
||||||
|
`readable_markdown=false`.
|
||||||
|
|
||||||
|
Root cause:
|
||||||
|
|
||||||
|
The committed contract existed in `prompts/working_protocol.md`, but there was
|
||||||
|
no reusable Working Protocol V2 renderer implementation with a strict output
|
||||||
|
validator. The model ignored explicit prompt instructions, and the pipeline did
|
||||||
|
not enforce the contract deterministically before preserving the final protocol
|
||||||
|
file.
|
||||||
|
|
||||||
|
Fix:
|
||||||
|
|
||||||
|
`src/meeting_lab/protocol/render_working_protocol.py` now implements the
|
||||||
|
Working Protocol V2 renderer path. It:
|
||||||
|
|
||||||
|
- preserves the raw Ollama JSON response and raw response text
|
||||||
|
- removes leading prose only when a real `# Working Protocol` heading exists
|
||||||
|
- normalizes harmless heading whitespace
|
||||||
|
- rejects output without the required heading
|
||||||
|
- rejects category-summary framing that lacks topic-oriented body structure
|
||||||
|
- rejects malformed Markdown such as unclosed fenced code blocks
|
||||||
|
- writes `working_protocol.md` only after validation succeeds
|
||||||
|
|
||||||
|
Verification:
|
||||||
|
|
||||||
|
Focused renderer tests cover valid output, deterministic preamble cleanup,
|
||||||
|
leading whitespace cleanup, missing heading rejection, categorized-summary
|
||||||
|
rejection, malformed Markdown rejection, raw response preservation and ensuring
|
||||||
|
`working_protocol.md` contains only validated content.
|
||||||
|
|
||||||
|
RC1 downstream demo:
|
||||||
|
|
||||||
|
- Existing RC1 raw renderer output contained no embedded valid `# Working
|
||||||
|
Protocol` body.
|
||||||
|
- The new renderer path was run once against the existing RC1 consolidated
|
||||||
|
input with current committed settings.
|
||||||
|
- The raw response and cleaned candidate were preserved.
|
||||||
|
- Validation correctly rejected the output and did not write a final
|
||||||
|
`working_protocol.md`.
|
||||||
|
|
||||||
|
Related files:
|
||||||
|
|
||||||
|
- `src/meeting_lab/protocol/render_working_protocol.py`
|
||||||
|
- `tests/test_render_working_protocol.py`
|
||||||
|
- `docs/output-views.md`
|
||||||
|
- `samples/benchmarks/progeo_meeting_rc1_20260804_121502/working_protocol_bug011/`
|
||||||
|
|
||||||
|
Regression test available (yes/no): yes
|
||||||
|
|
||||||
|
Current status:
|
||||||
|
|
||||||
|
Verified. The renderer contract is now enforced deterministically. This does
|
||||||
|
not improve semantic quality of the generated prose; it prevents invalid
|
||||||
|
renderer output from being accepted as a final Working Protocol.
|
||||||
|
|||||||
@@ -2,29 +2,124 @@ schema_version: "1"
|
|||||||
|
|
||||||
meeting:
|
meeting:
|
||||||
meeting_id: "progeo-meeting"
|
meeting_id: "progeo-meeting"
|
||||||
title: "Progeo meeting"
|
title: "Progeo Meeting"
|
||||||
language: "de"
|
language: "de"
|
||||||
date: null
|
date: 2026-07-27
|
||||||
objective: ""
|
objective: >-
|
||||||
notes: "Meeting Context scaffold for Real-Life Reference Meeting #2. Participant and entity metadata require manual completion."
|
Ergebnisbericht zur abgeschlossenen Versuchsproduktion eines Geogitters
|
||||||
|
aus recyceltem Polypropylen und Abstimmung weiterer Großversuche.
|
||||||
|
notes: >-
|
||||||
|
Real-Life Reference Meeting #2. Beteiligte externe Organisationen:
|
||||||
|
Fraunhofer, MAS, Universität Leeds, FH Münster, Kiwa und Kockmann.
|
||||||
|
|
||||||
participants: []
|
participants:
|
||||||
|
- participant_id: "martin"
|
||||||
|
display_name: "Martin"
|
||||||
|
aliases: [Martin Tazl, Herr Tazl]
|
||||||
|
role: "Entwicklungsleiter Naue"
|
||||||
|
department: "naue"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
mentioned_people: []
|
- participant_id: "james"
|
||||||
|
display_name: "James"
|
||||||
|
aliases: [James Nachtigall, Herr Nachtigall]
|
||||||
|
role: "Entwicklungsingenieur"
|
||||||
|
department: "naue"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "david"
|
||||||
|
display_name: "David"
|
||||||
|
aliases: []
|
||||||
|
role: "Entwicklungsingenieur"
|
||||||
|
department: "naue"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "gotthard"
|
||||||
|
display_name: "Gotthard"
|
||||||
|
aliases: [Gothard, Gotthardt, Gothardt, Herr Walter]
|
||||||
|
role: "Projektleiter"
|
||||||
|
department: "fh-muenster"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "tim"
|
||||||
|
display_name: "Tim"
|
||||||
|
aliases: [Timm, Tim Schulte-Uebbing, Herr Schulte-Uebbing]
|
||||||
|
role: "PhD"
|
||||||
|
department: "fh-muenster"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "antonius"
|
||||||
|
display_name: "Antonius"
|
||||||
|
aliases: []
|
||||||
|
role: "PhD"
|
||||||
|
department: "fh-muenster"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "marleen"
|
||||||
|
display_name: "Marleen"
|
||||||
|
aliases: [Marlene, Frau Wever, Marleen Wever]
|
||||||
|
role: "Entwicklungsingenieurin"
|
||||||
|
department: "huesker"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
mentioned_people:
|
||||||
|
- person_id: "henning"
|
||||||
|
display_name: "Henning"
|
||||||
|
aliases: []
|
||||||
|
role: null
|
||||||
|
department: null
|
||||||
|
attendance_status: "not_present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
organization:
|
organization:
|
||||||
name: null
|
name: null
|
||||||
|
|
||||||
departments: []
|
departments:
|
||||||
|
- id: "naue"
|
||||||
|
name: "Naue"
|
||||||
|
aliases: [Naue GmbH & Co. KG]
|
||||||
|
|
||||||
abbreviations: {}
|
- id: "huesker"
|
||||||
|
name: "Huesker"
|
||||||
|
aliases: [HÜSKER]
|
||||||
|
|
||||||
|
- id: "fh-muenster"
|
||||||
|
name: "FH Münster"
|
||||||
|
aliases: [Fachhochschule Münster]
|
||||||
|
|
||||||
|
abbreviations:
|
||||||
|
RC: "Recycling"
|
||||||
|
PP: "Polypropylen"
|
||||||
|
PET: "Polyethylenterephthalat"
|
||||||
|
MFI: "Melt Flow Index"
|
||||||
|
IV: "intrinsische Viskosität"
|
||||||
|
SSP: "Solid-State Polymerization"
|
||||||
|
|
||||||
known_entities:
|
known_entities:
|
||||||
projects: []
|
projects: [Progeo]
|
||||||
products: []
|
products: [Secugrid, Combigrid]
|
||||||
systems: []
|
systems: []
|
||||||
locations: []
|
locations: [Stettin]
|
||||||
technical_terms: []
|
technical_terms:
|
||||||
|
- Rezyklat
|
||||||
|
- Virgin-Material
|
||||||
|
- Glührückstand
|
||||||
|
- Verstrecken
|
||||||
|
- Streckwerk
|
||||||
|
- Bruchspannung
|
||||||
|
- Zugfestigkeit
|
||||||
|
- Dehnung
|
||||||
|
- Trommelsieb
|
||||||
|
- chemisches Recycling
|
||||||
|
- mechanisches Recycling
|
||||||
|
- Nachkondensation
|
||||||
|
|
||||||
context_rules:
|
context_rules:
|
||||||
participant_list_is_authoritative: true
|
participant_list_is_authoritative: true
|
||||||
@@ -32,4 +127,4 @@ context_rules:
|
|||||||
do_not_infer_departments: true
|
do_not_infer_departments: true
|
||||||
do_not_infer_responsibilities: true
|
do_not_infer_responsibilities: true
|
||||||
do_not_infer_attendance: true
|
do_not_infer_attendance: true
|
||||||
mentioned_people_are_not_participants: true
|
mentioned_people_are_not_participants: true
|
||||||
@@ -8,9 +8,12 @@ import json
|
|||||||
import re
|
import re
|
||||||
import sys
|
import sys
|
||||||
from collections import Counter
|
from collections import Counter
|
||||||
|
from dataclasses import dataclass
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
|
from src.meeting_lab.models.meeting_context import load_meeting_context
|
||||||
|
|
||||||
|
|
||||||
SCHEMA_VERSION = "1"
|
SCHEMA_VERSION = "1"
|
||||||
|
|
||||||
@@ -36,6 +39,51 @@ CANONICAL_CATEGORIES = tuple(CATEGORY_NAMES.values())
|
|||||||
|
|
||||||
CHUNK_EXTRACTION_RE = re.compile(r"^chunk_(\d+)_extraction\.json$")
|
CHUNK_EXTRACTION_RE = re.compile(r"^chunk_(\d+)_extraction\.json$")
|
||||||
WHITESPACE_RE = re.compile(r"\s+")
|
WHITESPACE_RE = re.compile(r"\s+")
|
||||||
|
RESPONSIBLE_KEY_RE = re.compile(r"\s+")
|
||||||
|
NUMERIC_DATE_RE = re.compile(r"(?i)\b\d{1,2}\s*[./]\s*(?:\d{1,2}|[a-zäöü]+)?\b")
|
||||||
|
MONTH_DATE_RE = re.compile(
|
||||||
|
r"(?i)\b\d{1,2}\.?\s*(?:oder\s+\d{1,2}\.?\s*)?"
|
||||||
|
r"(januar|februar|märz|maerz|april|mai|juni|juli|august|september|oktober|november|dezember)\b"
|
||||||
|
)
|
||||||
|
TIME_RE = re.compile(r"(?i)\b\d{1,2}[:.]\d{2}\s*(?:uhr)?\b|\b\d{1,2}\s*uhr\b")
|
||||||
|
DATE_FRAGMENT_RE = re.compile(r"(?i)\b(?:am|zum|bis|vor|nach)\s+\d{1,2}\.?\b")
|
||||||
|
|
||||||
|
DATE_WORDS = {
|
||||||
|
"heute",
|
||||||
|
"morgen",
|
||||||
|
"übermorgen",
|
||||||
|
"uebermorgen",
|
||||||
|
"gestern",
|
||||||
|
"vorgestern",
|
||||||
|
}
|
||||||
|
WEEKDAY_WORDS = {
|
||||||
|
"montag",
|
||||||
|
"dienstag",
|
||||||
|
"mittwoch",
|
||||||
|
"donnerstag",
|
||||||
|
"freitag",
|
||||||
|
"samstag",
|
||||||
|
"sonntag",
|
||||||
|
}
|
||||||
|
RELATIVE_DATE_PHRASES = {
|
||||||
|
"nächste woche",
|
||||||
|
"naechste woche",
|
||||||
|
"diese woche",
|
||||||
|
"kommende woche",
|
||||||
|
"nächsten monat",
|
||||||
|
"naechsten monat",
|
||||||
|
}
|
||||||
|
GENERIC_RESPONSIBLE_WORDS = {
|
||||||
|
"team",
|
||||||
|
"projektteam",
|
||||||
|
"projektleitung",
|
||||||
|
"logistik",
|
||||||
|
"alle",
|
||||||
|
"autor",
|
||||||
|
"labor",
|
||||||
|
"partner",
|
||||||
|
"gruppe",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def parse_args() -> argparse.Namespace:
|
def parse_args() -> argparse.Namespace:
|
||||||
@@ -59,9 +107,110 @@ def parse_args() -> argparse.Namespace:
|
|||||||
action="store_true",
|
action="store_true",
|
||||||
help="Preserve exact duplicate items instead of merging them.",
|
help="Preserve exact duplicate items instead of merging them.",
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--meeting-context",
|
||||||
|
type=Path,
|
||||||
|
help=(
|
||||||
|
"Optional Meeting Context V1 YAML used to validate and normalize "
|
||||||
|
"action-item responsible fields. If omitted, canonicalizer uses "
|
||||||
|
"context provenance from extraction JSON when available."
|
||||||
|
),
|
||||||
|
)
|
||||||
return parser.parse_args()
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ResponsiblePartyNormalizer:
|
||||||
|
allowed_names: dict[str, str]
|
||||||
|
invalid_values: dict[str, str]
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def from_meeting_context_data(cls, data: dict[str, Any]) -> "ResponsiblePartyNormalizer":
|
||||||
|
allowed: dict[str, str] = {}
|
||||||
|
invalid: dict[str, str] = {}
|
||||||
|
|
||||||
|
for collection, id_key in (
|
||||||
|
("participants", "participant_id"),
|
||||||
|
("mentioned_people", "person_id"),
|
||||||
|
):
|
||||||
|
for person in data.get(collection, []) or []:
|
||||||
|
if not isinstance(person, dict):
|
||||||
|
continue
|
||||||
|
display_name = clean_text(person.get("display_name"))
|
||||||
|
if not display_name:
|
||||||
|
continue
|
||||||
|
for value in [display_name, *(person.get("aliases") or [])]:
|
||||||
|
text = clean_text(value)
|
||||||
|
if text:
|
||||||
|
allowed[responsible_key(text)] = display_name
|
||||||
|
|
||||||
|
organization = data.get("organization")
|
||||||
|
if isinstance(organization, dict):
|
||||||
|
organization_name = clean_text(organization.get("name"))
|
||||||
|
if organization_name:
|
||||||
|
allowed[responsible_key(organization_name)] = organization_name
|
||||||
|
|
||||||
|
for department in organization.get("departments") or []:
|
||||||
|
if not isinstance(department, dict):
|
||||||
|
continue
|
||||||
|
department_name = clean_text(department.get("name"))
|
||||||
|
if department_name:
|
||||||
|
allowed[responsible_key(department_name)] = department_name
|
||||||
|
for alias in department.get("aliases") or []:
|
||||||
|
text = clean_text(alias)
|
||||||
|
if text and department_name:
|
||||||
|
allowed[responsible_key(text)] = department_name
|
||||||
|
|
||||||
|
known_entities = data.get("known_entities")
|
||||||
|
if isinstance(known_entities, dict):
|
||||||
|
for collection in (
|
||||||
|
"projects",
|
||||||
|
"products",
|
||||||
|
"systems",
|
||||||
|
"locations",
|
||||||
|
"technical_terms",
|
||||||
|
):
|
||||||
|
for value in known_entities.get(collection) or []:
|
||||||
|
text = clean_text(value)
|
||||||
|
if text:
|
||||||
|
invalid[responsible_key(text)] = collection
|
||||||
|
|
||||||
|
return cls(allowed_names=allowed, invalid_values=invalid)
|
||||||
|
|
||||||
|
def normalize(
|
||||||
|
self,
|
||||||
|
value: str | None,
|
||||||
|
item_id: str,
|
||||||
|
) -> tuple[str | None, dict[str, Any] | None]:
|
||||||
|
original = clean_text(value)
|
||||||
|
if original is None:
|
||||||
|
return None, None
|
||||||
|
|
||||||
|
key = responsible_key(original)
|
||||||
|
canonical = self.allowed_names.get(key)
|
||||||
|
if canonical is not None:
|
||||||
|
return canonical, {
|
||||||
|
"action_item_id": item_id,
|
||||||
|
"field": "responsible",
|
||||||
|
"original_value": original,
|
||||||
|
"normalized_value": canonical,
|
||||||
|
"status": "accepted",
|
||||||
|
"reason": "resolved_to_meeting_context_entity",
|
||||||
|
"cleared_to_null": False,
|
||||||
|
}
|
||||||
|
|
||||||
|
reason = invalid_responsible_reason(original, self.invalid_values)
|
||||||
|
return None, {
|
||||||
|
"action_item_id": item_id,
|
||||||
|
"field": "responsible",
|
||||||
|
"original_value": original,
|
||||||
|
"normalized_value": None,
|
||||||
|
"status": "rejected",
|
||||||
|
"reason": reason,
|
||||||
|
"cleared_to_null": True,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
def chunk_sort_key(path: Path) -> tuple[int, str]:
|
def chunk_sort_key(path: Path) -> tuple[int, str]:
|
||||||
match = CHUNK_EXTRACTION_RE.match(path.name)
|
match = CHUNK_EXTRACTION_RE.match(path.name)
|
||||||
if not match:
|
if not match:
|
||||||
@@ -109,6 +258,30 @@ def clean_text(value: Any) -> str | None:
|
|||||||
return text if text else None
|
return text if text else None
|
||||||
|
|
||||||
|
|
||||||
|
def responsible_key(value: str) -> str:
|
||||||
|
return RESPONSIBLE_KEY_RE.sub(" ", value.casefold()).strip(" .,:;")
|
||||||
|
|
||||||
|
|
||||||
|
def invalid_responsible_reason(
|
||||||
|
value: str,
|
||||||
|
invalid_values: dict[str, str],
|
||||||
|
) -> str:
|
||||||
|
key = responsible_key(value)
|
||||||
|
if key in invalid_values:
|
||||||
|
return f"known_{invalid_values[key]}_not_responsible_entity"
|
||||||
|
if key in GENERIC_RESPONSIBLE_WORDS:
|
||||||
|
return "generic_process_word"
|
||||||
|
if key in DATE_WORDS or key in WEEKDAY_WORDS or key in RELATIVE_DATE_PHRASES:
|
||||||
|
return "date_or_relative_date"
|
||||||
|
if NUMERIC_DATE_RE.search(value) or MONTH_DATE_RE.search(value):
|
||||||
|
return "date_or_date_range"
|
||||||
|
if TIME_RE.search(value):
|
||||||
|
return "clock_time"
|
||||||
|
if DATE_FRAGMENT_RE.search(value):
|
||||||
|
return "date_fragment"
|
||||||
|
return "unknown_responsible_entity"
|
||||||
|
|
||||||
|
|
||||||
def split_legacy_string(value: str) -> list[str]:
|
def split_legacy_string(value: str) -> list[str]:
|
||||||
return [part.strip() for part in value.split("|")]
|
return [part.strip() for part in value.split("|")]
|
||||||
|
|
||||||
@@ -324,6 +497,7 @@ def canonicalize_value(
|
|||||||
source_file: str,
|
source_file: str,
|
||||||
source_index: int,
|
source_index: int,
|
||||||
counts: Counter[str],
|
counts: Counter[str],
|
||||||
|
responsible_normalizer: ResponsiblePartyNormalizer | None = None,
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
parsed = PARSERS[category](value)
|
parsed = PARSERS[category](value)
|
||||||
item: dict[str, Any] = {
|
item: dict[str, Any] = {
|
||||||
@@ -337,6 +511,14 @@ def canonicalize_value(
|
|||||||
}
|
}
|
||||||
for key, parsed_value in parsed.items():
|
for key, parsed_value in parsed.items():
|
||||||
item[key] = parsed_value
|
item[key] = parsed_value
|
||||||
|
if category == "action_item" and responsible_normalizer is not None:
|
||||||
|
normalized, validation = responsible_normalizer.normalize(
|
||||||
|
item.get("responsible"),
|
||||||
|
item["item_id"],
|
||||||
|
)
|
||||||
|
item["responsible"] = normalized
|
||||||
|
if validation is not None:
|
||||||
|
item["responsibility_validation"] = validation
|
||||||
item["source_references"] = [
|
item["source_references"] = [
|
||||||
source_reference(source_file, source_index, value, item["evidence"])
|
source_reference(source_file, source_index, value, item["evidence"])
|
||||||
]
|
]
|
||||||
@@ -367,6 +549,7 @@ def merge_exact_duplicates(items: list[dict[str, Any]]) -> tuple[list[dict[str,
|
|||||||
def canonicalize_extractions(
|
def canonicalize_extractions(
|
||||||
input_dir: Path,
|
input_dir: Path,
|
||||||
merge_duplicates: bool = True,
|
merge_duplicates: bool = True,
|
||||||
|
meeting_context_path: Path | None = None,
|
||||||
) -> dict[str, Any]:
|
) -> dict[str, Any]:
|
||||||
files = find_extraction_files(input_dir)
|
files = find_extraction_files(input_dir)
|
||||||
if not files:
|
if not files:
|
||||||
@@ -375,9 +558,22 @@ def canonicalize_extractions(
|
|||||||
counts: Counter[str] = Counter()
|
counts: Counter[str] = Counter()
|
||||||
input_counts: Counter[str] = Counter()
|
input_counts: Counter[str] = Counter()
|
||||||
items: list[dict[str, Any]] = []
|
items: list[dict[str, Any]] = []
|
||||||
|
responsible_normalizer: ResponsiblePartyNormalizer | None = None
|
||||||
|
responsibility_validations: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
loaded_files: list[tuple[Path, dict[str, Any]]] = []
|
||||||
for path in files:
|
for path in files:
|
||||||
data = load_json_object(path)
|
data = load_json_object(path)
|
||||||
|
loaded_files.append((path, data))
|
||||||
|
|
||||||
|
context_path = meeting_context_path or infer_meeting_context_path(loaded_files)
|
||||||
|
if context_path is not None:
|
||||||
|
meeting_context = load_meeting_context(context_path)
|
||||||
|
responsible_normalizer = ResponsiblePartyNormalizer.from_meeting_context_data(
|
||||||
|
meeting_context.data
|
||||||
|
)
|
||||||
|
|
||||||
|
for path, data in loaded_files:
|
||||||
validate_required_categories(data, path)
|
validate_required_categories(data, path)
|
||||||
for raw_category in REQUIRED_CATEGORIES:
|
for raw_category in REQUIRED_CATEGORIES:
|
||||||
category = CATEGORY_NAMES[raw_category]
|
category = CATEGORY_NAMES[raw_category]
|
||||||
@@ -391,8 +587,16 @@ def canonicalize_extractions(
|
|||||||
source_file=path.name,
|
source_file=path.name,
|
||||||
source_index=source_index,
|
source_index=source_index,
|
||||||
counts=counts,
|
counts=counts,
|
||||||
|
responsible_normalizer=responsible_normalizer,
|
||||||
)
|
)
|
||||||
)
|
)
|
||||||
|
if (
|
||||||
|
items[-1].get("category") == "action_item"
|
||||||
|
and "responsibility_validation" in items[-1]
|
||||||
|
):
|
||||||
|
responsibility_validations.append(
|
||||||
|
items[-1]["responsibility_validation"]
|
||||||
|
)
|
||||||
|
|
||||||
exact_duplicates = 0
|
exact_duplicates = 0
|
||||||
if merge_duplicates:
|
if merge_duplicates:
|
||||||
@@ -415,11 +619,40 @@ def canonicalize_extractions(
|
|||||||
"input_item_count_by_category": input_counts_by_category,
|
"input_item_count_by_category": input_counts_by_category,
|
||||||
"output_item_count_by_category": output_counts_by_category,
|
"output_item_count_by_category": output_counts_by_category,
|
||||||
"exact_duplicates_merged": exact_duplicates,
|
"exact_duplicates_merged": exact_duplicates,
|
||||||
|
"responsibility_validation_count": len(responsibility_validations),
|
||||||
|
"responsibility_rejection_count": sum(
|
||||||
|
1
|
||||||
|
for validation in responsibility_validations
|
||||||
|
if validation.get("status") == "rejected"
|
||||||
|
),
|
||||||
|
"responsibility_normalization_count": sum(
|
||||||
|
1
|
||||||
|
for validation in responsibility_validations
|
||||||
|
if validation.get("status") == "accepted"
|
||||||
|
and validation.get("original_value") != validation.get("normalized_value")
|
||||||
|
),
|
||||||
},
|
},
|
||||||
|
"responsibility_validations": responsibility_validations,
|
||||||
"items": items,
|
"items": items,
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def infer_meeting_context_path(
|
||||||
|
loaded_files: list[tuple[Path, dict[str, Any]]],
|
||||||
|
) -> Path | None:
|
||||||
|
paths: set[str] = set()
|
||||||
|
for _path, data in loaded_files:
|
||||||
|
context = data.get("context")
|
||||||
|
if not isinstance(context, dict):
|
||||||
|
continue
|
||||||
|
source_file = clean_text(context.get("source_file"))
|
||||||
|
if source_file:
|
||||||
|
paths.add(source_file)
|
||||||
|
if len(paths) != 1:
|
||||||
|
return None
|
||||||
|
return Path(next(iter(paths)))
|
||||||
|
|
||||||
|
|
||||||
def write_canonicalized(output: dict[str, Any], output_path: Path) -> Path:
|
def write_canonicalized(output: dict[str, Any], output_path: Path) -> Path:
|
||||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
output_path.write_text(
|
output_path.write_text(
|
||||||
@@ -435,6 +668,7 @@ def main() -> int:
|
|||||||
output = canonicalize_extractions(
|
output = canonicalize_extractions(
|
||||||
args.input_dir,
|
args.input_dir,
|
||||||
merge_duplicates=not args.no_merge_exact_duplicates,
|
merge_duplicates=not args.no_merge_exact_duplicates,
|
||||||
|
meeting_context_path=args.meeting_context,
|
||||||
)
|
)
|
||||||
output_path = write_canonicalized(output, args.output)
|
output_path = write_canonicalized(output, args.output)
|
||||||
except (OSError, UnicodeError, ValueError) as exc:
|
except (OSError, UnicodeError, ValueError) as exc:
|
||||||
|
|||||||
@@ -4,6 +4,7 @@
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import argparse
|
import argparse
|
||||||
|
import copy
|
||||||
import json
|
import json
|
||||||
import sys
|
import sys
|
||||||
import time
|
import time
|
||||||
@@ -25,8 +26,13 @@ except ModuleNotFoundError: # pragma: no cover - used by repository-root tests.
|
|||||||
DEFAULT_MODEL = "qwen3.5:9b"
|
DEFAULT_MODEL = "qwen3.5:9b"
|
||||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||||
DEFAULT_NUM_CTX = 32768
|
DEFAULT_NUM_CTX = 32768
|
||||||
DEFAULT_NUM_PREDICT = 4096
|
DEFAULT_MIN_NUM_PREDICT = 4096
|
||||||
|
DEFAULT_NUM_PREDICT = DEFAULT_MIN_NUM_PREDICT
|
||||||
DEFAULT_PROGRESS_INTERVAL = 30
|
DEFAULT_PROGRESS_INTERVAL = 30
|
||||||
|
OUTPUT_CONTEXT_RESERVE_TOKENS = 1024
|
||||||
|
OUTPUT_TOKEN_ESTIMATE_CHARS = 4
|
||||||
|
OUTPUT_GROUP_OVERHEAD_CHARS = 320
|
||||||
|
OUTPUT_SAFETY_MARGIN = 1.35
|
||||||
PROMPT_NAME = "consolidate_facts.md"
|
PROMPT_NAME = "consolidate_facts.md"
|
||||||
|
|
||||||
|
|
||||||
@@ -75,10 +81,11 @@ def parse_args() -> argparse.Namespace:
|
|||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--num-predict",
|
"--num-predict",
|
||||||
type=int,
|
type=int,
|
||||||
default=DEFAULT_NUM_PREDICT,
|
default=None,
|
||||||
help=(
|
help=(
|
||||||
"Maximum generated tokens. The default is bounded for the expected "
|
"Maximum generated tokens. By default this is estimated from the "
|
||||||
f"fact-group JSON while leaving truncation headroom (default: {DEFAULT_NUM_PREDICT})."
|
"fact payload size and bounded by the context window. Explicit "
|
||||||
|
"values preserve the previous fixed-budget behavior."
|
||||||
),
|
),
|
||||||
)
|
)
|
||||||
thinking = parser.add_mutually_exclusive_group()
|
thinking = parser.add_mutually_exclusive_group()
|
||||||
@@ -158,6 +165,49 @@ def build_consolidation_prompt(facts: list[dict[str, Any]]) -> str:
|
|||||||
return f"{task_prompt}\n\nFACT ITEMS:\n{payload}\n"
|
return f"{task_prompt}\n\nFACT ITEMS:\n{payload}\n"
|
||||||
|
|
||||||
|
|
||||||
|
def estimate_response_tokens(facts: list[dict[str, Any]]) -> int:
|
||||||
|
"""
|
||||||
|
Estimate the token budget needed for the model's grouping JSON.
|
||||||
|
|
||||||
|
Semantic Consolidator V0 asks the model to return one group per source fact
|
||||||
|
unless it finds a conservative duplicate. The response therefore scales with
|
||||||
|
the number and text size of fact items. The estimate intentionally includes
|
||||||
|
per-group JSON overhead and a safety margin; strict validation still decides
|
||||||
|
whether the actual response is usable.
|
||||||
|
"""
|
||||||
|
text_chars = 0
|
||||||
|
for item in facts:
|
||||||
|
text_chars += len(str(item.get("text", "")))
|
||||||
|
text_chars += len(str(item.get("evidence", "")))
|
||||||
|
|
||||||
|
estimated_chars = int(
|
||||||
|
(text_chars + len(facts) * OUTPUT_GROUP_OVERHEAD_CHARS)
|
||||||
|
* OUTPUT_SAFETY_MARGIN
|
||||||
|
)
|
||||||
|
return max(
|
||||||
|
DEFAULT_MIN_NUM_PREDICT,
|
||||||
|
(estimated_chars + OUTPUT_TOKEN_ESTIMATE_CHARS - 1)
|
||||||
|
// OUTPUT_TOKEN_ESTIMATE_CHARS,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_num_predict(
|
||||||
|
requested_num_predict: int | None,
|
||||||
|
facts: list[dict[str, Any]],
|
||||||
|
prompt_token_estimate: int,
|
||||||
|
num_ctx: int,
|
||||||
|
) -> int:
|
||||||
|
if requested_num_predict is not None:
|
||||||
|
return requested_num_predict
|
||||||
|
|
||||||
|
estimated = estimate_response_tokens(facts)
|
||||||
|
max_available = max(
|
||||||
|
DEFAULT_MIN_NUM_PREDICT,
|
||||||
|
num_ctx - prompt_token_estimate - OUTPUT_CONTEXT_RESERVE_TOKENS,
|
||||||
|
)
|
||||||
|
return min(estimated, max_available)
|
||||||
|
|
||||||
|
|
||||||
def response_text_from_ollama_data(data: dict[str, Any]) -> str | None:
|
def response_text_from_ollama_data(data: dict[str, Any]) -> str | None:
|
||||||
text = data.get("response")
|
text = data.get("response")
|
||||||
if isinstance(text, str) and text.strip():
|
if isinstance(text, str) and text.strip():
|
||||||
@@ -370,6 +420,96 @@ def validate_group_shapes(groups: list[dict[str, Any]]) -> None:
|
|||||||
raise ConsolidationValidationError("Merged groups need at least two IDs.")
|
raise ConsolidationValidationError("Merged groups need at least two IDs.")
|
||||||
|
|
||||||
|
|
||||||
|
def repair_model_group_coverage(
|
||||||
|
model_output: dict[str, Any],
|
||||||
|
facts: list[dict[str, Any]],
|
||||||
|
) -> tuple[dict[str, Any], list[dict[str, Any]]]:
|
||||||
|
"""
|
||||||
|
Apply deterministic source-coverage repairs to model grouping JSON.
|
||||||
|
|
||||||
|
The repair is intentionally conservative. Repeated source IDs are removed
|
||||||
|
after their first occurrence, empty groups created by that removal are
|
||||||
|
dropped, and missing facts are restored as singleton groups using the
|
||||||
|
original canonicalized fact text. No existing group text, merge reason or
|
||||||
|
semantic merge is rewritten.
|
||||||
|
"""
|
||||||
|
groups = model_output.get("groups")
|
||||||
|
if not isinstance(groups, list):
|
||||||
|
return model_output, []
|
||||||
|
|
||||||
|
repaired = copy.deepcopy(model_output)
|
||||||
|
repaired_groups = repaired["groups"]
|
||||||
|
facts_by_id = {str(item.get("item_id")): item for item in facts}
|
||||||
|
expected_ids = set(facts_by_id)
|
||||||
|
seen: set[str] = set()
|
||||||
|
changes: list[dict[str, Any]] = []
|
||||||
|
|
||||||
|
for group_index, group in enumerate(repaired_groups):
|
||||||
|
if not isinstance(group, dict):
|
||||||
|
continue
|
||||||
|
source_ids = group.get("source_item_ids")
|
||||||
|
if not isinstance(source_ids, list):
|
||||||
|
continue
|
||||||
|
kept_ids: list[str] = []
|
||||||
|
for id_index, item_id in enumerate(source_ids):
|
||||||
|
if not isinstance(item_id, str) or item_id not in expected_ids:
|
||||||
|
kept_ids.append(item_id)
|
||||||
|
continue
|
||||||
|
if item_id in seen:
|
||||||
|
changes.append(
|
||||||
|
{
|
||||||
|
"operation": "remove_duplicate_source_id",
|
||||||
|
"id": item_id,
|
||||||
|
"group_index": group_index,
|
||||||
|
"id_index": id_index,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
seen.add(item_id)
|
||||||
|
kept_ids.append(item_id)
|
||||||
|
group["source_item_ids"] = kept_ids
|
||||||
|
|
||||||
|
non_empty_groups: list[dict[str, Any]] = []
|
||||||
|
for group_index, group in enumerate(repaired_groups):
|
||||||
|
if (
|
||||||
|
isinstance(group, dict)
|
||||||
|
and isinstance(group.get("source_item_ids"), list)
|
||||||
|
and len(group["source_item_ids"]) == 0
|
||||||
|
):
|
||||||
|
changes.append(
|
||||||
|
{
|
||||||
|
"operation": "remove_empty_group",
|
||||||
|
"group_index": group_index,
|
||||||
|
"canonical_text": group.get("canonical_text"),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
non_empty_groups.append(group)
|
||||||
|
repaired["groups"] = non_empty_groups
|
||||||
|
|
||||||
|
missing_ids = sorted(expected_ids - seen)
|
||||||
|
for item_id in missing_ids:
|
||||||
|
fact = facts_by_id[item_id]
|
||||||
|
repaired["groups"].append(
|
||||||
|
{
|
||||||
|
"canonical_text": str(fact.get("text", "")).strip(),
|
||||||
|
"source_item_ids": [item_id],
|
||||||
|
"merge_reason": (
|
||||||
|
"Deterministic coverage repair: source fact was missing "
|
||||||
|
"from the model grouping and is preserved as a singleton."
|
||||||
|
),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
changes.append(
|
||||||
|
{
|
||||||
|
"operation": "restore_missing_source_id_as_singleton",
|
||||||
|
"id": item_id,
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
return repaired, changes
|
||||||
|
|
||||||
|
|
||||||
def build_consolidated_fact_item(
|
def build_consolidated_fact_item(
|
||||||
group: dict[str, Any],
|
group: dict[str, Any],
|
||||||
fact_by_id: dict[str, dict[str, Any]],
|
fact_by_id: dict[str, dict[str, Any]],
|
||||||
@@ -481,9 +621,11 @@ def write_report(
|
|||||||
runtime: float,
|
runtime: float,
|
||||||
prompt_chars: int,
|
prompt_chars: int,
|
||||||
prompt_token_estimate: int,
|
prompt_token_estimate: int,
|
||||||
|
num_predict: int,
|
||||||
fact_count: int,
|
fact_count: int,
|
||||||
groups: list[dict[str, Any]],
|
groups: list[dict[str, Any]],
|
||||||
output_path: Path,
|
output_path: Path,
|
||||||
|
repair_changes: list[dict[str, Any]] | None = None,
|
||||||
) -> None:
|
) -> None:
|
||||||
merged = [group for group in groups if len(group["source_item_ids"]) > 1]
|
merged = [group for group in groups if len(group["source_item_ids"]) > 1]
|
||||||
singletons = [group for group in groups if len(group["source_item_ids"]) == 1]
|
singletons = [group for group in groups if len(group["source_item_ids"]) == 1]
|
||||||
@@ -497,9 +639,11 @@ def write_report(
|
|||||||
f"- Fact item count: {fact_count}",
|
f"- Fact item count: {fact_count}",
|
||||||
f"- Prompt characters: {prompt_chars}",
|
f"- Prompt characters: {prompt_chars}",
|
||||||
f"- Estimated prompt tokens: {prompt_token_estimate}",
|
f"- Estimated prompt tokens: {prompt_token_estimate}",
|
||||||
|
f"- num_predict: {num_predict}",
|
||||||
f"- Merged fact groups: {len(merged)}",
|
f"- Merged fact groups: {len(merged)}",
|
||||||
f"- Source facts involved in merges: {sum(len(group['source_item_ids']) for group in merged)}",
|
f"- Source facts involved in merges: {sum(len(group['source_item_ids']) for group in merged)}",
|
||||||
f"- Singleton fact groups: {len(singletons)}",
|
f"- Singleton fact groups: {len(singletons)}",
|
||||||
|
f"- Deterministic repair changes: {len(repair_changes or [])}",
|
||||||
f"- Output path: `{output_path}`",
|
f"- Output path: `{output_path}`",
|
||||||
"",
|
"",
|
||||||
"## Actual Merges",
|
"## Actual Merges",
|
||||||
@@ -518,6 +662,10 @@ def write_report(
|
|||||||
"",
|
"",
|
||||||
]
|
]
|
||||||
)
|
)
|
||||||
|
if repair_changes:
|
||||||
|
lines.extend(["", "## Deterministic Coverage Repairs", ""])
|
||||||
|
for change in repair_changes:
|
||||||
|
lines.append(f"- `{change['operation']}`: {json.dumps(change, ensure_ascii=False, sort_keys=True)}")
|
||||||
path.write_text("\n".join(lines).rstrip() + "\n", encoding="utf-8")
|
path.write_text("\n".join(lines).rstrip() + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
@@ -527,6 +675,7 @@ def main() -> int:
|
|||||||
raw_response_path = args.output_dir / "raw_model_response.txt"
|
raw_response_path = args.output_dir / "raw_model_response.txt"
|
||||||
output_path = args.output_dir / "consolidated_extractions.json"
|
output_path = args.output_dir / "consolidated_extractions.json"
|
||||||
report_path = args.output_dir / "report.md"
|
report_path = args.output_dir / "report.md"
|
||||||
|
repair_metadata_path = args.output_dir / "repair_metadata.json"
|
||||||
|
|
||||||
try:
|
try:
|
||||||
canonicalized = load_json_object(args.canonicalized_input)
|
canonicalized = load_json_object(args.canonicalized_input)
|
||||||
@@ -534,9 +683,16 @@ def main() -> int:
|
|||||||
prompt = build_consolidation_prompt(facts)
|
prompt = build_consolidation_prompt(facts)
|
||||||
prompt_chars = len(prompt)
|
prompt_chars = len(prompt)
|
||||||
prompt_token_estimate = (prompt_chars + 3) // 4
|
prompt_token_estimate = (prompt_chars + 3) // 4
|
||||||
|
num_predict = resolve_num_predict(
|
||||||
|
requested_num_predict=args.num_predict,
|
||||||
|
facts=facts,
|
||||||
|
prompt_token_estimate=prompt_token_estimate,
|
||||||
|
num_ctx=args.num_ctx,
|
||||||
|
)
|
||||||
print(f"Fact item count: {len(facts)}")
|
print(f"Fact item count: {len(facts)}")
|
||||||
print(f"Estimated prompt size chars: {prompt_chars}")
|
print(f"Estimated prompt size chars: {prompt_chars}")
|
||||||
print(f"Estimated prompt tokens: {prompt_token_estimate}")
|
print(f"Estimated prompt tokens: {prompt_token_estimate}")
|
||||||
|
print(f"Resolved num_predict: {num_predict}")
|
||||||
print("Expected LLM call count: 1")
|
print("Expected LLM call count: 1")
|
||||||
print("Expected runtime: 5-10 minutes on current local benchmark basis")
|
print("Expected runtime: 5-10 minutes on current local benchmark basis")
|
||||||
|
|
||||||
@@ -546,12 +702,23 @@ def main() -> int:
|
|||||||
prompt=prompt,
|
prompt=prompt,
|
||||||
timeout=args.timeout,
|
timeout=args.timeout,
|
||||||
num_ctx=args.num_ctx,
|
num_ctx=args.num_ctx,
|
||||||
num_predict=args.num_predict,
|
num_predict=num_predict,
|
||||||
think=args.think,
|
think=args.think,
|
||||||
progress_interval=args.progress_interval,
|
progress_interval=args.progress_interval,
|
||||||
)
|
)
|
||||||
raw_response_path.write_text(raw_text + "\n", encoding="utf-8")
|
raw_response_path.write_text(raw_text + "\n", encoding="utf-8")
|
||||||
model_output = parse_model_json(raw_text)
|
model_output = parse_model_json(raw_text)
|
||||||
|
model_output, repair_changes = repair_model_group_coverage(model_output, facts)
|
||||||
|
if repair_changes:
|
||||||
|
write_json(
|
||||||
|
repair_metadata_path,
|
||||||
|
{
|
||||||
|
"scope": "semantic_consolidator_v0_source_coverage",
|
||||||
|
"llm_used": False,
|
||||||
|
"repair_count": len(repair_changes),
|
||||||
|
"repairs": repair_changes,
|
||||||
|
},
|
||||||
|
)
|
||||||
expected_fact_ids = {item["item_id"] for item in facts}
|
expected_fact_ids = {item["item_id"] for item in facts}
|
||||||
groups = validate_model_groups(model_output, expected_fact_ids)
|
groups = validate_model_groups(model_output, expected_fact_ids)
|
||||||
validate_group_shapes(groups)
|
validate_group_shapes(groups)
|
||||||
@@ -564,9 +731,11 @@ def main() -> int:
|
|||||||
runtime=runtime,
|
runtime=runtime,
|
||||||
prompt_chars=prompt_chars,
|
prompt_chars=prompt_chars,
|
||||||
prompt_token_estimate=prompt_token_estimate,
|
prompt_token_estimate=prompt_token_estimate,
|
||||||
|
num_predict=num_predict,
|
||||||
fact_count=len(facts),
|
fact_count=len(facts),
|
||||||
groups=groups,
|
groups=groups,
|
||||||
output_path=output_path,
|
output_path=output_path,
|
||||||
|
repair_changes=repair_changes,
|
||||||
)
|
)
|
||||||
except requests.ConnectionError as exc:
|
except requests.ConnectionError as exc:
|
||||||
print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr)
|
print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr)
|
||||||
|
|||||||
@@ -0,0 +1,519 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""LLM-backed Working Protocol V2 renderer with deterministic contract checks."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
import time
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
import requests
|
||||||
|
|
||||||
|
try:
|
||||||
|
from meeting_lab.llm.prompts import load_prompt
|
||||||
|
except ModuleNotFoundError: # pragma: no cover - used by repository-root tests.
|
||||||
|
from src.meeting_lab.llm.prompts import load_prompt
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_MODEL = "qwen3.5:9b"
|
||||||
|
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||||
|
DEFAULT_NUM_CTX = 32768
|
||||||
|
DEFAULT_NUM_PREDICT = 4096
|
||||||
|
DEFAULT_TIMEOUT = 1800
|
||||||
|
PROMPT_NAME = "working_protocol.md"
|
||||||
|
REQUIRED_TITLE = "# Working Protocol"
|
||||||
|
ALLOWED_TOPIC_SECTIONS = {
|
||||||
|
"Background",
|
||||||
|
"Decisions",
|
||||||
|
"Action Items",
|
||||||
|
"Open Questions",
|
||||||
|
}
|
||||||
|
GENERIC_SUMMARY_HEADINGS = {
|
||||||
|
"Entscheidungen",
|
||||||
|
"Decisions",
|
||||||
|
"Handlungsaufträge",
|
||||||
|
"Handlungsauftraege",
|
||||||
|
"Action Items",
|
||||||
|
"Offene Fragen",
|
||||||
|
"Open Questions",
|
||||||
|
"Technische Details",
|
||||||
|
"Technical Details",
|
||||||
|
"Fakten",
|
||||||
|
"Facts",
|
||||||
|
"Zusammenfassung",
|
||||||
|
"Summary",
|
||||||
|
"Konsolidierter Projektstatusbericht",
|
||||||
|
}
|
||||||
|
HEADING_RE = re.compile(r"^(#{1,6})\s+(.+?)\s*$")
|
||||||
|
|
||||||
|
|
||||||
|
class WorkingProtocolValidationError(ValueError):
|
||||||
|
"""Raised when rendered Markdown violates the Working Protocol V2 contract."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class WorkingProtocolValidationReport:
|
||||||
|
valid: bool
|
||||||
|
violations: list[dict[str, Any]]
|
||||||
|
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return {"valid": self.valid, "violations": self.violations}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser(
|
||||||
|
description="Render a consolidated meeting representation as Working Protocol V2."
|
||||||
|
)
|
||||||
|
parser.add_argument("input", type=Path, help="Consolidated meeting JSON input.")
|
||||||
|
parser.add_argument(
|
||||||
|
"-o",
|
||||||
|
"--output-dir",
|
||||||
|
type=Path,
|
||||||
|
required=True,
|
||||||
|
help="Directory for working_protocol.md, raw response, validation and metadata.",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--model",
|
||||||
|
default=DEFAULT_MODEL,
|
||||||
|
help=f"Ollama model name (default: {DEFAULT_MODEL}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--endpoint",
|
||||||
|
default=DEFAULT_ENDPOINT,
|
||||||
|
help=f"Ollama generate endpoint (default: {DEFAULT_ENDPOINT}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--timeout",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_TIMEOUT,
|
||||||
|
help=f"HTTP timeout in seconds (default: {DEFAULT_TIMEOUT}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--num-ctx",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_NUM_CTX,
|
||||||
|
help=f"Context window tokens (default: {DEFAULT_NUM_CTX}).",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--num-predict",
|
||||||
|
type=int,
|
||||||
|
default=DEFAULT_NUM_PREDICT,
|
||||||
|
help=f"Maximum generated tokens (default: {DEFAULT_NUM_PREDICT}).",
|
||||||
|
)
|
||||||
|
thinking = parser.add_mutually_exclusive_group()
|
||||||
|
thinking.add_argument("--think", dest="think", action="store_true")
|
||||||
|
thinking.add_argument("--no-think", dest="think", action="store_false")
|
||||||
|
parser.set_defaults(think=False)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def build_renderer_prompt(input_text: str) -> str:
|
||||||
|
return f"{load_prompt(PROMPT_NAME)}\n\nINPUT JSON:\n{input_text}\n"
|
||||||
|
|
||||||
|
|
||||||
|
def response_text_from_ollama_data(data: dict[str, Any]) -> str:
|
||||||
|
text = data.get("response")
|
||||||
|
if isinstance(text, str):
|
||||||
|
return text
|
||||||
|
message = data.get("message")
|
||||||
|
if isinstance(message, dict) and isinstance(message.get("content"), str):
|
||||||
|
return message["content"]
|
||||||
|
return ""
|
||||||
|
|
||||||
|
|
||||||
|
def build_ollama_payload(
|
||||||
|
model: str,
|
||||||
|
prompt: str,
|
||||||
|
num_ctx: int,
|
||||||
|
num_predict: int,
|
||||||
|
think: bool,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
return {
|
||||||
|
"model": model,
|
||||||
|
"prompt": prompt,
|
||||||
|
"think": think,
|
||||||
|
"stream": False,
|
||||||
|
"options": {
|
||||||
|
"temperature": 0.0,
|
||||||
|
"num_ctx": num_ctx,
|
||||||
|
"num_predict": num_predict,
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def call_ollama(
|
||||||
|
endpoint: str,
|
||||||
|
payload: dict[str, Any],
|
||||||
|
timeout: int,
|
||||||
|
) -> tuple[dict[str, Any], float]:
|
||||||
|
start = time.perf_counter()
|
||||||
|
response = requests.post(endpoint, json=payload, timeout=timeout)
|
||||||
|
runtime = time.perf_counter() - start
|
||||||
|
response.raise_for_status()
|
||||||
|
data = response.json()
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise ValueError("Ollama response must be a JSON object.")
|
||||||
|
return data, runtime
|
||||||
|
|
||||||
|
|
||||||
|
def normalize_heading_whitespace(line: str) -> str:
|
||||||
|
match = HEADING_RE.match(line.strip())
|
||||||
|
if not match:
|
||||||
|
return line.rstrip()
|
||||||
|
return f"{match.group(1)} {match.group(2).strip()}"
|
||||||
|
|
||||||
|
|
||||||
|
def clean_working_protocol_markdown(text: str) -> str:
|
||||||
|
"""
|
||||||
|
Remove only non-semantic wrapper text before an existing protocol heading.
|
||||||
|
|
||||||
|
The cleanup does not create headings or transform a categorized summary into
|
||||||
|
a Working Protocol. It only makes an already present contract heading the
|
||||||
|
first byte of the candidate document.
|
||||||
|
"""
|
||||||
|
text = text.replace("\r\n", "\n").replace("\r", "\n")
|
||||||
|
lines = text.split("\n")
|
||||||
|
start_index: int | None = None
|
||||||
|
for index, line in enumerate(lines):
|
||||||
|
if normalize_heading_whitespace(line) == REQUIRED_TITLE:
|
||||||
|
start_index = index
|
||||||
|
break
|
||||||
|
if start_index is None:
|
||||||
|
return text.lstrip()
|
||||||
|
|
||||||
|
cleaned_lines = lines[start_index:]
|
||||||
|
if cleaned_lines:
|
||||||
|
cleaned_lines[0] = REQUIRED_TITLE
|
||||||
|
return "\n".join(normalize_heading_whitespace(line) for line in cleaned_lines).strip() + "\n"
|
||||||
|
|
||||||
|
|
||||||
|
def validate_markdown_shape(markdown: str) -> list[dict[str, Any]]:
|
||||||
|
violations: list[dict[str, Any]] = []
|
||||||
|
if markdown.count("```") % 2 != 0:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "malformed_markdown",
|
||||||
|
"reason": "unclosed_fenced_code_block",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
seen_title = False
|
||||||
|
seen_topic = False
|
||||||
|
current_topic_has_section = False
|
||||||
|
topic_count = 0
|
||||||
|
section_count = 0
|
||||||
|
|
||||||
|
for line_number, line in enumerate(markdown.splitlines(), start=1):
|
||||||
|
match = HEADING_RE.match(line)
|
||||||
|
if not match:
|
||||||
|
continue
|
||||||
|
level = len(match.group(1))
|
||||||
|
title = match.group(2).strip().strip("*")
|
||||||
|
|
||||||
|
if level == 1:
|
||||||
|
if line_number != 1 or title != "Working Protocol":
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "invalid_heading",
|
||||||
|
"line": line_number,
|
||||||
|
"heading": line,
|
||||||
|
"reason": "only the first line may be '# Working Protocol'",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
seen_title = True
|
||||||
|
continue
|
||||||
|
|
||||||
|
if not seen_title:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "invalid_heading_order",
|
||||||
|
"line": line_number,
|
||||||
|
"heading": line,
|
||||||
|
"reason": "heading appears before required title",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
|
||||||
|
if level == 2:
|
||||||
|
topic_count += 1
|
||||||
|
seen_topic = True
|
||||||
|
current_topic_has_section = False
|
||||||
|
if title in GENERIC_SUMMARY_HEADINGS:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "generic_summary_framing",
|
||||||
|
"line": line_number,
|
||||||
|
"heading": line,
|
||||||
|
"reason": "top-level category heading is not a topic",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
continue
|
||||||
|
|
||||||
|
if level == 3:
|
||||||
|
if not seen_topic:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "missing_topic",
|
||||||
|
"line": line_number,
|
||||||
|
"heading": line,
|
||||||
|
"reason": "section appears before any topic heading",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if title not in ALLOWED_TOPIC_SECTIONS:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "unknown_topic_section",
|
||||||
|
"line": line_number,
|
||||||
|
"heading": line,
|
||||||
|
"allowed": sorted(ALLOWED_TOPIC_SECTIONS),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
section_count += 1
|
||||||
|
current_topic_has_section = True
|
||||||
|
continue
|
||||||
|
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "unsupported_heading_level",
|
||||||
|
"line": line_number,
|
||||||
|
"heading": line,
|
||||||
|
"reason": "Working Protocol V2 uses h1 title, h2 topics and h3 sections only",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
if seen_topic and not current_topic_has_section:
|
||||||
|
# This catches the last topic; earlier empty topics are caught below by
|
||||||
|
# counting consecutive h2 headings.
|
||||||
|
pass
|
||||||
|
|
||||||
|
h2_without_section = _topic_headings_without_sections(markdown)
|
||||||
|
violations.extend(h2_without_section)
|
||||||
|
|
||||||
|
if topic_count == 0:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "missing_topic",
|
||||||
|
"reason": "document must contain at least one '## Topic title' section",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
if section_count == 0:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "missing_topic_sections",
|
||||||
|
"reason": "document must contain at least one allowed h3 topic section",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return violations
|
||||||
|
|
||||||
|
|
||||||
|
def _topic_headings_without_sections(markdown: str) -> list[dict[str, Any]]:
|
||||||
|
violations: list[dict[str, Any]] = []
|
||||||
|
current_topic: tuple[int, str] | None = None
|
||||||
|
current_has_section = False
|
||||||
|
|
||||||
|
for line_number, line in enumerate(markdown.splitlines(), start=1):
|
||||||
|
match = HEADING_RE.match(line)
|
||||||
|
if not match:
|
||||||
|
continue
|
||||||
|
level = len(match.group(1))
|
||||||
|
if level == 2:
|
||||||
|
if current_topic is not None and not current_has_section:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "empty_topic",
|
||||||
|
"line": current_topic[0],
|
||||||
|
"heading": current_topic[1],
|
||||||
|
"reason": "topic has no allowed h3 section",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
current_topic = (line_number, line)
|
||||||
|
current_has_section = False
|
||||||
|
elif level == 3 and current_topic is not None:
|
||||||
|
title = match.group(2).strip().strip("*")
|
||||||
|
if title in ALLOWED_TOPIC_SECTIONS:
|
||||||
|
current_has_section = True
|
||||||
|
|
||||||
|
if current_topic is not None and not current_has_section:
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "empty_topic",
|
||||||
|
"line": current_topic[0],
|
||||||
|
"heading": current_topic[1],
|
||||||
|
"reason": "topic has no allowed h3 section",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
return violations
|
||||||
|
|
||||||
|
|
||||||
|
def validate_working_protocol_markdown(markdown: str) -> WorkingProtocolValidationReport:
|
||||||
|
violations: list[dict[str, Any]] = []
|
||||||
|
if not markdown.startswith(REQUIRED_TITLE):
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "missing_required_heading",
|
||||||
|
"expected": REQUIRED_TITLE,
|
||||||
|
"reason": "document must begin exactly with '# Working Protocol'",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
elif not markdown.startswith(REQUIRED_TITLE + "\n"):
|
||||||
|
violations.append(
|
||||||
|
{
|
||||||
|
"type": "invalid_required_heading",
|
||||||
|
"expected": REQUIRED_TITLE,
|
||||||
|
"reason": "required heading must occupy the complete first line",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
if markdown.startswith(REQUIRED_TITLE):
|
||||||
|
violations.extend(validate_markdown_shape(markdown))
|
||||||
|
|
||||||
|
return WorkingProtocolValidationReport(
|
||||||
|
valid=len(violations) == 0,
|
||||||
|
violations=violations,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def write_json(path: Path, data: Any) -> None:
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
path.write_text(json.dumps(data, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def render_working_protocol(
|
||||||
|
input_path: Path,
|
||||||
|
output_dir: Path,
|
||||||
|
model: str = DEFAULT_MODEL,
|
||||||
|
endpoint: str = DEFAULT_ENDPOINT,
|
||||||
|
timeout: int = DEFAULT_TIMEOUT,
|
||||||
|
num_ctx: int = DEFAULT_NUM_CTX,
|
||||||
|
num_predict: int = DEFAULT_NUM_PREDICT,
|
||||||
|
think: bool = False,
|
||||||
|
) -> dict[str, Any]:
|
||||||
|
output_dir.mkdir(parents=True, exist_ok=True)
|
||||||
|
prompt_path = Path("prompts") / PROMPT_NAME
|
||||||
|
raw_response_path = output_dir / "raw_model_response.json"
|
||||||
|
raw_text_path = output_dir / "raw_model_response.txt"
|
||||||
|
candidate_path = output_dir / "cleaned_candidate.md"
|
||||||
|
validation_path = output_dir / "validation_report.json"
|
||||||
|
protocol_path = output_dir / "working_protocol.md"
|
||||||
|
metadata_path = output_dir / "metadata.json"
|
||||||
|
report_path = output_dir / "report.md"
|
||||||
|
|
||||||
|
input_text = input_path.read_text(encoding="utf-8-sig")
|
||||||
|
prompt = build_renderer_prompt(input_text)
|
||||||
|
payload = build_ollama_payload(model, prompt, num_ctx, num_predict, think)
|
||||||
|
|
||||||
|
data, runtime = call_ollama(endpoint, payload, timeout)
|
||||||
|
response_text = response_text_from_ollama_data(data)
|
||||||
|
raw_response_path.write_text(
|
||||||
|
json.dumps(data, ensure_ascii=False, indent=2) + "\n",
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
raw_text_path.write_text(response_text, encoding="utf-8")
|
||||||
|
|
||||||
|
cleaned = clean_working_protocol_markdown(response_text)
|
||||||
|
candidate_path.write_text(cleaned, encoding="utf-8")
|
||||||
|
validation = validate_working_protocol_markdown(cleaned)
|
||||||
|
write_json(validation_path, validation.to_dict())
|
||||||
|
|
||||||
|
if validation.valid:
|
||||||
|
protocol_path.write_text(cleaned, encoding="utf-8")
|
||||||
|
elif protocol_path.exists():
|
||||||
|
protocol_path.unlink()
|
||||||
|
|
||||||
|
metadata = {
|
||||||
|
"model": model,
|
||||||
|
"endpoint": endpoint,
|
||||||
|
"think": think,
|
||||||
|
"stream": False,
|
||||||
|
"temperature": 0.0,
|
||||||
|
"num_ctx": num_ctx,
|
||||||
|
"num_predict": num_predict,
|
||||||
|
"timeout_seconds": timeout,
|
||||||
|
"request_count": 1,
|
||||||
|
"runtime_seconds": runtime,
|
||||||
|
"prompt_path": str(prompt_path.resolve()),
|
||||||
|
"input_path": str(input_path.resolve()),
|
||||||
|
"output_path": str(protocol_path.resolve()) if validation.valid else None,
|
||||||
|
"raw_response_path": str(raw_response_path.resolve()),
|
||||||
|
"raw_text_path": str(raw_text_path.resolve()),
|
||||||
|
"cleaned_candidate_path": str(candidate_path.resolve()),
|
||||||
|
"validation_report_path": str(validation_path.resolve()),
|
||||||
|
"http_status_code": 200,
|
||||||
|
"response_text_length": len(response_text),
|
||||||
|
"candidate_text_length": len(cleaned),
|
||||||
|
"valid": validation.valid,
|
||||||
|
"readable_markdown": validation.valid,
|
||||||
|
"top_level_json_keys": sorted(data.keys()),
|
||||||
|
"done": data.get("done"),
|
||||||
|
"done_reason": data.get("done_reason"),
|
||||||
|
"total_duration": data.get("total_duration"),
|
||||||
|
"load_duration": data.get("load_duration"),
|
||||||
|
"prompt_eval_count": data.get("prompt_eval_count"),
|
||||||
|
"prompt_eval_duration": data.get("prompt_eval_duration"),
|
||||||
|
"eval_count": data.get("eval_count"),
|
||||||
|
"eval_duration": data.get("eval_duration"),
|
||||||
|
"created_at": datetime.now().isoformat(timespec="seconds"),
|
||||||
|
}
|
||||||
|
write_json(metadata_path, metadata)
|
||||||
|
|
||||||
|
lines = [
|
||||||
|
"# Working Protocol Renderer V2 Report",
|
||||||
|
"",
|
||||||
|
f"- Result: {'valid renderer run' if validation.valid else 'invalid renderer run'}",
|
||||||
|
f"- Model: `{model}`",
|
||||||
|
f"- Runtime: {runtime:.3f} seconds",
|
||||||
|
f"- Request count: 1",
|
||||||
|
f"- Valid: {validation.valid}",
|
||||||
|
f"- Violations: {len(validation.violations)}",
|
||||||
|
f"- Raw response path: `{raw_response_path}`",
|
||||||
|
f"- Cleaned candidate path: `{candidate_path}`",
|
||||||
|
f"- Validation report path: `{validation_path}`",
|
||||||
|
f"- Output path: `{protocol_path if validation.valid else 'not written'}`",
|
||||||
|
]
|
||||||
|
report_path.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||||
|
return metadata
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
args = parse_args()
|
||||||
|
try:
|
||||||
|
metadata = render_working_protocol(
|
||||||
|
input_path=args.input,
|
||||||
|
output_dir=args.output_dir,
|
||||||
|
model=args.model,
|
||||||
|
endpoint=args.endpoint,
|
||||||
|
timeout=args.timeout,
|
||||||
|
num_ctx=args.num_ctx,
|
||||||
|
num_predict=args.num_predict,
|
||||||
|
think=args.think,
|
||||||
|
)
|
||||||
|
except requests.ConnectionError as exc:
|
||||||
|
print(f"Error: Ollama is not reachable at {args.endpoint}: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except requests.Timeout as exc:
|
||||||
|
print(f"Error: Ollama request timed out after {args.timeout} seconds: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except requests.HTTPError as exc:
|
||||||
|
print(f"Error: Ollama returned an HTTP error: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
except (OSError, UnicodeError, ValueError, json.JSONDecodeError) as exc:
|
||||||
|
print(f"Error: {exc}", file=sys.stderr)
|
||||||
|
return 1
|
||||||
|
|
||||||
|
print(f"Runtime seconds: {metadata['runtime_seconds']:.3f}")
|
||||||
|
print(f"Validation result: {'passed' if metadata['valid'] else 'failed'}")
|
||||||
|
print(f"Output: {metadata['output_path'] or 'not written'}")
|
||||||
|
print(f"Raw model response: {metadata['raw_response_path']}")
|
||||||
|
print(f"Validation report: {metadata['validation_report_path']}")
|
||||||
|
return 0 if metadata["valid"] else 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -19,6 +19,8 @@ EMPTY_EXTRACTION = {
|
|||||||
"technical": [],
|
"technical": [],
|
||||||
}
|
}
|
||||||
|
|
||||||
|
PROGEO_CONTEXT = Path("samples/real_live/progeo_meeting/meeting_context.yaml")
|
||||||
|
|
||||||
|
|
||||||
def write_extraction(directory: Path, name: str, data: dict) -> Path:
|
def write_extraction(directory: Path, name: str, data: dict) -> Path:
|
||||||
path = directory / name
|
path = directory / name
|
||||||
@@ -163,6 +165,102 @@ class CanonicalizeTests(unittest.TestCase):
|
|||||||
self.assertEqual(output["stats"]["input_item_count"], 0)
|
self.assertEqual(output["stats"]["input_item_count"], 0)
|
||||||
self.assertEqual(output["stats"]["output_item_count"], 0)
|
self.assertEqual(output["stats"]["output_item_count"], 0)
|
||||||
|
|
||||||
|
def test_responsible_aliases_normalize_with_meeting_context(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["todos"] = [
|
||||||
|
{"task": "A", "responsible": "Martin", "deadline": None, "evidence": "e"},
|
||||||
|
{"task": "B", "responsible": "Herr Tazl", "deadline": None, "evidence": "e"},
|
||||||
|
{"task": "C", "responsible": "Marlene", "deadline": None, "evidence": "e"},
|
||||||
|
{"task": "D", "responsible": "Gothard", "deadline": None, "evidence": "e"},
|
||||||
|
{"task": "E", "responsible": "Henning", "deadline": None, "evidence": "e"},
|
||||||
|
]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root, meeting_context_path=PROGEO_CONTEXT)
|
||||||
|
responsibles = [item["responsible"] for item in output["items"]]
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
responsibles,
|
||||||
|
["Martin", "Martin", "Marleen", "Gotthard", "Henning"],
|
||||||
|
)
|
||||||
|
self.assertEqual(output["stats"]["responsibility_rejection_count"], 0)
|
||||||
|
self.assertEqual(output["stats"]["responsibility_normalization_count"], 3)
|
||||||
|
|
||||||
|
def test_invalid_responsible_values_are_cleared_with_validation_report(self) -> None:
|
||||||
|
invalid_values = [
|
||||||
|
"31. August",
|
||||||
|
"15. oder 16. September",
|
||||||
|
"27.8.",
|
||||||
|
"am 31.",
|
||||||
|
"morgen",
|
||||||
|
"nächste Woche",
|
||||||
|
"Frankfurt",
|
||||||
|
"Progeo",
|
||||||
|
"Secugrid",
|
||||||
|
"8:30 Uhr",
|
||||||
|
"Unbekannte Person",
|
||||||
|
]
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["todos"] = [
|
||||||
|
{
|
||||||
|
"task": f"Task {index}",
|
||||||
|
"responsible": value,
|
||||||
|
"deadline": None,
|
||||||
|
"evidence": "e",
|
||||||
|
}
|
||||||
|
for index, value in enumerate(invalid_values, start=1)
|
||||||
|
]
|
||||||
|
data["todos"].append(
|
||||||
|
{
|
||||||
|
"task": "Null task",
|
||||||
|
"responsible": None,
|
||||||
|
"deadline": None,
|
||||||
|
"evidence": "e",
|
||||||
|
}
|
||||||
|
)
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root, meeting_context_path=PROGEO_CONTEXT)
|
||||||
|
|
||||||
|
self.assertTrue(all(item["responsible"] is None for item in output["items"]))
|
||||||
|
self.assertEqual(
|
||||||
|
output["stats"]["responsibility_rejection_count"],
|
||||||
|
len(invalid_values),
|
||||||
|
)
|
||||||
|
rejected = output["responsibility_validations"]
|
||||||
|
self.assertEqual([item["original_value"] for item in rejected], invalid_values)
|
||||||
|
self.assertTrue(all(item["cleared_to_null"] for item in rejected))
|
||||||
|
|
||||||
|
def test_null_responsible_remains_valid_without_warning(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["todos"] = [
|
||||||
|
{"task": "Task", "responsible": None, "deadline": None, "evidence": "e"}
|
||||||
|
]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root, meeting_context_path=PROGEO_CONTEXT)
|
||||||
|
|
||||||
|
self.assertIsNone(output["items"][0]["responsible"])
|
||||||
|
self.assertEqual(output["responsibility_validations"], [])
|
||||||
|
|
||||||
|
def test_responsible_validation_does_not_run_without_meeting_context(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
data = dict(EMPTY_EXTRACTION)
|
||||||
|
data["todos"] = ["Update docs | Mira | Friday | I will update docs"]
|
||||||
|
write_extraction(root, "chunk_01_extraction.json", data)
|
||||||
|
|
||||||
|
output = canonicalize_extractions(root)
|
||||||
|
|
||||||
|
self.assertEqual(output["items"][0]["responsible"], "Mira")
|
||||||
|
self.assertEqual(output["responsibility_validations"], [])
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
if __name__ == "__main__":
|
||||||
unittest.main()
|
unittest.main()
|
||||||
|
|||||||
@@ -1,11 +1,18 @@
|
|||||||
import unittest
|
import unittest
|
||||||
|
import json
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
from src.meeting_lab.consolidation.consolidate_facts import (
|
from src.meeting_lab.consolidation.consolidate_facts import (
|
||||||
ConsolidationValidationError,
|
ConsolidationValidationError,
|
||||||
|
DEFAULT_MIN_NUM_PREDICT,
|
||||||
DEFAULT_NUM_PREDICT,
|
DEFAULT_NUM_PREDICT,
|
||||||
build_consolidated_output,
|
build_consolidated_output,
|
||||||
build_ollama_payload,
|
build_ollama_payload,
|
||||||
|
estimate_response_tokens,
|
||||||
|
fact_items,
|
||||||
parse_model_json,
|
parse_model_json,
|
||||||
|
repair_model_group_coverage,
|
||||||
|
resolve_num_predict,
|
||||||
validate_consolidated_output,
|
validate_consolidated_output,
|
||||||
validate_model_groups,
|
validate_model_groups,
|
||||||
)
|
)
|
||||||
@@ -98,6 +105,54 @@ class ConsolidateFactsTests(unittest.TestCase):
|
|||||||
self.assertEqual(payload["options"]["num_ctx"], 16384)
|
self.assertEqual(payload["options"]["num_ctx"], 16384)
|
||||||
self.assertEqual(payload["options"]["num_predict"], 1024)
|
self.assertEqual(payload["options"]["num_predict"], 1024)
|
||||||
|
|
||||||
|
def test_explicit_num_predict_is_preserved(self):
|
||||||
|
facts = [canonicalized_fixture()["items"][0]]
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
resolve_num_predict(
|
||||||
|
requested_num_predict=1234,
|
||||||
|
facts=facts,
|
||||||
|
prompt_token_estimate=100,
|
||||||
|
num_ctx=32768,
|
||||||
|
),
|
||||||
|
1234,
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_adaptive_num_predict_scales_with_fact_payload(self):
|
||||||
|
fixture = canonicalized_fixture()
|
||||||
|
base_facts = fact_items(fixture)
|
||||||
|
larger_facts = []
|
||||||
|
for index in range(80):
|
||||||
|
item = dict(base_facts[index % len(base_facts)])
|
||||||
|
item["item_id"] = f"fact_{index + 1:04d}"
|
||||||
|
item["text"] = item["text"] + " " + ("detail " * 20)
|
||||||
|
item["evidence"] = item["evidence"] + " " + ("evidence " * 20)
|
||||||
|
larger_facts.append(item)
|
||||||
|
|
||||||
|
resolved = resolve_num_predict(
|
||||||
|
requested_num_predict=None,
|
||||||
|
facts=larger_facts,
|
||||||
|
prompt_token_estimate=9000,
|
||||||
|
num_ctx=32768,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertGreater(resolved, DEFAULT_MIN_NUM_PREDICT)
|
||||||
|
self.assertLessEqual(resolved, 32768 - 9000 - 1024)
|
||||||
|
|
||||||
|
def test_progeo_context_benchmark_needs_more_than_fixed_default_when_available(self):
|
||||||
|
path = Path(
|
||||||
|
"samples/benchmarks/progeo_meeting_context_v1_20260804_110913/"
|
||||||
|
"canonicalizer/canonicalized_extractions.json"
|
||||||
|
)
|
||||||
|
if not path.exists():
|
||||||
|
self.skipTest("Progeo context benchmark artifact is not available.")
|
||||||
|
|
||||||
|
canonicalized = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
facts = fact_items(canonicalized)
|
||||||
|
|
||||||
|
self.assertGreater(len(facts), 60)
|
||||||
|
self.assertGreater(estimate_response_tokens(facts), DEFAULT_NUM_PREDICT)
|
||||||
|
|
||||||
def test_grouping_validation_accepts_complete_singletons(self):
|
def test_grouping_validation_accepts_complete_singletons(self):
|
||||||
groups = validate_model_groups(
|
groups = validate_model_groups(
|
||||||
{
|
{
|
||||||
@@ -154,6 +209,61 @@ class ConsolidateFactsTests(unittest.TestCase):
|
|||||||
{"fact_0001"},
|
{"fact_0001"},
|
||||||
)
|
)
|
||||||
|
|
||||||
|
def test_source_coverage_repair_restores_missing_singletons(self):
|
||||||
|
fixture = canonicalized_fixture()
|
||||||
|
facts = fact_items(fixture)
|
||||||
|
repaired, changes = repair_model_group_coverage(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "The lead maintains the list.",
|
||||||
|
"source_item_ids": ["fact_0001"],
|
||||||
|
"merge_reason": "Singleton.",
|
||||||
|
}
|
||||||
|
]
|
||||||
|
},
|
||||||
|
facts,
|
||||||
|
)
|
||||||
|
|
||||||
|
groups = validate_model_groups(repaired, {"fact_0001", "fact_0002"})
|
||||||
|
|
||||||
|
self.assertEqual(len(groups), 2)
|
||||||
|
self.assertEqual(groups[1]["source_item_ids"], ["fact_0002"])
|
||||||
|
self.assertEqual(
|
||||||
|
changes[0]["operation"],
|
||||||
|
"restore_missing_source_id_as_singleton",
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_source_coverage_repair_removes_duplicate_occurrences(self):
|
||||||
|
fixture = canonicalized_fixture()
|
||||||
|
facts = fact_items(fixture)
|
||||||
|
repaired, changes = repair_model_group_coverage(
|
||||||
|
{
|
||||||
|
"groups": [
|
||||||
|
{
|
||||||
|
"canonical_text": "Merged.",
|
||||||
|
"source_item_ids": ["fact_0001", "fact_0002"],
|
||||||
|
"merge_reason": "Same.",
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"canonical_text": "Duplicate.",
|
||||||
|
"source_item_ids": ["fact_0002"],
|
||||||
|
"merge_reason": "Duplicate.",
|
||||||
|
},
|
||||||
|
]
|
||||||
|
},
|
||||||
|
facts,
|
||||||
|
)
|
||||||
|
|
||||||
|
groups = validate_model_groups(repaired, {"fact_0001", "fact_0002"})
|
||||||
|
|
||||||
|
self.assertEqual(len(groups), 1)
|
||||||
|
self.assertEqual(groups[0]["source_item_ids"], ["fact_0001", "fact_0002"])
|
||||||
|
self.assertEqual(
|
||||||
|
[change["operation"] for change in changes],
|
||||||
|
["remove_duplicate_source_id", "remove_empty_group"],
|
||||||
|
)
|
||||||
|
|
||||||
def test_merged_group_validation(self):
|
def test_merged_group_validation(self):
|
||||||
groups = validate_model_groups(
|
groups = validate_model_groups(
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -0,0 +1,179 @@
|
|||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from src.meeting_lab.protocol.render_working_protocol import (
|
||||||
|
REQUIRED_TITLE,
|
||||||
|
clean_working_protocol_markdown,
|
||||||
|
render_working_protocol,
|
||||||
|
validate_working_protocol_markdown,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
VALID_PROTOCOL = """# Working Protocol
|
||||||
|
|
||||||
|
## Materialversuche
|
||||||
|
|
||||||
|
### Background
|
||||||
|
|
||||||
|
Die Materialversuche wurden besprochen.
|
||||||
|
|
||||||
|
### Decisions
|
||||||
|
|
||||||
|
- Der nächste Versuch wird vorbereitet.
|
||||||
|
|
||||||
|
### Action Items
|
||||||
|
|
||||||
|
- Verantwortung offen: Materialstatus prüfen.
|
||||||
|
|
||||||
|
### Open Questions
|
||||||
|
|
||||||
|
- Wann liegen die Ergebnisse vor?
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
class WorkingProtocolRendererTests(unittest.TestCase):
|
||||||
|
def test_valid_protocol_beginning_with_required_heading_passes(self) -> None:
|
||||||
|
report = validate_working_protocol_markdown(VALID_PROTOCOL)
|
||||||
|
|
||||||
|
self.assertTrue(report.valid)
|
||||||
|
self.assertEqual(report.violations, [])
|
||||||
|
|
||||||
|
def test_explanatory_prose_before_valid_protocol_is_removed(self) -> None:
|
||||||
|
cleaned = clean_working_protocol_markdown(
|
||||||
|
"Hier ist das Protokoll:\n\n" + VALID_PROTOCOL
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertTrue(cleaned.startswith(REQUIRED_TITLE + "\n"))
|
||||||
|
self.assertNotIn("Hier ist das Protokoll", cleaned)
|
||||||
|
self.assertTrue(validate_working_protocol_markdown(cleaned).valid)
|
||||||
|
|
||||||
|
def test_leading_whitespace_before_valid_heading_is_removed(self) -> None:
|
||||||
|
cleaned = clean_working_protocol_markdown("\n \n # Working Protocol\n\n## Thema\n\n### Background\n\nText.\n")
|
||||||
|
|
||||||
|
self.assertTrue(cleaned.startswith(REQUIRED_TITLE + "\n"))
|
||||||
|
self.assertTrue(validate_working_protocol_markdown(cleaned).valid)
|
||||||
|
|
||||||
|
def test_output_without_required_heading_is_rejected(self) -> None:
|
||||||
|
report = validate_working_protocol_markdown("## Materialversuche\n\nText.\n")
|
||||||
|
|
||||||
|
self.assertFalse(report.valid)
|
||||||
|
self.assertEqual(report.violations[0]["type"], "missing_required_heading")
|
||||||
|
|
||||||
|
def test_categorized_summary_without_topic_body_is_rejected(self) -> None:
|
||||||
|
text = """Hier ist die Zusammenfassung:
|
||||||
|
|
||||||
|
### **Entscheidungen (Decisions)**
|
||||||
|
|
||||||
|
- Entscheidung.
|
||||||
|
|
||||||
|
### **Handlungsaufträge (Action Items)**
|
||||||
|
|
||||||
|
- Aufgabe.
|
||||||
|
"""
|
||||||
|
|
||||||
|
cleaned = clean_working_protocol_markdown(text)
|
||||||
|
report = validate_working_protocol_markdown(cleaned)
|
||||||
|
|
||||||
|
self.assertFalse(report.valid)
|
||||||
|
self.assertEqual(report.violations[0]["type"], "missing_required_heading")
|
||||||
|
|
||||||
|
def test_categorized_summary_after_heading_is_rejected(self) -> None:
|
||||||
|
text = """# Working Protocol
|
||||||
|
|
||||||
|
## Entscheidungen
|
||||||
|
|
||||||
|
- Entscheidung.
|
||||||
|
|
||||||
|
## Action Items
|
||||||
|
|
||||||
|
- Aufgabe.
|
||||||
|
"""
|
||||||
|
|
||||||
|
report = validate_working_protocol_markdown(text)
|
||||||
|
|
||||||
|
self.assertFalse(report.valid)
|
||||||
|
violation_types = {violation["type"] for violation in report.violations}
|
||||||
|
self.assertIn("generic_summary_framing", violation_types)
|
||||||
|
self.assertIn("missing_topic_sections", violation_types)
|
||||||
|
|
||||||
|
def test_malformed_markdown_is_rejected(self) -> None:
|
||||||
|
text = """# Working Protocol
|
||||||
|
|
||||||
|
## Materialversuche
|
||||||
|
|
||||||
|
### Background
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"unterbrochen": true}
|
||||||
|
"""
|
||||||
|
|
||||||
|
report = validate_working_protocol_markdown(text)
|
||||||
|
|
||||||
|
self.assertFalse(report.valid)
|
||||||
|
self.assertIn(
|
||||||
|
"malformed_markdown",
|
||||||
|
{violation["type"] for violation in report.violations},
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_raw_response_preserved_and_protocol_contains_only_validated_content(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
input_path = root / "input.json"
|
||||||
|
input_path.write_text('{"items": []}\n', encoding="utf-8")
|
||||||
|
output_dir = root / "working_protocol"
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.protocol.render_working_protocol.call_ollama",
|
||||||
|
return_value=(
|
||||||
|
{
|
||||||
|
"response": "Einleitung.\n\n" + VALID_PROTOCOL,
|
||||||
|
"done": True,
|
||||||
|
"done_reason": "stop",
|
||||||
|
},
|
||||||
|
1.25,
|
||||||
|
),
|
||||||
|
):
|
||||||
|
metadata = render_working_protocol(input_path, output_dir)
|
||||||
|
|
||||||
|
raw = json.loads((output_dir / "raw_model_response.json").read_text(encoding="utf-8"))
|
||||||
|
self.assertEqual(raw["response"], "Einleitung.\n\n" + VALID_PROTOCOL)
|
||||||
|
self.assertTrue((output_dir / "cleaned_candidate.md").exists())
|
||||||
|
self.assertTrue((output_dir / "validation_report.json").exists())
|
||||||
|
self.assertTrue(metadata["valid"])
|
||||||
|
|
||||||
|
protocol = (output_dir / "working_protocol.md").read_text(encoding="utf-8")
|
||||||
|
self.assertEqual(protocol, clean_working_protocol_markdown(raw["response"]))
|
||||||
|
self.assertNotIn("Einleitung.", protocol)
|
||||||
|
|
||||||
|
def test_invalid_output_preserves_raw_but_does_not_write_protocol(self) -> None:
|
||||||
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
|
root = Path(directory)
|
||||||
|
input_path = root / "input.json"
|
||||||
|
input_path.write_text('{"items": []}\n', encoding="utf-8")
|
||||||
|
output_dir = root / "working_protocol"
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.protocol.render_working_protocol.call_ollama",
|
||||||
|
return_value=(
|
||||||
|
{
|
||||||
|
"response": "Hier ist die Zusammenfassung:\n\n### Entscheidungen\n\n- X\n",
|
||||||
|
"done": True,
|
||||||
|
"done_reason": "stop",
|
||||||
|
},
|
||||||
|
1.25,
|
||||||
|
),
|
||||||
|
):
|
||||||
|
metadata = render_working_protocol(input_path, output_dir)
|
||||||
|
|
||||||
|
self.assertFalse(metadata["valid"])
|
||||||
|
self.assertTrue((output_dir / "raw_model_response.json").exists())
|
||||||
|
self.assertTrue((output_dir / "cleaned_candidate.md").exists())
|
||||||
|
self.assertTrue((output_dir / "validation_report.json").exists())
|
||||||
|
self.assertFalse((output_dir / "working_protocol.md").exists())
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
Reference in New Issue
Block a user