Stabilize Meeting Lab pipeline for RC1 evaluation

This commit significantly improves the robustness and determinism of the Meeting Lab processing pipeline and establishes the first Release Candidate baseline for end-to-end evaluation.

Highlights

- BUG-009
  - Implement deterministic responsible-party validation
  - Normalize participant aliases using Meeting Context
  - Reject invalid responsible values (dates, locations, technical terms, projects, products, unknown entities)
  - Record structured responsibility validation metadata
  - Add focused regression tests

- BUG-010
  - Implement adaptive num_predict estimation for Semantic Consolidator
  - Eliminate JSON truncation caused by fixed output limits
  - Add deterministic source coverage repair
  - Preserve strict post-repair validation
  - Add regression tests

- BUG-011
  - Implement Working Protocol V2 renderer contract enforcement
  - Preserve raw renderer responses
  - Reject invalid protocol output instead of accepting malformed documents
  - Add deterministic cleanup for harmless formatting deviations
  - Add focused renderer regression tests

- Meeting Context
  - Validate Meeting Context V1
  - Integrate authoritative participant alias normalization

- Documentation
  - Update architecture documentation
  - Update output documentation
  - Update regression bug tracker

The pipeline now fails safely instead of silently accepting invalid intermediate or final artifacts.

Remaining work focuses primarily on extraction quality and semantic classification (decisions, action items, protocol faithfulness), rather than pipeline robustness.
This commit is contained in:
2026-08-04 13:11:54 +02:00
parent 60a8acae91
commit 950284e236
10 changed files with 1816 additions and 18 deletions
+16
View File
@@ -265,6 +265,11 @@ Deterministic Canonicalizer:
- validates and normalizes extraction objects
- assigns stable source references and IDs
- normalizes category names and basic field structure
- validates and normalizes action-item responsible fields against Meeting
Context when available: known participant and mentioned-person aliases are
normalized to canonical display names, while dates, locations, projects,
products, technical terms, generic process words and unknown free text are
cleared with a structured validation record
- performs only safe deterministic cleanup
- may group exact duplicates
- preserves all source evidence
@@ -277,6 +282,12 @@ Semantic Consolidator:
- V0 merges semantically equivalent fact items conservatively
- V0 preserves source references and evidence
- V0 validates that every source fact appears exactly once
- V0 sizes its Ollama output budget from the actual fact payload instead of
using a fixed response cap for every meeting
- V0 may apply deterministic source-coverage repair after valid model JSON is
parsed: duplicate source IDs are removed after their first occurrence, empty
groups are removed and missing source facts are restored as singleton groups
from canonicalized input before strict validation runs
- V0 does not process non-fact categories semantically
- later versions should group content by topic, mark contradictions and
uncertainty, separate durable information from transient discussion and
@@ -295,6 +306,11 @@ It only reformulates the analysis results for a specific audience and purpose.
Depending on the output and maturity of the implementation, a renderer may be
deterministic, template-based or LLM-assisted.
LLM-assisted renderers preserve raw model output separately and write the final
output artifact only after deterministic contract validation succeeds. Renderer
post-processing may remove non-semantic wrapper text, but must not fabricate
missing semantic sections or relabel an invalid summary as a valid output view.
The planned output products are:
- Working Protocol (`working_protocol.md`, Arbeitsprotokoll)
+15
View File
@@ -110,6 +110,21 @@ Characteristics:
Completeness goal: optimize for recall and traceability.
Current Working Protocol V2 renderer contract:
- output must begin exactly with `# Working Protocol`
- no explanatory preamble may appear before that heading
- content is organized by topic with `##` topic headings
- each topic may use only the supported `###` sections: `Background`,
`Decisions`, `Action Items`, `Open Questions`
- category-level report framing such as a global `## Decisions` / `## Action
Items` summary is not a valid topic-oriented Working Protocol
- raw model responses are preserved separately
- deterministic cleanup may remove leading prose before an already valid
`# Working Protocol` heading and normalize harmless heading whitespace
- `working_protocol.md` is written only after the cleaned candidate passes the
renderer contract validator
## Concise Distribution Protocol
Suggested filename: `distribution_protocol.md`
+363
View File
@@ -476,3 +476,366 @@ Notes:
This generalizes BUG-002 beyond the specific Jovana assignment case and should
be evaluated against the responsibility attribution invariant.
## BUG-008
ID: BUG-008
Title: Semantic Consolidator emits duplicate source fact coverage
Pipeline stage: Semantic Consolidator V0 / Constraint Repair
Severity: High
Status: Verified
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_20260804_083849`
Description:
The Progeo end-to-end evaluation stopped during Semantic Consolidator V0
validation because the preserved model grouping output assigned the same source
fact ID to more than one group.
Expected behaviour:
Every source fact ID must appear exactly once in the Semantic Consolidator V0
fact grouping output. The validator must reject duplicate or missing source
coverage before a consolidated document is rendered.
Actual behaviour:
The validator correctly rejected the preserved raw consolidator response with:
```text
Source item IDs appear in multiple groups: ['fact_0012']
```
`fact_0012` appeared once in a merged group with `fact_0002` and again as a
singleton group. The generic validator report recorded this as one
`duplicate_id` violation with both structural occurrences.
Recovery result:
The preserved raw consolidator response was repaired with the generic
Constraint Repair V1 interface and deterministic structural operations:
- remove the repeated `fact_0012` occurrence from the later singleton group
- remove the now-empty singleton group
No Semantic Consolidator LLM call was rerun. No prompts, Meeting Context,
extraction outputs or canonicalizer output were modified. Second validation
succeeded: every source fact ID appears exactly once, no IDs are missing, no
empty groups remain, and non-fact categories remained unchanged.
This verifies the recovery path for this structural coverage failure. It does
not fix or change the Semantic Consolidator generation behaviour itself.
Related files:
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/raw_model_response.txt`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_before.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_model_groups.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_consolidated_extractions.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_after.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repair_metadata.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/working_protocol/working_protocol.md`
- `docs/constraint-repair.md`
Regression test available (yes/no): no
Current status:
Verified. The recovery path repaired and revalidated this benchmark artifact,
then allowed the renderer to run once on the repaired consolidated output.
Notes:
The renderer output was produced, but it starts with explanatory prose instead
of the required `# Working Protocol` heading. That is a renderer faithfulness
issue, not part of the Semantic Consolidator coverage repair verified here.
## BUG-009
ID: BUG-009
Title: Invalid responsible-party values survive canonicalization
Pipeline stage: Extraction / Deterministic Canonicalizer
Severity: High
Status: Verified
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_context_v1_20260804_110913`
Description:
The context-aware Progeo benchmark produced action items whose `responsible`
field contained dates or date fragments instead of responsible entities. The
canonicalizer accepted these values as plain strings.
Expected behaviour:
When Meeting Context is available, action-item responsible values should be
deterministically resolved to known participants, mentioned people or supported
organization entities. Dates, locations, projects, products, technical terms,
generic process words and unknown free text must not survive as responsible
parties.
Actual behaviour:
The previous canonicalized Progeo artifact contained invalid responsible
values such as:
- `31. August`
- `15. oder 16. September`
- `am 31.`
- `27.8.`
Root cause:
The extraction schema allowed `responsible` to be a free-form string or null.
Canonicalizer V1 parsed and preserved that string without checking it against
Meeting Context or obvious non-person/non-organization patterns.
Fix:
Canonicalizer V1 now validates responsible fields when Meeting Context is
available either through `--meeting-context` or extraction provenance. It:
- accepts exact participant and mentioned-person display names
- accepts participant and mentioned-person aliases
- normalizes accepted aliases to canonical display names
- accepts explicit organization/department names only when represented in the
current Meeting Context structure
- rejects dates, relative dates, weekdays, clock times, locations, projects,
products, systems, technical terms, generic process words and unknown free
text
- clears rejected responsible values to null
- records structured `responsibility_validation` data and an aggregate
`responsibility_validations` report
Verification:
The existing context-aware Progeo extraction artifacts were re-canonicalized
without rerunning chunking, normalization or extraction:
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
Action item count remained 28 and semantic content other than responsible
fields remained unchanged. Invalid date-like responsible values were cleared.
`Martin Tazl` normalized to `Martin`, and `Marleen` remained valid. `Marleen
Wever` was rejected because the current Progeo Meeting Context does not list it
as a display name or alias.
Related files:
- `src/meeting_lab/consolidation/canonicalize.py`
- `tests/test_canonicalize.py`
- `samples/real_live/progeo_meeting/meeting_context.yaml`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer/canonicalized_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
Regression test available (yes/no): yes
Current status:
Verified. Focused tests pass and the Progeo downstream-only regression demo
removes the observed invalid responsible values without changing action-item
count or non-responsible semantic content.
Notes:
This does not prove that remaining accepted responsibilities are semantically
supported by transcript evidence. It only prevents invalid entity values from
surviving in the `responsible` field.
## BUG-010
ID: BUG-010
Title: Semantic Consolidator JSON truncates on larger real-life meetings
Pipeline stage: Semantic Consolidator V0
Severity: High
Status: Fixed
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_context_v1_20260804_110913`
Description:
The context-aware Progeo benchmark produced substantially more canonical fact
items than the earlier run. Semantic Consolidator V0 used the fixed committed
`num_predict` value of 4096 for its single grouping response and the preserved
raw response stopped in the middle of a JSON group.
Expected behaviour:
The consolidator should allocate enough output budget for the expected
fact-group JSON, preserve the raw model response, parse only valid JSON and
validate exact source fact coverage before writing consolidated output.
Actual behaviour:
The raw response ended after `fact_0060`/start of the next group, while Ollama
reported `eval_count: 4096`, exactly matching the old response cap. Parsing
failed before source coverage validation:
```text
Invalid model JSON: Expecting property name enclosed in double quotes
```
Root cause:
Output was truncated at the fixed `num_predict` cap. The prompt asks the model
to emit one JSON group per fact unless duplicates are found, so output size
scales with fact count and fact text size. The fixed cap was adequate for the
previous smaller Progeo run but not for the context-aware run.
Fix:
Semantic Consolidator V0 now estimates the response budget from the actual fact
payload and context-window headroom when `--num-predict` is not explicitly set.
Explicit `--num-predict` values are still respected.
After valid model JSON is parsed, the consolidator also applies deterministic
source-coverage repair before strict validation:
- duplicate source IDs after the first occurrence are removed
- empty groups created by removal are dropped
- missing source fact IDs are restored as singleton groups from the
canonicalized input
This repair does not rewrite existing model group text, invent facts or create
new semantic merges. Strict validation still runs after repair and can still
reject the output.
Verification:
The Progeo context canonicalized benchmark was rerun through the fixed
Semantic Consolidator into:
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/`
The fixed run used `num_predict: 14347`, the model returned valid JSON with
`eval_count: 4568`, deterministic repair restored seven omitted source fact
IDs as singletons and final validation passed with 77 unique source fact IDs,
no duplicates and no missing IDs.
Related files:
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `tests/test_consolidate_facts.py`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator/raw_model_response.txt`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/consolidated_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/repair_metadata.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/report.md`
Regression test available (yes/no): yes
Current status:
Fixed for the observed truncation failure and protected by focused
consolidator tests. This does not improve semantic merge quality; it only
prevents fixed-budget truncation and enforces complete source coverage
deterministically.
## BUG-011
ID: BUG-011
Title: Working Protocol Renderer V2 writes contract-invalid output
Pipeline stage: Working Protocol Renderer V2
Severity: High
Status: Verified
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_rc1_20260804_121502`
Description:
The RC1 Progeo renderer completed its LLM call and wrote `working_protocol.md`,
but the generated document did not start with `# Working Protocol`. It started
with explanatory prose and used category-summary/report framing instead of the
committed topic-oriented Working Protocol V2 structure.
Expected behaviour:
The renderer must preserve raw model output separately, deterministically clean
only harmless wrapper text, validate the cleaned candidate against the committed
Working Protocol V2 contract and write `working_protocol.md` only when the
candidate is valid.
Actual behaviour:
The one-off renderer path used for the benchmark wrote the raw LLM text to
`working_protocol.md` even though metadata recorded `valid=false` and
`readable_markdown=false`.
Root cause:
The committed contract existed in `prompts/working_protocol.md`, but there was
no reusable Working Protocol V2 renderer implementation with a strict output
validator. The model ignored explicit prompt instructions, and the pipeline did
not enforce the contract deterministically before preserving the final protocol
file.
Fix:
`src/meeting_lab/protocol/render_working_protocol.py` now implements the
Working Protocol V2 renderer path. It:
- preserves the raw Ollama JSON response and raw response text
- removes leading prose only when a real `# Working Protocol` heading exists
- normalizes harmless heading whitespace
- rejects output without the required heading
- rejects category-summary framing that lacks topic-oriented body structure
- rejects malformed Markdown such as unclosed fenced code blocks
- writes `working_protocol.md` only after validation succeeds
Verification:
Focused renderer tests cover valid output, deterministic preamble cleanup,
leading whitespace cleanup, missing heading rejection, categorized-summary
rejection, malformed Markdown rejection, raw response preservation and ensuring
`working_protocol.md` contains only validated content.
RC1 downstream demo:
- Existing RC1 raw renderer output contained no embedded valid `# Working
Protocol` body.
- The new renderer path was run once against the existing RC1 consolidated
input with current committed settings.
- The raw response and cleaned candidate were preserved.
- Validation correctly rejected the output and did not write a final
`working_protocol.md`.
Related files:
- `src/meeting_lab/protocol/render_working_protocol.py`
- `tests/test_render_working_protocol.py`
- `docs/output-views.md`
- `samples/benchmarks/progeo_meeting_rc1_20260804_121502/working_protocol_bug011/`
Regression test available (yes/no): yes
Current status:
Verified. The renderer contract is now enforced deterministically. This does
not improve semantic quality of the generated prose; it prevents invalid
renderer output from being accepted as a final Working Protocol.