Files
meeting-lab/docs/regression-bugs.md
T

1205 lines
42 KiB
Markdown

# Regression Bug Tracker
Living tracker for real bugs discovered during end-to-end Meeting Lab
evaluation.
Purpose:
- prevent forgotten regressions
- document known root causes
- record fixes when they happen
- verify that bugs never silently return
## BUG-001
ID: BUG-001
Title: Renderer reverses the intended process direction
Pipeline stage: Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated working protocol phrases the process direction as if the
existing F&E process is broadly taken over for all project types. The intended
meaning from the evaluation is more constrained: the existing F&E process is a
framework that may be adapted or extended for broader project handling.
Expected behaviour:
The protocol should preserve the direction and uncertainty of the source
discussion: F&E process/framework adapted for other project types where
appropriate, without implying unconditional adoption for all project types.
Actual behaviour:
The protocol states that the existing F&E process is generally adopted for all
project types.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from end-to-end evaluation output.
Notes:
Do not treat this entry as a prompt-change instruction. It records the observed
bug only.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. The new protocol still says the existing F&E process is
basically taken over for all project types.
## BUG-002
ID: BUG-002
Title: Jovana todo is strengthened beyond the meeting content
Pipeline stage: Extraction / Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated output strengthens a proposal or discussion about Jovana
assembling criteria into an assigned todo. The evaluation identified this as
overstating the meeting content.
Expected behaviour:
The pipeline should preserve the weaker source meaning unless the transcript
explicitly assigns, accepts or confirms responsibility.
Actual behaviour:
The consolidated input and final protocol include an action item assigning
criterion compilation to Jovana.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/todos.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from end-to-end evaluation output.
Notes:
This bug is subject to the responsibility attribution invariant in
`AGENTS.md`.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. The consolidated output still contains `Zusammenstellen der
Kriterien durch Jovana` with `responsible: "Jovana"`, and the final protocol
renders it as an action item.
## BUG-003
ID: BUG-003
Title: Unexpected synthetic person names appear in generated protocols
Pipeline stage: Extraction / Working Protocol Renderer
Severity: Medium
Status: Root Cause Identified
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated protocol contains unexpected or risky names/aliases such as
`Guido` and `Noah`. These names were flagged during evaluation as synthetic or
inconsistent in the generated protocol context.
Expected behaviour:
The protocol should use only names supported by the transcript, consolidated
input or Meeting Context, and should preserve uncertainty when a name or alias
is unclear.
Actual behaviour:
The final protocol includes names or name variants that were flagged as
unexpected in evaluation.
Likely root cause:
The unexpected names already exist in the Whisper transcript before extraction.
They are propagated through the pipeline and are not introduced by prompts,
gold tests or the renderer.
Related files:
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/real_live/project_process_meeting/meeting_context.yaml`
Regression test available (yes/no): no
Current status:
Root Cause Identified. Planned resolution is deferred until Meeting Context V2
/ Entity Registry implementation.
Notes:
Resolution strategy:
Introduce an interactive entity verification step after Whisper transcription:
```text
Whisper
↓
Entity Detection
↓
User Verification
↓
Meeting Context Builder
↓
Extraction
```
Expected effect:
Unknown or suspicious person names are detected before extraction begins and
can be classified by the user as participant, mentioned person, external
person, transcription error or ignore.
Implementation status:
Architecture accepted. Implementation deferred.
Regression test:
Not yet possible before Meeting Context V2 exists.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. `Guido` and `Noah` both appear in the generated working
protocol.
## BUG-004
ID: BUG-004
Title: Rhetorical question extracted as an open question
Pipeline stage: Extraction
Severity: Medium
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
A rhetorical or discussion-framing question is extracted and later rendered as
an open question. This makes the protocol imply that the meeting left a real
follow-up question unresolved.
Expected behaviour:
Only genuine unresolved questions should be extracted as open questions.
Rhetorical questions or conversational framing should not become protocol open
questions.
Actual behaviour:
The final protocol includes at least one open question identified during
evaluation as rhetorical rather than genuinely open.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/questions.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from end-to-end evaluation output.
Notes:
No specific fix has been proposed.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains. The final protocol still contains discussion-framing or
underspecified questions as open protocol questions, including the question
about exactly who is responsible for defining filter criteria.
## BUG-005
ID: BUG-005
Title: Renderer treats Björn as absent although Meeting Context marks him present
Pipeline stage: Working Protocol Renderer
Severity: Medium
Status: Investigating
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/final_protocol_with_context`
Description:
The generated protocol asks how missing participants such as Jovana and Björn
should be included in the process, even though Meeting Context marks Björn as
present. This creates a misleading participant-status implication.
Expected behaviour:
The protocol should not mark or imply Björn as absent when Meeting Context
records him as present. If the source discussion concerns whether a person had
reviewed materials or been included in a process, that should not be converted
into meeting absence.
Actual behaviour:
The final protocol renders an open question about including missing
participants, including Björn.
Likely root cause:
The extraction model receives Meeting Context but still classifies Björn with
missing participants in chunk 08. Downstream artifacts preserve only minimal
Meeting Context provenance, not structured participant attendance metadata, so
Canonicalizer, Semantic Consolidator and Renderer cannot correct the conflict.
Related files:
- `samples/real_live/project_process_meeting/meeting_context.yaml`
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `docs/bug-005-root-cause.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Investigating. Root-cause evidence has been documented in
`docs/bug-005-root-cause.md`, but no fix has been implemented or verified.
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
the issue remains.
Notes:
This bug is distinct from responsibility attribution. It concerns participant
presence/status handling.
Recommended resolution remains a Meeting Context constraint validator over
structured participant metadata before rendering.
## BUG-006
ID: BUG-006
Title: Renderer drops required consolidated items and changes item categories
Pipeline stage: Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
Description:
The Working Protocol Renderer does not preserve every consolidated decision and
open question even though the renderer prompt requires preserving all
decisions, action items and open questions. It also renders some consolidated
facts as decisions.
Expected behaviour:
Every consolidated decision and open question should be represented in the
final protocol unless there is an explicit, validated reason to omit it.
Background facts must not be promoted into decisions.
Actual behaviour:
The consolidated input contains decisions such as `Festlegung eines
Reporting-Zyklus für abgelehnte Projekte`, `Zusammengetragen und beantwortete
Fragen führen zur Sitzung mit fünf Leuten...`, and `Die Diskussion wird
beendet und die Änderungen werden verschickt...`; these are not preserved as
decisions in the final protocol. The consolidated input also contains open
questions such as `Wer soll die Rolle des Gatekeepers übernehmen...` and
`Welche Prüfsteine sind relevant für die Fachabteilung?`; these are omitted
from the final protocol. Conversely, consolidated facts such as `Im Zweifel
gehen Projekte durch` and `Ein Projekt kann abgelehnt werden...` are rendered
as decisions.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from current end-to-end benchmark output.
Notes:
This is a renderer faithfulness problem independent from whether the upstream
extracted items are themselves correct.
## BUG-007
ID: BUG-007
Title: Non-committed discussion items are extracted and rendered as todos
Pipeline stage: Extraction / Working Protocol Renderer
Severity: High
Status: Open
Date discovered: 2026-08-03
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
Description:
The pipeline extracts and renders several action items that are not clearly
assigned, accepted or confirmed commitments in the meeting evidence.
Expected behaviour:
Only explicit assignments, accepted responsibilities or confirmed follow-up
actions should become todos. Discussion, examples, proposals, broad process
needs or unclear alternatives should remain facts/questions or be marked
unclear where supported.
Actual behaviour:
The current consolidated output and final protocol contain todos such as
`Feedback zu den Kriterien und Änderungen am Prozess abwarten`, `Einbinden der
Fachbereiche (stellvertretend durch Guido) zur Definition von Prüfsteinen`,
and `Mail an Jovana und Björn erneut senden oder Prozessanpassung vornehmen`.
The evidence for these items is discussion-level or ambiguous rather than a
clear todo commitment.
Likely root cause:
Unknown.
Related files:
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
- `prompts/todos.md`
- `prompts/working_protocol.md`
Regression test available (yes/no): no
Current status:
Open. Documented from current end-to-end benchmark output.
Notes:
This generalizes BUG-002 beyond the specific Jovana assignment case and should
be evaluated against the responsibility attribution invariant.
## BUG-008
ID: BUG-008
Title: Semantic Consolidator emits duplicate source fact coverage
Pipeline stage: Semantic Consolidator V0 / Constraint Repair
Severity: High
Status: Verified
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_20260804_083849`
Description:
The Progeo end-to-end evaluation stopped during Semantic Consolidator V0
validation because the preserved model grouping output assigned the same source
fact ID to more than one group.
Expected behaviour:
Every source fact ID must appear exactly once in the Semantic Consolidator V0
fact grouping output. The validator must reject duplicate or missing source
coverage before a consolidated document is rendered.
Actual behaviour:
The validator correctly rejected the preserved raw consolidator response with:
```text
Source item IDs appear in multiple groups: ['fact_0012']
```
`fact_0012` appeared once in a merged group with `fact_0002` and again as a
singleton group. The generic validator report recorded this as one
`duplicate_id` violation with both structural occurrences.
Recovery result:
The preserved raw consolidator response was repaired with the generic
Constraint Repair V1 interface and deterministic structural operations:
- remove the repeated `fact_0012` occurrence from the later singleton group
- remove the now-empty singleton group
No Semantic Consolidator LLM call was rerun. No prompts, Meeting Context,
extraction outputs or canonicalizer output were modified. Second validation
succeeded: every source fact ID appears exactly once, no IDs are missing, no
empty groups remain, and non-fact categories remained unchanged.
This verifies the recovery path for this structural coverage failure. It does
not fix or change the Semantic Consolidator generation behaviour itself.
Related files:
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/raw_model_response.txt`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_before.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_model_groups.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_consolidated_extractions.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_after.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repair_metadata.json`
- `samples/benchmarks/progeo_meeting_20260804_083849/working_protocol/working_protocol.md`
- `docs/constraint-repair.md`
Regression test available (yes/no): no
Current status:
Verified. The recovery path repaired and revalidated this benchmark artifact,
then allowed the renderer to run once on the repaired consolidated output.
Notes:
The renderer output was produced, but it starts with explanatory prose instead
of the required `# Working Protocol` heading. That is a renderer faithfulness
issue, not part of the Semantic Consolidator coverage repair verified here.
## BUG-009
ID: BUG-009
Title: Invalid responsible-party values survive canonicalization
Pipeline stage: Extraction / Deterministic Canonicalizer
Severity: High
Status: Verified
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_context_v1_20260804_110913`
Description:
The context-aware Progeo benchmark produced action items whose `responsible`
field contained dates or date fragments instead of responsible entities. The
canonicalizer accepted these values as plain strings.
Expected behaviour:
When Meeting Context is available, action-item responsible values should be
deterministically resolved to known participants, mentioned people or supported
organization entities. Dates, locations, projects, products, technical terms,
generic process words and unknown free text must not survive as responsible
parties.
Actual behaviour:
The previous canonicalized Progeo artifact contained invalid responsible
values such as:
- `31. August`
- `15. oder 16. September`
- `am 31.`
- `27.8.`
Root cause:
The extraction schema allowed `responsible` to be a free-form string or null.
Canonicalizer V1 parsed and preserved that string without checking it against
Meeting Context or obvious non-person/non-organization patterns.
Fix:
Canonicalizer V1 now validates responsible fields when Meeting Context is
available either through `--meeting-context` or extraction provenance. It:
- accepts exact participant and mentioned-person display names
- accepts participant and mentioned-person aliases
- normalizes accepted aliases to canonical display names
- accepts explicit organization/department names only when represented in the
current Meeting Context structure
- rejects dates, relative dates, weekdays, clock times, locations, projects,
products, systems, technical terms, generic process words and unknown free
text
- clears rejected responsible values to null
- records structured `responsibility_validation` data and an aggregate
`responsibility_validations` report
Verification:
The existing context-aware Progeo extraction artifacts were re-canonicalized
without rerunning chunking, normalization or extraction:
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
Action item count remained 28 and semantic content other than responsible
fields remained unchanged. Invalid date-like responsible values were cleared.
`Martin Tazl` normalized to `Martin`, and `Marleen` remained valid. `Marleen
Wever` was rejected because the current Progeo Meeting Context does not list it
as a display name or alias.
Related files:
- `src/meeting_lab/consolidation/canonicalize.py`
- `tests/test_canonicalize.py`
- `samples/real_live/progeo_meeting/meeting_context.yaml`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer/canonicalized_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
Regression test available (yes/no): yes
Current status:
Verified. Focused tests pass and the Progeo downstream-only regression demo
removes the observed invalid responsible values without changing action-item
count or non-responsible semantic content.
Notes:
This does not prove that remaining accepted responsibilities are semantically
supported by transcript evidence. It only prevents invalid entity values from
surviving in the `responsible` field.
## BUG-010
ID: BUG-010
Title: Semantic Consolidator JSON truncates on larger real-life meetings
Pipeline stage: Semantic Consolidator V0
Severity: High
Status: Fixed
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_context_v1_20260804_110913`
Description:
The context-aware Progeo benchmark produced substantially more canonical fact
items than the earlier run. Semantic Consolidator V0 used the fixed committed
`num_predict` value of 4096 for its single grouping response and the preserved
raw response stopped in the middle of a JSON group.
Expected behaviour:
The consolidator should allocate enough output budget for the expected
fact-group JSON, preserve the raw model response, parse only valid JSON and
validate exact source fact coverage before writing consolidated output.
Actual behaviour:
The raw response ended after `fact_0060`/start of the next group, while Ollama
reported `eval_count: 4096`, exactly matching the old response cap. Parsing
failed before source coverage validation:
```text
Invalid model JSON: Expecting property name enclosed in double quotes
```
Root cause:
Output was truncated at the fixed `num_predict` cap. The prompt asks the model
to emit one JSON group per fact unless duplicates are found, so output size
scales with fact count and fact text size. The fixed cap was adequate for the
previous smaller Progeo run but not for the context-aware run.
Fix:
Semantic Consolidator V0 now estimates the response budget from the actual fact
payload and context-window headroom when `--num-predict` is not explicitly set.
Explicit `--num-predict` values are still respected.
After valid model JSON is parsed, the consolidator also applies deterministic
source-coverage repair before strict validation:
- duplicate source IDs after the first occurrence are removed
- empty groups created by removal are dropped
- missing source fact IDs are restored as singleton groups from the
canonicalized input
This repair does not rewrite existing model group text, invent facts or create
new semantic merges. Strict validation still runs after repair and can still
reject the output.
Verification:
The Progeo context canonicalized benchmark was rerun through the fixed
Semantic Consolidator into:
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/`
The fixed run used `num_predict: 14347`, the model returned valid JSON with
`eval_count: 4568`, deterministic repair restored seven omitted source fact
IDs as singletons and final validation passed with 77 unique source fact IDs,
no duplicates and no missing IDs.
Related files:
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `tests/test_consolidate_facts.py`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator/raw_model_response.txt`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/consolidated_extractions.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/repair_metadata.json`
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/report.md`
Regression test available (yes/no): yes
Current status:
Fixed for the observed truncation failure and protected by focused
consolidator tests. This does not improve semantic merge quality; it only
prevents fixed-budget truncation and enforces complete source coverage
deterministically.
## BUG-011
ID: BUG-011
Title: Working Protocol Renderer V2 writes contract-invalid output
Pipeline stage: Working Protocol Renderer V2
Severity: High
Status: Verified
Date discovered: 2026-08-04
Version first observed: `progeo_meeting_rc1_20260804_121502`
Description:
The RC1 Progeo renderer completed its LLM call and wrote `working_protocol.md`,
but the generated document did not start with `# Working Protocol`. It started
with explanatory prose and used category-summary/report framing instead of the
committed topic-oriented Working Protocol V2 structure.
Expected behaviour:
The renderer must preserve raw model output separately, deterministically clean
only harmless wrapper text, validate the cleaned candidate against the committed
Working Protocol V2 contract and write `working_protocol.md` only when the
candidate is valid.
Actual behaviour:
The one-off renderer path used for the benchmark wrote the raw LLM text to
`working_protocol.md` even though metadata recorded `valid=false` and
`readable_markdown=false`.
Root cause:
The committed contract existed in `prompts/working_protocol.md`, but there was
no reusable Working Protocol V2 renderer implementation with a strict output
validator. The model ignored explicit prompt instructions, and the pipeline did
not enforce the contract deterministically before preserving the final protocol
file.
Fix:
`src/meeting_lab/protocol/render_working_protocol.py` now implements the
Working Protocol V2 renderer path. It:
- preserves the raw Ollama JSON response and raw response text
- removes leading prose only when a real `# Working Protocol` heading exists
- normalizes harmless heading whitespace
- rejects output without the required heading
- rejects category-summary framing that lacks topic-oriented body structure
- rejects malformed Markdown such as unclosed fenced code blocks
- writes `working_protocol.md` only after validation succeeds
Verification:
Focused renderer tests cover valid output, deterministic preamble cleanup,
leading whitespace cleanup, missing heading rejection, categorized-summary
rejection, malformed Markdown rejection, raw response preservation and ensuring
`working_protocol.md` contains only validated content.
RC1 downstream demo:
- Existing RC1 raw renderer output contained no embedded valid `# Working
Protocol` body.
- The new renderer path was run once against the existing RC1 consolidated
input with current committed settings.
- The raw response and cleaned candidate were preserved.
- Validation correctly rejected the output and did not write a final
`working_protocol.md`.
Related files:
- `src/meeting_lab/protocol/render_working_protocol.py`
- `tests/test_render_working_protocol.py`
- `docs/output-views.md`
- `samples/benchmarks/progeo_meeting_rc1_20260804_121502/working_protocol_bug011/`
Regression test available (yes/no): yes
Current status:
Verified. The renderer contract is now enforced deterministically. This does
not improve semantic quality of the generated prose; it prevents invalid
renderer output from being accepted as a final Working Protocol.
Production-blocker follow-up, 2026-08-09:
Later preserved renderer failures showed that enforcement alone did not make a
valid protocol reliably obtainable. Provenance-heavy consolidated JSON inputs
were 199-350 KB, while failed Ollama runs consistently evaluated 16,386 prompt
tokens. The leading handwritten renderer contract was therefore effectively
lost or underweighted: models echoed trailing JSON or produced generic category
summaries. Prompt and validator also duplicated the structure independently,
and the validator did not enforce the prompt's no-empty-section rule.
The renderer now:
- preserves every renderable item in a compact semantic projection while
removing source-reference and original-value bulk
- records and excludes structurally empty items instead of inventing content
- appends an authoritative contract generated from validator heading constants
- rejects empty emitted sections
- requires hidden, input-derived coverage markers for every decision, action
item and open question exactly once in the matching section
- accepts optional background markers only for real projected input items
- uses adaptive output sizing instead of the truncating fixed 4,096-token cap
- keeps explicit output-budget overrides authoritative and makes no automatic
renderer retry
Renderer-only verification used the validated consolidated input at
`/tmp/meeting-lab-bug014-regression/consolidated_extractions.json`. The final
run used `qwen3.5:9B`, temperature 0, `num_ctx=32768`, adaptive
`num_predict=8192`, and stopped normally with `eval_count=3709` and
`done_reason=stop` after 53.069 seconds. Strict validation reported no
violations and wrote:
- `/tmp/meeting-lab-bug011-renderer-final/working_protocol.md`
Coverage was complete for all renderable priority items: 10 decisions, 36
action items and 23 open questions, with no missing or duplicate markers. One
upstream action item containing only null task/responsibility/deadline/evidence
was recorded as structurally empty and not rendered. The protocol retained
substantial technical background and did not add a new named responsibility;
the only structured responsible person in the renderer input remained Marleen.
Status remains Verified for Working Protocol V2 structural validity and
priority-item coverage. This does not claim full semantic or editorial protocol
quality, topic quality, or resolution of the other renderer-related regression
bugs.
## BUG-012
ID: BUG-012
Title: Chunking sensitivity to Whisper segmentation density
Pipeline stage: Chunking / Extraction
Severity: High
Status: Verified hypothesis
Date discovered: 2026-08-05
Version first observed: `progeo_whisper_turbo_20260805_120903`
Description:
Experiment 1 compared two valid trimmed whisper.cpp transcriptions while
keeping Meeting Context, source code, prompts, LLM, model parameters, pipeline
stages, validation/repair behavior and renderer configuration identical.
The large-v3 transcription produced 3,822 input segments, 216,500 segment-text
characters and 25 chunks. It completed extraction, canonicalization, Semantic
Consolidator V0 and renderer invocation under the identical downstream
settings.
The large-v3-turbo transcription produced 1,211 input segments, 83,473
segment-text characters and 10 chunks. Extraction failed at `chunk_04`. The
preserved raw response repeated the same fact many times and truncated
mid-string, producing invalid JSON. Because the extraction output for
`chunk_04` was invalid, canonicalization, semantic consolidation and rendering
could not proceed.
Expected behaviour:
Chunk construction should provide stable extraction inputs across reasonable
transcription segment-density differences. A valid, shorter Turbo transcript
should not fail extraction solely because its segment boundaries create
different chunk shapes.
Actual behaviour:
The pipeline was materially more compatible with the large-v3 transcript than
with the Turbo transcript. The Turbo run failed before downstream stages even
though the downstream settings were unchanged.
Important interpretation:
This does not prove that whisper.cpp large-v3 is intrinsically better than
large-v3-turbo. It proves only that large-v3 was more compatible with the
current chunking/extraction path in this controlled run.
Suspected root cause:
The current chunker treats Whisper JSON segments as atomic blocks and builds
chunks by character budget around those source segments. Different Whisper
segment density therefore changes the effective extraction input shape. The
suspected failure mode is chunk/token budgeting and extraction stability, not
merely transcription readability or Whisper quality.
Related files:
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_large_v3_trimmed_converted.json`
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json`
- `samples/benchmarks/progeo_whisper_large_v3_20260805_120903/`
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/`
- `samples/benchmarks/progeo_whisper_turbo_small_chunks_20260805_124324/`
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/extractions/chunk_04_extraction.raw.txt`
- `samples/benchmarks/progeo_whisper_comparison_20260805_120903_metrics.json`
- `src/meeting_lab/chunking/chunk_transcript.py`
Regression test available (yes/no): no
Current status:
Verified hypothesis. Experiment 1B reran the same Turbo transcription through
the same downstream settings while changing only the chunk budget from the
default `target_chars=9000`, `max_chars=11000`, `min_chars=5000`,
`overlap_blocks=0` to `target_chars=4500`, `max_chars=5500`,
`min_chars=2500`, `overlap_blocks=0`.
The reduced-size Turbo run produced 19 chunks instead of 10. All extraction
chunks produced valid JSON, canonicalization completed, Semantic Consolidator
V0 completed with existing deterministic source-coverage repair, and the
renderer was invoked. Renderer contract validation still failed because the
candidate did not start with `# Working Protocol`; that remains covered by
BUG-011 and does not invalidate this chunking hypothesis.
This verifies the hypothesis for this transcript/configuration only. It does
not claim the general chunker is fixed.
## BUG-013
ID: BUG-013
Title: Semantic Consolidator repetitive output loop exhausts generation budget
Pipeline stage: Semantic Consolidator
Severity: High
Status: Verified
Date discovered: 2026-08-09
Version first observed: `progeo_north_linux_20260809_124954`
Description:
The Semantic Consolidator entered a pathological generation loop that emitted
the same complete semantic group 192 consecutive times. The repeated group had
the stable structural signature consisting of the same `canonical_text`, the
same ordered `source_item_ids` (`fact_0004`, `fact_0019`) and the same
`merge_reason`. Generation then truncated at an incomplete
`"canonical_text":` field.
Expected behaviour:
The consolidator should detect a structural repeated-group loop
deterministically. Invalid JSON should receive at most one controlled retry
only when that loop signature is present. The retry must use the same model,
temperature and context configuration and add only an instruction preventing
duplicate group emission. Unrelated malformed JSON must not be retried.
Actual behaviour:
The first model response was invalid JSON at line 974, column 24, character
71184. Semantic Consolidator runtime was 277.910 seconds. The raw response had
71,184 characters (71,568 UTF-8 bytes), ended at `"canonical_text":`, and
could not reach deterministic source-coverage repair because it was not
parseable JSON.
Model metadata:
- Resolved `num_predict`: 19,532
- `num_ctx`: 32,768
- `prompt_eval_count`: 14,114
- `eval_count`: 18,654
- `done_reason`: `length`
- `eval_count` did not equal resolved `num_predict`; it exactly consumed the
remaining evaluated context (`32768 - 14114 = 18654`)
- Complete groups before truncation: 194
- Consecutive identical groups: 192, starting at group index 2
Root cause:
The model produced a degenerate repeated semantic-group sequence until the
remaining context budget was exhausted. BUG-010 adaptive response sizing was
working as designed; increasing `num_predict` would not address this failure
mode.
Implemented handling:
- Extract complete group objects from a response even when its outer JSON is
truncated.
- Identify groups by normalized `canonical_text`, ordered `source_item_ids`
and normalized `merge_reason`.
- Classify three or more consecutive identical complete groups as a loop; two
identical occurrences remain an ordinary duplicate handled by deterministic
coverage repair when the JSON is valid.
- Preserve first-attempt raw text, raw Ollama JSON and repetition metadata.
- Retry at most once only when JSON parsing fails and a loop is detected.
- Preserve separate retry artifacts and fail normally if retry parsing fails.
Related files:
- `samples/benchmarks/progeo_north_linux_20260809_124954/`
- `samples/benchmarks/progeo_north_linux_20260809_124954/semantic_consolidator/raw_model_response.txt`
- `samples/benchmarks/progeo_north_linux_20260809_124954/semantic_consolidator/raw_ollama_response.json`
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `scripts/run_meeting.py`
- `tests/test_consolidate_facts.py`
Regression test available (yes/no): yes
Current status:
Verified for the repetitive-loop failure mode. A consolidator-only regression
used the preserved canonicalized input and did not run extraction,
canonicalization or rendering.
The first attempt reproduced the original failure signature: 71,184 response
characters, `eval_count=18654`, `done_reason=length`, invalid JSON at character
71184, and a detected 192-group consecutive repetition run. This triggered the
only permitted retry.
The retry used the same model and generation configuration, returned parseable
JSON after 7,362 evaluated tokens with `done_reason=stop`, and contained no
repetition loop (longest identical consecutive run: 1). Its source-coverage
validation before repair was invalid: 105 groups, 154 observed source-ID
occurrences, duplicate IDs, four missing IDs and one unknown ID. Deterministic
repair made 63 changes: 44 duplicate-ID removals, 15 empty-group removals and
four missing-ID singleton restorations. Validation after repair remained
invalid only because group 89 contained unknown `fact_0114`.
The stage therefore failed cleanly after the single retry and preserved both
attempts. No third LLM call occurred. The remaining unknown-ID failure is not a
repetitive generation loop; its deterministic repair is tracked separately as
BUG-014. This verification does not claim that general Semantic Consolidator
output quality is solved.
## BUG-014
ID: BUG-014
Title: Semantic Consolidator hallucinates unknown source item IDs
Pipeline stage: Semantic Consolidator deterministic coverage repair
Severity: High
Status: Verified
Date discovered: 2026-08-09
Version first observed: BUG-013 consolidator-only regression
Description:
The parseable BUG-013 retry response contained a semantic group at index 104
with `source_item_ids` equal to `fact_0113` and unknown `fact_0114`. The
canonicalized input contains 113 facts ending at `fact_0113`; no `fact_0114`
exists.
The group canonical text, "Ich habe ihn heute Morgen im Büro angetroffen.", is
directly supported by valid `fact_0113`. The unknown ID adds no identifiable
canonical source content and appears to be an additional hallucinated ID, not
a safely correctable typo or a reference that can be mapped to another input
fact.
Invariant:
Every `source_item_id` emitted by the Semantic Consolidator must refer to an
existing canonical input fact. Strict validation must continue to reject any
unknown ID that survives deterministic repair.
Deterministic repair rule:
- Remove every string source ID that is not present in the canonical fact-ID
set. Never map it to a numerically nearby or semantically guessed ID.
- Preserve all valid source IDs in the group.
- Preserve the group when at least one valid source ID remains.
- Remove the group when unknown-ID removal leaves it empty.
- Apply existing duplicate-ID removal to surviving valid IDs.
- Restore genuinely missing canonical facts as singletons through the existing
coverage repair.
- Preserve the original group text and merge reason; do not rewrite semantics.
Related files:
- `/tmp/meeting-lab-bug013-regression/raw_model_response_retry.txt`
- `samples/benchmarks/progeo_north_linux_20260809_124954/canonicalizer/canonicalized_extractions.json`
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `tests/test_consolidate_facts.py`
Regression test available (yes/no): yes
Current status:
Verified with a consolidator-only regression using the preserved canonicalized
input. Extraction, canonicalization, full benchmark orchestration and rendering
were not run.
The first attempt reproduced BUG-013: `eval_count=18654`,
`done_reason=length`, invalid JSON and a detected 192-group consecutive
repetition run. This triggered one controlled retry. The retry returned 105
parseable groups with 154 source-ID occurrences, `eval_count=7362`,
`done_reason=stop`, and no repetition loop.
Strict source-coverage inspection before repair found one unknown ID
(`fact_0114` in group 104), duplicate IDs and four missing canonical IDs.
Deterministic repair made 64 changes:
- 44 `remove_duplicate_source_id`
- 1 `remove_unknown_source_id`
- 15 `remove_empty_group`
- 4 `restore_missing_source_id_as_singleton`
Strict validation after repair passed with 94 groups containing exactly 113
source-ID occurrences and 113 unique canonical source IDs. There were no
missing IDs, unknown IDs or duplicate occurrences. BUG-014 is independent of
BUG-013: BUG-013 detects invalid JSON caused by runaway repeated groups,
whereas BUG-014 repairs unknown source IDs in parseable model grouping JSON.
## BUG-015
ID: BUG-015
Title: Extraction classification precision for Decisions, Action Items, and Open Questions
Status: Open
First observed: Progeo production benchmark and BUG-011 renderer-only output
The Progeo extraction classified proposals/options as Decisions, suggestions
and hypothetical work as Action Items, and uncertainty or discussion fragments
as Open Questions. The production prompt combined a weak shared German rule,
a standalone Decision prompt, a responsibility-only Action prompt, and no
Open-Question definition in one multi-category extraction request.
Regression fixtures now record explicit positive and negative evidence for all
three categories. Extraction uses one unified classification contract after the
transcript, with precision-first thresholds, category boundaries, evidence
requirements, and the responsibility-attribution invariant. In a seven-chunk
Progeo measurement, bare uncertainty in chunk 08 stopped becoming an Open
Question and the preference about real plant versus Technikum in chunk 16
stopped becoming a Decision. Supported examples including the explicit Dr.
Schlummer rejection and established CET work remained detectable.
BUG-015 is not Verified. With `qwen3.5:9B`, temperature 0, the focused Action
fixture still classified the unaccepted suggestion to contact Dirk Textor as
an Action Item. Earlier prompt variants also showed category leakage by turning
rejected candidates into Open Questions or Decisions. Prompt iterations were
stopped under the Gold Standard methodology rather than adding phrase-specific
filters. A general semantic classification architecture improvement remains
necessary before this bug can be marked Verified.