Document and preserve the controlled Progeo benchmark series. Whisper comparison: - Compare whisper.cpp large-v3 and large-v3-turbo - Identify and verify BUG-012: extraction stability depends on chunk size - Repeat both transcription variants with identical reduced chunk budgets - Select large-v3-turbo as the current production transcription model LLM comparison: - Compare qwen3.5:9b with qwen3.5:35B-A3B - Preserve identical transcript, Meeting Context, prompts and chunking - Retain qwen3.5:9b as the production recommendation - Record runtime, repair burden and semantic-quality findings Current production benchmark configuration: - whisper.cpp large-v3-turbo - target_chars=4500 - max_chars=5500 - min_chars=2500 - overlap_blocks=0 - qwen3.5:9b - num_ctx=32768 - think=false - temperature=0 BUG-012 remains verified but not yet fixed in the production chunker.
934 lines
31 KiB
Markdown
934 lines
31 KiB
Markdown
# Regression Bug Tracker
|
|
|
|
Living tracker for real bugs discovered during end-to-end Meeting Lab
|
|
evaluation.
|
|
|
|
Purpose:
|
|
|
|
- prevent forgotten regressions
|
|
- document known root causes
|
|
- record fixes when they happen
|
|
- verify that bugs never silently return
|
|
|
|
## BUG-001
|
|
|
|
ID: BUG-001
|
|
|
|
Title: Renderer reverses the intended process direction
|
|
|
|
Pipeline stage: Working Protocol Renderer
|
|
|
|
Severity: High
|
|
|
|
Status: Open
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
|
|
|
Description:
|
|
|
|
The generated working protocol phrases the process direction as if the
|
|
existing F&E process is broadly taken over for all project types. The intended
|
|
meaning from the evaluation is more constrained: the existing F&E process is a
|
|
framework that may be adapted or extended for broader project handling.
|
|
|
|
Expected behaviour:
|
|
|
|
The protocol should preserve the direction and uncertainty of the source
|
|
discussion: F&E process/framework adapted for other project types where
|
|
appropriate, without implying unconditional adoption for all project types.
|
|
|
|
Actual behaviour:
|
|
|
|
The protocol states that the existing F&E process is generally adopted for all
|
|
project types.
|
|
|
|
Likely root cause:
|
|
|
|
Unknown.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `prompts/working_protocol.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Open. Documented from end-to-end evaluation output.
|
|
|
|
Notes:
|
|
|
|
Do not treat this entry as a prompt-change instruction. It records the observed
|
|
bug only.
|
|
|
|
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
|
the issue remains. The new protocol still says the existing F&E process is
|
|
basically taken over for all project types.
|
|
|
|
## BUG-002
|
|
|
|
ID: BUG-002
|
|
|
|
Title: Jovana todo is strengthened beyond the meeting content
|
|
|
|
Pipeline stage: Extraction / Working Protocol Renderer
|
|
|
|
Severity: High
|
|
|
|
Status: Open
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
|
|
|
Description:
|
|
|
|
The generated output strengthens a proposal or discussion about Jovana
|
|
assembling criteria into an assigned todo. The evaluation identified this as
|
|
overstating the meeting content.
|
|
|
|
Expected behaviour:
|
|
|
|
The pipeline should preserve the weaker source meaning unless the transcript
|
|
explicitly assigns, accepts or confirms responsibility.
|
|
|
|
Actual behaviour:
|
|
|
|
The consolidated input and final protocol include an action item assigning
|
|
criterion compilation to Jovana.
|
|
|
|
Likely root cause:
|
|
|
|
Unknown.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `prompts/todos.md`
|
|
- `prompts/working_protocol.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Open. Documented from end-to-end evaluation output.
|
|
|
|
Notes:
|
|
|
|
This bug is subject to the responsibility attribution invariant in
|
|
`AGENTS.md`.
|
|
|
|
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
|
the issue remains. The consolidated output still contains `Zusammenstellen der
|
|
Kriterien durch Jovana` with `responsible: "Jovana"`, and the final protocol
|
|
renders it as an action item.
|
|
|
|
## BUG-003
|
|
|
|
ID: BUG-003
|
|
|
|
Title: Unexpected synthetic person names appear in generated protocols
|
|
|
|
Pipeline stage: Extraction / Working Protocol Renderer
|
|
|
|
Severity: Medium
|
|
|
|
Status: Root Cause Identified
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
|
|
|
Description:
|
|
|
|
The generated protocol contains unexpected or risky names/aliases such as
|
|
`Guido` and `Noah`. These names were flagged during evaluation as synthetic or
|
|
inconsistent in the generated protocol context.
|
|
|
|
Expected behaviour:
|
|
|
|
The protocol should use only names supported by the transcript, consolidated
|
|
input or Meeting Context, and should preserve uncertainty when a name or alias
|
|
is unclear.
|
|
|
|
Actual behaviour:
|
|
|
|
The final protocol includes names or name variants that were flagged as
|
|
unexpected in evaluation.
|
|
|
|
Likely root cause:
|
|
|
|
The unexpected names already exist in the Whisper transcript before extraction.
|
|
They are propagated through the pipeline and are not introduced by prompts,
|
|
gold tests or the renderer.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
|
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `samples/real_live/project_process_meeting/meeting_context.yaml`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Root Cause Identified. Planned resolution is deferred until Meeting Context V2
|
|
/ Entity Registry implementation.
|
|
|
|
Notes:
|
|
|
|
Resolution strategy:
|
|
|
|
Introduce an interactive entity verification step after Whisper transcription:
|
|
|
|
```text
|
|
Whisper
|
|
↓
|
|
Entity Detection
|
|
↓
|
|
User Verification
|
|
↓
|
|
Meeting Context Builder
|
|
↓
|
|
Extraction
|
|
```
|
|
|
|
Expected effect:
|
|
|
|
Unknown or suspicious person names are detected before extraction begins and
|
|
can be classified by the user as participant, mentioned person, external
|
|
person, transcription error or ignore.
|
|
|
|
Implementation status:
|
|
|
|
Architecture accepted. Implementation deferred.
|
|
|
|
Regression test:
|
|
|
|
Not yet possible before Meeting Context V2 exists.
|
|
|
|
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
|
the issue remains. `Guido` and `Noah` both appear in the generated working
|
|
protocol.
|
|
|
|
## BUG-004
|
|
|
|
ID: BUG-004
|
|
|
|
Title: Rhetorical question extracted as an open question
|
|
|
|
Pipeline stage: Extraction
|
|
|
|
Severity: Medium
|
|
|
|
Status: Open
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
|
|
|
Description:
|
|
|
|
A rhetorical or discussion-framing question is extracted and later rendered as
|
|
an open question. This makes the protocol imply that the meeting left a real
|
|
follow-up question unresolved.
|
|
|
|
Expected behaviour:
|
|
|
|
Only genuine unresolved questions should be extracted as open questions.
|
|
Rhetorical questions or conversational framing should not become protocol open
|
|
questions.
|
|
|
|
Actual behaviour:
|
|
|
|
The final protocol includes at least one open question identified during
|
|
evaluation as rhetorical rather than genuinely open.
|
|
|
|
Likely root cause:
|
|
|
|
Unknown.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `prompts/questions.md`
|
|
- `prompts/working_protocol.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Open. Documented from end-to-end evaluation output.
|
|
|
|
Notes:
|
|
|
|
No specific fix has been proposed.
|
|
|
|
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
|
the issue remains. The final protocol still contains discussion-framing or
|
|
underspecified questions as open protocol questions, including the question
|
|
about exactly who is responsible for defining filter criteria.
|
|
|
|
## BUG-005
|
|
|
|
ID: BUG-005
|
|
|
|
Title: Renderer treats Björn as absent although Meeting Context marks him present
|
|
|
|
Pipeline stage: Working Protocol Renderer
|
|
|
|
Severity: Medium
|
|
|
|
Status: Investigating
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/final_protocol_with_context`
|
|
|
|
Description:
|
|
|
|
The generated protocol asks how missing participants such as Jovana and Björn
|
|
should be included in the process, even though Meeting Context marks Björn as
|
|
present. This creates a misleading participant-status implication.
|
|
|
|
Expected behaviour:
|
|
|
|
The protocol should not mark or imply Björn as absent when Meeting Context
|
|
records him as present. If the source discussion concerns whether a person had
|
|
reviewed materials or been included in a process, that should not be converted
|
|
into meeting absence.
|
|
|
|
Actual behaviour:
|
|
|
|
The final protocol renders an open question about including missing
|
|
participants, including Björn.
|
|
|
|
Likely root cause:
|
|
|
|
The extraction model receives Meeting Context but still classifies Björn with
|
|
missing participants in chunk 08. Downstream artifacts preserve only minimal
|
|
Meeting Context provenance, not structured participant attendance metadata, so
|
|
Canonicalizer, Semantic Consolidator and Renderer cannot correct the conflict.
|
|
|
|
Related files:
|
|
|
|
- `samples/real_live/project_process_meeting/meeting_context.yaml`
|
|
- `samples/benchmarks/meeting_context_v1/semantic_consolidator_repair_v1/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/working_protocol.md`
|
|
- `samples/benchmarks/meeting_context_v1/final_protocol_with_context/comparison.md`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `docs/bug-005-root-cause.md`
|
|
- `prompts/working_protocol.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Investigating. Root-cause evidence has been documented in
|
|
`docs/bug-005-root-cause.md`, but no fix has been implemented or verified.
|
|
Current benchmark `meeting_context_v1/e2e_current_20260803_151000` confirms
|
|
the issue remains.
|
|
|
|
Notes:
|
|
|
|
This bug is distinct from responsibility attribution. It concerns participant
|
|
presence/status handling.
|
|
|
|
Recommended resolution remains a Meeting Context constraint validator over
|
|
structured participant metadata before rendering.
|
|
|
|
## BUG-006
|
|
|
|
ID: BUG-006
|
|
|
|
Title: Renderer drops required consolidated items and changes item categories
|
|
|
|
Pipeline stage: Working Protocol Renderer
|
|
|
|
Severity: High
|
|
|
|
Status: Open
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
|
|
|
|
Description:
|
|
|
|
The Working Protocol Renderer does not preserve every consolidated decision and
|
|
open question even though the renderer prompt requires preserving all
|
|
decisions, action items and open questions. It also renders some consolidated
|
|
facts as decisions.
|
|
|
|
Expected behaviour:
|
|
|
|
Every consolidated decision and open question should be represented in the
|
|
final protocol unless there is an explicit, validated reason to omit it.
|
|
Background facts must not be promoted into decisions.
|
|
|
|
Actual behaviour:
|
|
|
|
The consolidated input contains decisions such as `Festlegung eines
|
|
Reporting-Zyklus für abgelehnte Projekte`, `Zusammengetragen und beantwortete
|
|
Fragen führen zur Sitzung mit fünf Leuten...`, and `Die Diskussion wird
|
|
beendet und die Änderungen werden verschickt...`; these are not preserved as
|
|
decisions in the final protocol. The consolidated input also contains open
|
|
questions such as `Wer soll die Rolle des Gatekeepers übernehmen...` and
|
|
`Welche Prüfsteine sind relevant für die Fachabteilung?`; these are omitted
|
|
from the final protocol. Conversely, consolidated facts such as `Im Zweifel
|
|
gehen Projekte durch` and `Ein Projekt kann abgelehnt werden...` are rendered
|
|
as decisions.
|
|
|
|
Likely root cause:
|
|
|
|
Unknown.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `prompts/working_protocol.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Open. Documented from current end-to-end benchmark output.
|
|
|
|
Notes:
|
|
|
|
This is a renderer faithfulness problem independent from whether the upstream
|
|
extracted items are themselves correct.
|
|
|
|
## BUG-007
|
|
|
|
ID: BUG-007
|
|
|
|
Title: Non-committed discussion items are extracted and rendered as todos
|
|
|
|
Pipeline stage: Extraction / Working Protocol Renderer
|
|
|
|
Severity: High
|
|
|
|
Status: Open
|
|
|
|
Date discovered: 2026-08-03
|
|
|
|
Version first observed: `meeting_context_v1/e2e_current_20260803_151000`
|
|
|
|
Description:
|
|
|
|
The pipeline extracts and renders several action items that are not clearly
|
|
assigned, accepted or confirmed commitments in the meeting evidence.
|
|
|
|
Expected behaviour:
|
|
|
|
Only explicit assignments, accepted responsibilities or confirmed follow-up
|
|
actions should become todos. Discussion, examples, proposals, broad process
|
|
needs or unclear alternatives should remain facts/questions or be marked
|
|
unclear where supported.
|
|
|
|
Actual behaviour:
|
|
|
|
The current consolidated output and final protocol contain todos such as
|
|
`Feedback zu den Kriterien und Änderungen am Prozess abwarten`, `Einbinden der
|
|
Fachbereiche (stellvertretend durch Guido) zur Definition von Prüfsteinen`,
|
|
and `Mail an Jovana und Björn erneut senden oder Prozessanpassung vornehmen`.
|
|
The evidence for these items is discussion-level or ambiguous rather than a
|
|
clear todo commitment.
|
|
|
|
Likely root cause:
|
|
|
|
Unknown.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/semantic_consolidator/consolidated_extractions.json`
|
|
- `samples/benchmarks/meeting_context_v1/e2e_current_20260803_151000/working_protocol/working_protocol.md`
|
|
- `prompts/todos.md`
|
|
- `prompts/working_protocol.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Open. Documented from current end-to-end benchmark output.
|
|
|
|
Notes:
|
|
|
|
This generalizes BUG-002 beyond the specific Jovana assignment case and should
|
|
be evaluated against the responsibility attribution invariant.
|
|
|
|
## BUG-008
|
|
|
|
ID: BUG-008
|
|
|
|
Title: Semantic Consolidator emits duplicate source fact coverage
|
|
|
|
Pipeline stage: Semantic Consolidator V0 / Constraint Repair
|
|
|
|
Severity: High
|
|
|
|
Status: Verified
|
|
|
|
Date discovered: 2026-08-04
|
|
|
|
Version first observed: `progeo_meeting_20260804_083849`
|
|
|
|
Description:
|
|
|
|
The Progeo end-to-end evaluation stopped during Semantic Consolidator V0
|
|
validation because the preserved model grouping output assigned the same source
|
|
fact ID to more than one group.
|
|
|
|
Expected behaviour:
|
|
|
|
Every source fact ID must appear exactly once in the Semantic Consolidator V0
|
|
fact grouping output. The validator must reject duplicate or missing source
|
|
coverage before a consolidated document is rendered.
|
|
|
|
Actual behaviour:
|
|
|
|
The validator correctly rejected the preserved raw consolidator response with:
|
|
|
|
```text
|
|
Source item IDs appear in multiple groups: ['fact_0012']
|
|
```
|
|
|
|
`fact_0012` appeared once in a merged group with `fact_0002` and again as a
|
|
singleton group. The generic validator report recorded this as one
|
|
`duplicate_id` violation with both structural occurrences.
|
|
|
|
Recovery result:
|
|
|
|
The preserved raw consolidator response was repaired with the generic
|
|
Constraint Repair V1 interface and deterministic structural operations:
|
|
|
|
- remove the repeated `fact_0012` occurrence from the later singleton group
|
|
- remove the now-empty singleton group
|
|
|
|
No Semantic Consolidator LLM call was rerun. No prompts, Meeting Context,
|
|
extraction outputs or canonicalizer output were modified. Second validation
|
|
succeeded: every source fact ID appears exactly once, no IDs are missing, no
|
|
empty groups remain, and non-fact categories remained unchanged.
|
|
|
|
This verifies the recovery path for this structural coverage failure. It does
|
|
not fix or change the Semantic Consolidator generation behaviour itself.
|
|
|
|
Related files:
|
|
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/raw_model_response.txt`
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_before.json`
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_model_groups.json`
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repaired_consolidated_extractions.json`
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/validator_after.json`
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/semantic_consolidator/repair_metadata.json`
|
|
- `samples/benchmarks/progeo_meeting_20260804_083849/working_protocol/working_protocol.md`
|
|
- `docs/constraint-repair.md`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Verified. The recovery path repaired and revalidated this benchmark artifact,
|
|
then allowed the renderer to run once on the repaired consolidated output.
|
|
|
|
Notes:
|
|
|
|
The renderer output was produced, but it starts with explanatory prose instead
|
|
of the required `# Working Protocol` heading. That is a renderer faithfulness
|
|
issue, not part of the Semantic Consolidator coverage repair verified here.
|
|
|
|
## BUG-009
|
|
|
|
ID: BUG-009
|
|
|
|
Title: Invalid responsible-party values survive canonicalization
|
|
|
|
Pipeline stage: Extraction / Deterministic Canonicalizer
|
|
|
|
Severity: High
|
|
|
|
Status: Verified
|
|
|
|
Date discovered: 2026-08-04
|
|
|
|
Version first observed: `progeo_meeting_context_v1_20260804_110913`
|
|
|
|
Description:
|
|
|
|
The context-aware Progeo benchmark produced action items whose `responsible`
|
|
field contained dates or date fragments instead of responsible entities. The
|
|
canonicalizer accepted these values as plain strings.
|
|
|
|
Expected behaviour:
|
|
|
|
When Meeting Context is available, action-item responsible values should be
|
|
deterministically resolved to known participants, mentioned people or supported
|
|
organization entities. Dates, locations, projects, products, technical terms,
|
|
generic process words and unknown free text must not survive as responsible
|
|
parties.
|
|
|
|
Actual behaviour:
|
|
|
|
The previous canonicalized Progeo artifact contained invalid responsible
|
|
values such as:
|
|
|
|
- `31. August`
|
|
- `15. oder 16. September`
|
|
- `am 31.`
|
|
- `27.8.`
|
|
|
|
Root cause:
|
|
|
|
The extraction schema allowed `responsible` to be a free-form string or null.
|
|
Canonicalizer V1 parsed and preserved that string without checking it against
|
|
Meeting Context or obvious non-person/non-organization patterns.
|
|
|
|
Fix:
|
|
|
|
Canonicalizer V1 now validates responsible fields when Meeting Context is
|
|
available either through `--meeting-context` or extraction provenance. It:
|
|
|
|
- accepts exact participant and mentioned-person display names
|
|
- accepts participant and mentioned-person aliases
|
|
- normalizes accepted aliases to canonical display names
|
|
- accepts explicit organization/department names only when represented in the
|
|
current Meeting Context structure
|
|
- rejects dates, relative dates, weekdays, clock times, locations, projects,
|
|
products, systems, technical terms, generic process words and unknown free
|
|
text
|
|
- clears rejected responsible values to null
|
|
- records structured `responsibility_validation` data and an aggregate
|
|
`responsibility_validations` report
|
|
|
|
Verification:
|
|
|
|
The existing context-aware Progeo extraction artifacts were re-canonicalized
|
|
without rerunning chunking, normalization or extraction:
|
|
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
|
|
|
|
Action item count remained 28 and semantic content other than responsible
|
|
fields remained unchanged. Invalid date-like responsible values were cleared.
|
|
`Martin Tazl` normalized to `Martin`, and `Marleen` remained valid. `Marleen
|
|
Wever` was rejected because the current Progeo Meeting Context does not list it
|
|
as a display name or alias.
|
|
|
|
Related files:
|
|
|
|
- `src/meeting_lab/consolidation/canonicalize.py`
|
|
- `tests/test_canonicalize.py`
|
|
- `samples/real_live/progeo_meeting/meeting_context.yaml`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer/canonicalized_extractions.json`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/canonicalized_extractions.json`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/canonicalizer_bug009/responsibility_demo.json`
|
|
|
|
Regression test available (yes/no): yes
|
|
|
|
Current status:
|
|
|
|
Verified. Focused tests pass and the Progeo downstream-only regression demo
|
|
removes the observed invalid responsible values without changing action-item
|
|
count or non-responsible semantic content.
|
|
|
|
Notes:
|
|
|
|
This does not prove that remaining accepted responsibilities are semantically
|
|
supported by transcript evidence. It only prevents invalid entity values from
|
|
surviving in the `responsible` field.
|
|
|
|
## BUG-010
|
|
|
|
ID: BUG-010
|
|
|
|
Title: Semantic Consolidator JSON truncates on larger real-life meetings
|
|
|
|
Pipeline stage: Semantic Consolidator V0
|
|
|
|
Severity: High
|
|
|
|
Status: Fixed
|
|
|
|
Date discovered: 2026-08-04
|
|
|
|
Version first observed: `progeo_meeting_context_v1_20260804_110913`
|
|
|
|
Description:
|
|
|
|
The context-aware Progeo benchmark produced substantially more canonical fact
|
|
items than the earlier run. Semantic Consolidator V0 used the fixed committed
|
|
`num_predict` value of 4096 for its single grouping response and the preserved
|
|
raw response stopped in the middle of a JSON group.
|
|
|
|
Expected behaviour:
|
|
|
|
The consolidator should allocate enough output budget for the expected
|
|
fact-group JSON, preserve the raw model response, parse only valid JSON and
|
|
validate exact source fact coverage before writing consolidated output.
|
|
|
|
Actual behaviour:
|
|
|
|
The raw response ended after `fact_0060`/start of the next group, while Ollama
|
|
reported `eval_count: 4096`, exactly matching the old response cap. Parsing
|
|
failed before source coverage validation:
|
|
|
|
```text
|
|
Invalid model JSON: Expecting property name enclosed in double quotes
|
|
```
|
|
|
|
Root cause:
|
|
|
|
Output was truncated at the fixed `num_predict` cap. The prompt asks the model
|
|
to emit one JSON group per fact unless duplicates are found, so output size
|
|
scales with fact count and fact text size. The fixed cap was adequate for the
|
|
previous smaller Progeo run but not for the context-aware run.
|
|
|
|
Fix:
|
|
|
|
Semantic Consolidator V0 now estimates the response budget from the actual fact
|
|
payload and context-window headroom when `--num-predict` is not explicitly set.
|
|
Explicit `--num-predict` values are still respected.
|
|
|
|
After valid model JSON is parsed, the consolidator also applies deterministic
|
|
source-coverage repair before strict validation:
|
|
|
|
- duplicate source IDs after the first occurrence are removed
|
|
- empty groups created by removal are dropped
|
|
- missing source fact IDs are restored as singleton groups from the
|
|
canonicalized input
|
|
|
|
This repair does not rewrite existing model group text, invent facts or create
|
|
new semantic merges. Strict validation still runs after repair and can still
|
|
reject the output.
|
|
|
|
Verification:
|
|
|
|
The Progeo context canonicalized benchmark was rerun through the fixed
|
|
Semantic Consolidator into:
|
|
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/`
|
|
|
|
The fixed run used `num_predict: 14347`, the model returned valid JSON with
|
|
`eval_count: 4568`, deterministic repair restored seven omitted source fact
|
|
IDs as singletons and final validation passed with 77 unique source fact IDs,
|
|
no duplicates and no missing IDs.
|
|
|
|
Related files:
|
|
|
|
- `src/meeting_lab/consolidation/consolidate_facts.py`
|
|
- `tests/test_consolidate_facts.py`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator/raw_model_response.txt`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/consolidated_extractions.json`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/repair_metadata.json`
|
|
- `samples/benchmarks/progeo_meeting_context_v1_20260804_110913/semantic_consolidator_bug010_adaptive_repair/report.md`
|
|
|
|
Regression test available (yes/no): yes
|
|
|
|
Current status:
|
|
|
|
Fixed for the observed truncation failure and protected by focused
|
|
consolidator tests. This does not improve semantic merge quality; it only
|
|
prevents fixed-budget truncation and enforces complete source coverage
|
|
deterministically.
|
|
|
|
## BUG-011
|
|
|
|
ID: BUG-011
|
|
|
|
Title: Working Protocol Renderer V2 writes contract-invalid output
|
|
|
|
Pipeline stage: Working Protocol Renderer V2
|
|
|
|
Severity: High
|
|
|
|
Status: Verified
|
|
|
|
Date discovered: 2026-08-04
|
|
|
|
Version first observed: `progeo_meeting_rc1_20260804_121502`
|
|
|
|
Description:
|
|
|
|
The RC1 Progeo renderer completed its LLM call and wrote `working_protocol.md`,
|
|
but the generated document did not start with `# Working Protocol`. It started
|
|
with explanatory prose and used category-summary/report framing instead of the
|
|
committed topic-oriented Working Protocol V2 structure.
|
|
|
|
Expected behaviour:
|
|
|
|
The renderer must preserve raw model output separately, deterministically clean
|
|
only harmless wrapper text, validate the cleaned candidate against the committed
|
|
Working Protocol V2 contract and write `working_protocol.md` only when the
|
|
candidate is valid.
|
|
|
|
Actual behaviour:
|
|
|
|
The one-off renderer path used for the benchmark wrote the raw LLM text to
|
|
`working_protocol.md` even though metadata recorded `valid=false` and
|
|
`readable_markdown=false`.
|
|
|
|
Root cause:
|
|
|
|
The committed contract existed in `prompts/working_protocol.md`, but there was
|
|
no reusable Working Protocol V2 renderer implementation with a strict output
|
|
validator. The model ignored explicit prompt instructions, and the pipeline did
|
|
not enforce the contract deterministically before preserving the final protocol
|
|
file.
|
|
|
|
Fix:
|
|
|
|
`src/meeting_lab/protocol/render_working_protocol.py` now implements the
|
|
Working Protocol V2 renderer path. It:
|
|
|
|
- preserves the raw Ollama JSON response and raw response text
|
|
- removes leading prose only when a real `# Working Protocol` heading exists
|
|
- normalizes harmless heading whitespace
|
|
- rejects output without the required heading
|
|
- rejects category-summary framing that lacks topic-oriented body structure
|
|
- rejects malformed Markdown such as unclosed fenced code blocks
|
|
- writes `working_protocol.md` only after validation succeeds
|
|
|
|
Verification:
|
|
|
|
Focused renderer tests cover valid output, deterministic preamble cleanup,
|
|
leading whitespace cleanup, missing heading rejection, categorized-summary
|
|
rejection, malformed Markdown rejection, raw response preservation and ensuring
|
|
`working_protocol.md` contains only validated content.
|
|
|
|
RC1 downstream demo:
|
|
|
|
- Existing RC1 raw renderer output contained no embedded valid `# Working
|
|
Protocol` body.
|
|
- The new renderer path was run once against the existing RC1 consolidated
|
|
input with current committed settings.
|
|
- The raw response and cleaned candidate were preserved.
|
|
- Validation correctly rejected the output and did not write a final
|
|
`working_protocol.md`.
|
|
|
|
Related files:
|
|
|
|
- `src/meeting_lab/protocol/render_working_protocol.py`
|
|
- `tests/test_render_working_protocol.py`
|
|
- `docs/output-views.md`
|
|
- `samples/benchmarks/progeo_meeting_rc1_20260804_121502/working_protocol_bug011/`
|
|
|
|
Regression test available (yes/no): yes
|
|
|
|
Current status:
|
|
|
|
Verified. The renderer contract is now enforced deterministically. This does
|
|
not improve semantic quality of the generated prose; it prevents invalid
|
|
renderer output from being accepted as a final Working Protocol.
|
|
|
|
## BUG-012
|
|
|
|
ID: BUG-012
|
|
|
|
Title: Chunking sensitivity to Whisper segmentation density
|
|
|
|
Pipeline stage: Chunking / Extraction
|
|
|
|
Severity: High
|
|
|
|
Status: Verified hypothesis
|
|
|
|
Date discovered: 2026-08-05
|
|
|
|
Version first observed: `progeo_whisper_turbo_20260805_120903`
|
|
|
|
Description:
|
|
|
|
Experiment 1 compared two valid trimmed whisper.cpp transcriptions while
|
|
keeping Meeting Context, source code, prompts, LLM, model parameters, pipeline
|
|
stages, validation/repair behavior and renderer configuration identical.
|
|
|
|
The large-v3 transcription produced 3,822 input segments, 216,500 segment-text
|
|
characters and 25 chunks. It completed extraction, canonicalization, Semantic
|
|
Consolidator V0 and renderer invocation under the identical downstream
|
|
settings.
|
|
|
|
The large-v3-turbo transcription produced 1,211 input segments, 83,473
|
|
segment-text characters and 10 chunks. Extraction failed at `chunk_04`. The
|
|
preserved raw response repeated the same fact many times and truncated
|
|
mid-string, producing invalid JSON. Because the extraction output for
|
|
`chunk_04` was invalid, canonicalization, semantic consolidation and rendering
|
|
could not proceed.
|
|
|
|
Expected behaviour:
|
|
|
|
Chunk construction should provide stable extraction inputs across reasonable
|
|
transcription segment-density differences. A valid, shorter Turbo transcript
|
|
should not fail extraction solely because its segment boundaries create
|
|
different chunk shapes.
|
|
|
|
Actual behaviour:
|
|
|
|
The pipeline was materially more compatible with the large-v3 transcript than
|
|
with the Turbo transcript. The Turbo run failed before downstream stages even
|
|
though the downstream settings were unchanged.
|
|
|
|
Important interpretation:
|
|
|
|
This does not prove that whisper.cpp large-v3 is intrinsically better than
|
|
large-v3-turbo. It proves only that large-v3 was more compatible with the
|
|
current chunking/extraction path in this controlled run.
|
|
|
|
Suspected root cause:
|
|
|
|
The current chunker treats Whisper JSON segments as atomic blocks and builds
|
|
chunks by character budget around those source segments. Different Whisper
|
|
segment density therefore changes the effective extraction input shape. The
|
|
suspected failure mode is chunk/token budgeting and extraction stability, not
|
|
merely transcription readability or Whisper quality.
|
|
|
|
Related files:
|
|
|
|
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_large_v3_trimmed_converted.json`
|
|
- `samples/real_live/progeo_meeting/progeo_whispercpp_vulkan_turbo_trimmed_converted.json`
|
|
- `samples/benchmarks/progeo_whisper_large_v3_20260805_120903/`
|
|
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/`
|
|
- `samples/benchmarks/progeo_whisper_turbo_small_chunks_20260805_124324/`
|
|
- `samples/benchmarks/progeo_whisper_turbo_20260805_120903/extractions/chunk_04_extraction.raw.txt`
|
|
- `samples/benchmarks/progeo_whisper_comparison_20260805_120903_metrics.json`
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
|
|
Regression test available (yes/no): no
|
|
|
|
Current status:
|
|
|
|
Verified hypothesis. Experiment 1B reran the same Turbo transcription through
|
|
the same downstream settings while changing only the chunk budget from the
|
|
default `target_chars=9000`, `max_chars=11000`, `min_chars=5000`,
|
|
`overlap_blocks=0` to `target_chars=4500`, `max_chars=5500`,
|
|
`min_chars=2500`, `overlap_blocks=0`.
|
|
|
|
The reduced-size Turbo run produced 19 chunks instead of 10. All extraction
|
|
chunks produced valid JSON, canonicalization completed, Semantic Consolidator
|
|
V0 completed with existing deterministic source-coverage repair, and the
|
|
renderer was invoked. Renderer contract validation still failed because the
|
|
candidate did not start with `# Working Protocol`; that remains covered by
|
|
BUG-011 and does not invalidate this chunking hypothesis.
|
|
|
|
This verifies the hypothesis for this transcript/configuration only. It does
|
|
not claim the general chunker is fixed.
|