Handle repetitive Semantic Consolidator output loops

This commit is contained in:
2026-08-09 14:33:46 +02:00
parent 3a5850430b
commit 58effcafc5
5 changed files with 462 additions and 11 deletions
+110
View File
@@ -931,3 +931,113 @@ BUG-011 and does not invalidate this chunking hypothesis.
This verifies the hypothesis for this transcript/configuration only. It does
not claim the general chunker is fixed.
## BUG-013
ID: BUG-013
Title: Semantic Consolidator repetitive output loop exhausts generation budget
Pipeline stage: Semantic Consolidator
Severity: High
Status: Verified
Date discovered: 2026-08-09
Version first observed: `progeo_north_linux_20260809_124954`
Description:
The Semantic Consolidator entered a pathological generation loop that emitted
the same complete semantic group 192 consecutive times. The repeated group had
the stable structural signature consisting of the same `canonical_text`, the
same ordered `source_item_ids` (`fact_0004`, `fact_0019`) and the same
`merge_reason`. Generation then truncated at an incomplete
`"canonical_text":` field.
Expected behaviour:
The consolidator should detect a structural repeated-group loop
deterministically. Invalid JSON should receive at most one controlled retry
only when that loop signature is present. The retry must use the same model,
temperature and context configuration and add only an instruction preventing
duplicate group emission. Unrelated malformed JSON must not be retried.
Actual behaviour:
The first model response was invalid JSON at line 974, column 24, character
71184. Semantic Consolidator runtime was 277.910 seconds. The raw response had
71,184 characters (71,568 UTF-8 bytes), ended at `"canonical_text":`, and
could not reach deterministic source-coverage repair because it was not
parseable JSON.
Model metadata:
- Resolved `num_predict`: 19,532
- `num_ctx`: 32,768
- `prompt_eval_count`: 14,114
- `eval_count`: 18,654
- `done_reason`: `length`
- `eval_count` did not equal resolved `num_predict`; it exactly consumed the
remaining evaluated context (`32768 - 14114 = 18654`)
- Complete groups before truncation: 194
- Consecutive identical groups: 192, starting at group index 2
Root cause:
The model produced a degenerate repeated semantic-group sequence until the
remaining context budget was exhausted. BUG-010 adaptive response sizing was
working as designed; increasing `num_predict` would not address this failure
mode.
Implemented handling:
- Extract complete group objects from a response even when its outer JSON is
truncated.
- Identify groups by normalized `canonical_text`, ordered `source_item_ids`
and normalized `merge_reason`.
- Classify three or more consecutive identical complete groups as a loop; two
identical occurrences remain an ordinary duplicate handled by deterministic
coverage repair when the JSON is valid.
- Preserve first-attempt raw text, raw Ollama JSON and repetition metadata.
- Retry at most once only when JSON parsing fails and a loop is detected.
- Preserve separate retry artifacts and fail normally if retry parsing fails.
Related files:
- `samples/benchmarks/progeo_north_linux_20260809_124954/`
- `samples/benchmarks/progeo_north_linux_20260809_124954/semantic_consolidator/raw_model_response.txt`
- `samples/benchmarks/progeo_north_linux_20260809_124954/semantic_consolidator/raw_ollama_response.json`
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `scripts/run_meeting.py`
- `tests/test_consolidate_facts.py`
Regression test available (yes/no): yes
Current status:
Verified for the repetitive-loop failure mode. A consolidator-only regression
used the preserved canonicalized input and did not run extraction,
canonicalization or rendering.
The first attempt reproduced the original failure signature: 71,184 response
characters, `eval_count=18654`, `done_reason=length`, invalid JSON at character
71184, and a detected 192-group consecutive repetition run. This triggered the
only permitted retry.
The retry used the same model and generation configuration, returned parseable
JSON after 7,362 evaluated tokens with `done_reason=stop`, and contained no
repetition loop (longest identical consecutive run: 1). Its source-coverage
validation before repair was invalid: 105 groups, 154 observed source-ID
occurrences, duplicate IDs, four missing IDs and one unknown ID. Deterministic
repair made 63 changes: 44 duplicate-ID removals, 15 empty-group removals and
four missing-ID singleton restorations. Validation after repair remained
invalid only because group 89 contained unknown `fact_0114`.
The stage therefore failed cleanly after the single retry and preserved both
attempts. No third LLM call occurred. The remaining unknown-ID failure is not a
repetitive generation loop and is not silently repaired by the existing
coverage repair. This verification does not claim that general Semantic
Consolidator output quality is solved.