Handle repetitive Semantic Consolidator output loops
This commit is contained in:
@@ -931,3 +931,113 @@ BUG-011 and does not invalidate this chunking hypothesis.
|
||||
|
||||
This verifies the hypothesis for this transcript/configuration only. It does
|
||||
not claim the general chunker is fixed.
|
||||
|
||||
## BUG-013
|
||||
|
||||
ID: BUG-013
|
||||
|
||||
Title: Semantic Consolidator repetitive output loop exhausts generation budget
|
||||
|
||||
Pipeline stage: Semantic Consolidator
|
||||
|
||||
Severity: High
|
||||
|
||||
Status: Verified
|
||||
|
||||
Date discovered: 2026-08-09
|
||||
|
||||
Version first observed: `progeo_north_linux_20260809_124954`
|
||||
|
||||
Description:
|
||||
|
||||
The Semantic Consolidator entered a pathological generation loop that emitted
|
||||
the same complete semantic group 192 consecutive times. The repeated group had
|
||||
the stable structural signature consisting of the same `canonical_text`, the
|
||||
same ordered `source_item_ids` (`fact_0004`, `fact_0019`) and the same
|
||||
`merge_reason`. Generation then truncated at an incomplete
|
||||
`"canonical_text":` field.
|
||||
|
||||
Expected behaviour:
|
||||
|
||||
The consolidator should detect a structural repeated-group loop
|
||||
deterministically. Invalid JSON should receive at most one controlled retry
|
||||
only when that loop signature is present. The retry must use the same model,
|
||||
temperature and context configuration and add only an instruction preventing
|
||||
duplicate group emission. Unrelated malformed JSON must not be retried.
|
||||
|
||||
Actual behaviour:
|
||||
|
||||
The first model response was invalid JSON at line 974, column 24, character
|
||||
71184. Semantic Consolidator runtime was 277.910 seconds. The raw response had
|
||||
71,184 characters (71,568 UTF-8 bytes), ended at `"canonical_text":`, and
|
||||
could not reach deterministic source-coverage repair because it was not
|
||||
parseable JSON.
|
||||
|
||||
Model metadata:
|
||||
|
||||
- Resolved `num_predict`: 19,532
|
||||
- `num_ctx`: 32,768
|
||||
- `prompt_eval_count`: 14,114
|
||||
- `eval_count`: 18,654
|
||||
- `done_reason`: `length`
|
||||
- `eval_count` did not equal resolved `num_predict`; it exactly consumed the
|
||||
remaining evaluated context (`32768 - 14114 = 18654`)
|
||||
- Complete groups before truncation: 194
|
||||
- Consecutive identical groups: 192, starting at group index 2
|
||||
|
||||
Root cause:
|
||||
|
||||
The model produced a degenerate repeated semantic-group sequence until the
|
||||
remaining context budget was exhausted. BUG-010 adaptive response sizing was
|
||||
working as designed; increasing `num_predict` would not address this failure
|
||||
mode.
|
||||
|
||||
Implemented handling:
|
||||
|
||||
- Extract complete group objects from a response even when its outer JSON is
|
||||
truncated.
|
||||
- Identify groups by normalized `canonical_text`, ordered `source_item_ids`
|
||||
and normalized `merge_reason`.
|
||||
- Classify three or more consecutive identical complete groups as a loop; two
|
||||
identical occurrences remain an ordinary duplicate handled by deterministic
|
||||
coverage repair when the JSON is valid.
|
||||
- Preserve first-attempt raw text, raw Ollama JSON and repetition metadata.
|
||||
- Retry at most once only when JSON parsing fails and a loop is detected.
|
||||
- Preserve separate retry artifacts and fail normally if retry parsing fails.
|
||||
|
||||
Related files:
|
||||
|
||||
- `samples/benchmarks/progeo_north_linux_20260809_124954/`
|
||||
- `samples/benchmarks/progeo_north_linux_20260809_124954/semantic_consolidator/raw_model_response.txt`
|
||||
- `samples/benchmarks/progeo_north_linux_20260809_124954/semantic_consolidator/raw_ollama_response.json`
|
||||
- `src/meeting_lab/consolidation/consolidate_facts.py`
|
||||
- `scripts/run_meeting.py`
|
||||
- `tests/test_consolidate_facts.py`
|
||||
|
||||
Regression test available (yes/no): yes
|
||||
|
||||
Current status:
|
||||
|
||||
Verified for the repetitive-loop failure mode. A consolidator-only regression
|
||||
used the preserved canonicalized input and did not run extraction,
|
||||
canonicalization or rendering.
|
||||
|
||||
The first attempt reproduced the original failure signature: 71,184 response
|
||||
characters, `eval_count=18654`, `done_reason=length`, invalid JSON at character
|
||||
71184, and a detected 192-group consecutive repetition run. This triggered the
|
||||
only permitted retry.
|
||||
|
||||
The retry used the same model and generation configuration, returned parseable
|
||||
JSON after 7,362 evaluated tokens with `done_reason=stop`, and contained no
|
||||
repetition loop (longest identical consecutive run: 1). Its source-coverage
|
||||
validation before repair was invalid: 105 groups, 154 observed source-ID
|
||||
occurrences, duplicate IDs, four missing IDs and one unknown ID. Deterministic
|
||||
repair made 63 changes: 44 duplicate-ID removals, 15 empty-group removals and
|
||||
four missing-ID singleton restorations. Validation after repair remained
|
||||
invalid only because group 89 contained unknown `fact_0114`.
|
||||
|
||||
The stage therefore failed cleanly after the single retry and preserved both
|
||||
attempts. No third LLM call occurred. The remaining unknown-ID failure is not a
|
||||
repetitive generation loop and is not silently repaired by the existing
|
||||
coverage repair. This verification does not claim that general Semantic
|
||||
Consolidator output quality is solved.
|
||||
|
||||
Reference in New Issue
Block a user