Improve semantic classification precision for BUG-015

This commit is contained in:
2026-08-09 16:15:35 +02:00
parent 0b24351127
commit fd7d5e1424
19 changed files with 464 additions and 15 deletions
+78
View File
@@ -840,6 +840,51 @@ Verified. The renderer contract is now enforced deterministically. This does
not improve semantic quality of the generated prose; it prevents invalid
renderer output from being accepted as a final Working Protocol.
Production-blocker follow-up, 2026-08-09:
Later preserved renderer failures showed that enforcement alone did not make a
valid protocol reliably obtainable. Provenance-heavy consolidated JSON inputs
were 199-350 KB, while failed Ollama runs consistently evaluated 16,386 prompt
tokens. The leading handwritten renderer contract was therefore effectively
lost or underweighted: models echoed trailing JSON or produced generic category
summaries. Prompt and validator also duplicated the structure independently,
and the validator did not enforce the prompt's no-empty-section rule.
The renderer now:
- preserves every renderable item in a compact semantic projection while
removing source-reference and original-value bulk
- records and excludes structurally empty items instead of inventing content
- appends an authoritative contract generated from validator heading constants
- rejects empty emitted sections
- requires hidden, input-derived coverage markers for every decision, action
item and open question exactly once in the matching section
- accepts optional background markers only for real projected input items
- uses adaptive output sizing instead of the truncating fixed 4,096-token cap
- keeps explicit output-budget overrides authoritative and makes no automatic
renderer retry
Renderer-only verification used the validated consolidated input at
`/tmp/meeting-lab-bug014-regression/consolidated_extractions.json`. The final
run used `qwen3.5:9B`, temperature 0, `num_ctx=32768`, adaptive
`num_predict=8192`, and stopped normally with `eval_count=3709` and
`done_reason=stop` after 53.069 seconds. Strict validation reported no
violations and wrote:
- `/tmp/meeting-lab-bug011-renderer-final/working_protocol.md`
Coverage was complete for all renderable priority items: 10 decisions, 36
action items and 23 open questions, with no missing or duplicate markers. One
upstream action item containing only null task/responsibility/deadline/evidence
was recorded as structurally empty and not rendered. The protocol retained
substantial technical background and did not add a new named responsibility;
the only structured responsible person in the renderer input remained Marleen.
Status remains Verified for Working Protocol V2 structural validity and
priority-item coverage. This does not claim full semantic or editorial protocol
quality, topic quality, or resolution of the other renderer-related regression
bugs.
## BUG-012
ID: BUG-012
@@ -1124,3 +1169,36 @@ source-ID occurrences and 113 unique canonical source IDs. There were no
missing IDs, unknown IDs or duplicate occurrences. BUG-014 is independent of
BUG-013: BUG-013 detects invalid JSON caused by runaway repeated groups,
whereas BUG-014 repairs unknown source IDs in parseable model grouping JSON.
## BUG-015
ID: BUG-015
Title: Extraction classification precision for Decisions, Action Items, and Open Questions
Status: Open
First observed: Progeo production benchmark and BUG-011 renderer-only output
The Progeo extraction classified proposals/options as Decisions, suggestions
and hypothetical work as Action Items, and uncertainty or discussion fragments
as Open Questions. The production prompt combined a weak shared German rule,
a standalone Decision prompt, a responsibility-only Action prompt, and no
Open-Question definition in one multi-category extraction request.
Regression fixtures now record explicit positive and negative evidence for all
three categories. Extraction uses one unified classification contract after the
transcript, with precision-first thresholds, category boundaries, evidence
requirements, and the responsibility-attribution invariant. In a seven-chunk
Progeo measurement, bare uncertainty in chunk 08 stopped becoming an Open
Question and the preference about real plant versus Technikum in chunk 16
stopped becoming a Decision. Supported examples including the explicit Dr.
Schlummer rejection and established CET work remained detectable.
BUG-015 is not Verified. With `qwen3.5:9B`, temperature 0, the focused Action
fixture still classified the unaccepted suggestion to contact Dirk Textor as
an Action Item. Earlier prompt variants also showed category leakage by turning
rejected candidates into Open Questions or Decisions. Prompt iterations were
stopped under the Gold Standard methodology rather than adding phrase-specific
filters. A general semantic classification architecture improvement remains
necessary before this bug can be marked Verified.