Improve semantic classification precision for BUG-015
This commit is contained in:
@@ -840,6 +840,51 @@ Verified. The renderer contract is now enforced deterministically. This does
|
||||
not improve semantic quality of the generated prose; it prevents invalid
|
||||
renderer output from being accepted as a final Working Protocol.
|
||||
|
||||
Production-blocker follow-up, 2026-08-09:
|
||||
|
||||
Later preserved renderer failures showed that enforcement alone did not make a
|
||||
valid protocol reliably obtainable. Provenance-heavy consolidated JSON inputs
|
||||
were 199-350 KB, while failed Ollama runs consistently evaluated 16,386 prompt
|
||||
tokens. The leading handwritten renderer contract was therefore effectively
|
||||
lost or underweighted: models echoed trailing JSON or produced generic category
|
||||
summaries. Prompt and validator also duplicated the structure independently,
|
||||
and the validator did not enforce the prompt's no-empty-section rule.
|
||||
|
||||
The renderer now:
|
||||
|
||||
- preserves every renderable item in a compact semantic projection while
|
||||
removing source-reference and original-value bulk
|
||||
- records and excludes structurally empty items instead of inventing content
|
||||
- appends an authoritative contract generated from validator heading constants
|
||||
- rejects empty emitted sections
|
||||
- requires hidden, input-derived coverage markers for every decision, action
|
||||
item and open question exactly once in the matching section
|
||||
- accepts optional background markers only for real projected input items
|
||||
- uses adaptive output sizing instead of the truncating fixed 4,096-token cap
|
||||
- keeps explicit output-budget overrides authoritative and makes no automatic
|
||||
renderer retry
|
||||
|
||||
Renderer-only verification used the validated consolidated input at
|
||||
`/tmp/meeting-lab-bug014-regression/consolidated_extractions.json`. The final
|
||||
run used `qwen3.5:9B`, temperature 0, `num_ctx=32768`, adaptive
|
||||
`num_predict=8192`, and stopped normally with `eval_count=3709` and
|
||||
`done_reason=stop` after 53.069 seconds. Strict validation reported no
|
||||
violations and wrote:
|
||||
|
||||
- `/tmp/meeting-lab-bug011-renderer-final/working_protocol.md`
|
||||
|
||||
Coverage was complete for all renderable priority items: 10 decisions, 36
|
||||
action items and 23 open questions, with no missing or duplicate markers. One
|
||||
upstream action item containing only null task/responsibility/deadline/evidence
|
||||
was recorded as structurally empty and not rendered. The protocol retained
|
||||
substantial technical background and did not add a new named responsibility;
|
||||
the only structured responsible person in the renderer input remained Marleen.
|
||||
|
||||
Status remains Verified for Working Protocol V2 structural validity and
|
||||
priority-item coverage. This does not claim full semantic or editorial protocol
|
||||
quality, topic quality, or resolution of the other renderer-related regression
|
||||
bugs.
|
||||
|
||||
## BUG-012
|
||||
|
||||
ID: BUG-012
|
||||
@@ -1124,3 +1169,36 @@ source-ID occurrences and 113 unique canonical source IDs. There were no
|
||||
missing IDs, unknown IDs or duplicate occurrences. BUG-014 is independent of
|
||||
BUG-013: BUG-013 detects invalid JSON caused by runaway repeated groups,
|
||||
whereas BUG-014 repairs unknown source IDs in parseable model grouping JSON.
|
||||
|
||||
## BUG-015
|
||||
|
||||
ID: BUG-015
|
||||
|
||||
Title: Extraction classification precision for Decisions, Action Items, and Open Questions
|
||||
|
||||
Status: Open
|
||||
|
||||
First observed: Progeo production benchmark and BUG-011 renderer-only output
|
||||
|
||||
The Progeo extraction classified proposals/options as Decisions, suggestions
|
||||
and hypothetical work as Action Items, and uncertainty or discussion fragments
|
||||
as Open Questions. The production prompt combined a weak shared German rule,
|
||||
a standalone Decision prompt, a responsibility-only Action prompt, and no
|
||||
Open-Question definition in one multi-category extraction request.
|
||||
|
||||
Regression fixtures now record explicit positive and negative evidence for all
|
||||
three categories. Extraction uses one unified classification contract after the
|
||||
transcript, with precision-first thresholds, category boundaries, evidence
|
||||
requirements, and the responsibility-attribution invariant. In a seven-chunk
|
||||
Progeo measurement, bare uncertainty in chunk 08 stopped becoming an Open
|
||||
Question and the preference about real plant versus Technikum in chunk 16
|
||||
stopped becoming a Decision. Supported examples including the explicit Dr.
|
||||
Schlummer rejection and established CET work remained detectable.
|
||||
|
||||
BUG-015 is not Verified. With `qwen3.5:9B`, temperature 0, the focused Action
|
||||
fixture still classified the unaccepted suggestion to contact Dirk Textor as
|
||||
an Action Item. Earlier prompt variants also showed category leakage by turning
|
||||
rejected candidates into Open Questions or Decisions. Prompt iterations were
|
||||
stopped under the Gold Standard methodology rather than adding phrase-specific
|
||||
filters. A general semantic classification architecture improvement remains
|
||||
necessary before this bug can be marked Verified.
|
||||
|
||||
Reference in New Issue
Block a user