Improve semantic classification precision for BUG-015

This commit is contained in:
2026-08-09 16:15:35 +02:00
parent 0b24351127
commit fd7d5e1424
19 changed files with 464 additions and 15 deletions
+36
View File
@@ -375,6 +375,20 @@ extraction results into `meeting_protocol.md` for technical validation. The
planned architecture separates Canonical Meeting Knowledge from the final Output
Views documented in `output-views.md`.
## Extraction classification contract
Decision, Action Item and Open Question extraction shares one evidence-oriented
classification contract. It is placed after the transcript so it remains the
final classification instruction in the single multi-category extraction call.
Decisions require a settled outcome; Action Items require established work;
Open Questions require a concrete unresolved need. Unsupported candidates must
not be moved into another category.
Action existence and responsibility attribution are separate checks. A valid
Action Item may have no known owner, while a named owner requires explicit
assignment, volunteering or acceptance. These are semantic LLM classifications;
deterministic validation must not guess intent from keywords.
## Semantic Consolidator failure handling
Semantic Consolidator V0 preserves every raw model response before parsing.
@@ -396,6 +410,28 @@ an instruction not to emit an identical group more than once. Both attempts
and the detected repetition metadata are preserved. If the retry also fails,
the stage fails normally; it does not make another LLM call.
## Working Protocol V2 renderer contract
The renderer deterministically projects consolidated items to the semantic
fields required for presentation and omits bulky provenance fields from the
LLM request. Every renderable item remains represented; structurally empty
items are recorded separately rather than turned into invented prose. The
compact renderer input is preserved as an artifact.
The exact Markdown structure is generated from the same heading constants used
by the validator and appended after the renderer input so it remains visible
within the evaluated context. Decisions, action items and open questions carry
input-derived hidden coverage markers. Strict validation requires every such
renderable priority item exactly once in its matching section and rejects
missing, duplicate, wrong-section or invented markers. Facts and technical
details remain condensable as background.
Renderer output budgeting is adaptive to required priority content and prompt
size, while an explicit `num_predict` override remains authoritative. Raw model
output is always preserved, and `working_protocol.md` is written only after
strict structure and coverage validation passes. The renderer does not retry
automatically.
---
# Next Milestone
+120 -1
View File
@@ -724,7 +724,7 @@ Evidence:
## EXP-0015 - Difficult synthetic meeting
Status: Running
Status: Accepted
Date or period: 2026-07-30
@@ -1231,3 +1231,122 @@ Evidence:
- `docs/output-views.md`
- `prompts/working_protocol.md`
- `tests/gold/responsibility_attribution_negative/`
## EXP-0024 - Working Protocol V2 contract visibility
Status: Running
Date or period: 2026-08-09
Target:
BUG-011 renderer-only regression using an already validated Semantic
Consolidator artifact.
Hypothesis:
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
placing the leading prompt contract outside the model's effective evaluated
context. The preserved failures all report `prompt_eval_count=16386`, while
their outputs either echo trailing JSON or produce an unconstrained generic
category summary instead of the requested Working Protocol.
Iteration 1 change:
- Project every consolidated item to rendering-relevant semantic fields while
retaining every item and its category/text/responsibility/deadline/status
information.
- Generate the exact structural contract from renderer validator constants and
append it after the compact INPUT JSON.
- Replace the independently handwritten prompt skeleton with a reference to
that authoritative appended contract.
- Enforce the existing prompt rule that emitted sections must not be empty.
This is one renderer-contract prompt iteration. It does not change extraction,
canonicalization, semantic consolidation or responsibility semantics.
Validation before LLM run:
- 15 focused renderer tests pass.
- Tests cover contract generation, compact input projection, valid and invalid
headings, missing topic sections, empty sections, wrapper cleanup, malformed
Markdown and final-file write gating.
Decision:
Iteration 1 passed structural validation and wrote `working_protocol.md`, but
the quality sanity check found that the model omitted most of the ten supplied
decisions, two open questions and several action items. The structurally valid
result therefore was not accepted as BUG-011 verification.
Iteration 2 change:
- Add input-derived hidden coverage markers for every decision, action item and
open question.
- Require every priority item exactly once in its matching section.
- Validate missing, duplicate, unknown and wrong-section markers
deterministically.
- Keep facts and technical details condensable as background.
This is the second single prompt iteration. It responds to the concrete
omission failure observed in Iteration 1 without changing upstream semantics or
inventing renderer content.
Iteration 2 pre-run validation:
- 18 focused renderer tests pass, including exact required-item coverage and
wrong-section rejection.
Decision:
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget
after emitting all decisions and most action items. Adaptive renderer budgeting
resolved to 8,192 tokens for this input. The final run stopped normally after
3,709 evaluated output tokens.
The final renderer-only regression passed strict validation and wrote
`working_protocol.md`. Exact coverage was 10/10 decisions, 36/36 renderable
action items and 23/23 open questions, each once in its matching section. One
structurally empty action item whose task, responsible, deadline and evidence
were all null was recorded and excluded rather than fabricated. Optional
background markers were accepted only when they referred to real projected
input items.
Accept the compact renderer input, validator-derived trailing contract,
priority-item coverage markers, empty-section validation and adaptive renderer
output sizing as the BUG-011 baseline. This establishes structural reliability
and priority-item coverage, not complete protocol prose quality.
Evidence:
- `prompts/working_protocol.md`
- `src/meeting_lab/protocol/render_working_protocol.py`
- `tests/test_render_working_protocol.py`
- `samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/`
- `samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/`
## EXP-0025 — BUG-015 classification precision
Date: 2026-08-09
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
Model/configuration: `qwen3.5:9B`, temperature 0, `num_ctx=32768`.
Tests were created before prompt changes. A Decision-only evidence threshold
kept the explicit Dr. Schlummer rejection and omitted an option and preference.
Adding Action and Open Question definitions improved several negatives but was
not stable: the model alternately promoted an unaccepted Textor suggestion or
moved rejected candidates into Open Questions. Moving the standalone category
prompts after the transcript made the partial Decision schema dominate and
misclassified true Action Items as Decisions in two consecutive runs.
The final iteration replaced the competing standalone category prompts with a
single unified classification contract after the transcript. It preserved the
assigned Nina task and ownerless established CET work, and prevented
cross-category leakage in the focused case, but still emitted the unaccepted
Textor suggestion as an Action Item. Further prompt iterations were stopped in
accordance with the Gold Standard methodology.
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
remains Open; no phrase-specific deterministic filter was introduced.
+78
View File
@@ -840,6 +840,51 @@ Verified. The renderer contract is now enforced deterministically. This does
not improve semantic quality of the generated prose; it prevents invalid
renderer output from being accepted as a final Working Protocol.
Production-blocker follow-up, 2026-08-09:
Later preserved renderer failures showed that enforcement alone did not make a
valid protocol reliably obtainable. Provenance-heavy consolidated JSON inputs
were 199-350 KB, while failed Ollama runs consistently evaluated 16,386 prompt
tokens. The leading handwritten renderer contract was therefore effectively
lost or underweighted: models echoed trailing JSON or produced generic category
summaries. Prompt and validator also duplicated the structure independently,
and the validator did not enforce the prompt's no-empty-section rule.
The renderer now:
- preserves every renderable item in a compact semantic projection while
removing source-reference and original-value bulk
- records and excludes structurally empty items instead of inventing content
- appends an authoritative contract generated from validator heading constants
- rejects empty emitted sections
- requires hidden, input-derived coverage markers for every decision, action
item and open question exactly once in the matching section
- accepts optional background markers only for real projected input items
- uses adaptive output sizing instead of the truncating fixed 4,096-token cap
- keeps explicit output-budget overrides authoritative and makes no automatic
renderer retry
Renderer-only verification used the validated consolidated input at
`/tmp/meeting-lab-bug014-regression/consolidated_extractions.json`. The final
run used `qwen3.5:9B`, temperature 0, `num_ctx=32768`, adaptive
`num_predict=8192`, and stopped normally with `eval_count=3709` and
`done_reason=stop` after 53.069 seconds. Strict validation reported no
violations and wrote:
- `/tmp/meeting-lab-bug011-renderer-final/working_protocol.md`
Coverage was complete for all renderable priority items: 10 decisions, 36
action items and 23 open questions, with no missing or duplicate markers. One
upstream action item containing only null task/responsibility/deadline/evidence
was recorded as structurally empty and not rendered. The protocol retained
substantial technical background and did not add a new named responsibility;
the only structured responsible person in the renderer input remained Marleen.
Status remains Verified for Working Protocol V2 structural validity and
priority-item coverage. This does not claim full semantic or editorial protocol
quality, topic quality, or resolution of the other renderer-related regression
bugs.
## BUG-012
ID: BUG-012
@@ -1124,3 +1169,36 @@ source-ID occurrences and 113 unique canonical source IDs. There were no
missing IDs, unknown IDs or duplicate occurrences. BUG-014 is independent of
BUG-013: BUG-013 detects invalid JSON caused by runaway repeated groups,
whereas BUG-014 repairs unknown source IDs in parseable model grouping JSON.
## BUG-015
ID: BUG-015
Title: Extraction classification precision for Decisions, Action Items, and Open Questions
Status: Open
First observed: Progeo production benchmark and BUG-011 renderer-only output
The Progeo extraction classified proposals/options as Decisions, suggestions
and hypothetical work as Action Items, and uncertainty or discussion fragments
as Open Questions. The production prompt combined a weak shared German rule,
a standalone Decision prompt, a responsibility-only Action prompt, and no
Open-Question definition in one multi-category extraction request.
Regression fixtures now record explicit positive and negative evidence for all
three categories. Extraction uses one unified classification contract after the
transcript, with precision-first thresholds, category boundaries, evidence
requirements, and the responsibility-attribution invariant. In a seven-chunk
Progeo measurement, bare uncertainty in chunk 08 stopped becoming an Open
Question and the preference about real plant versus Technikum in chunk 16
stopped becoming a Decision. Supported examples including the explicit Dr.
Schlummer rejection and established CET work remained detectable.
BUG-015 is not Verified. With `qwen3.5:9B`, temperature 0, the focused Action
fixture still classified the unaccepted suggestion to contact Dirk Textor as
an Action Item. Earlier prompt variants also showed category leakage by turning
rejected candidates into Open Questions or Decisions. Prompt iterations were
stopped under the Gold Standard methodology rather than adding phrase-specific
filters. A general semantic classification architecture improvement remains
necessary before this bug can be marked Verified.