Improve semantic classification precision for BUG-015

This commit is contained in:
2026-08-09 16:15:35 +02:00
parent 0b24351127
commit fd7d5e1424
19 changed files with 464 additions and 15 deletions
+120 -1
View File
@@ -724,7 +724,7 @@ Evidence:
## EXP-0015 - Difficult synthetic meeting
Status: Running
Status: Accepted
Date or period: 2026-07-30
@@ -1231,3 +1231,122 @@ Evidence:
- `docs/output-views.md`
- `prompts/working_protocol.md`
- `tests/gold/responsibility_attribution_negative/`
## EXP-0024 - Working Protocol V2 contract visibility
Status: Running
Date or period: 2026-08-09
Target:
BUG-011 renderer-only regression using an already validated Semantic
Consolidator artifact.
Hypothesis:
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
placing the leading prompt contract outside the model's effective evaluated
context. The preserved failures all report `prompt_eval_count=16386`, while
their outputs either echo trailing JSON or produce an unconstrained generic
category summary instead of the requested Working Protocol.
Iteration 1 change:
- Project every consolidated item to rendering-relevant semantic fields while
retaining every item and its category/text/responsibility/deadline/status
information.
- Generate the exact structural contract from renderer validator constants and
append it after the compact INPUT JSON.
- Replace the independently handwritten prompt skeleton with a reference to
that authoritative appended contract.
- Enforce the existing prompt rule that emitted sections must not be empty.
This is one renderer-contract prompt iteration. It does not change extraction,
canonicalization, semantic consolidation or responsibility semantics.
Validation before LLM run:
- 15 focused renderer tests pass.
- Tests cover contract generation, compact input projection, valid and invalid
headings, missing topic sections, empty sections, wrapper cleanup, malformed
Markdown and final-file write gating.
Decision:
Iteration 1 passed structural validation and wrote `working_protocol.md`, but
the quality sanity check found that the model omitted most of the ten supplied
decisions, two open questions and several action items. The structurally valid
result therefore was not accepted as BUG-011 verification.
Iteration 2 change:
- Add input-derived hidden coverage markers for every decision, action item and
open question.
- Require every priority item exactly once in its matching section.
- Validate missing, duplicate, unknown and wrong-section markers
deterministically.
- Keep facts and technical details condensable as background.
This is the second single prompt iteration. It responds to the concrete
omission failure observed in Iteration 1 without changing upstream semantics or
inventing renderer content.
Iteration 2 pre-run validation:
- 18 focused renderer tests pass, including exact required-item coverage and
wrong-section rejection.
Decision:
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget
after emitting all decisions and most action items. Adaptive renderer budgeting
resolved to 8,192 tokens for this input. The final run stopped normally after
3,709 evaluated output tokens.
The final renderer-only regression passed strict validation and wrote
`working_protocol.md`. Exact coverage was 10/10 decisions, 36/36 renderable
action items and 23/23 open questions, each once in its matching section. One
structurally empty action item whose task, responsible, deadline and evidence
were all null was recorded and excluded rather than fabricated. Optional
background markers were accepted only when they referred to real projected
input items.
Accept the compact renderer input, validator-derived trailing contract,
priority-item coverage markers, empty-section validation and adaptive renderer
output sizing as the BUG-011 baseline. This establishes structural reliability
and priority-item coverage, not complete protocol prose quality.
Evidence:
- `prompts/working_protocol.md`
- `src/meeting_lab/protocol/render_working_protocol.py`
- `tests/test_render_working_protocol.py`
- `samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/`
- `samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/`
## EXP-0025 — BUG-015 classification precision
Date: 2026-08-09
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
Model/configuration: `qwen3.5:9B`, temperature 0, `num_ctx=32768`.
Tests were created before prompt changes. A Decision-only evidence threshold
kept the explicit Dr. Schlummer rejection and omitted an option and preference.
Adding Action and Open Question definitions improved several negatives but was
not stable: the model alternately promoted an unaccepted Textor suggestion or
moved rejected candidates into Open Questions. Moving the standalone category
prompts after the transcript made the partial Decision schema dominate and
misclassified true Action Items as Decisions in two consecutive runs.
The final iteration replaced the competing standalone category prompts with a
single unified classification contract after the transcript. It preserved the
assigned Nina task and ownerless established CET work, and prevented
cross-category leakage in the focused case, but still emitted the unaccepted
Textor suggestion as an Action Item. Further prompt iterations were stopped in
accordance with the Gold Standard methodology.
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
remains Open; no phrase-specific deterministic filter was introduced.