Improve semantic classification precision for BUG-015
This commit is contained in:
+120
-1
@@ -724,7 +724,7 @@ Evidence:
|
||||
|
||||
## EXP-0015 - Difficult synthetic meeting
|
||||
|
||||
Status: Running
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
@@ -1231,3 +1231,122 @@ Evidence:
|
||||
- `docs/output-views.md`
|
||||
- `prompts/working_protocol.md`
|
||||
- `tests/gold/responsibility_attribution_negative/`
|
||||
|
||||
## EXP-0024 - Working Protocol V2 contract visibility
|
||||
|
||||
Status: Running
|
||||
|
||||
Date or period: 2026-08-09
|
||||
|
||||
Target:
|
||||
|
||||
BUG-011 renderer-only regression using an already validated Semantic
|
||||
Consolidator artifact.
|
||||
|
||||
Hypothesis:
|
||||
|
||||
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
|
||||
placing the leading prompt contract outside the model's effective evaluated
|
||||
context. The preserved failures all report `prompt_eval_count=16386`, while
|
||||
their outputs either echo trailing JSON or produce an unconstrained generic
|
||||
category summary instead of the requested Working Protocol.
|
||||
|
||||
Iteration 1 change:
|
||||
|
||||
- Project every consolidated item to rendering-relevant semantic fields while
|
||||
retaining every item and its category/text/responsibility/deadline/status
|
||||
information.
|
||||
- Generate the exact structural contract from renderer validator constants and
|
||||
append it after the compact INPUT JSON.
|
||||
- Replace the independently handwritten prompt skeleton with a reference to
|
||||
that authoritative appended contract.
|
||||
- Enforce the existing prompt rule that emitted sections must not be empty.
|
||||
|
||||
This is one renderer-contract prompt iteration. It does not change extraction,
|
||||
canonicalization, semantic consolidation or responsibility semantics.
|
||||
|
||||
Validation before LLM run:
|
||||
|
||||
- 15 focused renderer tests pass.
|
||||
- Tests cover contract generation, compact input projection, valid and invalid
|
||||
headings, missing topic sections, empty sections, wrapper cleanup, malformed
|
||||
Markdown and final-file write gating.
|
||||
|
||||
Decision:
|
||||
|
||||
Iteration 1 passed structural validation and wrote `working_protocol.md`, but
|
||||
the quality sanity check found that the model omitted most of the ten supplied
|
||||
decisions, two open questions and several action items. The structurally valid
|
||||
result therefore was not accepted as BUG-011 verification.
|
||||
|
||||
Iteration 2 change:
|
||||
|
||||
- Add input-derived hidden coverage markers for every decision, action item and
|
||||
open question.
|
||||
- Require every priority item exactly once in its matching section.
|
||||
- Validate missing, duplicate, unknown and wrong-section markers
|
||||
deterministically.
|
||||
- Keep facts and technical details condensable as background.
|
||||
|
||||
This is the second single prompt iteration. It responds to the concrete
|
||||
omission failure observed in Iteration 1 without changing upstream semantics or
|
||||
inventing renderer content.
|
||||
|
||||
Iteration 2 pre-run validation:
|
||||
|
||||
- 18 focused renderer tests pass, including exact required-item coverage and
|
||||
wrong-section rejection.
|
||||
|
||||
Decision:
|
||||
|
||||
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget
|
||||
after emitting all decisions and most action items. Adaptive renderer budgeting
|
||||
resolved to 8,192 tokens for this input. The final run stopped normally after
|
||||
3,709 evaluated output tokens.
|
||||
|
||||
The final renderer-only regression passed strict validation and wrote
|
||||
`working_protocol.md`. Exact coverage was 10/10 decisions, 36/36 renderable
|
||||
action items and 23/23 open questions, each once in its matching section. One
|
||||
structurally empty action item whose task, responsible, deadline and evidence
|
||||
were all null was recorded and excluded rather than fabricated. Optional
|
||||
background markers were accepted only when they referred to real projected
|
||||
input items.
|
||||
|
||||
Accept the compact renderer input, validator-derived trailing contract,
|
||||
priority-item coverage markers, empty-section validation and adaptive renderer
|
||||
output sizing as the BUG-011 baseline. This establishes structural reliability
|
||||
and priority-item coverage, not complete protocol prose quality.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `prompts/working_protocol.md`
|
||||
- `src/meeting_lab/protocol/render_working_protocol.py`
|
||||
- `tests/test_render_working_protocol.py`
|
||||
- `samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/`
|
||||
- `samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/`
|
||||
|
||||
## EXP-0025 — BUG-015 classification precision
|
||||
|
||||
Date: 2026-08-09
|
||||
|
||||
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
|
||||
|
||||
Model/configuration: `qwen3.5:9B`, temperature 0, `num_ctx=32768`.
|
||||
|
||||
Tests were created before prompt changes. A Decision-only evidence threshold
|
||||
kept the explicit Dr. Schlummer rejection and omitted an option and preference.
|
||||
Adding Action and Open Question definitions improved several negatives but was
|
||||
not stable: the model alternately promoted an unaccepted Textor suggestion or
|
||||
moved rejected candidates into Open Questions. Moving the standalone category
|
||||
prompts after the transcript made the partial Decision schema dominate and
|
||||
misclassified true Action Items as Decisions in two consecutive runs.
|
||||
|
||||
The final iteration replaced the competing standalone category prompts with a
|
||||
single unified classification contract after the transcript. It preserved the
|
||||
assigned Nina task and ownerless established CET work, and prevented
|
||||
cross-category leakage in the focused case, but still emitted the unaccepted
|
||||
Textor suggestion as an Action Item. Further prompt iterations were stopped in
|
||||
accordance with the Gold Standard methodology.
|
||||
|
||||
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
|
||||
remains Open; no phrase-specific deterministic filter was introduced.
|
||||
|
||||
Reference in New Issue
Block a user