Improve semantic classification precision for BUG-015
This commit is contained in:
@@ -375,6 +375,20 @@ extraction results into `meeting_protocol.md` for technical validation. The
|
||||
planned architecture separates Canonical Meeting Knowledge from the final Output
|
||||
Views documented in `output-views.md`.
|
||||
|
||||
## Extraction classification contract
|
||||
|
||||
Decision, Action Item and Open Question extraction shares one evidence-oriented
|
||||
classification contract. It is placed after the transcript so it remains the
|
||||
final classification instruction in the single multi-category extraction call.
|
||||
Decisions require a settled outcome; Action Items require established work;
|
||||
Open Questions require a concrete unresolved need. Unsupported candidates must
|
||||
not be moved into another category.
|
||||
|
||||
Action existence and responsibility attribution are separate checks. A valid
|
||||
Action Item may have no known owner, while a named owner requires explicit
|
||||
assignment, volunteering or acceptance. These are semantic LLM classifications;
|
||||
deterministic validation must not guess intent from keywords.
|
||||
|
||||
## Semantic Consolidator failure handling
|
||||
|
||||
Semantic Consolidator V0 preserves every raw model response before parsing.
|
||||
@@ -396,6 +410,28 @@ an instruction not to emit an identical group more than once. Both attempts
|
||||
and the detected repetition metadata are preserved. If the retry also fails,
|
||||
the stage fails normally; it does not make another LLM call.
|
||||
|
||||
## Working Protocol V2 renderer contract
|
||||
|
||||
The renderer deterministically projects consolidated items to the semantic
|
||||
fields required for presentation and omits bulky provenance fields from the
|
||||
LLM request. Every renderable item remains represented; structurally empty
|
||||
items are recorded separately rather than turned into invented prose. The
|
||||
compact renderer input is preserved as an artifact.
|
||||
|
||||
The exact Markdown structure is generated from the same heading constants used
|
||||
by the validator and appended after the renderer input so it remains visible
|
||||
within the evaluated context. Decisions, action items and open questions carry
|
||||
input-derived hidden coverage markers. Strict validation requires every such
|
||||
renderable priority item exactly once in its matching section and rejects
|
||||
missing, duplicate, wrong-section or invented markers. Facts and technical
|
||||
details remain condensable as background.
|
||||
|
||||
Renderer output budgeting is adaptive to required priority content and prompt
|
||||
size, while an explicit `num_predict` override remains authoritative. Raw model
|
||||
output is always preserved, and `working_protocol.md` is written only after
|
||||
strict structure and coverage validation passes. The renderer does not retry
|
||||
automatically.
|
||||
|
||||
---
|
||||
|
||||
# Next Milestone
|
||||
|
||||
+120
-1
@@ -724,7 +724,7 @@ Evidence:
|
||||
|
||||
## EXP-0015 - Difficult synthetic meeting
|
||||
|
||||
Status: Running
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
@@ -1231,3 +1231,122 @@ Evidence:
|
||||
- `docs/output-views.md`
|
||||
- `prompts/working_protocol.md`
|
||||
- `tests/gold/responsibility_attribution_negative/`
|
||||
|
||||
## EXP-0024 - Working Protocol V2 contract visibility
|
||||
|
||||
Status: Running
|
||||
|
||||
Date or period: 2026-08-09
|
||||
|
||||
Target:
|
||||
|
||||
BUG-011 renderer-only regression using an already validated Semantic
|
||||
Consolidator artifact.
|
||||
|
||||
Hypothesis:
|
||||
|
||||
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
|
||||
placing the leading prompt contract outside the model's effective evaluated
|
||||
context. The preserved failures all report `prompt_eval_count=16386`, while
|
||||
their outputs either echo trailing JSON or produce an unconstrained generic
|
||||
category summary instead of the requested Working Protocol.
|
||||
|
||||
Iteration 1 change:
|
||||
|
||||
- Project every consolidated item to rendering-relevant semantic fields while
|
||||
retaining every item and its category/text/responsibility/deadline/status
|
||||
information.
|
||||
- Generate the exact structural contract from renderer validator constants and
|
||||
append it after the compact INPUT JSON.
|
||||
- Replace the independently handwritten prompt skeleton with a reference to
|
||||
that authoritative appended contract.
|
||||
- Enforce the existing prompt rule that emitted sections must not be empty.
|
||||
|
||||
This is one renderer-contract prompt iteration. It does not change extraction,
|
||||
canonicalization, semantic consolidation or responsibility semantics.
|
||||
|
||||
Validation before LLM run:
|
||||
|
||||
- 15 focused renderer tests pass.
|
||||
- Tests cover contract generation, compact input projection, valid and invalid
|
||||
headings, missing topic sections, empty sections, wrapper cleanup, malformed
|
||||
Markdown and final-file write gating.
|
||||
|
||||
Decision:
|
||||
|
||||
Iteration 1 passed structural validation and wrote `working_protocol.md`, but
|
||||
the quality sanity check found that the model omitted most of the ten supplied
|
||||
decisions, two open questions and several action items. The structurally valid
|
||||
result therefore was not accepted as BUG-011 verification.
|
||||
|
||||
Iteration 2 change:
|
||||
|
||||
- Add input-derived hidden coverage markers for every decision, action item and
|
||||
open question.
|
||||
- Require every priority item exactly once in its matching section.
|
||||
- Validate missing, duplicate, unknown and wrong-section markers
|
||||
deterministically.
|
||||
- Keep facts and technical details condensable as background.
|
||||
|
||||
This is the second single prompt iteration. It responds to the concrete
|
||||
omission failure observed in Iteration 1 without changing upstream semantics or
|
||||
inventing renderer content.
|
||||
|
||||
Iteration 2 pre-run validation:
|
||||
|
||||
- 18 focused renderer tests pass, including exact required-item coverage and
|
||||
wrong-section rejection.
|
||||
|
||||
Decision:
|
||||
|
||||
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget
|
||||
after emitting all decisions and most action items. Adaptive renderer budgeting
|
||||
resolved to 8,192 tokens for this input. The final run stopped normally after
|
||||
3,709 evaluated output tokens.
|
||||
|
||||
The final renderer-only regression passed strict validation and wrote
|
||||
`working_protocol.md`. Exact coverage was 10/10 decisions, 36/36 renderable
|
||||
action items and 23/23 open questions, each once in its matching section. One
|
||||
structurally empty action item whose task, responsible, deadline and evidence
|
||||
were all null was recorded and excluded rather than fabricated. Optional
|
||||
background markers were accepted only when they referred to real projected
|
||||
input items.
|
||||
|
||||
Accept the compact renderer input, validator-derived trailing contract,
|
||||
priority-item coverage markers, empty-section validation and adaptive renderer
|
||||
output sizing as the BUG-011 baseline. This establishes structural reliability
|
||||
and priority-item coverage, not complete protocol prose quality.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `prompts/working_protocol.md`
|
||||
- `src/meeting_lab/protocol/render_working_protocol.py`
|
||||
- `tests/test_render_working_protocol.py`
|
||||
- `samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/`
|
||||
- `samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/`
|
||||
|
||||
## EXP-0025 — BUG-015 classification precision
|
||||
|
||||
Date: 2026-08-09
|
||||
|
||||
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
|
||||
|
||||
Model/configuration: `qwen3.5:9B`, temperature 0, `num_ctx=32768`.
|
||||
|
||||
Tests were created before prompt changes. A Decision-only evidence threshold
|
||||
kept the explicit Dr. Schlummer rejection and omitted an option and preference.
|
||||
Adding Action and Open Question definitions improved several negatives but was
|
||||
not stable: the model alternately promoted an unaccepted Textor suggestion or
|
||||
moved rejected candidates into Open Questions. Moving the standalone category
|
||||
prompts after the transcript made the partial Decision schema dominate and
|
||||
misclassified true Action Items as Decisions in two consecutive runs.
|
||||
|
||||
The final iteration replaced the competing standalone category prompts with a
|
||||
single unified classification contract after the transcript. It preserved the
|
||||
assigned Nina task and ownerless established CET work, and prevented
|
||||
cross-category leakage in the focused case, but still emitted the unaccepted
|
||||
Textor suggestion as an Action Item. Further prompt iterations were stopped in
|
||||
accordance with the Gold Standard methodology.
|
||||
|
||||
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
|
||||
remains Open; no phrase-specific deterministic filter was introduced.
|
||||
|
||||
@@ -840,6 +840,51 @@ Verified. The renderer contract is now enforced deterministically. This does
|
||||
not improve semantic quality of the generated prose; it prevents invalid
|
||||
renderer output from being accepted as a final Working Protocol.
|
||||
|
||||
Production-blocker follow-up, 2026-08-09:
|
||||
|
||||
Later preserved renderer failures showed that enforcement alone did not make a
|
||||
valid protocol reliably obtainable. Provenance-heavy consolidated JSON inputs
|
||||
were 199-350 KB, while failed Ollama runs consistently evaluated 16,386 prompt
|
||||
tokens. The leading handwritten renderer contract was therefore effectively
|
||||
lost or underweighted: models echoed trailing JSON or produced generic category
|
||||
summaries. Prompt and validator also duplicated the structure independently,
|
||||
and the validator did not enforce the prompt's no-empty-section rule.
|
||||
|
||||
The renderer now:
|
||||
|
||||
- preserves every renderable item in a compact semantic projection while
|
||||
removing source-reference and original-value bulk
|
||||
- records and excludes structurally empty items instead of inventing content
|
||||
- appends an authoritative contract generated from validator heading constants
|
||||
- rejects empty emitted sections
|
||||
- requires hidden, input-derived coverage markers for every decision, action
|
||||
item and open question exactly once in the matching section
|
||||
- accepts optional background markers only for real projected input items
|
||||
- uses adaptive output sizing instead of the truncating fixed 4,096-token cap
|
||||
- keeps explicit output-budget overrides authoritative and makes no automatic
|
||||
renderer retry
|
||||
|
||||
Renderer-only verification used the validated consolidated input at
|
||||
`/tmp/meeting-lab-bug014-regression/consolidated_extractions.json`. The final
|
||||
run used `qwen3.5:9B`, temperature 0, `num_ctx=32768`, adaptive
|
||||
`num_predict=8192`, and stopped normally with `eval_count=3709` and
|
||||
`done_reason=stop` after 53.069 seconds. Strict validation reported no
|
||||
violations and wrote:
|
||||
|
||||
- `/tmp/meeting-lab-bug011-renderer-final/working_protocol.md`
|
||||
|
||||
Coverage was complete for all renderable priority items: 10 decisions, 36
|
||||
action items and 23 open questions, with no missing or duplicate markers. One
|
||||
upstream action item containing only null task/responsibility/deadline/evidence
|
||||
was recorded as structurally empty and not rendered. The protocol retained
|
||||
substantial technical background and did not add a new named responsibility;
|
||||
the only structured responsible person in the renderer input remained Marleen.
|
||||
|
||||
Status remains Verified for Working Protocol V2 structural validity and
|
||||
priority-item coverage. This does not claim full semantic or editorial protocol
|
||||
quality, topic quality, or resolution of the other renderer-related regression
|
||||
bugs.
|
||||
|
||||
## BUG-012
|
||||
|
||||
ID: BUG-012
|
||||
@@ -1124,3 +1169,36 @@ source-ID occurrences and 113 unique canonical source IDs. There were no
|
||||
missing IDs, unknown IDs or duplicate occurrences. BUG-014 is independent of
|
||||
BUG-013: BUG-013 detects invalid JSON caused by runaway repeated groups,
|
||||
whereas BUG-014 repairs unknown source IDs in parseable model grouping JSON.
|
||||
|
||||
## BUG-015
|
||||
|
||||
ID: BUG-015
|
||||
|
||||
Title: Extraction classification precision for Decisions, Action Items, and Open Questions
|
||||
|
||||
Status: Open
|
||||
|
||||
First observed: Progeo production benchmark and BUG-011 renderer-only output
|
||||
|
||||
The Progeo extraction classified proposals/options as Decisions, suggestions
|
||||
and hypothetical work as Action Items, and uncertainty or discussion fragments
|
||||
as Open Questions. The production prompt combined a weak shared German rule,
|
||||
a standalone Decision prompt, a responsibility-only Action prompt, and no
|
||||
Open-Question definition in one multi-category extraction request.
|
||||
|
||||
Regression fixtures now record explicit positive and negative evidence for all
|
||||
three categories. Extraction uses one unified classification contract after the
|
||||
transcript, with precision-first thresholds, category boundaries, evidence
|
||||
requirements, and the responsibility-attribution invariant. In a seven-chunk
|
||||
Progeo measurement, bare uncertainty in chunk 08 stopped becoming an Open
|
||||
Question and the preference about real plant versus Technikum in chunk 16
|
||||
stopped becoming a Decision. Supported examples including the explicit Dr.
|
||||
Schlummer rejection and established CET work remained detectable.
|
||||
|
||||
BUG-015 is not Verified. With `qwen3.5:9B`, temperature 0, the focused Action
|
||||
fixture still classified the unaccepted suggestion to contact Dirk Textor as
|
||||
an Action Item. Earlier prompt variants also showed category leakage by turning
|
||||
rejected candidates into Open Questions or Decisions. Prompt iterations were
|
||||
stopped under the Gold Standard methodology rather than adding phrase-specific
|
||||
filters. A general semantic classification architecture improvement remains
|
||||
necessary before this bug can be marked Verified.
|
||||
|
||||
Reference in New Issue
Block a user