Implement Meeting Context V1 and extraction improvements

Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
This commit is contained in:
2026-08-01 16:49:48 +02:00
parent 63075eaca9
commit 06f0e7e651
25 changed files with 1453 additions and 22 deletions
+40 -4
View File
@@ -23,9 +23,13 @@ Implemented:
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
semantic duplicate detection.
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Meeting Context V1 loading, validation and optional extraction prompt
injection with minimal extraction JSON provenance.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation.
- Meeting Context V1 scaffold and documentation for manually maintained
meeting metadata.
Experimental/prototype:
@@ -38,6 +42,8 @@ Planned:
- Broader semantic consolidation for topic grouping, contradiction handling,
uncertainty marking and durable/transient separation.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Meeting Context integration with Canonicalizer, Semantic Consolidator,
Canonical Meeting Knowledge and output renderers.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
@@ -59,8 +65,10 @@ src/meeting_lab/
Supporting areas:
- `docs/`: architecture, pipeline, data models and output-view concepts.
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain
substantive extraction prompt text in the current tree.
- `docs/meeting-context.md`: Meeting Context V1 scaffold, fields and future
integration rules.
- `prompts/`: active extraction prompt files. The shared extraction prompt is
assembled from `common.md`, `decisions.md` and `todos.md`.
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
- `samples/`: sample inputs and generated or experimental artifacts.
- `scripts/`: operational scripts for cleanup and gold-test execution.
@@ -111,6 +119,13 @@ Accepted decision semantics:
- A process decision to defer a substantive decision is still a decision.
- "No decision was reached" is different from "the group decided to defer the
decision."
- A personal commitment to perform concrete future work is normally a todo,
not a decision, unless the group also establishes a separate binding outcome,
rule, approval, rejection, deferral, selection, process state or
responsibility policy.
- The same proposition should not be duplicated under decisions and todos.
Extract both only when the transcript contains a group-level decision and a
semantically separate resulting action item.
## Responsibility Attribution
@@ -118,6 +133,13 @@ Meeting Lab distinguishes mentioned people, speakers, participants,
responsible people, departments, owners and assignees. These concepts must not
be collapsed into one field.
Meeting Context V1 reinforces this distinction by separating actual
participants from mentioned non-participants and by storing aliases, roles and
departments only when they are explicitly supplied as metadata. It must not be
used to infer responsibilities. In the current implementation this context can
be injected into chunk extraction prompts as authoritative metadata, and only
minimal provenance is written to extraction JSON.
A `responsible` or future `owner` / `assignee` value may be recorded only when
source evidence explicitly assigns, accepts or confirms responsibility. If the
evidence is incomplete or ambiguous, the responsible person remains `null` or
@@ -125,6 +147,16 @@ unset and the evidence is preserved. Future schema work may add
`responsibility_status` values such as `explicit`, `accepted`, `proposed` and
`unclear`, plus `attribution_evidence`.
Meeting Context may validate identity, role, department and attendance, but it
never establishes responsibility.
Focused Gold coverage now separates these concerns:
- `responsibility_attribution_negative`: tests that discussion, objection and
department proximity do not create an owner.
- `position_explicit_objection`: tests explicit position extraction separately
from responsibility attribution.
Current Prompt Version 2 decision baseline:
- `decision_simple`: passing.
@@ -208,10 +240,14 @@ language is requested.
are the intended architecture.
- Semantic Consolidator V0 is implemented only for facts-only duplicate
detection.
- Meeting Context V1 is implemented only through chunk extraction; later-stage
integration remains planned.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts.
- Gold tests currently emphasize extraction semantics, especially decisions.
- Some prompt files remain placeholders; `common.md`, `decisions.md` and
`todos.md` are active in the shared extraction prompt.
- Gold tests currently emphasize extraction semantics, especially decisions,
todos, responsibility attribution and focused position extraction.
## Next Recommended Engineering Step