Implement Meeting Context V1 and extraction improvements
Introduce Meeting Context V1 with YAML schema, validation and template. Support optional --meeting-context during chunk extraction. Inject authoritative Meeting Context into extraction prompts. Record Meeting Context provenance in extraction output. Activate todos.md in shared prompt assembly. Strengthen responsibility attribution and decision/todo boundaries. Add focused Gold scenarios and validation tests. Update architecture and pipeline documentation.
This commit is contained in:
@@ -188,6 +188,24 @@ Each extractor has exactly one task and one prompt.
|
||||
|
||||
---
|
||||
|
||||
## Meeting Context
|
||||
|
||||
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
|
||||
metadata such as title, language, participants, aliases, departments,
|
||||
abbreviations and known entities.
|
||||
|
||||
It is documented in `docs/meeting-context.md` and templated at
|
||||
`samples/templates/meeting_context.template.yaml`. It is implemented for
|
||||
loading, validation and optional injection into chunk extraction prompts.
|
||||
Extraction results record only minimal context provenance. It is not yet
|
||||
connected to consolidation, Canonical Meeting Knowledge or output rendering.
|
||||
|
||||
The context can help prevent non-participants from being interpreted as
|
||||
attendees and can normalize known aliases for extraction. It must not infer
|
||||
roles, departments, responsibilities or decisions.
|
||||
|
||||
---
|
||||
|
||||
## consolidation/
|
||||
|
||||
Planned area for canonicalization and consolidation.
|
||||
@@ -283,6 +301,7 @@ Implemented:
|
||||
- Transcript normalization
|
||||
- Technical chunk generation
|
||||
- Experimental LLM-based information extraction
|
||||
- Meeting Context V1 loading, validation and extraction prompt integration
|
||||
- Canonicalizer V1 deterministic extraction canonicalization
|
||||
|
||||
The current extraction step still performs multiple tasks simultaneously.
|
||||
|
||||
@@ -218,6 +218,54 @@ Owner remains empty if unknown.
|
||||
|
||||
---
|
||||
|
||||
# Meeting Context V1
|
||||
|
||||
Manually maintained YAML metadata scaffold with an implemented Python loader,
|
||||
validator and deterministic prompt renderer for chunk extraction.
|
||||
|
||||
Top-level structure:
|
||||
|
||||
```yaml
|
||||
schema_version: "1"
|
||||
meeting: {}
|
||||
participants: []
|
||||
mentioned_people: []
|
||||
organization: {}
|
||||
known_entities: {}
|
||||
context_rules: {}
|
||||
```
|
||||
|
||||
Meeting Context separates participants, mentioned people, transcript speakers,
|
||||
responsible people, roles and departments. It is authoritative only for
|
||||
explicitly supplied metadata. It must not be used to infer responsibilities,
|
||||
decisions or commitments.
|
||||
|
||||
Template:
|
||||
|
||||
- `samples/templates/meeting_context.template.yaml`
|
||||
|
||||
Documentation:
|
||||
|
||||
- `docs/meeting-context.md`
|
||||
|
||||
When `--meeting-context` is supplied to extraction, output JSON receives only
|
||||
minimal provenance:
|
||||
|
||||
```json
|
||||
{
|
||||
"context": {
|
||||
"meeting_id": "...",
|
||||
"source_file": "...",
|
||||
"schema_version": "1"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Later Canonicalizer, Semantic Consolidator, Canonical Meeting Knowledge and
|
||||
renderer integration remains planned.
|
||||
|
||||
---
|
||||
|
||||
# Topic Result
|
||||
|
||||
After extraction, every topic contains the collected information.
|
||||
|
||||
@@ -0,0 +1,300 @@
|
||||
# Meeting Context V1
|
||||
|
||||
## Purpose
|
||||
|
||||
Meeting Context V1 is a manually maintained YAML file for reliable meeting
|
||||
metadata. It gives later pipeline stages known names, aliases, organizational
|
||||
terms and vocabulary without asking an LLM to infer them from a transcript.
|
||||
|
||||
The context is authoritative only for metadata that is explicitly supplied in
|
||||
the file. It must not be used to infer responsibilities, decisions or
|
||||
commitments.
|
||||
|
||||
## Scope
|
||||
|
||||
V1 is implemented for loading, validation and optional injection into the
|
||||
chunk extraction prompt. It is not yet integrated with the Canonicalizer,
|
||||
Semantic Consolidator, Canonical Meeting Knowledge or output renderers.
|
||||
|
||||
Supported metadata:
|
||||
|
||||
- meeting title, language, date, objective and notes
|
||||
- actual participants and participant aliases
|
||||
- participant role and department when known
|
||||
- known non-participants mentioned during the meeting
|
||||
- known departments and aliases
|
||||
- abbreviations
|
||||
- relevant products, projects, systems, locations and technical terms
|
||||
|
||||
## File Locations
|
||||
|
||||
- Generic template: `samples/templates/meeting_context.template.yaml`
|
||||
- Real-Life Sample context:
|
||||
`samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||
- Future meeting-specific contexts should live next to the meeting input data.
|
||||
|
||||
## Field Descriptions
|
||||
|
||||
`schema_version`: Version of the Meeting Context file shape.
|
||||
|
||||
`meeting.title`: Required for future GUI entry. Human-readable meeting title.
|
||||
|
||||
`meeting.meeting_id`: Required stable meeting identifier used for extraction
|
||||
context provenance.
|
||||
|
||||
`meeting.language`: Required for future GUI entry. Dominant meeting language,
|
||||
for example `de` or `en`.
|
||||
|
||||
`meeting.date`: Optional ISO date or `null`.
|
||||
|
||||
`meeting.objective`: Optional objective entered by the user.
|
||||
|
||||
`meeting.notes`: Optional neutral notes about context or scope.
|
||||
|
||||
`participants`: Actual meeting attendees. Participant names are required for
|
||||
future GUI entry.
|
||||
|
||||
`participant_id`: Stable identifier. It should not change when display names
|
||||
or aliases are corrected.
|
||||
|
||||
`display_name`: Preferred display name.
|
||||
|
||||
`aliases`: Alternative spellings, short forms or Whisper variants.
|
||||
|
||||
`role`: Organizational function. Optional and nullable.
|
||||
|
||||
`department`: Organizational unit. Optional and nullable.
|
||||
|
||||
`attendance_status`: `present` for participants. This distinguishes attendees
|
||||
from mentioned people.
|
||||
|
||||
`mentioned_people`: People discussed or referenced but not present. They are
|
||||
not participants and must not be treated as speakers.
|
||||
|
||||
`organization.name`: Optional organization name.
|
||||
|
||||
`organization.departments`: Known departments with stable ids, names and
|
||||
aliases.
|
||||
|
||||
`organization.abbreviations`: Known abbreviation expansions. Empty strings mean
|
||||
the expansion is not yet confirmed.
|
||||
|
||||
`known_entities`: Meeting vocabulary for projects, products, systems,
|
||||
locations and technical terms. These lists do not imply responsibility.
|
||||
|
||||
`context_rules`: Conservative defaults for later integrations.
|
||||
|
||||
## Filled Example
|
||||
|
||||
```yaml
|
||||
schema_version: "1"
|
||||
|
||||
meeting:
|
||||
title: "Projektprozess fuer neue Initiativen"
|
||||
language: "de"
|
||||
date: "2026-08-01"
|
||||
objective: "Klaeren, wie der bestehende Projektprozess angepasst wird."
|
||||
notes: ""
|
||||
|
||||
participants:
|
||||
- participant_id: "martin"
|
||||
display_name: "Martin"
|
||||
aliases: ["Martin T."]
|
||||
role: null
|
||||
department: null
|
||||
attendance_status: "present"
|
||||
notes: null
|
||||
|
||||
mentioned_people:
|
||||
- person_id: "alex"
|
||||
display_name: "Alex"
|
||||
aliases: []
|
||||
role: null
|
||||
department: null
|
||||
attendance_status: "not_present"
|
||||
notes: "Wurde erwaehnt, war aber nicht anwesend."
|
||||
|
||||
organization:
|
||||
name: null
|
||||
departments:
|
||||
- id: "pm"
|
||||
name: "PM"
|
||||
aliases: ["Projektmanagement"]
|
||||
abbreviations:
|
||||
PM: "Projektmanagement"
|
||||
BD: "Business Development"
|
||||
MK: ""
|
||||
GF: ""
|
||||
|
||||
known_entities:
|
||||
projects: []
|
||||
products: []
|
||||
systems: []
|
||||
locations: []
|
||||
technical_terms: ["Lastenheft"]
|
||||
|
||||
context_rules:
|
||||
participant_list_is_authoritative: true
|
||||
do_not_infer_roles: true
|
||||
do_not_infer_departments: true
|
||||
do_not_infer_responsibilities: true
|
||||
mentioned_people_are_not_participants: true
|
||||
```
|
||||
|
||||
## Concept Distinctions
|
||||
|
||||
Participant: a person who actually attended the meeting.
|
||||
|
||||
Mentioned person: a person discussed or referenced but not present.
|
||||
|
||||
Speaker: transcript-level attribution, which may be unknown or unreliable.
|
||||
|
||||
Responsible person: a person explicitly assigned to or accepting an action
|
||||
item.
|
||||
|
||||
Role: the person's organizational function.
|
||||
|
||||
Department: the organizational unit to which the person belongs.
|
||||
|
||||
These concepts must never be collapsed automatically. A participant may discuss
|
||||
a topic outside their own department. Discussing, objecting to or suggesting
|
||||
work does not establish responsibility.
|
||||
|
||||
## Editing Guidance
|
||||
|
||||
Enter only objective context that is known independently or confirmed by the
|
||||
user. Do not guess uncertain names, roles, departments or abbreviation
|
||||
expansions. Leave unknown values as `null`, an empty string or an empty list.
|
||||
|
||||
Use aliases for spelling variants, shortened names and Whisper variants. Keep
|
||||
stable ids unchanged after a context file has been used in experiments.
|
||||
|
||||
Do not copy responsibilities from generated protocols into Meeting Context.
|
||||
Responsibilities belong to evidence-backed extraction and later action-item
|
||||
models, not to this metadata file.
|
||||
|
||||
## Extraction Integration
|
||||
|
||||
The chunk extraction CLI accepts an optional Meeting Context file:
|
||||
|
||||
```text
|
||||
PYTHONPATH=src .venv/bin/python -m meeting_lab.extraction.extract_chunks \
|
||||
samples/chunks/chunk_01_normalized.txt \
|
||||
-o /tmp/chunk_01_extraction.json \
|
||||
--meeting-context samples/real_live/project_process_meeting/meeting_context.yaml
|
||||
```
|
||||
|
||||
When supplied, the file is validated and rendered as deterministic
|
||||
authoritative metadata inside the extraction prompt. When omitted, prompt
|
||||
construction and extraction output remain unchanged.
|
||||
|
||||
The shared extraction prompt is assembled from:
|
||||
|
||||
- `common.md`
|
||||
- `decisions.md`
|
||||
- `todos.md`
|
||||
|
||||
Extraction JSON receives only minimal context provenance:
|
||||
|
||||
```json
|
||||
{
|
||||
"context": {
|
||||
"meeting_id": "2026-07-27-projektprozess",
|
||||
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
|
||||
"schema_version": "1"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Full personal metadata is not copied into every extraction result.
|
||||
|
||||
Meeting Context loading uses PyYAML when installed. `PyYAML>=6.0` is declared
|
||||
as a project dependency.
|
||||
|
||||
## Prompt Behavior
|
||||
|
||||
Meeting Context validates identity, role, department and attendance, but never
|
||||
establishes responsibility.
|
||||
|
||||
The todo prompt records a named responsible person only when the transcript
|
||||
explicitly assigns the task, the person volunteers, or the person accepts or
|
||||
confirms the task. Addressing a person, requesting department input, discussing
|
||||
department-specific criteria, stating expertise, objecting, suggesting, or
|
||||
listing Meeting Context role metadata is not enough.
|
||||
|
||||
The decision prompt treats a personal commitment to concrete future work as a
|
||||
todo unless the group separately establishes a binding outcome, rule, approval,
|
||||
rejection, deferral, selection, process state or responsibility policy. The
|
||||
same proposition should not be duplicated under decisions and todos; extract
|
||||
both only when the transcript contains semantically separate propositions.
|
||||
|
||||
## Validation Rules
|
||||
|
||||
The current validator checks that:
|
||||
|
||||
- YAML syntax is valid
|
||||
- `schema_version` is supported
|
||||
- `meeting.meeting_id`, `meeting.title` and `meeting.language` are present
|
||||
- participant ids are unique and non-empty
|
||||
- mentioned-person ids are unique and non-empty
|
||||
- participant ids and mentioned-person ids do not collide
|
||||
- referenced departments exist in `organization.departments`
|
||||
- `attendance_status` values are valid
|
||||
- participants are marked `present`
|
||||
- mentioned people are not marked `present`
|
||||
|
||||
Invalid values are not inferred or repaired.
|
||||
|
||||
## Gold Tests
|
||||
|
||||
Two focused Gold scenarios cover the main responsibility boundary:
|
||||
|
||||
- `responsibility_attribution_negative`: discussion, objection and department
|
||||
proximity must not create a named owner.
|
||||
- `position_explicit_objection`: explicit objection extraction is tested
|
||||
separately as a position scenario.
|
||||
|
||||
Responsibility attribution and position extraction are intentionally tested
|
||||
independently so each scenario has unique ground truth.
|
||||
|
||||
## Confidentiality
|
||||
|
||||
Meeting Context can contain real names, departments, project names and internal
|
||||
terminology. Treat it as confidential meeting data. Keep private sample
|
||||
contexts inside the private repository and avoid exposing them in screenshots,
|
||||
logs or generated reports.
|
||||
|
||||
## Future GUI Entry
|
||||
|
||||
Future product UI should make these fields easy to enter before processing.
|
||||
|
||||
Required GUI fields:
|
||||
|
||||
- meeting title
|
||||
- language
|
||||
- participant names
|
||||
|
||||
Optional GUI fields:
|
||||
|
||||
- objective
|
||||
- aliases
|
||||
- role
|
||||
- department
|
||||
- mentioned non-participants
|
||||
- abbreviations
|
||||
- known entities
|
||||
- notes
|
||||
|
||||
## Future Pipeline Integration
|
||||
|
||||
Later stages may use Meeting Context to normalize names, recognize aliases,
|
||||
avoid treating absent mentioned people as speakers and avoid expanding
|
||||
abbreviations incorrectly. This later-stage integration is still planned.
|
||||
|
||||
Integration must remain conservative:
|
||||
|
||||
- Meeting Context may supply metadata only.
|
||||
- It must not infer decisions.
|
||||
- It must not infer responsibilities.
|
||||
- It must not override source evidence.
|
||||
- It must preserve uncertainty when transcript evidence is ambiguous.
|
||||
@@ -67,6 +67,12 @@ Output View Rendering
|
||||
|
||||
Each stage receives a well-defined input and produces a well-defined output.
|
||||
|
||||
Meeting Context V1 exists as a manually maintained YAML metadata scaffold. It
|
||||
is implemented for validation, optional `--meeting-context` use during chunk
|
||||
extraction, deterministic prompt injection and minimal extraction JSON
|
||||
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
|
||||
Knowledge and renderer integration remains future work.
|
||||
|
||||
---
|
||||
|
||||
# Stage 1 – Normalization
|
||||
@@ -284,6 +290,22 @@ Each extractor has:
|
||||
- one responsibility
|
||||
- one output schema
|
||||
|
||||
The current shared extraction prompt is assembled from:
|
||||
|
||||
- `common.md`
|
||||
- `decisions.md`
|
||||
- `todos.md`
|
||||
|
||||
The todo prompt requires explicit assignment, volunteering or acceptance before
|
||||
recording a named responsible person. Meeting Context may validate identity,
|
||||
role, department and attendance, but never establishes responsibility.
|
||||
|
||||
The decision prompt treats a personal commitment to concrete future work as a
|
||||
todo unless the group separately establishes a binding outcome, rule, approval,
|
||||
rejection, deferral, selection, process state or responsibility policy. The
|
||||
same proposition should not be duplicated under decisions and todos; extract
|
||||
both only for semantically separate propositions.
|
||||
|
||||
## Processing Type
|
||||
|
||||
LLM
|
||||
|
||||
Reference in New Issue
Block a user