Files
meeting-lab/docs/meeting-context.md
T
admin 06f0e7e651 Implement Meeting Context V1 and extraction improvements
Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
2026-08-01 16:49:48 +02:00

301 lines
9.1 KiB
Markdown

# Meeting Context V1
## Purpose
Meeting Context V1 is a manually maintained YAML file for reliable meeting
metadata. It gives later pipeline stages known names, aliases, organizational
terms and vocabulary without asking an LLM to infer them from a transcript.
The context is authoritative only for metadata that is explicitly supplied in
the file. It must not be used to infer responsibilities, decisions or
commitments.
## Scope
V1 is implemented for loading, validation and optional injection into the
chunk extraction prompt. It is not yet integrated with the Canonicalizer,
Semantic Consolidator, Canonical Meeting Knowledge or output renderers.
Supported metadata:
- meeting title, language, date, objective and notes
- actual participants and participant aliases
- participant role and department when known
- known non-participants mentioned during the meeting
- known departments and aliases
- abbreviations
- relevant products, projects, systems, locations and technical terms
## File Locations
- Generic template: `samples/templates/meeting_context.template.yaml`
- Real-Life Sample context:
`samples/real_live/project_process_meeting/meeting_context.yaml`
- Future meeting-specific contexts should live next to the meeting input data.
## Field Descriptions
`schema_version`: Version of the Meeting Context file shape.
`meeting.title`: Required for future GUI entry. Human-readable meeting title.
`meeting.meeting_id`: Required stable meeting identifier used for extraction
context provenance.
`meeting.language`: Required for future GUI entry. Dominant meeting language,
for example `de` or `en`.
`meeting.date`: Optional ISO date or `null`.
`meeting.objective`: Optional objective entered by the user.
`meeting.notes`: Optional neutral notes about context or scope.
`participants`: Actual meeting attendees. Participant names are required for
future GUI entry.
`participant_id`: Stable identifier. It should not change when display names
or aliases are corrected.
`display_name`: Preferred display name.
`aliases`: Alternative spellings, short forms or Whisper variants.
`role`: Organizational function. Optional and nullable.
`department`: Organizational unit. Optional and nullable.
`attendance_status`: `present` for participants. This distinguishes attendees
from mentioned people.
`mentioned_people`: People discussed or referenced but not present. They are
not participants and must not be treated as speakers.
`organization.name`: Optional organization name.
`organization.departments`: Known departments with stable ids, names and
aliases.
`organization.abbreviations`: Known abbreviation expansions. Empty strings mean
the expansion is not yet confirmed.
`known_entities`: Meeting vocabulary for projects, products, systems,
locations and technical terms. These lists do not imply responsibility.
`context_rules`: Conservative defaults for later integrations.
## Filled Example
```yaml
schema_version: "1"
meeting:
title: "Projektprozess fuer neue Initiativen"
language: "de"
date: "2026-08-01"
objective: "Klaeren, wie der bestehende Projektprozess angepasst wird."
notes: ""
participants:
- participant_id: "martin"
display_name: "Martin"
aliases: ["Martin T."]
role: null
department: null
attendance_status: "present"
notes: null
mentioned_people:
- person_id: "alex"
display_name: "Alex"
aliases: []
role: null
department: null
attendance_status: "not_present"
notes: "Wurde erwaehnt, war aber nicht anwesend."
organization:
name: null
departments:
- id: "pm"
name: "PM"
aliases: ["Projektmanagement"]
abbreviations:
PM: "Projektmanagement"
BD: "Business Development"
MK: ""
GF: ""
known_entities:
projects: []
products: []
systems: []
locations: []
technical_terms: ["Lastenheft"]
context_rules:
participant_list_is_authoritative: true
do_not_infer_roles: true
do_not_infer_departments: true
do_not_infer_responsibilities: true
mentioned_people_are_not_participants: true
```
## Concept Distinctions
Participant: a person who actually attended the meeting.
Mentioned person: a person discussed or referenced but not present.
Speaker: transcript-level attribution, which may be unknown or unreliable.
Responsible person: a person explicitly assigned to or accepting an action
item.
Role: the person's organizational function.
Department: the organizational unit to which the person belongs.
These concepts must never be collapsed automatically. A participant may discuss
a topic outside their own department. Discussing, objecting to or suggesting
work does not establish responsibility.
## Editing Guidance
Enter only objective context that is known independently or confirmed by the
user. Do not guess uncertain names, roles, departments or abbreviation
expansions. Leave unknown values as `null`, an empty string or an empty list.
Use aliases for spelling variants, shortened names and Whisper variants. Keep
stable ids unchanged after a context file has been used in experiments.
Do not copy responsibilities from generated protocols into Meeting Context.
Responsibilities belong to evidence-backed extraction and later action-item
models, not to this metadata file.
## Extraction Integration
The chunk extraction CLI accepts an optional Meeting Context file:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.extraction.extract_chunks \
samples/chunks/chunk_01_normalized.txt \
-o /tmp/chunk_01_extraction.json \
--meeting-context samples/real_live/project_process_meeting/meeting_context.yaml
```
When supplied, the file is validated and rendered as deterministic
authoritative metadata inside the extraction prompt. When omitted, prompt
construction and extraction output remain unchanged.
The shared extraction prompt is assembled from:
- `common.md`
- `decisions.md`
- `todos.md`
Extraction JSON receives only minimal context provenance:
```json
{
"context": {
"meeting_id": "2026-07-27-projektprozess",
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
"schema_version": "1"
}
}
```
Full personal metadata is not copied into every extraction result.
Meeting Context loading uses PyYAML when installed. `PyYAML>=6.0` is declared
as a project dependency.
## Prompt Behavior
Meeting Context validates identity, role, department and attendance, but never
establishes responsibility.
The todo prompt records a named responsible person only when the transcript
explicitly assigns the task, the person volunteers, or the person accepts or
confirms the task. Addressing a person, requesting department input, discussing
department-specific criteria, stating expertise, objecting, suggesting, or
listing Meeting Context role metadata is not enough.
The decision prompt treats a personal commitment to concrete future work as a
todo unless the group separately establishes a binding outcome, rule, approval,
rejection, deferral, selection, process state or responsibility policy. The
same proposition should not be duplicated under decisions and todos; extract
both only when the transcript contains semantically separate propositions.
## Validation Rules
The current validator checks that:
- YAML syntax is valid
- `schema_version` is supported
- `meeting.meeting_id`, `meeting.title` and `meeting.language` are present
- participant ids are unique and non-empty
- mentioned-person ids are unique and non-empty
- participant ids and mentioned-person ids do not collide
- referenced departments exist in `organization.departments`
- `attendance_status` values are valid
- participants are marked `present`
- mentioned people are not marked `present`
Invalid values are not inferred or repaired.
## Gold Tests
Two focused Gold scenarios cover the main responsibility boundary:
- `responsibility_attribution_negative`: discussion, objection and department
proximity must not create a named owner.
- `position_explicit_objection`: explicit objection extraction is tested
separately as a position scenario.
Responsibility attribution and position extraction are intentionally tested
independently so each scenario has unique ground truth.
## Confidentiality
Meeting Context can contain real names, departments, project names and internal
terminology. Treat it as confidential meeting data. Keep private sample
contexts inside the private repository and avoid exposing them in screenshots,
logs or generated reports.
## Future GUI Entry
Future product UI should make these fields easy to enter before processing.
Required GUI fields:
- meeting title
- language
- participant names
Optional GUI fields:
- objective
- aliases
- role
- department
- mentioned non-participants
- abbreviations
- known entities
- notes
## Future Pipeline Integration
Later stages may use Meeting Context to normalize names, recognize aliases,
avoid treating absent mentioned people as speakers and avoid expanding
abbreviations incorrectly. This later-stage integration is still planned.
Integration must remain conservative:
- Meeting Context may supply metadata only.
- It must not infer decisions.
- It must not infer responsibilities.
- It must not override source evidence.
- It must preserve uncertainty when transcript evidence is ambiguous.