Files
meeting-lab/docs/meeting-context.md
T
admin 06f0e7e651 Implement Meeting Context V1 and extraction improvements
Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
2026-08-01 16:49:48 +02:00

9.1 KiB

Meeting Context V1

Purpose

Meeting Context V1 is a manually maintained YAML file for reliable meeting metadata. It gives later pipeline stages known names, aliases, organizational terms and vocabulary without asking an LLM to infer them from a transcript.

The context is authoritative only for metadata that is explicitly supplied in the file. It must not be used to infer responsibilities, decisions or commitments.

Scope

V1 is implemented for loading, validation and optional injection into the chunk extraction prompt. It is not yet integrated with the Canonicalizer, Semantic Consolidator, Canonical Meeting Knowledge or output renderers.

Supported metadata:

  • meeting title, language, date, objective and notes
  • actual participants and participant aliases
  • participant role and department when known
  • known non-participants mentioned during the meeting
  • known departments and aliases
  • abbreviations
  • relevant products, projects, systems, locations and technical terms

File Locations

  • Generic template: samples/templates/meeting_context.template.yaml
  • Real-Life Sample context: samples/real_live/project_process_meeting/meeting_context.yaml
  • Future meeting-specific contexts should live next to the meeting input data.

Field Descriptions

schema_version: Version of the Meeting Context file shape.

meeting.title: Required for future GUI entry. Human-readable meeting title.

meeting.meeting_id: Required stable meeting identifier used for extraction context provenance.

meeting.language: Required for future GUI entry. Dominant meeting language, for example de or en.

meeting.date: Optional ISO date or null.

meeting.objective: Optional objective entered by the user.

meeting.notes: Optional neutral notes about context or scope.

participants: Actual meeting attendees. Participant names are required for future GUI entry.

participant_id: Stable identifier. It should not change when display names or aliases are corrected.

display_name: Preferred display name.

aliases: Alternative spellings, short forms or Whisper variants.

role: Organizational function. Optional and nullable.

department: Organizational unit. Optional and nullable.

attendance_status: present for participants. This distinguishes attendees from mentioned people.

mentioned_people: People discussed or referenced but not present. They are not participants and must not be treated as speakers.

organization.name: Optional organization name.

organization.departments: Known departments with stable ids, names and aliases.

organization.abbreviations: Known abbreviation expansions. Empty strings mean the expansion is not yet confirmed.

known_entities: Meeting vocabulary for projects, products, systems, locations and technical terms. These lists do not imply responsibility.

context_rules: Conservative defaults for later integrations.

Filled Example

schema_version: "1"

meeting:
  title: "Projektprozess fuer neue Initiativen"
  language: "de"
  date: "2026-08-01"
  objective: "Klaeren, wie der bestehende Projektprozess angepasst wird."
  notes: ""

participants:
  - participant_id: "martin"
    display_name: "Martin"
    aliases: ["Martin T."]
    role: null
    department: null
    attendance_status: "present"
    notes: null

mentioned_people:
  - person_id: "alex"
    display_name: "Alex"
    aliases: []
    role: null
    department: null
    attendance_status: "not_present"
    notes: "Wurde erwaehnt, war aber nicht anwesend."

organization:
  name: null
  departments:
    - id: "pm"
      name: "PM"
      aliases: ["Projektmanagement"]
  abbreviations:
    PM: "Projektmanagement"
    BD: "Business Development"
    MK: ""
    GF: ""

known_entities:
  projects: []
  products: []
  systems: []
  locations: []
  technical_terms: ["Lastenheft"]

context_rules:
  participant_list_is_authoritative: true
  do_not_infer_roles: true
  do_not_infer_departments: true
  do_not_infer_responsibilities: true
  mentioned_people_are_not_participants: true

Concept Distinctions

Participant: a person who actually attended the meeting.

Mentioned person: a person discussed or referenced but not present.

Speaker: transcript-level attribution, which may be unknown or unreliable.

Responsible person: a person explicitly assigned to or accepting an action item.

Role: the person's organizational function.

Department: the organizational unit to which the person belongs.

These concepts must never be collapsed automatically. A participant may discuss a topic outside their own department. Discussing, objecting to or suggesting work does not establish responsibility.

Editing Guidance

Enter only objective context that is known independently or confirmed by the user. Do not guess uncertain names, roles, departments or abbreviation expansions. Leave unknown values as null, an empty string or an empty list.

Use aliases for spelling variants, shortened names and Whisper variants. Keep stable ids unchanged after a context file has been used in experiments.

Do not copy responsibilities from generated protocols into Meeting Context. Responsibilities belong to evidence-backed extraction and later action-item models, not to this metadata file.

Extraction Integration

The chunk extraction CLI accepts an optional Meeting Context file:

PYTHONPATH=src .venv/bin/python -m meeting_lab.extraction.extract_chunks \
  samples/chunks/chunk_01_normalized.txt \
  -o /tmp/chunk_01_extraction.json \
  --meeting-context samples/real_live/project_process_meeting/meeting_context.yaml

When supplied, the file is validated and rendered as deterministic authoritative metadata inside the extraction prompt. When omitted, prompt construction and extraction output remain unchanged.

The shared extraction prompt is assembled from:

  • common.md
  • decisions.md
  • todos.md

Extraction JSON receives only minimal context provenance:

{
  "context": {
    "meeting_id": "2026-07-27-projektprozess",
    "source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
    "schema_version": "1"
  }
}

Full personal metadata is not copied into every extraction result.

Meeting Context loading uses PyYAML when installed. PyYAML>=6.0 is declared as a project dependency.

Prompt Behavior

Meeting Context validates identity, role, department and attendance, but never establishes responsibility.

The todo prompt records a named responsible person only when the transcript explicitly assigns the task, the person volunteers, or the person accepts or confirms the task. Addressing a person, requesting department input, discussing department-specific criteria, stating expertise, objecting, suggesting, or listing Meeting Context role metadata is not enough.

The decision prompt treats a personal commitment to concrete future work as a todo unless the group separately establishes a binding outcome, rule, approval, rejection, deferral, selection, process state or responsibility policy. The same proposition should not be duplicated under decisions and todos; extract both only when the transcript contains semantically separate propositions.

Validation Rules

The current validator checks that:

  • YAML syntax is valid
  • schema_version is supported
  • meeting.meeting_id, meeting.title and meeting.language are present
  • participant ids are unique and non-empty
  • mentioned-person ids are unique and non-empty
  • participant ids and mentioned-person ids do not collide
  • referenced departments exist in organization.departments
  • attendance_status values are valid
  • participants are marked present
  • mentioned people are not marked present

Invalid values are not inferred or repaired.

Gold Tests

Two focused Gold scenarios cover the main responsibility boundary:

  • responsibility_attribution_negative: discussion, objection and department proximity must not create a named owner.
  • position_explicit_objection: explicit objection extraction is tested separately as a position scenario.

Responsibility attribution and position extraction are intentionally tested independently so each scenario has unique ground truth.

Confidentiality

Meeting Context can contain real names, departments, project names and internal terminology. Treat it as confidential meeting data. Keep private sample contexts inside the private repository and avoid exposing them in screenshots, logs or generated reports.

Future GUI Entry

Future product UI should make these fields easy to enter before processing.

Required GUI fields:

  • meeting title
  • language
  • participant names

Optional GUI fields:

  • objective
  • aliases
  • role
  • department
  • mentioned non-participants
  • abbreviations
  • known entities
  • notes

Future Pipeline Integration

Later stages may use Meeting Context to normalize names, recognize aliases, avoid treating absent mentioned people as speakers and avoid expanding abbreviations incorrectly. This later-stage integration is still planned.

Integration must remain conservative:

  • Meeting Context may supply metadata only.
  • It must not infer decisions.
  • It must not infer responsibilities.
  • It must not override source evidence.
  • It must preserve uncertainty when transcript evidence is ambiguous.