Implement Meeting Context V1 and extraction improvements

Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
This commit is contained in:
2026-08-01 16:49:48 +02:00
parent 63075eaca9
commit 06f0e7e651
25 changed files with 1453 additions and 22 deletions
+2
View File
@@ -21,6 +21,8 @@ dist/
# Test # Test
.pytest_cache/ .pytest_cache/
.test-tmp/
.test-tmp-root/
.coverage .coverage
htmlcov/ htmlcov/
+14 -1
View File
@@ -28,6 +28,14 @@
validation and source fact coverage. validation and source fact coverage.
- Semantic Consolidator V0 benchmark report and consolidated extraction JSON. - Semantic Consolidator V0 benchmark report and consolidated extraction JSON.
- Gold regression scenario for responsibility attribution integrity. - Gold regression scenario for responsibility attribution integrity.
- Meeting Context V1 YAML scaffold, generic template and documentation for
manually maintained meeting metadata.
- Meeting Context V1 loader, validator, deterministic prompt representation,
optional `--meeting-context` extraction CLI integration and extraction
context provenance.
- PyYAML project dependency for Meeting Context YAML loading.
- Focused Gold scenarios for responsibility attribution and explicit position
extraction.
### Changed ### Changed
@@ -48,7 +56,12 @@
views. views.
- Updated decision extraction semantics to include explicit process decisions - Updated decision extraction semantics to include explicit process decisions
and deferrals. and deferrals.
- Simplified extraction prompt assembly around prompt files. - Shared extraction prompt assembly now loads `common.md`, `decisions.md` and
`todos.md`.
- Added explicit action-item responsibility attribution rules and clarified
the decision/todo boundary, including duplicate-classification handling.
- Marked Meeting Context integration with later pipeline stages as planned
while extraction-stage integration is implemented.
### Fixed ### Fixed
+40 -4
View File
@@ -23,9 +23,13 @@ Implemented:
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only `src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
semantic duplicate detection. semantic duplicate detection.
- Prompt loading from `src/meeting_lab/llm/prompts.py`. - Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Meeting Context V1 loading, validation and optional extraction prompt
injection with minimal extraction JSON provenance.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`. - Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and - Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation. gold-test runner validation.
- Meeting Context V1 scaffold and documentation for manually maintained
meeting metadata.
Experimental/prototype: Experimental/prototype:
@@ -38,6 +42,8 @@ Planned:
- Broader semantic consolidation for topic grouping, contradiction handling, - Broader semantic consolidation for topic grouping, contradiction handling,
uncertainty marking and durable/transient separation. uncertainty marking and durable/transient separation.
- Canonical Meeting Knowledge implementation as the semantic source of truth. - Canonical Meeting Knowledge implementation as the semantic source of truth.
- Meeting Context integration with Canonicalizer, Semantic Consolidator,
Canonical Meeting Knowledge and output renderers.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol / - Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers. Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
@@ -59,8 +65,10 @@ src/meeting_lab/
Supporting areas: Supporting areas:
- `docs/`: architecture, pipeline, data models and output-view concepts. - `docs/`: architecture, pipeline, data models and output-view concepts.
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain - `docs/meeting-context.md`: Meeting Context V1 scaffold, fields and future
substantive extraction prompt text in the current tree. integration rules.
- `prompts/`: active extraction prompt files. The shared extraction prompt is
assembled from `common.md`, `decisions.md` and `todos.md`.
- `tests/gold/`: semantic gold tests and prompt-engineering methodology. - `tests/gold/`: semantic gold tests and prompt-engineering methodology.
- `samples/`: sample inputs and generated or experimental artifacts. - `samples/`: sample inputs and generated or experimental artifacts.
- `scripts/`: operational scripts for cleanup and gold-test execution. - `scripts/`: operational scripts for cleanup and gold-test execution.
@@ -111,6 +119,13 @@ Accepted decision semantics:
- A process decision to defer a substantive decision is still a decision. - A process decision to defer a substantive decision is still a decision.
- "No decision was reached" is different from "the group decided to defer the - "No decision was reached" is different from "the group decided to defer the
decision." decision."
- A personal commitment to perform concrete future work is normally a todo,
not a decision, unless the group also establishes a separate binding outcome,
rule, approval, rejection, deferral, selection, process state or
responsibility policy.
- The same proposition should not be duplicated under decisions and todos.
Extract both only when the transcript contains a group-level decision and a
semantically separate resulting action item.
## Responsibility Attribution ## Responsibility Attribution
@@ -118,6 +133,13 @@ Meeting Lab distinguishes mentioned people, speakers, participants,
responsible people, departments, owners and assignees. These concepts must not responsible people, departments, owners and assignees. These concepts must not
be collapsed into one field. be collapsed into one field.
Meeting Context V1 reinforces this distinction by separating actual
participants from mentioned non-participants and by storing aliases, roles and
departments only when they are explicitly supplied as metadata. It must not be
used to infer responsibilities. In the current implementation this context can
be injected into chunk extraction prompts as authoritative metadata, and only
minimal provenance is written to extraction JSON.
A `responsible` or future `owner` / `assignee` value may be recorded only when A `responsible` or future `owner` / `assignee` value may be recorded only when
source evidence explicitly assigns, accepts or confirms responsibility. If the source evidence explicitly assigns, accepts or confirms responsibility. If the
evidence is incomplete or ambiguous, the responsible person remains `null` or evidence is incomplete or ambiguous, the responsible person remains `null` or
@@ -125,6 +147,16 @@ unset and the evidence is preserved. Future schema work may add
`responsibility_status` values such as `explicit`, `accepted`, `proposed` and `responsibility_status` values such as `explicit`, `accepted`, `proposed` and
`unclear`, plus `attribution_evidence`. `unclear`, plus `attribution_evidence`.
Meeting Context may validate identity, role, department and attendance, but it
never establishes responsibility.
Focused Gold coverage now separates these concerns:
- `responsibility_attribution_negative`: tests that discussion, objection and
department proximity do not create an owner.
- `position_explicit_objection`: tests explicit position extraction separately
from responsibility attribution.
Current Prompt Version 2 decision baseline: Current Prompt Version 2 decision baseline:
- `decision_simple`: passing. - `decision_simple`: passing.
@@ -208,10 +240,14 @@ language is requested.
are the intended architecture. are the intended architecture.
- Semantic Consolidator V0 is implemented only for facts-only duplicate - Semantic Consolidator V0 is implemented only for facts-only duplicate
detection. detection.
- Meeting Context V1 is implemented only through chunk extraction; later-stage
integration remains planned.
- Canonical Meeting Knowledge is documented but not implemented. - Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented. - Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts. - Some prompt files remain placeholders; `common.md`, `decisions.md` and
- Gold tests currently emphasize extraction semantics, especially decisions. `todos.md` are active in the shared extraction prompt.
- Gold tests currently emphasize extraction semantics, especially decisions,
todos, responsibility attribution and focused position extraction.
## Next Recommended Engineering Step ## Next Recommended Engineering Step
+21 -4
View File
@@ -108,6 +108,21 @@ und normalisiert Extraktionsobjekte ohne LLM. Semantic Consolidator V0 nutzt
das lokale LLM nur fuer konservative facts-only Duplikaterkennung, erhaelt das lokale LLM nur fuer konservative facts-only Duplikaterkennung, erhaelt
Evidenz und erzeugt noch keine Canonical Meeting Knowledge. Evidenz und erzeugt noch keine Canonical Meeting Knowledge.
Meeting Context V1 ist als manuell gepflegtes YAML-Geruest dokumentiert und
fuer die Chunk-Extraktion implementiert. Die Extraktion kann den Kontext
optional validieren, als autoritative Metadaten in den Prompt aufnehmen und
minimale Kontext-Provenienz im Extraction JSON speichern. Konsolidierung,
Canonical Meeting Knowledge und Rendering sind noch nicht daran angeschlossen.
Der gemeinsame Extraktionsprompt besteht aktuell aus `common.md`,
`decisions.md` und `todos.md`. Die Todo-Regeln verlangen explizite Zuweisung,
Freiwilligenmeldung oder Annahme, bevor eine verantwortliche Person gesetzt
wird. Meeting Context kann Identitaet, Rolle, Abteilung und Anwesenheit
validieren, begruendet aber niemals Verantwortung. Entscheidungen und Todos
werden nicht aus derselben Proposition doppelt extrahiert; beides wird nur
ausgegeben, wenn eine gruppenweite Entscheidung und eine davon getrennte
Folgeaufgabe vorliegen.
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
Knowledge**, also die kanonische semantische Repräsentation eines Meetings und Knowledge**, also die kanonische semantische Repräsentation eines Meetings und
@@ -138,10 +153,12 @@ werden, sofern keine explizite Ausgabesprache angefordert wurde.
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technisches Chunking, Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technisches Chunking,
lokale Chunk-Extraktion, Canonicalizer V1 als deterministische Vorbereitung der lokale Chunk-Extraktion, Meeting Context V1 fuer die Extraktionsstufe,
Extraktionsergebnisse und Semantic Consolidator V0 fuer facts-only Canonicalizer V1 als deterministische Vorbereitung der Extraktionsergebnisse
Duplikaterkennung. Canonical Meeting Knowledge, breitere semantische Synthese und Semantic Consolidator V0 fuer facts-only Duplikaterkennung. Canonical
und finale Output-View-Renderer sind geplante nächste Schritte. Meeting Knowledge, breitere semantische Synthese, Meeting-Context-Integration
in spaetere Stufen und finale Output-View-Renderer sind geplante naechste
Schritte.
Canonicalizer V1 kann aus dem Repository heraus so ausgeführt werden: Canonicalizer V1 kann aus dem Repository heraus so ausgeführt werden:
+4
View File
@@ -193,6 +193,10 @@ Deliverables:
- FFmpeg integration. - FFmpeg integration.
- Meeting metadata capture. - Meeting metadata capture.
- Participant entry. - Participant entry.
- Meeting Context V1 exists as a manually maintained YAML structure with
validation, optional extraction prompt integration and extraction provenance;
future work should add GUI entry and conservative integration with later
pipeline stages.
- GUI. - GUI.
- Stable deployment process. - Stable deployment process.
- Export workflows. - Export workflows.
+19
View File
@@ -188,6 +188,24 @@ Each extractor has exactly one task and one prompt.
--- ---
## Meeting Context
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
metadata such as title, language, participants, aliases, departments,
abbreviations and known entities.
It is documented in `docs/meeting-context.md` and templated at
`samples/templates/meeting_context.template.yaml`. It is implemented for
loading, validation and optional injection into chunk extraction prompts.
Extraction results record only minimal context provenance. It is not yet
connected to consolidation, Canonical Meeting Knowledge or output rendering.
The context can help prevent non-participants from being interpreted as
attendees and can normalize known aliases for extraction. It must not infer
roles, departments, responsibilities or decisions.
---
## consolidation/ ## consolidation/
Planned area for canonicalization and consolidation. Planned area for canonicalization and consolidation.
@@ -283,6 +301,7 @@ Implemented:
- Transcript normalization - Transcript normalization
- Technical chunk generation - Technical chunk generation
- Experimental LLM-based information extraction - Experimental LLM-based information extraction
- Meeting Context V1 loading, validation and extraction prompt integration
- Canonicalizer V1 deterministic extraction canonicalization - Canonicalizer V1 deterministic extraction canonicalization
The current extraction step still performs multiple tasks simultaneously. The current extraction step still performs multiple tasks simultaneously.
+48
View File
@@ -218,6 +218,54 @@ Owner remains empty if unknown.
--- ---
# Meeting Context V1
Manually maintained YAML metadata scaffold with an implemented Python loader,
validator and deterministic prompt renderer for chunk extraction.
Top-level structure:
```yaml
schema_version: "1"
meeting: {}
participants: []
mentioned_people: []
organization: {}
known_entities: {}
context_rules: {}
```
Meeting Context separates participants, mentioned people, transcript speakers,
responsible people, roles and departments. It is authoritative only for
explicitly supplied metadata. It must not be used to infer responsibilities,
decisions or commitments.
Template:
- `samples/templates/meeting_context.template.yaml`
Documentation:
- `docs/meeting-context.md`
When `--meeting-context` is supplied to extraction, output JSON receives only
minimal provenance:
```json
{
"context": {
"meeting_id": "...",
"source_file": "...",
"schema_version": "1"
}
}
```
Later Canonicalizer, Semantic Consolidator, Canonical Meeting Knowledge and
renderer integration remains planned.
---
# Topic Result # Topic Result
After extraction, every topic contains the collected information. After extraction, every topic contains the collected information.
+300
View File
@@ -0,0 +1,300 @@
# Meeting Context V1
## Purpose
Meeting Context V1 is a manually maintained YAML file for reliable meeting
metadata. It gives later pipeline stages known names, aliases, organizational
terms and vocabulary without asking an LLM to infer them from a transcript.
The context is authoritative only for metadata that is explicitly supplied in
the file. It must not be used to infer responsibilities, decisions or
commitments.
## Scope
V1 is implemented for loading, validation and optional injection into the
chunk extraction prompt. It is not yet integrated with the Canonicalizer,
Semantic Consolidator, Canonical Meeting Knowledge or output renderers.
Supported metadata:
- meeting title, language, date, objective and notes
- actual participants and participant aliases
- participant role and department when known
- known non-participants mentioned during the meeting
- known departments and aliases
- abbreviations
- relevant products, projects, systems, locations and technical terms
## File Locations
- Generic template: `samples/templates/meeting_context.template.yaml`
- Real-Life Sample context:
`samples/real_live/project_process_meeting/meeting_context.yaml`
- Future meeting-specific contexts should live next to the meeting input data.
## Field Descriptions
`schema_version`: Version of the Meeting Context file shape.
`meeting.title`: Required for future GUI entry. Human-readable meeting title.
`meeting.meeting_id`: Required stable meeting identifier used for extraction
context provenance.
`meeting.language`: Required for future GUI entry. Dominant meeting language,
for example `de` or `en`.
`meeting.date`: Optional ISO date or `null`.
`meeting.objective`: Optional objective entered by the user.
`meeting.notes`: Optional neutral notes about context or scope.
`participants`: Actual meeting attendees. Participant names are required for
future GUI entry.
`participant_id`: Stable identifier. It should not change when display names
or aliases are corrected.
`display_name`: Preferred display name.
`aliases`: Alternative spellings, short forms or Whisper variants.
`role`: Organizational function. Optional and nullable.
`department`: Organizational unit. Optional and nullable.
`attendance_status`: `present` for participants. This distinguishes attendees
from mentioned people.
`mentioned_people`: People discussed or referenced but not present. They are
not participants and must not be treated as speakers.
`organization.name`: Optional organization name.
`organization.departments`: Known departments with stable ids, names and
aliases.
`organization.abbreviations`: Known abbreviation expansions. Empty strings mean
the expansion is not yet confirmed.
`known_entities`: Meeting vocabulary for projects, products, systems,
locations and technical terms. These lists do not imply responsibility.
`context_rules`: Conservative defaults for later integrations.
## Filled Example
```yaml
schema_version: "1"
meeting:
title: "Projektprozess fuer neue Initiativen"
language: "de"
date: "2026-08-01"
objective: "Klaeren, wie der bestehende Projektprozess angepasst wird."
notes: ""
participants:
- participant_id: "martin"
display_name: "Martin"
aliases: ["Martin T."]
role: null
department: null
attendance_status: "present"
notes: null
mentioned_people:
- person_id: "alex"
display_name: "Alex"
aliases: []
role: null
department: null
attendance_status: "not_present"
notes: "Wurde erwaehnt, war aber nicht anwesend."
organization:
name: null
departments:
- id: "pm"
name: "PM"
aliases: ["Projektmanagement"]
abbreviations:
PM: "Projektmanagement"
BD: "Business Development"
MK: ""
GF: ""
known_entities:
projects: []
products: []
systems: []
locations: []
technical_terms: ["Lastenheft"]
context_rules:
participant_list_is_authoritative: true
do_not_infer_roles: true
do_not_infer_departments: true
do_not_infer_responsibilities: true
mentioned_people_are_not_participants: true
```
## Concept Distinctions
Participant: a person who actually attended the meeting.
Mentioned person: a person discussed or referenced but not present.
Speaker: transcript-level attribution, which may be unknown or unreliable.
Responsible person: a person explicitly assigned to or accepting an action
item.
Role: the person's organizational function.
Department: the organizational unit to which the person belongs.
These concepts must never be collapsed automatically. A participant may discuss
a topic outside their own department. Discussing, objecting to or suggesting
work does not establish responsibility.
## Editing Guidance
Enter only objective context that is known independently or confirmed by the
user. Do not guess uncertain names, roles, departments or abbreviation
expansions. Leave unknown values as `null`, an empty string or an empty list.
Use aliases for spelling variants, shortened names and Whisper variants. Keep
stable ids unchanged after a context file has been used in experiments.
Do not copy responsibilities from generated protocols into Meeting Context.
Responsibilities belong to evidence-backed extraction and later action-item
models, not to this metadata file.
## Extraction Integration
The chunk extraction CLI accepts an optional Meeting Context file:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.extraction.extract_chunks \
samples/chunks/chunk_01_normalized.txt \
-o /tmp/chunk_01_extraction.json \
--meeting-context samples/real_live/project_process_meeting/meeting_context.yaml
```
When supplied, the file is validated and rendered as deterministic
authoritative metadata inside the extraction prompt. When omitted, prompt
construction and extraction output remain unchanged.
The shared extraction prompt is assembled from:
- `common.md`
- `decisions.md`
- `todos.md`
Extraction JSON receives only minimal context provenance:
```json
{
"context": {
"meeting_id": "2026-07-27-projektprozess",
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
"schema_version": "1"
}
}
```
Full personal metadata is not copied into every extraction result.
Meeting Context loading uses PyYAML when installed. `PyYAML>=6.0` is declared
as a project dependency.
## Prompt Behavior
Meeting Context validates identity, role, department and attendance, but never
establishes responsibility.
The todo prompt records a named responsible person only when the transcript
explicitly assigns the task, the person volunteers, or the person accepts or
confirms the task. Addressing a person, requesting department input, discussing
department-specific criteria, stating expertise, objecting, suggesting, or
listing Meeting Context role metadata is not enough.
The decision prompt treats a personal commitment to concrete future work as a
todo unless the group separately establishes a binding outcome, rule, approval,
rejection, deferral, selection, process state or responsibility policy. The
same proposition should not be duplicated under decisions and todos; extract
both only when the transcript contains semantically separate propositions.
## Validation Rules
The current validator checks that:
- YAML syntax is valid
- `schema_version` is supported
- `meeting.meeting_id`, `meeting.title` and `meeting.language` are present
- participant ids are unique and non-empty
- mentioned-person ids are unique and non-empty
- participant ids and mentioned-person ids do not collide
- referenced departments exist in `organization.departments`
- `attendance_status` values are valid
- participants are marked `present`
- mentioned people are not marked `present`
Invalid values are not inferred or repaired.
## Gold Tests
Two focused Gold scenarios cover the main responsibility boundary:
- `responsibility_attribution_negative`: discussion, objection and department
proximity must not create a named owner.
- `position_explicit_objection`: explicit objection extraction is tested
separately as a position scenario.
Responsibility attribution and position extraction are intentionally tested
independently so each scenario has unique ground truth.
## Confidentiality
Meeting Context can contain real names, departments, project names and internal
terminology. Treat it as confidential meeting data. Keep private sample
contexts inside the private repository and avoid exposing them in screenshots,
logs or generated reports.
## Future GUI Entry
Future product UI should make these fields easy to enter before processing.
Required GUI fields:
- meeting title
- language
- participant names
Optional GUI fields:
- objective
- aliases
- role
- department
- mentioned non-participants
- abbreviations
- known entities
- notes
## Future Pipeline Integration
Later stages may use Meeting Context to normalize names, recognize aliases,
avoid treating absent mentioned people as speakers and avoid expanding
abbreviations incorrectly. This later-stage integration is still planned.
Integration must remain conservative:
- Meeting Context may supply metadata only.
- It must not infer decisions.
- It must not infer responsibilities.
- It must not override source evidence.
- It must preserve uncertainty when transcript evidence is ambiguous.
+22
View File
@@ -67,6 +67,12 @@ Output View Rendering
Each stage receives a well-defined input and produces a well-defined output. Each stage receives a well-defined input and produces a well-defined output.
Meeting Context V1 exists as a manually maintained YAML metadata scaffold. It
is implemented for validation, optional `--meeting-context` use during chunk
extraction, deterministic prompt injection and minimal extraction JSON
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
Knowledge and renderer integration remains future work.
--- ---
# Stage 1 – Normalization # Stage 1 – Normalization
@@ -284,6 +290,22 @@ Each extractor has:
- one responsibility - one responsibility
- one output schema - one output schema
The current shared extraction prompt is assembled from:
- `common.md`
- `decisions.md`
- `todos.md`
The todo prompt requires explicit assignment, volunteering or acceptance before
recording a named responsible person. Meeting Context may validate identity,
role, department and attendance, but never establishes responsibility.
The decision prompt treats a personal commitment to concrete future work as a
todo unless the group separately establishes a binding outcome, rule, approval,
rejection, deferral, selection, process state or responsibility policy. The
same proposition should not be duplicated under decisions and todos; extract
both only for semantically separate propositions.
## Processing Type ## Processing Type
LLM LLM
+19
View File
@@ -58,6 +58,25 @@ If a statement is only a proposal or suggestion, do not extract it.
If participants discuss something but do not explicitly agree to it, do not If participants discuss something but do not explicitly agree to it, do not
extract it. extract it.
A personal commitment to perform concrete future work is normally an action
item, not a decision. Do not extract it as a decision unless the group also
establishes a binding outcome, rule, approval, rejection, deferral, selection,
process state, or responsibility policy beyond the personal work assignment
itself.
Do not duplicate the same proposition under decisions and todos.
A single transcript passage may nevertheless contain both:
- a group-level decision, rule, approval, rejection, deferral or process state
- and a distinct resulting action item
Extract both when they are semantically separate propositions, even if they
occur in the same sentence or evidence passage.
A concrete personal work commitment without a separate group-level outcome
belongs only under todos.
If the transcript describes an existing process, rule, template, document, or If the transcript describes an existing process, rule, template, document, or
workflow, do not extract it unless the participants explicitly adopt or change it workflow, do not extract it unless the participants explicitly adopt or change it
in this meeting. in this meeting.
+16
View File
@@ -0,0 +1,16 @@
Action-item responsibility rule:
A named responsible person may be extracted only when the transcript contains
clear evidence that the person was explicitly assigned the task, explicitly
volunteered, or explicitly accepted/confirmed the task.
The following are not assignments: addressing a person while discussing a
topic; saying that a person or department will have different criteria;
requesting input from a department; stating expertise, competence, or
organizational role; suggestions, expectations, objections, or preferences;
Meeting Context role or department metadata by itself.
When work is clearly needed but no person accepted it, keep the action item
without a responsible person if the schema allows it; otherwise do not create a
named assignment. Meeting Context may validate identity, role, and attendance,
but never establishes responsibility.
+7
View File
@@ -0,0 +1,7 @@
[project]
name = "meeting-lab"
version = "0.1.0"
requires-python = ">=3.11"
dependencies = [
"PyYAML>=6.0",
]
@@ -31,6 +31,9 @@ Original input files:
Reference material: Reference material:
- `meeting_context.yaml`: manually maintained Meeting Context V1 scaffold for
reliable metadata. It can be validated and injected into chunk extraction
with `--meeting-context`; later pipeline stages are not connected yet.
- `reference/human_reference_protocol.md`: not currently included. No exact - `reference/human_reference_protocol.md`: not currently included. No exact
existing human-written protocol file was found in the repository workspace. existing human-written protocol file was found in the repository workspace.
@@ -0,0 +1,135 @@
schema_version: "1"
meeting:
meeting_id: "2026-07-27-projektprozess"
title: "Projektprozess für neue Projektideen in F&E, Marketing und Business Development"
language: "de"
date: 2026-07-27
objective: ""
notes: "Manuell gepflegter Kontext; noch nicht an die Pipeline angeschlossen."
participants:
- participant_id: "martin"
display_name: "Martin"
aliases:
- "Martin Tazl"
- "Herr Tazl"
role: "Leiter F&E und Technik"
department_id: "fe"
attendance_status: "present"
notes: null
- participant_id: "lars"
display_name: "Lars"
aliases:
- "Lars Vollmert"
- "Herr Vollmert"
roles:
- "Leiter EDD"
- "Leiter Product Management"
department_id: "pm"
attendance_status: "present"
notes: null
- participant_id: "malte"
display_name: "Malte"
aliases:
- "Jan-Malte"
- "Herr Schnau"
role: "Product Manager GreenLine"
department_id: "pm"
attendance_status: "present"
notes: null
- participant_id: "bjoern"
display_name: "Björn"
aliases:
- "Björn-Erik"
- "Herr Falkenau"
role: "Leiter Marketing"
department_id: "mk"
attendance_status: "present"
notes: null
mentioned_people:
- person_id: "jovana"
display_name: "Jovana"
aliases:
- "Jovana Husemann"
- "Frau Husemann"
- "Giovanna"
- "Jovanna"
- "Giovana"
role: "Leiterin Business Development"
department_id: "bd"
attendance_status: "not_present"
notes: null
organization:
name: "Naue GmbH & Co. KG"
departments:
- id: "fe"
name: "F&E"
aliases:
- "Forschung und Entwicklung"
- "Research and Development"
- id: "pm"
name: "Product Management"
aliases: []
- id: "mk"
name: "Marketing"
aliases: []
- id: "bd"
name: "Business Development"
aliases: []
abbreviations:
PM: "Product Management"
BD: "Business Development"
MK: "Marketing"
GF: "Geschäftsführung"
known_entities:
projects: []
products:
- "Bentofix"
- "Carbofol"
- "Combigrid"
- "Secugrid"
- "Secutex"
- "SoftRock"
systems: []
locations:
- "Adorf"
- "Bückeburg"
- "Espelkamp"
- "Malaysia"
technical_terms: []
context_rules:
participant_list_is_authoritative: true
do_not_infer_roles: true
do_not_infer_departments: true
do_not_infer_responsibilities: true
do_not_infer_attendance: true
mentioned_people_are_not_participants: true
@@ -0,0 +1,74 @@
# Meeting Context V1 template.
# This file is manually maintained. Enter only objective context that is known
# before or independently of model output.
schema_version: "1"
meeting:
# Human-readable title for the meeting.
title: ""
# Dominant meeting language, for example "de" or "en".
language: "de"
# Optional ISO date. Leave null if unknown.
date: null
# Optional meeting objective. Do not infer it from generated summaries.
objective: ""
# Optional neutral notes about context, scope or source material.
notes: ""
participants:
# participant_id must remain stable across corrections and later runs.
# Use aliases for alternative spellings, short names or Whisper variants.
# attendance_status distinguishes actual participants from mentioned persons.
# role and department may remain null; uncertain values must not be guessed.
- participant_id: ""
display_name: ""
aliases: []
role: null
department: null
attendance_status: "present"
notes: null
mentioned_people:
# People discussed or referenced but not present in the meeting.
# Mentioned people are not speakers and must not become responsible persons
# unless the meeting evidence explicitly assigns or confirms responsibility.
- person_id: ""
display_name: ""
aliases: []
role: null
department: null
attendance_status: "not_present"
notes: null
organization:
name: null
departments:
# Department ids should be stable. aliases capture spelling variants.
- id: ""
name: ""
aliases: []
abbreviations:
# Fill in only abbreviations that are known for this meeting context.
# Leave values empty when expansion is unknown.
PM: ""
BD: ""
MK: ""
GF: ""
known_entities:
# Relevant projects, products, systems, locations and technical terms.
# These lists provide vocabulary only; they must not imply responsibility.
projects: []
products: []
systems: []
locations: []
technical_terms: []
context_rules:
participant_list_is_authoritative: true
do_not_infer_roles: true
do_not_infer_departments: true
do_not_infer_responsibilities: true
mentioned_people_are_not_participants: true
+45 -4
View File
@@ -20,6 +20,11 @@ from typing import Any
import requests import requests
from src.meeting_lab.llm.prompts import build_extraction_prompt from src.meeting_lab.llm.prompts import build_extraction_prompt
from src.meeting_lab.models.meeting_context import (
MeetingContext,
load_meeting_context,
render_meeting_context_for_prompt,
)
DEFAULT_MODEL = "qwen3:8b" DEFAULT_MODEL = "qwen3:8b"
@@ -36,6 +41,8 @@ EXTRACTION_CATEGORIES = (
NORMALIZED_CHUNK_RE = re.compile(r"^(chunk_\d+)_normalized\.txt$") NORMALIZED_CHUNK_RE = re.compile(r"^(chunk_\d+)_normalized\.txt$")
EXTRACTION_TASK_PROMPT_NAMES = ("decisions.md", "todos.md")
OUTPUT_SCHEMA = { OUTPUT_SCHEMA = {
"chunk": { "chunk": {
@@ -138,11 +145,31 @@ def parse_args() -> argparse.Namespace:
default=32768, default=32768,
help="Context window tokens per chunk (default: 32768)", help="Context window tokens per chunk (default: 32768)",
) )
parser.add_argument(
"--meeting-context",
type=Path,
help="Optional Meeting Context V1 YAML file to inject into extraction prompts.",
)
return parser.parse_args() return parser.parse_args()
def build_prompt(source_name: str, transcript: str) -> str: def build_prompt(
return build_extraction_prompt(source_name, transcript, OUTPUT_SCHEMA) source_name: str,
transcript: str,
meeting_context: MeetingContext | None = None,
) -> str:
context_text = (
render_meeting_context_for_prompt(meeting_context)
if meeting_context is not None
else None
)
return build_extraction_prompt(
source_name,
transcript,
OUTPUT_SCHEMA,
meeting_context=context_text,
task_prompt_names=EXTRACTION_TASK_PROMPT_NAMES,
)
def call_ollama( def call_ollama(
@@ -376,12 +403,13 @@ def extract_chunk(
temperature: float, temperature: float,
num_predict: int | None, num_predict: int | None,
num_ctx: int | None, num_ctx: int | None,
) -> dict[str, list[str]]: meeting_context: MeetingContext | None = None,
) -> dict[str, Any]:
transcript = chunk_path.read_text(encoding="utf-8-sig").strip() transcript = chunk_path.read_text(encoding="utf-8-sig").strip()
if not transcript: if not transcript:
raise ValueError(f"The input file is empty: {chunk_path}") raise ValueError(f"The input file is empty: {chunk_path}")
prompt = build_prompt(chunk_path.name, transcript) prompt = build_prompt(chunk_path.name, transcript, meeting_context=meeting_context)
raw_text, _metadata = call_ollama( raw_text, _metadata = call_ollama(
endpoint=endpoint, endpoint=endpoint,
model=model, model=model,
@@ -402,6 +430,8 @@ def extract_chunk(
) from exc ) from exc
extraction = normalize_current_schema(parsed) extraction = normalize_current_schema(parsed)
if meeting_context is not None:
extraction["context"] = meeting_context.provenance()
output_path.parent.mkdir(parents=True, exist_ok=True) output_path.parent.mkdir(parents=True, exist_ok=True)
output_path.write_text( output_path.write_text(
json.dumps(extraction, ensure_ascii=False, indent=2) + "\n", json.dumps(extraction, ensure_ascii=False, indent=2) + "\n",
@@ -419,6 +449,7 @@ def extract_input(
temperature: float, temperature: float,
num_predict: int | None, num_predict: int | None,
num_ctx: int | None, num_ctx: int | None,
meeting_context: MeetingContext | None = None,
) -> list[Path]: ) -> list[Path]:
if input_path.is_file(): if input_path.is_file():
output_path = output or extraction_path_for_chunk(input_path) output_path = output or extraction_path_for_chunk(input_path)
@@ -431,6 +462,7 @@ def extract_input(
temperature, temperature,
num_predict, num_predict,
num_ctx, num_ctx,
meeting_context=meeting_context,
) )
return [output_path] return [output_path]
@@ -459,6 +491,7 @@ def extract_input(
temperature, temperature,
num_predict, num_predict,
num_ctx, num_ctx,
meeting_context=meeting_context,
) )
output_paths.append(output_path) output_paths.append(output_path)
@@ -469,6 +502,11 @@ def main() -> int:
args = parse_args() args = parse_args()
try: try:
meeting_context = (
load_meeting_context(args.meeting_context)
if args.meeting_context is not None
else None
)
output_paths = extract_input( output_paths = extract_input(
input_path=args.input, input_path=args.input,
output=args.output, output=args.output,
@@ -478,6 +516,7 @@ def main() -> int:
temperature=args.temperature, temperature=args.temperature,
num_predict=args.num_predict, num_predict=args.num_predict,
num_ctx=args.num_ctx, num_ctx=args.num_ctx,
meeting_context=meeting_context,
) )
except requests.ConnectionError: except requests.ConnectionError:
print( print(
@@ -496,6 +535,8 @@ def main() -> int:
return 1 return 1
print(f"Input: {args.input}") print(f"Input: {args.input}")
if args.meeting_context is not None:
print(f"Meeting context: {args.meeting_context}")
print(f"Processed chunks: {len(output_paths)}") print(f"Processed chunks: {len(output_paths)}")
print(f"Extraction JSON files: {len(output_paths)}") print(f"Extraction JSON files: {len(output_paths)}")
for output_path in output_paths: for output_path in output_paths:
+2
View File
@@ -30,12 +30,14 @@ def build_extraction_prompt(
source_name: str, source_name: str,
transcript: str, transcript: str,
output_schema: dict[str, Any], output_schema: dict[str, Any],
meeting_context: str | None = None,
task_prompt_names: Iterable[str] = ("decisions.md",), task_prompt_names: Iterable[str] = ("decisions.md",),
prompts_dir: Path = PROMPTS_DIR, prompts_dir: Path = PROMPTS_DIR,
) -> str: ) -> str:
schema_text = json.dumps(output_schema, ensure_ascii=False, indent=2) schema_text = json.dumps(output_schema, ensure_ascii=False, indent=2)
prompt_parts = [ prompt_parts = [
load_prompt("common.md", prompts_dir), load_prompt("common.md", prompts_dir),
meeting_context,
*load_existing_prompts(task_prompt_names, prompts_dir), *load_existing_prompts(task_prompt_names, prompts_dir),
f"""Quelldatei: f"""Quelldatei:
{source_name} {source_name}
+396
View File
@@ -0,0 +1,396 @@
"""Meeting Context V1 loading, validation and prompt rendering."""
from __future__ import annotations
import ast
from dataclasses import dataclass
from pathlib import Path
from typing import Any
SUPPORTED_SCHEMA_VERSIONS = {"1"}
VALID_ATTENDANCE_STATUSES = {"present", "not_present", "absent"}
class MeetingContextValidationError(ValueError):
"""Raised when a Meeting Context file is structurally invalid."""
@dataclass(frozen=True)
class MeetingContext:
data: dict[str, Any]
source_file: Path
@property
def schema_version(self) -> str:
return str(self.data["schema_version"])
@property
def meeting_id(self) -> str:
return str(self.data["meeting"]["meeting_id"])
def provenance(self) -> dict[str, str]:
return {
"meeting_id": self.meeting_id,
"source_file": str(self.source_file),
"schema_version": self.schema_version,
}
def load_meeting_context(path: Path) -> MeetingContext:
loaded = _load_yaml(path)
if not isinstance(loaded, dict):
raise MeetingContextValidationError("Meeting Context must be a YAML object.")
validate_meeting_context(loaded)
return MeetingContext(data=loaded, source_file=path)
def validate_meeting_context(data: dict[str, Any]) -> None:
schema_version = str(data.get("schema_version", "")).strip()
if schema_version not in SUPPORTED_SCHEMA_VERSIONS:
raise MeetingContextValidationError(
f"Unsupported meeting context schema_version: {schema_version!r}."
)
meeting = _require_mapping(data, "meeting")
_require_non_empty_string(meeting, "meeting.meeting_id")
_require_non_empty_string(meeting, "meeting.title")
_require_non_empty_string(meeting, "meeting.language")
organization = _optional_mapping(data.get("organization"), "organization")
departments = _optional_list(organization.get("departments"), "organization.departments")
department_ids = _collect_department_ids(departments)
participants = _optional_list(data.get("participants"), "participants")
mentioned_people = _optional_list(data.get("mentioned_people"), "mentioned_people")
participant_ids = _collect_unique_ids(participants, "participant_id", "participants")
person_ids = _collect_unique_ids(mentioned_people, "person_id", "mentioned_people")
collisions = sorted(participant_ids & person_ids)
if collisions:
raise MeetingContextValidationError(
"Participant IDs and mentioned-person IDs must not collide: "
+ ", ".join(collisions)
)
for index, participant in enumerate(participants):
item_path = f"participants[{index}]"
_validate_attendance(participant, item_path)
if participant.get("attendance_status") != "present":
raise MeetingContextValidationError(
f"{item_path}.attendance_status must be 'present'."
)
_validate_department_reference(participant, item_path, department_ids)
for index, person in enumerate(mentioned_people):
item_path = f"mentioned_people[{index}]"
_validate_attendance(person, item_path)
if person.get("attendance_status") == "present":
raise MeetingContextValidationError(
f"{item_path}.attendance_status must not be 'present'."
)
_validate_department_reference(person, item_path, department_ids)
def render_meeting_context_for_prompt(context: MeetingContext) -> str:
data = context.data
meeting = data["meeting"]
organization = _optional_mapping(data.get("organization"), "organization")
departments_by_id = {
str(department.get("id")): str(department.get("name"))
for department in _optional_list(
organization.get("departments"), "organization.departments"
)
if isinstance(department, dict) and department.get("id") and department.get("name")
}
lines = [
"MEETING CONTEXT V1 (AUTHORITATIVE METADATA)",
"",
"Rules:",
"- The participant list is authoritative.",
"- Mentioned people did not attend this meeting.",
"- Roles and departments must not be inferred or changed.",
"- Discussion of a department does not establish responsibility.",
"- An action item may name a responsible person only when assignment or acceptance is explicit in the transcript.",
"- Objections, suggestions and expertise do not establish ownership.",
"",
"Meeting:",
f"- Title: {_text(meeting.get('title'))}",
f"- Language: {_text(meeting.get('language'))}",
]
objective = _text(meeting.get("objective"))
if objective:
lines.append(f"- Objective: {objective}")
participants = _optional_list(data.get("participants"), "participants")
if participants:
lines.extend(["", "Actual participants:"])
for participant in participants:
lines.append(_render_person_line(participant, "participant_id", departments_by_id))
mentioned_people = _optional_list(data.get("mentioned_people"), "mentioned_people")
if mentioned_people:
lines.extend(["", "Mentioned but absent people:"])
for person in mentioned_people:
lines.append(_render_person_line(person, "person_id", departments_by_id))
abbreviations = _optional_mapping(organization.get("abbreviations"), "organization.abbreviations")
abbreviation_lines = [
f"- {key}: {_text(value)}"
for key, value in sorted(abbreviations.items())
if _text(value)
]
if abbreviation_lines:
lines.extend(["", "Abbreviations:", *abbreviation_lines])
known_entities = _optional_mapping(data.get("known_entities"), "known_entities")
entity_lines = []
for key in sorted(known_entities):
values = [_text(value) for value in _optional_list(known_entities.get(key), key)]
values = [value for value in values if value]
if values:
entity_lines.append(f"- {key}: {', '.join(values)}")
if entity_lines:
lines.extend(["", "Relevant known entities:", *entity_lines])
context_rules = _optional_mapping(data.get("context_rules"), "context_rules")
rule_lines = [
f"- {key}: {str(value).lower() if isinstance(value, bool) else _text(value)}"
for key, value in sorted(context_rules.items())
]
if rule_lines:
lines.extend(["", "Context rules:", *rule_lines])
return "\n".join(lines).strip() + "\n"
def _render_person_line(
person: dict[str, Any],
id_key: str,
departments_by_id: dict[str, str],
) -> str:
parts = [_text(person.get("display_name"))]
aliases = [_text(alias) for alias in _optional_list(person.get("aliases"), "aliases")]
aliases = [alias for alias in aliases if alias]
if aliases:
parts.append(f"aliases: {', '.join(aliases)}")
roles = _roles(person)
if roles:
parts.append(f"roles: {', '.join(roles)}")
department_id = _text(person.get("department_id") or person.get("department"))
if department_id:
department_name = departments_by_id.get(department_id, department_id)
parts.append(f"department: {department_name}")
identifier = _text(person.get(id_key))
if identifier:
parts.append(f"id: {identifier}")
return "- " + "; ".join(part for part in parts if part)
def _roles(person: dict[str, Any]) -> list[str]:
roles = [_text(role) for role in _optional_list(person.get("roles"), "roles")]
role = _text(person.get("role"))
if role:
roles.insert(0, role)
return [role for role in roles if role]
def _load_yaml(path: Path) -> Any:
text = path.read_text(encoding="utf-8-sig")
try:
import yaml # type: ignore[import-not-found]
except ModuleNotFoundError:
return _parse_simple_yaml(text)
return yaml.safe_load(text)
def _collect_department_ids(departments: list[Any]) -> set[str]:
ids: set[str] = set()
for index, department in enumerate(departments):
if not isinstance(department, dict):
raise MeetingContextValidationError(
f"organization.departments[{index}] must be an object."
)
department_id = _text(department.get("id"))
if not department_id:
raise MeetingContextValidationError(
f"organization.departments[{index}].id must be non-empty."
)
if department_id in ids:
raise MeetingContextValidationError(
f"Duplicate department id: {department_id!r}."
)
ids.add(department_id)
return ids
def _collect_unique_ids(items: list[Any], key: str, path: str) -> set[str]:
ids: set[str] = set()
for index, item in enumerate(items):
if not isinstance(item, dict):
raise MeetingContextValidationError(f"{path}[{index}] must be an object.")
identifier = _text(item.get(key))
if not identifier:
raise MeetingContextValidationError(f"{path}[{index}].{key} must be non-empty.")
if identifier in ids:
raise MeetingContextValidationError(f"Duplicate {key}: {identifier!r}.")
ids.add(identifier)
return ids
def _validate_attendance(item: dict[str, Any], path: str) -> None:
status = item.get("attendance_status")
if status not in VALID_ATTENDANCE_STATUSES:
raise MeetingContextValidationError(
f"{path}.attendance_status has invalid value: {status!r}."
)
def _validate_department_reference(
item: dict[str, Any],
path: str,
department_ids: set[str],
) -> None:
department_id = _text(item.get("department_id") or item.get("department"))
if department_id and department_id not in department_ids:
raise MeetingContextValidationError(
f"{path}.department_id references unknown department: {department_id!r}."
)
def _require_mapping(data: dict[str, Any], key: str) -> dict[str, Any]:
value = data.get(key)
if not isinstance(value, dict):
raise MeetingContextValidationError(f"{key} must be an object.")
return value
def _optional_mapping(value: Any, path: str) -> dict[str, Any]:
if value is None:
return {}
if not isinstance(value, dict):
raise MeetingContextValidationError(f"{path} must be an object.")
return value
def _optional_list(value: Any, path: str) -> list[Any]:
if value is None:
return []
if not isinstance(value, list):
raise MeetingContextValidationError(f"{path} must be a list.")
return value
def _require_non_empty_string(data: dict[str, Any], key_path: str) -> None:
key = key_path.split(".")[-1]
if not _text(data.get(key)):
raise MeetingContextValidationError(f"{key_path} must be present and non-empty.")
def _text(value: Any) -> str:
if value is None:
return ""
return str(value).strip()
def _parse_simple_yaml(text: str) -> Any:
lines = []
for raw_line in text.splitlines():
stripped = _strip_yaml_comment(raw_line.rstrip())
if stripped.strip():
lines.append((len(stripped) - len(stripped.lstrip(" ")), stripped.lstrip(" ")))
if not lines:
return None
parsed, index = _parse_yaml_block(lines, 0, lines[0][0])
if index != len(lines):
raise MeetingContextValidationError("Could not parse Meeting Context YAML.")
return parsed
def _parse_yaml_block(
lines: list[tuple[int, str]],
index: int,
indent: int,
) -> tuple[Any, int]:
if lines[index][1].startswith("- "):
result = []
while index < len(lines) and lines[index][0] == indent and lines[index][1].startswith("- "):
content = lines[index][1][2:].strip()
index += 1
if not content:
value, index = _parse_yaml_block(lines, index, lines[index][0])
result.append(value)
continue
if ":" in content:
key, raw_value = content.split(":", 1)
item = {key.strip(): _parse_scalar(raw_value.strip()) if raw_value.strip() else {}}
while index < len(lines) and lines[index][0] > indent:
child_indent, child_content = lines[index]
if child_content.startswith("- "):
break
child_key, child_raw_value = child_content.split(":", 1)
index += 1
child_raw_value = child_raw_value.strip()
if child_raw_value:
item[child_key.strip()] = _parse_scalar(child_raw_value)
elif index < len(lines) and lines[index][0] > child_indent:
item[child_key.strip()], index = _parse_yaml_block(lines, index, lines[index][0])
else:
item[child_key.strip()] = None
result.append(item)
else:
result.append(_parse_scalar(content))
return result, index
result = {}
while index < len(lines) and lines[index][0] == indent and not lines[index][1].startswith("- "):
key, raw_value = lines[index][1].split(":", 1)
index += 1
raw_value = raw_value.strip()
if raw_value:
result[key.strip()] = _parse_scalar(raw_value)
elif index < len(lines) and lines[index][0] > indent:
result[key.strip()], index = _parse_yaml_block(lines, index, lines[index][0])
else:
result[key.strip()] = None
return result, index
def _parse_scalar(value: str) -> Any:
if value in {"null", "Null", "NULL", "~"}:
return None
if value in {"true", "True", "TRUE"}:
return True
if value in {"false", "False", "FALSE"}:
return False
if value in {"[]", "{}"} or (
value.startswith("[") and value.endswith("]")
):
return ast.literal_eval(value)
if (
(value.startswith('"') and value.endswith('"'))
or (value.startswith("'") and value.endswith("'"))
):
return ast.literal_eval(value)
return value
def _strip_yaml_comment(line: str) -> str:
in_single = False
in_double = False
for index, char in enumerate(line):
if char == "'" and not in_double:
in_single = not in_single
elif char == '"' and not in_single:
in_double = not in_double
elif char == "#" and not in_single and not in_double:
return line[:index].rstrip()
return line
@@ -0,0 +1,18 @@
# position_explicit_objection
Tests extraction of one explicit objection as a position.
Scope:
- Extract Tom's stated objection as exactly one position.
- Do not create a todo for Tom.
- Do not derive a decision from the objection.
Exclusions:
- This scenario does not test responsibility attribution for an open owner.
- Responsibility attribution is covered by `responsibility_attribution_negative`.
The transcript gives unique ground truth because Tom explicitly says "I object"
and "My position is", then explicitly refuses responsibility. The group also
states that no decision and no task for Tom were created.
@@ -0,0 +1,14 @@
{
"facts": [],
"decisions": [],
"todos": [],
"questions": [],
"positions": [
{
"speaker": "Tom",
"position": "Tom objects to using one generic intake checklist because generic criteria will not work for analytics pilots and hardware trials.",
"evidence": "Tom: I object to that. My position is that generic criteria will not work for these two types of work."
}
],
"technical": []
}
@@ -0,0 +1,11 @@
Iris: We could use one generic intake checklist for analytics pilots and hardware trials.
Tom: I object to that. My position is that generic criteria will not work for these two types of work.
Iris: Understood. Are you taking responsibility for rewriting the checklist?
Tom: No. I am not taking that on, and I am not proposing an owner. I am only stating my position.
Uma: Then we are not deciding the checklist today.
Iris: Correct. No decision and no task for Tom.
@@ -7,9 +7,14 @@ The scenario includes a Marketing participant who comments critically on
Business Development criteria. No one assigns that participant responsibility Business Development criteria. No one assigns that participant responsibility
for defining the criteria, and the participant does not accept such a task. for defining the criteria, and the participant does not accept such a task.
Expected behavior: Scope:
- Do not assign Business Development criteria to the Marketing participant. - Do not assign Business Development criteria to the Marketing participant.
- Do not reclassify the Marketing participant as Business Development. - Do not reclassify the Marketing participant as Business Development.
- Preserve the objection as a position.
- Keep responsibility open unless explicitly assigned. - Keep responsibility open unless explicitly assigned.
- Preserve the agreed next step to ask Business Development for an owner.
Exclusions:
- This scenario does not test whether the objection is extracted as a position.
- Position extraction is covered by `position_explicit_objection`.
@@ -28,12 +28,6 @@
} }
], ],
"questions": [], "questions": [],
"positions": [ "positions": [],
{
"speaker": "Noah",
"position": "Noah says generic criteria will not work for different Marketing and product contexts.",
"evidence": "Noah: From Marketing, I can tell you that generic criteria will not work. A digital campaign and a physical product launch need different checks."
}
],
"technical": [] "technical": []
} }
+38
View File
@@ -5,12 +5,15 @@ from pathlib import Path
from src.meeting_lab.extraction.extract_chunks import ( from src.meeting_lab.extraction.extract_chunks import (
EXTRACTION_CATEGORIES, EXTRACTION_CATEGORIES,
EXTRACTION_TASK_PROMPT_NAMES,
build_prompt, build_prompt,
extraction_path_for_chunk, extraction_path_for_chunk,
normalize_current_schema, normalize_current_schema,
parse_json_response, parse_json_response,
) )
from src.meeting_lab.models.meeting_context import load_meeting_context
from src.meeting_lab.protocol.build_protocol import build_protocol from src.meeting_lab.protocol.build_protocol import build_protocol
from scripts import run_gold_test
class ExtractionProtocolTests(unittest.TestCase): class ExtractionProtocolTests(unittest.TestCase):
@@ -102,6 +105,41 @@ Final answer:
self.assertIn("Extract each decision as one atomic commitment.", prompt) self.assertIn("Extract each decision as one atomic commitment.", prompt)
self.assertIn("Anna: Agreed.", prompt) self.assertIn("Anna: Agreed.", prompt)
def test_build_prompt_includes_todo_prompt_file(self) -> None:
prompt = build_prompt("transcript.txt", "Nina: I will update it.")
self.assertEqual(EXTRACTION_TASK_PROMPT_NAMES, ("decisions.md", "todos.md"))
self.assertIn("Action-item responsibility rule:", prompt)
self.assertIn(
"A named responsible person may be extracted only when the transcript contains",
prompt,
)
def test_gold_runner_uses_shared_production_prompt_assembly(self) -> None:
self.assertIs(run_gold_test.build_prompt, build_prompt)
def test_no_context_and_meeting_context_prompts_use_same_task_prompt_set(self) -> None:
context = load_meeting_context(
Path("samples/real_live/project_process_meeting/meeting_context.yaml")
)
no_context_prompt = build_prompt("chunk_01_normalized.txt", "Anna: Agreed.")
context_prompt = build_prompt(
"chunk_01_normalized.txt",
"Anna: Agreed.",
meeting_context=context,
)
for prompt in (no_context_prompt, context_prompt):
self.assertIn("You extract decisions from meeting transcript text.", prompt)
self.assertIn("Action-item responsibility rule:", prompt)
def test_prompt_assembly_is_deterministic(self) -> None:
first = build_prompt("transcript.txt", "Nina: I will update it.")
second = build_prompt("transcript.txt", "Nina: I will update it.")
self.assertEqual(first, second)
def test_build_protocol_groups_extraction_items(self) -> None: def test_build_protocol_groups_extraction_items(self) -> None:
with tempfile.TemporaryDirectory() as directory: with tempfile.TemporaryDirectory() as directory:
input_dir = Path(directory) input_dir = Path(directory)
+197
View File
@@ -0,0 +1,197 @@
import copy
import json
import unittest
from pathlib import Path
from unittest.mock import patch
from src.meeting_lab.extraction.extract_chunks import (
EXTRACTION_CATEGORIES,
build_prompt,
extract_chunk,
)
from src.meeting_lab.models.meeting_context import (
MeetingContextValidationError,
load_meeting_context,
render_meeting_context_for_prompt,
validate_meeting_context,
)
CONTEXT_PATH = Path("samples/real_live/project_process_meeting/meeting_context.yaml")
SCRATCH_DIR = Path(".test-tmp")
class MeetingContextTests(unittest.TestCase):
def setUp(self) -> None:
self.context = load_meeting_context(CONTEXT_PATH)
def tearDown(self) -> None:
if not SCRATCH_DIR.exists():
return
for path in SCRATCH_DIR.glob("chunk_0*_*.txt"):
path.unlink()
for path in SCRATCH_DIR.glob("chunk_0*_*.json"):
path.unlink()
def test_valid_context_loading(self) -> None:
self.assertEqual(self.context.schema_version, "1")
self.assertEqual(self.context.meeting_id, "2026-07-27-projektprozess")
self.assertEqual(self.context.data["meeting"]["language"], "de")
def test_duplicate_participant_ids_are_invalid(self) -> None:
data = copy.deepcopy(self.context.data)
data["participants"][1]["participant_id"] = data["participants"][0][
"participant_id"
]
with self.assertRaisesRegex(MeetingContextValidationError, "Duplicate"):
validate_meeting_context(data)
def test_invalid_department_references_are_invalid(self) -> None:
data = copy.deepcopy(self.context.data)
data["participants"][0]["department_id"] = "unknown"
with self.assertRaisesRegex(MeetingContextValidationError, "unknown department"):
validate_meeting_context(data)
def test_participant_and_mentioned_person_id_collision_is_invalid(self) -> None:
data = copy.deepcopy(self.context.data)
data["mentioned_people"][0]["person_id"] = data["participants"][0][
"participant_id"
]
with self.assertRaisesRegex(MeetingContextValidationError, "collide"):
validate_meeting_context(data)
def test_invalid_attendance_status_is_invalid(self) -> None:
data = copy.deepcopy(self.context.data)
data["participants"][0]["attendance_status"] = "remote"
with self.assertRaisesRegex(MeetingContextValidationError, "invalid value"):
validate_meeting_context(data)
def test_prompt_representation_is_deterministic(self) -> None:
first = render_meeting_context_for_prompt(self.context)
second = render_meeting_context_for_prompt(self.context)
self.assertEqual(first, second)
self.assertIn("MEETING CONTEXT V1", first)
self.assertIn("- Language: de", first)
def test_authoritative_rules_appear_in_prompt(self) -> None:
prompt = build_prompt(
"chunk_01_normalized.txt",
"Martin: Wir besprechen Marketing.",
meeting_context=self.context,
)
self.assertIn("The participant list is authoritative.", prompt)
self.assertIn("Mentioned people did not attend this meeting.", prompt)
self.assertIn("Roles and departments must not be inferred or changed.", prompt)
self.assertIn("Discussion of a department does not establish responsibility.", prompt)
self.assertIn(
"An action item may name a responsible person only when assignment or acceptance is explicit",
prompt,
)
self.assertIn("Objections, suggestions and expertise do not establish ownership.", prompt)
def test_extraction_behavior_is_unchanged_without_context(self) -> None:
response = json.dumps(
{
"facts": [],
"decisions": [],
"todos": [],
"open_questions": [],
"positions": [],
"technical_details": [],
}
)
SCRATCH_DIR.mkdir(exist_ok=True)
chunk_path = SCRATCH_DIR / "chunk_01_normalized.txt"
output_path = SCRATCH_DIR / "chunk_01_extraction.json"
chunk_path.write_text("Anna: Keine Entscheidung.", encoding="utf-8")
captured_prompts = []
def fake_call_ollama(**kwargs):
captured_prompts.append(kwargs["prompt"])
return response, {}
with patch(
"src.meeting_lab.extraction.extract_chunks.call_ollama",
side_effect=fake_call_ollama,
):
extraction = extract_chunk(
chunk_path,
output_path,
model="test",
endpoint="http://example.invalid",
timeout=1,
temperature=0.0,
num_predict=None,
num_ctx=None,
)
self.assertEqual(set(extraction), set(EXTRACTION_CATEGORIES))
self.assertNotIn("context", extraction)
self.assertNotIn("MEETING CONTEXT V1", captured_prompts[0])
def test_extraction_result_records_meeting_context_provenance(self) -> None:
response = json.dumps(
{
"facts": [],
"decisions": [],
"todos": [],
"open_questions": [],
"positions": [],
"technical_details": [],
}
)
SCRATCH_DIR.mkdir(exist_ok=True)
chunk_path = SCRATCH_DIR / "chunk_02_normalized.txt"
output_path = SCRATCH_DIR / "chunk_02_extraction.json"
chunk_path.write_text("Martin: Hallo.", encoding="utf-8")
with patch(
"src.meeting_lab.extraction.extract_chunks.call_ollama",
return_value=(response, {}),
):
extraction = extract_chunk(
chunk_path,
output_path,
model="test",
endpoint="http://example.invalid",
timeout=1,
temperature=0.0,
num_predict=None,
num_ctx=None,
meeting_context=self.context,
)
written = json.loads(output_path.read_text(encoding="utf-8"))
self.assertEqual(
extraction["context"]["meeting_id"],
"2026-07-27-projektprozess",
)
self.assertEqual(extraction["context"]["schema_version"], "1")
self.assertEqual(written["context"], extraction["context"])
self.assertNotIn("participants", written["context"])
def test_real_context_keeps_metadata_separate_from_assignment(self) -> None:
prompt_context = render_meeting_context_for_prompt(self.context)
self.assertIn("Björn", prompt_context)
self.assertIn("department: Marketing", prompt_context)
self.assertIn("Mentioned but absent people:", prompt_context)
self.assertIn("Jovana", prompt_context)
self.assertIn("Discussion of a department does not establish responsibility.", prompt_context)
self.assertIn("assignment or acceptance is explicit", prompt_context)
self.assertNotIn("responsible: Björn", prompt_context)
self.assertNotIn("responsible: Jovana", prompt_context)
if __name__ == "__main__":
unittest.main()