Implement Meeting Context V1 and extraction improvements
Introduce Meeting Context V1 with YAML schema, validation and template. Support optional --meeting-context during chunk extraction. Inject authoritative Meeting Context into extraction prompts. Record Meeting Context provenance in extraction output. Activate todos.md in shared prompt assembly. Strengthen responsibility attribution and decision/todo boundaries. Add focused Gold scenarios and validation tests. Update architecture and pipeline documentation.
This commit is contained in:
@@ -21,6 +21,8 @@ dist/
|
|||||||
|
|
||||||
# Test
|
# Test
|
||||||
.pytest_cache/
|
.pytest_cache/
|
||||||
|
.test-tmp/
|
||||||
|
.test-tmp-root/
|
||||||
.coverage
|
.coverage
|
||||||
htmlcov/
|
htmlcov/
|
||||||
|
|
||||||
|
|||||||
+14
-1
@@ -28,6 +28,14 @@
|
|||||||
validation and source fact coverage.
|
validation and source fact coverage.
|
||||||
- Semantic Consolidator V0 benchmark report and consolidated extraction JSON.
|
- Semantic Consolidator V0 benchmark report and consolidated extraction JSON.
|
||||||
- Gold regression scenario for responsibility attribution integrity.
|
- Gold regression scenario for responsibility attribution integrity.
|
||||||
|
- Meeting Context V1 YAML scaffold, generic template and documentation for
|
||||||
|
manually maintained meeting metadata.
|
||||||
|
- Meeting Context V1 loader, validator, deterministic prompt representation,
|
||||||
|
optional `--meeting-context` extraction CLI integration and extraction
|
||||||
|
context provenance.
|
||||||
|
- PyYAML project dependency for Meeting Context YAML loading.
|
||||||
|
- Focused Gold scenarios for responsibility attribution and explicit position
|
||||||
|
extraction.
|
||||||
|
|
||||||
### Changed
|
### Changed
|
||||||
|
|
||||||
@@ -48,7 +56,12 @@
|
|||||||
views.
|
views.
|
||||||
- Updated decision extraction semantics to include explicit process decisions
|
- Updated decision extraction semantics to include explicit process decisions
|
||||||
and deferrals.
|
and deferrals.
|
||||||
- Simplified extraction prompt assembly around prompt files.
|
- Shared extraction prompt assembly now loads `common.md`, `decisions.md` and
|
||||||
|
`todos.md`.
|
||||||
|
- Added explicit action-item responsibility attribution rules and clarified
|
||||||
|
the decision/todo boundary, including duplicate-classification handling.
|
||||||
|
- Marked Meeting Context integration with later pipeline stages as planned
|
||||||
|
while extraction-stage integration is implemented.
|
||||||
|
|
||||||
### Fixed
|
### Fixed
|
||||||
|
|
||||||
|
|||||||
+40
-4
@@ -23,9 +23,13 @@ Implemented:
|
|||||||
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
|
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
|
||||||
semantic duplicate detection.
|
semantic duplicate detection.
|
||||||
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||||
|
- Meeting Context V1 loading, validation and optional extraction prompt
|
||||||
|
injection with minimal extraction JSON provenance.
|
||||||
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||||
gold-test runner validation.
|
gold-test runner validation.
|
||||||
|
- Meeting Context V1 scaffold and documentation for manually maintained
|
||||||
|
meeting metadata.
|
||||||
|
|
||||||
Experimental/prototype:
|
Experimental/prototype:
|
||||||
|
|
||||||
@@ -38,6 +42,8 @@ Planned:
|
|||||||
- Broader semantic consolidation for topic grouping, contradiction handling,
|
- Broader semantic consolidation for topic grouping, contradiction handling,
|
||||||
uncertainty marking and durable/transient separation.
|
uncertainty marking and durable/transient separation.
|
||||||
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
||||||
|
- Meeting Context integration with Canonicalizer, Semantic Consolidator,
|
||||||
|
Canonical Meeting Knowledge and output renderers.
|
||||||
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
||||||
|
|
||||||
@@ -59,8 +65,10 @@ src/meeting_lab/
|
|||||||
Supporting areas:
|
Supporting areas:
|
||||||
|
|
||||||
- `docs/`: architecture, pipeline, data models and output-view concepts.
|
- `docs/`: architecture, pipeline, data models and output-view concepts.
|
||||||
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain
|
- `docs/meeting-context.md`: Meeting Context V1 scaffold, fields and future
|
||||||
substantive extraction prompt text in the current tree.
|
integration rules.
|
||||||
|
- `prompts/`: active extraction prompt files. The shared extraction prompt is
|
||||||
|
assembled from `common.md`, `decisions.md` and `todos.md`.
|
||||||
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
|
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
|
||||||
- `samples/`: sample inputs and generated or experimental artifacts.
|
- `samples/`: sample inputs and generated or experimental artifacts.
|
||||||
- `scripts/`: operational scripts for cleanup and gold-test execution.
|
- `scripts/`: operational scripts for cleanup and gold-test execution.
|
||||||
@@ -111,6 +119,13 @@ Accepted decision semantics:
|
|||||||
- A process decision to defer a substantive decision is still a decision.
|
- A process decision to defer a substantive decision is still a decision.
|
||||||
- "No decision was reached" is different from "the group decided to defer the
|
- "No decision was reached" is different from "the group decided to defer the
|
||||||
decision."
|
decision."
|
||||||
|
- A personal commitment to perform concrete future work is normally a todo,
|
||||||
|
not a decision, unless the group also establishes a separate binding outcome,
|
||||||
|
rule, approval, rejection, deferral, selection, process state or
|
||||||
|
responsibility policy.
|
||||||
|
- The same proposition should not be duplicated under decisions and todos.
|
||||||
|
Extract both only when the transcript contains a group-level decision and a
|
||||||
|
semantically separate resulting action item.
|
||||||
|
|
||||||
## Responsibility Attribution
|
## Responsibility Attribution
|
||||||
|
|
||||||
@@ -118,6 +133,13 @@ Meeting Lab distinguishes mentioned people, speakers, participants,
|
|||||||
responsible people, departments, owners and assignees. These concepts must not
|
responsible people, departments, owners and assignees. These concepts must not
|
||||||
be collapsed into one field.
|
be collapsed into one field.
|
||||||
|
|
||||||
|
Meeting Context V1 reinforces this distinction by separating actual
|
||||||
|
participants from mentioned non-participants and by storing aliases, roles and
|
||||||
|
departments only when they are explicitly supplied as metadata. It must not be
|
||||||
|
used to infer responsibilities. In the current implementation this context can
|
||||||
|
be injected into chunk extraction prompts as authoritative metadata, and only
|
||||||
|
minimal provenance is written to extraction JSON.
|
||||||
|
|
||||||
A `responsible` or future `owner` / `assignee` value may be recorded only when
|
A `responsible` or future `owner` / `assignee` value may be recorded only when
|
||||||
source evidence explicitly assigns, accepts or confirms responsibility. If the
|
source evidence explicitly assigns, accepts or confirms responsibility. If the
|
||||||
evidence is incomplete or ambiguous, the responsible person remains `null` or
|
evidence is incomplete or ambiguous, the responsible person remains `null` or
|
||||||
@@ -125,6 +147,16 @@ unset and the evidence is preserved. Future schema work may add
|
|||||||
`responsibility_status` values such as `explicit`, `accepted`, `proposed` and
|
`responsibility_status` values such as `explicit`, `accepted`, `proposed` and
|
||||||
`unclear`, plus `attribution_evidence`.
|
`unclear`, plus `attribution_evidence`.
|
||||||
|
|
||||||
|
Meeting Context may validate identity, role, department and attendance, but it
|
||||||
|
never establishes responsibility.
|
||||||
|
|
||||||
|
Focused Gold coverage now separates these concerns:
|
||||||
|
|
||||||
|
- `responsibility_attribution_negative`: tests that discussion, objection and
|
||||||
|
department proximity do not create an owner.
|
||||||
|
- `position_explicit_objection`: tests explicit position extraction separately
|
||||||
|
from responsibility attribution.
|
||||||
|
|
||||||
Current Prompt Version 2 decision baseline:
|
Current Prompt Version 2 decision baseline:
|
||||||
|
|
||||||
- `decision_simple`: passing.
|
- `decision_simple`: passing.
|
||||||
@@ -208,10 +240,14 @@ language is requested.
|
|||||||
are the intended architecture.
|
are the intended architecture.
|
||||||
- Semantic Consolidator V0 is implemented only for facts-only duplicate
|
- Semantic Consolidator V0 is implemented only for facts-only duplicate
|
||||||
detection.
|
detection.
|
||||||
|
- Meeting Context V1 is implemented only through chunk extraction; later-stage
|
||||||
|
integration remains planned.
|
||||||
- Canonical Meeting Knowledge is documented but not implemented.
|
- Canonical Meeting Knowledge is documented but not implemented.
|
||||||
- Final output views are documented but not implemented.
|
- Final output views are documented but not implemented.
|
||||||
- Most prompt files are placeholders except the common and decision prompts.
|
- Some prompt files remain placeholders; `common.md`, `decisions.md` and
|
||||||
- Gold tests currently emphasize extraction semantics, especially decisions.
|
`todos.md` are active in the shared extraction prompt.
|
||||||
|
- Gold tests currently emphasize extraction semantics, especially decisions,
|
||||||
|
todos, responsibility attribution and focused position extraction.
|
||||||
|
|
||||||
## Next Recommended Engineering Step
|
## Next Recommended Engineering Step
|
||||||
|
|
||||||
|
|||||||
@@ -108,6 +108,21 @@ und normalisiert Extraktionsobjekte ohne LLM. Semantic Consolidator V0 nutzt
|
|||||||
das lokale LLM nur fuer konservative facts-only Duplikaterkennung, erhaelt
|
das lokale LLM nur fuer konservative facts-only Duplikaterkennung, erhaelt
|
||||||
Evidenz und erzeugt noch keine Canonical Meeting Knowledge.
|
Evidenz und erzeugt noch keine Canonical Meeting Knowledge.
|
||||||
|
|
||||||
|
Meeting Context V1 ist als manuell gepflegtes YAML-Geruest dokumentiert und
|
||||||
|
fuer die Chunk-Extraktion implementiert. Die Extraktion kann den Kontext
|
||||||
|
optional validieren, als autoritative Metadaten in den Prompt aufnehmen und
|
||||||
|
minimale Kontext-Provenienz im Extraction JSON speichern. Konsolidierung,
|
||||||
|
Canonical Meeting Knowledge und Rendering sind noch nicht daran angeschlossen.
|
||||||
|
|
||||||
|
Der gemeinsame Extraktionsprompt besteht aktuell aus `common.md`,
|
||||||
|
`decisions.md` und `todos.md`. Die Todo-Regeln verlangen explizite Zuweisung,
|
||||||
|
Freiwilligenmeldung oder Annahme, bevor eine verantwortliche Person gesetzt
|
||||||
|
wird. Meeting Context kann Identitaet, Rolle, Abteilung und Anwesenheit
|
||||||
|
validieren, begruendet aber niemals Verantwortung. Entscheidungen und Todos
|
||||||
|
werden nicht aus derselben Proposition doppelt extrahiert; beides wird nur
|
||||||
|
ausgegeben, wenn eine gruppenweite Entscheidung und eine davon getrennte
|
||||||
|
Folgeaufgabe vorliegen.
|
||||||
|
|
||||||
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
|
Das Meeting Lab behandelt "das Protokoll" nicht mehr als ein einzelnes
|
||||||
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
|
Endprodukt. Das konsolidierte Meeting-Wissen ist die **Canonical Meeting
|
||||||
Knowledge**, also die kanonische semantische Repräsentation eines Meetings und
|
Knowledge**, also die kanonische semantische Repräsentation eines Meetings und
|
||||||
@@ -138,10 +153,12 @@ werden, sofern keine explizite Ausgabesprache angefordert wurde.
|
|||||||
|
|
||||||
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
|
Aktuell liegt der Schwerpunkt auf der Entwicklung eines modularen
|
||||||
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technisches Chunking,
|
Diskussionsanalyzers. Implementiert sind Vorverarbeitung, technisches Chunking,
|
||||||
lokale Chunk-Extraktion, Canonicalizer V1 als deterministische Vorbereitung der
|
lokale Chunk-Extraktion, Meeting Context V1 fuer die Extraktionsstufe,
|
||||||
Extraktionsergebnisse und Semantic Consolidator V0 fuer facts-only
|
Canonicalizer V1 als deterministische Vorbereitung der Extraktionsergebnisse
|
||||||
Duplikaterkennung. Canonical Meeting Knowledge, breitere semantische Synthese
|
und Semantic Consolidator V0 fuer facts-only Duplikaterkennung. Canonical
|
||||||
und finale Output-View-Renderer sind geplante nächste Schritte.
|
Meeting Knowledge, breitere semantische Synthese, Meeting-Context-Integration
|
||||||
|
in spaetere Stufen und finale Output-View-Renderer sind geplante naechste
|
||||||
|
Schritte.
|
||||||
|
|
||||||
Canonicalizer V1 kann aus dem Repository heraus so ausgeführt werden:
|
Canonicalizer V1 kann aus dem Repository heraus so ausgeführt werden:
|
||||||
|
|
||||||
|
|||||||
@@ -193,6 +193,10 @@ Deliverables:
|
|||||||
- FFmpeg integration.
|
- FFmpeg integration.
|
||||||
- Meeting metadata capture.
|
- Meeting metadata capture.
|
||||||
- Participant entry.
|
- Participant entry.
|
||||||
|
- Meeting Context V1 exists as a manually maintained YAML structure with
|
||||||
|
validation, optional extraction prompt integration and extraction provenance;
|
||||||
|
future work should add GUI entry and conservative integration with later
|
||||||
|
pipeline stages.
|
||||||
- GUI.
|
- GUI.
|
||||||
- Stable deployment process.
|
- Stable deployment process.
|
||||||
- Export workflows.
|
- Export workflows.
|
||||||
|
|||||||
@@ -188,6 +188,24 @@ Each extractor has exactly one task and one prompt.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
## Meeting Context
|
||||||
|
|
||||||
|
Meeting Context V1 is a manually maintained YAML scaffold for reliable meeting
|
||||||
|
metadata such as title, language, participants, aliases, departments,
|
||||||
|
abbreviations and known entities.
|
||||||
|
|
||||||
|
It is documented in `docs/meeting-context.md` and templated at
|
||||||
|
`samples/templates/meeting_context.template.yaml`. It is implemented for
|
||||||
|
loading, validation and optional injection into chunk extraction prompts.
|
||||||
|
Extraction results record only minimal context provenance. It is not yet
|
||||||
|
connected to consolidation, Canonical Meeting Knowledge or output rendering.
|
||||||
|
|
||||||
|
The context can help prevent non-participants from being interpreted as
|
||||||
|
attendees and can normalize known aliases for extraction. It must not infer
|
||||||
|
roles, departments, responsibilities or decisions.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
## consolidation/
|
## consolidation/
|
||||||
|
|
||||||
Planned area for canonicalization and consolidation.
|
Planned area for canonicalization and consolidation.
|
||||||
@@ -283,6 +301,7 @@ Implemented:
|
|||||||
- Transcript normalization
|
- Transcript normalization
|
||||||
- Technical chunk generation
|
- Technical chunk generation
|
||||||
- Experimental LLM-based information extraction
|
- Experimental LLM-based information extraction
|
||||||
|
- Meeting Context V1 loading, validation and extraction prompt integration
|
||||||
- Canonicalizer V1 deterministic extraction canonicalization
|
- Canonicalizer V1 deterministic extraction canonicalization
|
||||||
|
|
||||||
The current extraction step still performs multiple tasks simultaneously.
|
The current extraction step still performs multiple tasks simultaneously.
|
||||||
|
|||||||
@@ -218,6 +218,54 @@ Owner remains empty if unknown.
|
|||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
# Meeting Context V1
|
||||||
|
|
||||||
|
Manually maintained YAML metadata scaffold with an implemented Python loader,
|
||||||
|
validator and deterministic prompt renderer for chunk extraction.
|
||||||
|
|
||||||
|
Top-level structure:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
schema_version: "1"
|
||||||
|
meeting: {}
|
||||||
|
participants: []
|
||||||
|
mentioned_people: []
|
||||||
|
organization: {}
|
||||||
|
known_entities: {}
|
||||||
|
context_rules: {}
|
||||||
|
```
|
||||||
|
|
||||||
|
Meeting Context separates participants, mentioned people, transcript speakers,
|
||||||
|
responsible people, roles and departments. It is authoritative only for
|
||||||
|
explicitly supplied metadata. It must not be used to infer responsibilities,
|
||||||
|
decisions or commitments.
|
||||||
|
|
||||||
|
Template:
|
||||||
|
|
||||||
|
- `samples/templates/meeting_context.template.yaml`
|
||||||
|
|
||||||
|
Documentation:
|
||||||
|
|
||||||
|
- `docs/meeting-context.md`
|
||||||
|
|
||||||
|
When `--meeting-context` is supplied to extraction, output JSON receives only
|
||||||
|
minimal provenance:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"context": {
|
||||||
|
"meeting_id": "...",
|
||||||
|
"source_file": "...",
|
||||||
|
"schema_version": "1"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Later Canonicalizer, Semantic Consolidator, Canonical Meeting Knowledge and
|
||||||
|
renderer integration remains planned.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
# Topic Result
|
# Topic Result
|
||||||
|
|
||||||
After extraction, every topic contains the collected information.
|
After extraction, every topic contains the collected information.
|
||||||
|
|||||||
@@ -0,0 +1,300 @@
|
|||||||
|
# Meeting Context V1
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Meeting Context V1 is a manually maintained YAML file for reliable meeting
|
||||||
|
metadata. It gives later pipeline stages known names, aliases, organizational
|
||||||
|
terms and vocabulary without asking an LLM to infer them from a transcript.
|
||||||
|
|
||||||
|
The context is authoritative only for metadata that is explicitly supplied in
|
||||||
|
the file. It must not be used to infer responsibilities, decisions or
|
||||||
|
commitments.
|
||||||
|
|
||||||
|
## Scope
|
||||||
|
|
||||||
|
V1 is implemented for loading, validation and optional injection into the
|
||||||
|
chunk extraction prompt. It is not yet integrated with the Canonicalizer,
|
||||||
|
Semantic Consolidator, Canonical Meeting Knowledge or output renderers.
|
||||||
|
|
||||||
|
Supported metadata:
|
||||||
|
|
||||||
|
- meeting title, language, date, objective and notes
|
||||||
|
- actual participants and participant aliases
|
||||||
|
- participant role and department when known
|
||||||
|
- known non-participants mentioned during the meeting
|
||||||
|
- known departments and aliases
|
||||||
|
- abbreviations
|
||||||
|
- relevant products, projects, systems, locations and technical terms
|
||||||
|
|
||||||
|
## File Locations
|
||||||
|
|
||||||
|
- Generic template: `samples/templates/meeting_context.template.yaml`
|
||||||
|
- Real-Life Sample context:
|
||||||
|
`samples/real_live/project_process_meeting/meeting_context.yaml`
|
||||||
|
- Future meeting-specific contexts should live next to the meeting input data.
|
||||||
|
|
||||||
|
## Field Descriptions
|
||||||
|
|
||||||
|
`schema_version`: Version of the Meeting Context file shape.
|
||||||
|
|
||||||
|
`meeting.title`: Required for future GUI entry. Human-readable meeting title.
|
||||||
|
|
||||||
|
`meeting.meeting_id`: Required stable meeting identifier used for extraction
|
||||||
|
context provenance.
|
||||||
|
|
||||||
|
`meeting.language`: Required for future GUI entry. Dominant meeting language,
|
||||||
|
for example `de` or `en`.
|
||||||
|
|
||||||
|
`meeting.date`: Optional ISO date or `null`.
|
||||||
|
|
||||||
|
`meeting.objective`: Optional objective entered by the user.
|
||||||
|
|
||||||
|
`meeting.notes`: Optional neutral notes about context or scope.
|
||||||
|
|
||||||
|
`participants`: Actual meeting attendees. Participant names are required for
|
||||||
|
future GUI entry.
|
||||||
|
|
||||||
|
`participant_id`: Stable identifier. It should not change when display names
|
||||||
|
or aliases are corrected.
|
||||||
|
|
||||||
|
`display_name`: Preferred display name.
|
||||||
|
|
||||||
|
`aliases`: Alternative spellings, short forms or Whisper variants.
|
||||||
|
|
||||||
|
`role`: Organizational function. Optional and nullable.
|
||||||
|
|
||||||
|
`department`: Organizational unit. Optional and nullable.
|
||||||
|
|
||||||
|
`attendance_status`: `present` for participants. This distinguishes attendees
|
||||||
|
from mentioned people.
|
||||||
|
|
||||||
|
`mentioned_people`: People discussed or referenced but not present. They are
|
||||||
|
not participants and must not be treated as speakers.
|
||||||
|
|
||||||
|
`organization.name`: Optional organization name.
|
||||||
|
|
||||||
|
`organization.departments`: Known departments with stable ids, names and
|
||||||
|
aliases.
|
||||||
|
|
||||||
|
`organization.abbreviations`: Known abbreviation expansions. Empty strings mean
|
||||||
|
the expansion is not yet confirmed.
|
||||||
|
|
||||||
|
`known_entities`: Meeting vocabulary for projects, products, systems,
|
||||||
|
locations and technical terms. These lists do not imply responsibility.
|
||||||
|
|
||||||
|
`context_rules`: Conservative defaults for later integrations.
|
||||||
|
|
||||||
|
## Filled Example
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
schema_version: "1"
|
||||||
|
|
||||||
|
meeting:
|
||||||
|
title: "Projektprozess fuer neue Initiativen"
|
||||||
|
language: "de"
|
||||||
|
date: "2026-08-01"
|
||||||
|
objective: "Klaeren, wie der bestehende Projektprozess angepasst wird."
|
||||||
|
notes: ""
|
||||||
|
|
||||||
|
participants:
|
||||||
|
- participant_id: "martin"
|
||||||
|
display_name: "Martin"
|
||||||
|
aliases: ["Martin T."]
|
||||||
|
role: null
|
||||||
|
department: null
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
mentioned_people:
|
||||||
|
- person_id: "alex"
|
||||||
|
display_name: "Alex"
|
||||||
|
aliases: []
|
||||||
|
role: null
|
||||||
|
department: null
|
||||||
|
attendance_status: "not_present"
|
||||||
|
notes: "Wurde erwaehnt, war aber nicht anwesend."
|
||||||
|
|
||||||
|
organization:
|
||||||
|
name: null
|
||||||
|
departments:
|
||||||
|
- id: "pm"
|
||||||
|
name: "PM"
|
||||||
|
aliases: ["Projektmanagement"]
|
||||||
|
abbreviations:
|
||||||
|
PM: "Projektmanagement"
|
||||||
|
BD: "Business Development"
|
||||||
|
MK: ""
|
||||||
|
GF: ""
|
||||||
|
|
||||||
|
known_entities:
|
||||||
|
projects: []
|
||||||
|
products: []
|
||||||
|
systems: []
|
||||||
|
locations: []
|
||||||
|
technical_terms: ["Lastenheft"]
|
||||||
|
|
||||||
|
context_rules:
|
||||||
|
participant_list_is_authoritative: true
|
||||||
|
do_not_infer_roles: true
|
||||||
|
do_not_infer_departments: true
|
||||||
|
do_not_infer_responsibilities: true
|
||||||
|
mentioned_people_are_not_participants: true
|
||||||
|
```
|
||||||
|
|
||||||
|
## Concept Distinctions
|
||||||
|
|
||||||
|
Participant: a person who actually attended the meeting.
|
||||||
|
|
||||||
|
Mentioned person: a person discussed or referenced but not present.
|
||||||
|
|
||||||
|
Speaker: transcript-level attribution, which may be unknown or unreliable.
|
||||||
|
|
||||||
|
Responsible person: a person explicitly assigned to or accepting an action
|
||||||
|
item.
|
||||||
|
|
||||||
|
Role: the person's organizational function.
|
||||||
|
|
||||||
|
Department: the organizational unit to which the person belongs.
|
||||||
|
|
||||||
|
These concepts must never be collapsed automatically. A participant may discuss
|
||||||
|
a topic outside their own department. Discussing, objecting to or suggesting
|
||||||
|
work does not establish responsibility.
|
||||||
|
|
||||||
|
## Editing Guidance
|
||||||
|
|
||||||
|
Enter only objective context that is known independently or confirmed by the
|
||||||
|
user. Do not guess uncertain names, roles, departments or abbreviation
|
||||||
|
expansions. Leave unknown values as `null`, an empty string or an empty list.
|
||||||
|
|
||||||
|
Use aliases for spelling variants, shortened names and Whisper variants. Keep
|
||||||
|
stable ids unchanged after a context file has been used in experiments.
|
||||||
|
|
||||||
|
Do not copy responsibilities from generated protocols into Meeting Context.
|
||||||
|
Responsibilities belong to evidence-backed extraction and later action-item
|
||||||
|
models, not to this metadata file.
|
||||||
|
|
||||||
|
## Extraction Integration
|
||||||
|
|
||||||
|
The chunk extraction CLI accepts an optional Meeting Context file:
|
||||||
|
|
||||||
|
```text
|
||||||
|
PYTHONPATH=src .venv/bin/python -m meeting_lab.extraction.extract_chunks \
|
||||||
|
samples/chunks/chunk_01_normalized.txt \
|
||||||
|
-o /tmp/chunk_01_extraction.json \
|
||||||
|
--meeting-context samples/real_live/project_process_meeting/meeting_context.yaml
|
||||||
|
```
|
||||||
|
|
||||||
|
When supplied, the file is validated and rendered as deterministic
|
||||||
|
authoritative metadata inside the extraction prompt. When omitted, prompt
|
||||||
|
construction and extraction output remain unchanged.
|
||||||
|
|
||||||
|
The shared extraction prompt is assembled from:
|
||||||
|
|
||||||
|
- `common.md`
|
||||||
|
- `decisions.md`
|
||||||
|
- `todos.md`
|
||||||
|
|
||||||
|
Extraction JSON receives only minimal context provenance:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"context": {
|
||||||
|
"meeting_id": "2026-07-27-projektprozess",
|
||||||
|
"source_file": "samples/real_live/project_process_meeting/meeting_context.yaml",
|
||||||
|
"schema_version": "1"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Full personal metadata is not copied into every extraction result.
|
||||||
|
|
||||||
|
Meeting Context loading uses PyYAML when installed. `PyYAML>=6.0` is declared
|
||||||
|
as a project dependency.
|
||||||
|
|
||||||
|
## Prompt Behavior
|
||||||
|
|
||||||
|
Meeting Context validates identity, role, department and attendance, but never
|
||||||
|
establishes responsibility.
|
||||||
|
|
||||||
|
The todo prompt records a named responsible person only when the transcript
|
||||||
|
explicitly assigns the task, the person volunteers, or the person accepts or
|
||||||
|
confirms the task. Addressing a person, requesting department input, discussing
|
||||||
|
department-specific criteria, stating expertise, objecting, suggesting, or
|
||||||
|
listing Meeting Context role metadata is not enough.
|
||||||
|
|
||||||
|
The decision prompt treats a personal commitment to concrete future work as a
|
||||||
|
todo unless the group separately establishes a binding outcome, rule, approval,
|
||||||
|
rejection, deferral, selection, process state or responsibility policy. The
|
||||||
|
same proposition should not be duplicated under decisions and todos; extract
|
||||||
|
both only when the transcript contains semantically separate propositions.
|
||||||
|
|
||||||
|
## Validation Rules
|
||||||
|
|
||||||
|
The current validator checks that:
|
||||||
|
|
||||||
|
- YAML syntax is valid
|
||||||
|
- `schema_version` is supported
|
||||||
|
- `meeting.meeting_id`, `meeting.title` and `meeting.language` are present
|
||||||
|
- participant ids are unique and non-empty
|
||||||
|
- mentioned-person ids are unique and non-empty
|
||||||
|
- participant ids and mentioned-person ids do not collide
|
||||||
|
- referenced departments exist in `organization.departments`
|
||||||
|
- `attendance_status` values are valid
|
||||||
|
- participants are marked `present`
|
||||||
|
- mentioned people are not marked `present`
|
||||||
|
|
||||||
|
Invalid values are not inferred or repaired.
|
||||||
|
|
||||||
|
## Gold Tests
|
||||||
|
|
||||||
|
Two focused Gold scenarios cover the main responsibility boundary:
|
||||||
|
|
||||||
|
- `responsibility_attribution_negative`: discussion, objection and department
|
||||||
|
proximity must not create a named owner.
|
||||||
|
- `position_explicit_objection`: explicit objection extraction is tested
|
||||||
|
separately as a position scenario.
|
||||||
|
|
||||||
|
Responsibility attribution and position extraction are intentionally tested
|
||||||
|
independently so each scenario has unique ground truth.
|
||||||
|
|
||||||
|
## Confidentiality
|
||||||
|
|
||||||
|
Meeting Context can contain real names, departments, project names and internal
|
||||||
|
terminology. Treat it as confidential meeting data. Keep private sample
|
||||||
|
contexts inside the private repository and avoid exposing them in screenshots,
|
||||||
|
logs or generated reports.
|
||||||
|
|
||||||
|
## Future GUI Entry
|
||||||
|
|
||||||
|
Future product UI should make these fields easy to enter before processing.
|
||||||
|
|
||||||
|
Required GUI fields:
|
||||||
|
|
||||||
|
- meeting title
|
||||||
|
- language
|
||||||
|
- participant names
|
||||||
|
|
||||||
|
Optional GUI fields:
|
||||||
|
|
||||||
|
- objective
|
||||||
|
- aliases
|
||||||
|
- role
|
||||||
|
- department
|
||||||
|
- mentioned non-participants
|
||||||
|
- abbreviations
|
||||||
|
- known entities
|
||||||
|
- notes
|
||||||
|
|
||||||
|
## Future Pipeline Integration
|
||||||
|
|
||||||
|
Later stages may use Meeting Context to normalize names, recognize aliases,
|
||||||
|
avoid treating absent mentioned people as speakers and avoid expanding
|
||||||
|
abbreviations incorrectly. This later-stage integration is still planned.
|
||||||
|
|
||||||
|
Integration must remain conservative:
|
||||||
|
|
||||||
|
- Meeting Context may supply metadata only.
|
||||||
|
- It must not infer decisions.
|
||||||
|
- It must not infer responsibilities.
|
||||||
|
- It must not override source evidence.
|
||||||
|
- It must preserve uncertainty when transcript evidence is ambiguous.
|
||||||
@@ -67,6 +67,12 @@ Output View Rendering
|
|||||||
|
|
||||||
Each stage receives a well-defined input and produces a well-defined output.
|
Each stage receives a well-defined input and produces a well-defined output.
|
||||||
|
|
||||||
|
Meeting Context V1 exists as a manually maintained YAML metadata scaffold. It
|
||||||
|
is implemented for validation, optional `--meeting-context` use during chunk
|
||||||
|
extraction, deterministic prompt injection and minimal extraction JSON
|
||||||
|
provenance. Later Canonicalizer, Semantic Consolidator, Canonical Meeting
|
||||||
|
Knowledge and renderer integration remains future work.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# Stage 1 – Normalization
|
# Stage 1 – Normalization
|
||||||
@@ -284,6 +290,22 @@ Each extractor has:
|
|||||||
- one responsibility
|
- one responsibility
|
||||||
- one output schema
|
- one output schema
|
||||||
|
|
||||||
|
The current shared extraction prompt is assembled from:
|
||||||
|
|
||||||
|
- `common.md`
|
||||||
|
- `decisions.md`
|
||||||
|
- `todos.md`
|
||||||
|
|
||||||
|
The todo prompt requires explicit assignment, volunteering or acceptance before
|
||||||
|
recording a named responsible person. Meeting Context may validate identity,
|
||||||
|
role, department and attendance, but never establishes responsibility.
|
||||||
|
|
||||||
|
The decision prompt treats a personal commitment to concrete future work as a
|
||||||
|
todo unless the group separately establishes a binding outcome, rule, approval,
|
||||||
|
rejection, deferral, selection, process state or responsibility policy. The
|
||||||
|
same proposition should not be duplicated under decisions and todos; extract
|
||||||
|
both only for semantically separate propositions.
|
||||||
|
|
||||||
## Processing Type
|
## Processing Type
|
||||||
|
|
||||||
LLM
|
LLM
|
||||||
|
|||||||
@@ -58,6 +58,25 @@ If a statement is only a proposal or suggestion, do not extract it.
|
|||||||
If participants discuss something but do not explicitly agree to it, do not
|
If participants discuss something but do not explicitly agree to it, do not
|
||||||
extract it.
|
extract it.
|
||||||
|
|
||||||
|
A personal commitment to perform concrete future work is normally an action
|
||||||
|
item, not a decision. Do not extract it as a decision unless the group also
|
||||||
|
establishes a binding outcome, rule, approval, rejection, deferral, selection,
|
||||||
|
process state, or responsibility policy beyond the personal work assignment
|
||||||
|
itself.
|
||||||
|
|
||||||
|
Do not duplicate the same proposition under decisions and todos.
|
||||||
|
|
||||||
|
A single transcript passage may nevertheless contain both:
|
||||||
|
|
||||||
|
- a group-level decision, rule, approval, rejection, deferral or process state
|
||||||
|
- and a distinct resulting action item
|
||||||
|
|
||||||
|
Extract both when they are semantically separate propositions, even if they
|
||||||
|
occur in the same sentence or evidence passage.
|
||||||
|
|
||||||
|
A concrete personal work commitment without a separate group-level outcome
|
||||||
|
belongs only under todos.
|
||||||
|
|
||||||
If the transcript describes an existing process, rule, template, document, or
|
If the transcript describes an existing process, rule, template, document, or
|
||||||
workflow, do not extract it unless the participants explicitly adopt or change it
|
workflow, do not extract it unless the participants explicitly adopt or change it
|
||||||
in this meeting.
|
in this meeting.
|
||||||
|
|||||||
@@ -0,0 +1,16 @@
|
|||||||
|
Action-item responsibility rule:
|
||||||
|
|
||||||
|
A named responsible person may be extracted only when the transcript contains
|
||||||
|
clear evidence that the person was explicitly assigned the task, explicitly
|
||||||
|
volunteered, or explicitly accepted/confirmed the task.
|
||||||
|
|
||||||
|
The following are not assignments: addressing a person while discussing a
|
||||||
|
topic; saying that a person or department will have different criteria;
|
||||||
|
requesting input from a department; stating expertise, competence, or
|
||||||
|
organizational role; suggestions, expectations, objections, or preferences;
|
||||||
|
Meeting Context role or department metadata by itself.
|
||||||
|
|
||||||
|
When work is clearly needed but no person accepted it, keep the action item
|
||||||
|
without a responsible person if the schema allows it; otherwise do not create a
|
||||||
|
named assignment. Meeting Context may validate identity, role, and attendance,
|
||||||
|
but never establishes responsibility.
|
||||||
|
|||||||
@@ -0,0 +1,7 @@
|
|||||||
|
[project]
|
||||||
|
name = "meeting-lab"
|
||||||
|
version = "0.1.0"
|
||||||
|
requires-python = ">=3.11"
|
||||||
|
dependencies = [
|
||||||
|
"PyYAML>=6.0",
|
||||||
|
]
|
||||||
|
|||||||
@@ -31,6 +31,9 @@ Original input files:
|
|||||||
|
|
||||||
Reference material:
|
Reference material:
|
||||||
|
|
||||||
|
- `meeting_context.yaml`: manually maintained Meeting Context V1 scaffold for
|
||||||
|
reliable metadata. It can be validated and injected into chunk extraction
|
||||||
|
with `--meeting-context`; later pipeline stages are not connected yet.
|
||||||
- `reference/human_reference_protocol.md`: not currently included. No exact
|
- `reference/human_reference_protocol.md`: not currently included. No exact
|
||||||
existing human-written protocol file was found in the repository workspace.
|
existing human-written protocol file was found in the repository workspace.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,135 @@
|
|||||||
|
schema_version: "1"
|
||||||
|
|
||||||
|
meeting:
|
||||||
|
meeting_id: "2026-07-27-projektprozess"
|
||||||
|
title: "Projektprozess für neue Projektideen in F&E, Marketing und Business Development"
|
||||||
|
language: "de"
|
||||||
|
date: 2026-07-27
|
||||||
|
objective: ""
|
||||||
|
notes: "Manuell gepflegter Kontext; noch nicht an die Pipeline angeschlossen."
|
||||||
|
|
||||||
|
participants:
|
||||||
|
|
||||||
|
- participant_id: "martin"
|
||||||
|
display_name: "Martin"
|
||||||
|
aliases:
|
||||||
|
- "Martin Tazl"
|
||||||
|
- "Herr Tazl"
|
||||||
|
role: "Leiter F&E und Technik"
|
||||||
|
department_id: "fe"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "lars"
|
||||||
|
display_name: "Lars"
|
||||||
|
aliases:
|
||||||
|
- "Lars Vollmert"
|
||||||
|
- "Herr Vollmert"
|
||||||
|
roles:
|
||||||
|
- "Leiter EDD"
|
||||||
|
- "Leiter Product Management"
|
||||||
|
department_id: "pm"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "malte"
|
||||||
|
display_name: "Malte"
|
||||||
|
aliases:
|
||||||
|
- "Jan-Malte"
|
||||||
|
- "Herr Schnau"
|
||||||
|
role: "Product Manager GreenLine"
|
||||||
|
department_id: "pm"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
- participant_id: "bjoern"
|
||||||
|
display_name: "Björn"
|
||||||
|
aliases:
|
||||||
|
- "Björn-Erik"
|
||||||
|
- "Herr Falkenau"
|
||||||
|
role: "Leiter Marketing"
|
||||||
|
department_id: "mk"
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
mentioned_people:
|
||||||
|
|
||||||
|
- person_id: "jovana"
|
||||||
|
display_name: "Jovana"
|
||||||
|
aliases:
|
||||||
|
- "Jovana Husemann"
|
||||||
|
- "Frau Husemann"
|
||||||
|
- "Giovanna"
|
||||||
|
- "Jovanna"
|
||||||
|
- "Giovana"
|
||||||
|
role: "Leiterin Business Development"
|
||||||
|
department_id: "bd"
|
||||||
|
attendance_status: "not_present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
organization:
|
||||||
|
|
||||||
|
name: "Naue GmbH & Co. KG"
|
||||||
|
|
||||||
|
departments:
|
||||||
|
|
||||||
|
- id: "fe"
|
||||||
|
name: "F&E"
|
||||||
|
aliases:
|
||||||
|
- "Forschung und Entwicklung"
|
||||||
|
- "Research and Development"
|
||||||
|
|
||||||
|
- id: "pm"
|
||||||
|
name: "Product Management"
|
||||||
|
aliases: []
|
||||||
|
|
||||||
|
- id: "mk"
|
||||||
|
name: "Marketing"
|
||||||
|
aliases: []
|
||||||
|
|
||||||
|
- id: "bd"
|
||||||
|
name: "Business Development"
|
||||||
|
aliases: []
|
||||||
|
|
||||||
|
abbreviations:
|
||||||
|
|
||||||
|
PM: "Product Management"
|
||||||
|
BD: "Business Development"
|
||||||
|
MK: "Marketing"
|
||||||
|
GF: "Geschäftsführung"
|
||||||
|
|
||||||
|
known_entities:
|
||||||
|
|
||||||
|
projects: []
|
||||||
|
|
||||||
|
products:
|
||||||
|
- "Bentofix"
|
||||||
|
- "Carbofol"
|
||||||
|
- "Combigrid"
|
||||||
|
- "Secugrid"
|
||||||
|
- "Secutex"
|
||||||
|
- "SoftRock"
|
||||||
|
|
||||||
|
systems: []
|
||||||
|
|
||||||
|
locations:
|
||||||
|
- "Adorf"
|
||||||
|
- "Bückeburg"
|
||||||
|
- "Espelkamp"
|
||||||
|
- "Malaysia"
|
||||||
|
|
||||||
|
technical_terms: []
|
||||||
|
|
||||||
|
context_rules:
|
||||||
|
|
||||||
|
participant_list_is_authoritative: true
|
||||||
|
|
||||||
|
do_not_infer_roles: true
|
||||||
|
|
||||||
|
do_not_infer_departments: true
|
||||||
|
|
||||||
|
do_not_infer_responsibilities: true
|
||||||
|
|
||||||
|
do_not_infer_attendance: true
|
||||||
|
|
||||||
|
mentioned_people_are_not_participants: true
|
||||||
@@ -0,0 +1,74 @@
|
|||||||
|
# Meeting Context V1 template.
|
||||||
|
# This file is manually maintained. Enter only objective context that is known
|
||||||
|
# before or independently of model output.
|
||||||
|
schema_version: "1"
|
||||||
|
|
||||||
|
meeting:
|
||||||
|
# Human-readable title for the meeting.
|
||||||
|
title: ""
|
||||||
|
# Dominant meeting language, for example "de" or "en".
|
||||||
|
language: "de"
|
||||||
|
# Optional ISO date. Leave null if unknown.
|
||||||
|
date: null
|
||||||
|
# Optional meeting objective. Do not infer it from generated summaries.
|
||||||
|
objective: ""
|
||||||
|
# Optional neutral notes about context, scope or source material.
|
||||||
|
notes: ""
|
||||||
|
|
||||||
|
participants:
|
||||||
|
# participant_id must remain stable across corrections and later runs.
|
||||||
|
# Use aliases for alternative spellings, short names or Whisper variants.
|
||||||
|
# attendance_status distinguishes actual participants from mentioned persons.
|
||||||
|
# role and department may remain null; uncertain values must not be guessed.
|
||||||
|
- participant_id: ""
|
||||||
|
display_name: ""
|
||||||
|
aliases: []
|
||||||
|
role: null
|
||||||
|
department: null
|
||||||
|
attendance_status: "present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
mentioned_people:
|
||||||
|
# People discussed or referenced but not present in the meeting.
|
||||||
|
# Mentioned people are not speakers and must not become responsible persons
|
||||||
|
# unless the meeting evidence explicitly assigns or confirms responsibility.
|
||||||
|
- person_id: ""
|
||||||
|
display_name: ""
|
||||||
|
aliases: []
|
||||||
|
role: null
|
||||||
|
department: null
|
||||||
|
attendance_status: "not_present"
|
||||||
|
notes: null
|
||||||
|
|
||||||
|
organization:
|
||||||
|
name: null
|
||||||
|
|
||||||
|
departments:
|
||||||
|
# Department ids should be stable. aliases capture spelling variants.
|
||||||
|
- id: ""
|
||||||
|
name: ""
|
||||||
|
aliases: []
|
||||||
|
|
||||||
|
abbreviations:
|
||||||
|
# Fill in only abbreviations that are known for this meeting context.
|
||||||
|
# Leave values empty when expansion is unknown.
|
||||||
|
PM: ""
|
||||||
|
BD: ""
|
||||||
|
MK: ""
|
||||||
|
GF: ""
|
||||||
|
|
||||||
|
known_entities:
|
||||||
|
# Relevant projects, products, systems, locations and technical terms.
|
||||||
|
# These lists provide vocabulary only; they must not imply responsibility.
|
||||||
|
projects: []
|
||||||
|
products: []
|
||||||
|
systems: []
|
||||||
|
locations: []
|
||||||
|
technical_terms: []
|
||||||
|
|
||||||
|
context_rules:
|
||||||
|
participant_list_is_authoritative: true
|
||||||
|
do_not_infer_roles: true
|
||||||
|
do_not_infer_departments: true
|
||||||
|
do_not_infer_responsibilities: true
|
||||||
|
mentioned_people_are_not_participants: true
|
||||||
@@ -20,6 +20,11 @@ from typing import Any
|
|||||||
import requests
|
import requests
|
||||||
|
|
||||||
from src.meeting_lab.llm.prompts import build_extraction_prompt
|
from src.meeting_lab.llm.prompts import build_extraction_prompt
|
||||||
|
from src.meeting_lab.models.meeting_context import (
|
||||||
|
MeetingContext,
|
||||||
|
load_meeting_context,
|
||||||
|
render_meeting_context_for_prompt,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
DEFAULT_MODEL = "qwen3:8b"
|
DEFAULT_MODEL = "qwen3:8b"
|
||||||
@@ -36,6 +41,8 @@ EXTRACTION_CATEGORIES = (
|
|||||||
|
|
||||||
NORMALIZED_CHUNK_RE = re.compile(r"^(chunk_\d+)_normalized\.txt$")
|
NORMALIZED_CHUNK_RE = re.compile(r"^(chunk_\d+)_normalized\.txt$")
|
||||||
|
|
||||||
|
EXTRACTION_TASK_PROMPT_NAMES = ("decisions.md", "todos.md")
|
||||||
|
|
||||||
|
|
||||||
OUTPUT_SCHEMA = {
|
OUTPUT_SCHEMA = {
|
||||||
"chunk": {
|
"chunk": {
|
||||||
@@ -138,11 +145,31 @@ def parse_args() -> argparse.Namespace:
|
|||||||
default=32768,
|
default=32768,
|
||||||
help="Context window tokens per chunk (default: 32768)",
|
help="Context window tokens per chunk (default: 32768)",
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--meeting-context",
|
||||||
|
type=Path,
|
||||||
|
help="Optional Meeting Context V1 YAML file to inject into extraction prompts.",
|
||||||
|
)
|
||||||
return parser.parse_args()
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
def build_prompt(source_name: str, transcript: str) -> str:
|
def build_prompt(
|
||||||
return build_extraction_prompt(source_name, transcript, OUTPUT_SCHEMA)
|
source_name: str,
|
||||||
|
transcript: str,
|
||||||
|
meeting_context: MeetingContext | None = None,
|
||||||
|
) -> str:
|
||||||
|
context_text = (
|
||||||
|
render_meeting_context_for_prompt(meeting_context)
|
||||||
|
if meeting_context is not None
|
||||||
|
else None
|
||||||
|
)
|
||||||
|
return build_extraction_prompt(
|
||||||
|
source_name,
|
||||||
|
transcript,
|
||||||
|
OUTPUT_SCHEMA,
|
||||||
|
meeting_context=context_text,
|
||||||
|
task_prompt_names=EXTRACTION_TASK_PROMPT_NAMES,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def call_ollama(
|
def call_ollama(
|
||||||
@@ -376,12 +403,13 @@ def extract_chunk(
|
|||||||
temperature: float,
|
temperature: float,
|
||||||
num_predict: int | None,
|
num_predict: int | None,
|
||||||
num_ctx: int | None,
|
num_ctx: int | None,
|
||||||
) -> dict[str, list[str]]:
|
meeting_context: MeetingContext | None = None,
|
||||||
|
) -> dict[str, Any]:
|
||||||
transcript = chunk_path.read_text(encoding="utf-8-sig").strip()
|
transcript = chunk_path.read_text(encoding="utf-8-sig").strip()
|
||||||
if not transcript:
|
if not transcript:
|
||||||
raise ValueError(f"The input file is empty: {chunk_path}")
|
raise ValueError(f"The input file is empty: {chunk_path}")
|
||||||
|
|
||||||
prompt = build_prompt(chunk_path.name, transcript)
|
prompt = build_prompt(chunk_path.name, transcript, meeting_context=meeting_context)
|
||||||
raw_text, _metadata = call_ollama(
|
raw_text, _metadata = call_ollama(
|
||||||
endpoint=endpoint,
|
endpoint=endpoint,
|
||||||
model=model,
|
model=model,
|
||||||
@@ -402,6 +430,8 @@ def extract_chunk(
|
|||||||
) from exc
|
) from exc
|
||||||
|
|
||||||
extraction = normalize_current_schema(parsed)
|
extraction = normalize_current_schema(parsed)
|
||||||
|
if meeting_context is not None:
|
||||||
|
extraction["context"] = meeting_context.provenance()
|
||||||
output_path.parent.mkdir(parents=True, exist_ok=True)
|
output_path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
output_path.write_text(
|
output_path.write_text(
|
||||||
json.dumps(extraction, ensure_ascii=False, indent=2) + "\n",
|
json.dumps(extraction, ensure_ascii=False, indent=2) + "\n",
|
||||||
@@ -419,6 +449,7 @@ def extract_input(
|
|||||||
temperature: float,
|
temperature: float,
|
||||||
num_predict: int | None,
|
num_predict: int | None,
|
||||||
num_ctx: int | None,
|
num_ctx: int | None,
|
||||||
|
meeting_context: MeetingContext | None = None,
|
||||||
) -> list[Path]:
|
) -> list[Path]:
|
||||||
if input_path.is_file():
|
if input_path.is_file():
|
||||||
output_path = output or extraction_path_for_chunk(input_path)
|
output_path = output or extraction_path_for_chunk(input_path)
|
||||||
@@ -431,6 +462,7 @@ def extract_input(
|
|||||||
temperature,
|
temperature,
|
||||||
num_predict,
|
num_predict,
|
||||||
num_ctx,
|
num_ctx,
|
||||||
|
meeting_context=meeting_context,
|
||||||
)
|
)
|
||||||
return [output_path]
|
return [output_path]
|
||||||
|
|
||||||
@@ -459,6 +491,7 @@ def extract_input(
|
|||||||
temperature,
|
temperature,
|
||||||
num_predict,
|
num_predict,
|
||||||
num_ctx,
|
num_ctx,
|
||||||
|
meeting_context=meeting_context,
|
||||||
)
|
)
|
||||||
output_paths.append(output_path)
|
output_paths.append(output_path)
|
||||||
|
|
||||||
@@ -469,6 +502,11 @@ def main() -> int:
|
|||||||
args = parse_args()
|
args = parse_args()
|
||||||
|
|
||||||
try:
|
try:
|
||||||
|
meeting_context = (
|
||||||
|
load_meeting_context(args.meeting_context)
|
||||||
|
if args.meeting_context is not None
|
||||||
|
else None
|
||||||
|
)
|
||||||
output_paths = extract_input(
|
output_paths = extract_input(
|
||||||
input_path=args.input,
|
input_path=args.input,
|
||||||
output=args.output,
|
output=args.output,
|
||||||
@@ -478,6 +516,7 @@ def main() -> int:
|
|||||||
temperature=args.temperature,
|
temperature=args.temperature,
|
||||||
num_predict=args.num_predict,
|
num_predict=args.num_predict,
|
||||||
num_ctx=args.num_ctx,
|
num_ctx=args.num_ctx,
|
||||||
|
meeting_context=meeting_context,
|
||||||
)
|
)
|
||||||
except requests.ConnectionError:
|
except requests.ConnectionError:
|
||||||
print(
|
print(
|
||||||
@@ -496,6 +535,8 @@ def main() -> int:
|
|||||||
return 1
|
return 1
|
||||||
|
|
||||||
print(f"Input: {args.input}")
|
print(f"Input: {args.input}")
|
||||||
|
if args.meeting_context is not None:
|
||||||
|
print(f"Meeting context: {args.meeting_context}")
|
||||||
print(f"Processed chunks: {len(output_paths)}")
|
print(f"Processed chunks: {len(output_paths)}")
|
||||||
print(f"Extraction JSON files: {len(output_paths)}")
|
print(f"Extraction JSON files: {len(output_paths)}")
|
||||||
for output_path in output_paths:
|
for output_path in output_paths:
|
||||||
|
|||||||
@@ -30,12 +30,14 @@ def build_extraction_prompt(
|
|||||||
source_name: str,
|
source_name: str,
|
||||||
transcript: str,
|
transcript: str,
|
||||||
output_schema: dict[str, Any],
|
output_schema: dict[str, Any],
|
||||||
|
meeting_context: str | None = None,
|
||||||
task_prompt_names: Iterable[str] = ("decisions.md",),
|
task_prompt_names: Iterable[str] = ("decisions.md",),
|
||||||
prompts_dir: Path = PROMPTS_DIR,
|
prompts_dir: Path = PROMPTS_DIR,
|
||||||
) -> str:
|
) -> str:
|
||||||
schema_text = json.dumps(output_schema, ensure_ascii=False, indent=2)
|
schema_text = json.dumps(output_schema, ensure_ascii=False, indent=2)
|
||||||
prompt_parts = [
|
prompt_parts = [
|
||||||
load_prompt("common.md", prompts_dir),
|
load_prompt("common.md", prompts_dir),
|
||||||
|
meeting_context,
|
||||||
*load_existing_prompts(task_prompt_names, prompts_dir),
|
*load_existing_prompts(task_prompt_names, prompts_dir),
|
||||||
f"""Quelldatei:
|
f"""Quelldatei:
|
||||||
{source_name}
|
{source_name}
|
||||||
|
|||||||
@@ -0,0 +1,396 @@
|
|||||||
|
"""Meeting Context V1 loading, validation and prompt rendering."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import ast
|
||||||
|
from dataclasses import dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
|
||||||
|
SUPPORTED_SCHEMA_VERSIONS = {"1"}
|
||||||
|
VALID_ATTENDANCE_STATUSES = {"present", "not_present", "absent"}
|
||||||
|
|
||||||
|
|
||||||
|
class MeetingContextValidationError(ValueError):
|
||||||
|
"""Raised when a Meeting Context file is structurally invalid."""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class MeetingContext:
|
||||||
|
data: dict[str, Any]
|
||||||
|
source_file: Path
|
||||||
|
|
||||||
|
@property
|
||||||
|
def schema_version(self) -> str:
|
||||||
|
return str(self.data["schema_version"])
|
||||||
|
|
||||||
|
@property
|
||||||
|
def meeting_id(self) -> str:
|
||||||
|
return str(self.data["meeting"]["meeting_id"])
|
||||||
|
|
||||||
|
def provenance(self) -> dict[str, str]:
|
||||||
|
return {
|
||||||
|
"meeting_id": self.meeting_id,
|
||||||
|
"source_file": str(self.source_file),
|
||||||
|
"schema_version": self.schema_version,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def load_meeting_context(path: Path) -> MeetingContext:
|
||||||
|
loaded = _load_yaml(path)
|
||||||
|
if not isinstance(loaded, dict):
|
||||||
|
raise MeetingContextValidationError("Meeting Context must be a YAML object.")
|
||||||
|
|
||||||
|
validate_meeting_context(loaded)
|
||||||
|
return MeetingContext(data=loaded, source_file=path)
|
||||||
|
|
||||||
|
|
||||||
|
def validate_meeting_context(data: dict[str, Any]) -> None:
|
||||||
|
schema_version = str(data.get("schema_version", "")).strip()
|
||||||
|
if schema_version not in SUPPORTED_SCHEMA_VERSIONS:
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"Unsupported meeting context schema_version: {schema_version!r}."
|
||||||
|
)
|
||||||
|
|
||||||
|
meeting = _require_mapping(data, "meeting")
|
||||||
|
_require_non_empty_string(meeting, "meeting.meeting_id")
|
||||||
|
_require_non_empty_string(meeting, "meeting.title")
|
||||||
|
_require_non_empty_string(meeting, "meeting.language")
|
||||||
|
|
||||||
|
organization = _optional_mapping(data.get("organization"), "organization")
|
||||||
|
departments = _optional_list(organization.get("departments"), "organization.departments")
|
||||||
|
department_ids = _collect_department_ids(departments)
|
||||||
|
|
||||||
|
participants = _optional_list(data.get("participants"), "participants")
|
||||||
|
mentioned_people = _optional_list(data.get("mentioned_people"), "mentioned_people")
|
||||||
|
|
||||||
|
participant_ids = _collect_unique_ids(participants, "participant_id", "participants")
|
||||||
|
person_ids = _collect_unique_ids(mentioned_people, "person_id", "mentioned_people")
|
||||||
|
collisions = sorted(participant_ids & person_ids)
|
||||||
|
if collisions:
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
"Participant IDs and mentioned-person IDs must not collide: "
|
||||||
|
+ ", ".join(collisions)
|
||||||
|
)
|
||||||
|
|
||||||
|
for index, participant in enumerate(participants):
|
||||||
|
item_path = f"participants[{index}]"
|
||||||
|
_validate_attendance(participant, item_path)
|
||||||
|
if participant.get("attendance_status") != "present":
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"{item_path}.attendance_status must be 'present'."
|
||||||
|
)
|
||||||
|
_validate_department_reference(participant, item_path, department_ids)
|
||||||
|
|
||||||
|
for index, person in enumerate(mentioned_people):
|
||||||
|
item_path = f"mentioned_people[{index}]"
|
||||||
|
_validate_attendance(person, item_path)
|
||||||
|
if person.get("attendance_status") == "present":
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"{item_path}.attendance_status must not be 'present'."
|
||||||
|
)
|
||||||
|
_validate_department_reference(person, item_path, department_ids)
|
||||||
|
|
||||||
|
|
||||||
|
def render_meeting_context_for_prompt(context: MeetingContext) -> str:
|
||||||
|
data = context.data
|
||||||
|
meeting = data["meeting"]
|
||||||
|
organization = _optional_mapping(data.get("organization"), "organization")
|
||||||
|
departments_by_id = {
|
||||||
|
str(department.get("id")): str(department.get("name"))
|
||||||
|
for department in _optional_list(
|
||||||
|
organization.get("departments"), "organization.departments"
|
||||||
|
)
|
||||||
|
if isinstance(department, dict) and department.get("id") and department.get("name")
|
||||||
|
}
|
||||||
|
|
||||||
|
lines = [
|
||||||
|
"MEETING CONTEXT V1 (AUTHORITATIVE METADATA)",
|
||||||
|
"",
|
||||||
|
"Rules:",
|
||||||
|
"- The participant list is authoritative.",
|
||||||
|
"- Mentioned people did not attend this meeting.",
|
||||||
|
"- Roles and departments must not be inferred or changed.",
|
||||||
|
"- Discussion of a department does not establish responsibility.",
|
||||||
|
"- An action item may name a responsible person only when assignment or acceptance is explicit in the transcript.",
|
||||||
|
"- Objections, suggestions and expertise do not establish ownership.",
|
||||||
|
"",
|
||||||
|
"Meeting:",
|
||||||
|
f"- Title: {_text(meeting.get('title'))}",
|
||||||
|
f"- Language: {_text(meeting.get('language'))}",
|
||||||
|
]
|
||||||
|
|
||||||
|
objective = _text(meeting.get("objective"))
|
||||||
|
if objective:
|
||||||
|
lines.append(f"- Objective: {objective}")
|
||||||
|
|
||||||
|
participants = _optional_list(data.get("participants"), "participants")
|
||||||
|
if participants:
|
||||||
|
lines.extend(["", "Actual participants:"])
|
||||||
|
for participant in participants:
|
||||||
|
lines.append(_render_person_line(participant, "participant_id", departments_by_id))
|
||||||
|
|
||||||
|
mentioned_people = _optional_list(data.get("mentioned_people"), "mentioned_people")
|
||||||
|
if mentioned_people:
|
||||||
|
lines.extend(["", "Mentioned but absent people:"])
|
||||||
|
for person in mentioned_people:
|
||||||
|
lines.append(_render_person_line(person, "person_id", departments_by_id))
|
||||||
|
|
||||||
|
abbreviations = _optional_mapping(organization.get("abbreviations"), "organization.abbreviations")
|
||||||
|
abbreviation_lines = [
|
||||||
|
f"- {key}: {_text(value)}"
|
||||||
|
for key, value in sorted(abbreviations.items())
|
||||||
|
if _text(value)
|
||||||
|
]
|
||||||
|
if abbreviation_lines:
|
||||||
|
lines.extend(["", "Abbreviations:", *abbreviation_lines])
|
||||||
|
|
||||||
|
known_entities = _optional_mapping(data.get("known_entities"), "known_entities")
|
||||||
|
entity_lines = []
|
||||||
|
for key in sorted(known_entities):
|
||||||
|
values = [_text(value) for value in _optional_list(known_entities.get(key), key)]
|
||||||
|
values = [value for value in values if value]
|
||||||
|
if values:
|
||||||
|
entity_lines.append(f"- {key}: {', '.join(values)}")
|
||||||
|
if entity_lines:
|
||||||
|
lines.extend(["", "Relevant known entities:", *entity_lines])
|
||||||
|
|
||||||
|
context_rules = _optional_mapping(data.get("context_rules"), "context_rules")
|
||||||
|
rule_lines = [
|
||||||
|
f"- {key}: {str(value).lower() if isinstance(value, bool) else _text(value)}"
|
||||||
|
for key, value in sorted(context_rules.items())
|
||||||
|
]
|
||||||
|
if rule_lines:
|
||||||
|
lines.extend(["", "Context rules:", *rule_lines])
|
||||||
|
|
||||||
|
return "\n".join(lines).strip() + "\n"
|
||||||
|
|
||||||
|
|
||||||
|
def _render_person_line(
|
||||||
|
person: dict[str, Any],
|
||||||
|
id_key: str,
|
||||||
|
departments_by_id: dict[str, str],
|
||||||
|
) -> str:
|
||||||
|
parts = [_text(person.get("display_name"))]
|
||||||
|
aliases = [_text(alias) for alias in _optional_list(person.get("aliases"), "aliases")]
|
||||||
|
aliases = [alias for alias in aliases if alias]
|
||||||
|
if aliases:
|
||||||
|
parts.append(f"aliases: {', '.join(aliases)}")
|
||||||
|
|
||||||
|
roles = _roles(person)
|
||||||
|
if roles:
|
||||||
|
parts.append(f"roles: {', '.join(roles)}")
|
||||||
|
|
||||||
|
department_id = _text(person.get("department_id") or person.get("department"))
|
||||||
|
if department_id:
|
||||||
|
department_name = departments_by_id.get(department_id, department_id)
|
||||||
|
parts.append(f"department: {department_name}")
|
||||||
|
|
||||||
|
identifier = _text(person.get(id_key))
|
||||||
|
if identifier:
|
||||||
|
parts.append(f"id: {identifier}")
|
||||||
|
|
||||||
|
return "- " + "; ".join(part for part in parts if part)
|
||||||
|
|
||||||
|
|
||||||
|
def _roles(person: dict[str, Any]) -> list[str]:
|
||||||
|
roles = [_text(role) for role in _optional_list(person.get("roles"), "roles")]
|
||||||
|
role = _text(person.get("role"))
|
||||||
|
if role:
|
||||||
|
roles.insert(0, role)
|
||||||
|
return [role for role in roles if role]
|
||||||
|
|
||||||
|
|
||||||
|
def _load_yaml(path: Path) -> Any:
|
||||||
|
text = path.read_text(encoding="utf-8-sig")
|
||||||
|
try:
|
||||||
|
import yaml # type: ignore[import-not-found]
|
||||||
|
except ModuleNotFoundError:
|
||||||
|
return _parse_simple_yaml(text)
|
||||||
|
return yaml.safe_load(text)
|
||||||
|
|
||||||
|
|
||||||
|
def _collect_department_ids(departments: list[Any]) -> set[str]:
|
||||||
|
ids: set[str] = set()
|
||||||
|
for index, department in enumerate(departments):
|
||||||
|
if not isinstance(department, dict):
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"organization.departments[{index}] must be an object."
|
||||||
|
)
|
||||||
|
department_id = _text(department.get("id"))
|
||||||
|
if not department_id:
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"organization.departments[{index}].id must be non-empty."
|
||||||
|
)
|
||||||
|
if department_id in ids:
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"Duplicate department id: {department_id!r}."
|
||||||
|
)
|
||||||
|
ids.add(department_id)
|
||||||
|
return ids
|
||||||
|
|
||||||
|
|
||||||
|
def _collect_unique_ids(items: list[Any], key: str, path: str) -> set[str]:
|
||||||
|
ids: set[str] = set()
|
||||||
|
for index, item in enumerate(items):
|
||||||
|
if not isinstance(item, dict):
|
||||||
|
raise MeetingContextValidationError(f"{path}[{index}] must be an object.")
|
||||||
|
identifier = _text(item.get(key))
|
||||||
|
if not identifier:
|
||||||
|
raise MeetingContextValidationError(f"{path}[{index}].{key} must be non-empty.")
|
||||||
|
if identifier in ids:
|
||||||
|
raise MeetingContextValidationError(f"Duplicate {key}: {identifier!r}.")
|
||||||
|
ids.add(identifier)
|
||||||
|
return ids
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_attendance(item: dict[str, Any], path: str) -> None:
|
||||||
|
status = item.get("attendance_status")
|
||||||
|
if status not in VALID_ATTENDANCE_STATUSES:
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"{path}.attendance_status has invalid value: {status!r}."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_department_reference(
|
||||||
|
item: dict[str, Any],
|
||||||
|
path: str,
|
||||||
|
department_ids: set[str],
|
||||||
|
) -> None:
|
||||||
|
department_id = _text(item.get("department_id") or item.get("department"))
|
||||||
|
if department_id and department_id not in department_ids:
|
||||||
|
raise MeetingContextValidationError(
|
||||||
|
f"{path}.department_id references unknown department: {department_id!r}."
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _require_mapping(data: dict[str, Any], key: str) -> dict[str, Any]:
|
||||||
|
value = data.get(key)
|
||||||
|
if not isinstance(value, dict):
|
||||||
|
raise MeetingContextValidationError(f"{key} must be an object.")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _optional_mapping(value: Any, path: str) -> dict[str, Any]:
|
||||||
|
if value is None:
|
||||||
|
return {}
|
||||||
|
if not isinstance(value, dict):
|
||||||
|
raise MeetingContextValidationError(f"{path} must be an object.")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _optional_list(value: Any, path: str) -> list[Any]:
|
||||||
|
if value is None:
|
||||||
|
return []
|
||||||
|
if not isinstance(value, list):
|
||||||
|
raise MeetingContextValidationError(f"{path} must be a list.")
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _require_non_empty_string(data: dict[str, Any], key_path: str) -> None:
|
||||||
|
key = key_path.split(".")[-1]
|
||||||
|
if not _text(data.get(key)):
|
||||||
|
raise MeetingContextValidationError(f"{key_path} must be present and non-empty.")
|
||||||
|
|
||||||
|
|
||||||
|
def _text(value: Any) -> str:
|
||||||
|
if value is None:
|
||||||
|
return ""
|
||||||
|
return str(value).strip()
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_simple_yaml(text: str) -> Any:
|
||||||
|
lines = []
|
||||||
|
for raw_line in text.splitlines():
|
||||||
|
stripped = _strip_yaml_comment(raw_line.rstrip())
|
||||||
|
if stripped.strip():
|
||||||
|
lines.append((len(stripped) - len(stripped.lstrip(" ")), stripped.lstrip(" ")))
|
||||||
|
if not lines:
|
||||||
|
return None
|
||||||
|
parsed, index = _parse_yaml_block(lines, 0, lines[0][0])
|
||||||
|
if index != len(lines):
|
||||||
|
raise MeetingContextValidationError("Could not parse Meeting Context YAML.")
|
||||||
|
return parsed
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_yaml_block(
|
||||||
|
lines: list[tuple[int, str]],
|
||||||
|
index: int,
|
||||||
|
indent: int,
|
||||||
|
) -> tuple[Any, int]:
|
||||||
|
if lines[index][1].startswith("- "):
|
||||||
|
result = []
|
||||||
|
while index < len(lines) and lines[index][0] == indent and lines[index][1].startswith("- "):
|
||||||
|
content = lines[index][1][2:].strip()
|
||||||
|
index += 1
|
||||||
|
if not content:
|
||||||
|
value, index = _parse_yaml_block(lines, index, lines[index][0])
|
||||||
|
result.append(value)
|
||||||
|
continue
|
||||||
|
|
||||||
|
if ":" in content:
|
||||||
|
key, raw_value = content.split(":", 1)
|
||||||
|
item = {key.strip(): _parse_scalar(raw_value.strip()) if raw_value.strip() else {}}
|
||||||
|
while index < len(lines) and lines[index][0] > indent:
|
||||||
|
child_indent, child_content = lines[index]
|
||||||
|
if child_content.startswith("- "):
|
||||||
|
break
|
||||||
|
child_key, child_raw_value = child_content.split(":", 1)
|
||||||
|
index += 1
|
||||||
|
child_raw_value = child_raw_value.strip()
|
||||||
|
if child_raw_value:
|
||||||
|
item[child_key.strip()] = _parse_scalar(child_raw_value)
|
||||||
|
elif index < len(lines) and lines[index][0] > child_indent:
|
||||||
|
item[child_key.strip()], index = _parse_yaml_block(lines, index, lines[index][0])
|
||||||
|
else:
|
||||||
|
item[child_key.strip()] = None
|
||||||
|
result.append(item)
|
||||||
|
else:
|
||||||
|
result.append(_parse_scalar(content))
|
||||||
|
return result, index
|
||||||
|
|
||||||
|
result = {}
|
||||||
|
while index < len(lines) and lines[index][0] == indent and not lines[index][1].startswith("- "):
|
||||||
|
key, raw_value = lines[index][1].split(":", 1)
|
||||||
|
index += 1
|
||||||
|
raw_value = raw_value.strip()
|
||||||
|
if raw_value:
|
||||||
|
result[key.strip()] = _parse_scalar(raw_value)
|
||||||
|
elif index < len(lines) and lines[index][0] > indent:
|
||||||
|
result[key.strip()], index = _parse_yaml_block(lines, index, lines[index][0])
|
||||||
|
else:
|
||||||
|
result[key.strip()] = None
|
||||||
|
return result, index
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_scalar(value: str) -> Any:
|
||||||
|
if value in {"null", "Null", "NULL", "~"}:
|
||||||
|
return None
|
||||||
|
if value in {"true", "True", "TRUE"}:
|
||||||
|
return True
|
||||||
|
if value in {"false", "False", "FALSE"}:
|
||||||
|
return False
|
||||||
|
if value in {"[]", "{}"} or (
|
||||||
|
value.startswith("[") and value.endswith("]")
|
||||||
|
):
|
||||||
|
return ast.literal_eval(value)
|
||||||
|
if (
|
||||||
|
(value.startswith('"') and value.endswith('"'))
|
||||||
|
or (value.startswith("'") and value.endswith("'"))
|
||||||
|
):
|
||||||
|
return ast.literal_eval(value)
|
||||||
|
return value
|
||||||
|
|
||||||
|
|
||||||
|
def _strip_yaml_comment(line: str) -> str:
|
||||||
|
in_single = False
|
||||||
|
in_double = False
|
||||||
|
for index, char in enumerate(line):
|
||||||
|
if char == "'" and not in_double:
|
||||||
|
in_single = not in_single
|
||||||
|
elif char == '"' and not in_single:
|
||||||
|
in_double = not in_double
|
||||||
|
elif char == "#" and not in_single and not in_double:
|
||||||
|
return line[:index].rstrip()
|
||||||
|
return line
|
||||||
@@ -0,0 +1,18 @@
|
|||||||
|
# position_explicit_objection
|
||||||
|
|
||||||
|
Tests extraction of one explicit objection as a position.
|
||||||
|
|
||||||
|
Scope:
|
||||||
|
|
||||||
|
- Extract Tom's stated objection as exactly one position.
|
||||||
|
- Do not create a todo for Tom.
|
||||||
|
- Do not derive a decision from the objection.
|
||||||
|
|
||||||
|
Exclusions:
|
||||||
|
|
||||||
|
- This scenario does not test responsibility attribution for an open owner.
|
||||||
|
- Responsibility attribution is covered by `responsibility_attribution_negative`.
|
||||||
|
|
||||||
|
The transcript gives unique ground truth because Tom explicitly says "I object"
|
||||||
|
and "My position is", then explicitly refuses responsibility. The group also
|
||||||
|
states that no decision and no task for Tom were created.
|
||||||
@@ -0,0 +1,14 @@
|
|||||||
|
{
|
||||||
|
"facts": [],
|
||||||
|
"decisions": [],
|
||||||
|
"todos": [],
|
||||||
|
"questions": [],
|
||||||
|
"positions": [
|
||||||
|
{
|
||||||
|
"speaker": "Tom",
|
||||||
|
"position": "Tom objects to using one generic intake checklist because generic criteria will not work for analytics pilots and hardware trials.",
|
||||||
|
"evidence": "Tom: I object to that. My position is that generic criteria will not work for these two types of work."
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"technical": []
|
||||||
|
}
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
Iris: We could use one generic intake checklist for analytics pilots and hardware trials.
|
||||||
|
|
||||||
|
Tom: I object to that. My position is that generic criteria will not work for these two types of work.
|
||||||
|
|
||||||
|
Iris: Understood. Are you taking responsibility for rewriting the checklist?
|
||||||
|
|
||||||
|
Tom: No. I am not taking that on, and I am not proposing an owner. I am only stating my position.
|
||||||
|
|
||||||
|
Uma: Then we are not deciding the checklist today.
|
||||||
|
|
||||||
|
Iris: Correct. No decision and no task for Tom.
|
||||||
@@ -7,9 +7,14 @@ The scenario includes a Marketing participant who comments critically on
|
|||||||
Business Development criteria. No one assigns that participant responsibility
|
Business Development criteria. No one assigns that participant responsibility
|
||||||
for defining the criteria, and the participant does not accept such a task.
|
for defining the criteria, and the participant does not accept such a task.
|
||||||
|
|
||||||
Expected behavior:
|
Scope:
|
||||||
|
|
||||||
- Do not assign Business Development criteria to the Marketing participant.
|
- Do not assign Business Development criteria to the Marketing participant.
|
||||||
- Do not reclassify the Marketing participant as Business Development.
|
- Do not reclassify the Marketing participant as Business Development.
|
||||||
- Preserve the objection as a position.
|
|
||||||
- Keep responsibility open unless explicitly assigned.
|
- Keep responsibility open unless explicitly assigned.
|
||||||
|
- Preserve the agreed next step to ask Business Development for an owner.
|
||||||
|
|
||||||
|
Exclusions:
|
||||||
|
|
||||||
|
- This scenario does not test whether the objection is extracted as a position.
|
||||||
|
- Position extraction is covered by `position_explicit_objection`.
|
||||||
|
|||||||
@@ -28,12 +28,6 @@
|
|||||||
}
|
}
|
||||||
],
|
],
|
||||||
"questions": [],
|
"questions": [],
|
||||||
"positions": [
|
"positions": [],
|
||||||
{
|
|
||||||
"speaker": "Noah",
|
|
||||||
"position": "Noah says generic criteria will not work for different Marketing and product contexts.",
|
|
||||||
"evidence": "Noah: From Marketing, I can tell you that generic criteria will not work. A digital campaign and a physical product launch need different checks."
|
|
||||||
}
|
|
||||||
],
|
|
||||||
"technical": []
|
"technical": []
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -5,12 +5,15 @@ from pathlib import Path
|
|||||||
|
|
||||||
from src.meeting_lab.extraction.extract_chunks import (
|
from src.meeting_lab.extraction.extract_chunks import (
|
||||||
EXTRACTION_CATEGORIES,
|
EXTRACTION_CATEGORIES,
|
||||||
|
EXTRACTION_TASK_PROMPT_NAMES,
|
||||||
build_prompt,
|
build_prompt,
|
||||||
extraction_path_for_chunk,
|
extraction_path_for_chunk,
|
||||||
normalize_current_schema,
|
normalize_current_schema,
|
||||||
parse_json_response,
|
parse_json_response,
|
||||||
)
|
)
|
||||||
|
from src.meeting_lab.models.meeting_context import load_meeting_context
|
||||||
from src.meeting_lab.protocol.build_protocol import build_protocol
|
from src.meeting_lab.protocol.build_protocol import build_protocol
|
||||||
|
from scripts import run_gold_test
|
||||||
|
|
||||||
|
|
||||||
class ExtractionProtocolTests(unittest.TestCase):
|
class ExtractionProtocolTests(unittest.TestCase):
|
||||||
@@ -102,6 +105,41 @@ Final answer:
|
|||||||
self.assertIn("Extract each decision as one atomic commitment.", prompt)
|
self.assertIn("Extract each decision as one atomic commitment.", prompt)
|
||||||
self.assertIn("Anna: Agreed.", prompt)
|
self.assertIn("Anna: Agreed.", prompt)
|
||||||
|
|
||||||
|
def test_build_prompt_includes_todo_prompt_file(self) -> None:
|
||||||
|
prompt = build_prompt("transcript.txt", "Nina: I will update it.")
|
||||||
|
|
||||||
|
self.assertEqual(EXTRACTION_TASK_PROMPT_NAMES, ("decisions.md", "todos.md"))
|
||||||
|
self.assertIn("Action-item responsibility rule:", prompt)
|
||||||
|
self.assertIn(
|
||||||
|
"A named responsible person may be extracted only when the transcript contains",
|
||||||
|
prompt,
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_gold_runner_uses_shared_production_prompt_assembly(self) -> None:
|
||||||
|
self.assertIs(run_gold_test.build_prompt, build_prompt)
|
||||||
|
|
||||||
|
def test_no_context_and_meeting_context_prompts_use_same_task_prompt_set(self) -> None:
|
||||||
|
context = load_meeting_context(
|
||||||
|
Path("samples/real_live/project_process_meeting/meeting_context.yaml")
|
||||||
|
)
|
||||||
|
|
||||||
|
no_context_prompt = build_prompt("chunk_01_normalized.txt", "Anna: Agreed.")
|
||||||
|
context_prompt = build_prompt(
|
||||||
|
"chunk_01_normalized.txt",
|
||||||
|
"Anna: Agreed.",
|
||||||
|
meeting_context=context,
|
||||||
|
)
|
||||||
|
|
||||||
|
for prompt in (no_context_prompt, context_prompt):
|
||||||
|
self.assertIn("You extract decisions from meeting transcript text.", prompt)
|
||||||
|
self.assertIn("Action-item responsibility rule:", prompt)
|
||||||
|
|
||||||
|
def test_prompt_assembly_is_deterministic(self) -> None:
|
||||||
|
first = build_prompt("transcript.txt", "Nina: I will update it.")
|
||||||
|
second = build_prompt("transcript.txt", "Nina: I will update it.")
|
||||||
|
|
||||||
|
self.assertEqual(first, second)
|
||||||
|
|
||||||
def test_build_protocol_groups_extraction_items(self) -> None:
|
def test_build_protocol_groups_extraction_items(self) -> None:
|
||||||
with tempfile.TemporaryDirectory() as directory:
|
with tempfile.TemporaryDirectory() as directory:
|
||||||
input_dir = Path(directory)
|
input_dir = Path(directory)
|
||||||
|
|||||||
@@ -0,0 +1,197 @@
|
|||||||
|
import copy
|
||||||
|
import json
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
from src.meeting_lab.extraction.extract_chunks import (
|
||||||
|
EXTRACTION_CATEGORIES,
|
||||||
|
build_prompt,
|
||||||
|
extract_chunk,
|
||||||
|
)
|
||||||
|
from src.meeting_lab.models.meeting_context import (
|
||||||
|
MeetingContextValidationError,
|
||||||
|
load_meeting_context,
|
||||||
|
render_meeting_context_for_prompt,
|
||||||
|
validate_meeting_context,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
CONTEXT_PATH = Path("samples/real_live/project_process_meeting/meeting_context.yaml")
|
||||||
|
SCRATCH_DIR = Path(".test-tmp")
|
||||||
|
|
||||||
|
|
||||||
|
class MeetingContextTests(unittest.TestCase):
|
||||||
|
def setUp(self) -> None:
|
||||||
|
self.context = load_meeting_context(CONTEXT_PATH)
|
||||||
|
|
||||||
|
def tearDown(self) -> None:
|
||||||
|
if not SCRATCH_DIR.exists():
|
||||||
|
return
|
||||||
|
for path in SCRATCH_DIR.glob("chunk_0*_*.txt"):
|
||||||
|
path.unlink()
|
||||||
|
for path in SCRATCH_DIR.glob("chunk_0*_*.json"):
|
||||||
|
path.unlink()
|
||||||
|
|
||||||
|
def test_valid_context_loading(self) -> None:
|
||||||
|
self.assertEqual(self.context.schema_version, "1")
|
||||||
|
self.assertEqual(self.context.meeting_id, "2026-07-27-projektprozess")
|
||||||
|
self.assertEqual(self.context.data["meeting"]["language"], "de")
|
||||||
|
|
||||||
|
def test_duplicate_participant_ids_are_invalid(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
data["participants"][1]["participant_id"] = data["participants"][0][
|
||||||
|
"participant_id"
|
||||||
|
]
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(MeetingContextValidationError, "Duplicate"):
|
||||||
|
validate_meeting_context(data)
|
||||||
|
|
||||||
|
def test_invalid_department_references_are_invalid(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
data["participants"][0]["department_id"] = "unknown"
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(MeetingContextValidationError, "unknown department"):
|
||||||
|
validate_meeting_context(data)
|
||||||
|
|
||||||
|
def test_participant_and_mentioned_person_id_collision_is_invalid(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
data["mentioned_people"][0]["person_id"] = data["participants"][0][
|
||||||
|
"participant_id"
|
||||||
|
]
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(MeetingContextValidationError, "collide"):
|
||||||
|
validate_meeting_context(data)
|
||||||
|
|
||||||
|
def test_invalid_attendance_status_is_invalid(self) -> None:
|
||||||
|
data = copy.deepcopy(self.context.data)
|
||||||
|
data["participants"][0]["attendance_status"] = "remote"
|
||||||
|
|
||||||
|
with self.assertRaisesRegex(MeetingContextValidationError, "invalid value"):
|
||||||
|
validate_meeting_context(data)
|
||||||
|
|
||||||
|
def test_prompt_representation_is_deterministic(self) -> None:
|
||||||
|
first = render_meeting_context_for_prompt(self.context)
|
||||||
|
second = render_meeting_context_for_prompt(self.context)
|
||||||
|
|
||||||
|
self.assertEqual(first, second)
|
||||||
|
self.assertIn("MEETING CONTEXT V1", first)
|
||||||
|
self.assertIn("- Language: de", first)
|
||||||
|
|
||||||
|
def test_authoritative_rules_appear_in_prompt(self) -> None:
|
||||||
|
prompt = build_prompt(
|
||||||
|
"chunk_01_normalized.txt",
|
||||||
|
"Martin: Wir besprechen Marketing.",
|
||||||
|
meeting_context=self.context,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertIn("The participant list is authoritative.", prompt)
|
||||||
|
self.assertIn("Mentioned people did not attend this meeting.", prompt)
|
||||||
|
self.assertIn("Roles and departments must not be inferred or changed.", prompt)
|
||||||
|
self.assertIn("Discussion of a department does not establish responsibility.", prompt)
|
||||||
|
self.assertIn(
|
||||||
|
"An action item may name a responsible person only when assignment or acceptance is explicit",
|
||||||
|
prompt,
|
||||||
|
)
|
||||||
|
self.assertIn("Objections, suggestions and expertise do not establish ownership.", prompt)
|
||||||
|
|
||||||
|
def test_extraction_behavior_is_unchanged_without_context(self) -> None:
|
||||||
|
response = json.dumps(
|
||||||
|
{
|
||||||
|
"facts": [],
|
||||||
|
"decisions": [],
|
||||||
|
"todos": [],
|
||||||
|
"open_questions": [],
|
||||||
|
"positions": [],
|
||||||
|
"technical_details": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
SCRATCH_DIR.mkdir(exist_ok=True)
|
||||||
|
chunk_path = SCRATCH_DIR / "chunk_01_normalized.txt"
|
||||||
|
output_path = SCRATCH_DIR / "chunk_01_extraction.json"
|
||||||
|
chunk_path.write_text("Anna: Keine Entscheidung.", encoding="utf-8")
|
||||||
|
|
||||||
|
captured_prompts = []
|
||||||
|
|
||||||
|
def fake_call_ollama(**kwargs):
|
||||||
|
captured_prompts.append(kwargs["prompt"])
|
||||||
|
return response, {}
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.extraction.extract_chunks.call_ollama",
|
||||||
|
side_effect=fake_call_ollama,
|
||||||
|
):
|
||||||
|
extraction = extract_chunk(
|
||||||
|
chunk_path,
|
||||||
|
output_path,
|
||||||
|
model="test",
|
||||||
|
endpoint="http://example.invalid",
|
||||||
|
timeout=1,
|
||||||
|
temperature=0.0,
|
||||||
|
num_predict=None,
|
||||||
|
num_ctx=None,
|
||||||
|
)
|
||||||
|
|
||||||
|
self.assertEqual(set(extraction), set(EXTRACTION_CATEGORIES))
|
||||||
|
self.assertNotIn("context", extraction)
|
||||||
|
self.assertNotIn("MEETING CONTEXT V1", captured_prompts[0])
|
||||||
|
|
||||||
|
def test_extraction_result_records_meeting_context_provenance(self) -> None:
|
||||||
|
response = json.dumps(
|
||||||
|
{
|
||||||
|
"facts": [],
|
||||||
|
"decisions": [],
|
||||||
|
"todos": [],
|
||||||
|
"open_questions": [],
|
||||||
|
"positions": [],
|
||||||
|
"technical_details": [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
SCRATCH_DIR.mkdir(exist_ok=True)
|
||||||
|
chunk_path = SCRATCH_DIR / "chunk_02_normalized.txt"
|
||||||
|
output_path = SCRATCH_DIR / "chunk_02_extraction.json"
|
||||||
|
chunk_path.write_text("Martin: Hallo.", encoding="utf-8")
|
||||||
|
|
||||||
|
with patch(
|
||||||
|
"src.meeting_lab.extraction.extract_chunks.call_ollama",
|
||||||
|
return_value=(response, {}),
|
||||||
|
):
|
||||||
|
extraction = extract_chunk(
|
||||||
|
chunk_path,
|
||||||
|
output_path,
|
||||||
|
model="test",
|
||||||
|
endpoint="http://example.invalid",
|
||||||
|
timeout=1,
|
||||||
|
temperature=0.0,
|
||||||
|
num_predict=None,
|
||||||
|
num_ctx=None,
|
||||||
|
meeting_context=self.context,
|
||||||
|
)
|
||||||
|
|
||||||
|
written = json.loads(output_path.read_text(encoding="utf-8"))
|
||||||
|
|
||||||
|
self.assertEqual(
|
||||||
|
extraction["context"]["meeting_id"],
|
||||||
|
"2026-07-27-projektprozess",
|
||||||
|
)
|
||||||
|
self.assertEqual(extraction["context"]["schema_version"], "1")
|
||||||
|
self.assertEqual(written["context"], extraction["context"])
|
||||||
|
self.assertNotIn("participants", written["context"])
|
||||||
|
|
||||||
|
def test_real_context_keeps_metadata_separate_from_assignment(self) -> None:
|
||||||
|
prompt_context = render_meeting_context_for_prompt(self.context)
|
||||||
|
|
||||||
|
self.assertIn("Björn", prompt_context)
|
||||||
|
self.assertIn("department: Marketing", prompt_context)
|
||||||
|
self.assertIn("Mentioned but absent people:", prompt_context)
|
||||||
|
self.assertIn("Jovana", prompt_context)
|
||||||
|
self.assertIn("Discussion of a department does not establish responsibility.", prompt_context)
|
||||||
|
self.assertIn("assignment or acceptance is explicit", prompt_context)
|
||||||
|
self.assertNotIn("responsible: Björn", prompt_context)
|
||||||
|
self.assertNotIn("responsible: Jovana", prompt_context)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
Reference in New Issue
Block a user