Introduce Meeting Context V1 with YAML schema, validation and template. Support optional --meeting-context during chunk extraction. Inject authoritative Meeting Context into extraction prompts. Record Meeting Context provenance in extraction output. Activate todos.md in shared prompt assembly. Strengthen responsibility attribution and decision/todo boundaries. Add focused Gold scenarios and validation tests. Update architecture and pipeline documentation.
2.1 KiB
Project Process Meeting Real-Life Sample
Purpose
This directory preserves the complete input data for the current real meeting experiment so the Meeting Lab pipeline can be reproduced on another computer.
This is a deliberately selected reference sample, not ordinary generated runtime data.
Confidentiality
This sample contains real meeting data. Keep the repository private. Do not copy, publish or redistribute these files outside the intended private development context.
Avoid exposing unnecessary personal details in derived documentation, reports or screenshots.
Language
The meeting language is German.
Contents
Original input files:
audio/meeting_speech.wav: source meeting audio.transcript/meeting_speech.json: raw Whisper transcript JSON.transcript/meeting_speech_cleaned.json: cleaned Whisper transcript JSON.
Reference material:
meeting_context.yaml: manually maintained Meeting Context V1 scaffold for reliable metadata. It can be validated and injected into chunk extraction with--meeting-context; later pipeline stages are not connected yet.reference/human_reference_protocol.md: not currently included. No exact existing human-written protocol file was found in the repository workspace.
Generated outputs are intentionally excluded. This sample does not include generated chunks, normalized chunks, segmentation outputs, extraction JSON files, consolidated outputs, generated protocols, raw model responses or temporary files.
Intended Pipeline Use
This sample may be used to reproduce and test:
- audio/transcription workflow setup
- Whisper JSON cleanup behavior
- transcript chunking
- normalization
- segmentation experiments
- extraction experiments
- canonicalization and future consolidation experiments
Generated artifacts should be written to normal experiment or benchmark output locations, not back into this input sample directory.
Git LFS
The WAV audio is intentionally versioned through Git LFS using the scoped
repository rule for samples/real_live/**/*.wav.
After cloning the repository, retrieve LFS files with:
git lfs pull