Introduce Meeting Context V1 with YAML schema, validation and template. Support optional --meeting-context during chunk extraction. Inject authoritative Meeting Context into extraction prompts. Record Meeting Context provenance in extraction output. Activate todos.md in shared prompt assembly. Strengthen responsibility attribution and decision/todo boundaries. Add focused Gold scenarios and validation tests. Update architecture and pipeline documentation.
71 lines
2.1 KiB
Markdown
71 lines
2.1 KiB
Markdown
# Project Process Meeting Real-Life Sample
|
|
|
|
## Purpose
|
|
|
|
This directory preserves the complete input data for the current real meeting
|
|
experiment so the Meeting Lab pipeline can be reproduced on another computer.
|
|
|
|
This is a deliberately selected reference sample, not ordinary generated
|
|
runtime data.
|
|
|
|
## Confidentiality
|
|
|
|
This sample contains real meeting data. Keep the repository private. Do not
|
|
copy, publish or redistribute these files outside the intended private
|
|
development context.
|
|
|
|
Avoid exposing unnecessary personal details in derived documentation, reports
|
|
or screenshots.
|
|
|
|
## Language
|
|
|
|
The meeting language is German.
|
|
|
|
## Contents
|
|
|
|
Original input files:
|
|
|
|
- `audio/meeting_speech.wav`: source meeting audio.
|
|
- `transcript/meeting_speech.json`: raw Whisper transcript JSON.
|
|
- `transcript/meeting_speech_cleaned.json`: cleaned Whisper transcript JSON.
|
|
|
|
Reference material:
|
|
|
|
- `meeting_context.yaml`: manually maintained Meeting Context V1 scaffold for
|
|
reliable metadata. It can be validated and injected into chunk extraction
|
|
with `--meeting-context`; later pipeline stages are not connected yet.
|
|
- `reference/human_reference_protocol.md`: not currently included. No exact
|
|
existing human-written protocol file was found in the repository workspace.
|
|
|
|
Generated outputs are intentionally excluded. This sample does not include
|
|
generated chunks, normalized chunks, segmentation outputs, extraction JSON
|
|
files, consolidated outputs, generated protocols, raw model responses or
|
|
temporary files.
|
|
|
|
## Intended Pipeline Use
|
|
|
|
This sample may be used to reproduce and test:
|
|
|
|
- audio/transcription workflow setup
|
|
- Whisper JSON cleanup behavior
|
|
- transcript chunking
|
|
- normalization
|
|
- segmentation experiments
|
|
- extraction experiments
|
|
- canonicalization and future consolidation experiments
|
|
|
|
Generated artifacts should be written to normal experiment or benchmark output
|
|
locations, not back into this input sample directory.
|
|
|
|
## Git LFS
|
|
|
|
The WAV audio is intentionally versioned through Git LFS using the scoped
|
|
repository rule for `samples/real_live/**/*.wav`.
|
|
|
|
After cloning the repository, retrieve LFS files with:
|
|
|
|
```text
|
|
git lfs pull
|
|
```
|
|
|