Files
meeting-lab/samples/real_live/project_process_meeting/README.md
T
admin 06f0e7e651 Implement Meeting Context V1 and extraction improvements
Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
2026-08-01 16:49:48 +02:00

71 lines
2.1 KiB
Markdown

# Project Process Meeting Real-Life Sample
## Purpose
This directory preserves the complete input data for the current real meeting
experiment so the Meeting Lab pipeline can be reproduced on another computer.
This is a deliberately selected reference sample, not ordinary generated
runtime data.
## Confidentiality
This sample contains real meeting data. Keep the repository private. Do not
copy, publish or redistribute these files outside the intended private
development context.
Avoid exposing unnecessary personal details in derived documentation, reports
or screenshots.
## Language
The meeting language is German.
## Contents
Original input files:
- `audio/meeting_speech.wav`: source meeting audio.
- `transcript/meeting_speech.json`: raw Whisper transcript JSON.
- `transcript/meeting_speech_cleaned.json`: cleaned Whisper transcript JSON.
Reference material:
- `meeting_context.yaml`: manually maintained Meeting Context V1 scaffold for
reliable metadata. It can be validated and injected into chunk extraction
with `--meeting-context`; later pipeline stages are not connected yet.
- `reference/human_reference_protocol.md`: not currently included. No exact
existing human-written protocol file was found in the repository workspace.
Generated outputs are intentionally excluded. This sample does not include
generated chunks, normalized chunks, segmentation outputs, extraction JSON
files, consolidated outputs, generated protocols, raw model responses or
temporary files.
## Intended Pipeline Use
This sample may be used to reproduce and test:
- audio/transcription workflow setup
- Whisper JSON cleanup behavior
- transcript chunking
- normalization
- segmentation experiments
- extraction experiments
- canonicalization and future consolidation experiments
Generated artifacts should be written to normal experiment or benchmark output
locations, not back into this input sample directory.
## Git LFS
The WAV audio is intentionally versioned through Git LFS using the scoped
repository rule for `samples/real_live/**/*.wav`.
After cloning the repository, retrieve LFS files with:
```text
git lfs pull
```