Add real-life reference meeting sample

- add reproducible real-world meeting sample
- store meeting audio using Git LFS
- include raw and cleaned Whisper transcripts
- add manifest with SHA-256 checksums
- document sample usage and confidentiality
- intentionally exclude generated pipeline artifacts
This commit is contained in:
2026-07-31 10:13:52 +02:00
parent 23bbc744f7
commit 6e34334506
6 changed files with 36141 additions and 0 deletions
@@ -0,0 +1,67 @@
# Project Process Meeting Real-Life Sample
## Purpose
This directory preserves the complete input data for the current real meeting
experiment so the Meeting Lab pipeline can be reproduced on another computer.
This is a deliberately selected reference sample, not ordinary generated
runtime data.
## Confidentiality
This sample contains real meeting data. Keep the repository private. Do not
copy, publish or redistribute these files outside the intended private
development context.
Avoid exposing unnecessary personal details in derived documentation, reports
or screenshots.
## Language
The meeting language is German.
## Contents
Original input files:
- `audio/meeting_speech.wav`: source meeting audio.
- `transcript/meeting_speech.json`: raw Whisper transcript JSON.
- `transcript/meeting_speech_cleaned.json`: cleaned Whisper transcript JSON.
Reference material:
- `reference/human_reference_protocol.md`: not currently included. No exact
existing human-written protocol file was found in the repository workspace.
Generated outputs are intentionally excluded. This sample does not include
generated chunks, normalized chunks, segmentation outputs, extraction JSON
files, consolidated outputs, generated protocols, raw model responses or
temporary files.
## Intended Pipeline Use
This sample may be used to reproduce and test:
- audio/transcription workflow setup
- Whisper JSON cleanup behavior
- transcript chunking
- normalization
- segmentation experiments
- extraction experiments
- canonicalization and future consolidation experiments
Generated artifacts should be written to normal experiment or benchmark output
locations, not back into this input sample directory.
## Git LFS
The WAV audio is intentionally versioned through Git LFS using the scoped
repository rule for `samples/real_live/**/*.wav`.
After cloning the repository, retrieve LFS files with:
```text
git lfs pull
```
@@ -0,0 +1,33 @@
{
"sample_id": "project_process_meeting",
"title": "Project Process Meeting",
"language": "de",
"audio_file": "audio/meeting_speech.wav",
"raw_transcript_file": "transcript/meeting_speech.json",
"cleaned_transcript_file": "transcript/meeting_speech_cleaned.json",
"reference_protocol_file": null,
"audio_format": "WAV PCM 16-bit mono 16000 Hz",
"transcript_format": "Whisper JSON",
"created_from_existing_project_data": true,
"notes": [
"Deliberately versioned real-life reference sample for reproducible Meeting Lab testing.",
"Contains real meeting data; repository must remain private.",
"Generated chunks, normalized chunks, segmentation outputs, extraction JSON, consolidated outputs and generated protocols are intentionally excluded.",
"No exact human-written reference protocol file was found in the repository workspace, so reference_protocol_file is null."
],
"checksums": {
"audio/meeting_speech.wav": {
"sha256": "d763b8886f089c49ed3bcc2813dfb3a78ae2fd03eee6d843bbee1c52686bf380",
"bytes": 179592270
},
"transcript/meeting_speech.json": {
"sha256": "55ea11571ce43ad299b00362bd1d9b92555aa9d1865779bf9b47ee89138e92ef",
"bytes": 542337
},
"transcript/meeting_speech_cleaned.json": {
"sha256": "22e8616a946a3f46f66ea56e8d80174b44f49a1c14dc297ea0405dee0afbdd3c",
"bytes": 800585
}
}
}
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long