Add second real-life reference meeting (Progeo)

- add second real-life Whisper transcript
- add cleaned transcript
- add Meeting Context scaffold
- document evaluation characteristics
- prepare reusable reference dataset
This commit is contained in:
2026-08-03 16:11:46 +02:00
parent 9446c6e0be
commit 60a8acae91
4 changed files with 41942 additions and 0 deletions
@@ -0,0 +1,46 @@
# Progeo Meeting
Real-Life Reference Meeting #2 for Meeting Lab.
This sample provides an additional real meeting dataset for evaluating the
pipeline beyond the project-process meeting. It is intended for quality
assessment, regression analysis and comparison of pipeline behaviour across
different meeting structures.
## Files
- `progeo_meeting_speech.json`: original Whisper JSON transcript.
- `progeo_meeting_speech_cleaned.json`: cleaned Whisper JSON transcript
generated with `scripts/clean_whisper_json.py`.
- `meeting_context.yaml`: Meeting Context V1 scaffold.
## Meeting Context Status
The Meeting Context is intentionally incomplete. It contains only known
meeting-level metadata and no inferred participants.
Participant, mentioned-person, organization, department, alias and known-entity
metadata require manual completion before context-aware extraction should be
treated as authoritative for this sample.
## Known Characteristics
- Strongly uneven speaking shares.
- Comparatively stable discussion topics.
- Intended as an additional evaluation dataset.
## Preparation Notes
The cleaned transcript was created with:
```text
.venv/bin/python scripts/clean_whisper_json.py \
samples/real_live/progeo_meeting/progeo_meeting_speech.json \
-o samples/real_live/progeo_meeting/progeo_meeting_speech_cleaned.json
```
Cleaning result:
- Segments before: 1448
- Segments after: 1448
- Segments removed: 0