- add second real-life Whisper transcript - add cleaned transcript - add Meeting Context scaffold - document evaluation characteristics - prepare reusable reference dataset
47 lines
1.4 KiB
Markdown
47 lines
1.4 KiB
Markdown
# Progeo Meeting
|
|
|
|
Real-Life Reference Meeting #2 for Meeting Lab.
|
|
|
|
This sample provides an additional real meeting dataset for evaluating the
|
|
pipeline beyond the project-process meeting. It is intended for quality
|
|
assessment, regression analysis and comparison of pipeline behaviour across
|
|
different meeting structures.
|
|
|
|
## Files
|
|
|
|
- `progeo_meeting_speech.json`: original Whisper JSON transcript.
|
|
- `progeo_meeting_speech_cleaned.json`: cleaned Whisper JSON transcript
|
|
generated with `scripts/clean_whisper_json.py`.
|
|
- `meeting_context.yaml`: Meeting Context V1 scaffold.
|
|
|
|
## Meeting Context Status
|
|
|
|
The Meeting Context is intentionally incomplete. It contains only known
|
|
meeting-level metadata and no inferred participants.
|
|
|
|
Participant, mentioned-person, organization, department, alias and known-entity
|
|
metadata require manual completion before context-aware extraction should be
|
|
treated as authoritative for this sample.
|
|
|
|
## Known Characteristics
|
|
|
|
- Strongly uneven speaking shares.
|
|
- Comparatively stable discussion topics.
|
|
- Intended as an additional evaluation dataset.
|
|
|
|
## Preparation Notes
|
|
|
|
The cleaned transcript was created with:
|
|
|
|
```text
|
|
.venv/bin/python scripts/clean_whisper_json.py \
|
|
samples/real_live/progeo_meeting/progeo_meeting_speech.json \
|
|
-o samples/real_live/progeo_meeting/progeo_meeting_speech_cleaned.json
|
|
```
|
|
|
|
Cleaning result:
|
|
|
|
- Segments before: 1448
|
|
- Segments after: 1448
|
|
- Segments removed: 0
|