Files
admin 06f0e7e651 Implement Meeting Context V1 and extraction improvements
Introduce Meeting Context V1 with YAML schema, validation and template.
Support optional --meeting-context during chunk extraction.
Inject authoritative Meeting Context into extraction prompts.
Record Meeting Context provenance in extraction output.
Activate todos.md in shared prompt assembly.
Strengthen responsibility attribution and decision/todo boundaries.
Add focused Gold scenarios and validation tests.
Update architecture and pipeline documentation.
2026-08-01 16:49:48 +02:00

2.1 KiB

Project Process Meeting Real-Life Sample

Purpose

This directory preserves the complete input data for the current real meeting experiment so the Meeting Lab pipeline can be reproduced on another computer.

This is a deliberately selected reference sample, not ordinary generated runtime data.

Confidentiality

This sample contains real meeting data. Keep the repository private. Do not copy, publish or redistribute these files outside the intended private development context.

Avoid exposing unnecessary personal details in derived documentation, reports or screenshots.

Language

The meeting language is German.

Contents

Original input files:

  • audio/meeting_speech.wav: source meeting audio.
  • transcript/meeting_speech.json: raw Whisper transcript JSON.
  • transcript/meeting_speech_cleaned.json: cleaned Whisper transcript JSON.

Reference material:

  • meeting_context.yaml: manually maintained Meeting Context V1 scaffold for reliable metadata. It can be validated and injected into chunk extraction with --meeting-context; later pipeline stages are not connected yet.
  • reference/human_reference_protocol.md: not currently included. No exact existing human-written protocol file was found in the repository workspace.

Generated outputs are intentionally excluded. This sample does not include generated chunks, normalized chunks, segmentation outputs, extraction JSON files, consolidated outputs, generated protocols, raw model responses or temporary files.

Intended Pipeline Use

This sample may be used to reproduce and test:

  • audio/transcription workflow setup
  • Whisper JSON cleanup behavior
  • transcript chunking
  • normalization
  • segmentation experiments
  • extraction experiments
  • canonicalization and future consolidation experiments

Generated artifacts should be written to normal experiment or benchmark output locations, not back into this input sample directory.

Git LFS

The WAV audio is intentionally versioned through Git LFS using the scoped repository rule for samples/real_live/**/*.wav.

After cloning the repository, retrieve LFS files with:

git lfs pull