Add benchmark artifacts created during the first RC1 robustness evaluation of the Meeting Lab pipeline. Included benchmark sets: - hardware_experimental_qwen35_9b - meeting_context_v1 - progeo_meeting_20260804_083849 - progeo_meeting_context_v1_20260804_110913 - progeo_meeting_rc1_20260804_121502 These artifacts document the evolution of the pipeline during the implementation and verification of BUG-009, BUG-010 and BUG-011. The benchmark data provide reproducible real-life regression cases for future development and allow quality comparisons across pipeline revisions. Current benchmark policy: During the active development phase, representative benchmark artifacts are intentionally versioned to preserve reproducibility and simplify regression analysis. Benchmark artifacts are considered part of the engineering evidence rather than temporary build output. Benchmark retention strategy will be revisited once the Meeting Assistant reaches production maturity.
8 lines
1.2 KiB
Plaintext
8 lines
1.2 KiB
Plaintext
Input directory: samples\benchmarks\progeo_meeting_20260804_083849\extractions
|
|
Input files processed: 10
|
|
Input item count by category: {'fact': 57, 'decision': 4, 'action_item': 32, 'open_question': 17, 'position': 0, 'technical_detail': 27}
|
|
Output item count by category: {'fact': 57, 'decision': 4, 'action_item': 31, 'open_question': 17, 'position': 0, 'technical_detail': 27}
|
|
Exact duplicates merged: 1
|
|
Output: samples\benchmarks\progeo_meeting_20260804_083849\canonicalizer\canonicalized_extractions.json
|
|
Output JSON size: 186499 bytes
|