Files
meeting-lab/samples/benchmarks/canonicalizer_v1/report.md
T
admin 90aa34d5d0 Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items
- add facts-only semantic consolidation using local Ollama
- preserve source evidence and validate complete fact coverage
- add conservative merge rules and non-LLM tests
- record the first validated real-life consolidation benchmark
- document current scope, limitations and next evaluation step
2026-07-31 11:25:46 +02:00

6.0 KiB

Canonicalizer V1 Evaluation Run

Scope

This run evaluates Canonicalizer V1 as deterministic input preparation for a future semantic consolidator. It does not evaluate protocol prose and does not run the semantic consolidator.

Input files:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_02_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_03_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_04_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_05_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_06_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_07_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_08_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_09_extraction.json

Output:

  • samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json

Run Metrics

  • Runtime: 0.04 seconds
  • Input file count: 9
  • Input byte size: 22,695 bytes
  • Output byte size: 115,924 bytes
  • Exact duplicate count: 0
  • Merged exact duplicate count: 0
  • Malformed or unparsed items: 0
  • Source-reference mismatches found during comparison: 0

Input item count by category:

  • fact: 33
  • decision: 14
  • action_item: 16
  • open_question: 10
  • position: 0
  • technical_detail: 14

Output item count by category:

  • fact: 33
  • decision: 14
  • action_item: 16
  • open_question: 10
  • position: 0
  • technical_detail: 14

Representative Inspection

Fact

  • Item: fact_0001
  • Source: chunk_01_extraction.json, index 0
  • Text: Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.
  • Speaker: Martin
  • Status: clear
  • Evidence preserved from the original extraction item.

Assessment: original information and evidence were preserved, and the source reference points back to the correct original item.

Decision

  • Item: decision_0001
  • Source: chunk_02_extraction.json, index 0
  • Text: Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.
  • Evidence preserved from the original extraction item.

Assessment: decision text and evidence were preserved without semantic rewriting.

Action Item

  • Item: action_item_0001
  • Source: chunk_01_extraction.json, index 0
  • Text: Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert.
  • Responsible: Giovanna
  • Deadline: null
  • Evidence preserved from the original extraction item.

Assessment: responsibility was parsed from the legacy string and no deadline was invented.

Open Question

  • Item: open_question_0001
  • Source: chunk_01_extraction.json, index 0
  • Text: Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?
  • Evidence preserved from the original extraction item.

Assessment: question text and evidence were preserved.

Position

The input extraction files contain zero positions items. The canonicalized output correctly reports zero position items.

Technical Detail

  • Item: technical_detail_0001
  • Source: chunk_01_extraction.json, index 0
  • Subject: Projektdeckblatt Struktur
  • Text: Das ist eine kurze Projektidee, Ziel, Gegenüber, Entwicklungshemmnisse (Patente), Budgetabfrage.
  • Status: clear
  • Evidence preserved from the original extraction item.

Assessment: subject, statement, status and evidence were parsed without adding technical interpretation.

Comparison Findings

  • Original information is preserved. All output items retain original_value.
  • Source references are correct. A comparison against the nine original JSON files found zero source-reference mismatches.
  • No semantic merges occurred. Output item counts match input item counts in every category.
  • Names and responsibilities were not invented. Missing optional values remain null; for example, action_item_0001 has deadline: null.
  • Similar but non-identical items remain separate. For example, fact_0025 and fact_0031 both discuss the F&E lead maintaining the project list, but they have different wording and remain separate items.
  • The output is suitable as deterministic input for a later LLM consolidator: it has stable IDs, normalized categories, source references, evidence and original values.

Conclusion

  1. Is the deterministic representation lossless enough?

Yes for the current extraction format. The canonicalized output preserves the original value, parsed text fields, evidence and source references for each item. The output is larger than the input because it adds deterministic metadata and source-reference structure.

  1. Are exact duplicates handled correctly?

Yes for this run. No exact duplicates were present, so no items were merged. The input and output category counts are identical.

  1. Are legacy extraction strings parsed reliably?

Yes for the inspected current files. Legacy pipe-delimited strings were parsed into category-specific fields such as speaker, status, responsible, deadline and subject where safely available. No malformed or unparsed items were detected.

  1. Does the result reduce avoidable work for the semantic consolidator?

Yes. The future consolidator can consume normalized categories, stable item IDs, source references, evidence and parsed optional fields instead of re-reading heterogeneous legacy strings directly.

  1. What unresolved transformations must remain an LLM task?
  • Merging semantically equivalent but differently worded statements.
  • Grouping items into coherent topics.
  • Reconciling facts, decisions, action items and questions across chunks.
  • Detecting contradictions and uncertainty.
  • Separating durable organizational knowledge from transient discussion.
  • Deciding whether similar items such as fact_0025 and fact_0031 should be merged, related or kept separate.