# Canonicalizer V1 Evaluation Run ## Scope This run evaluates Canonicalizer V1 as deterministic input preparation for a future semantic consolidator. It does not evaluate protocol prose and does not run the semantic consolidator. Input files: - `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_02_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_03_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_04_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_05_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_06_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_07_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_08_extraction.json` - `samples/whisper/meeting_speech_cleaned_chunks/chunk_09_extraction.json` Output: - `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json` ## Run Metrics - Runtime: 0.04 seconds - Input file count: 9 - Input byte size: 22,695 bytes - Output byte size: 115,924 bytes - Exact duplicate count: 0 - Merged exact duplicate count: 0 - Malformed or unparsed items: 0 - Source-reference mismatches found during comparison: 0 Input item count by category: - fact: 33 - decision: 14 - action_item: 16 - open_question: 10 - position: 0 - technical_detail: 14 Output item count by category: - fact: 33 - decision: 14 - action_item: 16 - open_question: 10 - position: 0 - technical_detail: 14 ## Representative Inspection ### Fact - Item: `fact_0001` - Source: `chunk_01_extraction.json`, index 0 - Text: `Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.` - Speaker: `Martin` - Status: `clear` - Evidence preserved from the original extraction item. Assessment: original information and evidence were preserved, and the source reference points back to the correct original item. ### Decision - Item: `decision_0001` - Source: `chunk_02_extraction.json`, index 0 - Text: `Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.` - Evidence preserved from the original extraction item. Assessment: decision text and evidence were preserved without semantic rewriting. ### Action Item - Item: `action_item_0001` - Source: `chunk_01_extraction.json`, index 0 - Text: `Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert.` - Responsible: `Giovanna` - Deadline: `null` - Evidence preserved from the original extraction item. Assessment: responsibility was parsed from the legacy string and no deadline was invented. ### Open Question - Item: `open_question_0001` - Source: `chunk_01_extraction.json`, index 0 - Text: `Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?` - Evidence preserved from the original extraction item. Assessment: question text and evidence were preserved. ### Position The input extraction files contain zero `positions` items. The canonicalized output correctly reports zero `position` items. ### Technical Detail - Item: `technical_detail_0001` - Source: `chunk_01_extraction.json`, index 0 - Subject: `Projektdeckblatt Struktur` - Text: `Das ist eine kurze Projektidee, Ziel, Gegenüber, Entwicklungshemmnisse (Patente), Budgetabfrage.` - Status: `clear` - Evidence preserved from the original extraction item. Assessment: subject, statement, status and evidence were parsed without adding technical interpretation. ## Comparison Findings - Original information is preserved. All output items retain `original_value`. - Source references are correct. A comparison against the nine original JSON files found zero source-reference mismatches. - No semantic merges occurred. Output item counts match input item counts in every category. - Names and responsibilities were not invented. Missing optional values remain `null`; for example, `action_item_0001` has `deadline: null`. - Similar but non-identical items remain separate. For example, `fact_0025` and `fact_0031` both discuss the F&E lead maintaining the project list, but they have different wording and remain separate items. - The output is suitable as deterministic input for a later LLM consolidator: it has stable IDs, normalized categories, source references, evidence and original values. ## Conclusion 1. Is the deterministic representation lossless enough? Yes for the current extraction format. The canonicalized output preserves the original value, parsed text fields, evidence and source references for each item. The output is larger than the input because it adds deterministic metadata and source-reference structure. 2. Are exact duplicates handled correctly? Yes for this run. No exact duplicates were present, so no items were merged. The input and output category counts are identical. 3. Are legacy extraction strings parsed reliably? Yes for the inspected current files. Legacy pipe-delimited strings were parsed into category-specific fields such as `speaker`, `status`, `responsible`, `deadline` and `subject` where safely available. No malformed or unparsed items were detected. 4. Does the result reduce avoidable work for the semantic consolidator? Yes. The future consolidator can consume normalized categories, stable item IDs, source references, evidence and parsed optional fields instead of re-reading heterogeneous legacy strings directly. 5. What unresolved transformations must remain an LLM task? - Merging semantically equivalent but differently worded statements. - Grouping items into coherent topics. - Reconciling facts, decisions, action items and questions across chunks. - Detecting contradictions and uncertainty. - Separating durable organizational knowledge from transient discussion. - Deciding whether similar items such as `fact_0025` and `fact_0031` should be merged, related or kept separate.