- add deterministic canonicalization support for extraction items - add facts-only semantic consolidation using local Ollama - preserve source evidence and validate complete fact coverage - add conservative merge rules and non-LLM tests - record the first validated real-life consolidation benchmark - document current scope, limitations and next evaluation step
169 lines
6.0 KiB
Markdown
169 lines
6.0 KiB
Markdown
# Canonicalizer V1 Evaluation Run
|
|
|
|
## Scope
|
|
|
|
This run evaluates Canonicalizer V1 as deterministic input preparation for a
|
|
future semantic consolidator. It does not evaluate protocol prose and does not
|
|
run the semantic consolidator.
|
|
|
|
Input files:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_02_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_03_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_04_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_05_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_06_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_07_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_08_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_09_extraction.json`
|
|
|
|
Output:
|
|
|
|
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
|
|
|
## Run Metrics
|
|
|
|
- Runtime: 0.04 seconds
|
|
- Input file count: 9
|
|
- Input byte size: 22,695 bytes
|
|
- Output byte size: 115,924 bytes
|
|
- Exact duplicate count: 0
|
|
- Merged exact duplicate count: 0
|
|
- Malformed or unparsed items: 0
|
|
- Source-reference mismatches found during comparison: 0
|
|
|
|
Input item count by category:
|
|
|
|
- fact: 33
|
|
- decision: 14
|
|
- action_item: 16
|
|
- open_question: 10
|
|
- position: 0
|
|
- technical_detail: 14
|
|
|
|
Output item count by category:
|
|
|
|
- fact: 33
|
|
- decision: 14
|
|
- action_item: 16
|
|
- open_question: 10
|
|
- position: 0
|
|
- technical_detail: 14
|
|
|
|
## Representative Inspection
|
|
|
|
### Fact
|
|
|
|
- Item: `fact_0001`
|
|
- Source: `chunk_01_extraction.json`, index 0
|
|
- Text: `Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.`
|
|
- Speaker: `Martin`
|
|
- Status: `clear`
|
|
- Evidence preserved from the original extraction item.
|
|
|
|
Assessment: original information and evidence were preserved, and the source
|
|
reference points back to the correct original item.
|
|
|
|
### Decision
|
|
|
|
- Item: `decision_0001`
|
|
- Source: `chunk_02_extraction.json`, index 0
|
|
- Text: `Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.`
|
|
- Evidence preserved from the original extraction item.
|
|
|
|
Assessment: decision text and evidence were preserved without semantic
|
|
rewriting.
|
|
|
|
### Action Item
|
|
|
|
- Item: `action_item_0001`
|
|
- Source: `chunk_01_extraction.json`, index 0
|
|
- Text: `Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert.`
|
|
- Responsible: `Giovanna`
|
|
- Deadline: `null`
|
|
- Evidence preserved from the original extraction item.
|
|
|
|
Assessment: responsibility was parsed from the legacy string and no deadline
|
|
was invented.
|
|
|
|
### Open Question
|
|
|
|
- Item: `open_question_0001`
|
|
- Source: `chunk_01_extraction.json`, index 0
|
|
- Text: `Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?`
|
|
- Evidence preserved from the original extraction item.
|
|
|
|
Assessment: question text and evidence were preserved.
|
|
|
|
### Position
|
|
|
|
The input extraction files contain zero `positions` items. The canonicalized
|
|
output correctly reports zero `position` items.
|
|
|
|
### Technical Detail
|
|
|
|
- Item: `technical_detail_0001`
|
|
- Source: `chunk_01_extraction.json`, index 0
|
|
- Subject: `Projektdeckblatt Struktur`
|
|
- Text: `Das ist eine kurze Projektidee, Ziel, Gegenüber, Entwicklungshemmnisse (Patente), Budgetabfrage.`
|
|
- Status: `clear`
|
|
- Evidence preserved from the original extraction item.
|
|
|
|
Assessment: subject, statement, status and evidence were parsed without adding
|
|
technical interpretation.
|
|
|
|
## Comparison Findings
|
|
|
|
- Original information is preserved. All output items retain `original_value`.
|
|
- Source references are correct. A comparison against the nine original JSON
|
|
files found zero source-reference mismatches.
|
|
- No semantic merges occurred. Output item counts match input item counts in
|
|
every category.
|
|
- Names and responsibilities were not invented. Missing optional values remain
|
|
`null`; for example, `action_item_0001` has `deadline: null`.
|
|
- Similar but non-identical items remain separate. For example, `fact_0025`
|
|
and `fact_0031` both discuss the F&E lead maintaining the project list, but
|
|
they have different wording and remain separate items.
|
|
- The output is suitable as deterministic input for a later LLM consolidator:
|
|
it has stable IDs, normalized categories, source references, evidence and
|
|
original values.
|
|
|
|
## Conclusion
|
|
|
|
1. Is the deterministic representation lossless enough?
|
|
|
|
Yes for the current extraction format. The canonicalized output preserves the
|
|
original value, parsed text fields, evidence and source references for each
|
|
item. The output is larger than the input because it adds deterministic
|
|
metadata and source-reference structure.
|
|
|
|
2. Are exact duplicates handled correctly?
|
|
|
|
Yes for this run. No exact duplicates were present, so no items were merged.
|
|
The input and output category counts are identical.
|
|
|
|
3. Are legacy extraction strings parsed reliably?
|
|
|
|
Yes for the inspected current files. Legacy pipe-delimited strings were parsed
|
|
into category-specific fields such as `speaker`, `status`, `responsible`,
|
|
`deadline` and `subject` where safely available. No malformed or unparsed items
|
|
were detected.
|
|
|
|
4. Does the result reduce avoidable work for the semantic consolidator?
|
|
|
|
Yes. The future consolidator can consume normalized categories, stable item
|
|
IDs, source references, evidence and parsed optional fields instead of
|
|
re-reading heterogeneous legacy strings directly.
|
|
|
|
5. What unresolved transformations must remain an LLM task?
|
|
|
|
- Merging semantically equivalent but differently worded statements.
|
|
- Grouping items into coherent topics.
|
|
- Reconciling facts, decisions, action items and questions across chunks.
|
|
- Detecting contradictions and uncertainty.
|
|
- Separating durable organizational knowledge from transient discussion.
|
|
- Deciding whether similar items such as `fact_0025` and `fact_0031` should be
|
|
merged, related or kept separate.
|
|
|