Implement Semantic Consolidator V0
- add deterministic canonicalization support for extraction items - add facts-only semantic consolidation using local Ollama - preserve source evidence and validate complete fact coverage - add conservative merge rules and non-LLM tests - record the first validated real-life consolidation benchmark - document current scope, limitations and next evaluation step
This commit is contained in:
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,168 @@
|
||||
# Canonicalizer V1 Evaluation Run
|
||||
|
||||
## Scope
|
||||
|
||||
This run evaluates Canonicalizer V1 as deterministic input preparation for a
|
||||
future semantic consolidator. It does not evaluate protocol prose and does not
|
||||
run the semantic consolidator.
|
||||
|
||||
Input files:
|
||||
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_02_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_03_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_04_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_05_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_06_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_07_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_08_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_09_extraction.json`
|
||||
|
||||
Output:
|
||||
|
||||
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
||||
|
||||
## Run Metrics
|
||||
|
||||
- Runtime: 0.04 seconds
|
||||
- Input file count: 9
|
||||
- Input byte size: 22,695 bytes
|
||||
- Output byte size: 115,924 bytes
|
||||
- Exact duplicate count: 0
|
||||
- Merged exact duplicate count: 0
|
||||
- Malformed or unparsed items: 0
|
||||
- Source-reference mismatches found during comparison: 0
|
||||
|
||||
Input item count by category:
|
||||
|
||||
- fact: 33
|
||||
- decision: 14
|
||||
- action_item: 16
|
||||
- open_question: 10
|
||||
- position: 0
|
||||
- technical_detail: 14
|
||||
|
||||
Output item count by category:
|
||||
|
||||
- fact: 33
|
||||
- decision: 14
|
||||
- action_item: 16
|
||||
- open_question: 10
|
||||
- position: 0
|
||||
- technical_detail: 14
|
||||
|
||||
## Representative Inspection
|
||||
|
||||
### Fact
|
||||
|
||||
- Item: `fact_0001`
|
||||
- Source: `chunk_01_extraction.json`, index 0
|
||||
- Text: `Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.`
|
||||
- Speaker: `Martin`
|
||||
- Status: `clear`
|
||||
- Evidence preserved from the original extraction item.
|
||||
|
||||
Assessment: original information and evidence were preserved, and the source
|
||||
reference points back to the correct original item.
|
||||
|
||||
### Decision
|
||||
|
||||
- Item: `decision_0001`
|
||||
- Source: `chunk_02_extraction.json`, index 0
|
||||
- Text: `Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.`
|
||||
- Evidence preserved from the original extraction item.
|
||||
|
||||
Assessment: decision text and evidence were preserved without semantic
|
||||
rewriting.
|
||||
|
||||
### Action Item
|
||||
|
||||
- Item: `action_item_0001`
|
||||
- Source: `chunk_01_extraction.json`, index 0
|
||||
- Text: `Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert.`
|
||||
- Responsible: `Giovanna`
|
||||
- Deadline: `null`
|
||||
- Evidence preserved from the original extraction item.
|
||||
|
||||
Assessment: responsibility was parsed from the legacy string and no deadline
|
||||
was invented.
|
||||
|
||||
### Open Question
|
||||
|
||||
- Item: `open_question_0001`
|
||||
- Source: `chunk_01_extraction.json`, index 0
|
||||
- Text: `Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?`
|
||||
- Evidence preserved from the original extraction item.
|
||||
|
||||
Assessment: question text and evidence were preserved.
|
||||
|
||||
### Position
|
||||
|
||||
The input extraction files contain zero `positions` items. The canonicalized
|
||||
output correctly reports zero `position` items.
|
||||
|
||||
### Technical Detail
|
||||
|
||||
- Item: `technical_detail_0001`
|
||||
- Source: `chunk_01_extraction.json`, index 0
|
||||
- Subject: `Projektdeckblatt Struktur`
|
||||
- Text: `Das ist eine kurze Projektidee, Ziel, Gegenüber, Entwicklungshemmnisse (Patente), Budgetabfrage.`
|
||||
- Status: `clear`
|
||||
- Evidence preserved from the original extraction item.
|
||||
|
||||
Assessment: subject, statement, status and evidence were parsed without adding
|
||||
technical interpretation.
|
||||
|
||||
## Comparison Findings
|
||||
|
||||
- Original information is preserved. All output items retain `original_value`.
|
||||
- Source references are correct. A comparison against the nine original JSON
|
||||
files found zero source-reference mismatches.
|
||||
- No semantic merges occurred. Output item counts match input item counts in
|
||||
every category.
|
||||
- Names and responsibilities were not invented. Missing optional values remain
|
||||
`null`; for example, `action_item_0001` has `deadline: null`.
|
||||
- Similar but non-identical items remain separate. For example, `fact_0025`
|
||||
and `fact_0031` both discuss the F&E lead maintaining the project list, but
|
||||
they have different wording and remain separate items.
|
||||
- The output is suitable as deterministic input for a later LLM consolidator:
|
||||
it has stable IDs, normalized categories, source references, evidence and
|
||||
original values.
|
||||
|
||||
## Conclusion
|
||||
|
||||
1. Is the deterministic representation lossless enough?
|
||||
|
||||
Yes for the current extraction format. The canonicalized output preserves the
|
||||
original value, parsed text fields, evidence and source references for each
|
||||
item. The output is larger than the input because it adds deterministic
|
||||
metadata and source-reference structure.
|
||||
|
||||
2. Are exact duplicates handled correctly?
|
||||
|
||||
Yes for this run. No exact duplicates were present, so no items were merged.
|
||||
The input and output category counts are identical.
|
||||
|
||||
3. Are legacy extraction strings parsed reliably?
|
||||
|
||||
Yes for the inspected current files. Legacy pipe-delimited strings were parsed
|
||||
into category-specific fields such as `speaker`, `status`, `responsible`,
|
||||
`deadline` and `subject` where safely available. No malformed or unparsed items
|
||||
were detected.
|
||||
|
||||
4. Does the result reduce avoidable work for the semantic consolidator?
|
||||
|
||||
Yes. The future consolidator can consume normalized categories, stable item
|
||||
IDs, source references, evidence and parsed optional fields instead of
|
||||
re-reading heterogeneous legacy strings directly.
|
||||
|
||||
5. What unresolved transformations must remain an LLM task?
|
||||
|
||||
- Merging semantically equivalent but differently worded statements.
|
||||
- Grouping items into coherent topics.
|
||||
- Reconciling facts, decisions, action items and questions across chunks.
|
||||
- Detecting contradictions and uncertainty.
|
||||
- Separating durable organizational knowledge from transient discussion.
|
||||
- Deciding whether similar items such as `fact_0025` and `fact_0031` should be
|
||||
merged, related or kept separate.
|
||||
|
||||
Reference in New Issue
Block a user