Introduce Meeting Context V1 with YAML schema, validation and template. Support optional --meeting-context during chunk extraction. Inject authoritative Meeting Context into extraction prompts. Record Meeting Context provenance in extraction output. Activate todos.md in shared prompt assembly. Strengthen responsibility attribution and decision/todo boundaries. Add focused Gold scenarios and validation tests. Update architecture and pipeline documentation.
6.2 KiB
6.2 KiB
Roadmap
No dates are assigned. Phases describe dependency order, not release promises.
Phase 1 - Stable Local Extraction
Goal:
- Establish reliable per-chunk extraction behavior for core meeting semantics.
Deliverables:
- Stronger Gold Standard coverage across facts, positions, decisions, todos, questions and technical details.
- Gold coverage for responsibility attribution: discussion, objection, role proximity and department mention must not become ownership.
- Improved category prompts.
- Repeatable evaluation workflow.
- Documented prompt experiment log.
Prerequisites:
- Existing chunk extraction flow.
- Existing Gold Standard runner and methodology.
Out of scope:
- Full-transcript LLM extraction.
- Larger context-window strategy changes without an explicit experiment.
- Canonicalization, semantic consolidation or final protocol rendering.
Phase 2 - Deterministic Canonicalization
Goal:
- Convert independent chunk extraction JSON into a validated, normalized, evidence-bearing intermediate representation without semantic guessing.
Deliverables:
- Canonicalizer V1 implemented in Python.
- Stable source references and IDs.
- Normalized category names and basic field structure.
- Safe deterministic cleanup.
- Exact duplicate grouping where unambiguous.
- Preservation of all source evidence.
Prerequisites:
- Stable local extraction baseline.
- Agreement on the extraction object shape that should be canonicalized.
Current status:
- Implemented as
meeting_lab.consolidation.canonicalize.
Out of scope:
- Uncertain semantic merging.
- Topic synthesis.
- Protocol writing.
- LLM calls.
Phase 3 - Semantic Consolidation
Goal:
- Merge canonicalized extraction objects into a coherent semantic meeting representation while preserving evidence and uncertainty.
Deliverables:
- Semantic Consolidator V0 using the local LLM for facts-only duplicate detection.
- Semantically equivalent fact statement merging.
- Evidence preserved from all contributing chunks.
- Complete source fact coverage validation.
- Later broader semantic consolidation with topic grouping, contradiction and uncertainty markers, durable/transient separation and Canonical Meeting Knowledge preparation.
Prerequisites:
- Canonicalizer V1 output with stable IDs and source references.
- Gold or benchmark cases that expose duplication and category shifts.
Current status:
- Semantic Consolidator V0 is implemented and experimentally validated for facts-only conservative merging.
- The first accepted benchmark merged one correct pair among 33 facts and left 31 singleton groups.
Out of scope:
- Direct protocol writing.
- Topic synthesis in V0.
- Processing decisions, action items, questions, positions or technical details in V0.
- Canonical Meeting Knowledge generation in V0.
- Deriving output views from one another.
- Retrieval or RAG integration.
Phase 4 - Canonical Meeting Knowledge
Goal:
- Define and implement the semantic intermediate model that becomes the source of truth for downstream outputs.
Deliverables:
- Canonical Meeting Knowledge schema.
- Source evidence and traceability fields.
- Clear distinction between durable knowledge and meeting-specific actions.
- Migration path from consolidated extraction JSON into the canonical model.
Prerequisites:
- Semantic consolidation behavior that preserves evidence and uncertainty.
- Agreement on required semantic categories.
Out of scope:
- GUI.
- Export formats beyond those needed to validate the model.
- Knowledge-system storage design.
Phase 5 - Output Views
Goal:
- Render purpose-specific outputs from Canonical Meeting Knowledge without changing meaning.
Deliverables:
- Working Protocol / Arbeitsprotokoll renderer.
- Distribution Protocol / Verteilerprotokoll renderer.
- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export.
- Later additional views such as action lists.
- Tests or checks showing that output views are parallel renderings of the same canonical model.
- Default output-language policy: rendered protocols normally match the dominant source language unless explicitly requested otherwise.
Prerequisites:
- Implemented Canonical Meeting Knowledge.
- Clear audience and completeness rules for each output view.
Out of scope:
- Additional analysis during rendering.
- Deriving one output view from another.
- Retrieval integration.
Phase 6 - Review and Quality Control
Goal:
- Add optional review stages that improve omission detection, consistency and model selection.
Deliverables:
- Optional whole-transcript review.
- Omission detection.
- Consistency checks.
- Model comparison workflow.
- Hardware and runtime benchmarks.
Prerequisites:
- Stable extraction, semantic consolidation and canonical model.
- Representative test meetings.
Out of scope:
- Automatic acceptance of review suggestions without evidence.
- Product UI work.
- Cloud deployment.
Phase 7 - Productization
Goal:
- Turn the validated pipeline into a usable local workflow.
Deliverables:
- Recording/transcription workflow.
- FFmpeg integration.
- Meeting metadata capture.
- Participant entry.
- Meeting Context V1 exists as a manually maintained YAML structure with validation, optional extraction prompt integration and extraction provenance; future work should add GUI entry and conservative integration with later pipeline stages.
- GUI.
- Stable deployment process.
- Export workflows.
Prerequisites:
- Stable pipeline stages and output views.
- Clear operational requirements for local use.
Out of scope:
- Enterprise knowledge retrieval.
- Future Meeting Assistant integration beyond export contracts.
- Cloud-first architecture.
Phase 8 - Knowledge-System Integration
Goal:
- Reuse durable meeting knowledge in broader knowledge systems.
Deliverables:
- Structured Knowledge Objects.
- Retrieval-ready storage format.
- Future RAG integration path.
- Reuse contracts for Meeting Assistant and other knowledge systems.
Prerequisites:
- Canonical Meeting Knowledge and Knowledge Objects are implemented and stable.
- Durable knowledge is separated from meeting-specific actions and discussion history.
Out of scope:
- Building a full enterprise search product inside Meeting Lab.
- Treating raw transcripts or generated protocols as the knowledge source of truth.