Files
meeting-lab/ROADMAP.md
T
admin 63075eaca9 Enforce explicit responsibility attribution
- document responsibility attribution as a project-wide invariant
- prevent inferred ownership in protocol rendering
- add negative gold regression for false responsibility assignment
- document future responsibility evidence and attribution model
- record the real-life benchmark finding
2026-07-31 13:01:11 +02:00

235 lines
6.0 KiB
Markdown

# Roadmap
No dates are assigned. Phases describe dependency order, not release promises.
## Phase 1 - Stable Local Extraction
Goal:
- Establish reliable per-chunk extraction behavior for core meeting semantics.
Deliverables:
- Stronger Gold Standard coverage across facts, positions, decisions, todos,
questions and technical details.
- Gold coverage for responsibility attribution: discussion, objection, role
proximity and department mention must not become ownership.
- Improved category prompts.
- Repeatable evaluation workflow.
- Documented prompt experiment log.
Prerequisites:
- Existing chunk extraction flow.
- Existing Gold Standard runner and methodology.
Out of scope:
- Full-transcript LLM extraction.
- Larger context-window strategy changes without an explicit experiment.
- Canonicalization, semantic consolidation or final protocol rendering.
## Phase 2 - Deterministic Canonicalization
Goal:
- Convert independent chunk extraction JSON into a validated, normalized,
evidence-bearing intermediate representation without semantic guessing.
Deliverables:
- Canonicalizer V1 implemented in Python.
- Stable source references and IDs.
- Normalized category names and basic field structure.
- Safe deterministic cleanup.
- Exact duplicate grouping where unambiguous.
- Preservation of all source evidence.
Prerequisites:
- Stable local extraction baseline.
- Agreement on the extraction object shape that should be canonicalized.
Current status:
- Implemented as `meeting_lab.consolidation.canonicalize`.
Out of scope:
- Uncertain semantic merging.
- Topic synthesis.
- Protocol writing.
- LLM calls.
## Phase 3 - Semantic Consolidation
Goal:
- Merge canonicalized extraction objects into a coherent semantic meeting
representation while preserving evidence and uncertainty.
Deliverables:
- Semantic Consolidator V0 using the local LLM for facts-only duplicate
detection.
- Semantically equivalent fact statement merging.
- Evidence preserved from all contributing chunks.
- Complete source fact coverage validation.
- Later broader semantic consolidation with topic grouping, contradiction and
uncertainty markers, durable/transient separation and Canonical Meeting
Knowledge preparation.
Prerequisites:
- Canonicalizer V1 output with stable IDs and source references.
- Gold or benchmark cases that expose duplication and category shifts.
Current status:
- Semantic Consolidator V0 is implemented and experimentally validated for
facts-only conservative merging.
- The first accepted benchmark merged one correct pair among 33 facts and left
31 singleton groups.
Out of scope:
- Direct protocol writing.
- Topic synthesis in V0.
- Processing decisions, action items, questions, positions or technical details
in V0.
- Canonical Meeting Knowledge generation in V0.
- Deriving output views from one another.
- Retrieval or RAG integration.
## Phase 4 - Canonical Meeting Knowledge
Goal:
- Define and implement the semantic intermediate model that becomes the source
of truth for downstream outputs.
Deliverables:
- Canonical Meeting Knowledge schema.
- Source evidence and traceability fields.
- Clear distinction between durable knowledge and meeting-specific actions.
- Migration path from consolidated extraction JSON into the canonical model.
Prerequisites:
- Semantic consolidation behavior that preserves evidence and uncertainty.
- Agreement on required semantic categories.
Out of scope:
- GUI.
- Export formats beyond those needed to validate the model.
- Knowledge-system storage design.
## Phase 5 - Output Views
Goal:
- Render purpose-specific outputs from Canonical Meeting Knowledge without
changing meaning.
Deliverables:
- Working Protocol / Arbeitsprotokoll renderer.
- Distribution Protocol / Verteilerprotokoll renderer.
- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export.
- Later additional views such as action lists.
- Tests or checks showing that output views are parallel renderings of the same
canonical model.
- Default output-language policy: rendered protocols normally match the
dominant source language unless explicitly requested otherwise.
Prerequisites:
- Implemented Canonical Meeting Knowledge.
- Clear audience and completeness rules for each output view.
Out of scope:
- Additional analysis during rendering.
- Deriving one output view from another.
- Retrieval integration.
## Phase 6 - Review and Quality Control
Goal:
- Add optional review stages that improve omission detection, consistency and
model selection.
Deliverables:
- Optional whole-transcript review.
- Omission detection.
- Consistency checks.
- Model comparison workflow.
- Hardware and runtime benchmarks.
Prerequisites:
- Stable extraction, semantic consolidation and canonical model.
- Representative test meetings.
Out of scope:
- Automatic acceptance of review suggestions without evidence.
- Product UI work.
- Cloud deployment.
## Phase 7 - Productization
Goal:
- Turn the validated pipeline into a usable local workflow.
Deliverables:
- Recording/transcription workflow.
- FFmpeg integration.
- Meeting metadata capture.
- Participant entry.
- GUI.
- Stable deployment process.
- Export workflows.
Prerequisites:
- Stable pipeline stages and output views.
- Clear operational requirements for local use.
Out of scope:
- Enterprise knowledge retrieval.
- Future Meeting Assistant integration beyond export contracts.
- Cloud-first architecture.
## Phase 8 - Knowledge-System Integration
Goal:
- Reuse durable meeting knowledge in broader knowledge systems.
Deliverables:
- Structured Knowledge Objects.
- Retrieval-ready storage format.
- Future RAG integration path.
- Reuse contracts for Meeting Assistant and other knowledge systems.
Prerequisites:
- Canonical Meeting Knowledge and Knowledge Objects are implemented and stable.
- Durable knowledge is separated from meeting-specific actions and discussion
history.
Out of scope:
- Building a full enterprise search product inside Meeting Lab.
- Treating raw transcripts or generated protocols as the knowledge source of
truth.