Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
490d557cb3 | ||
|
|
e028f1c82c |
@@ -1,5 +1,19 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
|
## [Unreleased]
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- Reconciled the documented MVP with the validated Meeting Lab backend.
|
||||||
|
- Selected `whisper.cpp` for transcription and made diarization optional.
|
||||||
|
- Defined the Python API, progress-event and Meeting Context integration
|
||||||
|
boundaries.
|
||||||
|
- Set the user-facing Meeting Assistant GUI as the next milestone.
|
||||||
|
|
||||||
|
### Added
|
||||||
|
|
||||||
|
- ADR 0011 for the Meeting Lab MVP backend architecture.
|
||||||
|
|
||||||
## [0.1.0] - 2026-07-14
|
## [0.1.0] - 2026-07-14
|
||||||
|
|
||||||
### Added
|
### Added
|
||||||
|
|||||||
+103
-153
@@ -2,162 +2,112 @@
|
|||||||
|
|
||||||
## Vision
|
## Vision
|
||||||
|
|
||||||
The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge.
|
Meeting Assistant turns recorded meetings into reviewed, user-facing protocols
|
||||||
|
and, over time, searchable organizational knowledge. It is the application
|
||||||
|
layer over the reusable Meeting Lab processing backend.
|
||||||
|
|
||||||
The transcript is the canonical source.
|
The source audio and canonical transcript are preserved. Generated protocols
|
||||||
|
are derived, reproducible artifacts and require human review.
|
||||||
|
|
||||||
AI-generated summaries are reproducible artefacts.
|
## System Boundary
|
||||||
|
|
||||||
---
|
### Meeting Lab
|
||||||
|
|
||||||
|
Meeting Lab owns reusable and experimental processing:
|
||||||
|
|
||||||
|
- FFmpeg audio preparation
|
||||||
|
- `whisper.cpp` transcription
|
||||||
|
- optional `pyannote.audio` diarization
|
||||||
|
- protocol generation
|
||||||
|
- pipeline orchestration and progress events
|
||||||
|
|
||||||
|
Its reusable integration surface is the Python API:
|
||||||
|
|
||||||
|
- `MvpMeetingConfig`
|
||||||
|
- `MvpRunResult`
|
||||||
|
- `run_mvp_meeting(...)`
|
||||||
|
|
||||||
|
The CLI is a thin adapter, not the Meeting Assistant integration boundary.
|
||||||
|
|
||||||
|
### Meeting Assistant
|
||||||
|
|
||||||
|
Meeting Assistant owns user interaction and product workflow:
|
||||||
|
|
||||||
|
- audio-file selection
|
||||||
|
- structured Meeting Context editing
|
||||||
|
- participant management
|
||||||
|
- optional explicit speaker mapping
|
||||||
|
- pipeline launch and progress display
|
||||||
|
- protocol review and editing
|
||||||
|
- export and presentation of protocol versions
|
||||||
|
|
||||||
|
Processing logic must not be duplicated in the application.
|
||||||
|
|
||||||
|
## Validated MVP Pipeline
|
||||||
|
|
||||||
|
```text
|
||||||
|
Imported audio
|
||||||
|
-> prepare as mono, 16 kHz PCM WAV with FFmpeg
|
||||||
|
-> transcribe with whisper.cpp and large-v3-turbo
|
||||||
|
-> optionally diarize with pyannote.audio Community-1
|
||||||
|
-> generate a protocol directly from the full transcript
|
||||||
|
-> review and edit by a human
|
||||||
|
```
|
||||||
|
|
||||||
|
Mandatory fixed five-minute chunking is not part of the productive MVP path.
|
||||||
|
Chunking or segmentation used internally by an engine does not shape the
|
||||||
|
Meeting Assistant storage or domain model.
|
||||||
|
|
||||||
|
On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070
|
||||||
|
transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through
|
||||||
|
PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These
|
||||||
|
are validation observations, not performance guarantees or hardware
|
||||||
|
requirements. CPU execution remains supported and may be substantially slower.
|
||||||
|
|
||||||
|
## Meeting Context and Speakers
|
||||||
|
|
||||||
|
`MeetingContext` is a structured domain object containing meeting metadata and
|
||||||
|
participants. It can also contain optional explicit mappings from
|
||||||
|
`SPEAKER_XX` labels to `participant_id` values.
|
||||||
|
|
||||||
|
A speaker mapping is authoritative only when a user explicitly confirms it.
|
||||||
|
Speakers otherwise remain anonymous. Automatic speaker-name inference is not
|
||||||
|
allowed. The GUI must create and edit Meeting Context; hand-written YAML is not
|
||||||
|
a product requirement.
|
||||||
|
|
||||||
|
## Progress Contract
|
||||||
|
|
||||||
|
Meeting Lab emits stage-based progress events for:
|
||||||
|
|
||||||
|
- `preparing`
|
||||||
|
- `transcription`
|
||||||
|
- `diarization`
|
||||||
|
- `protocol_generation`
|
||||||
|
- `completed`
|
||||||
|
- `failed`
|
||||||
|
|
||||||
|
The GUI consumes this event interface. Percentages are shown only when real
|
||||||
|
measurable progress is available.
|
||||||
|
|
||||||
|
## Protocol Policy
|
||||||
|
|
||||||
|
Direct full-transcript protocol generation is the practical MVP direction.
|
||||||
|
`qwen3.8:27B` has shown strong readability and contextual synthesis;
|
||||||
|
`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas.
|
||||||
|
Experimental dual-model and diarization-assisted hard-fact extraction has not
|
||||||
|
demonstrated reliably better strict attribution accuracy and is not mandatory.
|
||||||
|
|
||||||
|
The product should ultimately offer both a detailed contextual protocol and a
|
||||||
|
shorter participant/distribution version. The short version remains follow-up
|
||||||
|
work if it is not available for the first GUI milestone.
|
||||||
|
|
||||||
## Design Principles
|
## Design Principles
|
||||||
|
|
||||||
Offline-first where practical.
|
- offline-first where practical
|
||||||
|
- explicit application/backend boundaries
|
||||||
Replaceable AI components.
|
- replaceable AI components
|
||||||
|
- CPU-compatible operation with optional acceleration
|
||||||
Vendor independence.
|
- immutable source artifacts and traceable derived artifacts
|
||||||
|
- no automatic identity claims
|
||||||
Modular architecture.
|
- human review of generated protocols
|
||||||
|
- small modules and simple interfaces
|
||||||
Small focused modules.
|
|
||||||
|
|
||||||
Simple interfaces.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Core Pipeline
|
|
||||||
|
|
||||||
Recorder
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Transcription
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Speaker Identification
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
LLM Analysis
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Knowledge Extraction
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Storage
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Export
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## AI Components
|
|
||||||
|
|
||||||
Recorder
|
|
||||||
|
|
||||||
Responsible only for recording.
|
|
||||||
|
|
||||||
No AI.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Transcription
|
|
||||||
|
|
||||||
Responsible only for speech-to-text.
|
|
||||||
|
|
||||||
No summarization.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Speaker Identification
|
|
||||||
|
|
||||||
Responsible only for identifying speakers.
|
|
||||||
|
|
||||||
Must not modify transcript text.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
LLM Analysis
|
|
||||||
|
|
||||||
Responsible for:
|
|
||||||
|
|
||||||
- Summary
|
|
||||||
|
|
||||||
- Decisions
|
|
||||||
|
|
||||||
- Action Items
|
|
||||||
|
|
||||||
- Risks
|
|
||||||
|
|
||||||
- Questions
|
|
||||||
|
|
||||||
- Knowledge extraction
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Storage
|
|
||||||
|
|
||||||
Stores:
|
|
||||||
|
|
||||||
Audio
|
|
||||||
|
|
||||||
Transcript
|
|
||||||
|
|
||||||
Metadata
|
|
||||||
|
|
||||||
AI Results
|
|
||||||
|
|
||||||
Knowledge Objects
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Architectural Rule
|
|
||||||
|
|
||||||
Every module has exactly one responsibility.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Transcript Policy
|
|
||||||
|
|
||||||
Never modify the original transcript.
|
|
||||||
|
|
||||||
Corrections must create a new derived version.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## AI Output Policy
|
|
||||||
|
|
||||||
Every AI output should be reproducible.
|
|
||||||
|
|
||||||
Prompt version should be stored.
|
|
||||||
|
|
||||||
Model should be stored.
|
|
||||||
|
|
||||||
Timestamp should be stored.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Long-Term Goal
|
|
||||||
|
|
||||||
Every meeting becomes searchable organizational knowledge.
|
|
||||||
|
|
||||||
No information should be lost after the meeting.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Engineering Principles
|
|
||||||
|
|
||||||
The project follows an architecture-first development approach.
|
|
||||||
|
|
||||||
Before implementing a feature:
|
|
||||||
|
|
||||||
- define the domain model
|
|
||||||
- define module boundaries
|
|
||||||
- document architectural decisions
|
|
||||||
|
|
||||||
Implementation is intentionally delayed until the architecture is considered sufficiently stable.
|
|
||||||
|
|||||||
@@ -1,126 +1,65 @@
|
|||||||
# Meeting Knowledge Assistant
|
# Meeting Assistant
|
||||||
|
|
||||||
## Vision
|
Meeting Assistant is the user-facing application for turning recorded meetings
|
||||||
|
into reviewed, distributable protocols. It uses Meeting Lab as its reusable
|
||||||
|
local processing backend.
|
||||||
|
|
||||||
The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge.
|
## Current MVP Direction
|
||||||
|
|
||||||
Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings.
|
The first product milestone is a desktop GUI that lets a user:
|
||||||
Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication.
|
|
||||||
|
|
||||||
The project follows an "AI-first" architecture:
|
- select an existing audio recording
|
||||||
|
- create and edit structured meeting metadata and participants
|
||||||
|
- optionally map anonymous `SPEAKER_XX` labels to known participants
|
||||||
|
- start the Meeting Lab processing pipeline
|
||||||
|
- follow stage-based progress
|
||||||
|
- review and edit the generated protocol
|
||||||
|
- export the reviewed result
|
||||||
|
|
||||||
Audio
|
The MVP does not include integrated recording. OBS Studio remains a recommended
|
||||||
→ Transcription
|
reference recorder, but imported audio is not tied to OBS-specific behavior.
|
||||||
→ Speaker Identification
|
|
||||||
→ AI Analysis
|
|
||||||
→ Structured Knowledge
|
|
||||||
→ Search
|
|
||||||
|
|
||||||
---
|
## Processing Pipeline
|
||||||
|
|
||||||
## Core Features
|
|
||||||
|
|
||||||
- Import of existing meeting recordings
|
|
||||||
- OBS Studio as the recommended reference recorder
|
|
||||||
- High-quality local transcription
|
|
||||||
- Speaker diarization
|
|
||||||
- Editable working transcripts
|
|
||||||
- AI-generated meeting minutes
|
|
||||||
- Executive summaries
|
|
||||||
- Action items
|
|
||||||
- Decision tracking
|
|
||||||
- Knowledge extraction
|
|
||||||
- Local file-based storage
|
|
||||||
- Export to Markdown, PDF and DOCX
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Recording Workflow
|
|
||||||
|
|
||||||
The MVP does not include its own audio recorder.
|
|
||||||
|
|
||||||
OBS Studio is the recommended reference recorder for online meetings. It can
|
|
||||||
|
|
||||||
capture system audio and microphone audio without joining the meeting as a bot.
|
|
||||||
|
|
||||||
Initial workflow:
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
|
Audio file
|
||||||
|
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
|
||||||
|
-> whisper.cpp transcription (large-v3-turbo)
|
||||||
|
-> optional pyannote.audio Community-1 diarization
|
||||||
|
-> direct full-transcript protocol generation
|
||||||
|
-> human review and export
|
||||||
|
```
|
||||||
|
|
||||||
Online Meeting
|
Meeting Assistant calls Meeting Lab's reusable Python API
|
||||||
|
(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run
|
||||||
|
the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a
|
||||||
|
required part of this workflow.
|
||||||
|
|
||||||
↓
|
GPU acceleration is supported where available, but CPU execution remains a
|
||||||
|
compatibility path. Diarization is optional and produces anonymous speaker
|
||||||
|
labels. A label identifies a participant only when the user explicitly
|
||||||
|
confirms the mapping; automatic speaker-name inference is not allowed.
|
||||||
|
|
||||||
OBS Studio
|
## Product Outputs
|
||||||
|
|
||||||
↓
|
The product direction includes:
|
||||||
|
|
||||||
WAV Recording
|
- a detailed, contextual protocol for review and continued work
|
||||||
|
- a shorter version suitable for participants and distribution
|
||||||
|
|
||||||
↓
|
The detailed direct-protocol flow is the practical MVP direction. The shorter
|
||||||
|
version may be implemented after the initial GUI. Human review remains part of
|
||||||
|
the workflow for all generated protocols.
|
||||||
|
|
||||||
Meeting Assistant Import
|
## Project Boundary
|
||||||
|
|
||||||
↓
|
Meeting Lab owns reusable audio preparation, transcription, diarization,
|
||||||
|
orchestration and protocol-generation logic. Meeting Assistant owns the GUI,
|
||||||
|
context and participant editing, explicit speaker mapping, progress display,
|
||||||
|
protocol editing and export.
|
||||||
|
|
||||||
Transcription
|
The code currently contains the initial Meeting Assistant project and domain
|
||||||
|
foundation. The GUI has not yet been implemented.
|
||||||
|
|
||||||
↓
|
See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md),
|
||||||
|
[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md).
|
||||||
Speaker Diarization
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Analysis and Export
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Design Goals
|
|
||||||
|
|
||||||
- High transcription quality
|
|
||||||
- Modular architecture
|
|
||||||
- Offline-first whenever practical
|
|
||||||
- Replaceable AI components
|
|
||||||
- Vendor independence
|
|
||||||
- Long-term maintainability
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Philosophy
|
|
||||||
|
|
||||||
The transcript is the primary asset.
|
|
||||||
|
|
||||||
Everything else (summaries, reports, action items, knowledge extraction)
|
|
||||||
can always be regenerated using better AI models in the future.
|
|
||||||
|
|
||||||
Therefore:
|
|
||||||
|
|
||||||
Audio
|
|
||||||
→ Transcript
|
|
||||||
is considered immutable.
|
|
||||||
|
|
||||||
AI output is considered reproducible.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Planned Architecture
|
|
||||||
|
|
||||||
Recorder
|
|
||||||
↓
|
|
||||||
Transcription
|
|
||||||
↓
|
|
||||||
Speaker Identification
|
|
||||||
↓
|
|
||||||
LLM Processing
|
|
||||||
↓
|
|
||||||
Knowledge Database
|
|
||||||
↓
|
|
||||||
Export
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Current Status
|
|
||||||
|
|
||||||
Project planning.
|
|
||||||
|
|
||||||
No implementation has started yet.
|
|
||||||
|
|||||||
+38
-100
@@ -1,112 +1,50 @@
|
|||||||
# Roadmap
|
# Roadmap
|
||||||
|
|
||||||
## Phase 1: Minimum Viable Processing Pipeline
|
## Next Milestone: Meeting Assistant MVP GUI
|
||||||
|
|
||||||
- Import existing WAV recordings
|
Build the user-facing application over the validated Meeting Lab Python API.
|
||||||
- Validate recording metadata
|
|
||||||
- Create meeting directory
|
|
||||||
- Create canonical transcript
|
|
||||||
- Save transcript as JSON
|
|
||||||
- Export transcript as Markdown
|
|
||||||
- Archive confirmed recordings as FLAC
|
|
||||||
|
|
||||||
---
|
- select an existing audio file
|
||||||
|
- create and edit Meeting Context through structured fields
|
||||||
|
- add and manage participants
|
||||||
|
- optionally map `SPEAKER_XX` labels to participants with explicit confirmation
|
||||||
|
- configure and start `run_mvp_meeting(...)`
|
||||||
|
- display stage-based status and real progress when available
|
||||||
|
- handle completed and failed runs clearly
|
||||||
|
- display and edit the detailed generated protocol
|
||||||
|
- export the reviewed protocol
|
||||||
|
|
||||||
## Phase 2
|
The milestone must keep diarization, GPU acceleration and fixed-duration audio
|
||||||
|
chunking optional. It must not launch the Meeting Lab CLI as a subprocess or
|
||||||
|
duplicate Meeting Lab processing logic.
|
||||||
|
|
||||||
Speaker diarization
|
## Following Milestone: Distribution Protocol
|
||||||
|
|
||||||
- Speaker detection
|
- derive or generate a shorter participant-facing version
|
||||||
- Speaker naming
|
- allow review and editing before distribution
|
||||||
- Timeline view
|
- export both detailed and short versions
|
||||||
|
- retain provenance linking each version to its transcript, configuration,
|
||||||
|
prompt and model where practical
|
||||||
|
|
||||||
---
|
## Later Product Work
|
||||||
|
|
||||||
## Phase 3
|
- transcript viewing and correction workflows
|
||||||
|
- recording and artifact lifecycle management
|
||||||
|
- search, tags, projects and meeting history
|
||||||
|
- richer Markdown, PDF and DOCX export
|
||||||
|
- optional SQLite indexing and full-text search
|
||||||
|
- integrated recording controls
|
||||||
|
- calendar and conferencing integrations
|
||||||
|
- semantic search and organizational knowledge features
|
||||||
|
|
||||||
Meeting intelligence
|
## Research and Optional Extensions
|
||||||
|
|
||||||
- Executive Summary
|
|
||||||
- Action Items
|
|
||||||
- Decisions
|
|
||||||
- Risks
|
|
||||||
- Open Questions
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 4
|
|
||||||
|
|
||||||
Knowledge system
|
|
||||||
|
|
||||||
- SQLite database
|
|
||||||
- Full text search
|
|
||||||
- Semantic search
|
|
||||||
- Tags
|
|
||||||
- Projects
|
|
||||||
- Participants
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 5
|
|
||||||
|
|
||||||
Desktop application
|
|
||||||
|
|
||||||
- Recording UI
|
|
||||||
- Transcript viewer
|
|
||||||
- Search
|
|
||||||
- Export
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Phase 6
|
|
||||||
|
|
||||||
Integrations
|
|
||||||
|
|
||||||
- Outlook
|
|
||||||
- Teams
|
|
||||||
- Google Calendar
|
|
||||||
- Local LLM
|
|
||||||
- OpenAI
|
|
||||||
- Ollama
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Future Phase: Integrated Recording
|
|
||||||
|
|
||||||
- Capture microphone audio
|
|
||||||
|
|
||||||
- Capture system audio
|
|
||||||
|
|
||||||
- Support separate audio channels
|
|
||||||
|
|
||||||
- Provide recording controls in the desktop application
|
|
||||||
|
|
||||||
- Replace or complement the external OBS workflow
|
|
||||||
|
|
||||||
- Live transcription
|
|
||||||
|
|
||||||
- Real-time summaries
|
|
||||||
|
|
||||||
- Company glossary
|
|
||||||
|
|
||||||
- Custom vocabulary
|
|
||||||
|
|
||||||
- Automatic project detection
|
|
||||||
|
|
||||||
- Meeting templates
|
|
||||||
|
|
||||||
- Voice identification
|
|
||||||
|
|
||||||
- Audio cleanup
|
|
||||||
|
|
||||||
|
- live transcription and real-time summaries
|
||||||
|
- company glossary and custom vocabulary
|
||||||
|
- voice identification with explicit consent and confirmation
|
||||||
|
- audio cleanup
|
||||||
- OCR for shared screens
|
- OCR for shared screens
|
||||||
|
- RAG and knowledge-graph integration
|
||||||
- Timeline with bookmarks
|
- multi-language meetings and translation
|
||||||
|
- dual-model or evidence pipelines only if later validation demonstrates a
|
||||||
- RAG knowledge integration
|
material reliability benefit
|
||||||
|
|
||||||
- Multi-language meetings
|
|
||||||
|
|
||||||
- Translation
|
|
||||||
|
|
||||||
- Automatic follow-up generation
|
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Status
|
## Status
|
||||||
|
|
||||||
Accepted
|
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||||
|
|
||||||
## Context
|
## Context
|
||||||
|
|
||||||
@@ -29,4 +29,9 @@ The transcript is treated as the canonical source. AI-generated summaries and kn
|
|||||||
- The system must preserve original recordings and original transcripts.
|
- The system must preserve original recordings and original transcripts.
|
||||||
- AI outputs must be reproducible where practical.
|
- AI outputs must be reproducible where practical.
|
||||||
- Modules should remain replaceable.
|
- Modules should remain replaceable.
|
||||||
- The project should avoid unnecessary vendor lock-in.
|
- The project should avoid unnecessary vendor lock-in.
|
||||||
|
|
||||||
|
ADR 0011 narrows the practical MVP to a user-facing application over the
|
||||||
|
reusable Meeting Lab backend. Direct protocol generation is the validated MVP
|
||||||
|
path; a separate knowledge-extraction stage remains a possible future
|
||||||
|
extension rather than an MVP requirement.
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Status
|
## Status
|
||||||
|
|
||||||
Accepted
|
Superseded by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||||
|
|
||||||
## Context
|
## Context
|
||||||
|
|
||||||
@@ -10,60 +10,16 @@ The project requires high-quality speech-to-text transcription for German and En
|
|||||||
|
|
||||||
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
The transcription component should support local execution where practical and should remain replaceable in the future. The project should not depend on a meeting bot or on a proprietary meeting platform.
|
||||||
|
|
||||||
Practical testing showed that long continuous recordings may suffer from error propagation when the decoder continuously conditions on previously generated text.
|
|
||||||
|
|
||||||
Independent transcription of smaller audio chunks produced significantly more stable results.
|
|
||||||
|
|
||||||
## Decision
|
## Decision
|
||||||
|
|
||||||
The project will use `faster-whisper` as the initial transcription engine.
|
The project will use `faster-whisper` as the initial transcription engine.
|
||||||
|
|
||||||
The preferred first model is `large-v3-turbo`, with `large-v3` as a quality-oriented fallback if required.
|
The preferred first model is `large-v3-turbo`, with `large-v3` as quality-oriented fallback if required.
|
||||||
|
|
||||||
The transcription pipeline shall process recordings as independent audio chunks rather than as a single continuous recording.
|
|
||||||
|
|
||||||
The default processing strategy is:
|
|
||||||
|
|
||||||
```text
|
|
||||||
Recording (WAV)
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
Split into 5-minute chunks
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
Independent Whisper transcription
|
|
||||||
(condition_on_previous_text = False)
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
JSON transcript per chunk
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
Merge chunk transcripts
|
|
||||||
│
|
|
||||||
▼
|
|
||||||
Canonical transcript
|
|
||||||
```
|
|
||||||
|
|
||||||
Chunk duration shall be configurable, with **5 minutes** as the default value.
|
|
||||||
|
|
||||||
Each chunk shall be processed independently.
|
|
||||||
|
|
||||||
The decoder shall not use text generated from previous chunks as context.
|
|
||||||
|
|
||||||
Each chunk shall produce its own intermediate JSON transcript before merging.
|
|
||||||
|
|
||||||
The merge step shall preserve timestamps and chunk ordering.
|
|
||||||
|
|
||||||
## Consequences
|
## Consequences
|
||||||
|
|
||||||
- Long recordings become more robust.
|
|
||||||
- Error propagation across chunk boundaries is prevented.
|
|
||||||
- Failed chunks can be retranscribed independently.
|
|
||||||
- Parallel processing of multiple chunks becomes possible.
|
|
||||||
- Intermediate JSON files simplify debugging and future reprocessing.
|
|
||||||
- Chunk size can be tuned later without changing the overall architecture.
|
|
||||||
- Transcription can run locally on suitable hardware.
|
- Transcription can run locally on suitable hardware.
|
||||||
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
- NVIDIA CUDA acceleration can be used later on an office AI PC.
|
||||||
- The transcription module must hide the concrete engine behind an internal interface.
|
- The transcription module must hide the concrete engine behind an internal interface.
|
||||||
- Model name, language setting, timestamp mode, engine version, chunk duration and relevant decoding parameters shall be stored with every transcript.
|
- Model name, language setting, timestamp and engine version should be stored with every transcript.
|
||||||
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
- The decision can be revisited if another engine provides clearly better quality, speed or deployment characteristics.
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Status
|
## Status
|
||||||
|
|
||||||
Accepted
|
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||||
|
|
||||||
## Context
|
## Context
|
||||||
|
|
||||||
@@ -23,3 +23,11 @@ Speaker diarization will be treated as a separate pipeline step after transcript
|
|||||||
- Human correction of speaker names should be supported later.
|
- Human correction of speaker names should be supported later.
|
||||||
- The diarization module must hide the concrete engine behind an internal interface.
|
- The diarization module must hide the concrete engine behind an internal interface.
|
||||||
- Model name, engine version and timestamp should be stored with every diarization result.
|
- Model name, engine version and timestamp should be stored with every diarization result.
|
||||||
|
|
||||||
|
## Amendment
|
||||||
|
|
||||||
|
ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio`
|
||||||
|
Community-1 as the current preferred backend. CPU execution remains supported,
|
||||||
|
with GPU acceleration used when available. Speaker labels remain anonymous
|
||||||
|
unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the
|
||||||
|
Meeting Context. Automatic speaker-name inference is not allowed.
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Status
|
## Status
|
||||||
|
|
||||||
Accepted
|
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||||
|
|
||||||
## Context
|
## Context
|
||||||
|
|
||||||
@@ -18,8 +18,8 @@ removing temporary source files without risking data loss.
|
|||||||
Recordings shall initially be captured in a processing-friendly lossless
|
Recordings shall initially be captured in a processing-friendly lossless
|
||||||
format, normally WAV.
|
format, normally WAV.
|
||||||
|
|
||||||
After transcription and diarization have completed, the recording may be
|
After transcription and, when enabled, diarization have completed, the
|
||||||
converted to an archive format.
|
recording may be converted to an archive format.
|
||||||
|
|
||||||
The preferred archive formats are:
|
The preferred archive formats are:
|
||||||
|
|
||||||
|
|||||||
@@ -54,12 +54,7 @@ meeting/
|
|||||||
├── transcript/
|
├── transcript/
|
||||||
│ ├── canonical.json
|
│ ├── canonical.json
|
||||||
│ ├── working_v1.json
|
│ ├── working_v1.json
|
||||||
│
|
│ └── working_v2.json
|
||||||
├── chunks/
|
|
||||||
│ ├── chunk_000.json
|
|
||||||
│ ├── chunk_001.json
|
|
||||||
│ ├── chunk_002.json
|
|
||||||
│ ├── ...
|
|
||||||
│
|
│
|
||||||
├── artifacts/
|
├── artifacts/
|
||||||
│ ├── summary.md
|
│ ├── summary.md
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
## Status
|
## Status
|
||||||
|
|
||||||
Accepted
|
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||||
|
|
||||||
## Context
|
## Context
|
||||||
|
|
||||||
@@ -13,8 +13,9 @@ OBS Studio already provides stable, configurable and cross-platform audio
|
|||||||
capture. Reimplementing audio capture during the initial project phase would
|
capture. Reimplementing audio capture during the initial project phase would
|
||||||
add substantial complexity without directly improving transcription quality.
|
add substantial complexity without directly improving transcription quality.
|
||||||
|
|
||||||
The primary value of the Meeting Assistant lies in transcription, speaker
|
The primary value of the Meeting Assistant lies in transcription, optional
|
||||||
diarization, analysis, structured knowledge extraction and export.
|
speaker diarization, protocol generation, review and export. ADR 0011 defines
|
||||||
|
the validated processing boundary and current MVP protocol path.
|
||||||
|
|
||||||
## Decision
|
## Decision
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,94 @@
|
|||||||
|
# ADR 0011: Use Meeting Lab as the MVP Processing Backend
|
||||||
|
|
||||||
|
## Status
|
||||||
|
|
||||||
|
Accepted
|
||||||
|
|
||||||
|
Supersedes ADR 0002 and amends ADRs 0001 and 0003.
|
||||||
|
|
||||||
|
## Context
|
||||||
|
|
||||||
|
Meeting Lab experiments have validated a practical local pipeline for turning
|
||||||
|
a recorded meeting into an editable protocol. Earlier Meeting Assistant plans
|
||||||
|
selected `faster-whisper`, treated diarization as required and described a
|
||||||
|
separate knowledge-extraction stage as part of the primary pipeline. Work on an
|
||||||
|
old feature branch also proposed mandatory fixed five-minute audio chunks and
|
||||||
|
chunk-oriented storage.
|
||||||
|
|
||||||
|
The validated backend now has a reusable Python API, optional diarization and a
|
||||||
|
direct full-transcript protocol path. The product boundary must keep backend
|
||||||
|
processing in Meeting Lab and user interaction in Meeting Assistant.
|
||||||
|
|
||||||
|
## Decision
|
||||||
|
|
||||||
|
Meeting Lab is the reusable processing backend and experimental engine.
|
||||||
|
Meeting Assistant is the user-facing application and calls Meeting Lab through
|
||||||
|
its Python API rather than launching its CLI as a subprocess.
|
||||||
|
|
||||||
|
The MVP processing flow is:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Source audio
|
||||||
|
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
|
||||||
|
-> whisper.cpp transcription (large-v3-turbo)
|
||||||
|
-> optional pyannote.audio Community-1 diarization
|
||||||
|
-> direct full-transcript protocol generation
|
||||||
|
-> human review and editing
|
||||||
|
```
|
||||||
|
|
||||||
|
The productive path does not require fixed five-minute audio chunks. Internal
|
||||||
|
streaming or segmentation remains an implementation detail of Meeting Lab and
|
||||||
|
must not define Meeting Assistant's domain or storage model.
|
||||||
|
|
||||||
|
Meeting Assistant integrates with these public Meeting Lab interfaces:
|
||||||
|
|
||||||
|
- `MvpMeetingConfig`
|
||||||
|
- `MvpRunResult`
|
||||||
|
- `run_mvp_meeting(...)`
|
||||||
|
- stage-based progress events: `preparing`, `transcription`, `diarization`,
|
||||||
|
`protocol_generation`, `completed` and `failed`
|
||||||
|
|
||||||
|
Progress percentages are displayed only when the backend reports real
|
||||||
|
measurable progress.
|
||||||
|
|
||||||
|
`MeetingContext` is the structured input containing meeting metadata and
|
||||||
|
participants. It may contain explicitly confirmed `SPEAKER_XX -> participant_id`
|
||||||
|
mappings. Such mappings are authoritative only when confirmed by the user.
|
||||||
|
The system must not infer speaker names automatically.
|
||||||
|
|
||||||
|
The preferred local engines are:
|
||||||
|
|
||||||
|
- `whisper.cpp` with `large-v3-turbo` for transcription
|
||||||
|
- `pyannote.audio` Community-1 for optional diarization
|
||||||
|
|
||||||
|
GPU acceleration is optional. Vulkan transcription and PyTorch/ROCm
|
||||||
|
diarization are validated acceleration paths on North's AMD RX 9070. CPU
|
||||||
|
execution remains the compatibility fallback.
|
||||||
|
|
||||||
|
The direct full-transcript protocol path is the practical MVP direction.
|
||||||
|
`qwen3.8:27B` has produced strong readability and contextual synthesis, while
|
||||||
|
`qwen3.6:35B-A3B` has sometimes behaved more conservatively. Dual-model and
|
||||||
|
diarization-assisted hard-fact experiments have not established sufficiently
|
||||||
|
reliable attribution gains to justify mandatory MVP complexity. Model choice
|
||||||
|
remains backend configuration, and human review is required.
|
||||||
|
|
||||||
|
Meeting Assistant owns the GUI, structured context editor, participant and
|
||||||
|
speaker-mapping controls, progress presentation, protocol editor and export.
|
||||||
|
Meeting Lab owns audio preparation, transcription, diarization, orchestration
|
||||||
|
and protocol-generation logic.
|
||||||
|
|
||||||
|
The product should ultimately provide a detailed contextual protocol and a
|
||||||
|
shorter participant/distribution version. The shorter version may follow the
|
||||||
|
first GUI milestone.
|
||||||
|
|
||||||
|
## Consequences
|
||||||
|
|
||||||
|
- ADR 0002's `faster-whisper` selection is no longer current.
|
||||||
|
- Fixed-duration chunking and chunk-centric storage are not MVP requirements.
|
||||||
|
- Diarization and GPU acceleration remain optional.
|
||||||
|
- The Meeting Lab CLI remains useful for command-line operation but is not the
|
||||||
|
application integration boundary.
|
||||||
|
- Meeting Context is created through the future GUI; users are not expected to
|
||||||
|
hand-write YAML.
|
||||||
|
- Backend experiments can evolve without moving processing logic into the GUI.
|
||||||
|
- Protocols remain AI-assisted outputs that require human review.
|
||||||
+125
-233
@@ -2,254 +2,146 @@
|
|||||||
|
|
||||||
## Purpose
|
## Purpose
|
||||||
|
|
||||||
The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline.
|
Meeting Assistant is the user-facing application for preparing meeting
|
||||||
|
context, running the Meeting Lab pipeline and reviewing its results. Meeting
|
||||||
|
Lab is the reusable processing backend and experimental engine.
|
||||||
|
|
||||||
---
|
## System Boundary
|
||||||
|
|
||||||
# Design Principles
|
|
||||||
|
|
||||||
- Architecture first
|
|
||||||
- Offline first where practical
|
|
||||||
- Immutable source data
|
|
||||||
- Replaceable AI components
|
|
||||||
- Small, focused modules
|
|
||||||
- Explicit interfaces
|
|
||||||
- Reproducible AI outputs
|
|
||||||
- External recording and internal processing are separate responsibilities.
|
|
||||||
- The processing pipeline must not depend on OBS-specific metadata or behavior.
|
|
||||||
- Imported recordings must be treated like recordings from any other supported source.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## High-Level Pipeline
|
|
||||||
|
|
||||||
External Recorder
|
|
||||||
↓
|
|
||||||
Audio Import
|
|
||||||
↓
|
|
||||||
Recording Validation
|
|
||||||
↓
|
|
||||||
Meeting Storage
|
|
||||||
↓
|
|
||||||
Transcription
|
|
||||||
↓
|
|
||||||
Speaker Diarization
|
|
||||||
↓
|
|
||||||
Working Transcript
|
|
||||||
↓
|
|
||||||
LLM Analysis
|
|
||||||
↓
|
|
||||||
Knowledge Extraction
|
|
||||||
↓
|
|
||||||
Export
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Core Pipeline
|
|
||||||
|
|
||||||
External Recording
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Audio Import and Validation
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Transcription
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Speaker Diarization
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Working Transcript
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
LLM Analysis
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Knowledge Extraction
|
|
||||||
|
|
||||||
↓
|
|
||||||
|
|
||||||
Storage and Export
|
|
||||||
|
|
||||||
# Domain Model
|
|
||||||
|
|
||||||
Meeting
|
|
||||||
│
|
|
||||||
├── Recording
|
|
||||||
├── Transcript
|
|
||||||
│ ├── Canonical
|
|
||||||
│ └── Working
|
|
||||||
├── Speakers
|
|
||||||
├── Participants
|
|
||||||
├── AI Artifacts
|
|
||||||
├── Knowledge Objects
|
|
||||||
├── Attachments
|
|
||||||
└── Exports
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
# Module Responsibilities
|
|
||||||
|
|
||||||
## Input and Ingest
|
|
||||||
|
|
||||||
Responsible for importing existing recordings into the application.
|
|
||||||
|
|
||||||
Initial reference source:
|
|
||||||
|
|
||||||
- OBS Studio
|
|
||||||
|
|
||||||
Initial supported formats:
|
|
||||||
|
|
||||||
- WAV
|
|
||||||
|
|
||||||
- FLAC
|
|
||||||
|
|
||||||
Responsibilities:
|
|
||||||
|
|
||||||
- validate the input file
|
|
||||||
|
|
||||||
- collect technical metadata
|
|
||||||
|
|
||||||
- calculate checksums
|
|
||||||
|
|
||||||
- copy or move the recording into the meeting directory
|
|
||||||
|
|
||||||
- create the initial Recording domain object
|
|
||||||
|
|
||||||
The ingest component must not perform transcription or modify the audio content.
|
|
||||||
|
|
||||||
## Recorder
|
|
||||||
|
|
||||||
The recorder module is reserved for a future integrated recording implementation.
|
|
||||||
|
|
||||||
It is not required for the MVP.
|
|
||||||
|
|
||||||
The initial application workflow uses externally created recordings, with OBS
|
|
||||||
Studio as the recommended reference recorder.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Transcription
|
|
||||||
|
|
||||||
Responsible only for speech-to-text conversion.
|
|
||||||
|
|
||||||
Input:
|
|
||||||
Recording
|
|
||||||
|
|
||||||
Output:
|
|
||||||
Canonical Transcript
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Diarization
|
|
||||||
|
|
||||||
Responsible only for speaker identification.
|
|
||||||
|
|
||||||
Input:
|
|
||||||
Recording + Canonical Transcript
|
|
||||||
|
|
||||||
Output:
|
|
||||||
Working Transcript
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## LLM
|
|
||||||
|
|
||||||
Responsible for semantic analysis.
|
|
||||||
|
|
||||||
Input:
|
|
||||||
Transcript
|
|
||||||
|
|
||||||
Output:
|
|
||||||
AI Artifacts
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Knowledge Extraction
|
|
||||||
|
|
||||||
Responsible for creating structured knowledge.
|
|
||||||
|
|
||||||
Input:
|
|
||||||
AI Artifacts
|
|
||||||
|
|
||||||
Output:
|
|
||||||
Knowledge Objects
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Export
|
|
||||||
|
|
||||||
Responsible for creating user-facing documents.
|
|
||||||
|
|
||||||
Input:
|
|
||||||
Knowledge Objects
|
|
||||||
|
|
||||||
Output:
|
|
||||||
Markdown
|
|
||||||
PDF
|
|
||||||
DOCX
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Artifact Lifecycle
|
|
||||||
|
|
||||||
```text
|
```text
|
||||||
Imported Recording
|
Meeting Assistant Meeting Lab
|
||||||
↓
|
----------------- -----------
|
||||||
Validated Source Recording
|
Audio selection ------> Audio preparation
|
||||||
↓
|
Meeting Context editor ------> Transcription
|
||||||
Canonical Transcript
|
Participant management ------> Optional diarization
|
||||||
↓
|
Explicit speaker mapping ------> Protocol generation
|
||||||
Working Transcript
|
Progress presentation <------ Progress events
|
||||||
↓
|
Protocol editor and export <------ Run result and artifacts
|
||||||
AI Artifacts
|
```
|
||||||
↓
|
|
||||||
Knowledge Objects
|
|
||||||
↓
|
|
||||||
Exports
|
|
||||||
|
|
||||||
---
|
Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab
|
||||||
|
CLI is a thin adapter over that API and must not be launched as an application
|
||||||
|
subprocess.
|
||||||
|
|
||||||
## Initial Recording Strategy
|
## MVP Processing Flow
|
||||||
|
|
||||||
The MVP does not implement platform-specific audio capture.
|
```text
|
||||||
|
Source audio
|
||||||
|
-> FFmpeg normalization/preparation
|
||||||
|
-> mono, 16 kHz PCM WAV
|
||||||
|
-> whisper.cpp transcription with large-v3-turbo
|
||||||
|
-> optional pyannote.audio Community-1 diarization
|
||||||
|
-> direct full-transcript protocol generation
|
||||||
|
-> human review and editing
|
||||||
|
-> export
|
||||||
|
```
|
||||||
|
|
||||||
OBS Studio is the recommended reference recorder for online meetings.
|
The MVP does not require fixed five-minute audio chunks, semantic chunking, a
|
||||||
|
separate evidence-extraction pipeline or multiple LLMs. Any internal
|
||||||
|
segmentation remains a Meeting Lab implementation detail.
|
||||||
|
|
||||||
The application initially processes existing WAV or FLAC recordings. Integrated
|
## Backend Integration
|
||||||
|
|
||||||
recording remains a future extension and must not be required by transcription,
|
The reusable Meeting Lab interface consists of:
|
||||||
|
|
||||||
diarization or analysis modules.
|
- `MvpMeetingConfig`, the run configuration
|
||||||
|
- `run_mvp_meeting(...)`, the orchestration entry point
|
||||||
|
- `MvpRunResult`, the completed run result
|
||||||
|
- stage-based progress events
|
||||||
|
|
||||||
---
|
The known stages are:
|
||||||
|
|
||||||
# Storage Strategy
|
```text
|
||||||
|
preparing
|
||||||
|
transcription
|
||||||
|
diarization
|
||||||
|
protocol_generation
|
||||||
|
completed
|
||||||
|
failed
|
||||||
|
```
|
||||||
|
|
||||||
The domain model is independent of the storage backend.
|
The GUI shows the current stage. It shows a percentage only when the event
|
||||||
|
contains real measurable progress; stage changes must not be presented as
|
||||||
|
invented percentages.
|
||||||
|
|
||||||
Possible implementations:
|
## Meeting Context
|
||||||
|
|
||||||
- File System
|
`MeetingContext` is structured domain input rather than an informal prompt or a
|
||||||
- SQLite
|
file users must author manually. It contains meeting metadata and participants
|
||||||
- PostgreSQL
|
and can include optional mappings:
|
||||||
- Cloud Storage
|
|
||||||
|
|
||||||
---
|
```text
|
||||||
|
SPEAKER_XX -> participant_id
|
||||||
|
```
|
||||||
|
|
||||||
# Future Extensions
|
Mappings are authoritative only after explicit user confirmation. Diarization
|
||||||
|
labels otherwise remain anonymous, and the application must not infer speaker
|
||||||
|
names automatically. The Meeting Assistant GUI owns creation and editing of
|
||||||
|
this context and its mappings.
|
||||||
|
|
||||||
- Live transcription
|
## Processing Components
|
||||||
- Video processing
|
|
||||||
- OCR
|
### Audio Preparation
|
||||||
- Semantic search
|
|
||||||
- Knowledge graph
|
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
|
||||||
- Company glossary
|
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
|
||||||
- Multi-language meetings
|
separate source artifact.
|
||||||
- Local LLM support
|
|
||||||
|
### Transcription
|
||||||
|
|
||||||
|
`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a
|
||||||
|
validated acceleration path on North's AMD RX 9070, while CPU execution remains
|
||||||
|
a compatibility fallback. Engine and model details should be retained with the
|
||||||
|
result for traceability.
|
||||||
|
|
||||||
|
### Diarization
|
||||||
|
|
||||||
|
`pyannote.audio` Community-1 is preferred when diarization is enabled. It can
|
||||||
|
run on CPU and may use GPU acceleration such as PyTorch/ROCm where available.
|
||||||
|
Diarization is optional and its anonymous labels do not modify the canonical
|
||||||
|
transcript or assert participant identity.
|
||||||
|
|
||||||
|
### Protocol Generation
|
||||||
|
|
||||||
|
The direct full-transcript path is the practical MVP. Current model experiments
|
||||||
|
favor `qwen3.8:27B` for readability and contextual synthesis, while
|
||||||
|
`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a
|
||||||
|
specific dual-model arrangement nor an evidence pipeline is an application
|
||||||
|
requirement. Generated protocols require human review.
|
||||||
|
|
||||||
|
## Application Responsibilities
|
||||||
|
|
||||||
|
The first GUI milestone provides:
|
||||||
|
|
||||||
|
- audio-file selection
|
||||||
|
- structured Meeting Context and participant editing
|
||||||
|
- optional explicit speaker mapping
|
||||||
|
- pipeline start and configuration
|
||||||
|
- progress and failure presentation
|
||||||
|
- detailed protocol display and editing
|
||||||
|
- export of the reviewed result
|
||||||
|
|
||||||
|
The product direction also includes a shorter participant/distribution
|
||||||
|
protocol. It may be delivered after the first GUI milestone.
|
||||||
|
|
||||||
|
## Artifact and Storage Principles
|
||||||
|
|
||||||
|
The existing Meeting domain remains the application-level container for source
|
||||||
|
recordings, canonical and working transcripts, context, generated artifacts and
|
||||||
|
exports. Storage remains independent of processing segmentation. In particular,
|
||||||
|
chunks are not primary Meeting Assistant domain objects or the required unit of
|
||||||
|
MVP persistence.
|
||||||
|
|
||||||
|
Source audio and the canonical transcription output are preserved. Corrections
|
||||||
|
and reviewed protocols are derived versions. Generated artifacts should retain
|
||||||
|
their input version, backend/model configuration, prompt version and timestamp
|
||||||
|
where practical.
|
||||||
|
|
||||||
|
## Future Extensions
|
||||||
|
|
||||||
|
- shorter distribution protocols
|
||||||
|
- integrated recording
|
||||||
|
- searchable meeting history and optional database indexes
|
||||||
|
- live transcription
|
||||||
|
- OCR and video processing
|
||||||
|
- semantic search and organizational knowledge features
|
||||||
|
|||||||
+17
-1
@@ -2,7 +2,23 @@ Transcript
|
|||||||
Original speech-to-text output.
|
Original speech-to-text output.
|
||||||
|
|
||||||
Diarization
|
Diarization
|
||||||
Assignment of transcript segments to speakers.
|
Assignment of transcript segments to anonymous speaker labels. It does not by
|
||||||
|
itself identify participants.
|
||||||
|
|
||||||
|
Meeting Context
|
||||||
|
Structured meeting metadata, participants and optional explicitly confirmed
|
||||||
|
speaker-to-participant mappings supplied to the processing pipeline.
|
||||||
|
|
||||||
|
Speaker Mapping
|
||||||
|
An explicitly confirmed association from a diarization label such as
|
||||||
|
`SPEAKER_00` to a participant ID. It must not be inferred automatically.
|
||||||
|
|
||||||
|
Detailed Protocol
|
||||||
|
Contextual protocol generated from the full transcript for human review and
|
||||||
|
continued work.
|
||||||
|
|
||||||
|
Distribution Protocol
|
||||||
|
Shorter, participant-facing version of a reviewed meeting protocol.
|
||||||
|
|
||||||
Knowledge Object
|
Knowledge Object
|
||||||
Structured information extracted from meetings.
|
Structured information extracted from meetings.
|
||||||
|
|||||||
Reference in New Issue
Block a user