Reconcile Meeting Assistant architecture with Meeting Lab
This commit is contained in:
@@ -2,7 +2,7 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||
|
||||
## Context
|
||||
|
||||
@@ -29,4 +29,9 @@ The transcript is treated as the canonical source. AI-generated summaries and kn
|
||||
- The system must preserve original recordings and original transcripts.
|
||||
- AI outputs must be reproducible where practical.
|
||||
- Modules should remain replaceable.
|
||||
- The project should avoid unnecessary vendor lock-in.
|
||||
- The project should avoid unnecessary vendor lock-in.
|
||||
|
||||
ADR 0011 narrows the practical MVP to a user-facing application over the
|
||||
reusable Meeting Lab backend. Direct protocol generation is the validated MVP
|
||||
path; a separate knowledge-extraction stage remains a possible future
|
||||
extension rather than an MVP requirement.
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
Superseded by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||
|
||||
## Context
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||
|
||||
## Context
|
||||
|
||||
@@ -23,3 +23,11 @@ Speaker diarization will be treated as a separate pipeline step after transcript
|
||||
- Human correction of speaker names should be supported later.
|
||||
- The diarization module must hide the concrete engine behind an internal interface.
|
||||
- Model name, engine version and timestamp should be stored with every diarization result.
|
||||
|
||||
## Amendment
|
||||
|
||||
ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio`
|
||||
Community-1 as the current preferred backend. CPU execution remains supported,
|
||||
with GPU acceleration used when available. Speaker labels remain anonymous
|
||||
unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the
|
||||
Meeting Context. Automatic speaker-name inference is not allowed.
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||
|
||||
## Context
|
||||
|
||||
@@ -18,8 +18,8 @@ removing temporary source files without risking data loss.
|
||||
Recordings shall initially be captured in a processing-friendly lossless
|
||||
format, normally WAV.
|
||||
|
||||
After transcription and diarization have completed, the recording may be
|
||||
converted to an archive format.
|
||||
After transcription and, when enabled, diarization have completed, the
|
||||
recording may be converted to an archive format.
|
||||
|
||||
The preferred archive formats are:
|
||||
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md)
|
||||
|
||||
## Context
|
||||
|
||||
@@ -13,8 +13,9 @@ OBS Studio already provides stable, configurable and cross-platform audio
|
||||
capture. Reimplementing audio capture during the initial project phase would
|
||||
add substantial complexity without directly improving transcription quality.
|
||||
|
||||
The primary value of the Meeting Assistant lies in transcription, speaker
|
||||
diarization, analysis, structured knowledge extraction and export.
|
||||
The primary value of the Meeting Assistant lies in transcription, optional
|
||||
speaker diarization, protocol generation, review and export. ADR 0011 defines
|
||||
the validated processing boundary and current MVP protocol path.
|
||||
|
||||
## Decision
|
||||
|
||||
|
||||
@@ -0,0 +1,94 @@
|
||||
# ADR 0011: Use Meeting Lab as the MVP Processing Backend
|
||||
|
||||
## Status
|
||||
|
||||
Accepted
|
||||
|
||||
Supersedes ADR 0002 and amends ADRs 0001 and 0003.
|
||||
|
||||
## Context
|
||||
|
||||
Meeting Lab experiments have validated a practical local pipeline for turning
|
||||
a recorded meeting into an editable protocol. Earlier Meeting Assistant plans
|
||||
selected `faster-whisper`, treated diarization as required and described a
|
||||
separate knowledge-extraction stage as part of the primary pipeline. Work on an
|
||||
old feature branch also proposed mandatory fixed five-minute audio chunks and
|
||||
chunk-oriented storage.
|
||||
|
||||
The validated backend now has a reusable Python API, optional diarization and a
|
||||
direct full-transcript protocol path. The product boundary must keep backend
|
||||
processing in Meeting Lab and user interaction in Meeting Assistant.
|
||||
|
||||
## Decision
|
||||
|
||||
Meeting Lab is the reusable processing backend and experimental engine.
|
||||
Meeting Assistant is the user-facing application and calls Meeting Lab through
|
||||
its Python API rather than launching its CLI as a subprocess.
|
||||
|
||||
The MVP processing flow is:
|
||||
|
||||
```text
|
||||
Source audio
|
||||
-> FFmpeg preparation (mono, 16 kHz PCM WAV)
|
||||
-> whisper.cpp transcription (large-v3-turbo)
|
||||
-> optional pyannote.audio Community-1 diarization
|
||||
-> direct full-transcript protocol generation
|
||||
-> human review and editing
|
||||
```
|
||||
|
||||
The productive path does not require fixed five-minute audio chunks. Internal
|
||||
streaming or segmentation remains an implementation detail of Meeting Lab and
|
||||
must not define Meeting Assistant's domain or storage model.
|
||||
|
||||
Meeting Assistant integrates with these public Meeting Lab interfaces:
|
||||
|
||||
- `MvpMeetingConfig`
|
||||
- `MvpRunResult`
|
||||
- `run_mvp_meeting(...)`
|
||||
- stage-based progress events: `preparing`, `transcription`, `diarization`,
|
||||
`protocol_generation`, `completed` and `failed`
|
||||
|
||||
Progress percentages are displayed only when the backend reports real
|
||||
measurable progress.
|
||||
|
||||
`MeetingContext` is the structured input containing meeting metadata and
|
||||
participants. It may contain explicitly confirmed `SPEAKER_XX -> participant_id`
|
||||
mappings. Such mappings are authoritative only when confirmed by the user.
|
||||
The system must not infer speaker names automatically.
|
||||
|
||||
The preferred local engines are:
|
||||
|
||||
- `whisper.cpp` with `large-v3-turbo` for transcription
|
||||
- `pyannote.audio` Community-1 for optional diarization
|
||||
|
||||
GPU acceleration is optional. Vulkan transcription and PyTorch/ROCm
|
||||
diarization are validated acceleration paths on North's AMD RX 9070. CPU
|
||||
execution remains the compatibility fallback.
|
||||
|
||||
The direct full-transcript protocol path is the practical MVP direction.
|
||||
`qwen3.8:27B` has produced strong readability and contextual synthesis, while
|
||||
`qwen3.6:35B-A3B` has sometimes behaved more conservatively. Dual-model and
|
||||
diarization-assisted hard-fact experiments have not established sufficiently
|
||||
reliable attribution gains to justify mandatory MVP complexity. Model choice
|
||||
remains backend configuration, and human review is required.
|
||||
|
||||
Meeting Assistant owns the GUI, structured context editor, participant and
|
||||
speaker-mapping controls, progress presentation, protocol editor and export.
|
||||
Meeting Lab owns audio preparation, transcription, diarization, orchestration
|
||||
and protocol-generation logic.
|
||||
|
||||
The product should ultimately provide a detailed contextual protocol and a
|
||||
shorter participant/distribution version. The shorter version may follow the
|
||||
first GUI milestone.
|
||||
|
||||
## Consequences
|
||||
|
||||
- ADR 0002's `faster-whisper` selection is no longer current.
|
||||
- Fixed-duration chunking and chunk-centric storage are not MVP requirements.
|
||||
- Diarization and GPU acceleration remain optional.
|
||||
- The Meeting Lab CLI remains useful for command-line operation but is not the
|
||||
application integration boundary.
|
||||
- Meeting Context is created through the future GUI; users are not expected to
|
||||
hand-write YAML.
|
||||
- Backend experiments can evolve without moving processing logic into the GUI.
|
||||
- Protocols remain AI-assisted outputs that require human review.
|
||||
Reference in New Issue
Block a user