diff --git a/CHANGELOG.md b/CHANGELOG.md index 981df42..953f176 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,5 +1,19 @@ # Changelog +## [Unreleased] + +### Changed + +- Reconciled the documented MVP with the validated Meeting Lab backend. +- Selected `whisper.cpp` for transcription and made diarization optional. +- Defined the Python API, progress-event and Meeting Context integration + boundaries. +- Set the user-facing Meeting Assistant GUI as the next milestone. + +### Added + +- ADR 0011 for the Meeting Lab MVP backend architecture. + ## [0.1.0] - 2026-07-14 ### Added diff --git a/PROJECT_KNOWLEDGE.md b/PROJECT_KNOWLEDGE.md index dddb205..36738ca 100644 --- a/PROJECT_KNOWLEDGE.md +++ b/PROJECT_KNOWLEDGE.md @@ -2,162 +2,112 @@ ## Vision -The project creates a complete pipeline that transforms spoken meetings into structured organizational knowledge. +Meeting Assistant turns recorded meetings into reviewed, user-facing protocols +and, over time, searchable organizational knowledge. It is the application +layer over the reusable Meeting Lab processing backend. -The transcript is the canonical source. +The source audio and canonical transcript are preserved. Generated protocols +are derived, reproducible artifacts and require human review. -AI-generated summaries are reproducible artefacts. +## System Boundary ---- +### Meeting Lab + +Meeting Lab owns reusable and experimental processing: + +- FFmpeg audio preparation +- `whisper.cpp` transcription +- optional `pyannote.audio` diarization +- protocol generation +- pipeline orchestration and progress events + +Its reusable integration surface is the Python API: + +- `MvpMeetingConfig` +- `MvpRunResult` +- `run_mvp_meeting(...)` + +The CLI is a thin adapter, not the Meeting Assistant integration boundary. + +### Meeting Assistant + +Meeting Assistant owns user interaction and product workflow: + +- audio-file selection +- structured Meeting Context editing +- participant management +- optional explicit speaker mapping +- pipeline launch and progress display +- protocol review and editing +- export and presentation of protocol versions + +Processing logic must not be duplicated in the application. + +## Validated MVP Pipeline + +```text +Imported audio + -> prepare as mono, 16 kHz PCM WAV with FFmpeg + -> transcribe with whisper.cpp and large-v3-turbo + -> optionally diarize with pyannote.audio Community-1 + -> generate a protocol directly from the full transcript + -> review and edit by a human +``` + +Mandatory fixed five-minute chunking is not part of the productive MVP path. +Chunking or segmentation used internally by an engine does not shape the +Meeting Assistant storage or domain model. + +On the North reference machine, `whisper.cpp` with Vulkan on an AMD RX 9070 +transcribed approximately 93.5 minutes in about 100 seconds. Pyannote through +PyTorch/ROCm processed the same duration in approximately 98.6 seconds. These +are validation observations, not performance guarantees or hardware +requirements. CPU execution remains supported and may be substantially slower. + +## Meeting Context and Speakers + +`MeetingContext` is a structured domain object containing meeting metadata and +participants. It can also contain optional explicit mappings from +`SPEAKER_XX` labels to `participant_id` values. + +A speaker mapping is authoritative only when a user explicitly confirms it. +Speakers otherwise remain anonymous. Automatic speaker-name inference is not +allowed. The GUI must create and edit Meeting Context; hand-written YAML is not +a product requirement. + +## Progress Contract + +Meeting Lab emits stage-based progress events for: + +- `preparing` +- `transcription` +- `diarization` +- `protocol_generation` +- `completed` +- `failed` + +The GUI consumes this event interface. Percentages are shown only when real +measurable progress is available. + +## Protocol Policy + +Direct full-transcript protocol generation is the practical MVP direction. +`qwen3.8:27B` has shown strong readability and contextual synthesis; +`qwen3.6:35B-A3B` has shown somewhat more conservative behavior in some areas. +Experimental dual-model and diarization-assisted hard-fact extraction has not +demonstrated reliably better strict attribution accuracy and is not mandatory. + +The product should ultimately offer both a detailed contextual protocol and a +shorter participant/distribution version. The short version remains follow-up +work if it is not available for the first GUI milestone. ## Design Principles -Offline-first where practical. - -Replaceable AI components. - -Vendor independence. - -Modular architecture. - -Small focused modules. - -Simple interfaces. - ---- - -## Core Pipeline - -Recorder - -↓ - -Transcription - -↓ - -Speaker Identification - -↓ - -LLM Analysis - -↓ - -Knowledge Extraction - -↓ - -Storage - -↓ - -Export - ---- - -## AI Components - -Recorder - -Responsible only for recording. - -No AI. - ---- - -Transcription - -Responsible only for speech-to-text. - -No summarization. - ---- - -Speaker Identification - -Responsible only for identifying speakers. - -Must not modify transcript text. - ---- - -LLM Analysis - -Responsible for: - -- Summary - -- Decisions - -- Action Items - -- Risks - -- Questions - -- Knowledge extraction - ---- - -Storage - -Stores: - -Audio - -Transcript - -Metadata - -AI Results - -Knowledge Objects - ---- - -## Architectural Rule - -Every module has exactly one responsibility. - ---- - -## Transcript Policy - -Never modify the original transcript. - -Corrections must create a new derived version. - ---- - -## AI Output Policy - -Every AI output should be reproducible. - -Prompt version should be stored. - -Model should be stored. - -Timestamp should be stored. - ---- - -## Long-Term Goal - -Every meeting becomes searchable organizational knowledge. - -No information should be lost after the meeting. - ---- - -## Engineering Principles - -The project follows an architecture-first development approach. - -Before implementing a feature: - -- define the domain model -- define module boundaries -- document architectural decisions - -Implementation is intentionally delayed until the architecture is considered sufficiently stable. \ No newline at end of file +- offline-first where practical +- explicit application/backend boundaries +- replaceable AI components +- CPU-compatible operation with optional acceleration +- immutable source artifacts and traceable derived artifacts +- no automatic identity claims +- human review of generated protocols +- small modules and simple interfaces diff --git a/README.md b/README.md index cdc25ad..04c09ba 100644 --- a/README.md +++ b/README.md @@ -1,126 +1,65 @@ -# Meeting Knowledge Assistant +# Meeting Assistant -## Vision +Meeting Assistant is the user-facing application for turning recorded meetings +into reviewed, distributable protocols. It uses Meeting Lab as its reusable +local processing backend. -The Meeting Knowledge Assistant (MKA) is an AI-assisted desktop application that records meetings, creates high-quality transcripts and transforms them into structured knowledge. +## Current MVP Direction -Unlike traditional meeting assistants, MKA does not aim to replace human interaction during meetings. -Its goal is to create an accurate, searchable and long-term knowledge base from spoken communication. +The first product milestone is a desktop GUI that lets a user: -The project follows an "AI-first" architecture: +- select an existing audio recording +- create and edit structured meeting metadata and participants +- optionally map anonymous `SPEAKER_XX` labels to known participants +- start the Meeting Lab processing pipeline +- follow stage-based progress +- review and edit the generated protocol +- export the reviewed result -Audio -→ Transcription -→ Speaker Identification -→ AI Analysis -→ Structured Knowledge -→ Search +The MVP does not include integrated recording. OBS Studio remains a recommended +reference recorder, but imported audio is not tied to OBS-specific behavior. ---- - -## Core Features - -- Import of existing meeting recordings -- OBS Studio as the recommended reference recorder -- High-quality local transcription -- Speaker diarization -- Editable working transcripts -- AI-generated meeting minutes -- Executive summaries -- Action items -- Decision tracking -- Knowledge extraction -- Local file-based storage -- Export to Markdown, PDF and DOCX - ---- - -## Recording Workflow - -The MVP does not include its own audio recorder. - -OBS Studio is the recommended reference recorder for online meetings. It can - -capture system audio and microphone audio without joining the meeting as a bot. - -Initial workflow: +## Processing Pipeline ```text +Audio file + -> FFmpeg preparation (mono, 16 kHz PCM WAV) + -> whisper.cpp transcription (large-v3-turbo) + -> optional pyannote.audio Community-1 diarization + -> direct full-transcript protocol generation + -> human review and export +``` -Online Meeting +Meeting Assistant calls Meeting Lab's reusable Python API +(`MvpMeetingConfig`, `MvpRunResult` and `run_mvp_meeting(...)`). It does not run +the Meeting Lab CLI as a subprocess. Fixed five-minute audio chunking is not a +required part of this workflow. - ↓ +GPU acceleration is supported where available, but CPU execution remains a +compatibility path. Diarization is optional and produces anonymous speaker +labels. A label identifies a participant only when the user explicitly +confirms the mapping; automatic speaker-name inference is not allowed. -OBS Studio +## Product Outputs - ↓ +The product direction includes: -WAV Recording +- a detailed, contextual protocol for review and continued work +- a shorter version suitable for participants and distribution - ↓ +The detailed direct-protocol flow is the practical MVP direction. The shorter +version may be implemented after the initial GUI. Human review remains part of +the workflow for all generated protocols. -Meeting Assistant Import +## Project Boundary - ↓ +Meeting Lab owns reusable audio preparation, transcription, diarization, +orchestration and protocol-generation logic. Meeting Assistant owns the GUI, +context and participant editing, explicit speaker mapping, progress display, +protocol editing and export. -Transcription +The code currently contains the initial Meeting Assistant project and domain +foundation. The GUI has not yet been implemented. - ↓ - -Speaker Diarization - - ↓ - -Analysis and Export - ---- - -## Design Goals - -- High transcription quality -- Modular architecture -- Offline-first whenever practical -- Replaceable AI components -- Vendor independence -- Long-term maintainability - ---- - -## Philosophy - -The transcript is the primary asset. - -Everything else (summaries, reports, action items, knowledge extraction) -can always be regenerated using better AI models in the future. - -Therefore: - -Audio -→ Transcript -is considered immutable. - -AI output is considered reproducible. - ---- - -## Planned Architecture - -Recorder -↓ -Transcription -↓ -Speaker Identification -↓ -LLM Processing -↓ -Knowledge Database -↓ -Export - ---- - -## Current Status - -Project planning. - -No implementation has started yet. +See [Architecture](docs/architecture.md), [Project Knowledge](PROJECT_KNOWLEDGE.md), +[Roadmap](ROADMAP.md) and [ADR 0011](docs/adr/0011-use-meeting-lab-mvp-backend.md). diff --git a/ROADMAP.md b/ROADMAP.md index 092f243..da31cad 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -1,112 +1,50 @@ # Roadmap -## Phase 1: Minimum Viable Processing Pipeline +## Next Milestone: Meeting Assistant MVP GUI -- Import existing WAV recordings -- Validate recording metadata -- Create meeting directory -- Create canonical transcript -- Save transcript as JSON -- Export transcript as Markdown -- Archive confirmed recordings as FLAC +Build the user-facing application over the validated Meeting Lab Python API. ---- +- select an existing audio file +- create and edit Meeting Context through structured fields +- add and manage participants +- optionally map `SPEAKER_XX` labels to participants with explicit confirmation +- configure and start `run_mvp_meeting(...)` +- display stage-based status and real progress when available +- handle completed and failed runs clearly +- display and edit the detailed generated protocol +- export the reviewed protocol -## Phase 2 +The milestone must keep diarization, GPU acceleration and fixed-duration audio +chunking optional. It must not launch the Meeting Lab CLI as a subprocess or +duplicate Meeting Lab processing logic. -Speaker diarization +## Following Milestone: Distribution Protocol -- Speaker detection -- Speaker naming -- Timeline view +- derive or generate a shorter participant-facing version +- allow review and editing before distribution +- export both detailed and short versions +- retain provenance linking each version to its transcript, configuration, + prompt and model where practical ---- +## Later Product Work -## Phase 3 +- transcript viewing and correction workflows +- recording and artifact lifecycle management +- search, tags, projects and meeting history +- richer Markdown, PDF and DOCX export +- optional SQLite indexing and full-text search +- integrated recording controls +- calendar and conferencing integrations +- semantic search and organizational knowledge features -Meeting intelligence - -- Executive Summary -- Action Items -- Decisions -- Risks -- Open Questions - ---- - -## Phase 4 - -Knowledge system - -- SQLite database -- Full text search -- Semantic search -- Tags -- Projects -- Participants - ---- - -## Phase 5 - -Desktop application - -- Recording UI -- Transcript viewer -- Search -- Export - ---- - -## Phase 6 - -Integrations - -- Outlook -- Teams -- Google Calendar -- Local LLM -- OpenAI -- Ollama - ---- - -## Future Phase: Integrated Recording - -- Capture microphone audio - -- Capture system audio - -- Support separate audio channels - -- Provide recording controls in the desktop application - -- Replace or complement the external OBS workflow - -- Live transcription - -- Real-time summaries - -- Company glossary - -- Custom vocabulary - -- Automatic project detection - -- Meeting templates - -- Voice identification - -- Audio cleanup +## Research and Optional Extensions +- live transcription and real-time summaries +- company glossary and custom vocabulary +- voice identification with explicit consent and confirmation +- audio cleanup - OCR for shared screens - -- Timeline with bookmarks - -- RAG knowledge integration - -- Multi-language meetings - -- Translation - -- Automatic follow-up generation +- RAG and knowledge-graph integration +- multi-language meetings and translation +- dual-model or evidence pipelines only if later validation demonstrates a + material reliability benefit diff --git a/docs/adr/0001-project-vision.md b/docs/adr/0001-project-vision.md index cdc6d81..bcd63ef 100644 --- a/docs/adr/0001-project-vision.md +++ b/docs/adr/0001-project-vision.md @@ -2,7 +2,7 @@ ## Status -Accepted +Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md) ## Context @@ -29,4 +29,9 @@ The transcript is treated as the canonical source. AI-generated summaries and kn - The system must preserve original recordings and original transcripts. - AI outputs must be reproducible where practical. - Modules should remain replaceable. -- The project should avoid unnecessary vendor lock-in. \ No newline at end of file +- The project should avoid unnecessary vendor lock-in. + +ADR 0011 narrows the practical MVP to a user-facing application over the +reusable Meeting Lab backend. Direct protocol generation is the validated MVP +path; a separate knowledge-extraction stage remains a possible future +extension rather than an MVP requirement. diff --git a/docs/adr/0002-use-faster-whisper.md b/docs/adr/0002-use-faster-whisper.md index 30dab4d..f199da9 100644 --- a/docs/adr/0002-use-faster-whisper.md +++ b/docs/adr/0002-use-faster-whisper.md @@ -2,7 +2,7 @@ ## Status -Accepted +Superseded by [ADR 0011](0011-use-meeting-lab-mvp-backend.md) ## Context diff --git a/docs/adr/0003-use-pyannote.md b/docs/adr/0003-use-pyannote.md index f72e62d..02a76a0 100644 --- a/docs/adr/0003-use-pyannote.md +++ b/docs/adr/0003-use-pyannote.md @@ -2,7 +2,7 @@ ## Status -Accepted +Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md) ## Context @@ -23,3 +23,11 @@ Speaker diarization will be treated as a separate pipeline step after transcript - Human correction of speaker names should be supported later. - The diarization module must hide the concrete engine behind an internal interface. - Model name, engine version and timestamp should be stored with every diarization result. + +## Amendment + +ADR 0011 makes diarization optional for the MVP and selects `pyannote.audio` +Community-1 as the current preferred backend. CPU execution remains supported, +with GPU acceleration used when available. Speaker labels remain anonymous +unless a user explicitly confirms a `SPEAKER_XX` to participant mapping in the +Meeting Context. Automatic speaker-name inference is not allowed. diff --git a/docs/adr/0007-recording-archive-policy.md b/docs/adr/0007-recording-archive-policy.md index 5cd9c92..daee264 100644 --- a/docs/adr/0007-recording-archive-policy.md +++ b/docs/adr/0007-recording-archive-policy.md @@ -2,7 +2,7 @@ ## Status -Accepted +Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md) ## Context @@ -18,8 +18,8 @@ removing temporary source files without risking data loss. Recordings shall initially be captured in a processing-friendly lossless format, normally WAV. -After transcription and diarization have completed, the recording may be -converted to an archive format. +After transcription and, when enabled, diarization have completed, the +recording may be converted to an archive format. The preferred archive formats are: diff --git a/docs/adr/0009-use-external-recording-for-mvp.md b/docs/adr/0009-use-external-recording-for-mvp.md index b6aa978..0bdd8b9 100644 --- a/docs/adr/0009-use-external-recording-for-mvp.md +++ b/docs/adr/0009-use-external-recording-for-mvp.md @@ -2,7 +2,7 @@ ## Status -Accepted +Accepted, amended by [ADR 0011](0011-use-meeting-lab-mvp-backend.md) ## Context @@ -13,8 +13,9 @@ OBS Studio already provides stable, configurable and cross-platform audio capture. Reimplementing audio capture during the initial project phase would add substantial complexity without directly improving transcription quality. -The primary value of the Meeting Assistant lies in transcription, speaker -diarization, analysis, structured knowledge extraction and export. +The primary value of the Meeting Assistant lies in transcription, optional +speaker diarization, protocol generation, review and export. ADR 0011 defines +the validated processing boundary and current MVP protocol path. ## Decision diff --git a/docs/adr/0011-use-meeting-lab-mvp-backend.md b/docs/adr/0011-use-meeting-lab-mvp-backend.md new file mode 100644 index 0000000..65a08a2 --- /dev/null +++ b/docs/adr/0011-use-meeting-lab-mvp-backend.md @@ -0,0 +1,94 @@ +# ADR 0011: Use Meeting Lab as the MVP Processing Backend + +## Status + +Accepted + +Supersedes ADR 0002 and amends ADRs 0001 and 0003. + +## Context + +Meeting Lab experiments have validated a practical local pipeline for turning +a recorded meeting into an editable protocol. Earlier Meeting Assistant plans +selected `faster-whisper`, treated diarization as required and described a +separate knowledge-extraction stage as part of the primary pipeline. Work on an +old feature branch also proposed mandatory fixed five-minute audio chunks and +chunk-oriented storage. + +The validated backend now has a reusable Python API, optional diarization and a +direct full-transcript protocol path. The product boundary must keep backend +processing in Meeting Lab and user interaction in Meeting Assistant. + +## Decision + +Meeting Lab is the reusable processing backend and experimental engine. +Meeting Assistant is the user-facing application and calls Meeting Lab through +its Python API rather than launching its CLI as a subprocess. + +The MVP processing flow is: + +```text +Source audio + -> FFmpeg preparation (mono, 16 kHz PCM WAV) + -> whisper.cpp transcription (large-v3-turbo) + -> optional pyannote.audio Community-1 diarization + -> direct full-transcript protocol generation + -> human review and editing +``` + +The productive path does not require fixed five-minute audio chunks. Internal +streaming or segmentation remains an implementation detail of Meeting Lab and +must not define Meeting Assistant's domain or storage model. + +Meeting Assistant integrates with these public Meeting Lab interfaces: + +- `MvpMeetingConfig` +- `MvpRunResult` +- `run_mvp_meeting(...)` +- stage-based progress events: `preparing`, `transcription`, `diarization`, + `protocol_generation`, `completed` and `failed` + +Progress percentages are displayed only when the backend reports real +measurable progress. + +`MeetingContext` is the structured input containing meeting metadata and +participants. It may contain explicitly confirmed `SPEAKER_XX -> participant_id` +mappings. Such mappings are authoritative only when confirmed by the user. +The system must not infer speaker names automatically. + +The preferred local engines are: + +- `whisper.cpp` with `large-v3-turbo` for transcription +- `pyannote.audio` Community-1 for optional diarization + +GPU acceleration is optional. Vulkan transcription and PyTorch/ROCm +diarization are validated acceleration paths on North's AMD RX 9070. CPU +execution remains the compatibility fallback. + +The direct full-transcript protocol path is the practical MVP direction. +`qwen3.8:27B` has produced strong readability and contextual synthesis, while +`qwen3.6:35B-A3B` has sometimes behaved more conservatively. Dual-model and +diarization-assisted hard-fact experiments have not established sufficiently +reliable attribution gains to justify mandatory MVP complexity. Model choice +remains backend configuration, and human review is required. + +Meeting Assistant owns the GUI, structured context editor, participant and +speaker-mapping controls, progress presentation, protocol editor and export. +Meeting Lab owns audio preparation, transcription, diarization, orchestration +and protocol-generation logic. + +The product should ultimately provide a detailed contextual protocol and a +shorter participant/distribution version. The shorter version may follow the +first GUI milestone. + +## Consequences + +- ADR 0002's `faster-whisper` selection is no longer current. +- Fixed-duration chunking and chunk-centric storage are not MVP requirements. +- Diarization and GPU acceleration remain optional. +- The Meeting Lab CLI remains useful for command-line operation but is not the + application integration boundary. +- Meeting Context is created through the future GUI; users are not expected to + hand-write YAML. +- Backend experiments can evolve without moving processing logic into the GUI. +- Protocols remain AI-assisted outputs that require human review. diff --git a/docs/architecture.md b/docs/architecture.md index eabcc99..a835dc3 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -2,254 +2,146 @@ ## Purpose -The Meeting Knowledge Assistant transforms recorded meetings into structured organizational knowledge through a modular processing pipeline. +Meeting Assistant is the user-facing application for preparing meeting +context, running the Meeting Lab pipeline and reviewing its results. Meeting +Lab is the reusable processing backend and experimental engine. ---- - -# Design Principles - -- Architecture first -- Offline first where practical -- Immutable source data -- Replaceable AI components -- Small, focused modules -- Explicit interfaces -- Reproducible AI outputs -- External recording and internal processing are separate responsibilities. -- The processing pipeline must not depend on OBS-specific metadata or behavior. -- Imported recordings must be treated like recordings from any other supported source. - ---- - -## High-Level Pipeline - -External Recorder - ↓ -Audio Import - ↓ -Recording Validation - ↓ -Meeting Storage - ↓ -Transcription - ↓ -Speaker Diarization - ↓ -Working Transcript - ↓ -LLM Analysis - ↓ -Knowledge Extraction - ↓ -Export - ---- - -## Core Pipeline - -External Recording - -↓ - -Audio Import and Validation - -↓ - -Transcription - -↓ - -Speaker Diarization - -↓ - -Working Transcript - -↓ - -LLM Analysis - -↓ - -Knowledge Extraction - -↓ - -Storage and Export - -# Domain Model - -Meeting -│ -├── Recording -├── Transcript -│ ├── Canonical -│ └── Working -├── Speakers -├── Participants -├── AI Artifacts -├── Knowledge Objects -├── Attachments -└── Exports - ---- - -# Module Responsibilities - -## Input and Ingest - -Responsible for importing existing recordings into the application. - -Initial reference source: - -- OBS Studio - -Initial supported formats: - -- WAV - -- FLAC - -Responsibilities: - -- validate the input file - -- collect technical metadata - -- calculate checksums - -- copy or move the recording into the meeting directory - -- create the initial Recording domain object - -The ingest component must not perform transcription or modify the audio content. - -## Recorder - -The recorder module is reserved for a future integrated recording implementation. - -It is not required for the MVP. - -The initial application workflow uses externally created recordings, with OBS -Studio as the recommended reference recorder. - ---- - -## Transcription - -Responsible only for speech-to-text conversion. - -Input: -Recording - -Output: -Canonical Transcript - ---- - -## Diarization - -Responsible only for speaker identification. - -Input: -Recording + Canonical Transcript - -Output: -Working Transcript - ---- - -## LLM - -Responsible for semantic analysis. - -Input: -Transcript - -Output: -AI Artifacts - ---- - -## Knowledge Extraction - -Responsible for creating structured knowledge. - -Input: -AI Artifacts - -Output: -Knowledge Objects - ---- - -## Export - -Responsible for creating user-facing documents. - -Input: -Knowledge Objects - -Output: -Markdown -PDF -DOCX - ---- - -## Artifact Lifecycle +## System Boundary ```text -Imported Recording - ↓ -Validated Source Recording - ↓ -Canonical Transcript - ↓ -Working Transcript - ↓ -AI Artifacts - ↓ -Knowledge Objects - ↓ -Exports +Meeting Assistant Meeting Lab +----------------- ----------- +Audio selection ------> Audio preparation +Meeting Context editor ------> Transcription +Participant management ------> Optional diarization +Explicit speaker mapping ------> Protocol generation +Progress presentation <------ Progress events +Protocol editor and export <------ Run result and artifacts +``` ---- +Meeting Assistant calls the Meeting Lab Python API directly. The Meeting Lab +CLI is a thin adapter over that API and must not be launched as an application +subprocess. -## Initial Recording Strategy +## MVP Processing Flow -The MVP does not implement platform-specific audio capture. +```text +Source audio + -> FFmpeg normalization/preparation + -> mono, 16 kHz PCM WAV + -> whisper.cpp transcription with large-v3-turbo + -> optional pyannote.audio Community-1 diarization + -> direct full-transcript protocol generation + -> human review and editing + -> export +``` -OBS Studio is the recommended reference recorder for online meetings. +The MVP does not require fixed five-minute audio chunks, semantic chunking, a +separate evidence-extraction pipeline or multiple LLMs. Any internal +segmentation remains a Meeting Lab implementation detail. -The application initially processes existing WAV or FLAC recordings. Integrated +## Backend Integration -recording remains a future extension and must not be required by transcription, +The reusable Meeting Lab interface consists of: -diarization or analysis modules. +- `MvpMeetingConfig`, the run configuration +- `run_mvp_meeting(...)`, the orchestration entry point +- `MvpRunResult`, the completed run result +- stage-based progress events ---- +The known stages are: -# Storage Strategy +```text +preparing +transcription +diarization +protocol_generation +completed +failed +``` -The domain model is independent of the storage backend. +The GUI shows the current stage. It shows a percentage only when the event +contains real measurable progress; stage changes must not be presented as +invented percentages. -Possible implementations: +## Meeting Context -- File System -- SQLite -- PostgreSQL -- Cloud Storage +`MeetingContext` is structured domain input rather than an informal prompt or a +file users must author manually. It contains meeting metadata and participants +and can include optional mappings: ---- +```text +SPEAKER_XX -> participant_id +``` -# Future Extensions +Mappings are authoritative only after explicit user confirmation. Diarization +labels otherwise remain anonymous, and the application must not infer speaker +names automatically. The Meeting Assistant GUI owns creation and editing of +this context and its mappings. -- Live transcription -- Video processing -- OCR -- Semantic search -- Knowledge graph -- Company glossary -- Multi-language meetings -- Local LLM support \ No newline at end of file +## Processing Components + +### Audio Preparation + +Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The +current practical target is mono, 16 kHz PCM WAV. The imported source remains a +separate source artifact. + +### Transcription + +`whisper.cpp` with `large-v3-turbo` is the preferred local backend. Vulkan is a +validated acceleration path on North's AMD RX 9070, while CPU execution remains +a compatibility fallback. Engine and model details should be retained with the +result for traceability. + +### Diarization + +`pyannote.audio` Community-1 is preferred when diarization is enabled. It can +run on CPU and may use GPU acceleration such as PyTorch/ROCm where available. +Diarization is optional and its anonymous labels do not modify the canonical +transcript or assert participant identity. + +### Protocol Generation + +The direct full-transcript path is the practical MVP. Current model experiments +favor `qwen3.8:27B` for readability and contextual synthesis, while +`qwen3.6:35B-A3B` has shown more conservative behavior in some areas. Neither a +specific dual-model arrangement nor an evidence pipeline is an application +requirement. Generated protocols require human review. + +## Application Responsibilities + +The first GUI milestone provides: + +- audio-file selection +- structured Meeting Context and participant editing +- optional explicit speaker mapping +- pipeline start and configuration +- progress and failure presentation +- detailed protocol display and editing +- export of the reviewed result + +The product direction also includes a shorter participant/distribution +protocol. It may be delivered after the first GUI milestone. + +## Artifact and Storage Principles + +The existing Meeting domain remains the application-level container for source +recordings, canonical and working transcripts, context, generated artifacts and +exports. Storage remains independent of processing segmentation. In particular, +chunks are not primary Meeting Assistant domain objects or the required unit of +MVP persistence. + +Source audio and the canonical transcription output are preserved. Corrections +and reviewed protocols are derived versions. Generated artifacts should retain +their input version, backend/model configuration, prompt version and timestamp +where practical. + +## Future Extensions + +- shorter distribution protocols +- integrated recording +- searchable meeting history and optional database indexes +- live transcription +- OCR and video processing +- semantic search and organizational knowledge features diff --git a/docs/glossary.md b/docs/glossary.md index 1ba9bef..cf807ca 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -2,7 +2,23 @@ Transcript Original speech-to-text output. Diarization -Assignment of transcript segments to speakers. +Assignment of transcript segments to anonymous speaker labels. It does not by +itself identify participants. + +Meeting Context +Structured meeting metadata, participants and optional explicitly confirmed +speaker-to-participant mappings supplied to the processing pipeline. + +Speaker Mapping +An explicitly confirmed association from a diarization label such as +`SPEAKER_00` to a participant ID. It must not be inferred automatically. + +Detailed Protocol +Contextual protocol generated from the full transcript for human review and +continued work. + +Distribution Protocol +Shorter, participant-facing version of a reviewed meeting protocol. Knowledge Object Structured information extracted from meetings.