Add Streamlit meeting assistant MVP
This commit is contained in:
@@ -0,0 +1,67 @@
|
||||
# ADR 0012: Preserve Corrections as Post-Run Knowledge
|
||||
|
||||
## Status
|
||||
|
||||
Accepted as a future product direction; not required for the current MVP.
|
||||
|
||||
## Context
|
||||
|
||||
Real-meeting review exposes corrections that should not require repeating
|
||||
expensive processing. These include misspelled or recurring name variants,
|
||||
people mentioned but omitted from the initial Meeting Context, and confirmed
|
||||
`SPEAKER_XX -> participant_id` mappings. A destructive edit would lose useful
|
||||
provenance, while rerunning Whisper or diarization would add cost without
|
||||
improving a deterministic correction.
|
||||
|
||||
The product is also expected to provide both a detailed contextual protocol and
|
||||
a short participant/distribution protocol. Both should eventually use the same
|
||||
confirmed context and corrections.
|
||||
|
||||
## Decision
|
||||
|
||||
Original machine-generated artifacts remain immutable. Corrected or reviewed
|
||||
artifacts are separate derivatives. Human corrections should evolve into
|
||||
structured, meeting-specific knowledge rather than opaque destructive edits;
|
||||
the correction schema is deliberately not defined by this ADR.
|
||||
|
||||
An early correction tool may be simple deterministic search and replace, for
|
||||
example `Grossman` to `Herr Grossmann`. Applying such a confirmed correction to
|
||||
a suitable editable artifact requires neither Whisper, diarization nor an LLM
|
||||
call.
|
||||
|
||||
A later review stage may run after diarization or the complete initial run. It
|
||||
may propose likely person/name matches, allow addition of previously omitted
|
||||
mentioned people, and present anonymous speaker mappings for review. Every
|
||||
suggestion is non-authoritative: anonymous labels remain anonymous until the
|
||||
user explicitly confirms or corrects them, and `mentioned_only` people cannot
|
||||
be mapped as speakers.
|
||||
|
||||
After confirmation, the product should offer two paths:
|
||||
|
||||
1. A fast path applies confirmed speaker, name or text corrections
|
||||
deterministically where that is semantically safe.
|
||||
2. A quality path regenerates protocol output using the existing transcription,
|
||||
existing diarization, corrected Meeting Context and confirmed mappings or
|
||||
corrections. It reruns protocol generation only. Whisper and diarization run
|
||||
again only when separately requested or technically necessary.
|
||||
|
||||
The likely later workflow is therefore:
|
||||
|
||||
```text
|
||||
Audio -> transcription -> optional diarization -> initial protocol
|
||||
-> name/person/speaker review -> human confirmation
|
||||
-> deterministic correction or protocol-only regeneration
|
||||
-> final reviewed detailed and/or distribution protocol
|
||||
```
|
||||
|
||||
The exact ordering may evolve, and this review stage is not mandatory for the
|
||||
current Streamlit MVP.
|
||||
|
||||
## Consequences
|
||||
|
||||
- Human knowledge can improve outputs without unnecessary upstream work.
|
||||
- Provenance is retained because originals and corrected derivatives coexist.
|
||||
- Correction data can later be reused consistently across detailed and short
|
||||
protocol views.
|
||||
- Search/replace, correction storage, review UI, identity suggestions and
|
||||
protocol-only rerun controls remain future implementation work.
|
||||
+15
-2
@@ -27,7 +27,7 @@ subprocess.
|
||||
|
||||
```text
|
||||
Source audio
|
||||
-> FFmpeg normalization/preparation
|
||||
-> FFmpeg preparation (normalization optional, default on)
|
||||
-> mono, 16 kHz PCM WAV
|
||||
-> whisper.cpp transcription with large-v3-turbo
|
||||
-> optional pyannote.audio Community-1 diarization
|
||||
@@ -85,7 +85,11 @@ this context and its mappings.
|
||||
|
||||
Meeting Lab uses FFmpeg to prepare a consistent local-processing input. The
|
||||
current practical target is mono, 16 kHz PCM WAV. The imported source remains a
|
||||
separate source artifact.
|
||||
separate source artifact. Preparation always runs for WAV, FLAC and M4A,
|
||||
regardless of the normalization switch. When enabled, Meeting Lab currently
|
||||
uses `loudnorm=I=-16:LRA=11:TP=-1.5`, an isolated conservative default for
|
||||
speech recordings that may be revisited after empirical comparison. Meeting
|
||||
Assistant passes only an on/off choice and does not own filter parameters.
|
||||
|
||||
### Transcription
|
||||
|
||||
@@ -137,6 +141,15 @@ and reviewed protocols are derived versions. Generated artifacts should retain
|
||||
their input version, backend/model configuration, prompt version and timestamp
|
||||
where practical.
|
||||
|
||||
## Later Post-Run Correction Flow
|
||||
|
||||
Post-run corrections are planned as a separate, non-mandatory workflow after
|
||||
initial protocol generation. As detailed in [ADR 0012](adr/0012-post-run-corrections.md),
|
||||
confirmed name, person and anonymous-speaker corrections should become
|
||||
meeting-specific knowledge. They may then be applied deterministically to safe
|
||||
derived artifacts or used for protocol-only regeneration without needlessly
|
||||
rerunning transcription or diarization.
|
||||
|
||||
## Future Extensions
|
||||
|
||||
- shorter distribution protocols
|
||||
|
||||
Reference in New Issue
Block a user