Files
meeting-lab/PROJECT_KNOWLEDGE.md
T

302 lines
14 KiB
Markdown

# Project Knowledge
## Meeting language in direct protocols
Direct protocol prompts derive their explicit output language from persisted
Meeting Context `meeting.language`, including regeneration and diarized fallback.
Missing context/language explicitly defaults to `de`; omitted language is accepted
without modifying context data. Protocol runtime `output_language` is derived
provenance, not a separate setting. No transcript/context translation is performed.
Prompt evolution (2026-09-11): replaced unconditional German output with the
meeting-language instruction in the shared prompt builder. Names remain verbatim.
Validation uses mocked generation for German/English, plain/diarized inputs,
regeneration and legacy contexts; no live model or extraction gold run is involved.
This is a compact operational summary of the current Meeting Lab state.
## Objective
Meeting Lab develops and evaluates local methods for extracting structured
organizational knowledge from real meeting recordings and transcripts. The
project is a research and validation environment for a future Meeting
Assistant, not a finished product.
## Implemented Pipeline Stages
Implemented:
- Whisper JSON cleanup via `scripts/clean_whisper_json.py`.
- Transcript normalization in `src/meeting_lab/normalization/`.
- Technical chunking in `src/meeting_lab/chunking/`.
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
- Deterministic Canonicalizer V1 in
`src/meeting_lab/consolidation/canonicalize.py`.
- Semantic Consolidator V0 in
`src/meeting_lab/consolidation/consolidate_facts.py` for facts-only
semantic duplicate detection.
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Meeting Context V1 loading, validation and optional extraction prompt
injection with minimal extraction JSON provenance.
- FFmpeg-backed WAV, FLAC and M4A preparation into a per-run canonical mono
16 kHz signed PCM16 WAV artifact before transcription or diarization. Audio
preparation always runs. Optional loudness normalization defaults to on and
currently uses the isolated FFmpeg filter
`loudnorm=I=-16:LRA=11:TP=-1.5`. This is a conservative speech-recording
default and may be revisited after empirical comparison without changing the
orchestration API.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Direct protocol prompt input protection: diarized transcripts are rendered as
compact adjacent-speaker blocks without per-segment timestamps. Every source
segment remains represented in order. A deterministic heuristic enforces a
configurable safe input budget, falls back to complete plain transcript text
when necessary, and fails before any Ollama request if even that input is too
large. Silent head/tail truncation is prohibited.
- The `qwen3.8:27b` direct-protocol stage explicitly requests `num_ctx=32768`
and `think=false`; the practical prompt target is approximately 29,000 tokens.
A 31,038-token synthetic prompt passed, but larger prompts are not assumed safe
from the model's advertised 262,144-token native context alone.
- `regenerate_mvp_protocol` updates the run's validated Meeting Context and
regenerates protocol artifacts from the existing diarized transcript when
available. It never reruns audio preparation, Whisper or Pyannote, and it
preserves anonymous speaker labels in the source transcript.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation.
- Meeting Context V1 scaffold and documentation for manually maintained
meeting metadata.
Experimental/prototype:
- Topic segmentation in `src/meeting_lab/segmentation/`.
- Windowed segmentation and review output in `samples/chunks/`.
- Gold Standard extraction corpus under `tests/gold/`.
Planned:
- Broader semantic consolidation for topic grouping, contradiction handling,
uncertainty marking and durable/transient separation.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Meeting Context integration with Canonicalizer, Semantic Consolidator,
Canonical Meeting Knowledge and output renderers.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
## Source Tree
```text
src/meeting_lab/
chunking/ technical transcript chunking
consolidation/ planned canonicalization/consolidation area
extraction/ current local LLM extraction flow
io/ lightweight file and JSON helpers
llm/ Ollama and prompt support
models/ current lightweight model definitions
normalization/ deterministic transcript cleanup
protocol/ interim Markdown protocol builder
segmentation/ experimental topic segmentation tooling
```
Supporting areas:
- `docs/`: architecture, pipeline, data models and output-view concepts.
- `docs/meeting-context.md`: Meeting Context V1 scaffold, fields and future
integration rules.
- `prompts/`: active extraction prompt files. The shared extraction prompt is
assembled from `common.md`, `decisions.md` and `todos.md`.
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
- `samples/`: sample inputs and generated or experimental artifacts.
- `scripts/`: operational scripts for cleanup and gold-test execution.
## Current Model Strategy
The current extraction strategy is one normalized chunk per LLM call. This is
preferred over expanding context windows or asking one model call to analyze a
full meeting.
Known working models from current project notes and experiment practice:
- `qwen3:1.7b`: useful for smoke tests.
- `qwen3.5:9b`: useful for meaningful extraction and segmentation work.
LLM calls use Ollama locally. The current extractor defaults to `qwen3:8b`, but
validated work may specify another model explicitly.
## Important Findings
- Whisper JSON chunking must use `segments[*].text`, not only the top-level
`text` field.
- Independent chunk extraction is currently preferred.
- Larger context windows can change classification behavior and increase
instability.
- Extraction and consolidation are separate problems.
- Generation limits can truncate JSON.
- Qwen thinking may be returned separately by the Ollama API.
- Gold Standard tests are also a formal specification of meeting semantics.
- Raw model responses should be preserved when diagnosing parser or truncation
failures.
- Responsibility attribution is a critical correctness invariant: people,
teams and departments must not be assigned ownership from discussion,
expertise, objection, thematic proximity, speaker adjacency, role guesses or
world knowledge. Assignment requires explicit evidence.
## Decision Taxonomy
Accepted decision semantics:
- A decision is an explicit agreement that creates a binding change in action,
process, responsibility, approval status, timing or next step.
- Included: substantive decisions, organizational decisions, process decisions,
approvals, rejections, deferrals, explicit agreement not to decide yet, and
explicit agreement to gather more information before deciding.
- Excluded: opinions, preferences, proposals without agreement, open questions,
current-state descriptions and explanations without commitment.
- A process decision to defer a substantive decision is still a decision.
- "No decision was reached" is different from "the group decided to defer the
decision."
- A personal commitment to perform concrete future work is normally a todo,
not a decision, unless the group also establishes a separate binding outcome,
rule, approval, rejection, deferral, selection, process state or
responsibility policy.
- The same proposition should not be duplicated under decisions and todos.
Extract both only when the transcript contains a group-level decision and a
semantically separate resulting action item.
## Responsibility Attribution
Meeting Lab distinguishes mentioned people, speakers, participants,
responsible people, departments, owners and assignees. These concepts must not
be collapsed into one field.
Meeting Context V1 reinforces this distinction by separating actual
participants from mentioned non-participants and by storing aliases, roles and
departments only when they are explicitly supplied as metadata. It must not be
used to infer responsibilities. In the current implementation this context can
be injected into chunk extraction prompts as authoritative metadata, and only
minimal provenance is written to extraction JSON.
The implemented MVP statuses are exactly `present` and `mentioned_only`.
Legacy entries without a status receive collection-appropriate defaults. Only
present participants may be targets of explicit `SPEAKER_XX` mappings.
A `responsible` or future `owner` / `assignee` value may be recorded only when
source evidence explicitly assigns, accepts or confirms responsibility. If the
evidence is incomplete or ambiguous, the responsible person remains `null` or
unset and the evidence is preserved. Future schema work may add
`responsibility_status` values such as `explicit`, `accepted`, `proposed` and
`unclear`, plus `attribution_evidence`.
Meeting Context may validate identity, role, department and attendance, but it
never establishes responsibility.
Focused Gold coverage now separates these concerns:
- `responsibility_attribution_negative`: tests that discussion, objection and
department proximity do not create an owner.
- `position_explicit_objection`: tests explicit position extraction separately
from responsibility attribution.
Current Prompt Version 2 decision baseline:
- `decision_simple`: passing.
- `decision_deferred`: passing.
- `decision_none`: passing.
- Prompt Version 2 explicitly supports process decisions where the group agrees
to defer a substantive decision until more information is available.
## Canonical Knowledge Architecture
The next documented pipeline milestone is:
```text
Chunk Extractions
-> Deterministic Canonicalizer
-> Semantic Consolidator
-> Canonical Meeting Knowledge
-> Output View Renderers
```
Canonicalizer V1 is implemented Python code with no LLM. It validates and
normalizes extraction objects, assigns stable source references and IDs,
normalizes category names and basic field structure, performs only safe
deterministic cleanup, optionally groups exact duplicates, and preserves all
source evidence. It must not perform uncertain semantic merging.
CLI:
```text
PYTHONPATH=src .venv/bin/python -m meeting_lab.consolidation.canonicalize \
samples/whisper/meeting_speech_cleaned_chunks \
-o /tmp/canonicalized_extractions.json
```
Semantic Consolidator V0 is implemented as narrow local-LLM work for fact
items only. It merges semantically equivalent fact statements, preserves source
references and evidence, and prefers false negatives over false-positive
merges. It is not a summarizer, topic grouper, protocol renderer or complete
Canonical Meeting Knowledge stage.
The first accepted V0 benchmark used `qwen3.5:9B` in one Ollama call over 33
fact items. Runtime on the current machine was 390.119 seconds. One correct
merge was accepted, involving `fact_0025` and `fact_0031`; 31 facts remained
singletons, validation passed, no source fact was lost or duplicated, and
non-fact categories remained unchanged. This is a local benchmark, not a
general hardware claim.
Broader semantic consolidation remains planned. It should group content by
topic, mark contradictions and uncertainty, separate durable information from
transient discussion, and produce Canonical Meeting Knowledge. It does not
directly write a protocol.
Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should be structured, preferably JSON, and
preserve topics, facts, decisions, action items, open questions, positions,
technical details, rationale, uncertainty, contradictions and source evidence.
It is not itself a prose protocol.
Output views are planned as independent renderings from that canonical model:
- Working Protocol / Arbeitsprotokoll: relatively complete, optimized for
recall and traceability.
- Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented,
optimized for circulation.
- Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge
optimized for reuse.
The current `meeting_protocol.md` builder is an interim technical validation
tool, not the final output-view architecture.
Rendered protocol output should normally use the dominant language of the
source transcript or consolidated meeting knowledge unless an explicit output
language is requested.
## Current Limitations
- Discussion Blocks are documented as a stable semantic unit but are not yet a
separate implemented pipeline artifact.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors
are the intended architecture.
- Semantic Consolidator V0 is implemented only for facts-only duplicate
detection.
- Meeting Context V1 is implemented only through chunk extraction; later-stage
integration remains planned.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Some prompt files remain placeholders; `common.md`, `decisions.md` and
`todos.md` are active in the shared extraction prompt.
- Gold tests currently emphasize extraction semantics, especially decisions,
todos, responsibility attribution and focused position extraction.
## Next Recommended Engineering Step
Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks.
The next recommended evaluation step is to use the consolidated V0 result as
input for the unchanged Working Protocol renderer and compare that output
against the Working Protocol Synthesizer V0 baseline and the human reference
protocol.
Protocol calls accept an optional positive `protocol_num_thread` setting through
both initial processing and regeneration. With `None`, Ollama receives no
`num_thread` override and selects its own thread configuration.