diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..432575a --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,124 @@ +# AGENTS.md + +Practical instructions for coding agents working in Meeting Lab. + +## Project Purpose + +Meeting Lab extracts and structures organizational knowledge from meeting +recordings. It is an experimental local discussion analyzer, not merely a +one-step meeting-protocol generator. + +Successful approaches may later move into the Meeting Assistant project. + +## Current Pipeline + +Current and intended flow: + +```text +Audio +-> Whisper +-> cleanup +-> normalization +-> chunking +-> local chunk extraction +-> consolidation +-> Canonical Meeting Knowledge +-> Output Views +``` + +Status: + +- Implemented: Whisper JSON cleanup script, normalization, technical chunking, + local chunk extraction, interim Markdown protocol builder. +- Experimental/prototype: topic segmentation and review tooling. +- Planned: consolidation, Canonical Meeting Knowledge implementation, final + Output Views. + +## Architectural Principles + +- Canonical Meeting Knowledge is the intended semantic source of truth. +- Working Protocol / Arbeitsprotokoll, Distribution Protocol / + Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are + parallel output views. +- Output views must not silently change meaning. They may select, condense or + render information for an audience, but not invent new semantics. +- Extraction, consolidation, synthesis and rendering are separate concerns. +- Prefer small, testable processing stages over one monolithic LLM prompt. +- Current extraction strategy is one normalized chunk per LLM call. +- Do not expand context windows or redesign the extraction strategy without an + explicit experiment. +- Deterministic stages should remain deterministic where possible. + +## Prompt Engineering Rules + +The Gold Standard corpus is the reference specification. Follow Rules 1-11 from +`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`: + +1. Make only one prompt change per iteration. +2. Optimize only one target gold test case at a time. +3. Validate every prompt modification immediately. +4. Accept a prompt change only if it improves the target and causes no + regressions in previously passing gold tests. +5. Never modify `expected.json` merely to make a prompt pass. +6. Prompt engineering edits prompt files only; Python code changes require a + separate explicit task. +7. Maintain a prompt evolution log for every iteration. +8. Stop arbitrary iterations if small changes do not improve the test; analyze + the root cause. +9. Avoid gold-test overfitting. Prompt changes must generalize and must not + special-case one transcript. +10. Stop after two consecutive non-improving prompt iterations and classify the + root cause. +11. Verify whether the target gold test has objectively unique ground truth + before changing a prompt for unexpected behavior. + +Current documented Prompt Version 2 decision baseline: + +- `decision_simple`: passing +- `decision_deferred`: passing +- `decision_none`: passing + +## LLM Execution Safety + +- Never start a full multi-chunk LLM run unless explicitly requested. +- Before any LLM run, state the model, inputs, expected LLM-call count and + output location. +- Do not retry LLM calls automatically unless explicitly allowed. +- Do not download models automatically. +- Prefer small-scope validation runs. +- Never use generated output as committed source data. +- Preserve raw model responses when diagnosing parser or truncation failures. +- Do not run Ollama from unit tests. + +## Development Rules + +- Make small, focused changes. +- Preserve the existing architecture unless a redesign is explicitly requested. +- Add regression tests for bugs. +- Run non-LLM tests before committing when code changes are made. +- Do not commit generated transcripts, audio, extraction JSON, protocol output + or temporary files. +- Report files changed, tests run and assumptions. +- Do not commit or push unless explicitly requested. + +## Repository Conventions + +- `README.md`: project overview and current high-level status. +- `docs/`: architecture, pipeline, data model and output-view documentation. +- `prompts/`: extraction and segmentation prompts. Treat prompt edits as + controlled experiments. +- `tests/gold/`: Gold Standard corpus and semantic specification for extraction + behavior. +- `scripts/`: command-line support scripts such as Whisper cleanup and gold + test execution. +- `src/meeting_lab/normalization/`: deterministic transcript cleanup. +- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics. +- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling. +- `src/meeting_lab/extraction/`: local LLM extraction flow and category + extractor modules. +- `src/meeting_lab/consolidation/`: planned consolidation area. +- `src/meeting_lab/protocol/`: interim protocol rendering. +- `src/meeting_lab/models/`: current lightweight data models. +- `samples/`: sample inputs and generated/experimental artifacts; do not treat + sample output as canonical source data. + diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..11af922 --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,47 @@ +# Changelog + +## Unreleased + +### Added + +- Initial Meeting Lab project structure with source, docs, prompts, samples and + tests directories. +- Deterministic transcript normalization module. +- Technical transcript chunking with block-aligned chunk generation. +- Whisper JSON cleanup support. +- Local Ollama-based chunk extraction flow. +- Interim Markdown meeting protocol builder for technical validation. +- Initial and windowed topic segmentation prototypes plus review tooling. +- Gold Standard extraction corpus and gold-test runner. +- Prompt loading support and common/decision prompt baseline. +- Output-view architecture documentation for Canonical Meeting Knowledge, + Working Protocol / Arbeitsprotokoll, Distribution Protocol / + Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag. + +### Changed + +- Refined the architecture from a single meeting protocol toward Canonical + Meeting Knowledge as the planned semantic source of truth. +- Clarified that Working Protocol, Distribution Protocol and Knowledge Objects + are parallel renderings, not derived from one another. +- Updated decision extraction semantics to include explicit process decisions + and deferrals. +- Simplified extraction prompt assembly around prompt files. + +### Fixed + +- Fixed Whisper JSON chunk extraction to prefer `segments[*].text` over the + aggregate top-level `text` field. +- Added chunking tests to ensure chunks do not duplicate later blocks when no + overlap is requested. +- Added parser handling for model responses that contain thinking text before + the final JSON object. + +### Documentation + +- Added architecture, pipeline and data-model documentation. +- Added output-view documentation. +- Added Gold Standard prompt-engineering methodology. +- Added formal decision-definition documentation. +- Added scenario README files for the Gold Standard corpus. + diff --git a/PROJECT_KNOWLEDGE.md b/PROJECT_KNOWLEDGE.md new file mode 100644 index 0000000..82f1e34 --- /dev/null +++ b/PROJECT_KNOWLEDGE.md @@ -0,0 +1,151 @@ +# Project Knowledge + +This is a compact operational summary of the current Meeting Lab state. + +## Objective + +Meeting Lab develops and evaluates local methods for extracting structured +organizational knowledge from real meeting recordings and transcripts. The +project is a research and validation environment for a future Meeting +Assistant, not a finished product. + +## Implemented Pipeline Stages + +Implemented: + +- Whisper JSON cleanup via `scripts/clean_whisper_json.py`. +- Transcript normalization in `src/meeting_lab/normalization/`. +- Technical chunking in `src/meeting_lab/chunking/`. +- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`. +- Prompt loading from `src/meeting_lab/llm/prompts.py`. +- Interim Markdown protocol generation in `src/meeting_lab/protocol/`. +- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and + gold-test runner validation. + +Experimental/prototype: + +- Topic segmentation in `src/meeting_lab/segmentation/`. +- Windowed segmentation and review output in `samples/chunks/`. +- Gold Standard extraction corpus under `tests/gold/`. + +Planned: + +- Consolidation of extraction results. +- Canonical Meeting Knowledge implementation as the semantic source of truth. +- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol / + Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers. + +## Source Tree + +```text +src/meeting_lab/ + chunking/ technical transcript chunking + consolidation/ planned merge/consolidation area + extraction/ current local LLM extraction flow + io/ lightweight file and JSON helpers + llm/ Ollama and prompt support + models/ current lightweight model definitions + normalization/ deterministic transcript cleanup + protocol/ interim Markdown protocol builder + segmentation/ experimental topic segmentation tooling +``` + +Supporting areas: + +- `docs/`: architecture, pipeline, data models and output-view concepts. +- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain + substantive extraction prompt text in the current tree. +- `tests/gold/`: semantic gold tests and prompt-engineering methodology. +- `samples/`: sample inputs and generated or experimental artifacts. +- `scripts/`: operational scripts for cleanup and gold-test execution. + +## Current Model Strategy + +The current extraction strategy is one normalized chunk per LLM call. This is +preferred over expanding context windows or asking one model call to analyze a +full meeting. + +Known working models from current project notes and experiment practice: + +- `qwen3:1.7b`: useful for smoke tests. +- `qwen3.5:9b`: useful for meaningful extraction and segmentation work. + +LLM calls use Ollama locally. The current extractor defaults to `qwen3:8b`, but +validated work may specify another model explicitly. + +## Important Findings + +- Whisper JSON chunking must use `segments[*].text`, not only the top-level + `text` field. +- Independent chunk extraction is currently preferred. +- Larger context windows can change classification behavior and increase + instability. +- Extraction and consolidation are separate problems. +- Generation limits can truncate JSON. +- Qwen thinking may be returned separately by the Ollama API. +- Gold Standard tests are also a formal specification of meeting semantics. +- Raw model responses should be preserved when diagnosing parser or truncation + failures. + +## Decision Taxonomy + +Accepted decision semantics: + +- A decision is an explicit agreement that creates a binding change in action, + process, responsibility, approval status, timing or next step. +- Included: substantive decisions, organizational decisions, process decisions, + approvals, rejections, deferrals, explicit agreement not to decide yet, and + explicit agreement to gather more information before deciding. +- Excluded: opinions, preferences, proposals without agreement, open questions, + current-state descriptions and explanations without commitment. +- A process decision to defer a substantive decision is still a decision. +- "No decision was reached" is different from "the group decided to defer the + decision." + +Current Prompt Version 2 decision baseline: + +- `decision_simple`: passing. +- `decision_deferred`: passing. +- `decision_none`: passing. +- Prompt Version 2 explicitly supports process decisions where the group agrees + to defer a substantive decision until more information is available. + +## Canonical Knowledge Architecture + +Canonical Meeting Knowledge is the planned semantic intermediate model and +future single source of truth. It should preserve topics, facts, decisions, +action items, open questions, positions, technical details, rationale, +uncertainty, contradictions and source evidence. + +Output views are planned as independent renderings from that canonical model: + +- Working Protocol / Arbeitsprotokoll: relatively complete, optimized for + recall and traceability. +- Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented, + optimized for circulation. +- Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge + optimized for reuse. + +The current `meeting_protocol.md` builder is an interim technical validation +tool, not the final output-view architecture. + +## Current Limitations + +- Discussion Blocks are documented as a stable semantic unit but are not yet a + separate implemented pipeline artifact. +- Topic segmentation exists as prototype tooling, not a stable pipeline stage. +- Extraction is still a combined current flow, even though separate extractors + are the intended architecture. +- Consolidation is not implemented. +- Canonical Meeting Knowledge is documented but not implemented. +- Final output views are documented but not implemented. +- Most prompt files are placeholders except the common and decision prompts. +- Gold tests currently emphasize extraction semantics, especially decisions. + +## Next Recommended Engineering Step + +Stabilize repeatable local extraction evaluation before broadening the pipeline: +expand Gold Standard coverage by category, keep one-chunk extraction as the +baseline, and use small prompt experiments with immediate non-regression checks. +After extraction behavior is stable enough, implement consolidation with +evidence retention as the next major pipeline stage. diff --git a/ROADMAP.md b/ROADMAP.md new file mode 100644 index 0000000..fa0f6a4 --- /dev/null +++ b/ROADMAP.md @@ -0,0 +1,185 @@ +# Roadmap + +No dates are assigned. Phases describe dependency order, not release promises. + +## Phase 1 - Stable Local Extraction + +Goal: + +- Establish reliable per-chunk extraction behavior for core meeting semantics. + +Deliverables: + +- Stronger Gold Standard coverage across facts, positions, decisions, todos, + questions and technical details. +- Improved category prompts. +- Repeatable evaluation workflow. +- Documented prompt experiment log. + +Prerequisites: + +- Existing chunk extraction flow. +- Existing Gold Standard runner and methodology. + +Out of scope: + +- Full-transcript LLM extraction. +- Larger context-window strategy changes without an explicit experiment. +- Consolidation or final protocol rendering. + +## Phase 2 - Consolidation + +Goal: + +- Merge independent extraction results into a coherent meeting-level + representation without losing evidence. + +Deliverables: + +- Duplicate merging. +- Evidence retention. +- Category-shift reconciliation, especially facts versus positions and + positions versus decisions. +- Contradiction and uncertainty markers. +- Consolidated meeting representation. + +Prerequisites: + +- Stable local extraction baseline. +- Gold tests that expose cross-chunk duplication and category shifts. + +Out of scope: + +- Final Canonical Meeting Knowledge schema. +- User-facing protocol polish. +- Retrieval or RAG integration. + +## Phase 3 - Canonical Meeting Knowledge + +Goal: + +- Define and implement the semantic intermediate model that becomes the source + of truth for downstream outputs. + +Deliverables: + +- Canonical Meeting Knowledge schema. +- Source evidence and traceability fields. +- Clear distinction between durable knowledge and meeting-specific actions. +- Migration path from consolidated extraction JSON into the canonical model. + +Prerequisites: + +- Consolidation behavior that preserves evidence and uncertainty. +- Agreement on required semantic categories. + +Out of scope: + +- GUI. +- Export formats beyond those needed to validate the model. +- Knowledge-system storage design. + +## Phase 4 - Output Views + +Goal: + +- Render purpose-specific outputs from Canonical Meeting Knowledge without + changing meaning. + +Deliverables: + +- Working Protocol / Arbeitsprotokoll renderer. +- Distribution Protocol / Verteilerprotokoll renderer. +- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export. +- Tests or checks showing that output views are parallel renderings of the same + canonical model. + +Prerequisites: + +- Implemented Canonical Meeting Knowledge. +- Clear audience and completeness rules for each output view. + +Out of scope: + +- Additional analysis during rendering. +- Deriving one output view from another. +- Retrieval integration. + +## Phase 5 - Review and Quality Control + +Goal: + +- Add optional review stages that improve omission detection, consistency and + model selection. + +Deliverables: + +- Optional whole-transcript review. +- Omission detection. +- Consistency checks. +- Model comparison workflow. +- Hardware and runtime benchmarks. + +Prerequisites: + +- Stable extraction, consolidation and canonical model. +- Representative test meetings. + +Out of scope: + +- Automatic acceptance of review suggestions without evidence. +- Product UI work. +- Cloud deployment. + +## Phase 6 - Productization + +Goal: + +- Turn the validated pipeline into a usable local workflow. + +Deliverables: + +- Recording/transcription workflow. +- FFmpeg integration. +- Meeting metadata capture. +- Participant entry. +- GUI. +- Stable deployment process. +- Export workflows. + +Prerequisites: + +- Stable pipeline stages and output views. +- Clear operational requirements for local use. + +Out of scope: + +- Enterprise knowledge retrieval. +- Future Meeting Assistant integration beyond export contracts. +- Cloud-first architecture. + +## Phase 7 - Knowledge-System Integration + +Goal: + +- Reuse durable meeting knowledge in broader knowledge systems. + +Deliverables: + +- Structured Knowledge Objects. +- Retrieval-ready storage format. +- Future RAG integration path. +- Reuse contracts for Meeting Assistant and other knowledge systems. + +Prerequisites: + +- Canonical Meeting Knowledge and Knowledge Objects are implemented and stable. +- Durable knowledge is separated from meeting-specific actions and discussion + history. + +Out of scope: + +- Building a full enterprise search product inside Meeting Lab. +- Treating raw transcripts or generated protocols as the knowledge source of + truth. + diff --git a/docs/experiments.md b/docs/experiments.md index e69de29..3cb9315 100644 --- a/docs/experiments.md +++ b/docs/experiments.md @@ -0,0 +1,971 @@ +# Experiments + +This file records durable technical experiments and findings for Meeting Lab. +It is not a diary and does not replace commit history. + +## Status values + +- Proposed: experiment idea exists, but no result is recorded. +- Running: experiment is in progress and no decision has been made. +- Accepted: finding is the current baseline or design conclusion. +- Rejected: hypothesis was tested and should not be repeated as-is. +- Superseded: finding was useful but has been replaced by a newer baseline. + +## EXP-0001 - Whisper JSON interpretation + +Status: Accepted + +Date or period: 2026-07-29 + +Hypothesis: + +Whisper JSON should be chunked from its segment stream, not from the aggregate +top-level text field. + +Setup: + +`chunk_transcript.py` was updated to parse JSON input and prefer +`segments[*].text` when `segments` exists. A regression test supplies JSON with +both top-level `text` and separate segment texts. + +Inputs: + +- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`. +- Real Whisper artifacts under `samples/whisper/`. + +Model / configuration: + +- No LLM. + +Result: + +The test verifies that the block stream is `["alpha", "beta", "gamma"]` and +does not include the aggregate `"alpha beta gamma"` text. The repository history +records this as the fix for the earlier failure where the first chunk contained +the complete transcript. + +Decision: + +When `segments` exists, `segments[*].text` is the authoritative transcript +stream. The top-level `text` field is only a fallback. + +Lessons learned: + +Whisper JSON is structured input. Treating it like plain text can duplicate the +entire transcript and invalidate downstream chunking. + +Evidence: + +- `src/meeting_lab/chunking/chunk_transcript.py` +- `tests/test_chunking.py` +- Commit `4656523` - `Fix Whisper JSON chunk extraction` + +## EXP-0002 - Technical transcript chunking baseline + +Status: Accepted + +Date or period: 2026-07-29 to 2026-07-30 + +Hypothesis: + +Sequential technical chunks around the configured target size can preserve the +transcript while keeping extraction calls small enough for local models. + +Setup: + +The chunker splits block-aligned text with configurable target, minimum, +maximum and overlap settings. Tests verify no duplicate later blocks when +overlap is zero. Generated meeting artifacts contain a nine-chunk manifest. + +Inputs: + +- Synthetic block list in `tests/test_chunking.py`. +- `samples/whisper/meeting_speech_cleaned.json`. + +Model / configuration: + +- No LLM for chunking. +- Manifest uses default chunking behavior recorded in + `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`. + +Result: + +With overlap set to zero, tests verify that all blocks appear exactly once. The +real sample manifest contains nine chunks, mostly near the configured target +size, with a smaller final chunk. + +Decision: + +Independent sequential chunks are the current technical baseline. One +normalized chunk per extraction call is the preferred extraction strategy. + +Lessons learned: + +Chunking solves model-size constraints only. It must not perform topic +detection or semantic merging. + +Evidence: + +- `src/meeting_lab/chunking/chunk_transcript.py` +- `tests/test_chunking.py` +- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json` +- `AGENTS.md` +- `PROJECT_KNOWLEDGE.md` + +## EXP-0003 - Conservative transcript normalization + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Transcript cleanup should improve readability without changing meeting +semantics. + +Setup: + +The normalizer removes isolated filler sounds, immediate duplicate words or +short duplicate phrases, and redundant whitespace. It records changed blocks in +a JSON change log and explicitly preserves semantic content categories. + +Inputs: + +- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`. +- Change logs such as `chunk_01_changes.json`. + +Model / configuration: + +- No LLM. + +Result: + +The implementation and generated change logs show a conservative policy: +negations, qualifiers, dates, numbers, responsibilities, technical statements, +deadlines, decisions and commitments are preserved. + +Decision: + +Normalization remains deterministic and low-risk. When uncertain, leave text +unchanged. + +Lessons learned: + +Filler removal is useful only if it is tightly scoped. Broad cleanup can remove +semantic cues needed by extraction. + +Evidence: + +- `src/meeting_lab/normalization/normalize_transcript.py` +- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json` +- `docs/pipeline.md` + +## EXP-0004 - Full-context topic segmentation + +Status: Superseded + +Date or period: 2026-07-21 to 2026-07-22 + +Hypothesis: + +A single full-context topic segmentation call can identify topic boundaries in +a normalized transcript chunk. + +Setup: + +The initial segmentation prototype asked the model for topic changes and then +converted those boundaries into continuous, non-overlapping segments. + +Inputs: + +- `samples/chunks/chunk_01_normalized.txt`. + +Model / configuration: + +- Generated artifact records `qwen3:14b`. + +Result: + +The generated artifact contains 85 blocks, three topic-change boundaries and +four segments. The run metadata records a substantially longer elapsed time +than the later windowed artifact for the same input. + +Decision: + +Full-context segmentation was useful as a prototype, but it was superseded by +windowed segmentation and manual review tooling. + +Lessons learned: + +The prototype established the boundary-to-segment representation, but did not +settle segmentation quality. + +Evidence: + +- `src/meeting_lab/segmentation/segment_topics.py` +- `samples/chunks/chunk_01_normalized_segments.json` +- Commit `f234efc` - `Add initial topic segmentation prototype` +- Commit `889a4fe` - `Detect topic boundaries as continuous segments` + +## EXP-0005 - Windowed topic segmentation and review + +Status: Accepted + +Date or period: 2026-07-22 + +Hypothesis: + +Windowed topic segmentation can reduce runtime and make boundary evaluation +more inspectable than a single full-context call. + +Setup: + +`segment_topics_windowed.py` analyzes overlapping windows and reports only +boundaries from the decision range. Python merges boundaries into continuous, +non-overlapping segments. `review_segmentation.py` renders each boundary with +neighboring transcript context for manual classification. + +Inputs: + +- `samples/chunks/chunk_01_normalized.txt`. +- Full meeting normalized chunks under + `samples/whisper/meeting_speech_cleaned_chunks/`. + +Model / configuration: + +- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3. +- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3. + +Result: + +The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments +for 85 blocks, with recorded elapsed time lower than the full-context artifact. +Manual review output shows that some boundaries were assessed as subtopics +rather than full topic changes. Full-meeting artifacts show one window per +already-small normalized chunk and two segments per chunk. + +Decision: + +Windowed segmentation and review tooling are accepted as prototype tooling, not +as a stable production segmentation stage. + +Lessons learned: + +Windowing improves inspectability and can reduce runtime, but it can also +cluster boundaries and over-segment. Manual review remains necessary. + +Evidence: + +- `src/meeting_lab/segmentation/segment_topics_windowed.py` +- `src/meeting_lab/segmentation/review_segmentation.py` +- `samples/chunks/chunk_01_normalized_windowed_segments.json` +- `samples/chunks/chunk_01_normalized_windowed_segments_review.md` +- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json` +- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling` + +## EXP-0006 - Qwen model comparison + +Status: Accepted + +Hypothesis: + +Larger local Qwen-family models should improve meaningful extraction and +segmentation, but model size alone will not solve prompt or pipeline problems. + +Setup: + +Project work used smaller models for smoke checks and larger local models for +meaningful extraction or segmentation. Artifacts and project knowledge record +the currently useful model roles. + +Inputs: + +- Gold Standard scenarios under `tests/gold/`. +- Generated segmentation artifacts under `samples/`. +- Generated extraction artifacts under + `samples/whisper/meeting_speech_cleaned_chunks/`. + +Model / configuration: + +- `qwen3:1.7b`: smoke-test model according to project knowledge. +- `qwen3.5:9b`: current meaningful extraction and segmentation model according + to project knowledge and generated full-meeting segmentation artifacts. +- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts. + +Result: + +The repository supports the conclusion that `qwen3.5:9b` is the meaningful +current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger +models and longer contexts may increase runtime substantially, but no hardware +benchmark suite is recorded. + +Decision: + +Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current +experiments. Do not assume model size alone fixes prompt or pipeline design. + +Lessons learned: + +Evaluation must separate model capability from prompt clarity, context +strategy, extraction schema and consolidation. + +Evidence: + +- `PROJECT_KNOWLEDGE.md` +- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json` +- `samples/chunks/chunk_01_normalized_segments.json` +- `samples/chunks/chunk_01_normalized_windowed_segments.json` + +## EXP-0007 - Thinking output and Ollama API behavior + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Thinking-capable Qwen models may return reasoning separately from the final +answer, and extraction parsing should not fail merely because extra text or +multiple JSON objects appear. + +Setup: + +The Ollama response reader checks `response`, chat-style `message.content` and +then `thinking`. The JSON parser tries a full parse first, then scans JSON +object candidates and returns the final valid object. A regression test covers +thinking text before final JSON. + +Inputs: + +- Synthetic parser test in `tests/test_extraction_protocol.py`. + +Model / configuration: + +- No LLM run in the test. +- Code path is used by Ollama extraction. + +Result: + +The parser can handle additional text and multiple JSON objects where the final +valid object is the intended answer. Current code still falls back to `thinking` +only if no usable response or message content is present. + +Decision: + +Keep parser robustness, but do not treat thinking output as the root cause of +all extraction failures. + +Lessons learned: + +API response shape and model output shape are separate concerns. Preserve raw +responses when diagnosing failures. + +Evidence: + +- `src/meeting_lab/extraction/extract_chunks.py` +- `tests/test_extraction_protocol.py` +- `PROJECT_KNOWLEDGE.md` + +## EXP-0008 - JSON truncation and generation limits + +Status: Accepted + +Hypothesis: + +Some extraction failures are caused by generation limits truncating JSON rather +than by prompt wording or parser behavior. + +Setup: + +A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The +generation limit was increased for the successful path. Exact failing limit is +not recorded in the repository; the current extractor default is verifiably +`--num-predict 8192`. + +Inputs: + +- Local extraction runs referenced by project knowledge. +- Current extraction CLI. + +Model / configuration: + +- `qwen3.5:9b`. +- Current extractor default: `num_predict=8192`. + +Result: + +Increasing the generation limit fixed the technical JSON failure. This was not +primarily a parser or prompt problem. + +Decision: + +When JSON is truncated, inspect raw output and generation limits before editing +prompts. + +Lessons learned: + +Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong +first response if the model simply ran out of output tokens. + +Evidence: + +- `src/meeting_lab/extraction/extract_chunks.py` +- `PROJECT_KNOWLEDGE.md` +- `AGENTS.md` + +## EXP-0009 - Minimal end-to-end pipeline + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +A minimal local pipeline can transform Whisper output into chunk extractions +and an interim protocol, proving the technical path before the final +architecture exists. + +Setup: + +The repository added cleanup, normalization, chunking, extraction and protocol +builder scripts, with sample generated artifacts. + +Inputs: + +- `samples/whisper/meeting_speech.json` +- `samples/whisper/meeting_speech_cleaned.json` +- Generated chunks and normalized chunks under + `samples/whisper/meeting_speech_cleaned_chunks/`. + +Model / configuration: + +- Local Ollama extraction for chunk JSON. +- Windowed segmentation artifacts use `qwen3.5:9b`. + +Result: + +The repository contains cleaned input, nine chunks, nine normalized chunks, nine +extraction JSON files, windowed segmentation artifacts and +`meeting_protocol.md`. + +Decision: + +The minimal pipeline is technically validated. The first protocol builder is an +interim validation tool, not the final architecture. + +Lessons learned: + +End-to-end execution exposed the next limitation: extraction output needs +consolidation and purpose-specific rendering. + +Evidence: + +- `scripts/clean_whisper_json.py` +- `src/meeting_lab/normalization/normalize_transcript.py` +- `src/meeting_lab/chunking/chunk_transcript.py` +- `src/meeting_lab/extraction/extract_chunks.py` +- `src/meeting_lab/protocol/build_protocol.py` +- `samples/whisper/meeting_speech_cleaned_chunks/` +- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline` + +## EXP-0010 - Gold Standard corpus + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Reproducible prompt engineering requires synthetic transcripts with explicit +expected semantic outputs. + +Setup: + +The Gold Standard corpus defines scenario directories with `transcript.txt`, +`expected.json` and README files describing ground truth and common model +mistakes. The runner validates schema keys and writes `actual.json` for a +scenario. + +Inputs: + +- Gold scenarios under `tests/gold/`. + +Model / configuration: + +- Runner requires an explicit Ollama model for LLM evaluation. +- Existing unit tests for the runner do not invoke Ollama. + +Result: + +The corpus gives stable semantics for decisions, facts, positions, todos, +questions, technical details and difficult mixed cases. Initial structured +transcripts are Phase 1 and easier than raw Whisper-style transcripts. + +Decision: + +Use Gold Standard tests as both regression tests and formal meeting-semantics +specification. Raw or unlabelled transcript cases remain later-phase work. + +Lessons learned: + +Without expected outputs, prompt changes cannot be evaluated reproducibly. + +Evidence: + +- `tests/gold/` +- `scripts/run_gold_test.py` +- `tests/test_gold_runner.py` +- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests` + +## EXP-0011 - Gold-test quality and unique ground truth + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +If a prompt produces unexpected behavior, the gold test itself may be ambiguous +and should be reviewed before the prompt is changed. + +Setup: + +Decision-focused scenarios were clarified during baseline creation. The current +methodology requires checking unique ground truth before changing prompts. + +Inputs: + +- `decision_simple` +- `decision_deferred` +- `decision_none` +- Gold methodology document. + +Model / configuration: + +- Prompt Version 2 baseline work. + +Result: + +`decision_simple` required clarification around the explicit agreement and +nearby non-decision wording. The earlier negative/deferral ambiguity is now +represented by distinct `decision_none` and `decision_deferred` scenarios in +the repository. Punctuation is not reliable evidence for Whisper transcripts; +agreement language and wording must carry the semantics. + +Decision: + +Ambiguous gold tests must be reviewed before prompt changes. Do not treat +punctuation as reliable evidence in real Whisper-style transcripts. + +Lessons learned: + +Bad gold tests create false prompt failures and can encourage overfitting. + +Evidence: + +- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md` +- `tests/gold/decision_simple/README.md` +- `tests/gold/decision_deferred/README.md` +- `tests/gold/decision_none/README.md` +- Commit `f7ad9ba` + +## EXP-0012 - Decision taxonomy + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Decision extraction needs a formal taxonomy that distinguishes substantive +decisions from process decisions and non-decisions. + +Setup: + +The decision definition document and decision prompt define included and +excluded categories. Gold tests cover explicit decisions, true no-decision +cases and deferrals. + +Inputs: + +- `tests/gold/DECISION_DEFINITION.md` +- `prompts/decisions.md` +- Decision gold scenarios. + +Model / configuration: + +- Prompt Version 2 baseline. + +Result: + +Accepted decision categories include substantive decisions, organizational +decisions, process decisions, approvals, rejections, deferrals, explicit +decisions not to decide yet and explicit agreement to gather more information +before deciding. Opinions, preferences and proposals without agreement are not +decisions. + +Decision: + +"No decision was reached" and "the decision was deferred" are distinct semantic +outcomes. + +Lessons learned: + +Deferral can be a valid process decision even when the substantive topic remains +unresolved. + +Evidence: + +- `tests/gold/DECISION_DEFINITION.md` +- `tests/gold/decision_deferred/` +- `tests/gold/decision_none/` +- `prompts/decisions.md` + +## EXP-0013 - Prompt engineering methodology + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Prompt iteration needs strict experimental controls to prevent regression, +overfitting and arbitrary prompt churn. + +Setup: + +The methodology was documented alongside the Gold Standard corpus and later +summarized for agents. + +Inputs: + +- Gold scenarios. +- Prompt files. + +Model / configuration: + +- Applies to all prompt experiments. + +Result: + +The accepted method is one prompt change per iteration, one target test at a +time, immediate validation, no regressions, no `expected.json` edits merely to +force a pass, no test-specific prompt hacks, stopping after two consecutive +non-improving iterations, and verifying unique ground truth before prompt +changes. + +Decision: + +Prompt changes are controlled experiments. See `AGENTS.md` for agent operating +rules. + +Lessons learned: + +Most prompt changes are not isolated unless the experiment explicitly constrains +the target and regression set. + +Evidence: + +- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md` +- `AGENTS.md` +- Commit `f7ad9ba` + +## EXP-0014 - Decision Prompt Version 2 + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Adding explicit process-decision language to the decision prompt can preserve +true decision detection while recognizing deferrals. + +Setup: + +Prompt Version 2 added explicit support for deferrals and process decisions. +The baseline was validated on three decision scenarios. + +Inputs: + +- `decision_simple` +- `decision_deferred` +- `decision_none` + +Model / configuration: + +- Prompt Version 2. +- Model used for validation is not recorded in the committed methodology. + +Result: + +The committed methodology records all three baseline scenarios as passing. The +current generated `actual.json` files also show the expected decision count for +these decision scenarios, although some non-decision categories remain less +complete. + +Decision: + +Prompt Version 2 is the current decision baseline. Explicit deferrals are +recognized as process decisions. + +Lessons learned: + +Decision-count success does not imply all categories are solved. Category-level +evaluation must continue. + +Evidence: + +- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md` +- `tests/gold/decision_simple/actual.json` +- `tests/gold/decision_deferred/actual.json` +- `tests/gold/decision_none/actual.json` +- `prompts/decisions.md` + +## EXP-0015 - Difficult synthetic meeting + +Status: Running + +Date or period: 2026-07-30 + +Hypothesis: + +A deliberately adversarial synthetic meeting can expose extraction failures +that simple category tests miss. + +Setup: + +`evil_meeting` includes interruptions, corrections, absent referenced people, +near-decisions, changed positions and one expected explicit decision. + +Inputs: + +- `tests/gold/evil_meeting/`. + +Model / configuration: + +- Existing generated `actual.json`; exact model is not stored in the artifact. + +Result: + +The generated result found the expected FR-7 exclusion decision. It also +classified "do not migrate until the mapping table is checked" as an additional +decision. The current expected file treats that statement as a position, but +contextual review suggests it may be a valid process instruction or decision. + +Decision: + +Do not classify this as a simple model failure without reviewing the gold +standard. The scenario exposes a semantic gap in the expected output. + +Lessons learned: + +Difficult synthetic cases are valuable because they reveal ambiguity in the +specification as well as model mistakes. + +Evidence: + +- `tests/gold/evil_meeting/README.md` +- `tests/gold/evil_meeting/expected.json` +- `tests/gold/evil_meeting/actual.json` +- See EXP-0011 and EXP-0012. + +## EXP-0016 - Context-size extraction comparisons + +Status: Rejected + +Hypothesis: + +Increasing extraction context from one chunk to neighboring chunk groups should +make extraction more complete and therefore should become the baseline. + +Setup: + +Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3 +and 1+2+3. + +Inputs: + +- Normalized meeting chunks. + +Model / configuration: + +- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for + extraction experiments. + +Result: + +More context sometimes improved completeness, but it also shifted category +classification, added duplicates and reduced stability. No numeric winner is +recorded in the repository. + +Decision: + +Do not adopt larger extraction windows as the baseline. Independent chunk +extraction remains current strategy. Recover global context through +consolidation rather than continuously enlarging extraction windows. + +Lessons learned: + +Context size is not a monotonic quality knob. It changes the task the model is +performing. + +Evidence: + +- `PROJECT_KNOWLEDGE.md` +- `AGENTS.md` +- `ROADMAP.md` +- See EXP-0002 and EXP-0018. + +## EXP-0017 - Independent full-meeting chunk extraction + +Status: Accepted + +Hypothesis: + +Extracting every normalized chunk independently can produce enough structured +material for a useful protocol draft. + +Setup: + +Nine normalized chunks were extracted into separate JSON files and then +rendered by the interim protocol builder. + +Inputs: + +- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt` + through `chunk_09_normalized.txt`. + +Model / configuration: + +- Existing extraction artifacts do not record model metadata. + +Result: + +The nine extraction files contain facts, decisions, todos, questions and +technical details. `meeting_protocol.md` aggregates them into a readable draft. +Duplicates, category shifts and synthesis became the dominant limitations. + +Decision: + +Independent chunk extraction is useful enough to keep as the baseline, but it +requires a consolidation stage. + +Lessons learned: + +Per-chunk extraction gives recall-oriented raw material. It does not by itself +produce a polished or canonical meeting representation. + +Evidence: + +- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json` +- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md` +- `src/meeting_lab/protocol/build_protocol.py` +- See EXP-0018. + +## EXP-0018 - Human protocol comparison and output-view split + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +One generated protocol cannot satisfy every use case; protocol output should be +separated by purpose and audience. + +Setup: + +The interim machine protocol was compared against the desired human protocol +shape and then the architecture was revised toward parallel output views. + +Inputs: + +- Interim `meeting_protocol.md`. +- Architecture and output-view documentation. + +Model / configuration: + +- Not applicable; this is a design evaluation. + +Result: + +The human protocol target is denser and organized by purpose and topic rather +than extraction categories. The machine extraction retains more context and is +useful for recall, but it is not the right direct source for a concise +distribution artifact or durable knowledge entry. + +Decision: + +Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution +Protocol / Verteilerprotokoll and Knowledge Objects / +Wissensdatenbankeintrag. + +Lessons learned: + +Rendering is a separate concern from extraction and consolidation. Output views +must be parallel renderings of shared semantics, not transformations of one +another. + +Evidence: + +- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md` +- `docs/output-views.md` +- `docs/architecture.md` +- `docs/pipeline.md` +- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture` + +## EXP-0019 - Consolidation and Canonical Meeting Knowledge + +Status: Accepted + +Date or period: 2026-07-30 + +Hypothesis: + +Extraction, consolidation and output rendering are separate problems and should +not be collapsed into one LLM prompt or one protocol file. + +Setup: + +The architecture was refined after the minimal pipeline and protocol draft +showed duplicate, synthesis and audience-specific rendering limitations. + +Inputs: + +- Extraction JSON artifacts. +- Interim protocol draft. +- Architecture and data-model documentation. + +Model / configuration: + +- Not applicable; this is an architectural conclusion. + +Result: + +The accepted design is a planned Canonical Meeting Knowledge layer as the +semantic source of truth, with Working Protocol, Distribution Protocol and +Knowledge Objects as parallel output views. Consolidation must merge duplicates, +preserve evidence, reconcile category shifts and mark contradictions or +uncertainty. + +Decision: + +Consolidation is the next major engineering step after stable local extraction. +Canonical Meeting Knowledge and final output views are planned, not implemented. + +Lessons learned: + +Global meeting understanding should be recovered by consolidation over +evidence-bearing extractions, not by silently changing output views or expanding +LLM context indefinitely. + +Evidence: + +- `docs/architecture.md` +- `docs/pipeline.md` +- `docs/data-models.md` +- `docs/output-views.md` +- `PROJECT_KNOWLEDGE.md` +- `ROADMAP.md` +- Commit `5c03ed7`