Files
meeting-lab/docs/experiments.md
T
admin 09d125e54a Document project architecture and development methodology
- add AGENTS.md with development and prompt-engineering rules
- add PROJECT_KNOWLEDGE.md summarizing current architecture and findings
- add CHANGELOG.md
- add ROADMAP.md
- establish experiments.md as the project's experiment log
- document Canonical Meeting Knowledge architecture
- document Output Views and Knowledge Objects
- capture accepted experimental results and engineering methodology
2026-07-31 08:25:16 +02:00

26 KiB

Experiments

This file records durable technical experiments and findings for Meeting Lab. It is not a diary and does not replace commit history.

Status values

  • Proposed: experiment idea exists, but no result is recorded.
  • Running: experiment is in progress and no decision has been made.
  • Accepted: finding is the current baseline or design conclusion.
  • Rejected: hypothesis was tested and should not be repeated as-is.
  • Superseded: finding was useful but has been replaced by a newer baseline.

EXP-0001 - Whisper JSON interpretation

Status: Accepted

Date or period: 2026-07-29

Hypothesis:

Whisper JSON should be chunked from its segment stream, not from the aggregate top-level text field.

Setup:

chunk_transcript.py was updated to parse JSON input and prefer segments[*].text when segments exists. A regression test supplies JSON with both top-level text and separate segment texts.

Inputs:

  • Minimal synthetic Whisper-style JSON in tests/test_chunking.py.
  • Real Whisper artifacts under samples/whisper/.

Model / configuration:

  • No LLM.

Result:

The test verifies that the block stream is ["alpha", "beta", "gamma"] and does not include the aggregate "alpha beta gamma" text. The repository history records this as the fix for the earlier failure where the first chunk contained the complete transcript.

Decision:

When segments exists, segments[*].text is the authoritative transcript stream. The top-level text field is only a fallback.

Lessons learned:

Whisper JSON is structured input. Treating it like plain text can duplicate the entire transcript and invalidate downstream chunking.

Evidence:

  • src/meeting_lab/chunking/chunk_transcript.py
  • tests/test_chunking.py
  • Commit 4656523 - Fix Whisper JSON chunk extraction

EXP-0002 - Technical transcript chunking baseline

Status: Accepted

Date or period: 2026-07-29 to 2026-07-30

Hypothesis:

Sequential technical chunks around the configured target size can preserve the transcript while keeping extraction calls small enough for local models.

Setup:

The chunker splits block-aligned text with configurable target, minimum, maximum and overlap settings. Tests verify no duplicate later blocks when overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.

Inputs:

  • Synthetic block list in tests/test_chunking.py.
  • samples/whisper/meeting_speech_cleaned.json.

Model / configuration:

  • No LLM for chunking.
  • Manifest uses default chunking behavior recorded in samples/whisper/meeting_speech_cleaned_chunks/manifest.json.

Result:

With overlap set to zero, tests verify that all blocks appear exactly once. The real sample manifest contains nine chunks, mostly near the configured target size, with a smaller final chunk.

Decision:

Independent sequential chunks are the current technical baseline. One normalized chunk per extraction call is the preferred extraction strategy.

Lessons learned:

Chunking solves model-size constraints only. It must not perform topic detection or semantic merging.

Evidence:

  • src/meeting_lab/chunking/chunk_transcript.py
  • tests/test_chunking.py
  • samples/whisper/meeting_speech_cleaned_chunks/manifest.json
  • AGENTS.md
  • PROJECT_KNOWLEDGE.md

EXP-0003 - Conservative transcript normalization

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Transcript cleanup should improve readability without changing meeting semantics.

Setup:

The normalizer removes isolated filler sounds, immediate duplicate words or short duplicate phrases, and redundant whitespace. It records changed blocks in a JSON change log and explicitly preserves semantic content categories.

Inputs:

  • Chunk text files under samples/whisper/meeting_speech_cleaned_chunks/.
  • Change logs such as chunk_01_changes.json.

Model / configuration:

  • No LLM.

Result:

The implementation and generated change logs show a conservative policy: negations, qualifiers, dates, numbers, responsibilities, technical statements, deadlines, decisions and commitments are preserved.

Decision:

Normalization remains deterministic and low-risk. When uncertain, leave text unchanged.

Lessons learned:

Filler removal is useful only if it is tightly scoped. Broad cleanup can remove semantic cues needed by extraction.

Evidence:

  • src/meeting_lab/normalization/normalize_transcript.py
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json
  • docs/pipeline.md

EXP-0004 - Full-context topic segmentation

Status: Superseded

Date or period: 2026-07-21 to 2026-07-22

Hypothesis:

A single full-context topic segmentation call can identify topic boundaries in a normalized transcript chunk.

Setup:

The initial segmentation prototype asked the model for topic changes and then converted those boundaries into continuous, non-overlapping segments.

Inputs:

  • samples/chunks/chunk_01_normalized.txt.

Model / configuration:

  • Generated artifact records qwen3:14b.

Result:

The generated artifact contains 85 blocks, three topic-change boundaries and four segments. The run metadata records a substantially longer elapsed time than the later windowed artifact for the same input.

Decision:

Full-context segmentation was useful as a prototype, but it was superseded by windowed segmentation and manual review tooling.

Lessons learned:

The prototype established the boundary-to-segment representation, but did not settle segmentation quality.

Evidence:

  • src/meeting_lab/segmentation/segment_topics.py
  • samples/chunks/chunk_01_normalized_segments.json
  • Commit f234efc - Add initial topic segmentation prototype
  • Commit 889a4fe - Detect topic boundaries as continuous segments

EXP-0005 - Windowed topic segmentation and review

Status: Accepted

Date or period: 2026-07-22

Hypothesis:

Windowed topic segmentation can reduce runtime and make boundary evaluation more inspectable than a single full-context call.

Setup:

segment_topics_windowed.py analyzes overlapping windows and reports only boundaries from the decision range. Python merges boundaries into continuous, non-overlapping segments. review_segmentation.py renders each boundary with neighboring transcript context for manual classification.

Inputs:

  • samples/chunks/chunk_01_normalized.txt.
  • Full meeting normalized chunks under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • samples/chunks artifact: qwen3:8b, window size 20, overlap 3.
  • Full-meeting chunk artifacts: qwen3.5:9b, window size 20, overlap 3.

Result:

The samples/chunks windowed artifact produced 11 boundaries and 12 segments for 85 blocks, with recorded elapsed time lower than the full-context artifact. Manual review output shows that some boundaries were assessed as subtopics rather than full topic changes. Full-meeting artifacts show one window per already-small normalized chunk and two segments per chunk.

Decision:

Windowed segmentation and review tooling are accepted as prototype tooling, not as a stable production segmentation stage.

Lessons learned:

Windowing improves inspectability and can reduce runtime, but it can also cluster boundaries and over-segment. Manual review remains necessary.

Evidence:

  • src/meeting_lab/segmentation/segment_topics_windowed.py
  • src/meeting_lab/segmentation/review_segmentation.py
  • samples/chunks/chunk_01_normalized_windowed_segments.json
  • samples/chunks/chunk_01_normalized_windowed_segments_review.md
  • samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json
  • Commit 1a6d731 - Add windowed segmentation pipeline and review tooling

EXP-0006 - Qwen model comparison

Status: Accepted

Hypothesis:

Larger local Qwen-family models should improve meaningful extraction and segmentation, but model size alone will not solve prompt or pipeline problems.

Setup:

Project work used smaller models for smoke checks and larger local models for meaningful extraction or segmentation. Artifacts and project knowledge record the currently useful model roles.

Inputs:

  • Gold Standard scenarios under tests/gold/.
  • Generated segmentation artifacts under samples/.
  • Generated extraction artifacts under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • qwen3:1.7b: smoke-test model according to project knowledge.
  • qwen3.5:9b: current meaningful extraction and segmentation model according to project knowledge and generated full-meeting segmentation artifacts.
  • qwen3:8b and qwen3:14b: present in earlier segmentation artifacts.

Result:

The repository supports the conclusion that qwen3.5:9b is the meaningful current experiment model and qwen3:1.7b is useful for smoke tests. Larger models and longer contexts may increase runtime substantially, but no hardware benchmark suite is recorded.

Decision:

Use qwen3:1.7b for smoke tests and qwen3.5:9b for meaningful current experiments. Do not assume model size alone fixes prompt or pipeline design.

Lessons learned:

Evaluation must separate model capability from prompt clarity, context strategy, extraction schema and consolidation.

Evidence:

  • PROJECT_KNOWLEDGE.md
  • samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json
  • samples/chunks/chunk_01_normalized_segments.json
  • samples/chunks/chunk_01_normalized_windowed_segments.json

EXP-0007 - Thinking output and Ollama API behavior

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Thinking-capable Qwen models may return reasoning separately from the final answer, and extraction parsing should not fail merely because extra text or multiple JSON objects appear.

Setup:

The Ollama response reader checks response, chat-style message.content and then thinking. The JSON parser tries a full parse first, then scans JSON object candidates and returns the final valid object. A regression test covers thinking text before final JSON.

Inputs:

  • Synthetic parser test in tests/test_extraction_protocol.py.

Model / configuration:

  • No LLM run in the test.
  • Code path is used by Ollama extraction.

Result:

The parser can handle additional text and multiple JSON objects where the final valid object is the intended answer. Current code still falls back to thinking only if no usable response or message content is present.

Decision:

Keep parser robustness, but do not treat thinking output as the root cause of all extraction failures.

Lessons learned:

API response shape and model output shape are separate concerns. Preserve raw responses when diagnosing failures.

Evidence:

  • src/meeting_lab/extraction/extract_chunks.py
  • tests/test_extraction_protocol.py
  • PROJECT_KNOWLEDGE.md

EXP-0008 - JSON truncation and generation limits

Status: Accepted

Hypothesis:

Some extraction failures are caused by generation limits truncating JSON rather than by prompt wording or parser behavior.

Setup:

A qwen3.5:9b extraction failure was diagnosed as truncated JSON. The generation limit was increased for the successful path. Exact failing limit is not recorded in the repository; the current extractor default is verifiably --num-predict 8192.

Inputs:

  • Local extraction runs referenced by project knowledge.
  • Current extraction CLI.

Model / configuration:

  • qwen3.5:9b.
  • Current extractor default: num_predict=8192.

Result:

Increasing the generation limit fixed the technical JSON failure. This was not primarily a parser or prompt problem.

Decision:

When JSON is truncated, inspect raw output and generation limits before editing prompts.

Lessons learned:

Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong first response if the model simply ran out of output tokens.

Evidence:

  • src/meeting_lab/extraction/extract_chunks.py
  • PROJECT_KNOWLEDGE.md
  • AGENTS.md

EXP-0009 - Minimal end-to-end pipeline

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

A minimal local pipeline can transform Whisper output into chunk extractions and an interim protocol, proving the technical path before the final architecture exists.

Setup:

The repository added cleanup, normalization, chunking, extraction and protocol builder scripts, with sample generated artifacts.

Inputs:

  • samples/whisper/meeting_speech.json
  • samples/whisper/meeting_speech_cleaned.json
  • Generated chunks and normalized chunks under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • Local Ollama extraction for chunk JSON.
  • Windowed segmentation artifacts use qwen3.5:9b.

Result:

The repository contains cleaned input, nine chunks, nine normalized chunks, nine extraction JSON files, windowed segmentation artifacts and meeting_protocol.md.

Decision:

The minimal pipeline is technically validated. The first protocol builder is an interim validation tool, not the final architecture.

Lessons learned:

End-to-end execution exposed the next limitation: extraction output needs consolidation and purpose-specific rendering.

Evidence:

  • scripts/clean_whisper_json.py
  • src/meeting_lab/normalization/normalize_transcript.py
  • src/meeting_lab/chunking/chunk_transcript.py
  • src/meeting_lab/extraction/extract_chunks.py
  • src/meeting_lab/protocol/build_protocol.py
  • samples/whisper/meeting_speech_cleaned_chunks/
  • Commit 07b0d80 - Implement first end-to-end meeting analysis pipeline

EXP-0010 - Gold Standard corpus

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Reproducible prompt engineering requires synthetic transcripts with explicit expected semantic outputs.

Setup:

The Gold Standard corpus defines scenario directories with transcript.txt, expected.json and README files describing ground truth and common model mistakes. The runner validates schema keys and writes actual.json for a scenario.

Inputs:

  • Gold scenarios under tests/gold/.

Model / configuration:

  • Runner requires an explicit Ollama model for LLM evaluation.
  • Existing unit tests for the runner do not invoke Ollama.

Result:

The corpus gives stable semantics for decisions, facts, positions, todos, questions, technical details and difficult mixed cases. Initial structured transcripts are Phase 1 and easier than raw Whisper-style transcripts.

Decision:

Use Gold Standard tests as both regression tests and formal meeting-semantics specification. Raw or unlabelled transcript cases remain later-phase work.

Lessons learned:

Without expected outputs, prompt changes cannot be evaluated reproducibly.

Evidence:

  • tests/gold/
  • scripts/run_gold_test.py
  • tests/test_gold_runner.py
  • Commit f7ad9ba - Establish prompt engineering baseline with Gold Standard tests

EXP-0011 - Gold-test quality and unique ground truth

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

If a prompt produces unexpected behavior, the gold test itself may be ambiguous and should be reviewed before the prompt is changed.

Setup:

Decision-focused scenarios were clarified during baseline creation. The current methodology requires checking unique ground truth before changing prompts.

Inputs:

  • decision_simple
  • decision_deferred
  • decision_none
  • Gold methodology document.

Model / configuration:

  • Prompt Version 2 baseline work.

Result:

decision_simple required clarification around the explicit agreement and nearby non-decision wording. The earlier negative/deferral ambiguity is now represented by distinct decision_none and decision_deferred scenarios in the repository. Punctuation is not reliable evidence for Whisper transcripts; agreement language and wording must carry the semantics.

Decision:

Ambiguous gold tests must be reviewed before prompt changes. Do not treat punctuation as reliable evidence in real Whisper-style transcripts.

Lessons learned:

Bad gold tests create false prompt failures and can encourage overfitting.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • tests/gold/decision_simple/README.md
  • tests/gold/decision_deferred/README.md
  • tests/gold/decision_none/README.md
  • Commit f7ad9ba

EXP-0012 - Decision taxonomy

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Decision extraction needs a formal taxonomy that distinguishes substantive decisions from process decisions and non-decisions.

Setup:

The decision definition document and decision prompt define included and excluded categories. Gold tests cover explicit decisions, true no-decision cases and deferrals.

Inputs:

  • tests/gold/DECISION_DEFINITION.md
  • prompts/decisions.md
  • Decision gold scenarios.

Model / configuration:

  • Prompt Version 2 baseline.

Result:

Accepted decision categories include substantive decisions, organizational decisions, process decisions, approvals, rejections, deferrals, explicit decisions not to decide yet and explicit agreement to gather more information before deciding. Opinions, preferences and proposals without agreement are not decisions.

Decision:

"No decision was reached" and "the decision was deferred" are distinct semantic outcomes.

Lessons learned:

Deferral can be a valid process decision even when the substantive topic remains unresolved.

Evidence:

  • tests/gold/DECISION_DEFINITION.md
  • tests/gold/decision_deferred/
  • tests/gold/decision_none/
  • prompts/decisions.md

EXP-0013 - Prompt engineering methodology

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Prompt iteration needs strict experimental controls to prevent regression, overfitting and arbitrary prompt churn.

Setup:

The methodology was documented alongside the Gold Standard corpus and later summarized for agents.

Inputs:

  • Gold scenarios.
  • Prompt files.

Model / configuration:

  • Applies to all prompt experiments.

Result:

The accepted method is one prompt change per iteration, one target test at a time, immediate validation, no regressions, no expected.json edits merely to force a pass, no test-specific prompt hacks, stopping after two consecutive non-improving iterations, and verifying unique ground truth before prompt changes.

Decision:

Prompt changes are controlled experiments. See AGENTS.md for agent operating rules.

Lessons learned:

Most prompt changes are not isolated unless the experiment explicitly constrains the target and regression set.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • AGENTS.md
  • Commit f7ad9ba

EXP-0014 - Decision Prompt Version 2

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Adding explicit process-decision language to the decision prompt can preserve true decision detection while recognizing deferrals.

Setup:

Prompt Version 2 added explicit support for deferrals and process decisions. The baseline was validated on three decision scenarios.

Inputs:

  • decision_simple
  • decision_deferred
  • decision_none

Model / configuration:

  • Prompt Version 2.
  • Model used for validation is not recorded in the committed methodology.

Result:

The committed methodology records all three baseline scenarios as passing. The current generated actual.json files also show the expected decision count for these decision scenarios, although some non-decision categories remain less complete.

Decision:

Prompt Version 2 is the current decision baseline. Explicit deferrals are recognized as process decisions.

Lessons learned:

Decision-count success does not imply all categories are solved. Category-level evaluation must continue.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • tests/gold/decision_simple/actual.json
  • tests/gold/decision_deferred/actual.json
  • tests/gold/decision_none/actual.json
  • prompts/decisions.md

EXP-0015 - Difficult synthetic meeting

Status: Running

Date or period: 2026-07-30

Hypothesis:

A deliberately adversarial synthetic meeting can expose extraction failures that simple category tests miss.

Setup:

evil_meeting includes interruptions, corrections, absent referenced people, near-decisions, changed positions and one expected explicit decision.

Inputs:

  • tests/gold/evil_meeting/.

Model / configuration:

  • Existing generated actual.json; exact model is not stored in the artifact.

Result:

The generated result found the expected FR-7 exclusion decision. It also classified "do not migrate until the mapping table is checked" as an additional decision. The current expected file treats that statement as a position, but contextual review suggests it may be a valid process instruction or decision.

Decision:

Do not classify this as a simple model failure without reviewing the gold standard. The scenario exposes a semantic gap in the expected output.

Lessons learned:

Difficult synthetic cases are valuable because they reveal ambiguity in the specification as well as model mistakes.

Evidence:

  • tests/gold/evil_meeting/README.md
  • tests/gold/evil_meeting/expected.json
  • tests/gold/evil_meeting/actual.json
  • See EXP-0011 and EXP-0012.

EXP-0016 - Context-size extraction comparisons

Status: Rejected

Hypothesis:

Increasing extraction context from one chunk to neighboring chunk groups should make extraction more complete and therefore should become the baseline.

Setup:

Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3 and 1+2+3.

Inputs:

  • Normalized meeting chunks.

Model / configuration:

  • Current project knowledge identifies qwen3.5:9b as the meaningful model for extraction experiments.

Result:

More context sometimes improved completeness, but it also shifted category classification, added duplicates and reduced stability. No numeric winner is recorded in the repository.

Decision:

Do not adopt larger extraction windows as the baseline. Independent chunk extraction remains current strategy. Recover global context through consolidation rather than continuously enlarging extraction windows.

Lessons learned:

Context size is not a monotonic quality knob. It changes the task the model is performing.

Evidence:

  • PROJECT_KNOWLEDGE.md
  • AGENTS.md
  • ROADMAP.md
  • See EXP-0002 and EXP-0018.

EXP-0017 - Independent full-meeting chunk extraction

Status: Accepted

Hypothesis:

Extracting every normalized chunk independently can produce enough structured material for a useful protocol draft.

Setup:

Nine normalized chunks were extracted into separate JSON files and then rendered by the interim protocol builder.

Inputs:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt through chunk_09_normalized.txt.

Model / configuration:

  • Existing extraction artifacts do not record model metadata.

Result:

The nine extraction files contain facts, decisions, todos, questions and technical details. meeting_protocol.md aggregates them into a readable draft. Duplicates, category shifts and synthesis became the dominant limitations.

Decision:

Independent chunk extraction is useful enough to keep as the baseline, but it requires a consolidation stage.

Lessons learned:

Per-chunk extraction gives recall-oriented raw material. It does not by itself produce a polished or canonical meeting representation.

Evidence:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md
  • src/meeting_lab/protocol/build_protocol.py
  • See EXP-0018.

EXP-0018 - Human protocol comparison and output-view split

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

One generated protocol cannot satisfy every use case; protocol output should be separated by purpose and audience.

Setup:

The interim machine protocol was compared against the desired human protocol shape and then the architecture was revised toward parallel output views.

Inputs:

  • Interim meeting_protocol.md.
  • Architecture and output-view documentation.

Model / configuration:

  • Not applicable; this is a design evaluation.

Result:

The human protocol target is denser and organized by purpose and topic rather than extraction categories. The machine extraction retains more context and is useful for recall, but it is not the right direct source for a concise distribution artifact or durable knowledge entry.

Decision:

Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.

Lessons learned:

Rendering is a separate concern from extraction and consolidation. Output views must be parallel renderings of shared semantics, not transformations of one another.

Evidence:

  • samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md
  • docs/output-views.md
  • docs/architecture.md
  • docs/pipeline.md
  • Commit 5c03ed7 - Refine canonical meeting knowledge architecture

EXP-0019 - Consolidation and Canonical Meeting Knowledge

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Extraction, consolidation and output rendering are separate problems and should not be collapsed into one LLM prompt or one protocol file.

Setup:

The architecture was refined after the minimal pipeline and protocol draft showed duplicate, synthesis and audience-specific rendering limitations.

Inputs:

  • Extraction JSON artifacts.
  • Interim protocol draft.
  • Architecture and data-model documentation.

Model / configuration:

  • Not applicable; this is an architectural conclusion.

Result:

The accepted design is a planned Canonical Meeting Knowledge layer as the semantic source of truth, with Working Protocol, Distribution Protocol and Knowledge Objects as parallel output views. Consolidation must merge duplicates, preserve evidence, reconcile category shifts and mark contradictions or uncertainty.

Decision:

Consolidation is the next major engineering step after stable local extraction. Canonical Meeting Knowledge and final output views are planned, not implemented.

Lessons learned:

Global meeting understanding should be recovered by consolidation over evidence-bearing extractions, not by silently changing output views or expanding LLM context indefinitely.

Evidence:

  • docs/architecture.md
  • docs/pipeline.md
  • docs/data-models.md
  • docs/output-views.md
  • PROJECT_KNOWLEDGE.md
  • ROADMAP.md
  • Commit 5c03ed7