Files
meeting-lab/docs/experiments.md
T

39 KiB

Experiments

This file records durable technical experiments and findings for Meeting Lab. It is not a diary and does not replace commit history.

Status values

  • Proposed: experiment idea exists, but no result is recorded.
  • Running: experiment is in progress and no decision has been made.
  • Accepted: finding is the current baseline or design conclusion.
  • Rejected: hypothesis was tested and should not be repeated as-is.
  • Superseded: finding was useful but has been replaced by a newer baseline.

EXP-0001 - Whisper JSON interpretation

Status: Accepted

Date or period: 2026-07-29

Hypothesis:

Whisper JSON should be chunked from its segment stream, not from the aggregate top-level text field.

Setup:

chunk_transcript.py was updated to parse JSON input and prefer segments[*].text when segments exists. A regression test supplies JSON with both top-level text and separate segment texts.

Inputs:

  • Minimal synthetic Whisper-style JSON in tests/test_chunking.py.
  • Real Whisper artifacts under samples/whisper/.

Model / configuration:

  • No LLM.

Result:

The test verifies that the block stream is ["alpha", "beta", "gamma"] and does not include the aggregate "alpha beta gamma" text. The repository history records this as the fix for the earlier failure where the first chunk contained the complete transcript.

Decision:

When segments exists, segments[*].text is the authoritative transcript stream. The top-level text field is only a fallback.

Lessons learned:

Whisper JSON is structured input. Treating it like plain text can duplicate the entire transcript and invalidate downstream chunking.

Evidence:

  • src/meeting_lab/chunking/chunk_transcript.py
  • tests/test_chunking.py
  • Commit 4656523 - Fix Whisper JSON chunk extraction

EXP-0002 - Technical transcript chunking baseline

Status: Accepted

Date or period: 2026-07-29 to 2026-07-30

Hypothesis:

Sequential technical chunks around the configured target size can preserve the transcript while keeping extraction calls small enough for local models.

Setup:

The chunker splits block-aligned text with configurable target, minimum, maximum and overlap settings. Tests verify no duplicate later blocks when overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.

Inputs:

  • Synthetic block list in tests/test_chunking.py.
  • samples/whisper/meeting_speech_cleaned.json.

Model / configuration:

  • No LLM for chunking.
  • Manifest uses default chunking behavior recorded in samples/whisper/meeting_speech_cleaned_chunks/manifest.json.

Result:

With overlap set to zero, tests verify that all blocks appear exactly once. The real sample manifest contains nine chunks, mostly near the configured target size, with a smaller final chunk.

Decision:

Independent sequential chunks are the current technical baseline. One normalized chunk per extraction call is the preferred extraction strategy.

Lessons learned:

Chunking solves model-size constraints only. It must not perform topic detection or semantic merging.

Evidence:

  • src/meeting_lab/chunking/chunk_transcript.py
  • tests/test_chunking.py
  • samples/whisper/meeting_speech_cleaned_chunks/manifest.json
  • AGENTS.md
  • PROJECT_KNOWLEDGE.md

EXP-0003 - Conservative transcript normalization

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Transcript cleanup should improve readability without changing meeting semantics.

Setup:

The normalizer removes isolated filler sounds, immediate duplicate words or short duplicate phrases, and redundant whitespace. It records changed blocks in a JSON change log and explicitly preserves semantic content categories.

Inputs:

  • Chunk text files under samples/whisper/meeting_speech_cleaned_chunks/.
  • Change logs such as chunk_01_changes.json.

Model / configuration:

  • No LLM.

Result:

The implementation and generated change logs show a conservative policy: negations, qualifiers, dates, numbers, responsibilities, technical statements, deadlines, decisions and commitments are preserved.

Decision:

Normalization remains deterministic and low-risk. When uncertain, leave text unchanged.

Lessons learned:

Filler removal is useful only if it is tightly scoped. Broad cleanup can remove semantic cues needed by extraction.

Evidence:

  • src/meeting_lab/normalization/normalize_transcript.py
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json
  • docs/pipeline.md

EXP-0004 - Full-context topic segmentation

Status: Superseded

Date or period: 2026-07-21 to 2026-07-22

Hypothesis:

A single full-context topic segmentation call can identify topic boundaries in a normalized transcript chunk.

Setup:

The initial segmentation prototype asked the model for topic changes and then converted those boundaries into continuous, non-overlapping segments.

Inputs:

  • samples/chunks/chunk_01_normalized.txt.

Model / configuration:

  • Generated artifact records qwen3:14b.

Result:

The generated artifact contains 85 blocks, three topic-change boundaries and four segments. The run metadata records a substantially longer elapsed time than the later windowed artifact for the same input.

Decision:

Full-context segmentation was useful as a prototype, but it was superseded by windowed segmentation and manual review tooling.

Lessons learned:

The prototype established the boundary-to-segment representation, but did not settle segmentation quality.

Evidence:

  • src/meeting_lab/segmentation/segment_topics.py
  • samples/chunks/chunk_01_normalized_segments.json
  • Commit f234efc - Add initial topic segmentation prototype
  • Commit 889a4fe - Detect topic boundaries as continuous segments

EXP-0005 - Windowed topic segmentation and review

Status: Accepted

Date or period: 2026-07-22

Hypothesis:

Windowed topic segmentation can reduce runtime and make boundary evaluation more inspectable than a single full-context call.

Setup:

segment_topics_windowed.py analyzes overlapping windows and reports only boundaries from the decision range. Python merges boundaries into continuous, non-overlapping segments. review_segmentation.py renders each boundary with neighboring transcript context for manual classification.

Inputs:

  • samples/chunks/chunk_01_normalized.txt.
  • Full meeting normalized chunks under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • samples/chunks artifact: qwen3:8b, window size 20, overlap 3.
  • Full-meeting chunk artifacts: qwen3.5:9b, window size 20, overlap 3.

Result:

The samples/chunks windowed artifact produced 11 boundaries and 12 segments for 85 blocks, with recorded elapsed time lower than the full-context artifact. Manual review output shows that some boundaries were assessed as subtopics rather than full topic changes. Full-meeting artifacts show one window per already-small normalized chunk and two segments per chunk.

Decision:

Windowed segmentation and review tooling are accepted as prototype tooling, not as a stable production segmentation stage.

Lessons learned:

Windowing improves inspectability and can reduce runtime, but it can also cluster boundaries and over-segment. Manual review remains necessary.

Evidence:

  • src/meeting_lab/segmentation/segment_topics_windowed.py
  • src/meeting_lab/segmentation/review_segmentation.py
  • samples/chunks/chunk_01_normalized_windowed_segments.json
  • samples/chunks/chunk_01_normalized_windowed_segments_review.md
  • samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json
  • Commit 1a6d731 - Add windowed segmentation pipeline and review tooling

EXP-0006 - Qwen model comparison

Status: Accepted

Hypothesis:

Larger local Qwen-family models should improve meaningful extraction and segmentation, but model size alone will not solve prompt or pipeline problems.

Setup:

Project work used smaller models for smoke checks and larger local models for meaningful extraction or segmentation. Artifacts and project knowledge record the currently useful model roles.

Inputs:

  • Gold Standard scenarios under tests/gold/.
  • Generated segmentation artifacts under samples/.
  • Generated extraction artifacts under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • qwen3:1.7b: smoke-test model according to project knowledge.
  • qwen3.5:9b: current meaningful extraction and segmentation model according to project knowledge and generated full-meeting segmentation artifacts.
  • qwen3:8b and qwen3:14b: present in earlier segmentation artifacts.

Result:

The repository supports the conclusion that qwen3.5:9b is the meaningful current experiment model and qwen3:1.7b is useful for smoke tests. Larger models and longer contexts may increase runtime substantially, but no hardware benchmark suite is recorded.

Decision:

Use qwen3:1.7b for smoke tests and qwen3.5:9b for meaningful current experiments. Do not assume model size alone fixes prompt or pipeline design.

Lessons learned:

Evaluation must separate model capability from prompt clarity, context strategy, extraction schema and consolidation.

Evidence:

  • PROJECT_KNOWLEDGE.md
  • samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json
  • samples/chunks/chunk_01_normalized_segments.json
  • samples/chunks/chunk_01_normalized_windowed_segments.json

EXP-0007 - Thinking output and Ollama API behavior

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Thinking-capable Qwen models may return reasoning separately from the final answer, and extraction parsing should not fail merely because extra text or multiple JSON objects appear.

Setup:

The Ollama response reader checks response, chat-style message.content and then thinking. The JSON parser tries a full parse first, then scans JSON object candidates and returns the final valid object. A regression test covers thinking text before final JSON.

Inputs:

  • Synthetic parser test in tests/test_extraction_protocol.py.

Model / configuration:

  • No LLM run in the test.
  • Code path is used by Ollama extraction.

Result:

The parser can handle additional text and multiple JSON objects where the final valid object is the intended answer. Current code still falls back to thinking only if no usable response or message content is present.

Decision:

Keep parser robustness, but do not treat thinking output as the root cause of all extraction failures.

Lessons learned:

API response shape and model output shape are separate concerns. Preserve raw responses when diagnosing failures.

Evidence:

  • src/meeting_lab/extraction/extract_chunks.py
  • tests/test_extraction_protocol.py
  • PROJECT_KNOWLEDGE.md

EXP-0008 - JSON truncation and generation limits

Status: Accepted

Hypothesis:

Some extraction failures are caused by generation limits truncating JSON rather than by prompt wording or parser behavior.

Setup:

A qwen3.5:9b extraction failure was diagnosed as truncated JSON. The generation limit was increased for the successful path. Exact failing limit is not recorded in the repository; the current extractor default is verifiably --num-predict 8192.

Inputs:

  • Local extraction runs referenced by project knowledge.
  • Current extraction CLI.

Model / configuration:

  • qwen3.5:9b.
  • Current extractor default: num_predict=8192.

Result:

Increasing the generation limit fixed the technical JSON failure. This was not primarily a parser or prompt problem.

Decision:

When JSON is truncated, inspect raw output and generation limits before editing prompts.

Lessons learned:

Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong first response if the model simply ran out of output tokens.

Evidence:

  • src/meeting_lab/extraction/extract_chunks.py
  • PROJECT_KNOWLEDGE.md
  • AGENTS.md

EXP-0009 - Minimal end-to-end pipeline

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

A minimal local pipeline can transform Whisper output into chunk extractions and an interim protocol, proving the technical path before the final architecture exists.

Setup:

The repository added cleanup, normalization, chunking, extraction and protocol builder scripts, with sample generated artifacts.

Inputs:

  • samples/whisper/meeting_speech.json
  • samples/whisper/meeting_speech_cleaned.json
  • Generated chunks and normalized chunks under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • Local Ollama extraction for chunk JSON.
  • Windowed segmentation artifacts use qwen3.5:9b.

Result:

The repository contains cleaned input, nine chunks, nine normalized chunks, nine extraction JSON files, windowed segmentation artifacts and meeting_protocol.md.

Decision:

The minimal pipeline is technically validated. The first protocol builder is an interim validation tool, not the final architecture.

Lessons learned:

End-to-end execution exposed the next limitation: extraction output needs consolidation and purpose-specific rendering.

Evidence:

  • scripts/clean_whisper_json.py
  • src/meeting_lab/normalization/normalize_transcript.py
  • src/meeting_lab/chunking/chunk_transcript.py
  • src/meeting_lab/extraction/extract_chunks.py
  • src/meeting_lab/protocol/build_protocol.py
  • samples/whisper/meeting_speech_cleaned_chunks/
  • Commit 07b0d80 - Implement first end-to-end meeting analysis pipeline

EXP-0010 - Gold Standard corpus

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Reproducible prompt engineering requires synthetic transcripts with explicit expected semantic outputs.

Setup:

The Gold Standard corpus defines scenario directories with transcript.txt, expected.json and README files describing ground truth and common model mistakes. The runner validates schema keys and writes actual.json for a scenario.

Inputs:

  • Gold scenarios under tests/gold/.

Model / configuration:

  • Runner requires an explicit Ollama model for LLM evaluation.
  • Existing unit tests for the runner do not invoke Ollama.

Result:

The corpus gives stable semantics for decisions, facts, positions, todos, questions, technical details and difficult mixed cases. Initial structured transcripts are Phase 1 and easier than raw Whisper-style transcripts.

Decision:

Use Gold Standard tests as both regression tests and formal meeting-semantics specification. Raw or unlabelled transcript cases remain later-phase work.

Lessons learned:

Without expected outputs, prompt changes cannot be evaluated reproducibly.

Evidence:

  • tests/gold/
  • scripts/run_gold_test.py
  • tests/test_gold_runner.py
  • Commit f7ad9ba - Establish prompt engineering baseline with Gold Standard tests

EXP-0011 - Gold-test quality and unique ground truth

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

If a prompt produces unexpected behavior, the gold test itself may be ambiguous and should be reviewed before the prompt is changed.

Setup:

Decision-focused scenarios were clarified during baseline creation. The current methodology requires checking unique ground truth before changing prompts.

Inputs:

  • decision_simple
  • decision_deferred
  • decision_none
  • Gold methodology document.

Model / configuration:

  • Prompt Version 2 baseline work.

Result:

decision_simple required clarification around the explicit agreement and nearby non-decision wording. The earlier negative/deferral ambiguity is now represented by distinct decision_none and decision_deferred scenarios in the repository. Punctuation is not reliable evidence for Whisper transcripts; agreement language and wording must carry the semantics.

Decision:

Ambiguous gold tests must be reviewed before prompt changes. Do not treat punctuation as reliable evidence in real Whisper-style transcripts.

Lessons learned:

Bad gold tests create false prompt failures and can encourage overfitting.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • tests/gold/decision_simple/README.md
  • tests/gold/decision_deferred/README.md
  • tests/gold/decision_none/README.md
  • Commit f7ad9ba

EXP-0012 - Decision taxonomy

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Decision extraction needs a formal taxonomy that distinguishes substantive decisions from process decisions and non-decisions.

Setup:

The decision definition document and decision prompt define included and excluded categories. Gold tests cover explicit decisions, true no-decision cases and deferrals.

Inputs:

  • tests/gold/DECISION_DEFINITION.md
  • prompts/decisions.md
  • Decision gold scenarios.

Model / configuration:

  • Prompt Version 2 baseline.

Result:

Accepted decision categories include substantive decisions, organizational decisions, process decisions, approvals, rejections, deferrals, explicit decisions not to decide yet and explicit agreement to gather more information before deciding. Opinions, preferences and proposals without agreement are not decisions.

Decision:

"No decision was reached" and "the decision was deferred" are distinct semantic outcomes.

Lessons learned:

Deferral can be a valid process decision even when the substantive topic remains unresolved.

Evidence:

  • tests/gold/DECISION_DEFINITION.md
  • tests/gold/decision_deferred/
  • tests/gold/decision_none/
  • prompts/decisions.md

EXP-0013 - Prompt engineering methodology

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Prompt iteration needs strict experimental controls to prevent regression, overfitting and arbitrary prompt churn.

Setup:

The methodology was documented alongside the Gold Standard corpus and later summarized for agents.

Inputs:

  • Gold scenarios.
  • Prompt files.

Model / configuration:

  • Applies to all prompt experiments.

Result:

The accepted method is one prompt change per iteration, one target test at a time, immediate validation, no regressions, no expected.json edits merely to force a pass, no test-specific prompt hacks, stopping after two consecutive non-improving iterations, and verifying unique ground truth before prompt changes.

Decision:

Prompt changes are controlled experiments. See AGENTS.md for agent operating rules.

Lessons learned:

Most prompt changes are not isolated unless the experiment explicitly constrains the target and regression set.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • AGENTS.md
  • Commit f7ad9ba

EXP-0014 - Decision Prompt Version 2

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Adding explicit process-decision language to the decision prompt can preserve true decision detection while recognizing deferrals.

Setup:

Prompt Version 2 added explicit support for deferrals and process decisions. The baseline was validated on three decision scenarios.

Inputs:

  • decision_simple
  • decision_deferred
  • decision_none

Model / configuration:

  • Prompt Version 2.
  • Model used for validation is not recorded in the committed methodology.

Result:

The committed methodology records all three baseline scenarios as passing. The current generated actual.json files also show the expected decision count for these decision scenarios, although some non-decision categories remain less complete.

Decision:

Prompt Version 2 is the current decision baseline. Explicit deferrals are recognized as process decisions.

Lessons learned:

Decision-count success does not imply all categories are solved. Category-level evaluation must continue.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • tests/gold/decision_simple/actual.json
  • tests/gold/decision_deferred/actual.json
  • tests/gold/decision_none/actual.json
  • prompts/decisions.md

EXP-0015 - Difficult synthetic meeting

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

A deliberately adversarial synthetic meeting can expose extraction failures that simple category tests miss.

Setup:

evil_meeting includes interruptions, corrections, absent referenced people, near-decisions, changed positions and one expected explicit decision.

Inputs:

  • tests/gold/evil_meeting/.

Model / configuration:

  • Existing generated actual.json; exact model is not stored in the artifact.

Result:

The generated result found the expected FR-7 exclusion decision. It also classified "do not migrate until the mapping table is checked" as an additional decision. The current expected file treats that statement as a position, but contextual review suggests it may be a valid process instruction or decision.

Decision:

Do not classify this as a simple model failure without reviewing the gold standard. The scenario exposes a semantic gap in the expected output.

Lessons learned:

Difficult synthetic cases are valuable because they reveal ambiguity in the specification as well as model mistakes.

Evidence:

  • tests/gold/evil_meeting/README.md
  • tests/gold/evil_meeting/expected.json
  • tests/gold/evil_meeting/actual.json
  • See EXP-0011 and EXP-0012.

EXP-0016 - Context-size extraction comparisons

Status: Rejected

Hypothesis:

Increasing extraction context from one chunk to neighboring chunk groups should make extraction more complete and therefore should become the baseline.

Setup:

Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3 and 1+2+3.

Inputs:

  • Normalized meeting chunks.

Model / configuration:

  • Current project knowledge identifies qwen3.5:9b as the meaningful model for extraction experiments.

Result:

More context sometimes improved completeness, but it also shifted category classification, added duplicates and reduced stability. No numeric winner is recorded in the repository.

Decision:

Do not adopt larger extraction windows as the baseline. Independent chunk extraction remains current strategy. Recover global context through consolidation rather than continuously enlarging extraction windows.

Lessons learned:

Context size is not a monotonic quality knob. It changes the task the model is performing.

Evidence:

  • PROJECT_KNOWLEDGE.md
  • AGENTS.md
  • ROADMAP.md
  • See EXP-0002 and EXP-0018.

EXP-0017 - Independent full-meeting chunk extraction

Status: Accepted

Hypothesis:

Extracting every normalized chunk independently can produce enough structured material for a useful protocol draft.

Setup:

Nine normalized chunks were extracted into separate JSON files and then rendered by the interim protocol builder.

Inputs:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt through chunk_09_normalized.txt.

Model / configuration:

  • Existing extraction artifacts do not record model metadata.

Result:

The nine extraction files contain facts, decisions, todos, questions and technical details. meeting_protocol.md aggregates them into a readable draft. Duplicates, category shifts and synthesis became the dominant limitations.

Decision:

Independent chunk extraction is useful enough to keep as the baseline, but it requires a consolidation stage.

Lessons learned:

Per-chunk extraction gives recall-oriented raw material. It does not by itself produce a polished or canonical meeting representation.

Evidence:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md
  • src/meeting_lab/protocol/build_protocol.py
  • See EXP-0018.

EXP-0018 - Human protocol comparison and output-view split

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

One generated protocol cannot satisfy every use case; protocol output should be separated by purpose and audience.

Setup:

The interim machine protocol was compared against the desired human protocol shape and then the architecture was revised toward parallel output views.

Inputs:

  • Interim meeting_protocol.md.
  • Architecture and output-view documentation.

Model / configuration:

  • Not applicable; this is a design evaluation.

Result:

The human protocol target is denser and organized by purpose and topic rather than extraction categories. The machine extraction retains more context and is useful for recall, but it is not the right direct source for a concise distribution artifact or durable knowledge entry.

Decision:

Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.

Lessons learned:

Rendering is a separate concern from extraction and consolidation. Output views must be parallel renderings of shared semantics, not transformations of one another.

Evidence:

  • samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md
  • docs/output-views.md
  • docs/architecture.md
  • docs/pipeline.md
  • Commit 5c03ed7 - Refine canonical meeting knowledge architecture

EXP-0019 - Consolidation and Canonical Meeting Knowledge

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Extraction, consolidation and output rendering are separate problems and should not be collapsed into one LLM prompt or one protocol file.

Setup:

The architecture was refined after the minimal pipeline and protocol draft showed duplicate, synthesis and audience-specific rendering limitations.

Inputs:

  • Extraction JSON artifacts.
  • Interim protocol draft.
  • Architecture and data-model documentation.

Model / configuration:

  • Not applicable; this is an architectural conclusion.

Result:

The accepted design is a planned Canonical Meeting Knowledge layer as the semantic source of truth, with Working Protocol, Distribution Protocol and Knowledge Objects as parallel output views. The next consolidation architecture is split into a Deterministic Canonicalizer and a Semantic Consolidator. The canonicalizer prepares validated evidence-bearing objects without uncertain semantic merging. The consolidator then merges semantically equivalent statements, preserves evidence, reconciles category shifts where supported and marks contradictions or uncertainty.

Decision:

Deterministic canonicalization is implemented as Canonicalizer V1. The first semantic consolidation milestone is implemented as Semantic Consolidator V0 for facts-only duplicate detection. Broader semantic consolidation, Canonical Meeting Knowledge and final output views remain planned.

Lessons learned:

Global meeting understanding should be recovered by consolidation over evidence-bearing extractions, not by silently changing output views or expanding LLM context indefinitely.

Evidence:

  • docs/architecture.md
  • docs/pipeline.md
  • docs/data-models.md
  • docs/output-views.md
  • PROJECT_KNOWLEDGE.md
  • ROADMAP.md
  • Commit 5c03ed7

EXP-0020 - Working Protocol Synthesizer V0

Status: Accepted

Date or period: 2026-07-31

Hypothesis:

The current local synthesis model may be able to generate a useful detailed Working Protocol directly from the existing independent chunk extraction JSON files, before canonicalization or semantic consolidation exists.

Setup:

One synthesis prompt was constructed from exactly nine chunk extraction JSON files. The model was instructed to use only those extraction files, merge duplicates, group related information into topics, preserve useful discussion context and write a neutral technical Working Protocol.

Inputs:

  • chunk_01_extraction.json through chunk_09_extraction.json.
  • No original transcript, normalized chunks or Whisper output were used as synthesis input.

Model / configuration:

  • Model: qwen3.5:9b
  • Prompt characters: 24,979
  • Actual prompt eval tokens: 5,929
  • Output tokens: 1,486
  • Runtime: 294.204 seconds

Result:

The generated Working Protocol was readable, well structured and topic-oriented. It was still based directly on raw chunk extractions, without a separate deterministic canonicalization stage or semantic consolidation stage. The output language was English even though the source meeting material was German.

Decision:

Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline for later canonicalizer, consolidator and renderer comparisons. This selected generated artifact is intentionally versioned even though generated runtime artifacts are normally ignored.

Lessons learned:

Direct synthesis from chunk extractions can create a useful recall-oriented draft, but it does not replace Canonical Meeting Knowledge. The language mismatch also establishes a default renderer rule: protocol output should normally match the dominant source language unless an explicit output language is requested.

Evidence:

  • samples/benchmarks/working_protocol_synthesizer_v0/README.md
  • samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md
  • PROJECT_KNOWLEDGE.md
  • docs/output-views.md
  • See EXP-0017 and EXP-0019.

EXP-0021 - Canonicalizer V1

Status: Accepted

Date or period: 2026-07-31

Hypothesis:

Independent chunk extraction JSON can be converted into a stable deterministic intermediate format before any semantic LLM consolidation is attempted.

Setup:

Canonicalizer V1 discovers chunk_*_extraction.json files in stable chunk order, validates required categories, normalizes category names and basic field structure, parses existing legacy string formats where safe, trims redundant whitespace, assigns deterministic IDs, preserves original values and source references, and merges only exact duplicates when all semantic fields are identical.

Inputs:

  • Synthetic unit-test fixtures.
  • Existing nine extraction JSON files under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • No LLM.
  • CLI module: meeting_lab.consolidation.canonicalize.

Result:

Canonicalizer V1 produces schema_version, source_files, stats and items. It is deterministic preparation for the future Semantic Consolidator and is not Canonical Meeting Knowledge.

Decision:

Canonicalizer V1 is the current implemented deterministic canonicalization stage. Semantic Consolidator V0 now uses this representation for facts-only semantic duplicate detection; broader semantic consolidation and Canonical Meeting Knowledge remain planned.

Lessons learned:

Exact duplicate handling, source-reference preservation and legacy string parsing can be tested without model calls. Any uncertain semantic merge remains out of scope for this stage.

Evidence:

  • src/meeting_lab/consolidation/canonicalize.py
  • tests/test_canonicalize.py
  • docs/data-models.md
  • docs/pipeline.md

EXP-0022 - Semantic Consolidator V0 facts-only merge

Status: Accepted

Date or period: 2026-07-31

Hypothesis:

The Canonicalizer V1 output contains enough stable structure for a local LLM to identify semantically equivalent fact items without losing source coverage or changing non-fact categories.

Setup:

Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON, selected only items with category: "fact", and sent one bounded consolidation request to local Ollama. The merge rules required semantic equivalence, not topic similarity, and validation required every source fact ID to appear exactly once.

Inputs:

  • samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json
  • 33 fact items.

Model / configuration:

  • qwen3.5:9B
  • Ollama endpoint: http://127.0.0.1:11434/api/generate
  • Thinking disabled.
  • One LLM call.
  • num_ctx=32768
  • num_predict=4096

Result:

  • Runtime: 390.119 seconds on the current machine.
  • Merged fact groups: 1.
  • Source facts involved in merges: 2.
  • Singleton fact groups: 31.
  • Validation: passed.
  • No source fact was lost or duplicated.
  • Non-fact categories remained unchanged.

Accepted merge:

  • fact_0025 + fact_0031
  • Canonical statement: "Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up."

Decision:

Semantic Consolidator V0 is complete for its current narrow scope: conservative facts-only semantic duplicate detection with source evidence preserved. The selected report.md and consolidated_extractions.json benchmark artifacts should be versioned for later comparison. The raw model response remains a local diagnostic artifact and is not versioned.

Lessons learned:

Semantic duplicate consolidation is technically viable and conservative enough for continued evaluation, but broader semantic synthesis remains a separate future stage. The measured runtime is useful for this machine and run, but should not be generalized into a universal benchmark.

Evidence:

  • src/meeting_lab/consolidation/consolidate_facts.py
  • prompts/consolidate_facts.md
  • tests/test_consolidate_facts.py
  • samples/benchmarks/semantic_consolidator_v0/report.md
  • samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json
  • Local diagnostic only: samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt

EXP-0023 - Responsibility attribution integrity

Status: Running

Date or period: 2026-07-31

Hypothesis:

Protocol generation is operationally unsafe if responsibility, ownership or departmental role attribution is inferred from discussion context rather than explicit meeting evidence.

Setup:

The real-life Working Protocol Renderer V2 benchmark was inspected against the consolidated input. A false assignment connected a Marketing participant to Business Development criteria work even though the participant's contribution was critical or reluctant and did not establish acceptance of that task.

Inputs:

  • samples/benchmarks/working_protocol_renderer_v2/working_protocol.md
  • samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json
  • samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json

Model / configuration:

  • Not rerun for this finding.
  • Finding is based on existing benchmark artifacts.

Result:

The false responsibility attribution is visible in the structured input before rendering, so the issue is not merely stylistic renderer wording. The root cause may originate earlier in extraction and then be preserved by canonicalization and consolidation. Renderer guardrails are still required so output views do not strengthen ambiguous ownership.

Decision:

Responsibility attribution is now treated as a critical project-wide invariant. A person, team or department may be recorded as responsible only when the evidence explicitly assigns, accepts or confirms that responsibility. Discussion, expertise, objection, suggestion, thematic proximity, speaker adjacency, organizational assumptions and likely job roles do not establish ownership.

Lessons learned:

This class of error affects operational correctness, not only style. The pipeline needs traceable attribution evidence and future schema support for responsibility status such as explicit, accepted, proposed or unclear.

Evidence:

  • AGENTS.md
  • PROJECT_KNOWLEDGE.md
  • docs/data-models.md
  • docs/output-views.md
  • prompts/working_protocol.md
  • tests/gold/responsibility_attribution_negative/

EXP-0024 - Working Protocol V2 contract visibility

Status: Running

Date or period: 2026-08-09

Target:

BUG-011 renderer-only regression using an already validated Semantic Consolidator artifact.

Hypothesis:

The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs placing the leading prompt contract outside the model's effective evaluated context. The preserved failures all report prompt_eval_count=16386, while their outputs either echo trailing JSON or produce an unconstrained generic category summary instead of the requested Working Protocol.

Iteration 1 change:

  • Project every consolidated item to rendering-relevant semantic fields while retaining every item and its category/text/responsibility/deadline/status information.
  • Generate the exact structural contract from renderer validator constants and append it after the compact INPUT JSON.
  • Replace the independently handwritten prompt skeleton with a reference to that authoritative appended contract.
  • Enforce the existing prompt rule that emitted sections must not be empty.

This is one renderer-contract prompt iteration. It does not change extraction, canonicalization, semantic consolidation or responsibility semantics.

Validation before LLM run:

  • 15 focused renderer tests pass.
  • Tests cover contract generation, compact input projection, valid and invalid headings, missing topic sections, empty sections, wrapper cleanup, malformed Markdown and final-file write gating.

Decision:

Iteration 1 passed structural validation and wrote working_protocol.md, but the quality sanity check found that the model omitted most of the ten supplied decisions, two open questions and several action items. The structurally valid result therefore was not accepted as BUG-011 verification.

Iteration 2 change:

  • Add input-derived hidden coverage markers for every decision, action item and open question.
  • Require every priority item exactly once in its matching section.
  • Validate missing, duplicate, unknown and wrong-section markers deterministically.
  • Keep facts and technical details condensable as background.

This is the second single prompt iteration. It responds to the concrete omission failure observed in Iteration 1 without changing upstream semantics or inventing renderer content.

Iteration 2 pre-run validation:

  • 18 focused renderer tests pass, including exact required-item coverage and wrong-section rejection.

Decision:

Iteration 2 initially exhausted the fixed 4,096-token renderer output budget after emitting all decisions and most action items. Adaptive renderer budgeting resolved to 8,192 tokens for this input. The final run stopped normally after 3,709 evaluated output tokens.

The final renderer-only regression passed strict validation and wrote working_protocol.md. Exact coverage was 10/10 decisions, 36/36 renderable action items and 23/23 open questions, each once in its matching section. One structurally empty action item whose task, responsible, deadline and evidence were all null was recorded and excluded rather than fabricated. Optional background markers were accepted only when they referred to real projected input items.

Accept the compact renderer input, validator-derived trailing contract, priority-item coverage markers, empty-section validation and adaptive renderer output sizing as the BUG-011 baseline. This establishes structural reliability and priority-item coverage, not complete protocol prose quality.

Evidence:

  • prompts/working_protocol.md
  • src/meeting_lab/protocol/render_working_protocol.py
  • tests/test_render_working_protocol.py
  • samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/
  • samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/

EXP-0025 — BUG-015 classification precision

Date: 2026-08-09

Target: Progeo-derived Decision, Action Item and Open Question precision cases.

Model/configuration: qwen3.5:9B, temperature 0, num_ctx=32768.

Tests were created before prompt changes. A Decision-only evidence threshold kept the explicit Dr. Schlummer rejection and omitted an option and preference. Adding Action and Open Question definitions improved several negatives but was not stable: the model alternately promoted an unaccepted Textor suggestion or moved rejected candidates into Open Questions. Moving the standalone category prompts after the transcript made the partial Decision schema dominate and misclassified true Action Items as Decisions in two consecutive runs.

The final iteration replaced the competing standalone category prompts with a single unified classification contract after the transcript. It preserved the assigned Nina task and ownerless established CET work, and prevented cross-category leakage in the focused case, but still emitted the unaccepted Textor suggestion as an Action Item. Further prompt iterations were stopped in accordance with the Gold Standard methodology.

Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015 remains Open; no phrase-specific deterministic filter was introduced.