Files
meeting-lab/docs/experiments.md
T

70 KiB

Experiments

This file records durable technical experiments and findings for Meeting Lab. It is not a diary and does not replace commit history.

Status values

  • Proposed: experiment idea exists, but no result is recorded.
  • Running: experiment is in progress and no decision has been made.
  • Accepted: finding is the current baseline or design conclusion.
  • Rejected: hypothesis was tested and should not be repeated as-is.
  • Superseded: finding was useful but has been replaced by a newer baseline.

EXP-0001 - Whisper JSON interpretation

Status: Accepted

Date or period: 2026-07-29

Hypothesis:

Whisper JSON should be chunked from its segment stream, not from the aggregate top-level text field.

Setup:

chunk_transcript.py was updated to parse JSON input and prefer segments[*].text when segments exists. A regression test supplies JSON with both top-level text and separate segment texts.

Inputs:

  • Minimal synthetic Whisper-style JSON in tests/test_chunking.py.
  • Real Whisper artifacts under samples/whisper/.

Model / configuration:

  • No LLM.

Result:

The test verifies that the block stream is ["alpha", "beta", "gamma"] and does not include the aggregate "alpha beta gamma" text. The repository history records this as the fix for the earlier failure where the first chunk contained the complete transcript.

Decision:

When segments exists, segments[*].text is the authoritative transcript stream. The top-level text field is only a fallback.

Lessons learned:

Whisper JSON is structured input. Treating it like plain text can duplicate the entire transcript and invalidate downstream chunking.

Evidence:

  • src/meeting_lab/chunking/chunk_transcript.py
  • tests/test_chunking.py
  • Commit 4656523 - Fix Whisper JSON chunk extraction

EXP-0002 - Technical transcript chunking baseline

Status: Accepted

Date or period: 2026-07-29 to 2026-07-30

Hypothesis:

Sequential technical chunks around the configured target size can preserve the transcript while keeping extraction calls small enough for local models.

Setup:

The chunker splits block-aligned text with configurable target, minimum, maximum and overlap settings. Tests verify no duplicate later blocks when overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.

Inputs:

  • Synthetic block list in tests/test_chunking.py.
  • samples/whisper/meeting_speech_cleaned.json.

Model / configuration:

  • No LLM for chunking.
  • Manifest uses default chunking behavior recorded in samples/whisper/meeting_speech_cleaned_chunks/manifest.json.

Result:

With overlap set to zero, tests verify that all blocks appear exactly once. The real sample manifest contains nine chunks, mostly near the configured target size, with a smaller final chunk.

Decision:

Independent sequential chunks are the current technical baseline. One normalized chunk per extraction call is the preferred extraction strategy.

Lessons learned:

Chunking solves model-size constraints only. It must not perform topic detection or semantic merging.

Evidence:

  • src/meeting_lab/chunking/chunk_transcript.py
  • tests/test_chunking.py
  • samples/whisper/meeting_speech_cleaned_chunks/manifest.json
  • AGENTS.md
  • PROJECT_KNOWLEDGE.md

EXP-0003 - Conservative transcript normalization

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Transcript cleanup should improve readability without changing meeting semantics.

Setup:

The normalizer removes isolated filler sounds, immediate duplicate words or short duplicate phrases, and redundant whitespace. It records changed blocks in a JSON change log and explicitly preserves semantic content categories.

Inputs:

  • Chunk text files under samples/whisper/meeting_speech_cleaned_chunks/.
  • Change logs such as chunk_01_changes.json.

Model / configuration:

  • No LLM.

Result:

The implementation and generated change logs show a conservative policy: negations, qualifiers, dates, numbers, responsibilities, technical statements, deadlines, decisions and commitments are preserved.

Decision:

Normalization remains deterministic and low-risk. When uncertain, leave text unchanged.

Lessons learned:

Filler removal is useful only if it is tightly scoped. Broad cleanup can remove semantic cues needed by extraction.

Evidence:

  • src/meeting_lab/normalization/normalize_transcript.py
  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json
  • docs/pipeline.md

EXP-0004 - Full-context topic segmentation

Status: Superseded

Date or period: 2026-07-21 to 2026-07-22

Hypothesis:

A single full-context topic segmentation call can identify topic boundaries in a normalized transcript chunk.

Setup:

The initial segmentation prototype asked the model for topic changes and then converted those boundaries into continuous, non-overlapping segments.

Inputs:

  • samples/chunks/chunk_01_normalized.txt.

Model / configuration:

  • Generated artifact records qwen3:14b.

Result:

The generated artifact contains 85 blocks, three topic-change boundaries and four segments. The run metadata records a substantially longer elapsed time than the later windowed artifact for the same input.

Decision:

Full-context segmentation was useful as a prototype, but it was superseded by windowed segmentation and manual review tooling.

Lessons learned:

The prototype established the boundary-to-segment representation, but did not settle segmentation quality.

Evidence:

  • src/meeting_lab/segmentation/segment_topics.py
  • samples/chunks/chunk_01_normalized_segments.json
  • Commit f234efc - Add initial topic segmentation prototype
  • Commit 889a4fe - Detect topic boundaries as continuous segments

EXP-0005 - Windowed topic segmentation and review

Status: Accepted

Date or period: 2026-07-22

Hypothesis:

Windowed topic segmentation can reduce runtime and make boundary evaluation more inspectable than a single full-context call.

Setup:

segment_topics_windowed.py analyzes overlapping windows and reports only boundaries from the decision range. Python merges boundaries into continuous, non-overlapping segments. review_segmentation.py renders each boundary with neighboring transcript context for manual classification.

Inputs:

  • samples/chunks/chunk_01_normalized.txt.
  • Full meeting normalized chunks under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • samples/chunks artifact: qwen3:8b, window size 20, overlap 3.
  • Full-meeting chunk artifacts: qwen3.5:9b, window size 20, overlap 3.

Result:

The samples/chunks windowed artifact produced 11 boundaries and 12 segments for 85 blocks, with recorded elapsed time lower than the full-context artifact. Manual review output shows that some boundaries were assessed as subtopics rather than full topic changes. Full-meeting artifacts show one window per already-small normalized chunk and two segments per chunk.

Decision:

Windowed segmentation and review tooling are accepted as prototype tooling, not as a stable production segmentation stage.

Lessons learned:

Windowing improves inspectability and can reduce runtime, but it can also cluster boundaries and over-segment. Manual review remains necessary.

Evidence:

  • src/meeting_lab/segmentation/segment_topics_windowed.py
  • src/meeting_lab/segmentation/review_segmentation.py
  • samples/chunks/chunk_01_normalized_windowed_segments.json
  • samples/chunks/chunk_01_normalized_windowed_segments_review.md
  • samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json
  • Commit 1a6d731 - Add windowed segmentation pipeline and review tooling

EXP-0006 - Qwen model comparison

Status: Accepted

Hypothesis:

Larger local Qwen-family models should improve meaningful extraction and segmentation, but model size alone will not solve prompt or pipeline problems.

Setup:

Project work used smaller models for smoke checks and larger local models for meaningful extraction or segmentation. Artifacts and project knowledge record the currently useful model roles.

Inputs:

  • Gold Standard scenarios under tests/gold/.
  • Generated segmentation artifacts under samples/.
  • Generated extraction artifacts under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • qwen3:1.7b: smoke-test model according to project knowledge.
  • qwen3.5:9b: current meaningful extraction and segmentation model according to project knowledge and generated full-meeting segmentation artifacts.
  • qwen3:8b and qwen3:14b: present in earlier segmentation artifacts.

Result:

The repository supports the conclusion that qwen3.5:9b is the meaningful current experiment model and qwen3:1.7b is useful for smoke tests. Larger models and longer contexts may increase runtime substantially, but no hardware benchmark suite is recorded.

Decision:

Use qwen3:1.7b for smoke tests and qwen3.5:9b for meaningful current experiments. Do not assume model size alone fixes prompt or pipeline design.

Lessons learned:

Evaluation must separate model capability from prompt clarity, context strategy, extraction schema and consolidation.

Evidence:

  • PROJECT_KNOWLEDGE.md
  • samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json
  • samples/chunks/chunk_01_normalized_segments.json
  • samples/chunks/chunk_01_normalized_windowed_segments.json

EXP-0007 - Thinking output and Ollama API behavior

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Thinking-capable Qwen models may return reasoning separately from the final answer, and extraction parsing should not fail merely because extra text or multiple JSON objects appear.

Setup:

The Ollama response reader checks response, chat-style message.content and then thinking. The JSON parser tries a full parse first, then scans JSON object candidates and returns the final valid object. A regression test covers thinking text before final JSON.

Inputs:

  • Synthetic parser test in tests/test_extraction_protocol.py.

Model / configuration:

  • No LLM run in the test.
  • Code path is used by Ollama extraction.

Result:

The parser can handle additional text and multiple JSON objects where the final valid object is the intended answer. Current code still falls back to thinking only if no usable response or message content is present.

Decision:

Keep parser robustness, but do not treat thinking output as the root cause of all extraction failures.

Lessons learned:

API response shape and model output shape are separate concerns. Preserve raw responses when diagnosing failures.

Evidence:

  • src/meeting_lab/extraction/extract_chunks.py
  • tests/test_extraction_protocol.py
  • PROJECT_KNOWLEDGE.md

EXP-0008 - JSON truncation and generation limits

Status: Accepted

Hypothesis:

Some extraction failures are caused by generation limits truncating JSON rather than by prompt wording or parser behavior.

Setup:

A qwen3.5:9b extraction failure was diagnosed as truncated JSON. The generation limit was increased for the successful path. Exact failing limit is not recorded in the repository; the current extractor default is verifiably --num-predict 8192.

Inputs:

  • Local extraction runs referenced by project knowledge.
  • Current extraction CLI.

Model / configuration:

  • qwen3.5:9b.
  • Current extractor default: num_predict=8192.

Result:

Increasing the generation limit fixed the technical JSON failure. This was not primarily a parser or prompt problem.

Decision:

When JSON is truncated, inspect raw output and generation limits before editing prompts.

Lessons learned:

Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong first response if the model simply ran out of output tokens.

Evidence:

  • src/meeting_lab/extraction/extract_chunks.py
  • PROJECT_KNOWLEDGE.md
  • AGENTS.md

EXP-0009 - Minimal end-to-end pipeline

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

A minimal local pipeline can transform Whisper output into chunk extractions and an interim protocol, proving the technical path before the final architecture exists.

Setup:

The repository added cleanup, normalization, chunking, extraction and protocol builder scripts, with sample generated artifacts.

Inputs:

  • samples/whisper/meeting_speech.json
  • samples/whisper/meeting_speech_cleaned.json
  • Generated chunks and normalized chunks under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • Local Ollama extraction for chunk JSON.
  • Windowed segmentation artifacts use qwen3.5:9b.

Result:

The repository contains cleaned input, nine chunks, nine normalized chunks, nine extraction JSON files, windowed segmentation artifacts and meeting_protocol.md.

Decision:

The minimal pipeline is technically validated. The first protocol builder is an interim validation tool, not the final architecture.

Lessons learned:

End-to-end execution exposed the next limitation: extraction output needs consolidation and purpose-specific rendering.

Evidence:

  • scripts/clean_whisper_json.py
  • src/meeting_lab/normalization/normalize_transcript.py
  • src/meeting_lab/chunking/chunk_transcript.py
  • src/meeting_lab/extraction/extract_chunks.py
  • src/meeting_lab/protocol/build_protocol.py
  • samples/whisper/meeting_speech_cleaned_chunks/
  • Commit 07b0d80 - Implement first end-to-end meeting analysis pipeline

EXP-0010 - Gold Standard corpus

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Reproducible prompt engineering requires synthetic transcripts with explicit expected semantic outputs.

Setup:

The Gold Standard corpus defines scenario directories with transcript.txt, expected.json and README files describing ground truth and common model mistakes. The runner validates schema keys and writes actual.json for a scenario.

Inputs:

  • Gold scenarios under tests/gold/.

Model / configuration:

  • Runner requires an explicit Ollama model for LLM evaluation.
  • Existing unit tests for the runner do not invoke Ollama.

Result:

The corpus gives stable semantics for decisions, facts, positions, todos, questions, technical details and difficult mixed cases. Initial structured transcripts are Phase 1 and easier than raw Whisper-style transcripts.

Decision:

Use Gold Standard tests as both regression tests and formal meeting-semantics specification. Raw or unlabelled transcript cases remain later-phase work.

Lessons learned:

Without expected outputs, prompt changes cannot be evaluated reproducibly.

Evidence:

  • tests/gold/
  • scripts/run_gold_test.py
  • tests/test_gold_runner.py
  • Commit f7ad9ba - Establish prompt engineering baseline with Gold Standard tests

EXP-0011 - Gold-test quality and unique ground truth

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

If a prompt produces unexpected behavior, the gold test itself may be ambiguous and should be reviewed before the prompt is changed.

Setup:

Decision-focused scenarios were clarified during baseline creation. The current methodology requires checking unique ground truth before changing prompts.

Inputs:

  • decision_simple
  • decision_deferred
  • decision_none
  • Gold methodology document.

Model / configuration:

  • Prompt Version 2 baseline work.

Result:

decision_simple required clarification around the explicit agreement and nearby non-decision wording. The earlier negative/deferral ambiguity is now represented by distinct decision_none and decision_deferred scenarios in the repository. Punctuation is not reliable evidence for Whisper transcripts; agreement language and wording must carry the semantics.

Decision:

Ambiguous gold tests must be reviewed before prompt changes. Do not treat punctuation as reliable evidence in real Whisper-style transcripts.

Lessons learned:

Bad gold tests create false prompt failures and can encourage overfitting.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • tests/gold/decision_simple/README.md
  • tests/gold/decision_deferred/README.md
  • tests/gold/decision_none/README.md
  • Commit f7ad9ba

EXP-0012 - Decision taxonomy

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Decision extraction needs a formal taxonomy that distinguishes substantive decisions from process decisions and non-decisions.

Setup:

The decision definition document and decision prompt define included and excluded categories. Gold tests cover explicit decisions, true no-decision cases and deferrals.

Inputs:

  • tests/gold/DECISION_DEFINITION.md
  • prompts/decisions.md
  • Decision gold scenarios.

Model / configuration:

  • Prompt Version 2 baseline.

Result:

Accepted decision categories include substantive decisions, organizational decisions, process decisions, approvals, rejections, deferrals, explicit decisions not to decide yet and explicit agreement to gather more information before deciding. Opinions, preferences and proposals without agreement are not decisions.

Decision:

"No decision was reached" and "the decision was deferred" are distinct semantic outcomes.

Lessons learned:

Deferral can be a valid process decision even when the substantive topic remains unresolved.

Evidence:

  • tests/gold/DECISION_DEFINITION.md
  • tests/gold/decision_deferred/
  • tests/gold/decision_none/
  • prompts/decisions.md

EXP-0013 - Prompt engineering methodology

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Prompt iteration needs strict experimental controls to prevent regression, overfitting and arbitrary prompt churn.

Setup:

The methodology was documented alongside the Gold Standard corpus and later summarized for agents.

Inputs:

  • Gold scenarios.
  • Prompt files.

Model / configuration:

  • Applies to all prompt experiments.

Result:

The accepted method is one prompt change per iteration, one target test at a time, immediate validation, no regressions, no expected.json edits merely to force a pass, no test-specific prompt hacks, stopping after two consecutive non-improving iterations, and verifying unique ground truth before prompt changes.

Decision:

Prompt changes are controlled experiments. See AGENTS.md for agent operating rules.

Lessons learned:

Most prompt changes are not isolated unless the experiment explicitly constrains the target and regression set.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • AGENTS.md
  • Commit f7ad9ba

EXP-0014 - Decision Prompt Version 2

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Adding explicit process-decision language to the decision prompt can preserve true decision detection while recognizing deferrals.

Setup:

Prompt Version 2 added explicit support for deferrals and process decisions. The baseline was validated on three decision scenarios.

Inputs:

  • decision_simple
  • decision_deferred
  • decision_none

Model / configuration:

  • Prompt Version 2.
  • Model used for validation is not recorded in the committed methodology.

Result:

The committed methodology records all three baseline scenarios as passing. The current generated actual.json files also show the expected decision count for these decision scenarios, although some non-decision categories remain less complete.

Decision:

Prompt Version 2 is the current decision baseline. Explicit deferrals are recognized as process decisions.

Lessons learned:

Decision-count success does not imply all categories are solved. Category-level evaluation must continue.

Evidence:

  • tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md
  • tests/gold/decision_simple/actual.json
  • tests/gold/decision_deferred/actual.json
  • tests/gold/decision_none/actual.json
  • prompts/decisions.md

EXP-0015 - Difficult synthetic meeting

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

A deliberately adversarial synthetic meeting can expose extraction failures that simple category tests miss.

Setup:

evil_meeting includes interruptions, corrections, absent referenced people, near-decisions, changed positions and one expected explicit decision.

Inputs:

  • tests/gold/evil_meeting/.

Model / configuration:

  • Existing generated actual.json; exact model is not stored in the artifact.

Result:

The generated result found the expected FR-7 exclusion decision. It also classified "do not migrate until the mapping table is checked" as an additional decision. The current expected file treats that statement as a position, but contextual review suggests it may be a valid process instruction or decision.

Decision:

Do not classify this as a simple model failure without reviewing the gold standard. The scenario exposes a semantic gap in the expected output.

Lessons learned:

Difficult synthetic cases are valuable because they reveal ambiguity in the specification as well as model mistakes.

Evidence:

  • tests/gold/evil_meeting/README.md
  • tests/gold/evil_meeting/expected.json
  • tests/gold/evil_meeting/actual.json
  • See EXP-0011 and EXP-0012.

EXP-0016 - Context-size extraction comparisons

Status: Rejected

Hypothesis:

Increasing extraction context from one chunk to neighboring chunk groups should make extraction more complete and therefore should become the baseline.

Setup:

Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3 and 1+2+3.

Inputs:

  • Normalized meeting chunks.

Model / configuration:

  • Current project knowledge identifies qwen3.5:9b as the meaningful model for extraction experiments.

Result:

More context sometimes improved completeness, but it also shifted category classification, added duplicates and reduced stability. No numeric winner is recorded in the repository.

Decision:

Do not adopt larger extraction windows as the baseline. Independent chunk extraction remains current strategy. Recover global context through consolidation rather than continuously enlarging extraction windows.

Lessons learned:

Context size is not a monotonic quality knob. It changes the task the model is performing.

Evidence:

  • PROJECT_KNOWLEDGE.md
  • AGENTS.md
  • ROADMAP.md
  • See EXP-0002 and EXP-0018.

EXP-0017 - Independent full-meeting chunk extraction

Status: Accepted

Hypothesis:

Extracting every normalized chunk independently can produce enough structured material for a useful protocol draft.

Setup:

Nine normalized chunks were extracted into separate JSON files and then rendered by the interim protocol builder.

Inputs:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt through chunk_09_normalized.txt.

Model / configuration:

  • Existing extraction artifacts do not record model metadata.

Result:

The nine extraction files contain facts, decisions, todos, questions and technical details. meeting_protocol.md aggregates them into a readable draft. Duplicates, category shifts and synthesis became the dominant limitations.

Decision:

Independent chunk extraction is useful enough to keep as the baseline, but it requires a consolidation stage.

Lessons learned:

Per-chunk extraction gives recall-oriented raw material. It does not by itself produce a polished or canonical meeting representation.

Evidence:

  • samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json
  • samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md
  • src/meeting_lab/protocol/build_protocol.py
  • See EXP-0018.

EXP-0018 - Human protocol comparison and output-view split

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

One generated protocol cannot satisfy every use case; protocol output should be separated by purpose and audience.

Setup:

The interim machine protocol was compared against the desired human protocol shape and then the architecture was revised toward parallel output views.

Inputs:

  • Interim meeting_protocol.md.
  • Architecture and output-view documentation.

Model / configuration:

  • Not applicable; this is a design evaluation.

Result:

The human protocol target is denser and organized by purpose and topic rather than extraction categories. The machine extraction retains more context and is useful for recall, but it is not the right direct source for a concise distribution artifact or durable knowledge entry.

Decision:

Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.

Lessons learned:

Rendering is a separate concern from extraction and consolidation. Output views must be parallel renderings of shared semantics, not transformations of one another.

Evidence:

  • samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md
  • docs/output-views.md
  • docs/architecture.md
  • docs/pipeline.md
  • Commit 5c03ed7 - Refine canonical meeting knowledge architecture

EXP-0019 - Consolidation and Canonical Meeting Knowledge

Status: Accepted

Date or period: 2026-07-30

Hypothesis:

Extraction, consolidation and output rendering are separate problems and should not be collapsed into one LLM prompt or one protocol file.

Setup:

The architecture was refined after the minimal pipeline and protocol draft showed duplicate, synthesis and audience-specific rendering limitations.

Inputs:

  • Extraction JSON artifacts.
  • Interim protocol draft.
  • Architecture and data-model documentation.

Model / configuration:

  • Not applicable; this is an architectural conclusion.

Result:

The accepted design is a planned Canonical Meeting Knowledge layer as the semantic source of truth, with Working Protocol, Distribution Protocol and Knowledge Objects as parallel output views. The next consolidation architecture is split into a Deterministic Canonicalizer and a Semantic Consolidator. The canonicalizer prepares validated evidence-bearing objects without uncertain semantic merging. The consolidator then merges semantically equivalent statements, preserves evidence, reconciles category shifts where supported and marks contradictions or uncertainty.

Decision:

Deterministic canonicalization is implemented as Canonicalizer V1. The first semantic consolidation milestone is implemented as Semantic Consolidator V0 for facts-only duplicate detection. Broader semantic consolidation, Canonical Meeting Knowledge and final output views remain planned.

Lessons learned:

Global meeting understanding should be recovered by consolidation over evidence-bearing extractions, not by silently changing output views or expanding LLM context indefinitely.

Evidence:

  • docs/architecture.md
  • docs/pipeline.md
  • docs/data-models.md
  • docs/output-views.md
  • PROJECT_KNOWLEDGE.md
  • ROADMAP.md
  • Commit 5c03ed7

EXP-0020 - Working Protocol Synthesizer V0

Status: Accepted

Date or period: 2026-07-31

Hypothesis:

The current local synthesis model may be able to generate a useful detailed Working Protocol directly from the existing independent chunk extraction JSON files, before canonicalization or semantic consolidation exists.

Setup:

One synthesis prompt was constructed from exactly nine chunk extraction JSON files. The model was instructed to use only those extraction files, merge duplicates, group related information into topics, preserve useful discussion context and write a neutral technical Working Protocol.

Inputs:

  • chunk_01_extraction.json through chunk_09_extraction.json.
  • No original transcript, normalized chunks or Whisper output were used as synthesis input.

Model / configuration:

  • Model: qwen3.5:9b
  • Prompt characters: 24,979
  • Actual prompt eval tokens: 5,929
  • Output tokens: 1,486
  • Runtime: 294.204 seconds

Result:

The generated Working Protocol was readable, well structured and topic-oriented. It was still based directly on raw chunk extractions, without a separate deterministic canonicalization stage or semantic consolidation stage. The output language was English even though the source meeting material was German.

Decision:

Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline for later canonicalizer, consolidator and renderer comparisons. This selected generated artifact is intentionally versioned even though generated runtime artifacts are normally ignored.

Lessons learned:

Direct synthesis from chunk extractions can create a useful recall-oriented draft, but it does not replace Canonical Meeting Knowledge. The language mismatch also establishes a default renderer rule: protocol output should normally match the dominant source language unless an explicit output language is requested.

Evidence:

  • samples/benchmarks/working_protocol_synthesizer_v0/README.md
  • samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md
  • PROJECT_KNOWLEDGE.md
  • docs/output-views.md
  • See EXP-0017 and EXP-0019.

EXP-0021 - Canonicalizer V1

Status: Accepted

Date or period: 2026-07-31

Hypothesis:

Independent chunk extraction JSON can be converted into a stable deterministic intermediate format before any semantic LLM consolidation is attempted.

Setup:

Canonicalizer V1 discovers chunk_*_extraction.json files in stable chunk order, validates required categories, normalizes category names and basic field structure, parses existing legacy string formats where safe, trims redundant whitespace, assigns deterministic IDs, preserves original values and source references, and merges only exact duplicates when all semantic fields are identical.

Inputs:

  • Synthetic unit-test fixtures.
  • Existing nine extraction JSON files under samples/whisper/meeting_speech_cleaned_chunks/.

Model / configuration:

  • No LLM.
  • CLI module: meeting_lab.consolidation.canonicalize.

Result:

Canonicalizer V1 produces schema_version, source_files, stats and items. It is deterministic preparation for the future Semantic Consolidator and is not Canonical Meeting Knowledge.

Decision:

Canonicalizer V1 is the current implemented deterministic canonicalization stage. Semantic Consolidator V0 now uses this representation for facts-only semantic duplicate detection; broader semantic consolidation and Canonical Meeting Knowledge remain planned.

Lessons learned:

Exact duplicate handling, source-reference preservation and legacy string parsing can be tested without model calls. Any uncertain semantic merge remains out of scope for this stage.

Evidence:

  • src/meeting_lab/consolidation/canonicalize.py
  • tests/test_canonicalize.py
  • docs/data-models.md
  • docs/pipeline.md

EXP-0022 - Semantic Consolidator V0 facts-only merge

Status: Accepted

Date or period: 2026-07-31

Hypothesis:

The Canonicalizer V1 output contains enough stable structure for a local LLM to identify semantically equivalent fact items without losing source coverage or changing non-fact categories.

Setup:

Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON, selected only items with category: "fact", and sent one bounded consolidation request to local Ollama. The merge rules required semantic equivalence, not topic similarity, and validation required every source fact ID to appear exactly once.

Inputs:

  • samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json
  • 33 fact items.

Model / configuration:

  • qwen3.5:9B
  • Ollama endpoint: http://127.0.0.1:11434/api/generate
  • Thinking disabled.
  • One LLM call.
  • num_ctx=32768
  • num_predict=4096

Result:

  • Runtime: 390.119 seconds on the current machine.
  • Merged fact groups: 1.
  • Source facts involved in merges: 2.
  • Singleton fact groups: 31.
  • Validation: passed.
  • No source fact was lost or duplicated.
  • Non-fact categories remained unchanged.

Accepted merge:

  • fact_0025 + fact_0031
  • Canonical statement: "Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up."

Decision:

Semantic Consolidator V0 is complete for its current narrow scope: conservative facts-only semantic duplicate detection with source evidence preserved. The selected report.md and consolidated_extractions.json benchmark artifacts should be versioned for later comparison. The raw model response remains a local diagnostic artifact and is not versioned.

Lessons learned:

Semantic duplicate consolidation is technically viable and conservative enough for continued evaluation, but broader semantic synthesis remains a separate future stage. The measured runtime is useful for this machine and run, but should not be generalized into a universal benchmark.

Evidence:

  • src/meeting_lab/consolidation/consolidate_facts.py
  • prompts/consolidate_facts.md
  • tests/test_consolidate_facts.py
  • samples/benchmarks/semantic_consolidator_v0/report.md
  • samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json
  • Local diagnostic only: samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt

EXP-0023 - Responsibility attribution integrity

Status: Running

Date or period: 2026-07-31

Hypothesis:

Protocol generation is operationally unsafe if responsibility, ownership or departmental role attribution is inferred from discussion context rather than explicit meeting evidence.

Setup:

The real-life Working Protocol Renderer V2 benchmark was inspected against the consolidated input. A false assignment connected a Marketing participant to Business Development criteria work even though the participant's contribution was critical or reluctant and did not establish acceptance of that task.

Inputs:

  • samples/benchmarks/working_protocol_renderer_v2/working_protocol.md
  • samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json
  • samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json

Model / configuration:

  • Not rerun for this finding.
  • Finding is based on existing benchmark artifacts.

Result:

The false responsibility attribution is visible in the structured input before rendering, so the issue is not merely stylistic renderer wording. The root cause may originate earlier in extraction and then be preserved by canonicalization and consolidation. Renderer guardrails are still required so output views do not strengthen ambiguous ownership.

Decision:

Responsibility attribution is now treated as a critical project-wide invariant. A person, team or department may be recorded as responsible only when the evidence explicitly assigns, accepts or confirms that responsibility. Discussion, expertise, objection, suggestion, thematic proximity, speaker adjacency, organizational assumptions and likely job roles do not establish ownership.

Lessons learned:

This class of error affects operational correctness, not only style. The pipeline needs traceable attribution evidence and future schema support for responsibility status such as explicit, accepted, proposed or unclear.

Evidence:

  • AGENTS.md
  • PROJECT_KNOWLEDGE.md
  • docs/data-models.md
  • docs/output-views.md
  • prompts/working_protocol.md
  • tests/gold/responsibility_attribution_negative/

EXP-0024 - Working Protocol V2 contract visibility

Status: Running

Date or period: 2026-08-09

Target:

BUG-011 renderer-only regression using an already validated Semantic Consolidator artifact.

Hypothesis:

The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs placing the leading prompt contract outside the model's effective evaluated context. The preserved failures all report prompt_eval_count=16386, while their outputs either echo trailing JSON or produce an unconstrained generic category summary instead of the requested Working Protocol.

Iteration 1 change:

  • Project every consolidated item to rendering-relevant semantic fields while retaining every item and its category/text/responsibility/deadline/status information.
  • Generate the exact structural contract from renderer validator constants and append it after the compact INPUT JSON.
  • Replace the independently handwritten prompt skeleton with a reference to that authoritative appended contract.
  • Enforce the existing prompt rule that emitted sections must not be empty.

This is one renderer-contract prompt iteration. It does not change extraction, canonicalization, semantic consolidation or responsibility semantics.

Validation before LLM run:

  • 15 focused renderer tests pass.
  • Tests cover contract generation, compact input projection, valid and invalid headings, missing topic sections, empty sections, wrapper cleanup, malformed Markdown and final-file write gating.

Decision:

Iteration 1 passed structural validation and wrote working_protocol.md, but the quality sanity check found that the model omitted most of the ten supplied decisions, two open questions and several action items. The structurally valid result therefore was not accepted as BUG-011 verification.

Iteration 2 change:

  • Add input-derived hidden coverage markers for every decision, action item and open question.
  • Require every priority item exactly once in its matching section.
  • Validate missing, duplicate, unknown and wrong-section markers deterministically.
  • Keep facts and technical details condensable as background.

This is the second single prompt iteration. It responds to the concrete omission failure observed in Iteration 1 without changing upstream semantics or inventing renderer content.

Iteration 2 pre-run validation:

  • 18 focused renderer tests pass, including exact required-item coverage and wrong-section rejection.

Decision:

Iteration 2 initially exhausted the fixed 4,096-token renderer output budget after emitting all decisions and most action items. Adaptive renderer budgeting resolved to 8,192 tokens for this input. The final run stopped normally after 3,709 evaluated output tokens.

The final renderer-only regression passed strict validation and wrote working_protocol.md. Exact coverage was 10/10 decisions, 36/36 renderable action items and 23/23 open questions, each once in its matching section. One structurally empty action item whose task, responsible, deadline and evidence were all null was recorded and excluded rather than fabricated. Optional background markers were accepted only when they referred to real projected input items.

Accept the compact renderer input, validator-derived trailing contract, priority-item coverage markers, empty-section validation and adaptive renderer output sizing as the BUG-011 baseline. This establishes structural reliability and priority-item coverage, not complete protocol prose quality.

Evidence:

  • prompts/working_protocol.md
  • src/meeting_lab/protocol/render_working_protocol.py
  • tests/test_render_working_protocol.py
  • samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/
  • samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/

EXP-0025 — BUG-015 classification precision

Date: 2026-08-09

Target: Progeo-derived Decision, Action Item and Open Question precision cases.

Model/configuration: qwen3.5:9B, temperature 0, num_ctx=32768.

Tests were created before prompt changes. A Decision-only evidence threshold kept the explicit Dr. Schlummer rejection and omitted an option and preference. Adding Action and Open Question definitions improved several negatives but was not stable: the model alternately promoted an unaccepted Textor suggestion or moved rejected candidates into Open Questions. Moving the standalone category prompts after the transcript made the partial Decision schema dominate and misclassified true Action Items as Decisions in two consecutive runs.

The final iteration replaced the competing standalone category prompts with a single unified classification contract after the transcript. It preserved the assigned Nina task and ownerless established CET work, and prevented cross-category leakage in the focused case, but still emitted the unaccepted Textor suggestion as an Action Item. Further prompt iterations were stopped in accordance with the Gold Standard methodology.

Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015 remains Open; no phrase-specific deterministic filter was introduced.

EXP-0027 — Evidence-near observation extraction

Date: 2026-08-18

Hypothesis: qwen3.5:9B can more reliably extract evidence-near linguistic and semantic properties than directly synthesize protocol-level events, outcomes, actions and unresolved issues. This isolated experiment stops before semantic interpretation and does not connect to the production pipeline.

The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion Subjects from Semantic Synthesis Isolation. It defines atomic observations with source evidence, explicit targets, a five-value relation vocabulary, modality, temporality, evaluation, agreement, responsibility/person, uncertainty, clarification need and free-text scope. It contains no protocol-level category field. The validator requires sequential observation IDs, known evidence IDs, backward-only valid observation targets, closed categorical vocabularies, consistent responsibility/person pairs, and one canonical absence form: JSON null for person and absent for scope.

Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384, num_predict=4096, no retries. All nine cases ran exactly once, for nine LLM calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw responses, parsed observations, validation results, Ollama metadata and comparisons are preserved under /tmp/meeting-lab-evidence-observations-v1-20260818/.

Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL. Five cases failed structure because the model represented a single target as a one-element list, usually ["discussion_subject"]; the accepted schema permits a list only for two or more jointly referenced observations. Several responses also copied the relation label limits_scope into the free-text scope field. These were systematic model-output errors, not transport or parser failures. The prompt and run were not retried or tuned.

Human semantic review of the preserved raw responses:

Case Verdict Main result
A PARTIAL Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty.
B PARTIAL Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation.
C FAIL The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted.
D PARTIAL Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement.
E PARTIAL Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact.
F FAIL Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech.
G FAIL Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty.
H PARTIAL Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding.
I PARTIAL Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain.

Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override strict structural failures; it separates useful semantic signal from schema compliance.

Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving closer to evidence reduced some direct promotion behavior: the geometry mention did not become work, the washing disadvantage did not become an unresolved issue, both alternatives in B remained uncommitted, and the publication query did not become an assignment. However, the important promotion errors did not disappear. C acquired unsupported acceptance, and G still promoted a personal preference into rejection. Positive cases were only partly preserved: explicit rejection was recognized but incorrectly linked; trial-only language was kept but responsibility was invented; Nina's request and commitment were recognized but responsibility states were wrong; and the publication issue was recognized but its uncertainty was contradicted by rejection.

Result: B — evidence-near extraction is promising, but specific observation dimensions remain unreliable. Target/relation selection, scope attachment, responsibility state/person attribution, and agreement versus uncertainty are not reliable enough to justify designing the later interpretation stage yet. No production integration or later interpretation stage was implemented.

EXP-0028 — Evidence-Near Observation Extraction V2

Date: 2026-08-19

V2 tested whether qwen3.5:9B preserves the evidence needed by a later controlled interpretation stage when direct responsibility, agreement and semantic graph relations are removed. Responsibility was replaced by explicit participant/discourse facts (speaker, named_person, addressee, singular self-reference, collective we, and impersonal person reference). Agreement was replaced by explicit affirmation, explicit negation and determination statement signals. Graph relations were reduced to nullable scalar refers_to; scope became free-text qualifier plus nullable scalar limits_target. No later derivation stage was implemented.

The V2 Gold fixture preserves the unchanged A-I source evidence and intended human interpretations. It contains no responsibility, agreement, action, decision, open-question, accepted-trial, rejected-alternative or protocol eligibility fields. Validation enforces known evidence IDs, sequential unique observation IDs, backward-only scalar references, closed vocabularies, boolean participant flags, JSON-nullable participant/qualifier/reference fields and no string "null".

Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384, num_predict=4096, no retries or voting. A launch-path defect was corrected before the live run; the failed launch made zero model calls. A sandbox-blocked localhost attempt also made zero model calls. The completed run called the model exactly once for each of A-I: nine calls total, in 125.645 seconds. Persistent prompts, Gold input and expectations, raw and parsed model output, validation, automatic comparison, Ollama metadata and human evaluation are in artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/.

Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases were schema-invalid. The dominant serialization pattern was use of present instead of the specified explicit for affirmation/negation; E additionally used none instead of absent for a determination signal, while D emitted the separate uncertainty concept as an invalid modality. These errors are contract violations, although most present/explicit differences are deterministically normalizable without changing meaning. A and C were valid JSON/schema outputs but had critical semantic mismatches.

Human semantic review:

Case Verdict Main result
A PARTIAL Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification.
B PARTIAL Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; present caused schema failure.
C FAIL Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement.
D PARTIAL Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality.
E PARTIAL Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection.
F PARTIAL Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments.
G FAIL Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity.
H PARTIAL Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier.
I PARTIAL Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity.

Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced V1 promotion errors: collective speech and speaker identity no longer became individual responsibility; personal preference no longer became a group-level rejection field; an information question did not become work; and explicit negation/affirmation survived as separate evidence. Useful participant evidence also survived strongly in H and collective-speech evidence in F.

Simplification did not make all evidence-near dimensions reliable. Scalar references and limits_target were almost entirely omitted, qualifiers were usually omitted, committed modality was missed in E, F and H, and C/G repeated important modality, uncertainty and participant-feature errors. Some positive semantic information therefore survived only in free-text content, not in the structural signals a controlled derivation stage would need.

Result: B — V2 is materially better, but specific evidence-near dimensions still require refinement. Direct responsibility, agreement and graph-relation classification should remain excluded. Before designing the derivation stage, the next work should examine the minimal reliable representation of explicit reference/scope limitation, commitment modality and participant deixis. No production integration, Progeo run or derivation implementation was performed.

EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation

Date: 2026-08-19

Hypothesis: qwen3.5:9B is substantially more reliable when the first semantic stage preserves meeting meaning as atomic natural-language observations with provenance and only simple participant information, without classifying or deriving higher-level meeting semantics.

V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each observation contains exactly observation_id, evidence_id, content, speaker, nullable named_person, and nullable addressee. It contains no modality, temporality, evaluation, affirmation, negation, determination, uncertainty, clarification, responsibility, agreement, relation, reference, qualifier, scope, limit, protocol-category or protocol-eligibility fields. Instead, the prompt asks for conservative atomic content that retains hedges, conditions, personal/collective/impersonal language, requests, acceptances, rejections, quantities, deadlines and boundaries in natural language.

Structural validation is intentionally small: exact schema keys, non-empty observations/content, unique obs_N identifiers, known evidence IDs, speaker matching its evidence, explicit named people/addressees, and no string "null". Human semantic preservation against per-case requirements is the primary evaluation; wording differences do not fail a case.

Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384, num_predict=4096, no retries, voting or per-case tuning. One sandbox-blocked localhost launch made zero model calls. The completed run made exactly nine calls, one for each A-I case, in 40.074 seconds. All nine outputs passed structural validation. Persistent source evidence, semantic requirements, exact prompts, raw and parsed responses, validation, Ollama metadata and human evaluation are stored under artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/.

Human semantic preservation results:

Case Verdict Main result
A PASS Preserved kann, vielleicht, tentative follow-up, and explicit Dann sequence without commitment.
B PASS Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection.
C PASS Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work.
D PASS Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task.
E PASS Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided.
F PASS Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership.
G PARTIAL Preserved hypothetical risk, Martin's personal stance, wenn überhaupt, impersonal checking need and no decision, but dropped collective wir from who would receive contaminated material.
H PASS Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field.
I PASS Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment.

Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed “that we receive contaminated material back” into an impersonal passive phrase; the risk itself remained hypothetical. H translated Freitag to Friday, a harmless wording difference. I retained two compound observations rather than splitting every proposition, but all required semantic boundaries and dependencies remained explicit.

Compared with V2, categorical-field removal improved content preservation in A, G and I: A retained Dann; G retained wenn überhaupt, personal Ich and impersonal Man; I retained publication purpose and all boundaries. It also reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic strengthening into responsibility, group rejection, established work or assigned clarification. F and H remain sufficiently complete in natural language for a later interpretation experiment. No useful meaning was shown to depend on the removed fields; the V3 content retained the useful signals that V2's fields had attempted to encode.

Result: A — MINIMAL FIRST STAGE ACCEPTED. On A-I, minimal atomic content with evidence provenance and simple participants is sufficiently reliable to be the candidate first semantic stage. A later bounded experiment may examine controlled semantic interpretation, but no derivation stage, production integration or Progeo run was implemented here.

EXP-0030 — V3 model comparison: qwen3.5:9B vs qwen3.6:35B-A3B

Date: 2026-08-19

This controlled comparison reran the accepted V3 minimal semantic-preservation experiment unchanged with the locally installed qwen3.6:35B-A3B. It used the same implementation, A-I fixture, evidence, prompt, minimal schema, temperature 0, think=false, num_ctx=16384, num_predict=4096, no retries, no voting and one call per case. The completed run made exactly nine calls in 96.392 seconds. Artifacts are preserved under artifacts/experiments/evidence_observations_v3/20260819_v3_qwen36_35b_a3b_single_run/.

Structural validation passed 8/9 cases. D was semantically faithful but invalid because the model copied transcript speakers Antonius and Martin into named_person, although those names were not explicitly named within their utterances. Human semantic preservation produced 7 PASS, 1 PARTIAL and 1 FAIL:

Case Verdict Main result
A PASS Preserved can, perhaps, tentative would, explicit then and no commitment, but changed the content language to English.
B PARTIAL Preserved both approaches overall, but removed Oder from the 20-20 observation and locally strengthened it into collective planned conduct.
C PASS Preserved Tim's tentative personal Textor contact and possible follow-up without established work.
D PASS Preserved possibility, uncertainty, process and energy consequence; structural failure was confined to invalid speaker-as-named-person values.
E PASS Preserved cost, collective rejection/non-pursuit and later determination, though the confirmation dropped explicit Ja.
F FAIL Preserved quantity, timing and trial/series boundaries, but changed collective wir into “Martin suggests” and “Tim agrees,” inventing individual proposal/agreement meaning.
G PASS Preserved hypothetical risk, collective recipient wir, Martin's personal stance, wenn überhaupt, conditional Technikum, impersonal man müsste and no decision/owner.
H PASS Preserved request, addressee, Friday, explicit acceptance and future personal commitment without a responsibility field; content was English.
I PASS Preserved local/pure-production and comparison boundaries, five-degree difference, upstream effort, publication purpose, unresolved permission and clarification without assignment.

Direct comparison:

Measure qwen3.5:9B qwen3.6:35B-A3B
Structurally valid 9/9 8/9
Human PASS 8 7
Human PARTIAL 1 1
Human FAIL 0 1
Observations 28 30
Runtime 40.074 s 96.392 s
LLM calls 9 9
Prompt-evaluation tokens 6,268 6,268
Evaluation tokens 2,590 2,718

The larger model fixed the 9B weakness in G by preserving collective wir, and it split I's compound production/effort observations more cleanly. Those gains did not offset regressions: B was locally strengthened, F materially converted collective conduct into individual agreement, D violated the participant schema, observation count increased, and runtime was 2.4 times higher. Both models preserved German consistently in six of nine cases, but in different cases; the 35B-A3B model changed A, F and H to English, while 9B changed C, F and H wholly or partly to English.

Result: D — REGRESSION. qwen3.6:35B-A3B does not materially improve the accepted minimal V3 first-stage preservation over qwen3.5:9B; it is worse on the A-I comparison because of the F ownership-adjacent strengthening and lower structural validity. This conclusion applies only to the minimal V3 first stage and does not determine model choice for any later semantic derivation. No production integration, derivation implementation or Progeo run occurred.

EXP-0031 — Controlled Semantic Derivation H V0

Date: 2026-08-19

This isolated experiment tested the first controlled second-stage derivation using only the accepted qwen3.5:9B V3 observations for case H. The derivation LLM received the two V3 observations, not the transcript or Gold expectation. Its deliberately narrow task was limited to recognizing whether obs_1 is a concrete request and whether obs_2 explicitly commits its speaker to substantially the same work. Its strict output schema forbids responsibility, requested actor, establishment/status, Action Item, protocol, confidence and generic relation/graph fields.

Deterministic code validates observation/evidence provenance, obtains the requested actor only from the request observation's addressee, requires the acceptance to follow the request, requires the accepting speaker to equal that addressee, and establishes responsibility only after all semantic and structural gates pass. A bounded weekday normalizer reconciles Friday and Freitag, rejects conflicting weekdays, and separates the supported due date from the normalized action text. No general temporal or action ontology was introduced.

Twenty focused deterministic tests cover the positive H path and the required negative invariants: request alone, acknowledgement/non-commitment, tentative acceptance, different response speaker, different work, reversed order, speaker/name/addressee alone, conflicting deadlines, unknown observation IDs, inconsistent evidence provenance, forbidden semantic fields, malformed JSON and persistent artifacts. The complete non-LLM suite passed 192/192.

Configuration: one qwen3.5:9B call, temperature 0, think=false, num_ctx=16384, num_predict=1024, no retries or voting. The call took 11.765 seconds, with 469 prompt-evaluation and 124 evaluation tokens. The model returned a valid recognition object: obs_1 is a concrete request, obs_2 is an explicit commitment, and both concern substantially the same work. It returned no responsibility or establishment judgment.

All deterministic gates passed. The final derived result is an established action Prüfung der Messdaten, requested from and assigned to Nina, due Freitag, supported by request obs_1/e1 and acceptance obs_2/e2. The model included bis Friday in its normalized request text; after the single call, a deterministic-only bounded correction separated that already-recognized due phrase from action content without changing the prompt, recognition schema, semantic result or call count. Focused and complete non-LLM suites still passed after this correction.

Artifacts are preserved under artifacts/experiments/controlled_semantic_derivation_h/20260819_h_qwen35_9b_single_run/. Result: the H mechanism succeeded. This establishes only that the narrow request-plus-explicit-acceptance pattern can be recognized and gated for H; it does not generalize the derivation architecture to other cases or semantic categories. No production integration, other case run, semantic graph, protocol derivation or Progeo run occurred.

EXP-0032 — Request / Acceptance Gold V0

Status: Experimental; promising with semantic precision gaps

Date: 2026-08-20

This isolated regression experiment tested whether the EXP-0031 mechanism generalizes beyond H. It used ten short synthetic cases containing only V3-style observations. Evidence Observation V3 was neither called nor changed, and the model received no raw transcript or expected result. The fixed recognition schema permits only a nullable concrete request and nullable later explicit personal commitment, plus the same-requested-work judgment and normalized action text. Responsibility, requested actor, established status, Action Item, protocol, confidence and generic graph fields remain forbidden.

Cases:

  • RA-01 explicit positive acceptance: PASS.
  • RA-02 paraphrased positive acceptance: PASS.
  • RA-03 acknowledgement only: PASS.
  • RA-04 tentative response: PASS.
  • RA-05 different responder without personal acceptance: PASS.
  • RA-06 explicit commitment to different work: PARTIAL. The model returned no acceptance instead of recognizing a commitment with same_requested_work false. The requested action correctly remained unestablished.
  • RA-07 request without response: PASS.
  • RA-08 collective commitment: PARTIAL. The model over-recognized the collective wir statement as an explicit commitment, but no request existed and deterministic gates prevented individual responsibility.
  • RA-09 impersonal necessity: PARTIAL. The model over-recognized the impersonal necessity as a concrete request, but the observation had no addressee and deterministic gates prevented establishment.
  • RA-10 tentative personal suggestion: PASS.

Configuration: exactly ten sequential qwen3.5:9B calls, one per case, temperature 0, think=false, num_ctx=16384, num_predict=1024, no retries, no voting and no prompt change between cases. Summed call time was 23.754 seconds, with 4,826 prompt-evaluation tokens and 793 evaluation tokens. The strict schema validated every response and no responsibility or establishment field leaked into model output.

Both positive cases recognized the request, explicit commitment and same-work relationship, including the paraphrased acceptance, and deterministically established Clara as responsible with due date Dienstag. The model rendered the normalized action in semantically equivalent English; evaluation therefore checks the structural deterministic result exactly while treating normalized action wording as evidence-near semantic text rather than requiring lexical identity. Acknowledgement and tentative response were not promoted. Every negative case remained unestablished, and no individual responsibility was invented.

Recognition-level errors were two false positives (RA-08 commitment and RA-09 request) and one false negative (RA-06 different-work commitment). Final established-action false positives and false negatives were both zero. The overall result was seven PASS, three PARTIAL and zero FAIL.

Conclusion: the narrow request-plus-acceptance architecture remains promising for established individual actions because deterministic addressee, ordering, speaker, same-work, provenance and deadline gates contained all recognition errors. The recognition layer is not yet precise enough to generalize: its handling of collective commitment, impersonal necessity and commitments to different work needs further isolated study. No production integration or additional semantic category is justified by this result.

Artifacts are preserved under artifacts/experiments/request_acceptance_gold_v0/20260820_qwen35_9b_single_run/.

EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype

Date: 2026-08-11

Hypothesis: the primary protocol should be a topic-oriented reconstruction of the meeting rather than a category-oriented list of extracted information.

This first isolated prototype does not replace or connect to the production pipeline or Working Protocol renderer. It sends small evidence-ID-tagged transcript excerpts to qwen3.5:9B and requests Discussion Subjects. Each subject may contain supported discourse events, an outcome with mandatory scope, resulting actions and unresolved issues. Optional structures must be omitted when absent. Every semantic object must reference known evidence IDs.

The strict experimental schema validates:

  • non-empty subjects and globally unique semantic identifiers;
  • a closed discourse-event vocabulary;
  • non-empty, known and non-duplicated evidence references;
  • outcome text, scope, certainty and evidence;
  • action text, JSON-nullable responsibility/deadline and evidence;
  • unresolved-issue text and evidence;
  • omission rather than null or empty optional structures.

Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only, multiple options, unaccepted proposal, proposal with objection, rejected alternative, trial-scoped acceptance, no-decision discussion, resulting Action Item, and outcome plus unresolved issue. Evaluation targets semantic identity, development, outcome scope, actions, unresolved issues, traceability and absence of invented commitments rather than exact wording.

Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384, num_predict=4096. Each case received exactly one model call; there were no model retries or prompt iterations. The nine completed calls took 59.251 seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274 evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and failure artifacts were preserved under /tmp/meeting-lab-topic-reconstruction-v2-gold-run2/ and /tmp/meeting-lab-topic-reconstruction-v2-gold-run3/. Two earlier launch attempts made zero LLM calls: one failed on the script import path and one was blocked by sandbox networking.

Human-reviewed results after correcting two objectively wrong Gold assumptions without another model call:

Case Verdict Reason
A — idea only PARTIAL Correct subject and no invented outcome/action, but the isolated idea was labeled considered_option rather than introduced_idea.
B — multiple options FAIL Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues.
C — unaccepted proposal FAIL Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item.
D — proposal with objection FAIL Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue.
E — rejected alternative FAIL Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists.
F — trial-only acceptance PARTIAL Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event.
G — no decision FAIL Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome.
H — resulting action PASS Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references.
I — outcome plus unresolved FAIL Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list.

Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case F: the model distinguished acceptance for a bounded trial from acceptance as a final solution. It also handled the explicit action in case H well. However, the experiment failed systematically on sparse structured output, subject grouping and restraint around absent outcomes/actions/unresolved issues. The model frequently mirrored optional schema fields as empty/null values, treated alternatives as outcomes, split one discussion into multiple subjects, or invented open issues from mere non-selection.

The focused experiment is not promising enough to justify a real Progeo chunk sanity check. No such run was performed, and no architecture is accepted on the basis of this prototype. Further work should first analyze whether the failure comes from the schema/prompt representation, the model's sparse-output reliability, or the boundary between subject grouping and semantic synthesis. It should not proceed through repeated prompt tuning against these nine cases.