77 KiB
Experiments
This file records durable technical experiments and findings for Meeting Lab. It is not a diary and does not replace commit history.
Status values
- Proposed: experiment idea exists, but no result is recorded.
- Running: experiment is in progress and no decision has been made.
- Accepted: finding is the current baseline or design conclusion.
- Rejected: hypothesis was tested and should not be repeated as-is.
- Superseded: finding was useful but has been replaced by a newer baseline.
EXP-0001 - Whisper JSON interpretation
Status: Accepted
Date or period: 2026-07-29
Hypothesis:
Whisper JSON should be chunked from its segment stream, not from the aggregate top-level text field.
Setup:
chunk_transcript.py was updated to parse JSON input and prefer
segments[*].text when segments exists. A regression test supplies JSON with
both top-level text and separate segment texts.
Inputs:
- Minimal synthetic Whisper-style JSON in
tests/test_chunking.py. - Real Whisper artifacts under
samples/whisper/.
Model / configuration:
- No LLM.
Result:
The test verifies that the block stream is ["alpha", "beta", "gamma"] and
does not include the aggregate "alpha beta gamma" text. The repository history
records this as the fix for the earlier failure where the first chunk contained
the complete transcript.
Decision:
When segments exists, segments[*].text is the authoritative transcript
stream. The top-level text field is only a fallback.
Lessons learned:
Whisper JSON is structured input. Treating it like plain text can duplicate the entire transcript and invalidate downstream chunking.
Evidence:
src/meeting_lab/chunking/chunk_transcript.pytests/test_chunking.py- Commit
4656523-Fix Whisper JSON chunk extraction
EXP-0002 - Technical transcript chunking baseline
Status: Accepted
Date or period: 2026-07-29 to 2026-07-30
Hypothesis:
Sequential technical chunks around the configured target size can preserve the transcript while keeping extraction calls small enough for local models.
Setup:
The chunker splits block-aligned text with configurable target, minimum, maximum and overlap settings. Tests verify no duplicate later blocks when overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
Inputs:
- Synthetic block list in
tests/test_chunking.py. samples/whisper/meeting_speech_cleaned.json.
Model / configuration:
- No LLM for chunking.
- Manifest uses default chunking behavior recorded in
samples/whisper/meeting_speech_cleaned_chunks/manifest.json.
Result:
With overlap set to zero, tests verify that all blocks appear exactly once. The real sample manifest contains nine chunks, mostly near the configured target size, with a smaller final chunk.
Decision:
Independent sequential chunks are the current technical baseline. One normalized chunk per extraction call is the preferred extraction strategy.
Lessons learned:
Chunking solves model-size constraints only. It must not perform topic detection or semantic merging.
Evidence:
src/meeting_lab/chunking/chunk_transcript.pytests/test_chunking.pysamples/whisper/meeting_speech_cleaned_chunks/manifest.jsonAGENTS.mdPROJECT_KNOWLEDGE.md
EXP-0003 - Conservative transcript normalization
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Transcript cleanup should improve readability without changing meeting semantics.
Setup:
The normalizer removes isolated filler sounds, immediate duplicate words or short duplicate phrases, and redundant whitespace. It records changed blocks in a JSON change log and explicitly preserves semantic content categories.
Inputs:
- Chunk text files under
samples/whisper/meeting_speech_cleaned_chunks/. - Change logs such as
chunk_01_changes.json.
Model / configuration:
- No LLM.
Result:
The implementation and generated change logs show a conservative policy: negations, qualifiers, dates, numbers, responsibilities, technical statements, deadlines, decisions and commitments are preserved.
Decision:
Normalization remains deterministic and low-risk. When uncertain, leave text unchanged.
Lessons learned:
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove semantic cues needed by extraction.
Evidence:
src/meeting_lab/normalization/normalize_transcript.pysamples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.jsondocs/pipeline.md
EXP-0004 - Full-context topic segmentation
Status: Superseded
Date or period: 2026-07-21 to 2026-07-22
Hypothesis:
A single full-context topic segmentation call can identify topic boundaries in a normalized transcript chunk.
Setup:
The initial segmentation prototype asked the model for topic changes and then converted those boundaries into continuous, non-overlapping segments.
Inputs:
samples/chunks/chunk_01_normalized.txt.
Model / configuration:
- Generated artifact records
qwen3:14b.
Result:
The generated artifact contains 85 blocks, three topic-change boundaries and four segments. The run metadata records a substantially longer elapsed time than the later windowed artifact for the same input.
Decision:
Full-context segmentation was useful as a prototype, but it was superseded by windowed segmentation and manual review tooling.
Lessons learned:
The prototype established the boundary-to-segment representation, but did not settle segmentation quality.
Evidence:
src/meeting_lab/segmentation/segment_topics.pysamples/chunks/chunk_01_normalized_segments.json- Commit
f234efc-Add initial topic segmentation prototype - Commit
889a4fe-Detect topic boundaries as continuous segments
EXP-0005 - Windowed topic segmentation and review
Status: Accepted
Date or period: 2026-07-22
Hypothesis:
Windowed topic segmentation can reduce runtime and make boundary evaluation more inspectable than a single full-context call.
Setup:
segment_topics_windowed.py analyzes overlapping windows and reports only
boundaries from the decision range. Python merges boundaries into continuous,
non-overlapping segments. review_segmentation.py renders each boundary with
neighboring transcript context for manual classification.
Inputs:
samples/chunks/chunk_01_normalized.txt.- Full meeting normalized chunks under
samples/whisper/meeting_speech_cleaned_chunks/.
Model / configuration:
samples/chunksartifact:qwen3:8b, window size 20, overlap 3.- Full-meeting chunk artifacts:
qwen3.5:9b, window size 20, overlap 3.
Result:
The samples/chunks windowed artifact produced 11 boundaries and 12 segments
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
Manual review output shows that some boundaries were assessed as subtopics
rather than full topic changes. Full-meeting artifacts show one window per
already-small normalized chunk and two segments per chunk.
Decision:
Windowed segmentation and review tooling are accepted as prototype tooling, not as a stable production segmentation stage.
Lessons learned:
Windowing improves inspectability and can reduce runtime, but it can also cluster boundaries and over-segment. Manual review remains necessary.
Evidence:
src/meeting_lab/segmentation/segment_topics_windowed.pysrc/meeting_lab/segmentation/review_segmentation.pysamples/chunks/chunk_01_normalized_windowed_segments.jsonsamples/chunks/chunk_01_normalized_windowed_segments_review.mdsamples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json- Commit
1a6d731-Add windowed segmentation pipeline and review tooling
EXP-0006 - Qwen model comparison
Status: Accepted
Hypothesis:
Larger local Qwen-family models should improve meaningful extraction and segmentation, but model size alone will not solve prompt or pipeline problems.
Setup:
Project work used smaller models for smoke checks and larger local models for meaningful extraction or segmentation. Artifacts and project knowledge record the currently useful model roles.
Inputs:
- Gold Standard scenarios under
tests/gold/. - Generated segmentation artifacts under
samples/. - Generated extraction artifacts under
samples/whisper/meeting_speech_cleaned_chunks/.
Model / configuration:
qwen3:1.7b: smoke-test model according to project knowledge.qwen3.5:9b: current meaningful extraction and segmentation model according to project knowledge and generated full-meeting segmentation artifacts.qwen3:8bandqwen3:14b: present in earlier segmentation artifacts.
Result:
The repository supports the conclusion that qwen3.5:9b is the meaningful
current experiment model and qwen3:1.7b is useful for smoke tests. Larger
models and longer contexts may increase runtime substantially, but no hardware
benchmark suite is recorded.
Decision:
Use qwen3:1.7b for smoke tests and qwen3.5:9b for meaningful current
experiments. Do not assume model size alone fixes prompt or pipeline design.
Lessons learned:
Evaluation must separate model capability from prompt clarity, context strategy, extraction schema and consolidation.
Evidence:
PROJECT_KNOWLEDGE.mdsamples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.jsonsamples/chunks/chunk_01_normalized_segments.jsonsamples/chunks/chunk_01_normalized_windowed_segments.json
EXP-0007 - Thinking output and Ollama API behavior
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Thinking-capable Qwen models may return reasoning separately from the final answer, and extraction parsing should not fail merely because extra text or multiple JSON objects appear.
Setup:
The Ollama response reader checks response, chat-style message.content and
then thinking. The JSON parser tries a full parse first, then scans JSON
object candidates and returns the final valid object. A regression test covers
thinking text before final JSON.
Inputs:
- Synthetic parser test in
tests/test_extraction_protocol.py.
Model / configuration:
- No LLM run in the test.
- Code path is used by Ollama extraction.
Result:
The parser can handle additional text and multiple JSON objects where the final
valid object is the intended answer. Current code still falls back to thinking
only if no usable response or message content is present.
Decision:
Keep parser robustness, but do not treat thinking output as the root cause of all extraction failures.
Lessons learned:
API response shape and model output shape are separate concerns. Preserve raw responses when diagnosing failures.
Evidence:
src/meeting_lab/extraction/extract_chunks.pytests/test_extraction_protocol.pyPROJECT_KNOWLEDGE.md
EXP-0008 - JSON truncation and generation limits
Status: Accepted
Hypothesis:
Some extraction failures are caused by generation limits truncating JSON rather than by prompt wording or parser behavior.
Setup:
A qwen3.5:9b extraction failure was diagnosed as truncated JSON. The
generation limit was increased for the successful path. Exact failing limit is
not recorded in the repository; the current extractor default is verifiably
--num-predict 8192.
Inputs:
- Local extraction runs referenced by project knowledge.
- Current extraction CLI.
Model / configuration:
qwen3.5:9b.- Current extractor default:
num_predict=8192.
Result:
Increasing the generation limit fixed the technical JSON failure. This was not primarily a parser or prompt problem.
Decision:
When JSON is truncated, inspect raw output and generation limits before editing prompts.
Lessons learned:
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong first response if the model simply ran out of output tokens.
Evidence:
src/meeting_lab/extraction/extract_chunks.pyPROJECT_KNOWLEDGE.mdAGENTS.md
EXP-0009 - Minimal end-to-end pipeline
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
A minimal local pipeline can transform Whisper output into chunk extractions and an interim protocol, proving the technical path before the final architecture exists.
Setup:
The repository added cleanup, normalization, chunking, extraction and protocol builder scripts, with sample generated artifacts.
Inputs:
samples/whisper/meeting_speech.jsonsamples/whisper/meeting_speech_cleaned.json- Generated chunks and normalized chunks under
samples/whisper/meeting_speech_cleaned_chunks/.
Model / configuration:
- Local Ollama extraction for chunk JSON.
- Windowed segmentation artifacts use
qwen3.5:9b.
Result:
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
extraction JSON files, windowed segmentation artifacts and
meeting_protocol.md.
Decision:
The minimal pipeline is technically validated. The first protocol builder is an interim validation tool, not the final architecture.
Lessons learned:
End-to-end execution exposed the next limitation: extraction output needs consolidation and purpose-specific rendering.
Evidence:
scripts/clean_whisper_json.pysrc/meeting_lab/normalization/normalize_transcript.pysrc/meeting_lab/chunking/chunk_transcript.pysrc/meeting_lab/extraction/extract_chunks.pysrc/meeting_lab/protocol/build_protocol.pysamples/whisper/meeting_speech_cleaned_chunks/- Commit
07b0d80-Implement first end-to-end meeting analysis pipeline
EXP-0010 - Gold Standard corpus
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Reproducible prompt engineering requires synthetic transcripts with explicit expected semantic outputs.
Setup:
The Gold Standard corpus defines scenario directories with transcript.txt,
expected.json and README files describing ground truth and common model
mistakes. The runner validates schema keys and writes actual.json for a
scenario.
Inputs:
- Gold scenarios under
tests/gold/.
Model / configuration:
- Runner requires an explicit Ollama model for LLM evaluation.
- Existing unit tests for the runner do not invoke Ollama.
Result:
The corpus gives stable semantics for decisions, facts, positions, todos, questions, technical details and difficult mixed cases. Initial structured transcripts are Phase 1 and easier than raw Whisper-style transcripts.
Decision:
Use Gold Standard tests as both regression tests and formal meeting-semantics specification. Raw or unlabelled transcript cases remain later-phase work.
Lessons learned:
Without expected outputs, prompt changes cannot be evaluated reproducibly.
Evidence:
tests/gold/scripts/run_gold_test.pytests/test_gold_runner.py- Commit
f7ad9ba-Establish prompt engineering baseline with Gold Standard tests
EXP-0011 - Gold-test quality and unique ground truth
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
If a prompt produces unexpected behavior, the gold test itself may be ambiguous and should be reviewed before the prompt is changed.
Setup:
Decision-focused scenarios were clarified during baseline creation. The current methodology requires checking unique ground truth before changing prompts.
Inputs:
decision_simpledecision_deferreddecision_none- Gold methodology document.
Model / configuration:
- Prompt Version 2 baseline work.
Result:
decision_simple required clarification around the explicit agreement and
nearby non-decision wording. The earlier negative/deferral ambiguity is now
represented by distinct decision_none and decision_deferred scenarios in
the repository. Punctuation is not reliable evidence for Whisper transcripts;
agreement language and wording must carry the semantics.
Decision:
Ambiguous gold tests must be reviewed before prompt changes. Do not treat punctuation as reliable evidence in real Whisper-style transcripts.
Lessons learned:
Bad gold tests create false prompt failures and can encourage overfitting.
Evidence:
tests/gold/PROMPT_ENGINEERING_METHODOLOGY.mdtests/gold/decision_simple/README.mdtests/gold/decision_deferred/README.mdtests/gold/decision_none/README.md- Commit
f7ad9ba
EXP-0012 - Decision taxonomy
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Decision extraction needs a formal taxonomy that distinguishes substantive decisions from process decisions and non-decisions.
Setup:
The decision definition document and decision prompt define included and excluded categories. Gold tests cover explicit decisions, true no-decision cases and deferrals.
Inputs:
tests/gold/DECISION_DEFINITION.mdprompts/decisions.md- Decision gold scenarios.
Model / configuration:
- Prompt Version 2 baseline.
Result:
Accepted decision categories include substantive decisions, organizational decisions, process decisions, approvals, rejections, deferrals, explicit decisions not to decide yet and explicit agreement to gather more information before deciding. Opinions, preferences and proposals without agreement are not decisions.
Decision:
"No decision was reached" and "the decision was deferred" are distinct semantic outcomes.
Lessons learned:
Deferral can be a valid process decision even when the substantive topic remains unresolved.
Evidence:
tests/gold/DECISION_DEFINITION.mdtests/gold/decision_deferred/tests/gold/decision_none/prompts/decisions.md
EXP-0013 - Prompt engineering methodology
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Prompt iteration needs strict experimental controls to prevent regression, overfitting and arbitrary prompt churn.
Setup:
The methodology was documented alongside the Gold Standard corpus and later summarized for agents.
Inputs:
- Gold scenarios.
- Prompt files.
Model / configuration:
- Applies to all prompt experiments.
Result:
The accepted method is one prompt change per iteration, one target test at a
time, immediate validation, no regressions, no expected.json edits merely to
force a pass, no test-specific prompt hacks, stopping after two consecutive
non-improving iterations, and verifying unique ground truth before prompt
changes.
Decision:
Prompt changes are controlled experiments. See AGENTS.md for agent operating
rules.
Lessons learned:
Most prompt changes are not isolated unless the experiment explicitly constrains the target and regression set.
Evidence:
tests/gold/PROMPT_ENGINEERING_METHODOLOGY.mdAGENTS.md- Commit
f7ad9ba
EXP-0014 - Decision Prompt Version 2
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Adding explicit process-decision language to the decision prompt can preserve true decision detection while recognizing deferrals.
Setup:
Prompt Version 2 added explicit support for deferrals and process decisions. The baseline was validated on three decision scenarios.
Inputs:
decision_simpledecision_deferreddecision_none
Model / configuration:
- Prompt Version 2.
- Model used for validation is not recorded in the committed methodology.
Result:
The committed methodology records all three baseline scenarios as passing. The
current generated actual.json files also show the expected decision count for
these decision scenarios, although some non-decision categories remain less
complete.
Decision:
Prompt Version 2 is the current decision baseline. Explicit deferrals are recognized as process decisions.
Lessons learned:
Decision-count success does not imply all categories are solved. Category-level evaluation must continue.
Evidence:
tests/gold/PROMPT_ENGINEERING_METHODOLOGY.mdtests/gold/decision_simple/actual.jsontests/gold/decision_deferred/actual.jsontests/gold/decision_none/actual.jsonprompts/decisions.md
EXP-0015 - Difficult synthetic meeting
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
A deliberately adversarial synthetic meeting can expose extraction failures that simple category tests miss.
Setup:
evil_meeting includes interruptions, corrections, absent referenced people,
near-decisions, changed positions and one expected explicit decision.
Inputs:
tests/gold/evil_meeting/.
Model / configuration:
- Existing generated
actual.json; exact model is not stored in the artifact.
Result:
The generated result found the expected FR-7 exclusion decision. It also classified "do not migrate until the mapping table is checked" as an additional decision. The current expected file treats that statement as a position, but contextual review suggests it may be a valid process instruction or decision.
Decision:
Do not classify this as a simple model failure without reviewing the gold standard. The scenario exposes a semantic gap in the expected output.
Lessons learned:
Difficult synthetic cases are valuable because they reveal ambiguity in the specification as well as model mistakes.
Evidence:
tests/gold/evil_meeting/README.mdtests/gold/evil_meeting/expected.jsontests/gold/evil_meeting/actual.json- See EXP-0011 and EXP-0012.
EXP-0016 - Context-size extraction comparisons
Status: Rejected
Hypothesis:
Increasing extraction context from one chunk to neighboring chunk groups should make extraction more complete and therefore should become the baseline.
Setup:
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3 and 1+2+3.
Inputs:
- Normalized meeting chunks.
Model / configuration:
- Current project knowledge identifies
qwen3.5:9bas the meaningful model for extraction experiments.
Result:
More context sometimes improved completeness, but it also shifted category classification, added duplicates and reduced stability. No numeric winner is recorded in the repository.
Decision:
Do not adopt larger extraction windows as the baseline. Independent chunk extraction remains current strategy. Recover global context through consolidation rather than continuously enlarging extraction windows.
Lessons learned:
Context size is not a monotonic quality knob. It changes the task the model is performing.
Evidence:
PROJECT_KNOWLEDGE.mdAGENTS.mdROADMAP.md- See EXP-0002 and EXP-0018.
EXP-0017 - Independent full-meeting chunk extraction
Status: Accepted
Hypothesis:
Extracting every normalized chunk independently can produce enough structured material for a useful protocol draft.
Setup:
Nine normalized chunks were extracted into separate JSON files and then rendered by the interim protocol builder.
Inputs:
samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txtthroughchunk_09_normalized.txt.
Model / configuration:
- Existing extraction artifacts do not record model metadata.
Result:
The nine extraction files contain facts, decisions, todos, questions and
technical details. meeting_protocol.md aggregates them into a readable draft.
Duplicates, category shifts and synthesis became the dominant limitations.
Decision:
Independent chunk extraction is useful enough to keep as the baseline, but it requires a consolidation stage.
Lessons learned:
Per-chunk extraction gives recall-oriented raw material. It does not by itself produce a polished or canonical meeting representation.
Evidence:
samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.jsonsamples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.mdsrc/meeting_lab/protocol/build_protocol.py- See EXP-0018.
EXP-0018 - Human protocol comparison and output-view split
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
One generated protocol cannot satisfy every use case; protocol output should be separated by purpose and audience.
Setup:
The interim machine protocol was compared against the desired human protocol shape and then the architecture was revised toward parallel output views.
Inputs:
- Interim
meeting_protocol.md. - Architecture and output-view documentation.
Model / configuration:
- Not applicable; this is a design evaluation.
Result:
The human protocol target is denser and organized by purpose and topic rather than extraction categories. The machine extraction retains more context and is useful for recall, but it is not the right direct source for a concise distribution artifact or durable knowledge entry.
Decision:
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution Protocol / Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
Lessons learned:
Rendering is a separate concern from extraction and consolidation. Output views must be parallel renderings of shared semantics, not transformations of one another.
Evidence:
samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.mddocs/output-views.mddocs/architecture.mddocs/pipeline.md- Commit
5c03ed7-Refine canonical meeting knowledge architecture
EXP-0019 - Consolidation and Canonical Meeting Knowledge
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Extraction, consolidation and output rendering are separate problems and should not be collapsed into one LLM prompt or one protocol file.
Setup:
The architecture was refined after the minimal pipeline and protocol draft showed duplicate, synthesis and audience-specific rendering limitations.
Inputs:
- Extraction JSON artifacts.
- Interim protocol draft.
- Architecture and data-model documentation.
Model / configuration:
- Not applicable; this is an architectural conclusion.
Result:
The accepted design is a planned Canonical Meeting Knowledge layer as the semantic source of truth, with Working Protocol, Distribution Protocol and Knowledge Objects as parallel output views. The next consolidation architecture is split into a Deterministic Canonicalizer and a Semantic Consolidator. The canonicalizer prepares validated evidence-bearing objects without uncertain semantic merging. The consolidator then merges semantically equivalent statements, preserves evidence, reconciles category shifts where supported and marks contradictions or uncertainty.
Decision:
Deterministic canonicalization is implemented as Canonicalizer V1. The first semantic consolidation milestone is implemented as Semantic Consolidator V0 for facts-only duplicate detection. Broader semantic consolidation, Canonical Meeting Knowledge and final output views remain planned.
Lessons learned:
Global meeting understanding should be recovered by consolidation over evidence-bearing extractions, not by silently changing output views or expanding LLM context indefinitely.
Evidence:
docs/architecture.mddocs/pipeline.mddocs/data-models.mddocs/output-views.mdPROJECT_KNOWLEDGE.mdROADMAP.md- Commit
5c03ed7
EXP-0020 - Working Protocol Synthesizer V0
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The current local synthesis model may be able to generate a useful detailed Working Protocol directly from the existing independent chunk extraction JSON files, before canonicalization or semantic consolidation exists.
Setup:
One synthesis prompt was constructed from exactly nine chunk extraction JSON files. The model was instructed to use only those extraction files, merge duplicates, group related information into topics, preserve useful discussion context and write a neutral technical Working Protocol.
Inputs:
chunk_01_extraction.jsonthroughchunk_09_extraction.json.- No original transcript, normalized chunks or Whisper output were used as synthesis input.
Model / configuration:
- Model:
qwen3.5:9b - Prompt characters: 24,979
- Actual prompt eval tokens: 5,929
- Output tokens: 1,486
- Runtime: 294.204 seconds
Result:
The generated Working Protocol was readable, well structured and topic-oriented. It was still based directly on raw chunk extractions, without a separate deterministic canonicalization stage or semantic consolidation stage. The output language was English even though the source meeting material was German.
Decision:
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline for later canonicalizer, consolidator and renderer comparisons. This selected generated artifact is intentionally versioned even though generated runtime artifacts are normally ignored.
Lessons learned:
Direct synthesis from chunk extractions can create a useful recall-oriented draft, but it does not replace Canonical Meeting Knowledge. The language mismatch also establishes a default renderer rule: protocol output should normally match the dominant source language unless an explicit output language is requested.
Evidence:
samples/benchmarks/working_protocol_synthesizer_v0/README.mdsamples/benchmarks/working_protocol_synthesizer_v0/working_protocol.mdPROJECT_KNOWLEDGE.mddocs/output-views.md- See EXP-0017 and EXP-0019.
EXP-0021 - Canonicalizer V1
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
Independent chunk extraction JSON can be converted into a stable deterministic intermediate format before any semantic LLM consolidation is attempted.
Setup:
Canonicalizer V1 discovers chunk_*_extraction.json files in stable chunk
order, validates required categories, normalizes category names and basic field
structure, parses existing legacy string formats where safe, trims redundant
whitespace, assigns deterministic IDs, preserves original values and source
references, and merges only exact duplicates when all semantic fields are
identical.
Inputs:
- Synthetic unit-test fixtures.
- Existing nine extraction JSON files under
samples/whisper/meeting_speech_cleaned_chunks/.
Model / configuration:
- No LLM.
- CLI module:
meeting_lab.consolidation.canonicalize.
Result:
Canonicalizer V1 produces schema_version, source_files, stats and
items. It is deterministic preparation for the future Semantic Consolidator
and is not Canonical Meeting Knowledge.
Decision:
Canonicalizer V1 is the current implemented deterministic canonicalization stage. Semantic Consolidator V0 now uses this representation for facts-only semantic duplicate detection; broader semantic consolidation and Canonical Meeting Knowledge remain planned.
Lessons learned:
Exact duplicate handling, source-reference preservation and legacy string parsing can be tested without model calls. Any uncertain semantic merge remains out of scope for this stage.
Evidence:
src/meeting_lab/consolidation/canonicalize.pytests/test_canonicalize.pydocs/data-models.mddocs/pipeline.md
EXP-0022 - Semantic Consolidator V0 facts-only merge
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The Canonicalizer V1 output contains enough stable structure for a local LLM to identify semantically equivalent fact items without losing source coverage or changing non-fact categories.
Setup:
Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON,
selected only items with category: "fact", and sent one bounded consolidation
request to local Ollama. The merge rules required semantic equivalence, not
topic similarity, and validation required every source fact ID to appear
exactly once.
Inputs:
samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json- 33 fact items.
Model / configuration:
qwen3.5:9B- Ollama endpoint:
http://127.0.0.1:11434/api/generate - Thinking disabled.
- One LLM call.
num_ctx=32768num_predict=4096
Result:
- Runtime: 390.119 seconds on the current machine.
- Merged fact groups: 1.
- Source facts involved in merges: 2.
- Singleton fact groups: 31.
- Validation: passed.
- No source fact was lost or duplicated.
- Non-fact categories remained unchanged.
Accepted merge:
fact_0025+fact_0031- Canonical statement: "Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up."
Decision:
Semantic Consolidator V0 is complete for its current narrow scope:
conservative facts-only semantic duplicate detection with source evidence
preserved. The selected report.md and consolidated_extractions.json
benchmark artifacts should be versioned for later comparison. The raw model
response remains a local diagnostic artifact and is not versioned.
Lessons learned:
Semantic duplicate consolidation is technically viable and conservative enough for continued evaluation, but broader semantic synthesis remains a separate future stage. The measured runtime is useful for this machine and run, but should not be generalized into a universal benchmark.
Evidence:
src/meeting_lab/consolidation/consolidate_facts.pyprompts/consolidate_facts.mdtests/test_consolidate_facts.pysamples/benchmarks/semantic_consolidator_v0/report.mdsamples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json- Local diagnostic only:
samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt
EXP-0023 - Responsibility attribution integrity
Status: Running
Date or period: 2026-07-31
Hypothesis:
Protocol generation is operationally unsafe if responsibility, ownership or departmental role attribution is inferred from discussion context rather than explicit meeting evidence.
Setup:
The real-life Working Protocol Renderer V2 benchmark was inspected against the consolidated input. A false assignment connected a Marketing participant to Business Development criteria work even though the participant's contribution was critical or reluctant and did not establish acceptance of that task.
Inputs:
samples/benchmarks/working_protocol_renderer_v2/working_protocol.mdsamples/benchmarks/semantic_consolidator_v0/consolidated_extractions.jsonsamples/benchmarks/canonicalizer_v1/canonicalized_extractions.json
Model / configuration:
- Not rerun for this finding.
- Finding is based on existing benchmark artifacts.
Result:
The false responsibility attribution is visible in the structured input before rendering, so the issue is not merely stylistic renderer wording. The root cause may originate earlier in extraction and then be preserved by canonicalization and consolidation. Renderer guardrails are still required so output views do not strengthen ambiguous ownership.
Decision:
Responsibility attribution is now treated as a critical project-wide invariant. A person, team or department may be recorded as responsible only when the evidence explicitly assigns, accepts or confirms that responsibility. Discussion, expertise, objection, suggestion, thematic proximity, speaker adjacency, organizational assumptions and likely job roles do not establish ownership.
Lessons learned:
This class of error affects operational correctness, not only style. The pipeline needs traceable attribution evidence and future schema support for responsibility status such as explicit, accepted, proposed or unclear.
Evidence:
AGENTS.mdPROJECT_KNOWLEDGE.mddocs/data-models.mddocs/output-views.mdprompts/working_protocol.mdtests/gold/responsibility_attribution_negative/
EXP-0024 - Working Protocol V2 contract visibility
Status: Running
Date or period: 2026-08-09
Target:
BUG-011 renderer-only regression using an already validated Semantic Consolidator artifact.
Hypothesis:
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
placing the leading prompt contract outside the model's effective evaluated
context. The preserved failures all report prompt_eval_count=16386, while
their outputs either echo trailing JSON or produce an unconstrained generic
category summary instead of the requested Working Protocol.
Iteration 1 change:
- Project every consolidated item to rendering-relevant semantic fields while retaining every item and its category/text/responsibility/deadline/status information.
- Generate the exact structural contract from renderer validator constants and append it after the compact INPUT JSON.
- Replace the independently handwritten prompt skeleton with a reference to that authoritative appended contract.
- Enforce the existing prompt rule that emitted sections must not be empty.
This is one renderer-contract prompt iteration. It does not change extraction, canonicalization, semantic consolidation or responsibility semantics.
Validation before LLM run:
- 15 focused renderer tests pass.
- Tests cover contract generation, compact input projection, valid and invalid headings, missing topic sections, empty sections, wrapper cleanup, malformed Markdown and final-file write gating.
Decision:
Iteration 1 passed structural validation and wrote working_protocol.md, but
the quality sanity check found that the model omitted most of the ten supplied
decisions, two open questions and several action items. The structurally valid
result therefore was not accepted as BUG-011 verification.
Iteration 2 change:
- Add input-derived hidden coverage markers for every decision, action item and open question.
- Require every priority item exactly once in its matching section.
- Validate missing, duplicate, unknown and wrong-section markers deterministically.
- Keep facts and technical details condensable as background.
This is the second single prompt iteration. It responds to the concrete omission failure observed in Iteration 1 without changing upstream semantics or inventing renderer content.
Iteration 2 pre-run validation:
- 18 focused renderer tests pass, including exact required-item coverage and wrong-section rejection.
Decision:
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget after emitting all decisions and most action items. Adaptive renderer budgeting resolved to 8,192 tokens for this input. The final run stopped normally after 3,709 evaluated output tokens.
The final renderer-only regression passed strict validation and wrote
working_protocol.md. Exact coverage was 10/10 decisions, 36/36 renderable
action items and 23/23 open questions, each once in its matching section. One
structurally empty action item whose task, responsible, deadline and evidence
were all null was recorded and excluded rather than fabricated. Optional
background markers were accepted only when they referred to real projected
input items.
Accept the compact renderer input, validator-derived trailing contract, priority-item coverage markers, empty-section validation and adaptive renderer output sizing as the BUG-011 baseline. This establishes structural reliability and priority-item coverage, not complete protocol prose quality.
Evidence:
prompts/working_protocol.mdsrc/meeting_lab/protocol/render_working_protocol.pytests/test_render_working_protocol.pysamples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/
EXP-0025 — BUG-015 classification precision
Date: 2026-08-09
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
Model/configuration: qwen3.5:9B, temperature 0, num_ctx=32768.
Tests were created before prompt changes. A Decision-only evidence threshold kept the explicit Dr. Schlummer rejection and omitted an option and preference. Adding Action and Open Question definitions improved several negatives but was not stable: the model alternately promoted an unaccepted Textor suggestion or moved rejected candidates into Open Questions. Moving the standalone category prompts after the transcript made the partial Decision schema dominate and misclassified true Action Items as Decisions in two consecutive runs.
The final iteration replaced the competing standalone category prompts with a single unified classification contract after the transcript. It preserved the assigned Nina task and ownerless established CET work, and prevented cross-category leakage in the focused case, but still emitted the unaccepted Textor suggestion as an Action Item. Further prompt iterations were stopped in accordance with the Gold Standard methodology.
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015 remains Open; no phrase-specific deterministic filter was introduced.
EXP-0027 — Evidence-near observation extraction
Date: 2026-08-18
Hypothesis: qwen3.5:9B can more reliably extract evidence-near linguistic and
semantic properties than directly synthesize protocol-level events, outcomes,
actions and unresolved issues. This isolated experiment stops before semantic
interpretation and does not connect to the production pipeline.
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
source evidence, explicit targets, a five-value relation vocabulary, modality,
temporality, evaluation, agreement, responsibility/person, uncertainty,
clarification need and free-text scope. It contains no protocol-level category
field. The validator requires sequential observation IDs, known evidence IDs,
backward-only valid observation targets, closed categorical vocabularies,
consistent responsibility/person pairs, and one canonical absence form: JSON
null for person and absent for scope.
Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384,
num_predict=4096, no retries. All nine cases ran exactly once, for nine LLM
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
responses, parsed observations, validation results, Ollama metadata and
comparisons are preserved under
/tmp/meeting-lab-evidence-observations-v1-20260818/.
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
Five cases failed structure because the model represented a single target as a
one-element list, usually ["discussion_subject"]; the accepted schema permits
a list only for two or more jointly referenced observations. Several responses
also copied the relation label limits_scope into the free-text scope field.
These were systematic model-output errors, not transport or parser failures.
The prompt and run were not retried or tuned.
Human semantic review of the preserved raw responses:
| Case | Verdict | Main result |
|---|---|---|
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override strict structural failures; it separates useful semantic signal from schema compliance.
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving closer to evidence reduced some direct promotion behavior: the geometry mention did not become work, the washing disadvantage did not become an unresolved issue, both alternatives in B remained uncommitted, and the publication query did not become an assignment. However, the important promotion errors did not disappear. C acquired unsupported acceptance, and G still promoted a personal preference into rejection. Positive cases were only partly preserved: explicit rejection was recognized but incorrectly linked; trial-only language was kept but responsibility was invented; Nina's request and commitment were recognized but responsibility states were wrong; and the publication issue was recognized but its uncertainty was contradicted by rejection.
Result: B — evidence-near extraction is promising, but specific observation dimensions remain unreliable. Target/relation selection, scope attachment, responsibility state/person attribution, and agreement versus uncertainty are not reliable enough to justify designing the later interpretation stage yet. No production integration or later interpretation stage was implemented.
EXP-0028 — Evidence-Near Observation Extraction V2
Date: 2026-08-19
V2 tested whether qwen3.5:9B preserves the evidence needed by a later
controlled interpretation stage when direct responsibility, agreement and
semantic graph relations are removed. Responsibility was replaced by explicit
participant/discourse facts (speaker, named_person, addressee, singular
self-reference, collective we, and impersonal person reference). Agreement
was replaced by explicit affirmation, explicit negation and determination
statement signals. Graph relations were reduced to nullable scalar
refers_to; scope became free-text qualifier plus nullable scalar
limits_target. No later derivation stage was implemented.
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
human interpretations. It contains no responsibility, agreement, action,
decision, open-question, accepted-trial, rejected-alternative or protocol
eligibility fields. Validation enforces known evidence IDs, sequential unique
observation IDs, backward-only scalar references, closed vocabularies, boolean
participant flags, JSON-nullable participant/qualifier/reference fields and no
string "null".
Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384,
num_predict=4096, no retries or voting. A launch-path defect was corrected
before the live run; the failed launch made zero model calls. A sandbox-blocked
localhost attempt also made zero model calls. The completed run called the
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
Persistent prompts, Gold input and expectations, raw and parsed model output,
validation, automatic comparison, Ollama metadata and human evaluation are in
artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/.
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
were schema-invalid. The dominant serialization pattern was use of present
instead of the specified explicit for affirmation/negation; E additionally
used none instead of absent for a determination signal, while D emitted the
separate uncertainty concept as an invalid modality. These errors are
contract violations, although most present/explicit differences are
deterministically normalizable without changing meaning. A and C were valid
JSON/schema outputs but had critical semantic mismatches.
Human semantic review:
| Case | Verdict | Main result |
|---|---|---|
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; present caused schema failure. |
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced V1 promotion errors: collective speech and speaker identity no longer became individual responsibility; personal preference no longer became a group-level rejection field; an information question did not become work; and explicit negation/affirmation survived as separate evidence. Useful participant evidence also survived strongly in H and collective-speech evidence in F.
Simplification did not make all evidence-near dimensions reliable. Scalar
references and limits_target were almost entirely omitted, qualifiers were
usually omitted, committed modality was missed in E, F and H, and C/G repeated
important modality, uncertainty and participant-feature errors. Some positive
semantic information therefore survived only in free-text content, not in
the structural signals a controlled derivation stage would need.
Result: B — V2 is materially better, but specific evidence-near dimensions still require refinement. Direct responsibility, agreement and graph-relation classification should remain excluded. Before designing the derivation stage, the next work should examine the minimal reliable representation of explicit reference/scope limitation, commitment modality and participant deixis. No production integration, Progeo run or derivation implementation was performed.
EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
Date: 2026-08-19
Hypothesis: qwen3.5:9B is substantially more reliable when the first semantic
stage preserves meeting meaning as atomic natural-language observations with
provenance and only simple participant information, without classifying or
deriving higher-level meeting semantics.
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
observation contains exactly observation_id, evidence_id, content,
speaker, nullable named_person, and nullable addressee. It contains no
modality, temporality, evaluation, affirmation, negation, determination,
uncertainty, clarification, responsibility, agreement, relation, reference,
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
Instead, the prompt asks for conservative atomic content that retains hedges,
conditions, personal/collective/impersonal language, requests, acceptances,
rejections, quantities, deadlines and boundaries in natural language.
Structural validation is intentionally small: exact schema keys, non-empty
observations/content, unique obs_N identifiers, known evidence IDs, speaker
matching its evidence, explicit named people/addressees, and no string
"null". Human semantic preservation against per-case requirements is the
primary evaluation; wording differences do not fail a case.
Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384,
num_predict=4096, no retries, voting or per-case tuning. One sandbox-blocked
localhost launch made zero model calls. The completed run made exactly nine
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
structural validation. Persistent source evidence, semantic requirements,
exact prompts, raw and parsed responses, validation, Ollama metadata and human
evaluation are stored under
artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/.
Human semantic preservation results:
| Case | Verdict | Main result |
|---|---|---|
| A | PASS | Preserved kann, vielleicht, tentative follow-up, and explicit Dann sequence without commitment. |
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, wenn überhaupt, impersonal checking need and no decision, but dropped collective wir from who would receive contaminated material. |
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
“that we receive contaminated material back” into an impersonal passive phrase;
the risk itself remained hypothetical. H translated Freitag to Friday, a
harmless wording difference. I retained two compound observations rather than
splitting every proposition, but all required semantic boundaries and
dependencies remained explicit.
Compared with V2, categorical-field removal improved content preservation in
A, G and I: A retained Dann; G retained wenn überhaupt, personal Ich and
impersonal Man; I retained publication purpose and all boundaries. It also
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
strengthening into responsibility, group rejection, established work or
assigned clarification. F and H remain sufficiently complete in natural
language for a later interpretation experiment. No useful meaning was shown to
depend on the removed fields; the V3 content retained the useful signals that
V2's fields had attempted to encode.
Result: A — MINIMAL FIRST STAGE ACCEPTED. On A-I, minimal atomic content with evidence provenance and simple participants is sufficiently reliable to be the candidate first semantic stage. A later bounded experiment may examine controlled semantic interpretation, but no derivation stage, production integration or Progeo run was implemented here.
EXP-0030 — V3 model comparison: qwen3.5:9B vs qwen3.6:35B-A3B
Date: 2026-08-19
This controlled comparison reran the accepted V3 minimal semantic-preservation
experiment unchanged with the locally installed qwen3.6:35B-A3B. It used the
same implementation, A-I fixture, evidence, prompt, minimal schema, temperature
0, think=false, num_ctx=16384, num_predict=4096, no retries, no voting and
one call per case. The completed run made exactly nine calls in 96.392 seconds.
Artifacts are preserved under
artifacts/experiments/evidence_observations_v3/20260819_v3_qwen36_35b_a3b_single_run/.
Structural validation passed 8/9 cases. D was semantically faithful but invalid
because the model copied transcript speakers Antonius and Martin into
named_person, although those names were not explicitly named within their
utterances. Human semantic preservation produced 7 PASS, 1 PARTIAL and 1 FAIL:
| Case | Verdict | Main result |
|---|---|---|
| A | PASS | Preserved can, perhaps, tentative would, explicit then and no commitment, but changed the content language to English. |
| B | PARTIAL | Preserved both approaches overall, but removed Oder from the 20-20 observation and locally strengthened it into collective planned conduct. |
| C | PASS | Preserved Tim's tentative personal Textor contact and possible follow-up without established work. |
| D | PASS | Preserved possibility, uncertainty, process and energy consequence; structural failure was confined to invalid speaker-as-named-person values. |
| E | PASS | Preserved cost, collective rejection/non-pursuit and later determination, though the confirmation dropped explicit Ja. |
| F | FAIL | Preserved quantity, timing and trial/series boundaries, but changed collective wir into “Martin suggests” and “Tim agrees,” inventing individual proposal/agreement meaning. |
| G | PASS | Preserved hypothetical risk, collective recipient wir, Martin's personal stance, wenn überhaupt, conditional Technikum, impersonal man müsste and no decision/owner. |
| H | PASS | Preserved request, addressee, Friday, explicit acceptance and future personal commitment without a responsibility field; content was English. |
| I | PASS | Preserved local/pure-production and comparison boundaries, five-degree difference, upstream effort, publication purpose, unresolved permission and clarification without assignment. |
Direct comparison:
| Measure | qwen3.5:9B |
qwen3.6:35B-A3B |
|---|---|---|
| Structurally valid | 9/9 | 8/9 |
| Human PASS | 8 | 7 |
| Human PARTIAL | 1 | 1 |
| Human FAIL | 0 | 1 |
| Observations | 28 | 30 |
| Runtime | 40.074 s | 96.392 s |
| LLM calls | 9 | 9 |
| Prompt-evaluation tokens | 6,268 | 6,268 |
| Evaluation tokens | 2,590 | 2,718 |
The larger model fixed the 9B weakness in G by preserving collective wir, and
it split I's compound production/effort observations more cleanly. Those gains
did not offset regressions: B was locally strengthened, F materially converted
collective conduct into individual agreement, D violated the participant
schema, observation count increased, and runtime was 2.4 times higher. Both
models preserved German consistently in six of nine cases, but in different
cases; the 35B-A3B model changed A, F and H to English, while 9B changed C, F
and H wholly or partly to English.
Result: D — REGRESSION. qwen3.6:35B-A3B does not materially improve the
accepted minimal V3 first-stage preservation over qwen3.5:9B; it is worse on
the A-I comparison because of the F ownership-adjacent strengthening and lower
structural validity. This conclusion applies only to the minimal V3 first
stage and does not determine model choice for any later semantic derivation.
No production integration, derivation implementation or Progeo run occurred.
EXP-0031 — Controlled Semantic Derivation H V0
Date: 2026-08-19
This isolated experiment tested the first controlled second-stage derivation
using only the accepted qwen3.5:9B V3 observations for case H. The derivation
LLM received the two V3 observations, not the transcript or Gold expectation.
Its deliberately narrow task was limited to recognizing whether obs_1 is a
concrete request and whether obs_2 explicitly commits its speaker to
substantially the same work. Its strict output schema forbids responsibility,
requested actor, establishment/status, Action Item, protocol, confidence and
generic relation/graph fields.
Deterministic code validates observation/evidence provenance, obtains the
requested actor only from the request observation's addressee, requires the
acceptance to follow the request, requires the accepting speaker to equal that
addressee, and establishes responsibility only after all semantic and
structural gates pass. A bounded weekday normalizer reconciles Friday and
Freitag, rejects conflicting weekdays, and separates the supported due date
from the normalized action text. No general temporal or action ontology was
introduced.
Twenty focused deterministic tests cover the positive H path and the required negative invariants: request alone, acknowledgement/non-commitment, tentative acceptance, different response speaker, different work, reversed order, speaker/name/addressee alone, conflicting deadlines, unknown observation IDs, inconsistent evidence provenance, forbidden semantic fields, malformed JSON and persistent artifacts. The complete non-LLM suite passed 192/192.
Configuration: one qwen3.5:9B call, temperature 0, think=false,
num_ctx=16384, num_predict=1024, no retries or voting. The call took 11.765
seconds, with 469 prompt-evaluation and 124 evaluation tokens. The model
returned a valid recognition object: obs_1 is a concrete request, obs_2 is
an explicit commitment, and both concern substantially the same work. It
returned no responsibility or establishment judgment.
All deterministic gates passed. The final derived result is an established
action Prüfung der Messdaten, requested from and assigned to Nina, due
Freitag, supported by request obs_1/e1 and acceptance obs_2/e2. The model
included bis Friday in its normalized request text; after the single call, a
deterministic-only bounded correction separated that already-recognized due
phrase from action content without changing the prompt, recognition schema,
semantic result or call count. Focused and complete non-LLM suites still
passed after this correction.
Artifacts are preserved under
artifacts/experiments/controlled_semantic_derivation_h/20260819_h_qwen35_9b_single_run/.
Result: the H mechanism succeeded. This establishes only that the narrow
request-plus-explicit-acceptance pattern can be recognized and gated for H; it
does not generalize the derivation architecture to other cases or semantic
categories. No production integration, other case run, semantic graph,
protocol derivation or Progeo run occurred.
EXP-0032 — Request / Acceptance Gold V0
Status: Experimental; promising with semantic precision gaps
Date: 2026-08-20
This isolated regression experiment tested whether the EXP-0031 mechanism generalizes beyond H. It used ten short synthetic cases containing only V3-style observations. Evidence Observation V3 was neither called nor changed, and the model received no raw transcript or expected result. The fixed recognition schema permits only a nullable concrete request and nullable later explicit personal commitment, plus the same-requested-work judgment and normalized action text. Responsibility, requested actor, established status, Action Item, protocol, confidence and generic graph fields remain forbidden.
Cases:
- RA-01 explicit positive acceptance: PASS.
- RA-02 paraphrased positive acceptance: PASS.
- RA-03 acknowledgement only: PASS.
- RA-04 tentative response: PASS.
- RA-05 different responder without personal acceptance: PASS.
- RA-06 explicit commitment to different work: PARTIAL. The model returned no
acceptance instead of recognizing a commitment with
same_requested_workfalse. The requested action correctly remained unestablished. - RA-07 request without response: PASS.
- RA-08 collective commitment: PARTIAL. The model over-recognized the
collective
wirstatement as an explicit commitment, but no request existed and deterministic gates prevented individual responsibility. - RA-09 impersonal necessity: PARTIAL. The model over-recognized the impersonal necessity as a concrete request, but the observation had no addressee and deterministic gates prevented establishment.
- RA-10 tentative personal suggestion: PASS.
Configuration: exactly ten sequential qwen3.5:9B calls, one per case,
temperature 0, think=false, num_ctx=16384, num_predict=1024, no retries,
no voting and no prompt change between cases. Summed call time was 23.754
seconds, with 4,826 prompt-evaluation tokens and 793 evaluation tokens. The
strict schema validated every response and no responsibility or establishment
field leaked into model output.
Both positive cases recognized the request, explicit commitment and same-work
relationship, including the paraphrased acceptance, and deterministically
established Clara as responsible with due date Dienstag. The model rendered
the normalized action in semantically equivalent English; evaluation therefore
checks the structural deterministic result exactly while treating normalized
action wording as evidence-near semantic text rather than requiring lexical
identity. Acknowledgement and tentative response were not promoted. Every
negative case remained unestablished, and no individual responsibility was
invented.
Recognition-level errors were two false positives (RA-08 commitment and RA-09 request) and one false negative (RA-06 different-work commitment). Final established-action false positives and false negatives were both zero. The overall result was seven PASS, three PARTIAL and zero FAIL.
Conclusion: the narrow request-plus-acceptance architecture remains promising for established individual actions because deterministic addressee, ordering, speaker, same-work, provenance and deadline gates contained all recognition errors. The recognition layer is not yet precise enough to generalize: its handling of collective commitment, impersonal necessity and commitments to different work needs further isolated study. No production integration or additional semantic category is justified by this result.
Artifacts are preserved under
artifacts/experiments/request_acceptance_gold_v0/20260820_qwen35_9b_single_run/.
EXP-0033 — Collective Commitment Gold V0
Status: Experimental; architecturally successful with one contained recognition false positive
Date: 2026-08-20
This isolated second-stage experiment tested whether an explicit collective first-person commitment can establish an action without inventing an individual owner. It used ten synthetic cases containing one minimal V3-style observation each. Evidence Observation V3 was neither called nor changed, and the accepted Request/Acceptance mechanism remained unchanged and independent.
The strict semantic schema contains exactly observation_id,
commitment_form and normalized_action_text. commitment_form is closed to
individual_first_person, collective_first_person and none. The model
cannot output responsibility, ownership, requested actor, establishment,
Action Item, protocol, confidence, relations, graphs, decisions or unresolved
issues. Deterministic code validates schema and provenance, requires collective
commitment plus non-empty action text, applies bounded deadline consistency and
explicit-negation gates, and only then sets status: established,
commitment_scope: collective and responsible_person: null.
Gold results:
- CC-01 explicit collective commitment: PASS; established, due
nächste Woche, no person. - CC-02 individual commitment: PASS; correctly routed out of the collective path.
- CC-03 tentative collective possibility: PASS; unestablished.
- CC-04 collective suggestion: PASS; unestablished.
- CC-05 impersonal necessity: PASS; unestablished.
- CC-06 passive future statement: PASS; unestablished.
- CC-07 collective rejection: PARTIAL. The model incorrectly returned
collective_first_person, but the deterministic negation gate detectednichtand prevented establishment. - CC-08 qualified collective commitment: PASS; established with
nur im Technikumpreserved, null due and no person. - CC-09 collective commitment without deadline: PASS; established with null due and no person.
- CC-10 speaker ownership trap: PASS; established collectively while Martin remained only the speaker and was not assigned ownership.
Configuration: exactly ten successful sequential qwen3.5:9B calls, one per
case, temperature 0, think=false, num_ctx=16384, num_predict=1024, no
retries, no voting and no prompt change. There were zero technical failed
calls. Aggregate runner time was 10.504 seconds; summed per-call time was 10.500
seconds, with 4,267 prompt-evaluation tokens and 415 evaluation tokens.
The outcome was nine PASS, one PARTIAL and zero FAIL. There was one recognition
false positive and no recognition false negatives. No qualifier was lost, no
individual owner was invented, and no responsibility or status field leaked
into recognition. Bounded due handling preserved nächste Woche verbatim and
returned null when no deadline was present.
Conclusion: the collective-commitment path is architecturally successful for
this narrow Gold set. The deterministic negation gate contained the only model
error, and every successful collective result necessarily retained
responsible_person: null. This does not justify a generic commitment system,
production integration, group identity inference or another semantic category.
Artifacts are preserved under
artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/.
EXP-0034 — Explicit Rejection Gold V0
Status: Failed architecturally
Date: 2026-08-20
This isolated Stage-2 experiment tested the narrow evidence fact that a concrete action, option, proposal or future course was explicitly rejected, abandoned, discontinued or ruled out. It used twelve synthetic cases containing one self-contained observation or one local target/rejection pair. Evidence Observation V3 was not called or changed. The accepted Request/Acceptance and Collective Commitment paths remained unchanged and were not invoked.
The strict semantic schema contains exactly rejection_observation_id,
target_observation_id, rejection_form and
normalized_rejected_action_text. rejection_form is closed to
explicit_action_rejection and none. A positive recognition requires a
known local target and non-empty normalized target; none requires both target
and normalized text to be null. Decision, outcome, topic-closure,
responsibility, ownership, protocol, confidence and graph fields are forbidden.
Target resolution is limited to the same observation or one earlier supplied
observation. Deterministic code validates schema, IDs, ordering and complete
provenance before emitting the narrow status explicitly_rejected.
explicitly_rejected means rejected by the cited evidence only. It is not yet
a final meeting decision or final topic outcome, does not close a topic, and
does not supersede an earlier commitment.
Gold results:
- RJ-01 explicit collective rejection with local target: PASS.
- RJ-02 explicit non-pursuit with paired target: PASS.
- RJ-03 self-contained collaboration rejection: FAIL. The model returned
none, producing one recognition false negative. - RJ-04 personal preference: FAIL. The model promoted the preference to an explicit rejection and derived an unsupported rejection.
- RJ-05 concern: PASS; remained a non-rejection.
- RJ-06 uncertainty: PASS; remained a non-rejection.
- RJ-07 negative recommendation: FAIL. The model promoted advice to an explicit rejection and derived an unsupported rejection.
- RJ-08 deferral: PASS; remained a non-rejection.
- RJ-09 factual negation: PASS; remained a non-rejection.
- RJ-10 temporary non-action: FAIL. The model treated
erstmal noch nichtas abandonment and derived an unsupported rejection. - RJ-11 explicit rejection with material scope: PASS. Real-plant and Druckversuch scope were preserved.
- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option was rejected; the Technikum alternative was not absorbed.
Configuration: exactly twelve successful sequential qwen3.5:9B calls, one
per case, temperature 0, think=false, num_ctx=16384,
num_predict=1024, no retries, no voting and no prompt changes. There were zero
technical failed calls. Aggregate runner time was 15.518 seconds; summed
per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681
evaluation tokens.
The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03). There were four strict target-field expectation mismatches: three were consequences of false-positive rejection objects populating otherwise locally correct antecedents, and one was the missing self-contained RJ-03 target. No derived positive selected the wrong concrete antecedent. Qualifier-loss count was zero, positive-alternative absorption count was zero, and no responsibility, decision, outcome or topic-closure field leaked into model output.
Conclusion: the experiment is not architecturally successful. Deterministic
structural gates cannot contain a semantically well-formed false-positive
rejection with valid local target and provenance. The model did distinguish
concern, uncertainty, deferral and factual negation, and it handled scoped and
alternative-bearing positives correctly, but it did not reliably separate
explicit rejection from personal preference, advice or temporary non-action.
The current binary recognition explicit_action_rejection | none is
insufficient for reliable generalization.
No production integration, generic rejection system, prompt tuning or
cross-pattern reconciliation is justified.
Artifacts are preserved under
artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/.
EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
Date: 2026-08-11
Hypothesis: the primary protocol should be a topic-oriented reconstruction of the meeting rather than a category-oriented list of extracted information.
This first isolated prototype does not replace or connect to the production
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
transcript excerpts to qwen3.5:9B and requests Discussion Subjects. Each
subject may contain supported discourse events, an outcome with mandatory
scope, resulting actions and unresolved issues. Optional structures must be
omitted when absent. Every semantic object must reference known evidence IDs.
The strict experimental schema validates:
- non-empty subjects and globally unique semantic identifiers;
- a closed discourse-event vocabulary;
- non-empty, known and non-duplicated evidence references;
- outcome text, scope, certainty and evidence;
- action text, JSON-nullable responsibility/deadline and evidence;
- unresolved-issue text and evidence;
- omission rather than null or empty optional structures.
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only, multiple options, unaccepted proposal, proposal with objection, rejected alternative, trial-scoped acceptance, no-decision discussion, resulting Action Item, and outcome plus unresolved issue. Evaluation targets semantic identity, development, outcome scope, actions, unresolved issues, traceability and absence of invented commitments rather than exact wording.
Configuration: qwen3.5:9B, temperature 0, think=false, num_ctx=16384,
num_predict=4096. Each case received exactly one model call; there were no
model retries or prompt iterations. The nine completed calls took 59.251
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
failure artifacts were preserved under
/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/ and
/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/. Two earlier launch
attempts made zero LLM calls: one failed on the script import path and one was
blocked by sandbox networking.
Human-reviewed results after correcting two objectively wrong Gold assumptions without another model call:
| Case | Verdict | Reason |
|---|---|---|
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled considered_option rather than introduced_idea. |
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case F: the model distinguished acceptance for a bounded trial from acceptance as a final solution. It also handled the explicit action in case H well. However, the experiment failed systematically on sparse structured output, subject grouping and restraint around absent outcomes/actions/unresolved issues. The model frequently mirrored optional schema fields as empty/null values, treated alternatives as outcomes, split one discussion into multiple subjects, or invented open issues from mere non-selection.
The focused experiment is not promising enough to justify a real Progeo chunk sanity check. No such run was performed, and no architecture is accepted on the basis of this prototype. Further work should first analyze whether the failure comes from the schema/prompt representation, the model's sparse-output reliability, or the boundary between subject grouping and semantic synthesis. It should not proceed through repeated prompt tuning against these nine cases.