Files
meeting-lab/docs/experiments.md
T

1886 lines
73 KiB
Markdown

# Experiments
This file records durable technical experiments and findings for Meeting Lab.
It is not a diary and does not replace commit history.
## Status values
- Proposed: experiment idea exists, but no result is recorded.
- Running: experiment is in progress and no decision has been made.
- Accepted: finding is the current baseline or design conclusion.
- Rejected: hypothesis was tested and should not be repeated as-is.
- Superseded: finding was useful but has been replaced by a newer baseline.
## EXP-0001 - Whisper JSON interpretation
Status: Accepted
Date or period: 2026-07-29
Hypothesis:
Whisper JSON should be chunked from its segment stream, not from the aggregate
top-level text field.
Setup:
`chunk_transcript.py` was updated to parse JSON input and prefer
`segments[*].text` when `segments` exists. A regression test supplies JSON with
both top-level `text` and separate segment texts.
Inputs:
- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`.
- Real Whisper artifacts under `samples/whisper/`.
Model / configuration:
- No LLM.
Result:
The test verifies that the block stream is `["alpha", "beta", "gamma"]` and
does not include the aggregate `"alpha beta gamma"` text. The repository history
records this as the fix for the earlier failure where the first chunk contained
the complete transcript.
Decision:
When `segments` exists, `segments[*].text` is the authoritative transcript
stream. The top-level `text` field is only a fallback.
Lessons learned:
Whisper JSON is structured input. Treating it like plain text can duplicate the
entire transcript and invalidate downstream chunking.
Evidence:
- `src/meeting_lab/chunking/chunk_transcript.py`
- `tests/test_chunking.py`
- Commit `4656523` - `Fix Whisper JSON chunk extraction`
## EXP-0002 - Technical transcript chunking baseline
Status: Accepted
Date or period: 2026-07-29 to 2026-07-30
Hypothesis:
Sequential technical chunks around the configured target size can preserve the
transcript while keeping extraction calls small enough for local models.
Setup:
The chunker splits block-aligned text with configurable target, minimum,
maximum and overlap settings. Tests verify no duplicate later blocks when
overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
Inputs:
- Synthetic block list in `tests/test_chunking.py`.
- `samples/whisper/meeting_speech_cleaned.json`.
Model / configuration:
- No LLM for chunking.
- Manifest uses default chunking behavior recorded in
`samples/whisper/meeting_speech_cleaned_chunks/manifest.json`.
Result:
With overlap set to zero, tests verify that all blocks appear exactly once. The
real sample manifest contains nine chunks, mostly near the configured target
size, with a smaller final chunk.
Decision:
Independent sequential chunks are the current technical baseline. One
normalized chunk per extraction call is the preferred extraction strategy.
Lessons learned:
Chunking solves model-size constraints only. It must not perform topic
detection or semantic merging.
Evidence:
- `src/meeting_lab/chunking/chunk_transcript.py`
- `tests/test_chunking.py`
- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`
- `AGENTS.md`
- `PROJECT_KNOWLEDGE.md`
## EXP-0003 - Conservative transcript normalization
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Transcript cleanup should improve readability without changing meeting
semantics.
Setup:
The normalizer removes isolated filler sounds, immediate duplicate words or
short duplicate phrases, and redundant whitespace. It records changed blocks in
a JSON change log and explicitly preserves semantic content categories.
Inputs:
- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`.
- Change logs such as `chunk_01_changes.json`.
Model / configuration:
- No LLM.
Result:
The implementation and generated change logs show a conservative policy:
negations, qualifiers, dates, numbers, responsibilities, technical statements,
deadlines, decisions and commitments are preserved.
Decision:
Normalization remains deterministic and low-risk. When uncertain, leave text
unchanged.
Lessons learned:
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove
semantic cues needed by extraction.
Evidence:
- `src/meeting_lab/normalization/normalize_transcript.py`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json`
- `docs/pipeline.md`
## EXP-0004 - Full-context topic segmentation
Status: Superseded
Date or period: 2026-07-21 to 2026-07-22
Hypothesis:
A single full-context topic segmentation call can identify topic boundaries in
a normalized transcript chunk.
Setup:
The initial segmentation prototype asked the model for topic changes and then
converted those boundaries into continuous, non-overlapping segments.
Inputs:
- `samples/chunks/chunk_01_normalized.txt`.
Model / configuration:
- Generated artifact records `qwen3:14b`.
Result:
The generated artifact contains 85 blocks, three topic-change boundaries and
four segments. The run metadata records a substantially longer elapsed time
than the later windowed artifact for the same input.
Decision:
Full-context segmentation was useful as a prototype, but it was superseded by
windowed segmentation and manual review tooling.
Lessons learned:
The prototype established the boundary-to-segment representation, but did not
settle segmentation quality.
Evidence:
- `src/meeting_lab/segmentation/segment_topics.py`
- `samples/chunks/chunk_01_normalized_segments.json`
- Commit `f234efc` - `Add initial topic segmentation prototype`
- Commit `889a4fe` - `Detect topic boundaries as continuous segments`
## EXP-0005 - Windowed topic segmentation and review
Status: Accepted
Date or period: 2026-07-22
Hypothesis:
Windowed topic segmentation can reduce runtime and make boundary evaluation
more inspectable than a single full-context call.
Setup:
`segment_topics_windowed.py` analyzes overlapping windows and reports only
boundaries from the decision range. Python merges boundaries into continuous,
non-overlapping segments. `review_segmentation.py` renders each boundary with
neighboring transcript context for manual classification.
Inputs:
- `samples/chunks/chunk_01_normalized.txt`.
- Full meeting normalized chunks under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3.
- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3.
Result:
The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
Manual review output shows that some boundaries were assessed as subtopics
rather than full topic changes. Full-meeting artifacts show one window per
already-small normalized chunk and two segments per chunk.
Decision:
Windowed segmentation and review tooling are accepted as prototype tooling, not
as a stable production segmentation stage.
Lessons learned:
Windowing improves inspectability and can reduce runtime, but it can also
cluster boundaries and over-segment. Manual review remains necessary.
Evidence:
- `src/meeting_lab/segmentation/segment_topics_windowed.py`
- `src/meeting_lab/segmentation/review_segmentation.py`
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
- `samples/chunks/chunk_01_normalized_windowed_segments_review.md`
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling`
## EXP-0006 - Qwen model comparison
Status: Accepted
Hypothesis:
Larger local Qwen-family models should improve meaningful extraction and
segmentation, but model size alone will not solve prompt or pipeline problems.
Setup:
Project work used smaller models for smoke checks and larger local models for
meaningful extraction or segmentation. Artifacts and project knowledge record
the currently useful model roles.
Inputs:
- Gold Standard scenarios under `tests/gold/`.
- Generated segmentation artifacts under `samples/`.
- Generated extraction artifacts under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- `qwen3:1.7b`: smoke-test model according to project knowledge.
- `qwen3.5:9b`: current meaningful extraction and segmentation model according
to project knowledge and generated full-meeting segmentation artifacts.
- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts.
Result:
The repository supports the conclusion that `qwen3.5:9b` is the meaningful
current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger
models and longer contexts may increase runtime substantially, but no hardware
benchmark suite is recorded.
Decision:
Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current
experiments. Do not assume model size alone fixes prompt or pipeline design.
Lessons learned:
Evaluation must separate model capability from prompt clarity, context
strategy, extraction schema and consolidation.
Evidence:
- `PROJECT_KNOWLEDGE.md`
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
- `samples/chunks/chunk_01_normalized_segments.json`
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
## EXP-0007 - Thinking output and Ollama API behavior
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Thinking-capable Qwen models may return reasoning separately from the final
answer, and extraction parsing should not fail merely because extra text or
multiple JSON objects appear.
Setup:
The Ollama response reader checks `response`, chat-style `message.content` and
then `thinking`. The JSON parser tries a full parse first, then scans JSON
object candidates and returns the final valid object. A regression test covers
thinking text before final JSON.
Inputs:
- Synthetic parser test in `tests/test_extraction_protocol.py`.
Model / configuration:
- No LLM run in the test.
- Code path is used by Ollama extraction.
Result:
The parser can handle additional text and multiple JSON objects where the final
valid object is the intended answer. Current code still falls back to `thinking`
only if no usable response or message content is present.
Decision:
Keep parser robustness, but do not treat thinking output as the root cause of
all extraction failures.
Lessons learned:
API response shape and model output shape are separate concerns. Preserve raw
responses when diagnosing failures.
Evidence:
- `src/meeting_lab/extraction/extract_chunks.py`
- `tests/test_extraction_protocol.py`
- `PROJECT_KNOWLEDGE.md`
## EXP-0008 - JSON truncation and generation limits
Status: Accepted
Hypothesis:
Some extraction failures are caused by generation limits truncating JSON rather
than by prompt wording or parser behavior.
Setup:
A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The
generation limit was increased for the successful path. Exact failing limit is
not recorded in the repository; the current extractor default is verifiably
`--num-predict 8192`.
Inputs:
- Local extraction runs referenced by project knowledge.
- Current extraction CLI.
Model / configuration:
- `qwen3.5:9b`.
- Current extractor default: `num_predict=8192`.
Result:
Increasing the generation limit fixed the technical JSON failure. This was not
primarily a parser or prompt problem.
Decision:
When JSON is truncated, inspect raw output and generation limits before editing
prompts.
Lessons learned:
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong
first response if the model simply ran out of output tokens.
Evidence:
- `src/meeting_lab/extraction/extract_chunks.py`
- `PROJECT_KNOWLEDGE.md`
- `AGENTS.md`
## EXP-0009 - Minimal end-to-end pipeline
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
A minimal local pipeline can transform Whisper output into chunk extractions
and an interim protocol, proving the technical path before the final
architecture exists.
Setup:
The repository added cleanup, normalization, chunking, extraction and protocol
builder scripts, with sample generated artifacts.
Inputs:
- `samples/whisper/meeting_speech.json`
- `samples/whisper/meeting_speech_cleaned.json`
- Generated chunks and normalized chunks under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- Local Ollama extraction for chunk JSON.
- Windowed segmentation artifacts use `qwen3.5:9b`.
Result:
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
extraction JSON files, windowed segmentation artifacts and
`meeting_protocol.md`.
Decision:
The minimal pipeline is technically validated. The first protocol builder is an
interim validation tool, not the final architecture.
Lessons learned:
End-to-end execution exposed the next limitation: extraction output needs
consolidation and purpose-specific rendering.
Evidence:
- `scripts/clean_whisper_json.py`
- `src/meeting_lab/normalization/normalize_transcript.py`
- `src/meeting_lab/chunking/chunk_transcript.py`
- `src/meeting_lab/extraction/extract_chunks.py`
- `src/meeting_lab/protocol/build_protocol.py`
- `samples/whisper/meeting_speech_cleaned_chunks/`
- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline`
## EXP-0010 - Gold Standard corpus
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Reproducible prompt engineering requires synthetic transcripts with explicit
expected semantic outputs.
Setup:
The Gold Standard corpus defines scenario directories with `transcript.txt`,
`expected.json` and README files describing ground truth and common model
mistakes. The runner validates schema keys and writes `actual.json` for a
scenario.
Inputs:
- Gold scenarios under `tests/gold/`.
Model / configuration:
- Runner requires an explicit Ollama model for LLM evaluation.
- Existing unit tests for the runner do not invoke Ollama.
Result:
The corpus gives stable semantics for decisions, facts, positions, todos,
questions, technical details and difficult mixed cases. Initial structured
transcripts are Phase 1 and easier than raw Whisper-style transcripts.
Decision:
Use Gold Standard tests as both regression tests and formal meeting-semantics
specification. Raw or unlabelled transcript cases remain later-phase work.
Lessons learned:
Without expected outputs, prompt changes cannot be evaluated reproducibly.
Evidence:
- `tests/gold/`
- `scripts/run_gold_test.py`
- `tests/test_gold_runner.py`
- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests`
## EXP-0011 - Gold-test quality and unique ground truth
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
If a prompt produces unexpected behavior, the gold test itself may be ambiguous
and should be reviewed before the prompt is changed.
Setup:
Decision-focused scenarios were clarified during baseline creation. The current
methodology requires checking unique ground truth before changing prompts.
Inputs:
- `decision_simple`
- `decision_deferred`
- `decision_none`
- Gold methodology document.
Model / configuration:
- Prompt Version 2 baseline work.
Result:
`decision_simple` required clarification around the explicit agreement and
nearby non-decision wording. The earlier negative/deferral ambiguity is now
represented by distinct `decision_none` and `decision_deferred` scenarios in
the repository. Punctuation is not reliable evidence for Whisper transcripts;
agreement language and wording must carry the semantics.
Decision:
Ambiguous gold tests must be reviewed before prompt changes. Do not treat
punctuation as reliable evidence in real Whisper-style transcripts.
Lessons learned:
Bad gold tests create false prompt failures and can encourage overfitting.
Evidence:
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
- `tests/gold/decision_simple/README.md`
- `tests/gold/decision_deferred/README.md`
- `tests/gold/decision_none/README.md`
- Commit `f7ad9ba`
## EXP-0012 - Decision taxonomy
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Decision extraction needs a formal taxonomy that distinguishes substantive
decisions from process decisions and non-decisions.
Setup:
The decision definition document and decision prompt define included and
excluded categories. Gold tests cover explicit decisions, true no-decision
cases and deferrals.
Inputs:
- `tests/gold/DECISION_DEFINITION.md`
- `prompts/decisions.md`
- Decision gold scenarios.
Model / configuration:
- Prompt Version 2 baseline.
Result:
Accepted decision categories include substantive decisions, organizational
decisions, process decisions, approvals, rejections, deferrals, explicit
decisions not to decide yet and explicit agreement to gather more information
before deciding. Opinions, preferences and proposals without agreement are not
decisions.
Decision:
"No decision was reached" and "the decision was deferred" are distinct semantic
outcomes.
Lessons learned:
Deferral can be a valid process decision even when the substantive topic remains
unresolved.
Evidence:
- `tests/gold/DECISION_DEFINITION.md`
- `tests/gold/decision_deferred/`
- `tests/gold/decision_none/`
- `prompts/decisions.md`
## EXP-0013 - Prompt engineering methodology
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Prompt iteration needs strict experimental controls to prevent regression,
overfitting and arbitrary prompt churn.
Setup:
The methodology was documented alongside the Gold Standard corpus and later
summarized for agents.
Inputs:
- Gold scenarios.
- Prompt files.
Model / configuration:
- Applies to all prompt experiments.
Result:
The accepted method is one prompt change per iteration, one target test at a
time, immediate validation, no regressions, no `expected.json` edits merely to
force a pass, no test-specific prompt hacks, stopping after two consecutive
non-improving iterations, and verifying unique ground truth before prompt
changes.
Decision:
Prompt changes are controlled experiments. See `AGENTS.md` for agent operating
rules.
Lessons learned:
Most prompt changes are not isolated unless the experiment explicitly constrains
the target and regression set.
Evidence:
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
- `AGENTS.md`
- Commit `f7ad9ba`
## EXP-0014 - Decision Prompt Version 2
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Adding explicit process-decision language to the decision prompt can preserve
true decision detection while recognizing deferrals.
Setup:
Prompt Version 2 added explicit support for deferrals and process decisions.
The baseline was validated on three decision scenarios.
Inputs:
- `decision_simple`
- `decision_deferred`
- `decision_none`
Model / configuration:
- Prompt Version 2.
- Model used for validation is not recorded in the committed methodology.
Result:
The committed methodology records all three baseline scenarios as passing. The
current generated `actual.json` files also show the expected decision count for
these decision scenarios, although some non-decision categories remain less
complete.
Decision:
Prompt Version 2 is the current decision baseline. Explicit deferrals are
recognized as process decisions.
Lessons learned:
Decision-count success does not imply all categories are solved. Category-level
evaluation must continue.
Evidence:
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
- `tests/gold/decision_simple/actual.json`
- `tests/gold/decision_deferred/actual.json`
- `tests/gold/decision_none/actual.json`
- `prompts/decisions.md`
## EXP-0015 - Difficult synthetic meeting
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
A deliberately adversarial synthetic meeting can expose extraction failures
that simple category tests miss.
Setup:
`evil_meeting` includes interruptions, corrections, absent referenced people,
near-decisions, changed positions and one expected explicit decision.
Inputs:
- `tests/gold/evil_meeting/`.
Model / configuration:
- Existing generated `actual.json`; exact model is not stored in the artifact.
Result:
The generated result found the expected FR-7 exclusion decision. It also
classified "do not migrate until the mapping table is checked" as an additional
decision. The current expected file treats that statement as a position, but
contextual review suggests it may be a valid process instruction or decision.
Decision:
Do not classify this as a simple model failure without reviewing the gold
standard. The scenario exposes a semantic gap in the expected output.
Lessons learned:
Difficult synthetic cases are valuable because they reveal ambiguity in the
specification as well as model mistakes.
Evidence:
- `tests/gold/evil_meeting/README.md`
- `tests/gold/evil_meeting/expected.json`
- `tests/gold/evil_meeting/actual.json`
- See EXP-0011 and EXP-0012.
## EXP-0016 - Context-size extraction comparisons
Status: Rejected
Hypothesis:
Increasing extraction context from one chunk to neighboring chunk groups should
make extraction more complete and therefore should become the baseline.
Setup:
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3
and 1+2+3.
Inputs:
- Normalized meeting chunks.
Model / configuration:
- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for
extraction experiments.
Result:
More context sometimes improved completeness, but it also shifted category
classification, added duplicates and reduced stability. No numeric winner is
recorded in the repository.
Decision:
Do not adopt larger extraction windows as the baseline. Independent chunk
extraction remains current strategy. Recover global context through
consolidation rather than continuously enlarging extraction windows.
Lessons learned:
Context size is not a monotonic quality knob. It changes the task the model is
performing.
Evidence:
- `PROJECT_KNOWLEDGE.md`
- `AGENTS.md`
- `ROADMAP.md`
- See EXP-0002 and EXP-0018.
## EXP-0017 - Independent full-meeting chunk extraction
Status: Accepted
Hypothesis:
Extracting every normalized chunk independently can produce enough structured
material for a useful protocol draft.
Setup:
Nine normalized chunks were extracted into separate JSON files and then
rendered by the interim protocol builder.
Inputs:
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt`
through `chunk_09_normalized.txt`.
Model / configuration:
- Existing extraction artifacts do not record model metadata.
Result:
The nine extraction files contain facts, decisions, todos, questions and
technical details. `meeting_protocol.md` aggregates them into a readable draft.
Duplicates, category shifts and synthesis became the dominant limitations.
Decision:
Independent chunk extraction is useful enough to keep as the baseline, but it
requires a consolidation stage.
Lessons learned:
Per-chunk extraction gives recall-oriented raw material. It does not by itself
produce a polished or canonical meeting representation.
Evidence:
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
- `src/meeting_lab/protocol/build_protocol.py`
- See EXP-0018.
## EXP-0018 - Human protocol comparison and output-view split
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
One generated protocol cannot satisfy every use case; protocol output should be
separated by purpose and audience.
Setup:
The interim machine protocol was compared against the desired human protocol
shape and then the architecture was revised toward parallel output views.
Inputs:
- Interim `meeting_protocol.md`.
- Architecture and output-view documentation.
Model / configuration:
- Not applicable; this is a design evaluation.
Result:
The human protocol target is denser and organized by purpose and topic rather
than extraction categories. The machine extraction retains more context and is
useful for recall, but it is not the right direct source for a concise
distribution artifact or durable knowledge entry.
Decision:
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution
Protocol / Verteilerprotokoll and Knowledge Objects /
Wissensdatenbankeintrag.
Lessons learned:
Rendering is a separate concern from extraction and consolidation. Output views
must be parallel renderings of shared semantics, not transformations of one
another.
Evidence:
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
- `docs/output-views.md`
- `docs/architecture.md`
- `docs/pipeline.md`
- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture`
## EXP-0019 - Consolidation and Canonical Meeting Knowledge
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Extraction, consolidation and output rendering are separate problems and should
not be collapsed into one LLM prompt or one protocol file.
Setup:
The architecture was refined after the minimal pipeline and protocol draft
showed duplicate, synthesis and audience-specific rendering limitations.
Inputs:
- Extraction JSON artifacts.
- Interim protocol draft.
- Architecture and data-model documentation.
Model / configuration:
- Not applicable; this is an architectural conclusion.
Result:
The accepted design is a planned Canonical Meeting Knowledge layer as the
semantic source of truth, with Working Protocol, Distribution Protocol and
Knowledge Objects as parallel output views. The next consolidation architecture
is split into a Deterministic Canonicalizer and a Semantic Consolidator. The
canonicalizer prepares validated evidence-bearing objects without uncertain
semantic merging. The consolidator then merges semantically equivalent
statements, preserves evidence, reconciles category shifts where supported and
marks contradictions or uncertainty.
Decision:
Deterministic canonicalization is implemented as Canonicalizer V1. The first
semantic consolidation milestone is implemented as Semantic Consolidator V0 for
facts-only duplicate detection. Broader semantic consolidation, Canonical
Meeting Knowledge and final output views remain planned.
Lessons learned:
Global meeting understanding should be recovered by consolidation over
evidence-bearing extractions, not by silently changing output views or expanding
LLM context indefinitely.
Evidence:
- `docs/architecture.md`
- `docs/pipeline.md`
- `docs/data-models.md`
- `docs/output-views.md`
- `PROJECT_KNOWLEDGE.md`
- `ROADMAP.md`
- Commit `5c03ed7`
## EXP-0020 - Working Protocol Synthesizer V0
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The current local synthesis model may be able to generate a useful detailed
Working Protocol directly from the existing independent chunk extraction JSON
files, before canonicalization or semantic consolidation exists.
Setup:
One synthesis prompt was constructed from exactly nine chunk extraction JSON
files. The model was instructed to use only those extraction files, merge
duplicates, group related information into topics, preserve useful discussion
context and write a neutral technical Working Protocol.
Inputs:
- `chunk_01_extraction.json` through `chunk_09_extraction.json`.
- No original transcript, normalized chunks or Whisper output were used as
synthesis input.
Model / configuration:
- Model: `qwen3.5:9b`
- Prompt characters: 24,979
- Actual prompt eval tokens: 5,929
- Output tokens: 1,486
- Runtime: 294.204 seconds
Result:
The generated Working Protocol was readable, well structured and
topic-oriented. It was still based directly on raw chunk extractions, without a
separate deterministic canonicalization stage or semantic consolidation stage.
The output language was English even though the source meeting material was
German.
Decision:
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline
for later canonicalizer, consolidator and renderer comparisons. This selected
generated artifact is intentionally versioned even though generated runtime
artifacts are normally ignored.
Lessons learned:
Direct synthesis from chunk extractions can create a useful recall-oriented
draft, but it does not replace Canonical Meeting Knowledge. The language
mismatch also establishes a default renderer rule: protocol output should
normally match the dominant source language unless an explicit output language
is requested.
Evidence:
- `samples/benchmarks/working_protocol_synthesizer_v0/README.md`
- `samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md`
- `PROJECT_KNOWLEDGE.md`
- `docs/output-views.md`
- See EXP-0017 and EXP-0019.
## EXP-0021 - Canonicalizer V1
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
Independent chunk extraction JSON can be converted into a stable deterministic
intermediate format before any semantic LLM consolidation is attempted.
Setup:
Canonicalizer V1 discovers `chunk_*_extraction.json` files in stable chunk
order, validates required categories, normalizes category names and basic field
structure, parses existing legacy string formats where safe, trims redundant
whitespace, assigns deterministic IDs, preserves original values and source
references, and merges only exact duplicates when all semantic fields are
identical.
Inputs:
- Synthetic unit-test fixtures.
- Existing nine extraction JSON files under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- No LLM.
- CLI module: `meeting_lab.consolidation.canonicalize`.
Result:
Canonicalizer V1 produces `schema_version`, `source_files`, `stats` and
`items`. It is deterministic preparation for the future Semantic Consolidator
and is not Canonical Meeting Knowledge.
Decision:
Canonicalizer V1 is the current implemented deterministic canonicalization
stage. Semantic Consolidator V0 now uses this representation for facts-only
semantic duplicate detection; broader semantic consolidation and Canonical
Meeting Knowledge remain planned.
Lessons learned:
Exact duplicate handling, source-reference preservation and legacy string
parsing can be tested without model calls. Any uncertain semantic merge remains
out of scope for this stage.
Evidence:
- `src/meeting_lab/consolidation/canonicalize.py`
- `tests/test_canonicalize.py`
- `docs/data-models.md`
- `docs/pipeline.md`
## EXP-0022 - Semantic Consolidator V0 facts-only merge
Status: Accepted
Date or period: 2026-07-31
Hypothesis:
The Canonicalizer V1 output contains enough stable structure for a local LLM to
identify semantically equivalent fact items without losing source coverage or
changing non-fact categories.
Setup:
Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON,
selected only items with `category: "fact"`, and sent one bounded consolidation
request to local Ollama. The merge rules required semantic equivalence, not
topic similarity, and validation required every source fact ID to appear
exactly once.
Inputs:
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
- 33 fact items.
Model / configuration:
- `qwen3.5:9B`
- Ollama endpoint: `http://127.0.0.1:11434/api/generate`
- Thinking disabled.
- One LLM call.
- `num_ctx=32768`
- `num_predict=4096`
Result:
- Runtime: 390.119 seconds on the current machine.
- Merged fact groups: 1.
- Source facts involved in merges: 2.
- Singleton fact groups: 31.
- Validation: passed.
- No source fact was lost or duplicated.
- Non-fact categories remained unchanged.
Accepted merge:
- `fact_0025` + `fact_0031`
- Canonical statement: "Der Leiter F&E führt die Projektliste auf dem
zweiwöchentlichen Schnittstellen-Stand-Up."
Decision:
Semantic Consolidator V0 is complete for its current narrow scope:
conservative facts-only semantic duplicate detection with source evidence
preserved. The selected `report.md` and `consolidated_extractions.json`
benchmark artifacts should be versioned for later comparison. The raw model
response remains a local diagnostic artifact and is not versioned.
Lessons learned:
Semantic duplicate consolidation is technically viable and conservative enough
for continued evaluation, but broader semantic synthesis remains a separate
future stage. The measured runtime is useful for this machine and run, but
should not be generalized into a universal benchmark.
Evidence:
- `src/meeting_lab/consolidation/consolidate_facts.py`
- `prompts/consolidate_facts.md`
- `tests/test_consolidate_facts.py`
- `samples/benchmarks/semantic_consolidator_v0/report.md`
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
- Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt`
## EXP-0023 - Responsibility attribution integrity
Status: Running
Date or period: 2026-07-31
Hypothesis:
Protocol generation is operationally unsafe if responsibility, ownership or
departmental role attribution is inferred from discussion context rather than
explicit meeting evidence.
Setup:
The real-life Working Protocol Renderer V2 benchmark was inspected against the
consolidated input. A false assignment connected a Marketing participant to
Business Development criteria work even though the participant's contribution
was critical or reluctant and did not establish acceptance of that task.
Inputs:
- `samples/benchmarks/working_protocol_renderer_v2/working_protocol.md`
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
Model / configuration:
- Not rerun for this finding.
- Finding is based on existing benchmark artifacts.
Result:
The false responsibility attribution is visible in the structured input before
rendering, so the issue is not merely stylistic renderer wording. The root
cause may originate earlier in extraction and then be preserved by
canonicalization and consolidation. Renderer guardrails are still required so
output views do not strengthen ambiguous ownership.
Decision:
Responsibility attribution is now treated as a critical project-wide
invariant. A person, team or department may be recorded as responsible only
when the evidence explicitly assigns, accepts or confirms that responsibility.
Discussion, expertise, objection, suggestion, thematic proximity, speaker
adjacency, organizational assumptions and likely job roles do not establish
ownership.
Lessons learned:
This class of error affects operational correctness, not only style. The
pipeline needs traceable attribution evidence and future schema support for
responsibility status such as explicit, accepted, proposed or unclear.
Evidence:
- `AGENTS.md`
- `PROJECT_KNOWLEDGE.md`
- `docs/data-models.md`
- `docs/output-views.md`
- `prompts/working_protocol.md`
- `tests/gold/responsibility_attribution_negative/`
## EXP-0024 - Working Protocol V2 contract visibility
Status: Running
Date or period: 2026-08-09
Target:
BUG-011 renderer-only regression using an already validated Semantic
Consolidator artifact.
Hypothesis:
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
placing the leading prompt contract outside the model's effective evaluated
context. The preserved failures all report `prompt_eval_count=16386`, while
their outputs either echo trailing JSON or produce an unconstrained generic
category summary instead of the requested Working Protocol.
Iteration 1 change:
- Project every consolidated item to rendering-relevant semantic fields while
retaining every item and its category/text/responsibility/deadline/status
information.
- Generate the exact structural contract from renderer validator constants and
append it after the compact INPUT JSON.
- Replace the independently handwritten prompt skeleton with a reference to
that authoritative appended contract.
- Enforce the existing prompt rule that emitted sections must not be empty.
This is one renderer-contract prompt iteration. It does not change extraction,
canonicalization, semantic consolidation or responsibility semantics.
Validation before LLM run:
- 15 focused renderer tests pass.
- Tests cover contract generation, compact input projection, valid and invalid
headings, missing topic sections, empty sections, wrapper cleanup, malformed
Markdown and final-file write gating.
Decision:
Iteration 1 passed structural validation and wrote `working_protocol.md`, but
the quality sanity check found that the model omitted most of the ten supplied
decisions, two open questions and several action items. The structurally valid
result therefore was not accepted as BUG-011 verification.
Iteration 2 change:
- Add input-derived hidden coverage markers for every decision, action item and
open question.
- Require every priority item exactly once in its matching section.
- Validate missing, duplicate, unknown and wrong-section markers
deterministically.
- Keep facts and technical details condensable as background.
This is the second single prompt iteration. It responds to the concrete
omission failure observed in Iteration 1 without changing upstream semantics or
inventing renderer content.
Iteration 2 pre-run validation:
- 18 focused renderer tests pass, including exact required-item coverage and
wrong-section rejection.
Decision:
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget
after emitting all decisions and most action items. Adaptive renderer budgeting
resolved to 8,192 tokens for this input. The final run stopped normally after
3,709 evaluated output tokens.
The final renderer-only regression passed strict validation and wrote
`working_protocol.md`. Exact coverage was 10/10 decisions, 36/36 renderable
action items and 23/23 open questions, each once in its matching section. One
structurally empty action item whose task, responsible, deadline and evidence
were all null was recorded and excluded rather than fabricated. Optional
background markers were accepted only when they referred to real projected
input items.
Accept the compact renderer input, validator-derived trailing contract,
priority-item coverage markers, empty-section validation and adaptive renderer
output sizing as the BUG-011 baseline. This establishes structural reliability
and priority-item coverage, not complete protocol prose quality.
Evidence:
- `prompts/working_protocol.md`
- `src/meeting_lab/protocol/render_working_protocol.py`
- `tests/test_render_working_protocol.py`
- `samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/`
- `samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/`
## EXP-0025 — BUG-015 classification precision
Date: 2026-08-09
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
Model/configuration: `qwen3.5:9B`, temperature 0, `num_ctx=32768`.
Tests were created before prompt changes. A Decision-only evidence threshold
kept the explicit Dr. Schlummer rejection and omitted an option and preference.
Adding Action and Open Question definitions improved several negatives but was
not stable: the model alternately promoted an unaccepted Textor suggestion or
moved rejected candidates into Open Questions. Moving the standalone category
prompts after the transcript made the partial Decision schema dominate and
misclassified true Action Items as Decisions in two consecutive runs.
The final iteration replaced the competing standalone category prompts with a
single unified classification contract after the transcript. It preserved the
assigned Nina task and ownerless established CET work, and prevented
cross-category leakage in the focused case, but still emitted the unaccepted
Textor suggestion as an Action Item. Further prompt iterations were stopped in
accordance with the Gold Standard methodology.
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
remains Open; no phrase-specific deterministic filter was introduced.
## EXP-0027 — Evidence-near observation extraction
Date: 2026-08-18
Hypothesis: `qwen3.5:9B` can more reliably extract evidence-near linguistic and
semantic properties than directly synthesize protocol-level events, outcomes,
actions and unresolved issues. This isolated experiment stops before semantic
interpretation and does not connect to the production pipeline.
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
source evidence, explicit targets, a five-value relation vocabulary, modality,
temporality, evaluation, agreement, responsibility/person, uncertainty,
clarification need and free-text scope. It contains no protocol-level category
field. The validator requires sequential observation IDs, known evidence IDs,
backward-only valid observation targets, closed categorical vocabularies,
consistent responsibility/person pairs, and one canonical absence form: JSON
null for person and `absent` for scope.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries. All nine cases ran exactly once, for nine LLM
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
responses, parsed observations, validation results, Ollama metadata and
comparisons are preserved under
`/tmp/meeting-lab-evidence-observations-v1-20260818/`.
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
Five cases failed structure because the model represented a single target as a
one-element list, usually `["discussion_subject"]`; the accepted schema permits
a list only for two or more jointly referenced observations. Several responses
also copied the relation label `limits_scope` into the free-text scope field.
These were systematic model-output errors, not transport or parser failures.
The prompt and run were not retried or tuned.
Human semantic review of the preserved raw responses:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override
strict structural failures; it separates useful semantic signal from schema
compliance.
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving
closer to evidence reduced some direct promotion behavior: the geometry mention
did not become work, the washing disadvantage did not become an unresolved
issue, both alternatives in B remained uncommitted, and the publication query
did not become an assignment. However, the important promotion errors did not
disappear. C acquired unsupported acceptance, and G still promoted a personal
preference into rejection. Positive cases were only partly preserved: explicit
rejection was recognized but incorrectly linked; trial-only language was kept
but responsibility was invented; Nina's request and commitment were recognized
but responsibility states were wrong; and the publication issue was recognized
but its uncertainty was contradicted by rejection.
Result: **B — evidence-near extraction is promising, but specific observation
dimensions remain unreliable.** Target/relation selection, scope attachment,
responsibility state/person attribution, and agreement versus uncertainty are
not reliable enough to justify designing the later interpretation stage yet.
No production integration or later interpretation stage was implemented.
## EXP-0028 — Evidence-Near Observation Extraction V2
Date: 2026-08-19
V2 tested whether `qwen3.5:9B` preserves the evidence needed by a later
controlled interpretation stage when direct responsibility, agreement and
semantic graph relations are removed. Responsibility was replaced by explicit
participant/discourse facts (`speaker`, `named_person`, `addressee`, singular
self-reference, collective `we`, and impersonal person reference). Agreement
was replaced by explicit affirmation, explicit negation and determination
statement signals. Graph relations were reduced to nullable scalar
`refers_to`; scope became free-text `qualifier` plus nullable scalar
`limits_target`. No later derivation stage was implemented.
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
human interpretations. It contains no responsibility, agreement, action,
decision, open-question, accepted-trial, rejected-alternative or protocol
eligibility fields. Validation enforces known evidence IDs, sequential unique
observation IDs, backward-only scalar references, closed vocabularies, boolean
participant flags, JSON-nullable participant/qualifier/reference fields and no
string `"null"`.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries or voting. A launch-path defect was corrected
before the live run; the failed launch made zero model calls. A sandbox-blocked
localhost attempt also made zero model calls. The completed run called the
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
Persistent prompts, Gold input and expectations, raw and parsed model output,
validation, automatic comparison, Ollama metadata and human evaluation are in
`artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/`.
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
were schema-invalid. The dominant serialization pattern was use of `present`
instead of the specified `explicit` for affirmation/negation; E additionally
used `none` instead of `absent` for a determination signal, while D emitted the
separate uncertainty concept as an invalid modality. These errors are
contract violations, although most `present`/`explicit` differences are
deterministically normalizable without changing meaning. A and C were valid
JSON/schema outputs but had critical semantic mismatches.
Human semantic review:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; `present` caused schema failure. |
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced
V1 promotion errors: collective speech and speaker identity no longer became
individual responsibility; personal preference no longer became a group-level
rejection field; an information question did not become work; and explicit
negation/affirmation survived as separate evidence. Useful participant evidence
also survived strongly in H and collective-speech evidence in F.
Simplification did not make all evidence-near dimensions reliable. Scalar
references and `limits_target` were almost entirely omitted, qualifiers were
usually omitted, committed modality was missed in E, F and H, and C/G repeated
important modality, uncertainty and participant-feature errors. Some positive
semantic information therefore survived only in free-text `content`, not in
the structural signals a controlled derivation stage would need.
Result: **B — V2 is materially better, but specific evidence-near dimensions
still require refinement.** Direct responsibility, agreement and graph-relation
classification should remain excluded. Before designing the derivation stage,
the next work should examine the minimal reliable representation of explicit
reference/scope limitation, commitment modality and participant deixis. No
production integration, Progeo run or derivation implementation was performed.
## EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
Date: 2026-08-19
Hypothesis: `qwen3.5:9B` is substantially more reliable when the first semantic
stage preserves meeting meaning as atomic natural-language observations with
provenance and only simple participant information, without classifying or
deriving higher-level meeting semantics.
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
observation contains exactly `observation_id`, `evidence_id`, `content`,
`speaker`, nullable `named_person`, and nullable `addressee`. It contains no
modality, temporality, evaluation, affirmation, negation, determination,
uncertainty, clarification, responsibility, agreement, relation, reference,
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
Instead, the prompt asks for conservative atomic content that retains hedges,
conditions, personal/collective/impersonal language, requests, acceptances,
rejections, quantities, deadlines and boundaries in natural language.
Structural validation is intentionally small: exact schema keys, non-empty
observations/content, unique `obs_N` identifiers, known evidence IDs, speaker
matching its evidence, explicit named people/addressees, and no string
`"null"`. Human semantic preservation against per-case requirements is the
primary evaluation; wording differences do not fail a case.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries, voting or per-case tuning. One sandbox-blocked
localhost launch made zero model calls. The completed run made exactly nine
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
structural validation. Persistent source evidence, semantic requirements,
exact prompts, raw and parsed responses, validation, Ollama metadata and human
evaluation are stored under
`artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/`.
Human semantic preservation results:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PASS | Preserved `kann`, `vielleicht`, tentative follow-up, and explicit `Dann` sequence without commitment. |
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, `wenn überhaupt`, impersonal checking need and no decision, but dropped collective `wir` from who would receive contaminated material. |
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
“that we receive contaminated material back” into an impersonal passive phrase;
the risk itself remained hypothetical. H translated `Freitag` to `Friday`, a
harmless wording difference. I retained two compound observations rather than
splitting every proposition, but all required semantic boundaries and
dependencies remained explicit.
Compared with V2, categorical-field removal improved content preservation in
A, G and I: A retained `Dann`; G retained `wenn überhaupt`, personal `Ich` and
impersonal `Man`; I retained publication purpose and all boundaries. It also
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
strengthening into responsibility, group rejection, established work or
assigned clarification. F and H remain sufficiently complete in natural
language for a later interpretation experiment. No useful meaning was shown to
depend on the removed fields; the V3 content retained the useful signals that
V2's fields had attempted to encode.
Result: **A — MINIMAL FIRST STAGE ACCEPTED.** On A-I, minimal atomic content
with evidence provenance and simple participants is sufficiently reliable to
be the candidate first semantic stage. A later bounded experiment may examine
controlled semantic interpretation, but no derivation stage, production
integration or Progeo run was implemented here.
## EXP-0030 — V3 model comparison: qwen3.5:9B vs qwen3.6:35B-A3B
Date: 2026-08-19
This controlled comparison reran the accepted V3 minimal semantic-preservation
experiment unchanged with the locally installed `qwen3.6:35B-A3B`. It used the
same implementation, A-I fixture, evidence, prompt, minimal schema, temperature
0, `think=false`, `num_ctx=16384`, `num_predict=4096`, no retries, no voting and
one call per case. The completed run made exactly nine calls in 96.392 seconds.
Artifacts are preserved under
`artifacts/experiments/evidence_observations_v3/20260819_v3_qwen36_35b_a3b_single_run/`.
Structural validation passed 8/9 cases. D was semantically faithful but invalid
because the model copied transcript speakers Antonius and Martin into
`named_person`, although those names were not explicitly named within their
utterances. Human semantic preservation produced 7 PASS, 1 PARTIAL and 1 FAIL:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PASS | Preserved `can`, `perhaps`, tentative `would`, explicit `then` and no commitment, but changed the content language to English. |
| B | PARTIAL | Preserved both approaches overall, but removed `Oder` from the 20-20 observation and locally strengthened it into collective planned conduct. |
| C | PASS | Preserved Tim's tentative personal Textor contact and possible follow-up without established work. |
| D | PASS | Preserved possibility, uncertainty, process and energy consequence; structural failure was confined to invalid speaker-as-named-person values. |
| E | PASS | Preserved cost, collective rejection/non-pursuit and later determination, though the confirmation dropped explicit `Ja`. |
| F | FAIL | Preserved quantity, timing and trial/series boundaries, but changed collective `wir` into “Martin suggests” and “Tim agrees,” inventing individual proposal/agreement meaning. |
| G | PASS | Preserved hypothetical risk, collective recipient `wir`, Martin's personal stance, `wenn überhaupt`, conditional Technikum, impersonal `man müsste` and no decision/owner. |
| H | PASS | Preserved request, addressee, Friday, explicit acceptance and future personal commitment without a responsibility field; content was English. |
| I | PASS | Preserved local/pure-production and comparison boundaries, five-degree difference, upstream effort, publication purpose, unresolved permission and clarification without assignment. |
Direct comparison:
| Measure | `qwen3.5:9B` | `qwen3.6:35B-A3B` |
| --- | ---: | ---: |
| Structurally valid | 9/9 | 8/9 |
| Human PASS | 8 | 7 |
| Human PARTIAL | 1 | 1 |
| Human FAIL | 0 | 1 |
| Observations | 28 | 30 |
| Runtime | 40.074 s | 96.392 s |
| LLM calls | 9 | 9 |
| Prompt-evaluation tokens | 6,268 | 6,268 |
| Evaluation tokens | 2,590 | 2,718 |
The larger model fixed the 9B weakness in G by preserving collective `wir`, and
it split I's compound production/effort observations more cleanly. Those gains
did not offset regressions: B was locally strengthened, F materially converted
collective conduct into individual agreement, D violated the participant
schema, observation count increased, and runtime was 2.4 times higher. Both
models preserved German consistently in six of nine cases, but in different
cases; the 35B-A3B model changed A, F and H to English, while 9B changed C, F
and H wholly or partly to English.
Result: **D — REGRESSION.** `qwen3.6:35B-A3B` does not materially improve the
accepted minimal V3 first-stage preservation over `qwen3.5:9B`; it is worse on
the A-I comparison because of the F ownership-adjacent strengthening and lower
structural validity. This conclusion applies only to the minimal V3 first
stage and does not determine model choice for any later semantic derivation.
No production integration, derivation implementation or Progeo run occurred.
## EXP-0031 — Controlled Semantic Derivation H V0
Date: 2026-08-19
This isolated experiment tested the first controlled second-stage derivation
using only the accepted `qwen3.5:9B` V3 observations for case H. The derivation
LLM received the two V3 observations, not the transcript or Gold expectation.
Its deliberately narrow task was limited to recognizing whether `obs_1` is a
concrete request and whether `obs_2` explicitly commits its speaker to
substantially the same work. Its strict output schema forbids responsibility,
requested actor, establishment/status, Action Item, protocol, confidence and
generic relation/graph fields.
Deterministic code validates observation/evidence provenance, obtains the
requested actor only from the request observation's addressee, requires the
acceptance to follow the request, requires the accepting speaker to equal that
addressee, and establishes responsibility only after all semantic and
structural gates pass. A bounded weekday normalizer reconciles `Friday` and
`Freitag`, rejects conflicting weekdays, and separates the supported due date
from the normalized action text. No general temporal or action ontology was
introduced.
Twenty focused deterministic tests cover the positive H path and the required
negative invariants: request alone, acknowledgement/non-commitment, tentative
acceptance, different response speaker, different work, reversed order,
speaker/name/addressee alone, conflicting deadlines, unknown observation IDs,
inconsistent evidence provenance, forbidden semantic fields, malformed JSON
and persistent artifacts. The complete non-LLM suite passed 192/192.
Configuration: one `qwen3.5:9B` call, temperature 0, `think=false`,
`num_ctx=16384`, `num_predict=1024`, no retries or voting. The call took 11.765
seconds, with 469 prompt-evaluation and 124 evaluation tokens. The model
returned a valid recognition object: `obs_1` is a concrete request, `obs_2` is
an explicit commitment, and both concern substantially the same work. It
returned no responsibility or establishment judgment.
All deterministic gates passed. The final derived result is an established
action `Prüfung der Messdaten`, requested from and assigned to Nina, due
`Freitag`, supported by request `obs_1/e1` and acceptance `obs_2/e2`. The model
included `bis Friday` in its normalized request text; after the single call, a
deterministic-only bounded correction separated that already-recognized due
phrase from action content without changing the prompt, recognition schema,
semantic result or call count. Focused and complete non-LLM suites still
passed after this correction.
Artifacts are preserved under
`artifacts/experiments/controlled_semantic_derivation_h/20260819_h_qwen35_9b_single_run/`.
Result: the H mechanism succeeded. This establishes only that the narrow
request-plus-explicit-acceptance pattern can be recognized and gated for H; it
does not generalize the derivation architecture to other cases or semantic
categories. No production integration, other case run, semantic graph,
protocol derivation or Progeo run occurred.
## EXP-0032 — Request / Acceptance Gold V0
Status: Experimental; promising with semantic precision gaps
Date: 2026-08-20
This isolated regression experiment tested whether the EXP-0031 mechanism
generalizes beyond H. It used ten short synthetic cases containing only
V3-style observations. Evidence Observation V3 was neither called nor changed,
and the model received no raw transcript or expected result. The fixed
recognition schema permits only a nullable concrete request and nullable later
explicit personal commitment, plus the same-requested-work judgment and
normalized action text. Responsibility, requested actor, established status,
Action Item, protocol, confidence and generic graph fields remain forbidden.
Cases:
- RA-01 explicit positive acceptance: PASS.
- RA-02 paraphrased positive acceptance: PASS.
- RA-03 acknowledgement only: PASS.
- RA-04 tentative response: PASS.
- RA-05 different responder without personal acceptance: PASS.
- RA-06 explicit commitment to different work: PARTIAL. The model returned no
acceptance instead of recognizing a commitment with `same_requested_work`
false. The requested action correctly remained unestablished.
- RA-07 request without response: PASS.
- RA-08 collective commitment: PARTIAL. The model over-recognized the
collective `wir` statement as an explicit commitment, but no request existed
and deterministic gates prevented individual responsibility.
- RA-09 impersonal necessity: PARTIAL. The model over-recognized the impersonal
necessity as a concrete request, but the observation had no addressee and
deterministic gates prevented establishment.
- RA-10 tentative personal suggestion: PASS.
Configuration: exactly ten sequential `qwen3.5:9B` calls, one per case,
temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no retries,
no voting and no prompt change between cases. Summed call time was 23.754
seconds, with 4,826 prompt-evaluation tokens and 793 evaluation tokens. The
strict schema validated every response and no responsibility or establishment
field leaked into model output.
Both positive cases recognized the request, explicit commitment and same-work
relationship, including the paraphrased acceptance, and deterministically
established Clara as responsible with due date `Dienstag`. The model rendered
the normalized action in semantically equivalent English; evaluation therefore
checks the structural deterministic result exactly while treating normalized
action wording as evidence-near semantic text rather than requiring lexical
identity. Acknowledgement and tentative response were not promoted. Every
negative case remained unestablished, and no individual responsibility was
invented.
Recognition-level errors were two false positives (RA-08 commitment and RA-09
request) and one false negative (RA-06 different-work commitment). Final
established-action false positives and false negatives were both zero. The
overall result was seven PASS, three PARTIAL and zero FAIL.
Conclusion: the narrow request-plus-acceptance architecture remains promising
for established individual actions because deterministic addressee, ordering,
speaker, same-work, provenance and deadline gates contained all recognition
errors. The recognition layer is not yet precise enough to generalize: its
handling of collective commitment, impersonal necessity and commitments to
different work needs further isolated study. No production integration or
additional semantic category is justified by this result.
Artifacts are preserved under
`artifacts/experiments/request_acceptance_gold_v0/20260820_qwen35_9b_single_run/`.
## EXP-0033 — Collective Commitment Gold V0
Status: Experimental; architecturally successful with one contained
recognition false positive
Date: 2026-08-20
This isolated second-stage experiment tested whether an explicit collective
first-person commitment can establish an action without inventing an individual
owner. It used ten synthetic cases containing one minimal V3-style observation
each. Evidence Observation V3 was neither called nor changed, and the accepted
Request/Acceptance mechanism remained unchanged and independent.
The strict semantic schema contains exactly `observation_id`,
`commitment_form` and `normalized_action_text`. `commitment_form` is closed to
`individual_first_person`, `collective_first_person` and `none`. The model
cannot output responsibility, ownership, requested actor, establishment,
Action Item, protocol, confidence, relations, graphs, decisions or unresolved
issues. Deterministic code validates schema and provenance, requires collective
commitment plus non-empty action text, applies bounded deadline consistency and
explicit-negation gates, and only then sets `status: established`,
`commitment_scope: collective` and `responsible_person: null`.
Gold results:
- CC-01 explicit collective commitment: PASS; established, due `nächste
Woche`, no person.
- CC-02 individual commitment: PASS; correctly routed out of the collective
path.
- CC-03 tentative collective possibility: PASS; unestablished.
- CC-04 collective suggestion: PASS; unestablished.
- CC-05 impersonal necessity: PASS; unestablished.
- CC-06 passive future statement: PASS; unestablished.
- CC-07 collective rejection: PARTIAL. The model incorrectly returned
`collective_first_person`, but the deterministic negation gate detected
`nicht` and prevented establishment.
- CC-08 qualified collective commitment: PASS; established with `nur im
Technikum` preserved, null due and no person.
- CC-09 collective commitment without deadline: PASS; established with null
due and no person.
- CC-10 speaker ownership trap: PASS; established collectively while Martin
remained only the speaker and was not assigned ownership.
Configuration: exactly ten successful sequential `qwen3.5:9B` calls, one per
case, temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no
retries, no voting and no prompt change. There were zero technical failed
calls. Aggregate runner time was 10.504 seconds; summed per-call time was 10.500
seconds, with 4,267 prompt-evaluation tokens and 415 evaluation tokens.
The outcome was nine PASS, one PARTIAL and zero FAIL. There was one recognition
false positive and no recognition false negatives. No qualifier was lost, no
individual owner was invented, and no responsibility or status field leaked
into recognition. Bounded due handling preserved `nächste Woche` verbatim and
returned null when no deadline was present.
Conclusion: the collective-commitment path is architecturally successful for
this narrow Gold set. The deterministic negation gate contained the only model
error, and every successful collective result necessarily retained
`responsible_person: null`. This does not justify a generic commitment system,
production integration, group identity inference or another semantic category.
Artifacts are preserved under
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
Date: 2026-08-11
Hypothesis: the primary protocol should be a topic-oriented reconstruction of
the meeting rather than a category-oriented list of extracted information.
This first isolated prototype does not replace or connect to the production
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
transcript excerpts to `qwen3.5:9B` and requests Discussion Subjects. Each
subject may contain supported discourse events, an outcome with mandatory
scope, resulting actions and unresolved issues. Optional structures must be
omitted when absent. Every semantic object must reference known evidence IDs.
The strict experimental schema validates:
- non-empty subjects and globally unique semantic identifiers;
- a closed discourse-event vocabulary;
- non-empty, known and non-duplicated evidence references;
- outcome text, scope, certainty and evidence;
- action text, JSON-nullable responsibility/deadline and evidence;
- unresolved-issue text and evidence;
- omission rather than null or empty optional structures.
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only,
multiple options, unaccepted proposal, proposal with objection, rejected
alternative, trial-scoped acceptance, no-decision discussion, resulting Action
Item, and outcome plus unresolved issue. Evaluation targets semantic identity,
development, outcome scope, actions, unresolved issues, traceability and
absence of invented commitments rather than exact wording.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`. Each case received exactly one model call; there were no
model retries or prompt iterations. The nine completed calls took 59.251
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
failure artifacts were preserved under
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/` and
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/`. Two earlier launch
attempts made zero LLM calls: one failed on the script import path and one was
blocked by sandbox networking.
Human-reviewed results after correcting two objectively wrong Gold assumptions
without another model call:
| Case | Verdict | Reason |
| --- | --- | --- |
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled `considered_option` rather than `introduced_idea`. |
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case
F: the model distinguished acceptance for a bounded trial from acceptance as a
final solution. It also handled the explicit action in case H well. However,
the experiment failed systematically on sparse structured output, subject
grouping and restraint around absent outcomes/actions/unresolved issues. The
model frequently mirrored optional schema fields as empty/null values, treated
alternatives as outcomes, split one discussion into multiple subjects, or
invented open issues from mere non-selection.
The focused experiment is not promising enough to justify a real Progeo chunk
sanity check. No such run was performed, and no architecture is accepted on
the basis of this prototype. Further work should first analyze whether the
failure comes from the schema/prompt representation, the model's sparse-output
reliability, or the boundary between subject grouping and semantic synthesis.
It should not proceed through repeated prompt tuning against these nine cases.