2037 lines
81 KiB
Markdown
2037 lines
81 KiB
Markdown
# Experiments
|
|
|
|
This file records durable technical experiments and findings for Meeting Lab.
|
|
It is not a diary and does not replace commit history.
|
|
|
|
## Status values
|
|
|
|
- Proposed: experiment idea exists, but no result is recorded.
|
|
- Running: experiment is in progress and no decision has been made.
|
|
- Accepted: finding is the current baseline or design conclusion.
|
|
- Rejected: hypothesis was tested and should not be repeated as-is.
|
|
- Superseded: finding was useful but has been replaced by a newer baseline.
|
|
|
|
## EXP-0001 - Whisper JSON interpretation
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-29
|
|
|
|
Hypothesis:
|
|
|
|
Whisper JSON should be chunked from its segment stream, not from the aggregate
|
|
top-level text field.
|
|
|
|
Setup:
|
|
|
|
`chunk_transcript.py` was updated to parse JSON input and prefer
|
|
`segments[*].text` when `segments` exists. A regression test supplies JSON with
|
|
both top-level `text` and separate segment texts.
|
|
|
|
Inputs:
|
|
|
|
- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`.
|
|
- Real Whisper artifacts under `samples/whisper/`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM.
|
|
|
|
Result:
|
|
|
|
The test verifies that the block stream is `["alpha", "beta", "gamma"]` and
|
|
does not include the aggregate `"alpha beta gamma"` text. The repository history
|
|
records this as the fix for the earlier failure where the first chunk contained
|
|
the complete transcript.
|
|
|
|
Decision:
|
|
|
|
When `segments` exists, `segments[*].text` is the authoritative transcript
|
|
stream. The top-level `text` field is only a fallback.
|
|
|
|
Lessons learned:
|
|
|
|
Whisper JSON is structured input. Treating it like plain text can duplicate the
|
|
entire transcript and invalidate downstream chunking.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
- `tests/test_chunking.py`
|
|
- Commit `4656523` - `Fix Whisper JSON chunk extraction`
|
|
|
|
## EXP-0002 - Technical transcript chunking baseline
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-29 to 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Sequential technical chunks around the configured target size can preserve the
|
|
transcript while keeping extraction calls small enough for local models.
|
|
|
|
Setup:
|
|
|
|
The chunker splits block-aligned text with configurable target, minimum,
|
|
maximum and overlap settings. Tests verify no duplicate later blocks when
|
|
overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
|
|
|
|
Inputs:
|
|
|
|
- Synthetic block list in `tests/test_chunking.py`.
|
|
- `samples/whisper/meeting_speech_cleaned.json`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM for chunking.
|
|
- Manifest uses default chunking behavior recorded in
|
|
`samples/whisper/meeting_speech_cleaned_chunks/manifest.json`.
|
|
|
|
Result:
|
|
|
|
With overlap set to zero, tests verify that all blocks appear exactly once. The
|
|
real sample manifest contains nine chunks, mostly near the configured target
|
|
size, with a smaller final chunk.
|
|
|
|
Decision:
|
|
|
|
Independent sequential chunks are the current technical baseline. One
|
|
normalized chunk per extraction call is the preferred extraction strategy.
|
|
|
|
Lessons learned:
|
|
|
|
Chunking solves model-size constraints only. It must not perform topic
|
|
detection or semantic merging.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
- `tests/test_chunking.py`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`
|
|
- `AGENTS.md`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
|
|
## EXP-0003 - Conservative transcript normalization
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Transcript cleanup should improve readability without changing meeting
|
|
semantics.
|
|
|
|
Setup:
|
|
|
|
The normalizer removes isolated filler sounds, immediate duplicate words or
|
|
short duplicate phrases, and redundant whitespace. It records changed blocks in
|
|
a JSON change log and explicitly preserves semantic content categories.
|
|
|
|
Inputs:
|
|
|
|
- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
- Change logs such as `chunk_01_changes.json`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM.
|
|
|
|
Result:
|
|
|
|
The implementation and generated change logs show a conservative policy:
|
|
negations, qualifiers, dates, numbers, responsibilities, technical statements,
|
|
deadlines, decisions and commitments are preserved.
|
|
|
|
Decision:
|
|
|
|
Normalization remains deterministic and low-risk. When uncertain, leave text
|
|
unchanged.
|
|
|
|
Lessons learned:
|
|
|
|
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove
|
|
semantic cues needed by extraction.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/normalization/normalize_transcript.py`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json`
|
|
- `docs/pipeline.md`
|
|
|
|
## EXP-0004 - Full-context topic segmentation
|
|
|
|
Status: Superseded
|
|
|
|
Date or period: 2026-07-21 to 2026-07-22
|
|
|
|
Hypothesis:
|
|
|
|
A single full-context topic segmentation call can identify topic boundaries in
|
|
a normalized transcript chunk.
|
|
|
|
Setup:
|
|
|
|
The initial segmentation prototype asked the model for topic changes and then
|
|
converted those boundaries into continuous, non-overlapping segments.
|
|
|
|
Inputs:
|
|
|
|
- `samples/chunks/chunk_01_normalized.txt`.
|
|
|
|
Model / configuration:
|
|
|
|
- Generated artifact records `qwen3:14b`.
|
|
|
|
Result:
|
|
|
|
The generated artifact contains 85 blocks, three topic-change boundaries and
|
|
four segments. The run metadata records a substantially longer elapsed time
|
|
than the later windowed artifact for the same input.
|
|
|
|
Decision:
|
|
|
|
Full-context segmentation was useful as a prototype, but it was superseded by
|
|
windowed segmentation and manual review tooling.
|
|
|
|
Lessons learned:
|
|
|
|
The prototype established the boundary-to-segment representation, but did not
|
|
settle segmentation quality.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/segmentation/segment_topics.py`
|
|
- `samples/chunks/chunk_01_normalized_segments.json`
|
|
- Commit `f234efc` - `Add initial topic segmentation prototype`
|
|
- Commit `889a4fe` - `Detect topic boundaries as continuous segments`
|
|
|
|
## EXP-0005 - Windowed topic segmentation and review
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-22
|
|
|
|
Hypothesis:
|
|
|
|
Windowed topic segmentation can reduce runtime and make boundary evaluation
|
|
more inspectable than a single full-context call.
|
|
|
|
Setup:
|
|
|
|
`segment_topics_windowed.py` analyzes overlapping windows and reports only
|
|
boundaries from the decision range. Python merges boundaries into continuous,
|
|
non-overlapping segments. `review_segmentation.py` renders each boundary with
|
|
neighboring transcript context for manual classification.
|
|
|
|
Inputs:
|
|
|
|
- `samples/chunks/chunk_01_normalized.txt`.
|
|
- Full meeting normalized chunks under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3.
|
|
- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3.
|
|
|
|
Result:
|
|
|
|
The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments
|
|
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
|
|
Manual review output shows that some boundaries were assessed as subtopics
|
|
rather than full topic changes. Full-meeting artifacts show one window per
|
|
already-small normalized chunk and two segments per chunk.
|
|
|
|
Decision:
|
|
|
|
Windowed segmentation and review tooling are accepted as prototype tooling, not
|
|
as a stable production segmentation stage.
|
|
|
|
Lessons learned:
|
|
|
|
Windowing improves inspectability and can reduce runtime, but it can also
|
|
cluster boundaries and over-segment. Manual review remains necessary.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/segmentation/segment_topics_windowed.py`
|
|
- `src/meeting_lab/segmentation/review_segmentation.py`
|
|
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
|
- `samples/chunks/chunk_01_normalized_windowed_segments_review.md`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
|
- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling`
|
|
|
|
## EXP-0006 - Qwen model comparison
|
|
|
|
Status: Accepted
|
|
|
|
Hypothesis:
|
|
|
|
Larger local Qwen-family models should improve meaningful extraction and
|
|
segmentation, but model size alone will not solve prompt or pipeline problems.
|
|
|
|
Setup:
|
|
|
|
Project work used smaller models for smoke checks and larger local models for
|
|
meaningful extraction or segmentation. Artifacts and project knowledge record
|
|
the currently useful model roles.
|
|
|
|
Inputs:
|
|
|
|
- Gold Standard scenarios under `tests/gold/`.
|
|
- Generated segmentation artifacts under `samples/`.
|
|
- Generated extraction artifacts under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- `qwen3:1.7b`: smoke-test model according to project knowledge.
|
|
- `qwen3.5:9b`: current meaningful extraction and segmentation model according
|
|
to project knowledge and generated full-meeting segmentation artifacts.
|
|
- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts.
|
|
|
|
Result:
|
|
|
|
The repository supports the conclusion that `qwen3.5:9b` is the meaningful
|
|
current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger
|
|
models and longer contexts may increase runtime substantially, but no hardware
|
|
benchmark suite is recorded.
|
|
|
|
Decision:
|
|
|
|
Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current
|
|
experiments. Do not assume model size alone fixes prompt or pipeline design.
|
|
|
|
Lessons learned:
|
|
|
|
Evaluation must separate model capability from prompt clarity, context
|
|
strategy, extraction schema and consolidation.
|
|
|
|
Evidence:
|
|
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
|
- `samples/chunks/chunk_01_normalized_segments.json`
|
|
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
|
|
|
## EXP-0007 - Thinking output and Ollama API behavior
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Thinking-capable Qwen models may return reasoning separately from the final
|
|
answer, and extraction parsing should not fail merely because extra text or
|
|
multiple JSON objects appear.
|
|
|
|
Setup:
|
|
|
|
The Ollama response reader checks `response`, chat-style `message.content` and
|
|
then `thinking`. The JSON parser tries a full parse first, then scans JSON
|
|
object candidates and returns the final valid object. A regression test covers
|
|
thinking text before final JSON.
|
|
|
|
Inputs:
|
|
|
|
- Synthetic parser test in `tests/test_extraction_protocol.py`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM run in the test.
|
|
- Code path is used by Ollama extraction.
|
|
|
|
Result:
|
|
|
|
The parser can handle additional text and multiple JSON objects where the final
|
|
valid object is the intended answer. Current code still falls back to `thinking`
|
|
only if no usable response or message content is present.
|
|
|
|
Decision:
|
|
|
|
Keep parser robustness, but do not treat thinking output as the root cause of
|
|
all extraction failures.
|
|
|
|
Lessons learned:
|
|
|
|
API response shape and model output shape are separate concerns. Preserve raw
|
|
responses when diagnosing failures.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
|
- `tests/test_extraction_protocol.py`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
|
|
## EXP-0008 - JSON truncation and generation limits
|
|
|
|
Status: Accepted
|
|
|
|
Hypothesis:
|
|
|
|
Some extraction failures are caused by generation limits truncating JSON rather
|
|
than by prompt wording or parser behavior.
|
|
|
|
Setup:
|
|
|
|
A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The
|
|
generation limit was increased for the successful path. Exact failing limit is
|
|
not recorded in the repository; the current extractor default is verifiably
|
|
`--num-predict 8192`.
|
|
|
|
Inputs:
|
|
|
|
- Local extraction runs referenced by project knowledge.
|
|
- Current extraction CLI.
|
|
|
|
Model / configuration:
|
|
|
|
- `qwen3.5:9b`.
|
|
- Current extractor default: `num_predict=8192`.
|
|
|
|
Result:
|
|
|
|
Increasing the generation limit fixed the technical JSON failure. This was not
|
|
primarily a parser or prompt problem.
|
|
|
|
Decision:
|
|
|
|
When JSON is truncated, inspect raw output and generation limits before editing
|
|
prompts.
|
|
|
|
Lessons learned:
|
|
|
|
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong
|
|
first response if the model simply ran out of output tokens.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `AGENTS.md`
|
|
|
|
## EXP-0009 - Minimal end-to-end pipeline
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
A minimal local pipeline can transform Whisper output into chunk extractions
|
|
and an interim protocol, proving the technical path before the final
|
|
architecture exists.
|
|
|
|
Setup:
|
|
|
|
The repository added cleanup, normalization, chunking, extraction and protocol
|
|
builder scripts, with sample generated artifacts.
|
|
|
|
Inputs:
|
|
|
|
- `samples/whisper/meeting_speech.json`
|
|
- `samples/whisper/meeting_speech_cleaned.json`
|
|
- Generated chunks and normalized chunks under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- Local Ollama extraction for chunk JSON.
|
|
- Windowed segmentation artifacts use `qwen3.5:9b`.
|
|
|
|
Result:
|
|
|
|
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
|
|
extraction JSON files, windowed segmentation artifacts and
|
|
`meeting_protocol.md`.
|
|
|
|
Decision:
|
|
|
|
The minimal pipeline is technically validated. The first protocol builder is an
|
|
interim validation tool, not the final architecture.
|
|
|
|
Lessons learned:
|
|
|
|
End-to-end execution exposed the next limitation: extraction output needs
|
|
consolidation and purpose-specific rendering.
|
|
|
|
Evidence:
|
|
|
|
- `scripts/clean_whisper_json.py`
|
|
- `src/meeting_lab/normalization/normalize_transcript.py`
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
|
- `src/meeting_lab/protocol/build_protocol.py`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/`
|
|
- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline`
|
|
|
|
## EXP-0010 - Gold Standard corpus
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Reproducible prompt engineering requires synthetic transcripts with explicit
|
|
expected semantic outputs.
|
|
|
|
Setup:
|
|
|
|
The Gold Standard corpus defines scenario directories with `transcript.txt`,
|
|
`expected.json` and README files describing ground truth and common model
|
|
mistakes. The runner validates schema keys and writes `actual.json` for a
|
|
scenario.
|
|
|
|
Inputs:
|
|
|
|
- Gold scenarios under `tests/gold/`.
|
|
|
|
Model / configuration:
|
|
|
|
- Runner requires an explicit Ollama model for LLM evaluation.
|
|
- Existing unit tests for the runner do not invoke Ollama.
|
|
|
|
Result:
|
|
|
|
The corpus gives stable semantics for decisions, facts, positions, todos,
|
|
questions, technical details and difficult mixed cases. Initial structured
|
|
transcripts are Phase 1 and easier than raw Whisper-style transcripts.
|
|
|
|
Decision:
|
|
|
|
Use Gold Standard tests as both regression tests and formal meeting-semantics
|
|
specification. Raw or unlabelled transcript cases remain later-phase work.
|
|
|
|
Lessons learned:
|
|
|
|
Without expected outputs, prompt changes cannot be evaluated reproducibly.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/`
|
|
- `scripts/run_gold_test.py`
|
|
- `tests/test_gold_runner.py`
|
|
- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests`
|
|
|
|
## EXP-0011 - Gold-test quality and unique ground truth
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
If a prompt produces unexpected behavior, the gold test itself may be ambiguous
|
|
and should be reviewed before the prompt is changed.
|
|
|
|
Setup:
|
|
|
|
Decision-focused scenarios were clarified during baseline creation. The current
|
|
methodology requires checking unique ground truth before changing prompts.
|
|
|
|
Inputs:
|
|
|
|
- `decision_simple`
|
|
- `decision_deferred`
|
|
- `decision_none`
|
|
- Gold methodology document.
|
|
|
|
Model / configuration:
|
|
|
|
- Prompt Version 2 baseline work.
|
|
|
|
Result:
|
|
|
|
`decision_simple` required clarification around the explicit agreement and
|
|
nearby non-decision wording. The earlier negative/deferral ambiguity is now
|
|
represented by distinct `decision_none` and `decision_deferred` scenarios in
|
|
the repository. Punctuation is not reliable evidence for Whisper transcripts;
|
|
agreement language and wording must carry the semantics.
|
|
|
|
Decision:
|
|
|
|
Ambiguous gold tests must be reviewed before prompt changes. Do not treat
|
|
punctuation as reliable evidence in real Whisper-style transcripts.
|
|
|
|
Lessons learned:
|
|
|
|
Bad gold tests create false prompt failures and can encourage overfitting.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
|
- `tests/gold/decision_simple/README.md`
|
|
- `tests/gold/decision_deferred/README.md`
|
|
- `tests/gold/decision_none/README.md`
|
|
- Commit `f7ad9ba`
|
|
|
|
## EXP-0012 - Decision taxonomy
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Decision extraction needs a formal taxonomy that distinguishes substantive
|
|
decisions from process decisions and non-decisions.
|
|
|
|
Setup:
|
|
|
|
The decision definition document and decision prompt define included and
|
|
excluded categories. Gold tests cover explicit decisions, true no-decision
|
|
cases and deferrals.
|
|
|
|
Inputs:
|
|
|
|
- `tests/gold/DECISION_DEFINITION.md`
|
|
- `prompts/decisions.md`
|
|
- Decision gold scenarios.
|
|
|
|
Model / configuration:
|
|
|
|
- Prompt Version 2 baseline.
|
|
|
|
Result:
|
|
|
|
Accepted decision categories include substantive decisions, organizational
|
|
decisions, process decisions, approvals, rejections, deferrals, explicit
|
|
decisions not to decide yet and explicit agreement to gather more information
|
|
before deciding. Opinions, preferences and proposals without agreement are not
|
|
decisions.
|
|
|
|
Decision:
|
|
|
|
"No decision was reached" and "the decision was deferred" are distinct semantic
|
|
outcomes.
|
|
|
|
Lessons learned:
|
|
|
|
Deferral can be a valid process decision even when the substantive topic remains
|
|
unresolved.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/DECISION_DEFINITION.md`
|
|
- `tests/gold/decision_deferred/`
|
|
- `tests/gold/decision_none/`
|
|
- `prompts/decisions.md`
|
|
|
|
## EXP-0013 - Prompt engineering methodology
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Prompt iteration needs strict experimental controls to prevent regression,
|
|
overfitting and arbitrary prompt churn.
|
|
|
|
Setup:
|
|
|
|
The methodology was documented alongside the Gold Standard corpus and later
|
|
summarized for agents.
|
|
|
|
Inputs:
|
|
|
|
- Gold scenarios.
|
|
- Prompt files.
|
|
|
|
Model / configuration:
|
|
|
|
- Applies to all prompt experiments.
|
|
|
|
Result:
|
|
|
|
The accepted method is one prompt change per iteration, one target test at a
|
|
time, immediate validation, no regressions, no `expected.json` edits merely to
|
|
force a pass, no test-specific prompt hacks, stopping after two consecutive
|
|
non-improving iterations, and verifying unique ground truth before prompt
|
|
changes.
|
|
|
|
Decision:
|
|
|
|
Prompt changes are controlled experiments. See `AGENTS.md` for agent operating
|
|
rules.
|
|
|
|
Lessons learned:
|
|
|
|
Most prompt changes are not isolated unless the experiment explicitly constrains
|
|
the target and regression set.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
|
- `AGENTS.md`
|
|
- Commit `f7ad9ba`
|
|
|
|
## EXP-0014 - Decision Prompt Version 2
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Adding explicit process-decision language to the decision prompt can preserve
|
|
true decision detection while recognizing deferrals.
|
|
|
|
Setup:
|
|
|
|
Prompt Version 2 added explicit support for deferrals and process decisions.
|
|
The baseline was validated on three decision scenarios.
|
|
|
|
Inputs:
|
|
|
|
- `decision_simple`
|
|
- `decision_deferred`
|
|
- `decision_none`
|
|
|
|
Model / configuration:
|
|
|
|
- Prompt Version 2.
|
|
- Model used for validation is not recorded in the committed methodology.
|
|
|
|
Result:
|
|
|
|
The committed methodology records all three baseline scenarios as passing. The
|
|
current generated `actual.json` files also show the expected decision count for
|
|
these decision scenarios, although some non-decision categories remain less
|
|
complete.
|
|
|
|
Decision:
|
|
|
|
Prompt Version 2 is the current decision baseline. Explicit deferrals are
|
|
recognized as process decisions.
|
|
|
|
Lessons learned:
|
|
|
|
Decision-count success does not imply all categories are solved. Category-level
|
|
evaluation must continue.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
|
- `tests/gold/decision_simple/actual.json`
|
|
- `tests/gold/decision_deferred/actual.json`
|
|
- `tests/gold/decision_none/actual.json`
|
|
- `prompts/decisions.md`
|
|
|
|
## EXP-0015 - Difficult synthetic meeting
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
A deliberately adversarial synthetic meeting can expose extraction failures
|
|
that simple category tests miss.
|
|
|
|
Setup:
|
|
|
|
`evil_meeting` includes interruptions, corrections, absent referenced people,
|
|
near-decisions, changed positions and one expected explicit decision.
|
|
|
|
Inputs:
|
|
|
|
- `tests/gold/evil_meeting/`.
|
|
|
|
Model / configuration:
|
|
|
|
- Existing generated `actual.json`; exact model is not stored in the artifact.
|
|
|
|
Result:
|
|
|
|
The generated result found the expected FR-7 exclusion decision. It also
|
|
classified "do not migrate until the mapping table is checked" as an additional
|
|
decision. The current expected file treats that statement as a position, but
|
|
contextual review suggests it may be a valid process instruction or decision.
|
|
|
|
Decision:
|
|
|
|
Do not classify this as a simple model failure without reviewing the gold
|
|
standard. The scenario exposes a semantic gap in the expected output.
|
|
|
|
Lessons learned:
|
|
|
|
Difficult synthetic cases are valuable because they reveal ambiguity in the
|
|
specification as well as model mistakes.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/evil_meeting/README.md`
|
|
- `tests/gold/evil_meeting/expected.json`
|
|
- `tests/gold/evil_meeting/actual.json`
|
|
- See EXP-0011 and EXP-0012.
|
|
|
|
## EXP-0016 - Context-size extraction comparisons
|
|
|
|
Status: Rejected
|
|
|
|
Hypothesis:
|
|
|
|
Increasing extraction context from one chunk to neighboring chunk groups should
|
|
make extraction more complete and therefore should become the baseline.
|
|
|
|
Setup:
|
|
|
|
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3
|
|
and 1+2+3.
|
|
|
|
Inputs:
|
|
|
|
- Normalized meeting chunks.
|
|
|
|
Model / configuration:
|
|
|
|
- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for
|
|
extraction experiments.
|
|
|
|
Result:
|
|
|
|
More context sometimes improved completeness, but it also shifted category
|
|
classification, added duplicates and reduced stability. No numeric winner is
|
|
recorded in the repository.
|
|
|
|
Decision:
|
|
|
|
Do not adopt larger extraction windows as the baseline. Independent chunk
|
|
extraction remains current strategy. Recover global context through
|
|
consolidation rather than continuously enlarging extraction windows.
|
|
|
|
Lessons learned:
|
|
|
|
Context size is not a monotonic quality knob. It changes the task the model is
|
|
performing.
|
|
|
|
Evidence:
|
|
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `AGENTS.md`
|
|
- `ROADMAP.md`
|
|
- See EXP-0002 and EXP-0018.
|
|
|
|
## EXP-0017 - Independent full-meeting chunk extraction
|
|
|
|
Status: Accepted
|
|
|
|
Hypothesis:
|
|
|
|
Extracting every normalized chunk independently can produce enough structured
|
|
material for a useful protocol draft.
|
|
|
|
Setup:
|
|
|
|
Nine normalized chunks were extracted into separate JSON files and then
|
|
rendered by the interim protocol builder.
|
|
|
|
Inputs:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt`
|
|
through `chunk_09_normalized.txt`.
|
|
|
|
Model / configuration:
|
|
|
|
- Existing extraction artifacts do not record model metadata.
|
|
|
|
Result:
|
|
|
|
The nine extraction files contain facts, decisions, todos, questions and
|
|
technical details. `meeting_protocol.md` aggregates them into a readable draft.
|
|
Duplicates, category shifts and synthesis became the dominant limitations.
|
|
|
|
Decision:
|
|
|
|
Independent chunk extraction is useful enough to keep as the baseline, but it
|
|
requires a consolidation stage.
|
|
|
|
Lessons learned:
|
|
|
|
Per-chunk extraction gives recall-oriented raw material. It does not by itself
|
|
produce a polished or canonical meeting representation.
|
|
|
|
Evidence:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
|
- `src/meeting_lab/protocol/build_protocol.py`
|
|
- See EXP-0018.
|
|
|
|
## EXP-0018 - Human protocol comparison and output-view split
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
One generated protocol cannot satisfy every use case; protocol output should be
|
|
separated by purpose and audience.
|
|
|
|
Setup:
|
|
|
|
The interim machine protocol was compared against the desired human protocol
|
|
shape and then the architecture was revised toward parallel output views.
|
|
|
|
Inputs:
|
|
|
|
- Interim `meeting_protocol.md`.
|
|
- Architecture and output-view documentation.
|
|
|
|
Model / configuration:
|
|
|
|
- Not applicable; this is a design evaluation.
|
|
|
|
Result:
|
|
|
|
The human protocol target is denser and organized by purpose and topic rather
|
|
than extraction categories. The machine extraction retains more context and is
|
|
useful for recall, but it is not the right direct source for a concise
|
|
distribution artifact or durable knowledge entry.
|
|
|
|
Decision:
|
|
|
|
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution
|
|
Protocol / Verteilerprotokoll and Knowledge Objects /
|
|
Wissensdatenbankeintrag.
|
|
|
|
Lessons learned:
|
|
|
|
Rendering is a separate concern from extraction and consolidation. Output views
|
|
must be parallel renderings of shared semantics, not transformations of one
|
|
another.
|
|
|
|
Evidence:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
|
- `docs/output-views.md`
|
|
- `docs/architecture.md`
|
|
- `docs/pipeline.md`
|
|
- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture`
|
|
|
|
## EXP-0019 - Consolidation and Canonical Meeting Knowledge
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Extraction, consolidation and output rendering are separate problems and should
|
|
not be collapsed into one LLM prompt or one protocol file.
|
|
|
|
Setup:
|
|
|
|
The architecture was refined after the minimal pipeline and protocol draft
|
|
showed duplicate, synthesis and audience-specific rendering limitations.
|
|
|
|
Inputs:
|
|
|
|
- Extraction JSON artifacts.
|
|
- Interim protocol draft.
|
|
- Architecture and data-model documentation.
|
|
|
|
Model / configuration:
|
|
|
|
- Not applicable; this is an architectural conclusion.
|
|
|
|
Result:
|
|
|
|
The accepted design is a planned Canonical Meeting Knowledge layer as the
|
|
semantic source of truth, with Working Protocol, Distribution Protocol and
|
|
Knowledge Objects as parallel output views. The next consolidation architecture
|
|
is split into a Deterministic Canonicalizer and a Semantic Consolidator. The
|
|
canonicalizer prepares validated evidence-bearing objects without uncertain
|
|
semantic merging. The consolidator then merges semantically equivalent
|
|
statements, preserves evidence, reconciles category shifts where supported and
|
|
marks contradictions or uncertainty.
|
|
|
|
Decision:
|
|
|
|
Deterministic canonicalization is implemented as Canonicalizer V1. The first
|
|
semantic consolidation milestone is implemented as Semantic Consolidator V0 for
|
|
facts-only duplicate detection. Broader semantic consolidation, Canonical
|
|
Meeting Knowledge and final output views remain planned.
|
|
|
|
Lessons learned:
|
|
|
|
Global meeting understanding should be recovered by consolidation over
|
|
evidence-bearing extractions, not by silently changing output views or expanding
|
|
LLM context indefinitely.
|
|
|
|
Evidence:
|
|
|
|
- `docs/architecture.md`
|
|
- `docs/pipeline.md`
|
|
- `docs/data-models.md`
|
|
- `docs/output-views.md`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `ROADMAP.md`
|
|
- Commit `5c03ed7`
|
|
|
|
## EXP-0020 - Working Protocol Synthesizer V0
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-31
|
|
|
|
Hypothesis:
|
|
|
|
The current local synthesis model may be able to generate a useful detailed
|
|
Working Protocol directly from the existing independent chunk extraction JSON
|
|
files, before canonicalization or semantic consolidation exists.
|
|
|
|
Setup:
|
|
|
|
One synthesis prompt was constructed from exactly nine chunk extraction JSON
|
|
files. The model was instructed to use only those extraction files, merge
|
|
duplicates, group related information into topics, preserve useful discussion
|
|
context and write a neutral technical Working Protocol.
|
|
|
|
Inputs:
|
|
|
|
- `chunk_01_extraction.json` through `chunk_09_extraction.json`.
|
|
- No original transcript, normalized chunks or Whisper output were used as
|
|
synthesis input.
|
|
|
|
Model / configuration:
|
|
|
|
- Model: `qwen3.5:9b`
|
|
- Prompt characters: 24,979
|
|
- Actual prompt eval tokens: 5,929
|
|
- Output tokens: 1,486
|
|
- Runtime: 294.204 seconds
|
|
|
|
Result:
|
|
|
|
The generated Working Protocol was readable, well structured and
|
|
topic-oriented. It was still based directly on raw chunk extractions, without a
|
|
separate deterministic canonicalization stage or semantic consolidation stage.
|
|
The output language was English even though the source meeting material was
|
|
German.
|
|
|
|
Decision:
|
|
|
|
Preserve this output as the Working Protocol Synthesizer V0 benchmark baseline
|
|
for later canonicalizer, consolidator and renderer comparisons. This selected
|
|
generated artifact is intentionally versioned even though generated runtime
|
|
artifacts are normally ignored.
|
|
|
|
Lessons learned:
|
|
|
|
Direct synthesis from chunk extractions can create a useful recall-oriented
|
|
draft, but it does not replace Canonical Meeting Knowledge. The language
|
|
mismatch also establishes a default renderer rule: protocol output should
|
|
normally match the dominant source language unless an explicit output language
|
|
is requested.
|
|
|
|
Evidence:
|
|
|
|
- `samples/benchmarks/working_protocol_synthesizer_v0/README.md`
|
|
- `samples/benchmarks/working_protocol_synthesizer_v0/working_protocol.md`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `docs/output-views.md`
|
|
- See EXP-0017 and EXP-0019.
|
|
|
|
## EXP-0021 - Canonicalizer V1
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-31
|
|
|
|
Hypothesis:
|
|
|
|
Independent chunk extraction JSON can be converted into a stable deterministic
|
|
intermediate format before any semantic LLM consolidation is attempted.
|
|
|
|
Setup:
|
|
|
|
Canonicalizer V1 discovers `chunk_*_extraction.json` files in stable chunk
|
|
order, validates required categories, normalizes category names and basic field
|
|
structure, parses existing legacy string formats where safe, trims redundant
|
|
whitespace, assigns deterministic IDs, preserves original values and source
|
|
references, and merges only exact duplicates when all semantic fields are
|
|
identical.
|
|
|
|
Inputs:
|
|
|
|
- Synthetic unit-test fixtures.
|
|
- Existing nine extraction JSON files under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM.
|
|
- CLI module: `meeting_lab.consolidation.canonicalize`.
|
|
|
|
Result:
|
|
|
|
Canonicalizer V1 produces `schema_version`, `source_files`, `stats` and
|
|
`items`. It is deterministic preparation for the future Semantic Consolidator
|
|
and is not Canonical Meeting Knowledge.
|
|
|
|
Decision:
|
|
|
|
Canonicalizer V1 is the current implemented deterministic canonicalization
|
|
stage. Semantic Consolidator V0 now uses this representation for facts-only
|
|
semantic duplicate detection; broader semantic consolidation and Canonical
|
|
Meeting Knowledge remain planned.
|
|
|
|
Lessons learned:
|
|
|
|
Exact duplicate handling, source-reference preservation and legacy string
|
|
parsing can be tested without model calls. Any uncertain semantic merge remains
|
|
out of scope for this stage.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/consolidation/canonicalize.py`
|
|
- `tests/test_canonicalize.py`
|
|
- `docs/data-models.md`
|
|
- `docs/pipeline.md`
|
|
|
|
## EXP-0022 - Semantic Consolidator V0 facts-only merge
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-31
|
|
|
|
Hypothesis:
|
|
|
|
The Canonicalizer V1 output contains enough stable structure for a local LLM to
|
|
identify semantically equivalent fact items without losing source coverage or
|
|
changing non-fact categories.
|
|
|
|
Setup:
|
|
|
|
Semantic Consolidator V0 consumed the existing Canonicalizer V1 benchmark JSON,
|
|
selected only items with `category: "fact"`, and sent one bounded consolidation
|
|
request to local Ollama. The merge rules required semantic equivalence, not
|
|
topic similarity, and validation required every source fact ID to appear
|
|
exactly once.
|
|
|
|
Inputs:
|
|
|
|
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
|
- 33 fact items.
|
|
|
|
Model / configuration:
|
|
|
|
- `qwen3.5:9B`
|
|
- Ollama endpoint: `http://127.0.0.1:11434/api/generate`
|
|
- Thinking disabled.
|
|
- One LLM call.
|
|
- `num_ctx=32768`
|
|
- `num_predict=4096`
|
|
|
|
Result:
|
|
|
|
- Runtime: 390.119 seconds on the current machine.
|
|
- Merged fact groups: 1.
|
|
- Source facts involved in merges: 2.
|
|
- Singleton fact groups: 31.
|
|
- Validation: passed.
|
|
- No source fact was lost or duplicated.
|
|
- Non-fact categories remained unchanged.
|
|
|
|
Accepted merge:
|
|
|
|
- `fact_0025` + `fact_0031`
|
|
- Canonical statement: "Der Leiter F&E führt die Projektliste auf dem
|
|
zweiwöchentlichen Schnittstellen-Stand-Up."
|
|
|
|
Decision:
|
|
|
|
Semantic Consolidator V0 is complete for its current narrow scope:
|
|
conservative facts-only semantic duplicate detection with source evidence
|
|
preserved. The selected `report.md` and `consolidated_extractions.json`
|
|
benchmark artifacts should be versioned for later comparison. The raw model
|
|
response remains a local diagnostic artifact and is not versioned.
|
|
|
|
Lessons learned:
|
|
|
|
Semantic duplicate consolidation is technically viable and conservative enough
|
|
for continued evaluation, but broader semantic synthesis remains a separate
|
|
future stage. The measured runtime is useful for this machine and run, but
|
|
should not be generalized into a universal benchmark.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/consolidation/consolidate_facts.py`
|
|
- `prompts/consolidate_facts.md`
|
|
- `tests/test_consolidate_facts.py`
|
|
- `samples/benchmarks/semantic_consolidator_v0/report.md`
|
|
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
|
|
- Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt`
|
|
|
|
## EXP-0023 - Responsibility attribution integrity
|
|
|
|
Status: Running
|
|
|
|
Date or period: 2026-07-31
|
|
|
|
Hypothesis:
|
|
|
|
Protocol generation is operationally unsafe if responsibility, ownership or
|
|
departmental role attribution is inferred from discussion context rather than
|
|
explicit meeting evidence.
|
|
|
|
Setup:
|
|
|
|
The real-life Working Protocol Renderer V2 benchmark was inspected against the
|
|
consolidated input. A false assignment connected a Marketing participant to
|
|
Business Development criteria work even though the participant's contribution
|
|
was critical or reluctant and did not establish acceptance of that task.
|
|
|
|
Inputs:
|
|
|
|
- `samples/benchmarks/working_protocol_renderer_v2/working_protocol.md`
|
|
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
|
|
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
|
|
|
Model / configuration:
|
|
|
|
- Not rerun for this finding.
|
|
- Finding is based on existing benchmark artifacts.
|
|
|
|
Result:
|
|
|
|
The false responsibility attribution is visible in the structured input before
|
|
rendering, so the issue is not merely stylistic renderer wording. The root
|
|
cause may originate earlier in extraction and then be preserved by
|
|
canonicalization and consolidation. Renderer guardrails are still required so
|
|
output views do not strengthen ambiguous ownership.
|
|
|
|
Decision:
|
|
|
|
Responsibility attribution is now treated as a critical project-wide
|
|
invariant. A person, team or department may be recorded as responsible only
|
|
when the evidence explicitly assigns, accepts or confirms that responsibility.
|
|
Discussion, expertise, objection, suggestion, thematic proximity, speaker
|
|
adjacency, organizational assumptions and likely job roles do not establish
|
|
ownership.
|
|
|
|
Lessons learned:
|
|
|
|
This class of error affects operational correctness, not only style. The
|
|
pipeline needs traceable attribution evidence and future schema support for
|
|
responsibility status such as explicit, accepted, proposed or unclear.
|
|
|
|
Evidence:
|
|
|
|
- `AGENTS.md`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `docs/data-models.md`
|
|
- `docs/output-views.md`
|
|
- `prompts/working_protocol.md`
|
|
- `tests/gold/responsibility_attribution_negative/`
|
|
|
|
## EXP-0024 - Working Protocol V2 contract visibility
|
|
|
|
Status: Running
|
|
|
|
Date or period: 2026-08-09
|
|
|
|
Target:
|
|
|
|
BUG-011 renderer-only regression using an already validated Semantic
|
|
Consolidator artifact.
|
|
|
|
Hypothesis:
|
|
|
|
The renderer failures are caused by provenance-heavy 199-350 KB JSON inputs
|
|
placing the leading prompt contract outside the model's effective evaluated
|
|
context. The preserved failures all report `prompt_eval_count=16386`, while
|
|
their outputs either echo trailing JSON or produce an unconstrained generic
|
|
category summary instead of the requested Working Protocol.
|
|
|
|
Iteration 1 change:
|
|
|
|
- Project every consolidated item to rendering-relevant semantic fields while
|
|
retaining every item and its category/text/responsibility/deadline/status
|
|
information.
|
|
- Generate the exact structural contract from renderer validator constants and
|
|
append it after the compact INPUT JSON.
|
|
- Replace the independently handwritten prompt skeleton with a reference to
|
|
that authoritative appended contract.
|
|
- Enforce the existing prompt rule that emitted sections must not be empty.
|
|
|
|
This is one renderer-contract prompt iteration. It does not change extraction,
|
|
canonicalization, semantic consolidation or responsibility semantics.
|
|
|
|
Validation before LLM run:
|
|
|
|
- 15 focused renderer tests pass.
|
|
- Tests cover contract generation, compact input projection, valid and invalid
|
|
headings, missing topic sections, empty sections, wrapper cleanup, malformed
|
|
Markdown and final-file write gating.
|
|
|
|
Decision:
|
|
|
|
Iteration 1 passed structural validation and wrote `working_protocol.md`, but
|
|
the quality sanity check found that the model omitted most of the ten supplied
|
|
decisions, two open questions and several action items. The structurally valid
|
|
result therefore was not accepted as BUG-011 verification.
|
|
|
|
Iteration 2 change:
|
|
|
|
- Add input-derived hidden coverage markers for every decision, action item and
|
|
open question.
|
|
- Require every priority item exactly once in its matching section.
|
|
- Validate missing, duplicate, unknown and wrong-section markers
|
|
deterministically.
|
|
- Keep facts and technical details condensable as background.
|
|
|
|
This is the second single prompt iteration. It responds to the concrete
|
|
omission failure observed in Iteration 1 without changing upstream semantics or
|
|
inventing renderer content.
|
|
|
|
Iteration 2 pre-run validation:
|
|
|
|
- 18 focused renderer tests pass, including exact required-item coverage and
|
|
wrong-section rejection.
|
|
|
|
Decision:
|
|
|
|
Iteration 2 initially exhausted the fixed 4,096-token renderer output budget
|
|
after emitting all decisions and most action items. Adaptive renderer budgeting
|
|
resolved to 8,192 tokens for this input. The final run stopped normally after
|
|
3,709 evaluated output tokens.
|
|
|
|
The final renderer-only regression passed strict validation and wrote
|
|
`working_protocol.md`. Exact coverage was 10/10 decisions, 36/36 renderable
|
|
action items and 23/23 open questions, each once in its matching section. One
|
|
structurally empty action item whose task, responsible, deadline and evidence
|
|
were all null was recorded and excluded rather than fabricated. Optional
|
|
background markers were accepted only when they referred to real projected
|
|
input items.
|
|
|
|
Accept the compact renderer input, validator-derived trailing contract,
|
|
priority-item coverage markers, empty-section validation and adaptive renderer
|
|
output sizing as the BUG-011 baseline. This establishes structural reliability
|
|
and priority-item coverage, not complete protocol prose quality.
|
|
|
|
Evidence:
|
|
|
|
- `prompts/working_protocol.md`
|
|
- `src/meeting_lab/protocol/render_working_protocol.py`
|
|
- `tests/test_render_working_protocol.py`
|
|
- `samples/benchmarks/progeo_qwen35_9b_20260805_133942/working_protocol/`
|
|
- `samples/benchmarks/progeo_qwen35_35b_a3b_20260805_133942/working_protocol/`
|
|
|
|
## EXP-0025 — BUG-015 classification precision
|
|
|
|
Date: 2026-08-09
|
|
|
|
Target: Progeo-derived Decision, Action Item and Open Question precision cases.
|
|
|
|
Model/configuration: `qwen3.5:9B`, temperature 0, `num_ctx=32768`.
|
|
|
|
Tests were created before prompt changes. A Decision-only evidence threshold
|
|
kept the explicit Dr. Schlummer rejection and omitted an option and preference.
|
|
Adding Action and Open Question definitions improved several negatives but was
|
|
not stable: the model alternately promoted an unaccepted Textor suggestion or
|
|
moved rejected candidates into Open Questions. Moving the standalone category
|
|
prompts after the transcript made the partial Decision schema dominate and
|
|
misclassified true Action Items as Decisions in two consecutive runs.
|
|
|
|
The final iteration replaced the competing standalone category prompts with a
|
|
single unified classification contract after the transcript. It preserved the
|
|
assigned Nina task and ownerless established CET work, and prevented
|
|
cross-category leakage in the focused case, but still emitted the unaccepted
|
|
Textor suggestion as an Action Item. Further prompt iterations were stopped in
|
|
accordance with the Gold Standard methodology.
|
|
|
|
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
|
|
remains Open; no phrase-specific deterministic filter was introduced.
|
|
|
|
## EXP-0027 — Evidence-near observation extraction
|
|
|
|
Date: 2026-08-18
|
|
|
|
Hypothesis: `qwen3.5:9B` can more reliably extract evidence-near linguistic and
|
|
semantic properties than directly synthesize protocol-level events, outcomes,
|
|
actions and unresolved issues. This isolated experiment stops before semantic
|
|
interpretation and does not connect to the production pipeline.
|
|
|
|
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
|
|
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
|
|
source evidence, explicit targets, a five-value relation vocabulary, modality,
|
|
temporality, evaluation, agreement, responsibility/person, uncertainty,
|
|
clarification need and free-text scope. It contains no protocol-level category
|
|
field. The validator requires sequential observation IDs, known evidence IDs,
|
|
backward-only valid observation targets, closed categorical vocabularies,
|
|
consistent responsibility/person pairs, and one canonical absence form: JSON
|
|
null for person and `absent` for scope.
|
|
|
|
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
|
`num_predict=4096`, no retries. All nine cases ran exactly once, for nine LLM
|
|
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
|
|
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
|
|
responses, parsed observations, validation results, Ollama metadata and
|
|
comparisons are preserved under
|
|
`/tmp/meeting-lab-evidence-observations-v1-20260818/`.
|
|
|
|
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
|
|
Five cases failed structure because the model represented a single target as a
|
|
one-element list, usually `["discussion_subject"]`; the accepted schema permits
|
|
a list only for two or more jointly referenced observations. Several responses
|
|
also copied the relation label `limits_scope` into the free-text scope field.
|
|
These were systematic model-output errors, not transport or parser failures.
|
|
The prompt and run were not retried or tuned.
|
|
|
|
Human semantic review of the preserved raw responses:
|
|
|
|
| Case | Verdict | Main result |
|
|
| --- | --- | --- |
|
|
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
|
|
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
|
|
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
|
|
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
|
|
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
|
|
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
|
|
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
|
|
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
|
|
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
|
|
|
|
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override
|
|
strict structural failures; it separates useful semantic signal from schema
|
|
compliance.
|
|
|
|
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving
|
|
closer to evidence reduced some direct promotion behavior: the geometry mention
|
|
did not become work, the washing disadvantage did not become an unresolved
|
|
issue, both alternatives in B remained uncommitted, and the publication query
|
|
did not become an assignment. However, the important promotion errors did not
|
|
disappear. C acquired unsupported acceptance, and G still promoted a personal
|
|
preference into rejection. Positive cases were only partly preserved: explicit
|
|
rejection was recognized but incorrectly linked; trial-only language was kept
|
|
but responsibility was invented; Nina's request and commitment were recognized
|
|
but responsibility states were wrong; and the publication issue was recognized
|
|
but its uncertainty was contradicted by rejection.
|
|
|
|
Result: **B — evidence-near extraction is promising, but specific observation
|
|
dimensions remain unreliable.** Target/relation selection, scope attachment,
|
|
responsibility state/person attribution, and agreement versus uncertainty are
|
|
not reliable enough to justify designing the later interpretation stage yet.
|
|
No production integration or later interpretation stage was implemented.
|
|
|
|
## EXP-0028 — Evidence-Near Observation Extraction V2
|
|
|
|
Date: 2026-08-19
|
|
|
|
V2 tested whether `qwen3.5:9B` preserves the evidence needed by a later
|
|
controlled interpretation stage when direct responsibility, agreement and
|
|
semantic graph relations are removed. Responsibility was replaced by explicit
|
|
participant/discourse facts (`speaker`, `named_person`, `addressee`, singular
|
|
self-reference, collective `we`, and impersonal person reference). Agreement
|
|
was replaced by explicit affirmation, explicit negation and determination
|
|
statement signals. Graph relations were reduced to nullable scalar
|
|
`refers_to`; scope became free-text `qualifier` plus nullable scalar
|
|
`limits_target`. No later derivation stage was implemented.
|
|
|
|
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
|
|
human interpretations. It contains no responsibility, agreement, action,
|
|
decision, open-question, accepted-trial, rejected-alternative or protocol
|
|
eligibility fields. Validation enforces known evidence IDs, sequential unique
|
|
observation IDs, backward-only scalar references, closed vocabularies, boolean
|
|
participant flags, JSON-nullable participant/qualifier/reference fields and no
|
|
string `"null"`.
|
|
|
|
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
|
`num_predict=4096`, no retries or voting. A launch-path defect was corrected
|
|
before the live run; the failed launch made zero model calls. A sandbox-blocked
|
|
localhost attempt also made zero model calls. The completed run called the
|
|
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
|
|
Persistent prompts, Gold input and expectations, raw and parsed model output,
|
|
validation, automatic comparison, Ollama metadata and human evaluation are in
|
|
`artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/`.
|
|
|
|
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
|
|
were schema-invalid. The dominant serialization pattern was use of `present`
|
|
instead of the specified `explicit` for affirmation/negation; E additionally
|
|
used `none` instead of `absent` for a determination signal, while D emitted the
|
|
separate uncertainty concept as an invalid modality. These errors are
|
|
contract violations, although most `present`/`explicit` differences are
|
|
deterministically normalizable without changing meaning. A and C were valid
|
|
JSON/schema outputs but had critical semantic mismatches.
|
|
|
|
Human semantic review:
|
|
|
|
| Case | Verdict | Main result |
|
|
| --- | --- | --- |
|
|
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
|
|
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; `present` caused schema failure. |
|
|
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
|
|
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
|
|
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
|
|
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
|
|
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
|
|
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
|
|
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
|
|
|
|
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced
|
|
V1 promotion errors: collective speech and speaker identity no longer became
|
|
individual responsibility; personal preference no longer became a group-level
|
|
rejection field; an information question did not become work; and explicit
|
|
negation/affirmation survived as separate evidence. Useful participant evidence
|
|
also survived strongly in H and collective-speech evidence in F.
|
|
|
|
Simplification did not make all evidence-near dimensions reliable. Scalar
|
|
references and `limits_target` were almost entirely omitted, qualifiers were
|
|
usually omitted, committed modality was missed in E, F and H, and C/G repeated
|
|
important modality, uncertainty and participant-feature errors. Some positive
|
|
semantic information therefore survived only in free-text `content`, not in
|
|
the structural signals a controlled derivation stage would need.
|
|
|
|
Result: **B — V2 is materially better, but specific evidence-near dimensions
|
|
still require refinement.** Direct responsibility, agreement and graph-relation
|
|
classification should remain excluded. Before designing the derivation stage,
|
|
the next work should examine the minimal reliable representation of explicit
|
|
reference/scope limitation, commitment modality and participant deixis. No
|
|
production integration, Progeo run or derivation implementation was performed.
|
|
|
|
## EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
|
|
|
|
Date: 2026-08-19
|
|
|
|
Hypothesis: `qwen3.5:9B` is substantially more reliable when the first semantic
|
|
stage preserves meeting meaning as atomic natural-language observations with
|
|
provenance and only simple participant information, without classifying or
|
|
deriving higher-level meeting semantics.
|
|
|
|
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
|
|
observation contains exactly `observation_id`, `evidence_id`, `content`,
|
|
`speaker`, nullable `named_person`, and nullable `addressee`. It contains no
|
|
modality, temporality, evaluation, affirmation, negation, determination,
|
|
uncertainty, clarification, responsibility, agreement, relation, reference,
|
|
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
|
|
Instead, the prompt asks for conservative atomic content that retains hedges,
|
|
conditions, personal/collective/impersonal language, requests, acceptances,
|
|
rejections, quantities, deadlines and boundaries in natural language.
|
|
|
|
Structural validation is intentionally small: exact schema keys, non-empty
|
|
observations/content, unique `obs_N` identifiers, known evidence IDs, speaker
|
|
matching its evidence, explicit named people/addressees, and no string
|
|
`"null"`. Human semantic preservation against per-case requirements is the
|
|
primary evaluation; wording differences do not fail a case.
|
|
|
|
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
|
`num_predict=4096`, no retries, voting or per-case tuning. One sandbox-blocked
|
|
localhost launch made zero model calls. The completed run made exactly nine
|
|
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
|
|
structural validation. Persistent source evidence, semantic requirements,
|
|
exact prompts, raw and parsed responses, validation, Ollama metadata and human
|
|
evaluation are stored under
|
|
`artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/`.
|
|
|
|
Human semantic preservation results:
|
|
|
|
| Case | Verdict | Main result |
|
|
| --- | --- | --- |
|
|
| A | PASS | Preserved `kann`, `vielleicht`, tentative follow-up, and explicit `Dann` sequence without commitment. |
|
|
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
|
|
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
|
|
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
|
|
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
|
|
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
|
|
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, `wenn überhaupt`, impersonal checking need and no decision, but dropped collective `wir` from who would receive contaminated material. |
|
|
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
|
|
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
|
|
|
|
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
|
|
“that we receive contaminated material back” into an impersonal passive phrase;
|
|
the risk itself remained hypothetical. H translated `Freitag` to `Friday`, a
|
|
harmless wording difference. I retained two compound observations rather than
|
|
splitting every proposition, but all required semantic boundaries and
|
|
dependencies remained explicit.
|
|
|
|
Compared with V2, categorical-field removal improved content preservation in
|
|
A, G and I: A retained `Dann`; G retained `wenn überhaupt`, personal `Ich` and
|
|
impersonal `Man`; I retained publication purpose and all boundaries. It also
|
|
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
|
|
strengthening into responsibility, group rejection, established work or
|
|
assigned clarification. F and H remain sufficiently complete in natural
|
|
language for a later interpretation experiment. No useful meaning was shown to
|
|
depend on the removed fields; the V3 content retained the useful signals that
|
|
V2's fields had attempted to encode.
|
|
|
|
Result: **A — MINIMAL FIRST STAGE ACCEPTED.** On A-I, minimal atomic content
|
|
with evidence provenance and simple participants is sufficiently reliable to
|
|
be the candidate first semantic stage. A later bounded experiment may examine
|
|
controlled semantic interpretation, but no derivation stage, production
|
|
integration or Progeo run was implemented here.
|
|
|
|
## EXP-0030 — V3 model comparison: qwen3.5:9B vs qwen3.6:35B-A3B
|
|
|
|
Date: 2026-08-19
|
|
|
|
This controlled comparison reran the accepted V3 minimal semantic-preservation
|
|
experiment unchanged with the locally installed `qwen3.6:35B-A3B`. It used the
|
|
same implementation, A-I fixture, evidence, prompt, minimal schema, temperature
|
|
0, `think=false`, `num_ctx=16384`, `num_predict=4096`, no retries, no voting and
|
|
one call per case. The completed run made exactly nine calls in 96.392 seconds.
|
|
Artifacts are preserved under
|
|
`artifacts/experiments/evidence_observations_v3/20260819_v3_qwen36_35b_a3b_single_run/`.
|
|
|
|
Structural validation passed 8/9 cases. D was semantically faithful but invalid
|
|
because the model copied transcript speakers Antonius and Martin into
|
|
`named_person`, although those names were not explicitly named within their
|
|
utterances. Human semantic preservation produced 7 PASS, 1 PARTIAL and 1 FAIL:
|
|
|
|
| Case | Verdict | Main result |
|
|
| --- | --- | --- |
|
|
| A | PASS | Preserved `can`, `perhaps`, tentative `would`, explicit `then` and no commitment, but changed the content language to English. |
|
|
| B | PARTIAL | Preserved both approaches overall, but removed `Oder` from the 20-20 observation and locally strengthened it into collective planned conduct. |
|
|
| C | PASS | Preserved Tim's tentative personal Textor contact and possible follow-up without established work. |
|
|
| D | PASS | Preserved possibility, uncertainty, process and energy consequence; structural failure was confined to invalid speaker-as-named-person values. |
|
|
| E | PASS | Preserved cost, collective rejection/non-pursuit and later determination, though the confirmation dropped explicit `Ja`. |
|
|
| F | FAIL | Preserved quantity, timing and trial/series boundaries, but changed collective `wir` into “Martin suggests” and “Tim agrees,” inventing individual proposal/agreement meaning. |
|
|
| G | PASS | Preserved hypothetical risk, collective recipient `wir`, Martin's personal stance, `wenn überhaupt`, conditional Technikum, impersonal `man müsste` and no decision/owner. |
|
|
| H | PASS | Preserved request, addressee, Friday, explicit acceptance and future personal commitment without a responsibility field; content was English. |
|
|
| I | PASS | Preserved local/pure-production and comparison boundaries, five-degree difference, upstream effort, publication purpose, unresolved permission and clarification without assignment. |
|
|
|
|
Direct comparison:
|
|
|
|
| Measure | `qwen3.5:9B` | `qwen3.6:35B-A3B` |
|
|
| --- | ---: | ---: |
|
|
| Structurally valid | 9/9 | 8/9 |
|
|
| Human PASS | 8 | 7 |
|
|
| Human PARTIAL | 1 | 1 |
|
|
| Human FAIL | 0 | 1 |
|
|
| Observations | 28 | 30 |
|
|
| Runtime | 40.074 s | 96.392 s |
|
|
| LLM calls | 9 | 9 |
|
|
| Prompt-evaluation tokens | 6,268 | 6,268 |
|
|
| Evaluation tokens | 2,590 | 2,718 |
|
|
|
|
The larger model fixed the 9B weakness in G by preserving collective `wir`, and
|
|
it split I's compound production/effort observations more cleanly. Those gains
|
|
did not offset regressions: B was locally strengthened, F materially converted
|
|
collective conduct into individual agreement, D violated the participant
|
|
schema, observation count increased, and runtime was 2.4 times higher. Both
|
|
models preserved German consistently in six of nine cases, but in different
|
|
cases; the 35B-A3B model changed A, F and H to English, while 9B changed C, F
|
|
and H wholly or partly to English.
|
|
|
|
Result: **D — REGRESSION.** `qwen3.6:35B-A3B` does not materially improve the
|
|
accepted minimal V3 first-stage preservation over `qwen3.5:9B`; it is worse on
|
|
the A-I comparison because of the F ownership-adjacent strengthening and lower
|
|
structural validity. This conclusion applies only to the minimal V3 first
|
|
stage and does not determine model choice for any later semantic derivation.
|
|
No production integration, derivation implementation or Progeo run occurred.
|
|
|
|
## EXP-0031 — Controlled Semantic Derivation H V0
|
|
|
|
Date: 2026-08-19
|
|
|
|
This isolated experiment tested the first controlled second-stage derivation
|
|
using only the accepted `qwen3.5:9B` V3 observations for case H. The derivation
|
|
LLM received the two V3 observations, not the transcript or Gold expectation.
|
|
Its deliberately narrow task was limited to recognizing whether `obs_1` is a
|
|
concrete request and whether `obs_2` explicitly commits its speaker to
|
|
substantially the same work. Its strict output schema forbids responsibility,
|
|
requested actor, establishment/status, Action Item, protocol, confidence and
|
|
generic relation/graph fields.
|
|
|
|
Deterministic code validates observation/evidence provenance, obtains the
|
|
requested actor only from the request observation's addressee, requires the
|
|
acceptance to follow the request, requires the accepting speaker to equal that
|
|
addressee, and establishes responsibility only after all semantic and
|
|
structural gates pass. A bounded weekday normalizer reconciles `Friday` and
|
|
`Freitag`, rejects conflicting weekdays, and separates the supported due date
|
|
from the normalized action text. No general temporal or action ontology was
|
|
introduced.
|
|
|
|
Twenty focused deterministic tests cover the positive H path and the required
|
|
negative invariants: request alone, acknowledgement/non-commitment, tentative
|
|
acceptance, different response speaker, different work, reversed order,
|
|
speaker/name/addressee alone, conflicting deadlines, unknown observation IDs,
|
|
inconsistent evidence provenance, forbidden semantic fields, malformed JSON
|
|
and persistent artifacts. The complete non-LLM suite passed 192/192.
|
|
|
|
Configuration: one `qwen3.5:9B` call, temperature 0, `think=false`,
|
|
`num_ctx=16384`, `num_predict=1024`, no retries or voting. The call took 11.765
|
|
seconds, with 469 prompt-evaluation and 124 evaluation tokens. The model
|
|
returned a valid recognition object: `obs_1` is a concrete request, `obs_2` is
|
|
an explicit commitment, and both concern substantially the same work. It
|
|
returned no responsibility or establishment judgment.
|
|
|
|
All deterministic gates passed. The final derived result is an established
|
|
action `Prüfung der Messdaten`, requested from and assigned to Nina, due
|
|
`Freitag`, supported by request `obs_1/e1` and acceptance `obs_2/e2`. The model
|
|
included `bis Friday` in its normalized request text; after the single call, a
|
|
deterministic-only bounded correction separated that already-recognized due
|
|
phrase from action content without changing the prompt, recognition schema,
|
|
semantic result or call count. Focused and complete non-LLM suites still
|
|
passed after this correction.
|
|
|
|
Artifacts are preserved under
|
|
`artifacts/experiments/controlled_semantic_derivation_h/20260819_h_qwen35_9b_single_run/`.
|
|
Result: the H mechanism succeeded. This establishes only that the narrow
|
|
request-plus-explicit-acceptance pattern can be recognized and gated for H; it
|
|
does not generalize the derivation architecture to other cases or semantic
|
|
categories. No production integration, other case run, semantic graph,
|
|
protocol derivation or Progeo run occurred.
|
|
|
|
## EXP-0032 — Request / Acceptance Gold V0
|
|
|
|
Status: Experimental; promising with semantic precision gaps
|
|
|
|
Date: 2026-08-20
|
|
|
|
This isolated regression experiment tested whether the EXP-0031 mechanism
|
|
generalizes beyond H. It used ten short synthetic cases containing only
|
|
V3-style observations. Evidence Observation V3 was neither called nor changed,
|
|
and the model received no raw transcript or expected result. The fixed
|
|
recognition schema permits only a nullable concrete request and nullable later
|
|
explicit personal commitment, plus the same-requested-work judgment and
|
|
normalized action text. Responsibility, requested actor, established status,
|
|
Action Item, protocol, confidence and generic graph fields remain forbidden.
|
|
|
|
Cases:
|
|
|
|
- RA-01 explicit positive acceptance: PASS.
|
|
- RA-02 paraphrased positive acceptance: PASS.
|
|
- RA-03 acknowledgement only: PASS.
|
|
- RA-04 tentative response: PASS.
|
|
- RA-05 different responder without personal acceptance: PASS.
|
|
- RA-06 explicit commitment to different work: PARTIAL. The model returned no
|
|
acceptance instead of recognizing a commitment with `same_requested_work`
|
|
false. The requested action correctly remained unestablished.
|
|
- RA-07 request without response: PASS.
|
|
- RA-08 collective commitment: PARTIAL. The model over-recognized the
|
|
collective `wir` statement as an explicit commitment, but no request existed
|
|
and deterministic gates prevented individual responsibility.
|
|
- RA-09 impersonal necessity: PARTIAL. The model over-recognized the impersonal
|
|
necessity as a concrete request, but the observation had no addressee and
|
|
deterministic gates prevented establishment.
|
|
- RA-10 tentative personal suggestion: PASS.
|
|
|
|
Configuration: exactly ten sequential `qwen3.5:9B` calls, one per case,
|
|
temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no retries,
|
|
no voting and no prompt change between cases. Summed call time was 23.754
|
|
seconds, with 4,826 prompt-evaluation tokens and 793 evaluation tokens. The
|
|
strict schema validated every response and no responsibility or establishment
|
|
field leaked into model output.
|
|
|
|
Both positive cases recognized the request, explicit commitment and same-work
|
|
relationship, including the paraphrased acceptance, and deterministically
|
|
established Clara as responsible with due date `Dienstag`. The model rendered
|
|
the normalized action in semantically equivalent English; evaluation therefore
|
|
checks the structural deterministic result exactly while treating normalized
|
|
action wording as evidence-near semantic text rather than requiring lexical
|
|
identity. Acknowledgement and tentative response were not promoted. Every
|
|
negative case remained unestablished, and no individual responsibility was
|
|
invented.
|
|
|
|
Recognition-level errors were two false positives (RA-08 commitment and RA-09
|
|
request) and one false negative (RA-06 different-work commitment). Final
|
|
established-action false positives and false negatives were both zero. The
|
|
overall result was seven PASS, three PARTIAL and zero FAIL.
|
|
|
|
Conclusion: the narrow request-plus-acceptance architecture remains promising
|
|
for established individual actions because deterministic addressee, ordering,
|
|
speaker, same-work, provenance and deadline gates contained all recognition
|
|
errors. The recognition layer is not yet precise enough to generalize: its
|
|
handling of collective commitment, impersonal necessity and commitments to
|
|
different work needs further isolated study. No production integration or
|
|
additional semantic category is justified by this result.
|
|
|
|
Artifacts are preserved under
|
|
`artifacts/experiments/request_acceptance_gold_v0/20260820_qwen35_9b_single_run/`.
|
|
|
|
## EXP-0033 — Collective Commitment Gold V0
|
|
|
|
Status: Experimental; architecturally successful with one contained
|
|
recognition false positive
|
|
|
|
Date: 2026-08-20
|
|
|
|
This isolated second-stage experiment tested whether an explicit collective
|
|
first-person commitment can establish an action without inventing an individual
|
|
owner. It used ten synthetic cases containing one minimal V3-style observation
|
|
each. Evidence Observation V3 was neither called nor changed, and the accepted
|
|
Request/Acceptance mechanism remained unchanged and independent.
|
|
|
|
The strict semantic schema contains exactly `observation_id`,
|
|
`commitment_form` and `normalized_action_text`. `commitment_form` is closed to
|
|
`individual_first_person`, `collective_first_person` and `none`. The model
|
|
cannot output responsibility, ownership, requested actor, establishment,
|
|
Action Item, protocol, confidence, relations, graphs, decisions or unresolved
|
|
issues. Deterministic code validates schema and provenance, requires collective
|
|
commitment plus non-empty action text, applies bounded deadline consistency and
|
|
explicit-negation gates, and only then sets `status: established`,
|
|
`commitment_scope: collective` and `responsible_person: null`.
|
|
|
|
Gold results:
|
|
|
|
- CC-01 explicit collective commitment: PASS; established, due `nächste
|
|
Woche`, no person.
|
|
- CC-02 individual commitment: PASS; correctly routed out of the collective
|
|
path.
|
|
- CC-03 tentative collective possibility: PASS; unestablished.
|
|
- CC-04 collective suggestion: PASS; unestablished.
|
|
- CC-05 impersonal necessity: PASS; unestablished.
|
|
- CC-06 passive future statement: PASS; unestablished.
|
|
- CC-07 collective rejection: PARTIAL. The model incorrectly returned
|
|
`collective_first_person`, but the deterministic negation gate detected
|
|
`nicht` and prevented establishment.
|
|
- CC-08 qualified collective commitment: PASS; established with `nur im
|
|
Technikum` preserved, null due and no person.
|
|
- CC-09 collective commitment without deadline: PASS; established with null
|
|
due and no person.
|
|
- CC-10 speaker ownership trap: PASS; established collectively while Martin
|
|
remained only the speaker and was not assigned ownership.
|
|
|
|
Configuration: exactly ten successful sequential `qwen3.5:9B` calls, one per
|
|
case, temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no
|
|
retries, no voting and no prompt change. There were zero technical failed
|
|
calls. Aggregate runner time was 10.504 seconds; summed per-call time was 10.500
|
|
seconds, with 4,267 prompt-evaluation tokens and 415 evaluation tokens.
|
|
|
|
The outcome was nine PASS, one PARTIAL and zero FAIL. There was one recognition
|
|
false positive and no recognition false negatives. No qualifier was lost, no
|
|
individual owner was invented, and no responsibility or status field leaked
|
|
into recognition. Bounded due handling preserved `nächste Woche` verbatim and
|
|
returned null when no deadline was present.
|
|
|
|
Conclusion: the collective-commitment path is architecturally successful for
|
|
this narrow Gold set. The deterministic negation gate contained the only model
|
|
error, and every successful collective result necessarily retained
|
|
`responsible_person: null`. This does not justify a generic commitment system,
|
|
production integration, group identity inference or another semantic category.
|
|
|
|
Artifacts are preserved under
|
|
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
|
|
|
|
## EXP-0034 — Explicit Rejection Gold V0
|
|
|
|
Status: Failed architecturally
|
|
|
|
Date: 2026-08-20
|
|
|
|
This isolated Stage-2 experiment tested the narrow evidence fact that a
|
|
concrete action, option, proposal or future course was explicitly rejected,
|
|
abandoned, discontinued or ruled out. It used twelve synthetic cases containing
|
|
one self-contained observation or one local target/rejection pair. Evidence
|
|
Observation V3 was not called or changed. The accepted Request/Acceptance and
|
|
Collective Commitment paths remained unchanged and were not invoked.
|
|
|
|
The strict semantic schema contains exactly `rejection_observation_id`,
|
|
`target_observation_id`, `rejection_form` and
|
|
`normalized_rejected_action_text`. `rejection_form` is closed to
|
|
`explicit_action_rejection` and `none`. A positive recognition requires a
|
|
known local target and non-empty normalized target; `none` requires both target
|
|
and normalized text to be null. Decision, outcome, topic-closure,
|
|
responsibility, ownership, protocol, confidence and graph fields are forbidden.
|
|
Target resolution is limited to the same observation or one earlier supplied
|
|
observation. Deterministic code validates schema, IDs, ordering and complete
|
|
provenance before emitting the narrow status `explicitly_rejected`.
|
|
|
|
`explicitly_rejected` means rejected by the cited evidence only. It is not yet
|
|
a final meeting decision or final topic outcome, does not close a topic, and
|
|
does not supersede an earlier commitment.
|
|
|
|
Gold results:
|
|
|
|
- RJ-01 explicit collective rejection with local target: PASS.
|
|
- RJ-02 explicit non-pursuit with paired target: PASS.
|
|
- RJ-03 self-contained collaboration rejection: FAIL. The model returned
|
|
`none`, producing one recognition false negative.
|
|
- RJ-04 personal preference: FAIL. The model promoted the preference to an
|
|
explicit rejection and derived an unsupported rejection.
|
|
- RJ-05 concern: PASS; remained a non-rejection.
|
|
- RJ-06 uncertainty: PASS; remained a non-rejection.
|
|
- RJ-07 negative recommendation: FAIL. The model promoted advice to an
|
|
explicit rejection and derived an unsupported rejection.
|
|
- RJ-08 deferral: PASS; remained a non-rejection.
|
|
- RJ-09 factual negation: PASS; remained a non-rejection.
|
|
- RJ-10 temporary non-action: FAIL. The model treated `erstmal noch nicht` as
|
|
abandonment and derived an unsupported rejection.
|
|
- RJ-11 explicit rejection with material scope: PASS. Real-plant and
|
|
Druckversuch scope were preserved.
|
|
- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option
|
|
was rejected; the Technikum alternative was not absorbed.
|
|
|
|
Configuration: exactly twelve successful sequential `qwen3.5:9B` calls, one
|
|
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
|
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
|
technical failed calls. Aggregate runner time was 15.518 seconds; summed
|
|
per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681
|
|
evaluation tokens.
|
|
|
|
The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced
|
|
three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03).
|
|
There were four strict target-field expectation mismatches: three were
|
|
consequences of false-positive rejection objects populating otherwise locally
|
|
correct antecedents, and one was the missing self-contained RJ-03 target. No
|
|
derived positive selected the wrong concrete antecedent. Qualifier-loss count
|
|
was zero, positive-alternative absorption count was zero, and no responsibility,
|
|
decision, outcome or topic-closure field leaked into model output.
|
|
|
|
Conclusion: the experiment is not architecturally successful. Deterministic
|
|
structural gates cannot contain a semantically well-formed false-positive
|
|
rejection with valid local target and provenance. The model did distinguish
|
|
concern, uncertainty, deferral and factual negation, and it handled scoped and
|
|
alternative-bearing positives correctly, but it did not reliably separate
|
|
explicit rejection from personal preference, advice or temporary non-action.
|
|
The current binary recognition `explicit_action_rejection | none` is
|
|
insufficient for reliable generalization.
|
|
No production integration, generic rejection system, prompt tuning or
|
|
cross-pattern reconciliation is justified.
|
|
|
|
Artifacts are preserved under
|
|
`artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/`.
|
|
|
|
## EXP-0035 — Negative Act Form V0
|
|
|
|
Status: Experimental; successful for form classification with normalization
|
|
limitations
|
|
|
|
Date: 2026-08-20
|
|
|
|
EXP-0034 failed because the binary `explicit_action_rejection | none` question
|
|
collapsed materially different negative acts. It missed self-contained
|
|
non-pursuit and promoted personal preference, recommendation and temporary
|
|
non-action to rejection. This isolated follow-up tested only whether those
|
|
evidence-near forms can be distinguished before any normative derivation. It
|
|
does not derive rejection, decision, outcome, topic closure, responsibility or
|
|
protocol status, and EXP-0034 remained unchanged.
|
|
|
|
The strict output schema contains exactly `observation_id`,
|
|
`negative_act_form` and `normalized_action_text`. The closed form vocabulary is
|
|
`explicit_non_pursuit`, `personal_preference`, `recommendation`,
|
|
`temporary_non_action` and `none`. Non-`none` forms require non-empty normalized
|
|
action text; `none` requires null. Rejection, status, decision, outcome,
|
|
responsibility and other normative fields are forbidden recursively. Local
|
|
context may resolve a candidate observation's pronoun, but the schema contains
|
|
no target relation and the experiment exposes no derivation function.
|
|
|
|
Gold results:
|
|
|
|
- NA-01 explicit non-pursuit: PARTIAL. The form was correct; `working with Dr.
|
|
Schlummer` omitted the continuation aspect from normalization.
|
|
- NA-02 paraphrased explicit non-pursuit: PASS.
|
|
- NA-03 personal preference: PARTIAL. The form was correct, but normalization
|
|
repeated `Ich würde das nicht machen` instead of resolving the real-plant
|
|
trial target.
|
|
- NA-04 negative recommendation: PARTIAL. The form was correct; the normalized
|
|
English action used the loose rendering `real asset` for `reale Anlage`.
|
|
- NA-05 temporary non-action: PASS.
|
|
- NA-06 concern only: PASS with `none` and null action text.
|
|
- NA-07 uncertainty: PASS with `none` and null action text.
|
|
- NA-08 factual negation: PASS with `none` and null action text.
|
|
|
|
Expected-versus-actual form confusion was entirely diagonal:
|
|
|
|
| Expected form | Actual form | Count |
|
|
| --- | --- | ---: |
|
|
| `explicit_non_pursuit` | `explicit_non_pursuit` | 2 |
|
|
| `personal_preference` | `personal_preference` | 1 |
|
|
| `recommendation` | `recommendation` | 1 |
|
|
| `temporary_non_action` | `temporary_non_action` | 1 |
|
|
| `none` | `none` | 3 |
|
|
|
|
Configuration: exactly eight successful sequential `qwen3.5:9B` calls, one
|
|
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
|
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
|
technical failures. Aggregate runner time was 8.688 seconds; summed per-call
|
|
time was 8.686 seconds, with 4,183 prompt-evaluation tokens and 310 evaluation
|
|
tokens.
|
|
|
|
The result was five PASS, three PARTIAL and zero FAIL. All eight
|
|
`negative_act_form` classifications matched Gold. There was no unsupported
|
|
semantic strengthening and no rejection, status, decision, outcome,
|
|
responsibility or topic-closure leakage. Normalized action meaning was fully
|
|
acceptable in five cases and imperfect in three.
|
|
|
|
Conclusion: the finer evidence-near form vocabulary successfully distinguished
|
|
the four semantic boundaries that defeated the binary rejection experiment in
|
|
this small Gold set. The result supports separating negative-act-form
|
|
recognition from later normative derivation, but local target normalization is
|
|
not yet uniformly reliable. It does not justify modifying EXP-0034, deriving
|
|
rejection, production integration or beginning cross-pattern reconciliation.
|
|
|
|
Artifacts are preserved under
|
|
`artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run/`.
|
|
|
|
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
|
|
|
|
Date: 2026-08-11
|
|
|
|
Hypothesis: the primary protocol should be a topic-oriented reconstruction of
|
|
the meeting rather than a category-oriented list of extracted information.
|
|
|
|
This first isolated prototype does not replace or connect to the production
|
|
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
|
|
transcript excerpts to `qwen3.5:9B` and requests Discussion Subjects. Each
|
|
subject may contain supported discourse events, an outcome with mandatory
|
|
scope, resulting actions and unresolved issues. Optional structures must be
|
|
omitted when absent. Every semantic object must reference known evidence IDs.
|
|
|
|
The strict experimental schema validates:
|
|
|
|
- non-empty subjects and globally unique semantic identifiers;
|
|
- a closed discourse-event vocabulary;
|
|
- non-empty, known and non-duplicated evidence references;
|
|
- outcome text, scope, certainty and evidence;
|
|
- action text, JSON-nullable responsibility/deadline and evidence;
|
|
- unresolved-issue text and evidence;
|
|
- omission rather than null or empty optional structures.
|
|
|
|
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only,
|
|
multiple options, unaccepted proposal, proposal with objection, rejected
|
|
alternative, trial-scoped acceptance, no-decision discussion, resulting Action
|
|
Item, and outcome plus unresolved issue. Evaluation targets semantic identity,
|
|
development, outcome scope, actions, unresolved issues, traceability and
|
|
absence of invented commitments rather than exact wording.
|
|
|
|
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
|
`num_predict=4096`. Each case received exactly one model call; there were no
|
|
model retries or prompt iterations. The nine completed calls took 59.251
|
|
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
|
|
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
|
|
failure artifacts were preserved under
|
|
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/` and
|
|
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/`. Two earlier launch
|
|
attempts made zero LLM calls: one failed on the script import path and one was
|
|
blocked by sandbox networking.
|
|
|
|
Human-reviewed results after correcting two objectively wrong Gold assumptions
|
|
without another model call:
|
|
|
|
| Case | Verdict | Reason |
|
|
| --- | --- | --- |
|
|
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled `considered_option` rather than `introduced_idea`. |
|
|
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
|
|
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
|
|
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
|
|
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
|
|
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
|
|
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
|
|
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
|
|
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
|
|
|
|
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case
|
|
F: the model distinguished acceptance for a bounded trial from acceptance as a
|
|
final solution. It also handled the explicit action in case H well. However,
|
|
the experiment failed systematically on sparse structured output, subject
|
|
grouping and restraint around absent outcomes/actions/unresolved issues. The
|
|
model frequently mirrored optional schema fields as empty/null values, treated
|
|
alternatives as outcomes, split one discussion into multiple subjects, or
|
|
invented open issues from mere non-selection.
|
|
|
|
The focused experiment is not promising enough to justify a real Progeo chunk
|
|
sanity check. No such run was performed, and no architecture is accepted on
|
|
the basis of this prototype. Further work should first analyze whether the
|
|
failure comes from the schema/prompt representation, the model's sparse-output
|
|
reliability, or the boundary between subject grouping and semantic synthesis.
|
|
It should not proceed through repeated prompt tuning against these nine cases.
|