- add AGENTS.md with development and prompt-engineering rules - add PROJECT_KNOWLEDGE.md summarizing current architecture and findings - add CHANGELOG.md - add ROADMAP.md - establish experiments.md as the project's experiment log - document Canonical Meeting Knowledge architecture - document Output Views and Knowledge Objects - capture accepted experimental results and engineering methodology
972 lines
26 KiB
Markdown
972 lines
26 KiB
Markdown
# Experiments
|
|
|
|
This file records durable technical experiments and findings for Meeting Lab.
|
|
It is not a diary and does not replace commit history.
|
|
|
|
## Status values
|
|
|
|
- Proposed: experiment idea exists, but no result is recorded.
|
|
- Running: experiment is in progress and no decision has been made.
|
|
- Accepted: finding is the current baseline or design conclusion.
|
|
- Rejected: hypothesis was tested and should not be repeated as-is.
|
|
- Superseded: finding was useful but has been replaced by a newer baseline.
|
|
|
|
## EXP-0001 - Whisper JSON interpretation
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-29
|
|
|
|
Hypothesis:
|
|
|
|
Whisper JSON should be chunked from its segment stream, not from the aggregate
|
|
top-level text field.
|
|
|
|
Setup:
|
|
|
|
`chunk_transcript.py` was updated to parse JSON input and prefer
|
|
`segments[*].text` when `segments` exists. A regression test supplies JSON with
|
|
both top-level `text` and separate segment texts.
|
|
|
|
Inputs:
|
|
|
|
- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`.
|
|
- Real Whisper artifacts under `samples/whisper/`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM.
|
|
|
|
Result:
|
|
|
|
The test verifies that the block stream is `["alpha", "beta", "gamma"]` and
|
|
does not include the aggregate `"alpha beta gamma"` text. The repository history
|
|
records this as the fix for the earlier failure where the first chunk contained
|
|
the complete transcript.
|
|
|
|
Decision:
|
|
|
|
When `segments` exists, `segments[*].text` is the authoritative transcript
|
|
stream. The top-level `text` field is only a fallback.
|
|
|
|
Lessons learned:
|
|
|
|
Whisper JSON is structured input. Treating it like plain text can duplicate the
|
|
entire transcript and invalidate downstream chunking.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
- `tests/test_chunking.py`
|
|
- Commit `4656523` - `Fix Whisper JSON chunk extraction`
|
|
|
|
## EXP-0002 - Technical transcript chunking baseline
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-29 to 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Sequential technical chunks around the configured target size can preserve the
|
|
transcript while keeping extraction calls small enough for local models.
|
|
|
|
Setup:
|
|
|
|
The chunker splits block-aligned text with configurable target, minimum,
|
|
maximum and overlap settings. Tests verify no duplicate later blocks when
|
|
overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
|
|
|
|
Inputs:
|
|
|
|
- Synthetic block list in `tests/test_chunking.py`.
|
|
- `samples/whisper/meeting_speech_cleaned.json`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM for chunking.
|
|
- Manifest uses default chunking behavior recorded in
|
|
`samples/whisper/meeting_speech_cleaned_chunks/manifest.json`.
|
|
|
|
Result:
|
|
|
|
With overlap set to zero, tests verify that all blocks appear exactly once. The
|
|
real sample manifest contains nine chunks, mostly near the configured target
|
|
size, with a smaller final chunk.
|
|
|
|
Decision:
|
|
|
|
Independent sequential chunks are the current technical baseline. One
|
|
normalized chunk per extraction call is the preferred extraction strategy.
|
|
|
|
Lessons learned:
|
|
|
|
Chunking solves model-size constraints only. It must not perform topic
|
|
detection or semantic merging.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
- `tests/test_chunking.py`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`
|
|
- `AGENTS.md`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
|
|
## EXP-0003 - Conservative transcript normalization
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Transcript cleanup should improve readability without changing meeting
|
|
semantics.
|
|
|
|
Setup:
|
|
|
|
The normalizer removes isolated filler sounds, immediate duplicate words or
|
|
short duplicate phrases, and redundant whitespace. It records changed blocks in
|
|
a JSON change log and explicitly preserves semantic content categories.
|
|
|
|
Inputs:
|
|
|
|
- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
- Change logs such as `chunk_01_changes.json`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM.
|
|
|
|
Result:
|
|
|
|
The implementation and generated change logs show a conservative policy:
|
|
negations, qualifiers, dates, numbers, responsibilities, technical statements,
|
|
deadlines, decisions and commitments are preserved.
|
|
|
|
Decision:
|
|
|
|
Normalization remains deterministic and low-risk. When uncertain, leave text
|
|
unchanged.
|
|
|
|
Lessons learned:
|
|
|
|
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove
|
|
semantic cues needed by extraction.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/normalization/normalize_transcript.py`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json`
|
|
- `docs/pipeline.md`
|
|
|
|
## EXP-0004 - Full-context topic segmentation
|
|
|
|
Status: Superseded
|
|
|
|
Date or period: 2026-07-21 to 2026-07-22
|
|
|
|
Hypothesis:
|
|
|
|
A single full-context topic segmentation call can identify topic boundaries in
|
|
a normalized transcript chunk.
|
|
|
|
Setup:
|
|
|
|
The initial segmentation prototype asked the model for topic changes and then
|
|
converted those boundaries into continuous, non-overlapping segments.
|
|
|
|
Inputs:
|
|
|
|
- `samples/chunks/chunk_01_normalized.txt`.
|
|
|
|
Model / configuration:
|
|
|
|
- Generated artifact records `qwen3:14b`.
|
|
|
|
Result:
|
|
|
|
The generated artifact contains 85 blocks, three topic-change boundaries and
|
|
four segments. The run metadata records a substantially longer elapsed time
|
|
than the later windowed artifact for the same input.
|
|
|
|
Decision:
|
|
|
|
Full-context segmentation was useful as a prototype, but it was superseded by
|
|
windowed segmentation and manual review tooling.
|
|
|
|
Lessons learned:
|
|
|
|
The prototype established the boundary-to-segment representation, but did not
|
|
settle segmentation quality.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/segmentation/segment_topics.py`
|
|
- `samples/chunks/chunk_01_normalized_segments.json`
|
|
- Commit `f234efc` - `Add initial topic segmentation prototype`
|
|
- Commit `889a4fe` - `Detect topic boundaries as continuous segments`
|
|
|
|
## EXP-0005 - Windowed topic segmentation and review
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-22
|
|
|
|
Hypothesis:
|
|
|
|
Windowed topic segmentation can reduce runtime and make boundary evaluation
|
|
more inspectable than a single full-context call.
|
|
|
|
Setup:
|
|
|
|
`segment_topics_windowed.py` analyzes overlapping windows and reports only
|
|
boundaries from the decision range. Python merges boundaries into continuous,
|
|
non-overlapping segments. `review_segmentation.py` renders each boundary with
|
|
neighboring transcript context for manual classification.
|
|
|
|
Inputs:
|
|
|
|
- `samples/chunks/chunk_01_normalized.txt`.
|
|
- Full meeting normalized chunks under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3.
|
|
- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3.
|
|
|
|
Result:
|
|
|
|
The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments
|
|
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
|
|
Manual review output shows that some boundaries were assessed as subtopics
|
|
rather than full topic changes. Full-meeting artifacts show one window per
|
|
already-small normalized chunk and two segments per chunk.
|
|
|
|
Decision:
|
|
|
|
Windowed segmentation and review tooling are accepted as prototype tooling, not
|
|
as a stable production segmentation stage.
|
|
|
|
Lessons learned:
|
|
|
|
Windowing improves inspectability and can reduce runtime, but it can also
|
|
cluster boundaries and over-segment. Manual review remains necessary.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/segmentation/segment_topics_windowed.py`
|
|
- `src/meeting_lab/segmentation/review_segmentation.py`
|
|
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
|
- `samples/chunks/chunk_01_normalized_windowed_segments_review.md`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
|
- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling`
|
|
|
|
## EXP-0006 - Qwen model comparison
|
|
|
|
Status: Accepted
|
|
|
|
Hypothesis:
|
|
|
|
Larger local Qwen-family models should improve meaningful extraction and
|
|
segmentation, but model size alone will not solve prompt or pipeline problems.
|
|
|
|
Setup:
|
|
|
|
Project work used smaller models for smoke checks and larger local models for
|
|
meaningful extraction or segmentation. Artifacts and project knowledge record
|
|
the currently useful model roles.
|
|
|
|
Inputs:
|
|
|
|
- Gold Standard scenarios under `tests/gold/`.
|
|
- Generated segmentation artifacts under `samples/`.
|
|
- Generated extraction artifacts under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- `qwen3:1.7b`: smoke-test model according to project knowledge.
|
|
- `qwen3.5:9b`: current meaningful extraction and segmentation model according
|
|
to project knowledge and generated full-meeting segmentation artifacts.
|
|
- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts.
|
|
|
|
Result:
|
|
|
|
The repository supports the conclusion that `qwen3.5:9b` is the meaningful
|
|
current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger
|
|
models and longer contexts may increase runtime substantially, but no hardware
|
|
benchmark suite is recorded.
|
|
|
|
Decision:
|
|
|
|
Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current
|
|
experiments. Do not assume model size alone fixes prompt or pipeline design.
|
|
|
|
Lessons learned:
|
|
|
|
Evaluation must separate model capability from prompt clarity, context
|
|
strategy, extraction schema and consolidation.
|
|
|
|
Evidence:
|
|
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
|
- `samples/chunks/chunk_01_normalized_segments.json`
|
|
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
|
|
|
## EXP-0007 - Thinking output and Ollama API behavior
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Thinking-capable Qwen models may return reasoning separately from the final
|
|
answer, and extraction parsing should not fail merely because extra text or
|
|
multiple JSON objects appear.
|
|
|
|
Setup:
|
|
|
|
The Ollama response reader checks `response`, chat-style `message.content` and
|
|
then `thinking`. The JSON parser tries a full parse first, then scans JSON
|
|
object candidates and returns the final valid object. A regression test covers
|
|
thinking text before final JSON.
|
|
|
|
Inputs:
|
|
|
|
- Synthetic parser test in `tests/test_extraction_protocol.py`.
|
|
|
|
Model / configuration:
|
|
|
|
- No LLM run in the test.
|
|
- Code path is used by Ollama extraction.
|
|
|
|
Result:
|
|
|
|
The parser can handle additional text and multiple JSON objects where the final
|
|
valid object is the intended answer. Current code still falls back to `thinking`
|
|
only if no usable response or message content is present.
|
|
|
|
Decision:
|
|
|
|
Keep parser robustness, but do not treat thinking output as the root cause of
|
|
all extraction failures.
|
|
|
|
Lessons learned:
|
|
|
|
API response shape and model output shape are separate concerns. Preserve raw
|
|
responses when diagnosing failures.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
|
- `tests/test_extraction_protocol.py`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
|
|
## EXP-0008 - JSON truncation and generation limits
|
|
|
|
Status: Accepted
|
|
|
|
Hypothesis:
|
|
|
|
Some extraction failures are caused by generation limits truncating JSON rather
|
|
than by prompt wording or parser behavior.
|
|
|
|
Setup:
|
|
|
|
A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The
|
|
generation limit was increased for the successful path. Exact failing limit is
|
|
not recorded in the repository; the current extractor default is verifiably
|
|
`--num-predict 8192`.
|
|
|
|
Inputs:
|
|
|
|
- Local extraction runs referenced by project knowledge.
|
|
- Current extraction CLI.
|
|
|
|
Model / configuration:
|
|
|
|
- `qwen3.5:9b`.
|
|
- Current extractor default: `num_predict=8192`.
|
|
|
|
Result:
|
|
|
|
Increasing the generation limit fixed the technical JSON failure. This was not
|
|
primarily a parser or prompt problem.
|
|
|
|
Decision:
|
|
|
|
When JSON is truncated, inspect raw output and generation limits before editing
|
|
prompts.
|
|
|
|
Lessons learned:
|
|
|
|
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong
|
|
first response if the model simply ran out of output tokens.
|
|
|
|
Evidence:
|
|
|
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `AGENTS.md`
|
|
|
|
## EXP-0009 - Minimal end-to-end pipeline
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
A minimal local pipeline can transform Whisper output into chunk extractions
|
|
and an interim protocol, proving the technical path before the final
|
|
architecture exists.
|
|
|
|
Setup:
|
|
|
|
The repository added cleanup, normalization, chunking, extraction and protocol
|
|
builder scripts, with sample generated artifacts.
|
|
|
|
Inputs:
|
|
|
|
- `samples/whisper/meeting_speech.json`
|
|
- `samples/whisper/meeting_speech_cleaned.json`
|
|
- Generated chunks and normalized chunks under
|
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
|
|
|
Model / configuration:
|
|
|
|
- Local Ollama extraction for chunk JSON.
|
|
- Windowed segmentation artifacts use `qwen3.5:9b`.
|
|
|
|
Result:
|
|
|
|
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
|
|
extraction JSON files, windowed segmentation artifacts and
|
|
`meeting_protocol.md`.
|
|
|
|
Decision:
|
|
|
|
The minimal pipeline is technically validated. The first protocol builder is an
|
|
interim validation tool, not the final architecture.
|
|
|
|
Lessons learned:
|
|
|
|
End-to-end execution exposed the next limitation: extraction output needs
|
|
consolidation and purpose-specific rendering.
|
|
|
|
Evidence:
|
|
|
|
- `scripts/clean_whisper_json.py`
|
|
- `src/meeting_lab/normalization/normalize_transcript.py`
|
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
|
- `src/meeting_lab/protocol/build_protocol.py`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/`
|
|
- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline`
|
|
|
|
## EXP-0010 - Gold Standard corpus
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Reproducible prompt engineering requires synthetic transcripts with explicit
|
|
expected semantic outputs.
|
|
|
|
Setup:
|
|
|
|
The Gold Standard corpus defines scenario directories with `transcript.txt`,
|
|
`expected.json` and README files describing ground truth and common model
|
|
mistakes. The runner validates schema keys and writes `actual.json` for a
|
|
scenario.
|
|
|
|
Inputs:
|
|
|
|
- Gold scenarios under `tests/gold/`.
|
|
|
|
Model / configuration:
|
|
|
|
- Runner requires an explicit Ollama model for LLM evaluation.
|
|
- Existing unit tests for the runner do not invoke Ollama.
|
|
|
|
Result:
|
|
|
|
The corpus gives stable semantics for decisions, facts, positions, todos,
|
|
questions, technical details and difficult mixed cases. Initial structured
|
|
transcripts are Phase 1 and easier than raw Whisper-style transcripts.
|
|
|
|
Decision:
|
|
|
|
Use Gold Standard tests as both regression tests and formal meeting-semantics
|
|
specification. Raw or unlabelled transcript cases remain later-phase work.
|
|
|
|
Lessons learned:
|
|
|
|
Without expected outputs, prompt changes cannot be evaluated reproducibly.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/`
|
|
- `scripts/run_gold_test.py`
|
|
- `tests/test_gold_runner.py`
|
|
- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests`
|
|
|
|
## EXP-0011 - Gold-test quality and unique ground truth
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
If a prompt produces unexpected behavior, the gold test itself may be ambiguous
|
|
and should be reviewed before the prompt is changed.
|
|
|
|
Setup:
|
|
|
|
Decision-focused scenarios were clarified during baseline creation. The current
|
|
methodology requires checking unique ground truth before changing prompts.
|
|
|
|
Inputs:
|
|
|
|
- `decision_simple`
|
|
- `decision_deferred`
|
|
- `decision_none`
|
|
- Gold methodology document.
|
|
|
|
Model / configuration:
|
|
|
|
- Prompt Version 2 baseline work.
|
|
|
|
Result:
|
|
|
|
`decision_simple` required clarification around the explicit agreement and
|
|
nearby non-decision wording. The earlier negative/deferral ambiguity is now
|
|
represented by distinct `decision_none` and `decision_deferred` scenarios in
|
|
the repository. Punctuation is not reliable evidence for Whisper transcripts;
|
|
agreement language and wording must carry the semantics.
|
|
|
|
Decision:
|
|
|
|
Ambiguous gold tests must be reviewed before prompt changes. Do not treat
|
|
punctuation as reliable evidence in real Whisper-style transcripts.
|
|
|
|
Lessons learned:
|
|
|
|
Bad gold tests create false prompt failures and can encourage overfitting.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
|
- `tests/gold/decision_simple/README.md`
|
|
- `tests/gold/decision_deferred/README.md`
|
|
- `tests/gold/decision_none/README.md`
|
|
- Commit `f7ad9ba`
|
|
|
|
## EXP-0012 - Decision taxonomy
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Decision extraction needs a formal taxonomy that distinguishes substantive
|
|
decisions from process decisions and non-decisions.
|
|
|
|
Setup:
|
|
|
|
The decision definition document and decision prompt define included and
|
|
excluded categories. Gold tests cover explicit decisions, true no-decision
|
|
cases and deferrals.
|
|
|
|
Inputs:
|
|
|
|
- `tests/gold/DECISION_DEFINITION.md`
|
|
- `prompts/decisions.md`
|
|
- Decision gold scenarios.
|
|
|
|
Model / configuration:
|
|
|
|
- Prompt Version 2 baseline.
|
|
|
|
Result:
|
|
|
|
Accepted decision categories include substantive decisions, organizational
|
|
decisions, process decisions, approvals, rejections, deferrals, explicit
|
|
decisions not to decide yet and explicit agreement to gather more information
|
|
before deciding. Opinions, preferences and proposals without agreement are not
|
|
decisions.
|
|
|
|
Decision:
|
|
|
|
"No decision was reached" and "the decision was deferred" are distinct semantic
|
|
outcomes.
|
|
|
|
Lessons learned:
|
|
|
|
Deferral can be a valid process decision even when the substantive topic remains
|
|
unresolved.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/DECISION_DEFINITION.md`
|
|
- `tests/gold/decision_deferred/`
|
|
- `tests/gold/decision_none/`
|
|
- `prompts/decisions.md`
|
|
|
|
## EXP-0013 - Prompt engineering methodology
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Prompt iteration needs strict experimental controls to prevent regression,
|
|
overfitting and arbitrary prompt churn.
|
|
|
|
Setup:
|
|
|
|
The methodology was documented alongside the Gold Standard corpus and later
|
|
summarized for agents.
|
|
|
|
Inputs:
|
|
|
|
- Gold scenarios.
|
|
- Prompt files.
|
|
|
|
Model / configuration:
|
|
|
|
- Applies to all prompt experiments.
|
|
|
|
Result:
|
|
|
|
The accepted method is one prompt change per iteration, one target test at a
|
|
time, immediate validation, no regressions, no `expected.json` edits merely to
|
|
force a pass, no test-specific prompt hacks, stopping after two consecutive
|
|
non-improving iterations, and verifying unique ground truth before prompt
|
|
changes.
|
|
|
|
Decision:
|
|
|
|
Prompt changes are controlled experiments. See `AGENTS.md` for agent operating
|
|
rules.
|
|
|
|
Lessons learned:
|
|
|
|
Most prompt changes are not isolated unless the experiment explicitly constrains
|
|
the target and regression set.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
|
- `AGENTS.md`
|
|
- Commit `f7ad9ba`
|
|
|
|
## EXP-0014 - Decision Prompt Version 2
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Adding explicit process-decision language to the decision prompt can preserve
|
|
true decision detection while recognizing deferrals.
|
|
|
|
Setup:
|
|
|
|
Prompt Version 2 added explicit support for deferrals and process decisions.
|
|
The baseline was validated on three decision scenarios.
|
|
|
|
Inputs:
|
|
|
|
- `decision_simple`
|
|
- `decision_deferred`
|
|
- `decision_none`
|
|
|
|
Model / configuration:
|
|
|
|
- Prompt Version 2.
|
|
- Model used for validation is not recorded in the committed methodology.
|
|
|
|
Result:
|
|
|
|
The committed methodology records all three baseline scenarios as passing. The
|
|
current generated `actual.json` files also show the expected decision count for
|
|
these decision scenarios, although some non-decision categories remain less
|
|
complete.
|
|
|
|
Decision:
|
|
|
|
Prompt Version 2 is the current decision baseline. Explicit deferrals are
|
|
recognized as process decisions.
|
|
|
|
Lessons learned:
|
|
|
|
Decision-count success does not imply all categories are solved. Category-level
|
|
evaluation must continue.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
|
- `tests/gold/decision_simple/actual.json`
|
|
- `tests/gold/decision_deferred/actual.json`
|
|
- `tests/gold/decision_none/actual.json`
|
|
- `prompts/decisions.md`
|
|
|
|
## EXP-0015 - Difficult synthetic meeting
|
|
|
|
Status: Running
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
A deliberately adversarial synthetic meeting can expose extraction failures
|
|
that simple category tests miss.
|
|
|
|
Setup:
|
|
|
|
`evil_meeting` includes interruptions, corrections, absent referenced people,
|
|
near-decisions, changed positions and one expected explicit decision.
|
|
|
|
Inputs:
|
|
|
|
- `tests/gold/evil_meeting/`.
|
|
|
|
Model / configuration:
|
|
|
|
- Existing generated `actual.json`; exact model is not stored in the artifact.
|
|
|
|
Result:
|
|
|
|
The generated result found the expected FR-7 exclusion decision. It also
|
|
classified "do not migrate until the mapping table is checked" as an additional
|
|
decision. The current expected file treats that statement as a position, but
|
|
contextual review suggests it may be a valid process instruction or decision.
|
|
|
|
Decision:
|
|
|
|
Do not classify this as a simple model failure without reviewing the gold
|
|
standard. The scenario exposes a semantic gap in the expected output.
|
|
|
|
Lessons learned:
|
|
|
|
Difficult synthetic cases are valuable because they reveal ambiguity in the
|
|
specification as well as model mistakes.
|
|
|
|
Evidence:
|
|
|
|
- `tests/gold/evil_meeting/README.md`
|
|
- `tests/gold/evil_meeting/expected.json`
|
|
- `tests/gold/evil_meeting/actual.json`
|
|
- See EXP-0011 and EXP-0012.
|
|
|
|
## EXP-0016 - Context-size extraction comparisons
|
|
|
|
Status: Rejected
|
|
|
|
Hypothesis:
|
|
|
|
Increasing extraction context from one chunk to neighboring chunk groups should
|
|
make extraction more complete and therefore should become the baseline.
|
|
|
|
Setup:
|
|
|
|
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3
|
|
and 1+2+3.
|
|
|
|
Inputs:
|
|
|
|
- Normalized meeting chunks.
|
|
|
|
Model / configuration:
|
|
|
|
- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for
|
|
extraction experiments.
|
|
|
|
Result:
|
|
|
|
More context sometimes improved completeness, but it also shifted category
|
|
classification, added duplicates and reduced stability. No numeric winner is
|
|
recorded in the repository.
|
|
|
|
Decision:
|
|
|
|
Do not adopt larger extraction windows as the baseline. Independent chunk
|
|
extraction remains current strategy. Recover global context through
|
|
consolidation rather than continuously enlarging extraction windows.
|
|
|
|
Lessons learned:
|
|
|
|
Context size is not a monotonic quality knob. It changes the task the model is
|
|
performing.
|
|
|
|
Evidence:
|
|
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `AGENTS.md`
|
|
- `ROADMAP.md`
|
|
- See EXP-0002 and EXP-0018.
|
|
|
|
## EXP-0017 - Independent full-meeting chunk extraction
|
|
|
|
Status: Accepted
|
|
|
|
Hypothesis:
|
|
|
|
Extracting every normalized chunk independently can produce enough structured
|
|
material for a useful protocol draft.
|
|
|
|
Setup:
|
|
|
|
Nine normalized chunks were extracted into separate JSON files and then
|
|
rendered by the interim protocol builder.
|
|
|
|
Inputs:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt`
|
|
through `chunk_09_normalized.txt`.
|
|
|
|
Model / configuration:
|
|
|
|
- Existing extraction artifacts do not record model metadata.
|
|
|
|
Result:
|
|
|
|
The nine extraction files contain facts, decisions, todos, questions and
|
|
technical details. `meeting_protocol.md` aggregates them into a readable draft.
|
|
Duplicates, category shifts and synthesis became the dominant limitations.
|
|
|
|
Decision:
|
|
|
|
Independent chunk extraction is useful enough to keep as the baseline, but it
|
|
requires a consolidation stage.
|
|
|
|
Lessons learned:
|
|
|
|
Per-chunk extraction gives recall-oriented raw material. It does not by itself
|
|
produce a polished or canonical meeting representation.
|
|
|
|
Evidence:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json`
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
|
- `src/meeting_lab/protocol/build_protocol.py`
|
|
- See EXP-0018.
|
|
|
|
## EXP-0018 - Human protocol comparison and output-view split
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
One generated protocol cannot satisfy every use case; protocol output should be
|
|
separated by purpose and audience.
|
|
|
|
Setup:
|
|
|
|
The interim machine protocol was compared against the desired human protocol
|
|
shape and then the architecture was revised toward parallel output views.
|
|
|
|
Inputs:
|
|
|
|
- Interim `meeting_protocol.md`.
|
|
- Architecture and output-view documentation.
|
|
|
|
Model / configuration:
|
|
|
|
- Not applicable; this is a design evaluation.
|
|
|
|
Result:
|
|
|
|
The human protocol target is denser and organized by purpose and topic rather
|
|
than extraction categories. The machine extraction retains more context and is
|
|
useful for recall, but it is not the right direct source for a concise
|
|
distribution artifact or durable knowledge entry.
|
|
|
|
Decision:
|
|
|
|
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution
|
|
Protocol / Verteilerprotokoll and Knowledge Objects /
|
|
Wissensdatenbankeintrag.
|
|
|
|
Lessons learned:
|
|
|
|
Rendering is a separate concern from extraction and consolidation. Output views
|
|
must be parallel renderings of shared semantics, not transformations of one
|
|
another.
|
|
|
|
Evidence:
|
|
|
|
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
|
- `docs/output-views.md`
|
|
- `docs/architecture.md`
|
|
- `docs/pipeline.md`
|
|
- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture`
|
|
|
|
## EXP-0019 - Consolidation and Canonical Meeting Knowledge
|
|
|
|
Status: Accepted
|
|
|
|
Date or period: 2026-07-30
|
|
|
|
Hypothesis:
|
|
|
|
Extraction, consolidation and output rendering are separate problems and should
|
|
not be collapsed into one LLM prompt or one protocol file.
|
|
|
|
Setup:
|
|
|
|
The architecture was refined after the minimal pipeline and protocol draft
|
|
showed duplicate, synthesis and audience-specific rendering limitations.
|
|
|
|
Inputs:
|
|
|
|
- Extraction JSON artifacts.
|
|
- Interim protocol draft.
|
|
- Architecture and data-model documentation.
|
|
|
|
Model / configuration:
|
|
|
|
- Not applicable; this is an architectural conclusion.
|
|
|
|
Result:
|
|
|
|
The accepted design is a planned Canonical Meeting Knowledge layer as the
|
|
semantic source of truth, with Working Protocol, Distribution Protocol and
|
|
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
|
|
preserve evidence, reconcile category shifts and mark contradictions or
|
|
uncertainty.
|
|
|
|
Decision:
|
|
|
|
Consolidation is the next major engineering step after stable local extraction.
|
|
Canonical Meeting Knowledge and final output views are planned, not implemented.
|
|
|
|
Lessons learned:
|
|
|
|
Global meeting understanding should be recovered by consolidation over
|
|
evidence-bearing extractions, not by silently changing output views or expanding
|
|
LLM context indefinitely.
|
|
|
|
Evidence:
|
|
|
|
- `docs/architecture.md`
|
|
- `docs/pipeline.md`
|
|
- `docs/data-models.md`
|
|
- `docs/output-views.md`
|
|
- `PROJECT_KNOWLEDGE.md`
|
|
- `ROADMAP.md`
|
|
- Commit `5c03ed7`
|