Document project architecture and development methodology

- add AGENTS.md with development and prompt-engineering rules
- add PROJECT_KNOWLEDGE.md summarizing current architecture and findings
- add CHANGELOG.md
- add ROADMAP.md
- establish experiments.md as the project's experiment log
- document Canonical Meeting Knowledge architecture
- document Output Views and Knowledge Objects
- capture accepted experimental results and engineering methodology
This commit is contained in:
2026-07-31 08:25:16 +02:00
parent 5c03ed7efd
commit 09d125e54a
5 changed files with 1478 additions and 0 deletions
+124
View File
@@ -0,0 +1,124 @@
# AGENTS.md
Practical instructions for coding agents working in Meeting Lab.
## Project Purpose
Meeting Lab extracts and structures organizational knowledge from meeting
recordings. It is an experimental local discussion analyzer, not merely a
one-step meeting-protocol generator.
Successful approaches may later move into the Meeting Assistant project.
## Current Pipeline
Current and intended flow:
```text
Audio
-> Whisper
-> cleanup
-> normalization
-> chunking
-> local chunk extraction
-> consolidation
-> Canonical Meeting Knowledge
-> Output Views
```
Status:
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
local chunk extraction, interim Markdown protocol builder.
- Experimental/prototype: topic segmentation and review tooling.
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
Output Views.
## Architectural Principles
- Canonical Meeting Knowledge is the intended semantic source of truth.
- Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are
parallel output views.
- Output views must not silently change meaning. They may select, condense or
render information for an audience, but not invent new semantics.
- Extraction, consolidation, synthesis and rendering are separate concerns.
- Prefer small, testable processing stages over one monolithic LLM prompt.
- Current extraction strategy is one normalized chunk per LLM call.
- Do not expand context windows or redesign the extraction strategy without an
explicit experiment.
- Deterministic stages should remain deterministic where possible.
## Prompt Engineering Rules
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`:
1. Make only one prompt change per iteration.
2. Optimize only one target gold test case at a time.
3. Validate every prompt modification immediately.
4. Accept a prompt change only if it improves the target and causes no
regressions in previously passing gold tests.
5. Never modify `expected.json` merely to make a prompt pass.
6. Prompt engineering edits prompt files only; Python code changes require a
separate explicit task.
7. Maintain a prompt evolution log for every iteration.
8. Stop arbitrary iterations if small changes do not improve the test; analyze
the root cause.
9. Avoid gold-test overfitting. Prompt changes must generalize and must not
special-case one transcript.
10. Stop after two consecutive non-improving prompt iterations and classify the
root cause.
11. Verify whether the target gold test has objectively unique ground truth
before changing a prompt for unexpected behavior.
Current documented Prompt Version 2 decision baseline:
- `decision_simple`: passing
- `decision_deferred`: passing
- `decision_none`: passing
## LLM Execution Safety
- Never start a full multi-chunk LLM run unless explicitly requested.
- Before any LLM run, state the model, inputs, expected LLM-call count and
output location.
- Do not retry LLM calls automatically unless explicitly allowed.
- Do not download models automatically.
- Prefer small-scope validation runs.
- Never use generated output as committed source data.
- Preserve raw model responses when diagnosing parser or truncation failures.
- Do not run Ollama from unit tests.
## Development Rules
- Make small, focused changes.
- Preserve the existing architecture unless a redesign is explicitly requested.
- Add regression tests for bugs.
- Run non-LLM tests before committing when code changes are made.
- Do not commit generated transcripts, audio, extraction JSON, protocol output
or temporary files.
- Report files changed, tests run and assumptions.
- Do not commit or push unless explicitly requested.
## Repository Conventions
- `README.md`: project overview and current high-level status.
- `docs/`: architecture, pipeline, data model and output-view documentation.
- `prompts/`: extraction and segmentation prompts. Treat prompt edits as
controlled experiments.
- `tests/gold/`: Gold Standard corpus and semantic specification for extraction
behavior.
- `scripts/`: command-line support scripts such as Whisper cleanup and gold
test execution.
- `src/meeting_lab/normalization/`: deterministic transcript cleanup.
- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics.
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
extractor modules.
- `src/meeting_lab/consolidation/`: planned consolidation area.
- `src/meeting_lab/protocol/`: interim protocol rendering.
- `src/meeting_lab/models/`: current lightweight data models.
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
sample output as canonical source data.
+47
View File
@@ -0,0 +1,47 @@
# Changelog
## Unreleased
### Added
- Initial Meeting Lab project structure with source, docs, prompts, samples and
tests directories.
- Deterministic transcript normalization module.
- Technical transcript chunking with block-aligned chunk generation.
- Whisper JSON cleanup support.
- Local Ollama-based chunk extraction flow.
- Interim Markdown meeting protocol builder for technical validation.
- Initial and windowed topic segmentation prototypes plus review tooling.
- Gold Standard extraction corpus and gold-test runner.
- Prompt loading support and common/decision prompt baseline.
- Output-view architecture documentation for Canonical Meeting Knowledge,
Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
### Changed
- Refined the architecture from a single meeting protocol toward Canonical
Meeting Knowledge as the planned semantic source of truth.
- Clarified that Working Protocol, Distribution Protocol and Knowledge Objects
are parallel renderings, not derived from one another.
- Updated decision extraction semantics to include explicit process decisions
and deferrals.
- Simplified extraction prompt assembly around prompt files.
### Fixed
- Fixed Whisper JSON chunk extraction to prefer `segments[*].text` over the
aggregate top-level `text` field.
- Added chunking tests to ensure chunks do not duplicate later blocks when no
overlap is requested.
- Added parser handling for model responses that contain thinking text before
the final JSON object.
### Documentation
- Added architecture, pipeline and data-model documentation.
- Added output-view documentation.
- Added Gold Standard prompt-engineering methodology.
- Added formal decision-definition documentation.
- Added scenario README files for the Gold Standard corpus.
+151
View File
@@ -0,0 +1,151 @@
# Project Knowledge
This is a compact operational summary of the current Meeting Lab state.
## Objective
Meeting Lab develops and evaluates local methods for extracting structured
organizational knowledge from real meeting recordings and transcripts. The
project is a research and validation environment for a future Meeting
Assistant, not a finished product.
## Implemented Pipeline Stages
Implemented:
- Whisper JSON cleanup via `scripts/clean_whisper_json.py`.
- Transcript normalization in `src/meeting_lab/normalization/`.
- Technical chunking in `src/meeting_lab/chunking/`.
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation.
Experimental/prototype:
- Topic segmentation in `src/meeting_lab/segmentation/`.
- Windowed segmentation and review output in `samples/chunks/`.
- Gold Standard extraction corpus under `tests/gold/`.
Planned:
- Consolidation of extraction results.
- Canonical Meeting Knowledge implementation as the semantic source of truth.
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
## Source Tree
```text
src/meeting_lab/
chunking/ technical transcript chunking
consolidation/ planned merge/consolidation area
extraction/ current local LLM extraction flow
io/ lightweight file and JSON helpers
llm/ Ollama and prompt support
models/ current lightweight model definitions
normalization/ deterministic transcript cleanup
protocol/ interim Markdown protocol builder
segmentation/ experimental topic segmentation tooling
```
Supporting areas:
- `docs/`: architecture, pipeline, data models and output-view concepts.
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain
substantive extraction prompt text in the current tree.
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
- `samples/`: sample inputs and generated or experimental artifacts.
- `scripts/`: operational scripts for cleanup and gold-test execution.
## Current Model Strategy
The current extraction strategy is one normalized chunk per LLM call. This is
preferred over expanding context windows or asking one model call to analyze a
full meeting.
Known working models from current project notes and experiment practice:
- `qwen3:1.7b`: useful for smoke tests.
- `qwen3.5:9b`: useful for meaningful extraction and segmentation work.
LLM calls use Ollama locally. The current extractor defaults to `qwen3:8b`, but
validated work may specify another model explicitly.
## Important Findings
- Whisper JSON chunking must use `segments[*].text`, not only the top-level
`text` field.
- Independent chunk extraction is currently preferred.
- Larger context windows can change classification behavior and increase
instability.
- Extraction and consolidation are separate problems.
- Generation limits can truncate JSON.
- Qwen thinking may be returned separately by the Ollama API.
- Gold Standard tests are also a formal specification of meeting semantics.
- Raw model responses should be preserved when diagnosing parser or truncation
failures.
## Decision Taxonomy
Accepted decision semantics:
- A decision is an explicit agreement that creates a binding change in action,
process, responsibility, approval status, timing or next step.
- Included: substantive decisions, organizational decisions, process decisions,
approvals, rejections, deferrals, explicit agreement not to decide yet, and
explicit agreement to gather more information before deciding.
- Excluded: opinions, preferences, proposals without agreement, open questions,
current-state descriptions and explanations without commitment.
- A process decision to defer a substantive decision is still a decision.
- "No decision was reached" is different from "the group decided to defer the
decision."
Current Prompt Version 2 decision baseline:
- `decision_simple`: passing.
- `decision_deferred`: passing.
- `decision_none`: passing.
- Prompt Version 2 explicitly supports process decisions where the group agrees
to defer a substantive decision until more information is available.
## Canonical Knowledge Architecture
Canonical Meeting Knowledge is the planned semantic intermediate model and
future single source of truth. It should preserve topics, facts, decisions,
action items, open questions, positions, technical details, rationale,
uncertainty, contradictions and source evidence.
Output views are planned as independent renderings from that canonical model:
- Working Protocol / Arbeitsprotokoll: relatively complete, optimized for
recall and traceability.
- Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented,
optimized for circulation.
- Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge
optimized for reuse.
The current `meeting_protocol.md` builder is an interim technical validation
tool, not the final output-view architecture.
## Current Limitations
- Discussion Blocks are documented as a stable semantic unit but are not yet a
separate implemented pipeline artifact.
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
- Extraction is still a combined current flow, even though separate extractors
are the intended architecture.
- Consolidation is not implemented.
- Canonical Meeting Knowledge is documented but not implemented.
- Final output views are documented but not implemented.
- Most prompt files are placeholders except the common and decision prompts.
- Gold tests currently emphasize extraction semantics, especially decisions.
## Next Recommended Engineering Step
Stabilize repeatable local extraction evaluation before broadening the pipeline:
expand Gold Standard coverage by category, keep one-chunk extraction as the
baseline, and use small prompt experiments with immediate non-regression checks.
After extraction behavior is stable enough, implement consolidation with
evidence retention as the next major pipeline stage.
+185
View File
@@ -0,0 +1,185 @@
# Roadmap
No dates are assigned. Phases describe dependency order, not release promises.
## Phase 1 - Stable Local Extraction
Goal:
- Establish reliable per-chunk extraction behavior for core meeting semantics.
Deliverables:
- Stronger Gold Standard coverage across facts, positions, decisions, todos,
questions and technical details.
- Improved category prompts.
- Repeatable evaluation workflow.
- Documented prompt experiment log.
Prerequisites:
- Existing chunk extraction flow.
- Existing Gold Standard runner and methodology.
Out of scope:
- Full-transcript LLM extraction.
- Larger context-window strategy changes without an explicit experiment.
- Consolidation or final protocol rendering.
## Phase 2 - Consolidation
Goal:
- Merge independent extraction results into a coherent meeting-level
representation without losing evidence.
Deliverables:
- Duplicate merging.
- Evidence retention.
- Category-shift reconciliation, especially facts versus positions and
positions versus decisions.
- Contradiction and uncertainty markers.
- Consolidated meeting representation.
Prerequisites:
- Stable local extraction baseline.
- Gold tests that expose cross-chunk duplication and category shifts.
Out of scope:
- Final Canonical Meeting Knowledge schema.
- User-facing protocol polish.
- Retrieval or RAG integration.
## Phase 3 - Canonical Meeting Knowledge
Goal:
- Define and implement the semantic intermediate model that becomes the source
of truth for downstream outputs.
Deliverables:
- Canonical Meeting Knowledge schema.
- Source evidence and traceability fields.
- Clear distinction between durable knowledge and meeting-specific actions.
- Migration path from consolidated extraction JSON into the canonical model.
Prerequisites:
- Consolidation behavior that preserves evidence and uncertainty.
- Agreement on required semantic categories.
Out of scope:
- GUI.
- Export formats beyond those needed to validate the model.
- Knowledge-system storage design.
## Phase 4 - Output Views
Goal:
- Render purpose-specific outputs from Canonical Meeting Knowledge without
changing meaning.
Deliverables:
- Working Protocol / Arbeitsprotokoll renderer.
- Distribution Protocol / Verteilerprotokoll renderer.
- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export.
- Tests or checks showing that output views are parallel renderings of the same
canonical model.
Prerequisites:
- Implemented Canonical Meeting Knowledge.
- Clear audience and completeness rules for each output view.
Out of scope:
- Additional analysis during rendering.
- Deriving one output view from another.
- Retrieval integration.
## Phase 5 - Review and Quality Control
Goal:
- Add optional review stages that improve omission detection, consistency and
model selection.
Deliverables:
- Optional whole-transcript review.
- Omission detection.
- Consistency checks.
- Model comparison workflow.
- Hardware and runtime benchmarks.
Prerequisites:
- Stable extraction, consolidation and canonical model.
- Representative test meetings.
Out of scope:
- Automatic acceptance of review suggestions without evidence.
- Product UI work.
- Cloud deployment.
## Phase 6 - Productization
Goal:
- Turn the validated pipeline into a usable local workflow.
Deliverables:
- Recording/transcription workflow.
- FFmpeg integration.
- Meeting metadata capture.
- Participant entry.
- GUI.
- Stable deployment process.
- Export workflows.
Prerequisites:
- Stable pipeline stages and output views.
- Clear operational requirements for local use.
Out of scope:
- Enterprise knowledge retrieval.
- Future Meeting Assistant integration beyond export contracts.
- Cloud-first architecture.
## Phase 7 - Knowledge-System Integration
Goal:
- Reuse durable meeting knowledge in broader knowledge systems.
Deliverables:
- Structured Knowledge Objects.
- Retrieval-ready storage format.
- Future RAG integration path.
- Reuse contracts for Meeting Assistant and other knowledge systems.
Prerequisites:
- Canonical Meeting Knowledge and Knowledge Objects are implemented and stable.
- Durable knowledge is separated from meeting-specific actions and discussion
history.
Out of scope:
- Building a full enterprise search product inside Meeting Lab.
- Treating raw transcripts or generated protocols as the knowledge source of
truth.
+971
View File
@@ -0,0 +1,971 @@
# Experiments
This file records durable technical experiments and findings for Meeting Lab.
It is not a diary and does not replace commit history.
## Status values
- Proposed: experiment idea exists, but no result is recorded.
- Running: experiment is in progress and no decision has been made.
- Accepted: finding is the current baseline or design conclusion.
- Rejected: hypothesis was tested and should not be repeated as-is.
- Superseded: finding was useful but has been replaced by a newer baseline.
## EXP-0001 - Whisper JSON interpretation
Status: Accepted
Date or period: 2026-07-29
Hypothesis:
Whisper JSON should be chunked from its segment stream, not from the aggregate
top-level text field.
Setup:
`chunk_transcript.py` was updated to parse JSON input and prefer
`segments[*].text` when `segments` exists. A regression test supplies JSON with
both top-level `text` and separate segment texts.
Inputs:
- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`.
- Real Whisper artifacts under `samples/whisper/`.
Model / configuration:
- No LLM.
Result:
The test verifies that the block stream is `["alpha", "beta", "gamma"]` and
does not include the aggregate `"alpha beta gamma"` text. The repository history
records this as the fix for the earlier failure where the first chunk contained
the complete transcript.
Decision:
When `segments` exists, `segments[*].text` is the authoritative transcript
stream. The top-level `text` field is only a fallback.
Lessons learned:
Whisper JSON is structured input. Treating it like plain text can duplicate the
entire transcript and invalidate downstream chunking.
Evidence:
- `src/meeting_lab/chunking/chunk_transcript.py`
- `tests/test_chunking.py`
- Commit `4656523` - `Fix Whisper JSON chunk extraction`
## EXP-0002 - Technical transcript chunking baseline
Status: Accepted
Date or period: 2026-07-29 to 2026-07-30
Hypothesis:
Sequential technical chunks around the configured target size can preserve the
transcript while keeping extraction calls small enough for local models.
Setup:
The chunker splits block-aligned text with configurable target, minimum,
maximum and overlap settings. Tests verify no duplicate later blocks when
overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
Inputs:
- Synthetic block list in `tests/test_chunking.py`.
- `samples/whisper/meeting_speech_cleaned.json`.
Model / configuration:
- No LLM for chunking.
- Manifest uses default chunking behavior recorded in
`samples/whisper/meeting_speech_cleaned_chunks/manifest.json`.
Result:
With overlap set to zero, tests verify that all blocks appear exactly once. The
real sample manifest contains nine chunks, mostly near the configured target
size, with a smaller final chunk.
Decision:
Independent sequential chunks are the current technical baseline. One
normalized chunk per extraction call is the preferred extraction strategy.
Lessons learned:
Chunking solves model-size constraints only. It must not perform topic
detection or semantic merging.
Evidence:
- `src/meeting_lab/chunking/chunk_transcript.py`
- `tests/test_chunking.py`
- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`
- `AGENTS.md`
- `PROJECT_KNOWLEDGE.md`
## EXP-0003 - Conservative transcript normalization
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Transcript cleanup should improve readability without changing meeting
semantics.
Setup:
The normalizer removes isolated filler sounds, immediate duplicate words or
short duplicate phrases, and redundant whitespace. It records changed blocks in
a JSON change log and explicitly preserves semantic content categories.
Inputs:
- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`.
- Change logs such as `chunk_01_changes.json`.
Model / configuration:
- No LLM.
Result:
The implementation and generated change logs show a conservative policy:
negations, qualifiers, dates, numbers, responsibilities, technical statements,
deadlines, decisions and commitments are preserved.
Decision:
Normalization remains deterministic and low-risk. When uncertain, leave text
unchanged.
Lessons learned:
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove
semantic cues needed by extraction.
Evidence:
- `src/meeting_lab/normalization/normalize_transcript.py`
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json`
- `docs/pipeline.md`
## EXP-0004 - Full-context topic segmentation
Status: Superseded
Date or period: 2026-07-21 to 2026-07-22
Hypothesis:
A single full-context topic segmentation call can identify topic boundaries in
a normalized transcript chunk.
Setup:
The initial segmentation prototype asked the model for topic changes and then
converted those boundaries into continuous, non-overlapping segments.
Inputs:
- `samples/chunks/chunk_01_normalized.txt`.
Model / configuration:
- Generated artifact records `qwen3:14b`.
Result:
The generated artifact contains 85 blocks, three topic-change boundaries and
four segments. The run metadata records a substantially longer elapsed time
than the later windowed artifact for the same input.
Decision:
Full-context segmentation was useful as a prototype, but it was superseded by
windowed segmentation and manual review tooling.
Lessons learned:
The prototype established the boundary-to-segment representation, but did not
settle segmentation quality.
Evidence:
- `src/meeting_lab/segmentation/segment_topics.py`
- `samples/chunks/chunk_01_normalized_segments.json`
- Commit `f234efc` - `Add initial topic segmentation prototype`
- Commit `889a4fe` - `Detect topic boundaries as continuous segments`
## EXP-0005 - Windowed topic segmentation and review
Status: Accepted
Date or period: 2026-07-22
Hypothesis:
Windowed topic segmentation can reduce runtime and make boundary evaluation
more inspectable than a single full-context call.
Setup:
`segment_topics_windowed.py` analyzes overlapping windows and reports only
boundaries from the decision range. Python merges boundaries into continuous,
non-overlapping segments. `review_segmentation.py` renders each boundary with
neighboring transcript context for manual classification.
Inputs:
- `samples/chunks/chunk_01_normalized.txt`.
- Full meeting normalized chunks under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3.
- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3.
Result:
The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
Manual review output shows that some boundaries were assessed as subtopics
rather than full topic changes. Full-meeting artifacts show one window per
already-small normalized chunk and two segments per chunk.
Decision:
Windowed segmentation and review tooling are accepted as prototype tooling, not
as a stable production segmentation stage.
Lessons learned:
Windowing improves inspectability and can reduce runtime, but it can also
cluster boundaries and over-segment. Manual review remains necessary.
Evidence:
- `src/meeting_lab/segmentation/segment_topics_windowed.py`
- `src/meeting_lab/segmentation/review_segmentation.py`
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
- `samples/chunks/chunk_01_normalized_windowed_segments_review.md`
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling`
## EXP-0006 - Qwen model comparison
Status: Accepted
Hypothesis:
Larger local Qwen-family models should improve meaningful extraction and
segmentation, but model size alone will not solve prompt or pipeline problems.
Setup:
Project work used smaller models for smoke checks and larger local models for
meaningful extraction or segmentation. Artifacts and project knowledge record
the currently useful model roles.
Inputs:
- Gold Standard scenarios under `tests/gold/`.
- Generated segmentation artifacts under `samples/`.
- Generated extraction artifacts under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- `qwen3:1.7b`: smoke-test model according to project knowledge.
- `qwen3.5:9b`: current meaningful extraction and segmentation model according
to project knowledge and generated full-meeting segmentation artifacts.
- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts.
Result:
The repository supports the conclusion that `qwen3.5:9b` is the meaningful
current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger
models and longer contexts may increase runtime substantially, but no hardware
benchmark suite is recorded.
Decision:
Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current
experiments. Do not assume model size alone fixes prompt or pipeline design.
Lessons learned:
Evaluation must separate model capability from prompt clarity, context
strategy, extraction schema and consolidation.
Evidence:
- `PROJECT_KNOWLEDGE.md`
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
- `samples/chunks/chunk_01_normalized_segments.json`
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
## EXP-0007 - Thinking output and Ollama API behavior
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Thinking-capable Qwen models may return reasoning separately from the final
answer, and extraction parsing should not fail merely because extra text or
multiple JSON objects appear.
Setup:
The Ollama response reader checks `response`, chat-style `message.content` and
then `thinking`. The JSON parser tries a full parse first, then scans JSON
object candidates and returns the final valid object. A regression test covers
thinking text before final JSON.
Inputs:
- Synthetic parser test in `tests/test_extraction_protocol.py`.
Model / configuration:
- No LLM run in the test.
- Code path is used by Ollama extraction.
Result:
The parser can handle additional text and multiple JSON objects where the final
valid object is the intended answer. Current code still falls back to `thinking`
only if no usable response or message content is present.
Decision:
Keep parser robustness, but do not treat thinking output as the root cause of
all extraction failures.
Lessons learned:
API response shape and model output shape are separate concerns. Preserve raw
responses when diagnosing failures.
Evidence:
- `src/meeting_lab/extraction/extract_chunks.py`
- `tests/test_extraction_protocol.py`
- `PROJECT_KNOWLEDGE.md`
## EXP-0008 - JSON truncation and generation limits
Status: Accepted
Hypothesis:
Some extraction failures are caused by generation limits truncating JSON rather
than by prompt wording or parser behavior.
Setup:
A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The
generation limit was increased for the successful path. Exact failing limit is
not recorded in the repository; the current extractor default is verifiably
`--num-predict 8192`.
Inputs:
- Local extraction runs referenced by project knowledge.
- Current extraction CLI.
Model / configuration:
- `qwen3.5:9b`.
- Current extractor default: `num_predict=8192`.
Result:
Increasing the generation limit fixed the technical JSON failure. This was not
primarily a parser or prompt problem.
Decision:
When JSON is truncated, inspect raw output and generation limits before editing
prompts.
Lessons learned:
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong
first response if the model simply ran out of output tokens.
Evidence:
- `src/meeting_lab/extraction/extract_chunks.py`
- `PROJECT_KNOWLEDGE.md`
- `AGENTS.md`
## EXP-0009 - Minimal end-to-end pipeline
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
A minimal local pipeline can transform Whisper output into chunk extractions
and an interim protocol, proving the technical path before the final
architecture exists.
Setup:
The repository added cleanup, normalization, chunking, extraction and protocol
builder scripts, with sample generated artifacts.
Inputs:
- `samples/whisper/meeting_speech.json`
- `samples/whisper/meeting_speech_cleaned.json`
- Generated chunks and normalized chunks under
`samples/whisper/meeting_speech_cleaned_chunks/`.
Model / configuration:
- Local Ollama extraction for chunk JSON.
- Windowed segmentation artifacts use `qwen3.5:9b`.
Result:
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
extraction JSON files, windowed segmentation artifacts and
`meeting_protocol.md`.
Decision:
The minimal pipeline is technically validated. The first protocol builder is an
interim validation tool, not the final architecture.
Lessons learned:
End-to-end execution exposed the next limitation: extraction output needs
consolidation and purpose-specific rendering.
Evidence:
- `scripts/clean_whisper_json.py`
- `src/meeting_lab/normalization/normalize_transcript.py`
- `src/meeting_lab/chunking/chunk_transcript.py`
- `src/meeting_lab/extraction/extract_chunks.py`
- `src/meeting_lab/protocol/build_protocol.py`
- `samples/whisper/meeting_speech_cleaned_chunks/`
- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline`
## EXP-0010 - Gold Standard corpus
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Reproducible prompt engineering requires synthetic transcripts with explicit
expected semantic outputs.
Setup:
The Gold Standard corpus defines scenario directories with `transcript.txt`,
`expected.json` and README files describing ground truth and common model
mistakes. The runner validates schema keys and writes `actual.json` for a
scenario.
Inputs:
- Gold scenarios under `tests/gold/`.
Model / configuration:
- Runner requires an explicit Ollama model for LLM evaluation.
- Existing unit tests for the runner do not invoke Ollama.
Result:
The corpus gives stable semantics for decisions, facts, positions, todos,
questions, technical details and difficult mixed cases. Initial structured
transcripts are Phase 1 and easier than raw Whisper-style transcripts.
Decision:
Use Gold Standard tests as both regression tests and formal meeting-semantics
specification. Raw or unlabelled transcript cases remain later-phase work.
Lessons learned:
Without expected outputs, prompt changes cannot be evaluated reproducibly.
Evidence:
- `tests/gold/`
- `scripts/run_gold_test.py`
- `tests/test_gold_runner.py`
- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests`
## EXP-0011 - Gold-test quality and unique ground truth
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
If a prompt produces unexpected behavior, the gold test itself may be ambiguous
and should be reviewed before the prompt is changed.
Setup:
Decision-focused scenarios were clarified during baseline creation. The current
methodology requires checking unique ground truth before changing prompts.
Inputs:
- `decision_simple`
- `decision_deferred`
- `decision_none`
- Gold methodology document.
Model / configuration:
- Prompt Version 2 baseline work.
Result:
`decision_simple` required clarification around the explicit agreement and
nearby non-decision wording. The earlier negative/deferral ambiguity is now
represented by distinct `decision_none` and `decision_deferred` scenarios in
the repository. Punctuation is not reliable evidence for Whisper transcripts;
agreement language and wording must carry the semantics.
Decision:
Ambiguous gold tests must be reviewed before prompt changes. Do not treat
punctuation as reliable evidence in real Whisper-style transcripts.
Lessons learned:
Bad gold tests create false prompt failures and can encourage overfitting.
Evidence:
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
- `tests/gold/decision_simple/README.md`
- `tests/gold/decision_deferred/README.md`
- `tests/gold/decision_none/README.md`
- Commit `f7ad9ba`
## EXP-0012 - Decision taxonomy
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Decision extraction needs a formal taxonomy that distinguishes substantive
decisions from process decisions and non-decisions.
Setup:
The decision definition document and decision prompt define included and
excluded categories. Gold tests cover explicit decisions, true no-decision
cases and deferrals.
Inputs:
- `tests/gold/DECISION_DEFINITION.md`
- `prompts/decisions.md`
- Decision gold scenarios.
Model / configuration:
- Prompt Version 2 baseline.
Result:
Accepted decision categories include substantive decisions, organizational
decisions, process decisions, approvals, rejections, deferrals, explicit
decisions not to decide yet and explicit agreement to gather more information
before deciding. Opinions, preferences and proposals without agreement are not
decisions.
Decision:
"No decision was reached" and "the decision was deferred" are distinct semantic
outcomes.
Lessons learned:
Deferral can be a valid process decision even when the substantive topic remains
unresolved.
Evidence:
- `tests/gold/DECISION_DEFINITION.md`
- `tests/gold/decision_deferred/`
- `tests/gold/decision_none/`
- `prompts/decisions.md`
## EXP-0013 - Prompt engineering methodology
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Prompt iteration needs strict experimental controls to prevent regression,
overfitting and arbitrary prompt churn.
Setup:
The methodology was documented alongside the Gold Standard corpus and later
summarized for agents.
Inputs:
- Gold scenarios.
- Prompt files.
Model / configuration:
- Applies to all prompt experiments.
Result:
The accepted method is one prompt change per iteration, one target test at a
time, immediate validation, no regressions, no `expected.json` edits merely to
force a pass, no test-specific prompt hacks, stopping after two consecutive
non-improving iterations, and verifying unique ground truth before prompt
changes.
Decision:
Prompt changes are controlled experiments. See `AGENTS.md` for agent operating
rules.
Lessons learned:
Most prompt changes are not isolated unless the experiment explicitly constrains
the target and regression set.
Evidence:
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
- `AGENTS.md`
- Commit `f7ad9ba`
## EXP-0014 - Decision Prompt Version 2
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Adding explicit process-decision language to the decision prompt can preserve
true decision detection while recognizing deferrals.
Setup:
Prompt Version 2 added explicit support for deferrals and process decisions.
The baseline was validated on three decision scenarios.
Inputs:
- `decision_simple`
- `decision_deferred`
- `decision_none`
Model / configuration:
- Prompt Version 2.
- Model used for validation is not recorded in the committed methodology.
Result:
The committed methodology records all three baseline scenarios as passing. The
current generated `actual.json` files also show the expected decision count for
these decision scenarios, although some non-decision categories remain less
complete.
Decision:
Prompt Version 2 is the current decision baseline. Explicit deferrals are
recognized as process decisions.
Lessons learned:
Decision-count success does not imply all categories are solved. Category-level
evaluation must continue.
Evidence:
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
- `tests/gold/decision_simple/actual.json`
- `tests/gold/decision_deferred/actual.json`
- `tests/gold/decision_none/actual.json`
- `prompts/decisions.md`
## EXP-0015 - Difficult synthetic meeting
Status: Running
Date or period: 2026-07-30
Hypothesis:
A deliberately adversarial synthetic meeting can expose extraction failures
that simple category tests miss.
Setup:
`evil_meeting` includes interruptions, corrections, absent referenced people,
near-decisions, changed positions and one expected explicit decision.
Inputs:
- `tests/gold/evil_meeting/`.
Model / configuration:
- Existing generated `actual.json`; exact model is not stored in the artifact.
Result:
The generated result found the expected FR-7 exclusion decision. It also
classified "do not migrate until the mapping table is checked" as an additional
decision. The current expected file treats that statement as a position, but
contextual review suggests it may be a valid process instruction or decision.
Decision:
Do not classify this as a simple model failure without reviewing the gold
standard. The scenario exposes a semantic gap in the expected output.
Lessons learned:
Difficult synthetic cases are valuable because they reveal ambiguity in the
specification as well as model mistakes.
Evidence:
- `tests/gold/evil_meeting/README.md`
- `tests/gold/evil_meeting/expected.json`
- `tests/gold/evil_meeting/actual.json`
- See EXP-0011 and EXP-0012.
## EXP-0016 - Context-size extraction comparisons
Status: Rejected
Hypothesis:
Increasing extraction context from one chunk to neighboring chunk groups should
make extraction more complete and therefore should become the baseline.
Setup:
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3
and 1+2+3.
Inputs:
- Normalized meeting chunks.
Model / configuration:
- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for
extraction experiments.
Result:
More context sometimes improved completeness, but it also shifted category
classification, added duplicates and reduced stability. No numeric winner is
recorded in the repository.
Decision:
Do not adopt larger extraction windows as the baseline. Independent chunk
extraction remains current strategy. Recover global context through
consolidation rather than continuously enlarging extraction windows.
Lessons learned:
Context size is not a monotonic quality knob. It changes the task the model is
performing.
Evidence:
- `PROJECT_KNOWLEDGE.md`
- `AGENTS.md`
- `ROADMAP.md`
- See EXP-0002 and EXP-0018.
## EXP-0017 - Independent full-meeting chunk extraction
Status: Accepted
Hypothesis:
Extracting every normalized chunk independently can produce enough structured
material for a useful protocol draft.
Setup:
Nine normalized chunks were extracted into separate JSON files and then
rendered by the interim protocol builder.
Inputs:
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt`
through `chunk_09_normalized.txt`.
Model / configuration:
- Existing extraction artifacts do not record model metadata.
Result:
The nine extraction files contain facts, decisions, todos, questions and
technical details. `meeting_protocol.md` aggregates them into a readable draft.
Duplicates, category shifts and synthesis became the dominant limitations.
Decision:
Independent chunk extraction is useful enough to keep as the baseline, but it
requires a consolidation stage.
Lessons learned:
Per-chunk extraction gives recall-oriented raw material. It does not by itself
produce a polished or canonical meeting representation.
Evidence:
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json`
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
- `src/meeting_lab/protocol/build_protocol.py`
- See EXP-0018.
## EXP-0018 - Human protocol comparison and output-view split
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
One generated protocol cannot satisfy every use case; protocol output should be
separated by purpose and audience.
Setup:
The interim machine protocol was compared against the desired human protocol
shape and then the architecture was revised toward parallel output views.
Inputs:
- Interim `meeting_protocol.md`.
- Architecture and output-view documentation.
Model / configuration:
- Not applicable; this is a design evaluation.
Result:
The human protocol target is denser and organized by purpose and topic rather
than extraction categories. The machine extraction retains more context and is
useful for recall, but it is not the right direct source for a concise
distribution artifact or durable knowledge entry.
Decision:
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution
Protocol / Verteilerprotokoll and Knowledge Objects /
Wissensdatenbankeintrag.
Lessons learned:
Rendering is a separate concern from extraction and consolidation. Output views
must be parallel renderings of shared semantics, not transformations of one
another.
Evidence:
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
- `docs/output-views.md`
- `docs/architecture.md`
- `docs/pipeline.md`
- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture`
## EXP-0019 - Consolidation and Canonical Meeting Knowledge
Status: Accepted
Date or period: 2026-07-30
Hypothesis:
Extraction, consolidation and output rendering are separate problems and should
not be collapsed into one LLM prompt or one protocol file.
Setup:
The architecture was refined after the minimal pipeline and protocol draft
showed duplicate, synthesis and audience-specific rendering limitations.
Inputs:
- Extraction JSON artifacts.
- Interim protocol draft.
- Architecture and data-model documentation.
Model / configuration:
- Not applicable; this is an architectural conclusion.
Result:
The accepted design is a planned Canonical Meeting Knowledge layer as the
semantic source of truth, with Working Protocol, Distribution Protocol and
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
preserve evidence, reconcile category shifts and mark contradictions or
uncertainty.
Decision:
Consolidation is the next major engineering step after stable local extraction.
Canonical Meeting Knowledge and final output views are planned, not implemented.
Lessons learned:
Global meeting understanding should be recovered by consolidation over
evidence-bearing extractions, not by silently changing output views or expanding
LLM context indefinitely.
Evidence:
- `docs/architecture.md`
- `docs/pipeline.md`
- `docs/data-models.md`
- `docs/output-views.md`
- `PROJECT_KNOWLEDGE.md`
- `ROADMAP.md`
- Commit `5c03ed7`