Document project architecture and development methodology
- add AGENTS.md with development and prompt-engineering rules - add PROJECT_KNOWLEDGE.md summarizing current architecture and findings - add CHANGELOG.md - add ROADMAP.md - establish experiments.md as the project's experiment log - document Canonical Meeting Knowledge architecture - document Output Views and Knowledge Objects - capture accepted experimental results and engineering methodology
This commit is contained in:
@@ -0,0 +1,124 @@
|
|||||||
|
# AGENTS.md
|
||||||
|
|
||||||
|
Practical instructions for coding agents working in Meeting Lab.
|
||||||
|
|
||||||
|
## Project Purpose
|
||||||
|
|
||||||
|
Meeting Lab extracts and structures organizational knowledge from meeting
|
||||||
|
recordings. It is an experimental local discussion analyzer, not merely a
|
||||||
|
one-step meeting-protocol generator.
|
||||||
|
|
||||||
|
Successful approaches may later move into the Meeting Assistant project.
|
||||||
|
|
||||||
|
## Current Pipeline
|
||||||
|
|
||||||
|
Current and intended flow:
|
||||||
|
|
||||||
|
```text
|
||||||
|
Audio
|
||||||
|
-> Whisper
|
||||||
|
-> cleanup
|
||||||
|
-> normalization
|
||||||
|
-> chunking
|
||||||
|
-> local chunk extraction
|
||||||
|
-> consolidation
|
||||||
|
-> Canonical Meeting Knowledge
|
||||||
|
-> Output Views
|
||||||
|
```
|
||||||
|
|
||||||
|
Status:
|
||||||
|
|
||||||
|
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
||||||
|
local chunk extraction, interim Markdown protocol builder.
|
||||||
|
- Experimental/prototype: topic segmentation and review tooling.
|
||||||
|
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
|
||||||
|
Output Views.
|
||||||
|
|
||||||
|
## Architectural Principles
|
||||||
|
|
||||||
|
- Canonical Meeting Knowledge is the intended semantic source of truth.
|
||||||
|
- Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||||
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are
|
||||||
|
parallel output views.
|
||||||
|
- Output views must not silently change meaning. They may select, condense or
|
||||||
|
render information for an audience, but not invent new semantics.
|
||||||
|
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
||||||
|
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
||||||
|
- Current extraction strategy is one normalized chunk per LLM call.
|
||||||
|
- Do not expand context windows or redesign the extraction strategy without an
|
||||||
|
explicit experiment.
|
||||||
|
- Deterministic stages should remain deterministic where possible.
|
||||||
|
|
||||||
|
## Prompt Engineering Rules
|
||||||
|
|
||||||
|
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
|
||||||
|
`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`:
|
||||||
|
|
||||||
|
1. Make only one prompt change per iteration.
|
||||||
|
2. Optimize only one target gold test case at a time.
|
||||||
|
3. Validate every prompt modification immediately.
|
||||||
|
4. Accept a prompt change only if it improves the target and causes no
|
||||||
|
regressions in previously passing gold tests.
|
||||||
|
5. Never modify `expected.json` merely to make a prompt pass.
|
||||||
|
6. Prompt engineering edits prompt files only; Python code changes require a
|
||||||
|
separate explicit task.
|
||||||
|
7. Maintain a prompt evolution log for every iteration.
|
||||||
|
8. Stop arbitrary iterations if small changes do not improve the test; analyze
|
||||||
|
the root cause.
|
||||||
|
9. Avoid gold-test overfitting. Prompt changes must generalize and must not
|
||||||
|
special-case one transcript.
|
||||||
|
10. Stop after two consecutive non-improving prompt iterations and classify the
|
||||||
|
root cause.
|
||||||
|
11. Verify whether the target gold test has objectively unique ground truth
|
||||||
|
before changing a prompt for unexpected behavior.
|
||||||
|
|
||||||
|
Current documented Prompt Version 2 decision baseline:
|
||||||
|
|
||||||
|
- `decision_simple`: passing
|
||||||
|
- `decision_deferred`: passing
|
||||||
|
- `decision_none`: passing
|
||||||
|
|
||||||
|
## LLM Execution Safety
|
||||||
|
|
||||||
|
- Never start a full multi-chunk LLM run unless explicitly requested.
|
||||||
|
- Before any LLM run, state the model, inputs, expected LLM-call count and
|
||||||
|
output location.
|
||||||
|
- Do not retry LLM calls automatically unless explicitly allowed.
|
||||||
|
- Do not download models automatically.
|
||||||
|
- Prefer small-scope validation runs.
|
||||||
|
- Never use generated output as committed source data.
|
||||||
|
- Preserve raw model responses when diagnosing parser or truncation failures.
|
||||||
|
- Do not run Ollama from unit tests.
|
||||||
|
|
||||||
|
## Development Rules
|
||||||
|
|
||||||
|
- Make small, focused changes.
|
||||||
|
- Preserve the existing architecture unless a redesign is explicitly requested.
|
||||||
|
- Add regression tests for bugs.
|
||||||
|
- Run non-LLM tests before committing when code changes are made.
|
||||||
|
- Do not commit generated transcripts, audio, extraction JSON, protocol output
|
||||||
|
or temporary files.
|
||||||
|
- Report files changed, tests run and assumptions.
|
||||||
|
- Do not commit or push unless explicitly requested.
|
||||||
|
|
||||||
|
## Repository Conventions
|
||||||
|
|
||||||
|
- `README.md`: project overview and current high-level status.
|
||||||
|
- `docs/`: architecture, pipeline, data model and output-view documentation.
|
||||||
|
- `prompts/`: extraction and segmentation prompts. Treat prompt edits as
|
||||||
|
controlled experiments.
|
||||||
|
- `tests/gold/`: Gold Standard corpus and semantic specification for extraction
|
||||||
|
behavior.
|
||||||
|
- `scripts/`: command-line support scripts such as Whisper cleanup and gold
|
||||||
|
test execution.
|
||||||
|
- `src/meeting_lab/normalization/`: deterministic transcript cleanup.
|
||||||
|
- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics.
|
||||||
|
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
||||||
|
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
||||||
|
extractor modules.
|
||||||
|
- `src/meeting_lab/consolidation/`: planned consolidation area.
|
||||||
|
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
||||||
|
- `src/meeting_lab/models/`: current lightweight data models.
|
||||||
|
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
||||||
|
sample output as canonical source data.
|
||||||
|
|
||||||
@@ -0,0 +1,47 @@
|
|||||||
|
# Changelog
|
||||||
|
|
||||||
|
## Unreleased
|
||||||
|
|
||||||
|
### Added
|
||||||
|
|
||||||
|
- Initial Meeting Lab project structure with source, docs, prompts, samples and
|
||||||
|
tests directories.
|
||||||
|
- Deterministic transcript normalization module.
|
||||||
|
- Technical transcript chunking with block-aligned chunk generation.
|
||||||
|
- Whisper JSON cleanup support.
|
||||||
|
- Local Ollama-based chunk extraction flow.
|
||||||
|
- Interim Markdown meeting protocol builder for technical validation.
|
||||||
|
- Initial and windowed topic segmentation prototypes plus review tooling.
|
||||||
|
- Gold Standard extraction corpus and gold-test runner.
|
||||||
|
- Prompt loading support and common/decision prompt baseline.
|
||||||
|
- Output-view architecture documentation for Canonical Meeting Knowledge,
|
||||||
|
Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||||
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
|
||||||
|
|
||||||
|
### Changed
|
||||||
|
|
||||||
|
- Refined the architecture from a single meeting protocol toward Canonical
|
||||||
|
Meeting Knowledge as the planned semantic source of truth.
|
||||||
|
- Clarified that Working Protocol, Distribution Protocol and Knowledge Objects
|
||||||
|
are parallel renderings, not derived from one another.
|
||||||
|
- Updated decision extraction semantics to include explicit process decisions
|
||||||
|
and deferrals.
|
||||||
|
- Simplified extraction prompt assembly around prompt files.
|
||||||
|
|
||||||
|
### Fixed
|
||||||
|
|
||||||
|
- Fixed Whisper JSON chunk extraction to prefer `segments[*].text` over the
|
||||||
|
aggregate top-level `text` field.
|
||||||
|
- Added chunking tests to ensure chunks do not duplicate later blocks when no
|
||||||
|
overlap is requested.
|
||||||
|
- Added parser handling for model responses that contain thinking text before
|
||||||
|
the final JSON object.
|
||||||
|
|
||||||
|
### Documentation
|
||||||
|
|
||||||
|
- Added architecture, pipeline and data-model documentation.
|
||||||
|
- Added output-view documentation.
|
||||||
|
- Added Gold Standard prompt-engineering methodology.
|
||||||
|
- Added formal decision-definition documentation.
|
||||||
|
- Added scenario README files for the Gold Standard corpus.
|
||||||
|
|
||||||
@@ -0,0 +1,151 @@
|
|||||||
|
# Project Knowledge
|
||||||
|
|
||||||
|
This is a compact operational summary of the current Meeting Lab state.
|
||||||
|
|
||||||
|
## Objective
|
||||||
|
|
||||||
|
Meeting Lab develops and evaluates local methods for extracting structured
|
||||||
|
organizational knowledge from real meeting recordings and transcripts. The
|
||||||
|
project is a research and validation environment for a future Meeting
|
||||||
|
Assistant, not a finished product.
|
||||||
|
|
||||||
|
## Implemented Pipeline Stages
|
||||||
|
|
||||||
|
Implemented:
|
||||||
|
|
||||||
|
- Whisper JSON cleanup via `scripts/clean_whisper_json.py`.
|
||||||
|
- Transcript normalization in `src/meeting_lab/normalization/`.
|
||||||
|
- Technical chunking in `src/meeting_lab/chunking/`.
|
||||||
|
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
|
||||||
|
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||||
|
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||||
|
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||||
|
gold-test runner validation.
|
||||||
|
|
||||||
|
Experimental/prototype:
|
||||||
|
|
||||||
|
- Topic segmentation in `src/meeting_lab/segmentation/`.
|
||||||
|
- Windowed segmentation and review output in `samples/chunks/`.
|
||||||
|
- Gold Standard extraction corpus under `tests/gold/`.
|
||||||
|
|
||||||
|
Planned:
|
||||||
|
|
||||||
|
- Consolidation of extraction results.
|
||||||
|
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
||||||
|
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||||
|
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
||||||
|
|
||||||
|
## Source Tree
|
||||||
|
|
||||||
|
```text
|
||||||
|
src/meeting_lab/
|
||||||
|
chunking/ technical transcript chunking
|
||||||
|
consolidation/ planned merge/consolidation area
|
||||||
|
extraction/ current local LLM extraction flow
|
||||||
|
io/ lightweight file and JSON helpers
|
||||||
|
llm/ Ollama and prompt support
|
||||||
|
models/ current lightweight model definitions
|
||||||
|
normalization/ deterministic transcript cleanup
|
||||||
|
protocol/ interim Markdown protocol builder
|
||||||
|
segmentation/ experimental topic segmentation tooling
|
||||||
|
```
|
||||||
|
|
||||||
|
Supporting areas:
|
||||||
|
|
||||||
|
- `docs/`: architecture, pipeline, data models and output-view concepts.
|
||||||
|
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain
|
||||||
|
substantive extraction prompt text in the current tree.
|
||||||
|
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
|
||||||
|
- `samples/`: sample inputs and generated or experimental artifacts.
|
||||||
|
- `scripts/`: operational scripts for cleanup and gold-test execution.
|
||||||
|
|
||||||
|
## Current Model Strategy
|
||||||
|
|
||||||
|
The current extraction strategy is one normalized chunk per LLM call. This is
|
||||||
|
preferred over expanding context windows or asking one model call to analyze a
|
||||||
|
full meeting.
|
||||||
|
|
||||||
|
Known working models from current project notes and experiment practice:
|
||||||
|
|
||||||
|
- `qwen3:1.7b`: useful for smoke tests.
|
||||||
|
- `qwen3.5:9b`: useful for meaningful extraction and segmentation work.
|
||||||
|
|
||||||
|
LLM calls use Ollama locally. The current extractor defaults to `qwen3:8b`, but
|
||||||
|
validated work may specify another model explicitly.
|
||||||
|
|
||||||
|
## Important Findings
|
||||||
|
|
||||||
|
- Whisper JSON chunking must use `segments[*].text`, not only the top-level
|
||||||
|
`text` field.
|
||||||
|
- Independent chunk extraction is currently preferred.
|
||||||
|
- Larger context windows can change classification behavior and increase
|
||||||
|
instability.
|
||||||
|
- Extraction and consolidation are separate problems.
|
||||||
|
- Generation limits can truncate JSON.
|
||||||
|
- Qwen thinking may be returned separately by the Ollama API.
|
||||||
|
- Gold Standard tests are also a formal specification of meeting semantics.
|
||||||
|
- Raw model responses should be preserved when diagnosing parser or truncation
|
||||||
|
failures.
|
||||||
|
|
||||||
|
## Decision Taxonomy
|
||||||
|
|
||||||
|
Accepted decision semantics:
|
||||||
|
|
||||||
|
- A decision is an explicit agreement that creates a binding change in action,
|
||||||
|
process, responsibility, approval status, timing or next step.
|
||||||
|
- Included: substantive decisions, organizational decisions, process decisions,
|
||||||
|
approvals, rejections, deferrals, explicit agreement not to decide yet, and
|
||||||
|
explicit agreement to gather more information before deciding.
|
||||||
|
- Excluded: opinions, preferences, proposals without agreement, open questions,
|
||||||
|
current-state descriptions and explanations without commitment.
|
||||||
|
- A process decision to defer a substantive decision is still a decision.
|
||||||
|
- "No decision was reached" is different from "the group decided to defer the
|
||||||
|
decision."
|
||||||
|
|
||||||
|
Current Prompt Version 2 decision baseline:
|
||||||
|
|
||||||
|
- `decision_simple`: passing.
|
||||||
|
- `decision_deferred`: passing.
|
||||||
|
- `decision_none`: passing.
|
||||||
|
- Prompt Version 2 explicitly supports process decisions where the group agrees
|
||||||
|
to defer a substantive decision until more information is available.
|
||||||
|
|
||||||
|
## Canonical Knowledge Architecture
|
||||||
|
|
||||||
|
Canonical Meeting Knowledge is the planned semantic intermediate model and
|
||||||
|
future single source of truth. It should preserve topics, facts, decisions,
|
||||||
|
action items, open questions, positions, technical details, rationale,
|
||||||
|
uncertainty, contradictions and source evidence.
|
||||||
|
|
||||||
|
Output views are planned as independent renderings from that canonical model:
|
||||||
|
|
||||||
|
- Working Protocol / Arbeitsprotokoll: relatively complete, optimized for
|
||||||
|
recall and traceability.
|
||||||
|
- Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented,
|
||||||
|
optimized for circulation.
|
||||||
|
- Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge
|
||||||
|
optimized for reuse.
|
||||||
|
|
||||||
|
The current `meeting_protocol.md` builder is an interim technical validation
|
||||||
|
tool, not the final output-view architecture.
|
||||||
|
|
||||||
|
## Current Limitations
|
||||||
|
|
||||||
|
- Discussion Blocks are documented as a stable semantic unit but are not yet a
|
||||||
|
separate implemented pipeline artifact.
|
||||||
|
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
|
||||||
|
- Extraction is still a combined current flow, even though separate extractors
|
||||||
|
are the intended architecture.
|
||||||
|
- Consolidation is not implemented.
|
||||||
|
- Canonical Meeting Knowledge is documented but not implemented.
|
||||||
|
- Final output views are documented but not implemented.
|
||||||
|
- Most prompt files are placeholders except the common and decision prompts.
|
||||||
|
- Gold tests currently emphasize extraction semantics, especially decisions.
|
||||||
|
|
||||||
|
## Next Recommended Engineering Step
|
||||||
|
|
||||||
|
Stabilize repeatable local extraction evaluation before broadening the pipeline:
|
||||||
|
expand Gold Standard coverage by category, keep one-chunk extraction as the
|
||||||
|
baseline, and use small prompt experiments with immediate non-regression checks.
|
||||||
|
After extraction behavior is stable enough, implement consolidation with
|
||||||
|
evidence retention as the next major pipeline stage.
|
||||||
+185
@@ -0,0 +1,185 @@
|
|||||||
|
# Roadmap
|
||||||
|
|
||||||
|
No dates are assigned. Phases describe dependency order, not release promises.
|
||||||
|
|
||||||
|
## Phase 1 - Stable Local Extraction
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Establish reliable per-chunk extraction behavior for core meeting semantics.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Stronger Gold Standard coverage across facts, positions, decisions, todos,
|
||||||
|
questions and technical details.
|
||||||
|
- Improved category prompts.
|
||||||
|
- Repeatable evaluation workflow.
|
||||||
|
- Documented prompt experiment log.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Existing chunk extraction flow.
|
||||||
|
- Existing Gold Standard runner and methodology.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Full-transcript LLM extraction.
|
||||||
|
- Larger context-window strategy changes without an explicit experiment.
|
||||||
|
- Consolidation or final protocol rendering.
|
||||||
|
|
||||||
|
## Phase 2 - Consolidation
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Merge independent extraction results into a coherent meeting-level
|
||||||
|
representation without losing evidence.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Duplicate merging.
|
||||||
|
- Evidence retention.
|
||||||
|
- Category-shift reconciliation, especially facts versus positions and
|
||||||
|
positions versus decisions.
|
||||||
|
- Contradiction and uncertainty markers.
|
||||||
|
- Consolidated meeting representation.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Stable local extraction baseline.
|
||||||
|
- Gold tests that expose cross-chunk duplication and category shifts.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Final Canonical Meeting Knowledge schema.
|
||||||
|
- User-facing protocol polish.
|
||||||
|
- Retrieval or RAG integration.
|
||||||
|
|
||||||
|
## Phase 3 - Canonical Meeting Knowledge
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Define and implement the semantic intermediate model that becomes the source
|
||||||
|
of truth for downstream outputs.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Canonical Meeting Knowledge schema.
|
||||||
|
- Source evidence and traceability fields.
|
||||||
|
- Clear distinction between durable knowledge and meeting-specific actions.
|
||||||
|
- Migration path from consolidated extraction JSON into the canonical model.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Consolidation behavior that preserves evidence and uncertainty.
|
||||||
|
- Agreement on required semantic categories.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- GUI.
|
||||||
|
- Export formats beyond those needed to validate the model.
|
||||||
|
- Knowledge-system storage design.
|
||||||
|
|
||||||
|
## Phase 4 - Output Views
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Render purpose-specific outputs from Canonical Meeting Knowledge without
|
||||||
|
changing meaning.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Working Protocol / Arbeitsprotokoll renderer.
|
||||||
|
- Distribution Protocol / Verteilerprotokoll renderer.
|
||||||
|
- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export.
|
||||||
|
- Tests or checks showing that output views are parallel renderings of the same
|
||||||
|
canonical model.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Implemented Canonical Meeting Knowledge.
|
||||||
|
- Clear audience and completeness rules for each output view.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Additional analysis during rendering.
|
||||||
|
- Deriving one output view from another.
|
||||||
|
- Retrieval integration.
|
||||||
|
|
||||||
|
## Phase 5 - Review and Quality Control
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Add optional review stages that improve omission detection, consistency and
|
||||||
|
model selection.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Optional whole-transcript review.
|
||||||
|
- Omission detection.
|
||||||
|
- Consistency checks.
|
||||||
|
- Model comparison workflow.
|
||||||
|
- Hardware and runtime benchmarks.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Stable extraction, consolidation and canonical model.
|
||||||
|
- Representative test meetings.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Automatic acceptance of review suggestions without evidence.
|
||||||
|
- Product UI work.
|
||||||
|
- Cloud deployment.
|
||||||
|
|
||||||
|
## Phase 6 - Productization
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Turn the validated pipeline into a usable local workflow.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Recording/transcription workflow.
|
||||||
|
- FFmpeg integration.
|
||||||
|
- Meeting metadata capture.
|
||||||
|
- Participant entry.
|
||||||
|
- GUI.
|
||||||
|
- Stable deployment process.
|
||||||
|
- Export workflows.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Stable pipeline stages and output views.
|
||||||
|
- Clear operational requirements for local use.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Enterprise knowledge retrieval.
|
||||||
|
- Future Meeting Assistant integration beyond export contracts.
|
||||||
|
- Cloud-first architecture.
|
||||||
|
|
||||||
|
## Phase 7 - Knowledge-System Integration
|
||||||
|
|
||||||
|
Goal:
|
||||||
|
|
||||||
|
- Reuse durable meeting knowledge in broader knowledge systems.
|
||||||
|
|
||||||
|
Deliverables:
|
||||||
|
|
||||||
|
- Structured Knowledge Objects.
|
||||||
|
- Retrieval-ready storage format.
|
||||||
|
- Future RAG integration path.
|
||||||
|
- Reuse contracts for Meeting Assistant and other knowledge systems.
|
||||||
|
|
||||||
|
Prerequisites:
|
||||||
|
|
||||||
|
- Canonical Meeting Knowledge and Knowledge Objects are implemented and stable.
|
||||||
|
- Durable knowledge is separated from meeting-specific actions and discussion
|
||||||
|
history.
|
||||||
|
|
||||||
|
Out of scope:
|
||||||
|
|
||||||
|
- Building a full enterprise search product inside Meeting Lab.
|
||||||
|
- Treating raw transcripts or generated protocols as the knowledge source of
|
||||||
|
truth.
|
||||||
|
|
||||||
@@ -0,0 +1,971 @@
|
|||||||
|
# Experiments
|
||||||
|
|
||||||
|
This file records durable technical experiments and findings for Meeting Lab.
|
||||||
|
It is not a diary and does not replace commit history.
|
||||||
|
|
||||||
|
## Status values
|
||||||
|
|
||||||
|
- Proposed: experiment idea exists, but no result is recorded.
|
||||||
|
- Running: experiment is in progress and no decision has been made.
|
||||||
|
- Accepted: finding is the current baseline or design conclusion.
|
||||||
|
- Rejected: hypothesis was tested and should not be repeated as-is.
|
||||||
|
- Superseded: finding was useful but has been replaced by a newer baseline.
|
||||||
|
|
||||||
|
## EXP-0001 - Whisper JSON interpretation
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-29
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Whisper JSON should be chunked from its segment stream, not from the aggregate
|
||||||
|
top-level text field.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
`chunk_transcript.py` was updated to parse JSON input and prefer
|
||||||
|
`segments[*].text` when `segments` exists. A regression test supplies JSON with
|
||||||
|
both top-level `text` and separate segment texts.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`.
|
||||||
|
- Real Whisper artifacts under `samples/whisper/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- No LLM.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The test verifies that the block stream is `["alpha", "beta", "gamma"]` and
|
||||||
|
does not include the aggregate `"alpha beta gamma"` text. The repository history
|
||||||
|
records this as the fix for the earlier failure where the first chunk contained
|
||||||
|
the complete transcript.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
When `segments` exists, `segments[*].text` is the authoritative transcript
|
||||||
|
stream. The top-level `text` field is only a fallback.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Whisper JSON is structured input. Treating it like plain text can duplicate the
|
||||||
|
entire transcript and invalidate downstream chunking.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||||
|
- `tests/test_chunking.py`
|
||||||
|
- Commit `4656523` - `Fix Whisper JSON chunk extraction`
|
||||||
|
|
||||||
|
## EXP-0002 - Technical transcript chunking baseline
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-29 to 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Sequential technical chunks around the configured target size can preserve the
|
||||||
|
transcript while keeping extraction calls small enough for local models.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The chunker splits block-aligned text with configurable target, minimum,
|
||||||
|
maximum and overlap settings. Tests verify no duplicate later blocks when
|
||||||
|
overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Synthetic block list in `tests/test_chunking.py`.
|
||||||
|
- `samples/whisper/meeting_speech_cleaned.json`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- No LLM for chunking.
|
||||||
|
- Manifest uses default chunking behavior recorded in
|
||||||
|
`samples/whisper/meeting_speech_cleaned_chunks/manifest.json`.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
With overlap set to zero, tests verify that all blocks appear exactly once. The
|
||||||
|
real sample manifest contains nine chunks, mostly near the configured target
|
||||||
|
size, with a smaller final chunk.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Independent sequential chunks are the current technical baseline. One
|
||||||
|
normalized chunk per extraction call is the preferred extraction strategy.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Chunking solves model-size constraints only. It must not perform topic
|
||||||
|
detection or semantic merging.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||||
|
- `tests/test_chunking.py`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`
|
||||||
|
- `AGENTS.md`
|
||||||
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
|
|
||||||
|
## EXP-0003 - Conservative transcript normalization
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Transcript cleanup should improve readability without changing meeting
|
||||||
|
semantics.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The normalizer removes isolated filler sounds, immediate duplicate words or
|
||||||
|
short duplicate phrases, and redundant whitespace. It records changed blocks in
|
||||||
|
a JSON change log and explicitly preserves semantic content categories.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||||
|
- Change logs such as `chunk_01_changes.json`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- No LLM.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The implementation and generated change logs show a conservative policy:
|
||||||
|
negations, qualifiers, dates, numbers, responsibilities, technical statements,
|
||||||
|
deadlines, decisions and commitments are preserved.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Normalization remains deterministic and low-risk. When uncertain, leave text
|
||||||
|
unchanged.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove
|
||||||
|
semantic cues needed by extraction.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/normalization/normalize_transcript.py`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json`
|
||||||
|
- `docs/pipeline.md`
|
||||||
|
|
||||||
|
## EXP-0004 - Full-context topic segmentation
|
||||||
|
|
||||||
|
Status: Superseded
|
||||||
|
|
||||||
|
Date or period: 2026-07-21 to 2026-07-22
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
A single full-context topic segmentation call can identify topic boundaries in
|
||||||
|
a normalized transcript chunk.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The initial segmentation prototype asked the model for topic changes and then
|
||||||
|
converted those boundaries into continuous, non-overlapping segments.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `samples/chunks/chunk_01_normalized.txt`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Generated artifact records `qwen3:14b`.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The generated artifact contains 85 blocks, three topic-change boundaries and
|
||||||
|
four segments. The run metadata records a substantially longer elapsed time
|
||||||
|
than the later windowed artifact for the same input.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Full-context segmentation was useful as a prototype, but it was superseded by
|
||||||
|
windowed segmentation and manual review tooling.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
The prototype established the boundary-to-segment representation, but did not
|
||||||
|
settle segmentation quality.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/segmentation/segment_topics.py`
|
||||||
|
- `samples/chunks/chunk_01_normalized_segments.json`
|
||||||
|
- Commit `f234efc` - `Add initial topic segmentation prototype`
|
||||||
|
- Commit `889a4fe` - `Detect topic boundaries as continuous segments`
|
||||||
|
|
||||||
|
## EXP-0005 - Windowed topic segmentation and review
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-22
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Windowed topic segmentation can reduce runtime and make boundary evaluation
|
||||||
|
more inspectable than a single full-context call.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
`segment_topics_windowed.py` analyzes overlapping windows and reports only
|
||||||
|
boundaries from the decision range. Python merges boundaries into continuous,
|
||||||
|
non-overlapping segments. `review_segmentation.py` renders each boundary with
|
||||||
|
neighboring transcript context for manual classification.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `samples/chunks/chunk_01_normalized.txt`.
|
||||||
|
- Full meeting normalized chunks under
|
||||||
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3.
|
||||||
|
- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments
|
||||||
|
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
|
||||||
|
Manual review output shows that some boundaries were assessed as subtopics
|
||||||
|
rather than full topic changes. Full-meeting artifacts show one window per
|
||||||
|
already-small normalized chunk and two segments per chunk.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Windowed segmentation and review tooling are accepted as prototype tooling, not
|
||||||
|
as a stable production segmentation stage.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Windowing improves inspectability and can reduce runtime, but it can also
|
||||||
|
cluster boundaries and over-segment. Manual review remains necessary.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/segmentation/segment_topics_windowed.py`
|
||||||
|
- `src/meeting_lab/segmentation/review_segmentation.py`
|
||||||
|
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
||||||
|
- `samples/chunks/chunk_01_normalized_windowed_segments_review.md`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
||||||
|
- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling`
|
||||||
|
|
||||||
|
## EXP-0006 - Qwen model comparison
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Larger local Qwen-family models should improve meaningful extraction and
|
||||||
|
segmentation, but model size alone will not solve prompt or pipeline problems.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Project work used smaller models for smoke checks and larger local models for
|
||||||
|
meaningful extraction or segmentation. Artifacts and project knowledge record
|
||||||
|
the currently useful model roles.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Gold Standard scenarios under `tests/gold/`.
|
||||||
|
- Generated segmentation artifacts under `samples/`.
|
||||||
|
- Generated extraction artifacts under
|
||||||
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- `qwen3:1.7b`: smoke-test model according to project knowledge.
|
||||||
|
- `qwen3.5:9b`: current meaningful extraction and segmentation model according
|
||||||
|
to project knowledge and generated full-meeting segmentation artifacts.
|
||||||
|
- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The repository supports the conclusion that `qwen3.5:9b` is the meaningful
|
||||||
|
current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger
|
||||||
|
models and longer contexts may increase runtime substantially, but no hardware
|
||||||
|
benchmark suite is recorded.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current
|
||||||
|
experiments. Do not assume model size alone fixes prompt or pipeline design.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Evaluation must separate model capability from prompt clarity, context
|
||||||
|
strategy, extraction schema and consolidation.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
||||||
|
- `samples/chunks/chunk_01_normalized_segments.json`
|
||||||
|
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
||||||
|
|
||||||
|
## EXP-0007 - Thinking output and Ollama API behavior
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Thinking-capable Qwen models may return reasoning separately from the final
|
||||||
|
answer, and extraction parsing should not fail merely because extra text or
|
||||||
|
multiple JSON objects appear.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The Ollama response reader checks `response`, chat-style `message.content` and
|
||||||
|
then `thinking`. The JSON parser tries a full parse first, then scans JSON
|
||||||
|
object candidates and returns the final valid object. A regression test covers
|
||||||
|
thinking text before final JSON.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Synthetic parser test in `tests/test_extraction_protocol.py`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- No LLM run in the test.
|
||||||
|
- Code path is used by Ollama extraction.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The parser can handle additional text and multiple JSON objects where the final
|
||||||
|
valid object is the intended answer. Current code still falls back to `thinking`
|
||||||
|
only if no usable response or message content is present.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Keep parser robustness, but do not treat thinking output as the root cause of
|
||||||
|
all extraction failures.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
API response shape and model output shape are separate concerns. Preserve raw
|
||||||
|
responses when diagnosing failures.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
||||||
|
- `tests/test_extraction_protocol.py`
|
||||||
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
|
|
||||||
|
## EXP-0008 - JSON truncation and generation limits
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Some extraction failures are caused by generation limits truncating JSON rather
|
||||||
|
than by prompt wording or parser behavior.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The
|
||||||
|
generation limit was increased for the successful path. Exact failing limit is
|
||||||
|
not recorded in the repository; the current extractor default is verifiably
|
||||||
|
`--num-predict 8192`.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Local extraction runs referenced by project knowledge.
|
||||||
|
- Current extraction CLI.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- `qwen3.5:9b`.
|
||||||
|
- Current extractor default: `num_predict=8192`.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
Increasing the generation limit fixed the technical JSON failure. This was not
|
||||||
|
primarily a parser or prompt problem.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
When JSON is truncated, inspect raw output and generation limits before editing
|
||||||
|
prompts.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong
|
||||||
|
first response if the model simply ran out of output tokens.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
||||||
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
|
- `AGENTS.md`
|
||||||
|
|
||||||
|
## EXP-0009 - Minimal end-to-end pipeline
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
A minimal local pipeline can transform Whisper output into chunk extractions
|
||||||
|
and an interim protocol, proving the technical path before the final
|
||||||
|
architecture exists.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The repository added cleanup, normalization, chunking, extraction and protocol
|
||||||
|
builder scripts, with sample generated artifacts.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `samples/whisper/meeting_speech.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned.json`
|
||||||
|
- Generated chunks and normalized chunks under
|
||||||
|
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Local Ollama extraction for chunk JSON.
|
||||||
|
- Windowed segmentation artifacts use `qwen3.5:9b`.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
|
||||||
|
extraction JSON files, windowed segmentation artifacts and
|
||||||
|
`meeting_protocol.md`.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
The minimal pipeline is technically validated. The first protocol builder is an
|
||||||
|
interim validation tool, not the final architecture.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
End-to-end execution exposed the next limitation: extraction output needs
|
||||||
|
consolidation and purpose-specific rendering.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `scripts/clean_whisper_json.py`
|
||||||
|
- `src/meeting_lab/normalization/normalize_transcript.py`
|
||||||
|
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||||
|
- `src/meeting_lab/extraction/extract_chunks.py`
|
||||||
|
- `src/meeting_lab/protocol/build_protocol.py`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/`
|
||||||
|
- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline`
|
||||||
|
|
||||||
|
## EXP-0010 - Gold Standard corpus
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Reproducible prompt engineering requires synthetic transcripts with explicit
|
||||||
|
expected semantic outputs.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The Gold Standard corpus defines scenario directories with `transcript.txt`,
|
||||||
|
`expected.json` and README files describing ground truth and common model
|
||||||
|
mistakes. The runner validates schema keys and writes `actual.json` for a
|
||||||
|
scenario.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Gold scenarios under `tests/gold/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Runner requires an explicit Ollama model for LLM evaluation.
|
||||||
|
- Existing unit tests for the runner do not invoke Ollama.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The corpus gives stable semantics for decisions, facts, positions, todos,
|
||||||
|
questions, technical details and difficult mixed cases. Initial structured
|
||||||
|
transcripts are Phase 1 and easier than raw Whisper-style transcripts.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Use Gold Standard tests as both regression tests and formal meeting-semantics
|
||||||
|
specification. Raw or unlabelled transcript cases remain later-phase work.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Without expected outputs, prompt changes cannot be evaluated reproducibly.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `tests/gold/`
|
||||||
|
- `scripts/run_gold_test.py`
|
||||||
|
- `tests/test_gold_runner.py`
|
||||||
|
- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests`
|
||||||
|
|
||||||
|
## EXP-0011 - Gold-test quality and unique ground truth
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
If a prompt produces unexpected behavior, the gold test itself may be ambiguous
|
||||||
|
and should be reviewed before the prompt is changed.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Decision-focused scenarios were clarified during baseline creation. The current
|
||||||
|
methodology requires checking unique ground truth before changing prompts.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `decision_simple`
|
||||||
|
- `decision_deferred`
|
||||||
|
- `decision_none`
|
||||||
|
- Gold methodology document.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Prompt Version 2 baseline work.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
`decision_simple` required clarification around the explicit agreement and
|
||||||
|
nearby non-decision wording. The earlier negative/deferral ambiguity is now
|
||||||
|
represented by distinct `decision_none` and `decision_deferred` scenarios in
|
||||||
|
the repository. Punctuation is not reliable evidence for Whisper transcripts;
|
||||||
|
agreement language and wording must carry the semantics.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Ambiguous gold tests must be reviewed before prompt changes. Do not treat
|
||||||
|
punctuation as reliable evidence in real Whisper-style transcripts.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Bad gold tests create false prompt failures and can encourage overfitting.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
||||||
|
- `tests/gold/decision_simple/README.md`
|
||||||
|
- `tests/gold/decision_deferred/README.md`
|
||||||
|
- `tests/gold/decision_none/README.md`
|
||||||
|
- Commit `f7ad9ba`
|
||||||
|
|
||||||
|
## EXP-0012 - Decision taxonomy
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Decision extraction needs a formal taxonomy that distinguishes substantive
|
||||||
|
decisions from process decisions and non-decisions.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The decision definition document and decision prompt define included and
|
||||||
|
excluded categories. Gold tests cover explicit decisions, true no-decision
|
||||||
|
cases and deferrals.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `tests/gold/DECISION_DEFINITION.md`
|
||||||
|
- `prompts/decisions.md`
|
||||||
|
- Decision gold scenarios.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Prompt Version 2 baseline.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
Accepted decision categories include substantive decisions, organizational
|
||||||
|
decisions, process decisions, approvals, rejections, deferrals, explicit
|
||||||
|
decisions not to decide yet and explicit agreement to gather more information
|
||||||
|
before deciding. Opinions, preferences and proposals without agreement are not
|
||||||
|
decisions.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
"No decision was reached" and "the decision was deferred" are distinct semantic
|
||||||
|
outcomes.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Deferral can be a valid process decision even when the substantive topic remains
|
||||||
|
unresolved.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `tests/gold/DECISION_DEFINITION.md`
|
||||||
|
- `tests/gold/decision_deferred/`
|
||||||
|
- `tests/gold/decision_none/`
|
||||||
|
- `prompts/decisions.md`
|
||||||
|
|
||||||
|
## EXP-0013 - Prompt engineering methodology
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Prompt iteration needs strict experimental controls to prevent regression,
|
||||||
|
overfitting and arbitrary prompt churn.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The methodology was documented alongside the Gold Standard corpus and later
|
||||||
|
summarized for agents.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Gold scenarios.
|
||||||
|
- Prompt files.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Applies to all prompt experiments.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The accepted method is one prompt change per iteration, one target test at a
|
||||||
|
time, immediate validation, no regressions, no `expected.json` edits merely to
|
||||||
|
force a pass, no test-specific prompt hacks, stopping after two consecutive
|
||||||
|
non-improving iterations, and verifying unique ground truth before prompt
|
||||||
|
changes.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Prompt changes are controlled experiments. See `AGENTS.md` for agent operating
|
||||||
|
rules.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Most prompt changes are not isolated unless the experiment explicitly constrains
|
||||||
|
the target and regression set.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
||||||
|
- `AGENTS.md`
|
||||||
|
- Commit `f7ad9ba`
|
||||||
|
|
||||||
|
## EXP-0014 - Decision Prompt Version 2
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Adding explicit process-decision language to the decision prompt can preserve
|
||||||
|
true decision detection while recognizing deferrals.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Prompt Version 2 added explicit support for deferrals and process decisions.
|
||||||
|
The baseline was validated on three decision scenarios.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `decision_simple`
|
||||||
|
- `decision_deferred`
|
||||||
|
- `decision_none`
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Prompt Version 2.
|
||||||
|
- Model used for validation is not recorded in the committed methodology.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The committed methodology records all three baseline scenarios as passing. The
|
||||||
|
current generated `actual.json` files also show the expected decision count for
|
||||||
|
these decision scenarios, although some non-decision categories remain less
|
||||||
|
complete.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Prompt Version 2 is the current decision baseline. Explicit deferrals are
|
||||||
|
recognized as process decisions.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Decision-count success does not imply all categories are solved. Category-level
|
||||||
|
evaluation must continue.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
||||||
|
- `tests/gold/decision_simple/actual.json`
|
||||||
|
- `tests/gold/decision_deferred/actual.json`
|
||||||
|
- `tests/gold/decision_none/actual.json`
|
||||||
|
- `prompts/decisions.md`
|
||||||
|
|
||||||
|
## EXP-0015 - Difficult synthetic meeting
|
||||||
|
|
||||||
|
Status: Running
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
A deliberately adversarial synthetic meeting can expose extraction failures
|
||||||
|
that simple category tests miss.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
`evil_meeting` includes interruptions, corrections, absent referenced people,
|
||||||
|
near-decisions, changed positions and one expected explicit decision.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `tests/gold/evil_meeting/`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Existing generated `actual.json`; exact model is not stored in the artifact.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The generated result found the expected FR-7 exclusion decision. It also
|
||||||
|
classified "do not migrate until the mapping table is checked" as an additional
|
||||||
|
decision. The current expected file treats that statement as a position, but
|
||||||
|
contextual review suggests it may be a valid process instruction or decision.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Do not classify this as a simple model failure without reviewing the gold
|
||||||
|
standard. The scenario exposes a semantic gap in the expected output.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Difficult synthetic cases are valuable because they reveal ambiguity in the
|
||||||
|
specification as well as model mistakes.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `tests/gold/evil_meeting/README.md`
|
||||||
|
- `tests/gold/evil_meeting/expected.json`
|
||||||
|
- `tests/gold/evil_meeting/actual.json`
|
||||||
|
- See EXP-0011 and EXP-0012.
|
||||||
|
|
||||||
|
## EXP-0016 - Context-size extraction comparisons
|
||||||
|
|
||||||
|
Status: Rejected
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Increasing extraction context from one chunk to neighboring chunk groups should
|
||||||
|
make extraction more complete and therefore should become the baseline.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3
|
||||||
|
and 1+2+3.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Normalized meeting chunks.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for
|
||||||
|
extraction experiments.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
More context sometimes improved completeness, but it also shifted category
|
||||||
|
classification, added duplicates and reduced stability. No numeric winner is
|
||||||
|
recorded in the repository.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Do not adopt larger extraction windows as the baseline. Independent chunk
|
||||||
|
extraction remains current strategy. Recover global context through
|
||||||
|
consolidation rather than continuously enlarging extraction windows.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Context size is not a monotonic quality knob. It changes the task the model is
|
||||||
|
performing.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
|
- `AGENTS.md`
|
||||||
|
- `ROADMAP.md`
|
||||||
|
- See EXP-0002 and EXP-0018.
|
||||||
|
|
||||||
|
## EXP-0017 - Independent full-meeting chunk extraction
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Extracting every normalized chunk independently can produce enough structured
|
||||||
|
material for a useful protocol draft.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
Nine normalized chunks were extracted into separate JSON files and then
|
||||||
|
rendered by the interim protocol builder.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt`
|
||||||
|
through `chunk_09_normalized.txt`.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Existing extraction artifacts do not record model metadata.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The nine extraction files contain facts, decisions, todos, questions and
|
||||||
|
technical details. `meeting_protocol.md` aggregates them into a readable draft.
|
||||||
|
Duplicates, category shifts and synthesis became the dominant limitations.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Independent chunk extraction is useful enough to keep as the baseline, but it
|
||||||
|
requires a consolidation stage.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Per-chunk extraction gives recall-oriented raw material. It does not by itself
|
||||||
|
produce a polished or canonical meeting representation.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json`
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
||||||
|
- `src/meeting_lab/protocol/build_protocol.py`
|
||||||
|
- See EXP-0018.
|
||||||
|
|
||||||
|
## EXP-0018 - Human protocol comparison and output-view split
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
One generated protocol cannot satisfy every use case; protocol output should be
|
||||||
|
separated by purpose and audience.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The interim machine protocol was compared against the desired human protocol
|
||||||
|
shape and then the architecture was revised toward parallel output views.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Interim `meeting_protocol.md`.
|
||||||
|
- Architecture and output-view documentation.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Not applicable; this is a design evaluation.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The human protocol target is denser and organized by purpose and topic rather
|
||||||
|
than extraction categories. The machine extraction retains more context and is
|
||||||
|
useful for recall, but it is not the right direct source for a concise
|
||||||
|
distribution artifact or durable knowledge entry.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution
|
||||||
|
Protocol / Verteilerprotokoll and Knowledge Objects /
|
||||||
|
Wissensdatenbankeintrag.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Rendering is a separate concern from extraction and consolidation. Output views
|
||||||
|
must be parallel renderings of shared semantics, not transformations of one
|
||||||
|
another.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
||||||
|
- `docs/output-views.md`
|
||||||
|
- `docs/architecture.md`
|
||||||
|
- `docs/pipeline.md`
|
||||||
|
- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture`
|
||||||
|
|
||||||
|
## EXP-0019 - Consolidation and Canonical Meeting Knowledge
|
||||||
|
|
||||||
|
Status: Accepted
|
||||||
|
|
||||||
|
Date or period: 2026-07-30
|
||||||
|
|
||||||
|
Hypothesis:
|
||||||
|
|
||||||
|
Extraction, consolidation and output rendering are separate problems and should
|
||||||
|
not be collapsed into one LLM prompt or one protocol file.
|
||||||
|
|
||||||
|
Setup:
|
||||||
|
|
||||||
|
The architecture was refined after the minimal pipeline and protocol draft
|
||||||
|
showed duplicate, synthesis and audience-specific rendering limitations.
|
||||||
|
|
||||||
|
Inputs:
|
||||||
|
|
||||||
|
- Extraction JSON artifacts.
|
||||||
|
- Interim protocol draft.
|
||||||
|
- Architecture and data-model documentation.
|
||||||
|
|
||||||
|
Model / configuration:
|
||||||
|
|
||||||
|
- Not applicable; this is an architectural conclusion.
|
||||||
|
|
||||||
|
Result:
|
||||||
|
|
||||||
|
The accepted design is a planned Canonical Meeting Knowledge layer as the
|
||||||
|
semantic source of truth, with Working Protocol, Distribution Protocol and
|
||||||
|
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
|
||||||
|
preserve evidence, reconcile category shifts and mark contradictions or
|
||||||
|
uncertainty.
|
||||||
|
|
||||||
|
Decision:
|
||||||
|
|
||||||
|
Consolidation is the next major engineering step after stable local extraction.
|
||||||
|
Canonical Meeting Knowledge and final output views are planned, not implemented.
|
||||||
|
|
||||||
|
Lessons learned:
|
||||||
|
|
||||||
|
Global meeting understanding should be recovered by consolidation over
|
||||||
|
evidence-bearing extractions, not by silently changing output views or expanding
|
||||||
|
LLM context indefinitely.
|
||||||
|
|
||||||
|
Evidence:
|
||||||
|
|
||||||
|
- `docs/architecture.md`
|
||||||
|
- `docs/pipeline.md`
|
||||||
|
- `docs/data-models.md`
|
||||||
|
- `docs/output-views.md`
|
||||||
|
- `PROJECT_KNOWLEDGE.md`
|
||||||
|
- `ROADMAP.md`
|
||||||
|
- Commit `5c03ed7`
|
||||||
|
|||||||
Reference in New Issue
Block a user