Document project architecture and development methodology
- add AGENTS.md with development and prompt-engineering rules - add PROJECT_KNOWLEDGE.md summarizing current architecture and findings - add CHANGELOG.md - add ROADMAP.md - establish experiments.md as the project's experiment log - document Canonical Meeting Knowledge architecture - document Output Views and Knowledge Objects - capture accepted experimental results and engineering methodology
This commit is contained in:
@@ -0,0 +1,124 @@
|
||||
# AGENTS.md
|
||||
|
||||
Practical instructions for coding agents working in Meeting Lab.
|
||||
|
||||
## Project Purpose
|
||||
|
||||
Meeting Lab extracts and structures organizational knowledge from meeting
|
||||
recordings. It is an experimental local discussion analyzer, not merely a
|
||||
one-step meeting-protocol generator.
|
||||
|
||||
Successful approaches may later move into the Meeting Assistant project.
|
||||
|
||||
## Current Pipeline
|
||||
|
||||
Current and intended flow:
|
||||
|
||||
```text
|
||||
Audio
|
||||
-> Whisper
|
||||
-> cleanup
|
||||
-> normalization
|
||||
-> chunking
|
||||
-> local chunk extraction
|
||||
-> consolidation
|
||||
-> Canonical Meeting Knowledge
|
||||
-> Output Views
|
||||
```
|
||||
|
||||
Status:
|
||||
|
||||
- Implemented: Whisper JSON cleanup script, normalization, technical chunking,
|
||||
local chunk extraction, interim Markdown protocol builder.
|
||||
- Experimental/prototype: topic segmentation and review tooling.
|
||||
- Planned: consolidation, Canonical Meeting Knowledge implementation, final
|
||||
Output Views.
|
||||
|
||||
## Architectural Principles
|
||||
|
||||
- Canonical Meeting Knowledge is the intended semantic source of truth.
|
||||
- Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag are
|
||||
parallel output views.
|
||||
- Output views must not silently change meaning. They may select, condense or
|
||||
render information for an audience, but not invent new semantics.
|
||||
- Extraction, consolidation, synthesis and rendering are separate concerns.
|
||||
- Prefer small, testable processing stages over one monolithic LLM prompt.
|
||||
- Current extraction strategy is one normalized chunk per LLM call.
|
||||
- Do not expand context windows or redesign the extraction strategy without an
|
||||
explicit experiment.
|
||||
- Deterministic stages should remain deterministic where possible.
|
||||
|
||||
## Prompt Engineering Rules
|
||||
|
||||
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
|
||||
`tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`:
|
||||
|
||||
1. Make only one prompt change per iteration.
|
||||
2. Optimize only one target gold test case at a time.
|
||||
3. Validate every prompt modification immediately.
|
||||
4. Accept a prompt change only if it improves the target and causes no
|
||||
regressions in previously passing gold tests.
|
||||
5. Never modify `expected.json` merely to make a prompt pass.
|
||||
6. Prompt engineering edits prompt files only; Python code changes require a
|
||||
separate explicit task.
|
||||
7. Maintain a prompt evolution log for every iteration.
|
||||
8. Stop arbitrary iterations if small changes do not improve the test; analyze
|
||||
the root cause.
|
||||
9. Avoid gold-test overfitting. Prompt changes must generalize and must not
|
||||
special-case one transcript.
|
||||
10. Stop after two consecutive non-improving prompt iterations and classify the
|
||||
root cause.
|
||||
11. Verify whether the target gold test has objectively unique ground truth
|
||||
before changing a prompt for unexpected behavior.
|
||||
|
||||
Current documented Prompt Version 2 decision baseline:
|
||||
|
||||
- `decision_simple`: passing
|
||||
- `decision_deferred`: passing
|
||||
- `decision_none`: passing
|
||||
|
||||
## LLM Execution Safety
|
||||
|
||||
- Never start a full multi-chunk LLM run unless explicitly requested.
|
||||
- Before any LLM run, state the model, inputs, expected LLM-call count and
|
||||
output location.
|
||||
- Do not retry LLM calls automatically unless explicitly allowed.
|
||||
- Do not download models automatically.
|
||||
- Prefer small-scope validation runs.
|
||||
- Never use generated output as committed source data.
|
||||
- Preserve raw model responses when diagnosing parser or truncation failures.
|
||||
- Do not run Ollama from unit tests.
|
||||
|
||||
## Development Rules
|
||||
|
||||
- Make small, focused changes.
|
||||
- Preserve the existing architecture unless a redesign is explicitly requested.
|
||||
- Add regression tests for bugs.
|
||||
- Run non-LLM tests before committing when code changes are made.
|
||||
- Do not commit generated transcripts, audio, extraction JSON, protocol output
|
||||
or temporary files.
|
||||
- Report files changed, tests run and assumptions.
|
||||
- Do not commit or push unless explicitly requested.
|
||||
|
||||
## Repository Conventions
|
||||
|
||||
- `README.md`: project overview and current high-level status.
|
||||
- `docs/`: architecture, pipeline, data model and output-view documentation.
|
||||
- `prompts/`: extraction and segmentation prompts. Treat prompt edits as
|
||||
controlled experiments.
|
||||
- `tests/gold/`: Gold Standard corpus and semantic specification for extraction
|
||||
behavior.
|
||||
- `scripts/`: command-line support scripts such as Whisper cleanup and gold
|
||||
test execution.
|
||||
- `src/meeting_lab/normalization/`: deterministic transcript cleanup.
|
||||
- `src/meeting_lab/chunking/`: technical chunk creation; chunks are not topics.
|
||||
- `src/meeting_lab/segmentation/`: experimental topic segmentation tooling.
|
||||
- `src/meeting_lab/extraction/`: local LLM extraction flow and category
|
||||
extractor modules.
|
||||
- `src/meeting_lab/consolidation/`: planned consolidation area.
|
||||
- `src/meeting_lab/protocol/`: interim protocol rendering.
|
||||
- `src/meeting_lab/models/`: current lightweight data models.
|
||||
- `samples/`: sample inputs and generated/experimental artifacts; do not treat
|
||||
sample output as canonical source data.
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
# Changelog
|
||||
|
||||
## Unreleased
|
||||
|
||||
### Added
|
||||
|
||||
- Initial Meeting Lab project structure with source, docs, prompts, samples and
|
||||
tests directories.
|
||||
- Deterministic transcript normalization module.
|
||||
- Technical transcript chunking with block-aligned chunk generation.
|
||||
- Whisper JSON cleanup support.
|
||||
- Local Ollama-based chunk extraction flow.
|
||||
- Interim Markdown meeting protocol builder for technical validation.
|
||||
- Initial and windowed topic segmentation prototypes plus review tooling.
|
||||
- Gold Standard extraction corpus and gold-test runner.
|
||||
- Prompt loading support and common/decision prompt baseline.
|
||||
- Output-view architecture documentation for Canonical Meeting Knowledge,
|
||||
Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag.
|
||||
|
||||
### Changed
|
||||
|
||||
- Refined the architecture from a single meeting protocol toward Canonical
|
||||
Meeting Knowledge as the planned semantic source of truth.
|
||||
- Clarified that Working Protocol, Distribution Protocol and Knowledge Objects
|
||||
are parallel renderings, not derived from one another.
|
||||
- Updated decision extraction semantics to include explicit process decisions
|
||||
and deferrals.
|
||||
- Simplified extraction prompt assembly around prompt files.
|
||||
|
||||
### Fixed
|
||||
|
||||
- Fixed Whisper JSON chunk extraction to prefer `segments[*].text` over the
|
||||
aggregate top-level `text` field.
|
||||
- Added chunking tests to ensure chunks do not duplicate later blocks when no
|
||||
overlap is requested.
|
||||
- Added parser handling for model responses that contain thinking text before
|
||||
the final JSON object.
|
||||
|
||||
### Documentation
|
||||
|
||||
- Added architecture, pipeline and data-model documentation.
|
||||
- Added output-view documentation.
|
||||
- Added Gold Standard prompt-engineering methodology.
|
||||
- Added formal decision-definition documentation.
|
||||
- Added scenario README files for the Gold Standard corpus.
|
||||
|
||||
@@ -0,0 +1,151 @@
|
||||
# Project Knowledge
|
||||
|
||||
This is a compact operational summary of the current Meeting Lab state.
|
||||
|
||||
## Objective
|
||||
|
||||
Meeting Lab develops and evaluates local methods for extracting structured
|
||||
organizational knowledge from real meeting recordings and transcripts. The
|
||||
project is a research and validation environment for a future Meeting
|
||||
Assistant, not a finished product.
|
||||
|
||||
## Implemented Pipeline Stages
|
||||
|
||||
Implemented:
|
||||
|
||||
- Whisper JSON cleanup via `scripts/clean_whisper_json.py`.
|
||||
- Transcript normalization in `src/meeting_lab/normalization/`.
|
||||
- Technical chunking in `src/meeting_lab/chunking/`.
|
||||
- Local per-chunk extraction in `src/meeting_lab/extraction/extract_chunks.py`.
|
||||
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||
gold-test runner validation.
|
||||
|
||||
Experimental/prototype:
|
||||
|
||||
- Topic segmentation in `src/meeting_lab/segmentation/`.
|
||||
- Windowed segmentation and review output in `samples/chunks/`.
|
||||
- Gold Standard extraction corpus under `tests/gold/`.
|
||||
|
||||
Planned:
|
||||
|
||||
- Consolidation of extraction results.
|
||||
- Canonical Meeting Knowledge implementation as the semantic source of truth.
|
||||
- Final Working Protocol / Arbeitsprotokoll, Distribution Protocol /
|
||||
Verteilerprotokoll and Knowledge Objects / Wissensdatenbankeintrag renderers.
|
||||
|
||||
## Source Tree
|
||||
|
||||
```text
|
||||
src/meeting_lab/
|
||||
chunking/ technical transcript chunking
|
||||
consolidation/ planned merge/consolidation area
|
||||
extraction/ current local LLM extraction flow
|
||||
io/ lightweight file and JSON helpers
|
||||
llm/ Ollama and prompt support
|
||||
models/ current lightweight model definitions
|
||||
normalization/ deterministic transcript cleanup
|
||||
protocol/ interim Markdown protocol builder
|
||||
segmentation/ experimental topic segmentation tooling
|
||||
```
|
||||
|
||||
Supporting areas:
|
||||
|
||||
- `docs/`: architecture, pipeline, data models and output-view concepts.
|
||||
- `prompts/`: active prompt files. Only `common.md` and `decisions.md` contain
|
||||
substantive extraction prompt text in the current tree.
|
||||
- `tests/gold/`: semantic gold tests and prompt-engineering methodology.
|
||||
- `samples/`: sample inputs and generated or experimental artifacts.
|
||||
- `scripts/`: operational scripts for cleanup and gold-test execution.
|
||||
|
||||
## Current Model Strategy
|
||||
|
||||
The current extraction strategy is one normalized chunk per LLM call. This is
|
||||
preferred over expanding context windows or asking one model call to analyze a
|
||||
full meeting.
|
||||
|
||||
Known working models from current project notes and experiment practice:
|
||||
|
||||
- `qwen3:1.7b`: useful for smoke tests.
|
||||
- `qwen3.5:9b`: useful for meaningful extraction and segmentation work.
|
||||
|
||||
LLM calls use Ollama locally. The current extractor defaults to `qwen3:8b`, but
|
||||
validated work may specify another model explicitly.
|
||||
|
||||
## Important Findings
|
||||
|
||||
- Whisper JSON chunking must use `segments[*].text`, not only the top-level
|
||||
`text` field.
|
||||
- Independent chunk extraction is currently preferred.
|
||||
- Larger context windows can change classification behavior and increase
|
||||
instability.
|
||||
- Extraction and consolidation are separate problems.
|
||||
- Generation limits can truncate JSON.
|
||||
- Qwen thinking may be returned separately by the Ollama API.
|
||||
- Gold Standard tests are also a formal specification of meeting semantics.
|
||||
- Raw model responses should be preserved when diagnosing parser or truncation
|
||||
failures.
|
||||
|
||||
## Decision Taxonomy
|
||||
|
||||
Accepted decision semantics:
|
||||
|
||||
- A decision is an explicit agreement that creates a binding change in action,
|
||||
process, responsibility, approval status, timing or next step.
|
||||
- Included: substantive decisions, organizational decisions, process decisions,
|
||||
approvals, rejections, deferrals, explicit agreement not to decide yet, and
|
||||
explicit agreement to gather more information before deciding.
|
||||
- Excluded: opinions, preferences, proposals without agreement, open questions,
|
||||
current-state descriptions and explanations without commitment.
|
||||
- A process decision to defer a substantive decision is still a decision.
|
||||
- "No decision was reached" is different from "the group decided to defer the
|
||||
decision."
|
||||
|
||||
Current Prompt Version 2 decision baseline:
|
||||
|
||||
- `decision_simple`: passing.
|
||||
- `decision_deferred`: passing.
|
||||
- `decision_none`: passing.
|
||||
- Prompt Version 2 explicitly supports process decisions where the group agrees
|
||||
to defer a substantive decision until more information is available.
|
||||
|
||||
## Canonical Knowledge Architecture
|
||||
|
||||
Canonical Meeting Knowledge is the planned semantic intermediate model and
|
||||
future single source of truth. It should preserve topics, facts, decisions,
|
||||
action items, open questions, positions, technical details, rationale,
|
||||
uncertainty, contradictions and source evidence.
|
||||
|
||||
Output views are planned as independent renderings from that canonical model:
|
||||
|
||||
- Working Protocol / Arbeitsprotokoll: relatively complete, optimized for
|
||||
recall and traceability.
|
||||
- Distribution Protocol / Verteilerprotokoll: concise and outcome-oriented,
|
||||
optimized for circulation.
|
||||
- Knowledge Objects / Wissensdatenbankeintrag: durable organizational knowledge
|
||||
optimized for reuse.
|
||||
|
||||
The current `meeting_protocol.md` builder is an interim technical validation
|
||||
tool, not the final output-view architecture.
|
||||
|
||||
## Current Limitations
|
||||
|
||||
- Discussion Blocks are documented as a stable semantic unit but are not yet a
|
||||
separate implemented pipeline artifact.
|
||||
- Topic segmentation exists as prototype tooling, not a stable pipeline stage.
|
||||
- Extraction is still a combined current flow, even though separate extractors
|
||||
are the intended architecture.
|
||||
- Consolidation is not implemented.
|
||||
- Canonical Meeting Knowledge is documented but not implemented.
|
||||
- Final output views are documented but not implemented.
|
||||
- Most prompt files are placeholders except the common and decision prompts.
|
||||
- Gold tests currently emphasize extraction semantics, especially decisions.
|
||||
|
||||
## Next Recommended Engineering Step
|
||||
|
||||
Stabilize repeatable local extraction evaluation before broadening the pipeline:
|
||||
expand Gold Standard coverage by category, keep one-chunk extraction as the
|
||||
baseline, and use small prompt experiments with immediate non-regression checks.
|
||||
After extraction behavior is stable enough, implement consolidation with
|
||||
evidence retention as the next major pipeline stage.
|
||||
+185
@@ -0,0 +1,185 @@
|
||||
# Roadmap
|
||||
|
||||
No dates are assigned. Phases describe dependency order, not release promises.
|
||||
|
||||
## Phase 1 - Stable Local Extraction
|
||||
|
||||
Goal:
|
||||
|
||||
- Establish reliable per-chunk extraction behavior for core meeting semantics.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Stronger Gold Standard coverage across facts, positions, decisions, todos,
|
||||
questions and technical details.
|
||||
- Improved category prompts.
|
||||
- Repeatable evaluation workflow.
|
||||
- Documented prompt experiment log.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Existing chunk extraction flow.
|
||||
- Existing Gold Standard runner and methodology.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Full-transcript LLM extraction.
|
||||
- Larger context-window strategy changes without an explicit experiment.
|
||||
- Consolidation or final protocol rendering.
|
||||
|
||||
## Phase 2 - Consolidation
|
||||
|
||||
Goal:
|
||||
|
||||
- Merge independent extraction results into a coherent meeting-level
|
||||
representation without losing evidence.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Duplicate merging.
|
||||
- Evidence retention.
|
||||
- Category-shift reconciliation, especially facts versus positions and
|
||||
positions versus decisions.
|
||||
- Contradiction and uncertainty markers.
|
||||
- Consolidated meeting representation.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Stable local extraction baseline.
|
||||
- Gold tests that expose cross-chunk duplication and category shifts.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Final Canonical Meeting Knowledge schema.
|
||||
- User-facing protocol polish.
|
||||
- Retrieval or RAG integration.
|
||||
|
||||
## Phase 3 - Canonical Meeting Knowledge
|
||||
|
||||
Goal:
|
||||
|
||||
- Define and implement the semantic intermediate model that becomes the source
|
||||
of truth for downstream outputs.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Canonical Meeting Knowledge schema.
|
||||
- Source evidence and traceability fields.
|
||||
- Clear distinction between durable knowledge and meeting-specific actions.
|
||||
- Migration path from consolidated extraction JSON into the canonical model.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Consolidation behavior that preserves evidence and uncertainty.
|
||||
- Agreement on required semantic categories.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- GUI.
|
||||
- Export formats beyond those needed to validate the model.
|
||||
- Knowledge-system storage design.
|
||||
|
||||
## Phase 4 - Output Views
|
||||
|
||||
Goal:
|
||||
|
||||
- Render purpose-specific outputs from Canonical Meeting Knowledge without
|
||||
changing meaning.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Working Protocol / Arbeitsprotokoll renderer.
|
||||
- Distribution Protocol / Verteilerprotokoll renderer.
|
||||
- Knowledge Objects / Wissensdatenbankeintrag renderer or structured export.
|
||||
- Tests or checks showing that output views are parallel renderings of the same
|
||||
canonical model.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Implemented Canonical Meeting Knowledge.
|
||||
- Clear audience and completeness rules for each output view.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Additional analysis during rendering.
|
||||
- Deriving one output view from another.
|
||||
- Retrieval integration.
|
||||
|
||||
## Phase 5 - Review and Quality Control
|
||||
|
||||
Goal:
|
||||
|
||||
- Add optional review stages that improve omission detection, consistency and
|
||||
model selection.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Optional whole-transcript review.
|
||||
- Omission detection.
|
||||
- Consistency checks.
|
||||
- Model comparison workflow.
|
||||
- Hardware and runtime benchmarks.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Stable extraction, consolidation and canonical model.
|
||||
- Representative test meetings.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Automatic acceptance of review suggestions without evidence.
|
||||
- Product UI work.
|
||||
- Cloud deployment.
|
||||
|
||||
## Phase 6 - Productization
|
||||
|
||||
Goal:
|
||||
|
||||
- Turn the validated pipeline into a usable local workflow.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Recording/transcription workflow.
|
||||
- FFmpeg integration.
|
||||
- Meeting metadata capture.
|
||||
- Participant entry.
|
||||
- GUI.
|
||||
- Stable deployment process.
|
||||
- Export workflows.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Stable pipeline stages and output views.
|
||||
- Clear operational requirements for local use.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Enterprise knowledge retrieval.
|
||||
- Future Meeting Assistant integration beyond export contracts.
|
||||
- Cloud-first architecture.
|
||||
|
||||
## Phase 7 - Knowledge-System Integration
|
||||
|
||||
Goal:
|
||||
|
||||
- Reuse durable meeting knowledge in broader knowledge systems.
|
||||
|
||||
Deliverables:
|
||||
|
||||
- Structured Knowledge Objects.
|
||||
- Retrieval-ready storage format.
|
||||
- Future RAG integration path.
|
||||
- Reuse contracts for Meeting Assistant and other knowledge systems.
|
||||
|
||||
Prerequisites:
|
||||
|
||||
- Canonical Meeting Knowledge and Knowledge Objects are implemented and stable.
|
||||
- Durable knowledge is separated from meeting-specific actions and discussion
|
||||
history.
|
||||
|
||||
Out of scope:
|
||||
|
||||
- Building a full enterprise search product inside Meeting Lab.
|
||||
- Treating raw transcripts or generated protocols as the knowledge source of
|
||||
truth.
|
||||
|
||||
@@ -0,0 +1,971 @@
|
||||
# Experiments
|
||||
|
||||
This file records durable technical experiments and findings for Meeting Lab.
|
||||
It is not a diary and does not replace commit history.
|
||||
|
||||
## Status values
|
||||
|
||||
- Proposed: experiment idea exists, but no result is recorded.
|
||||
- Running: experiment is in progress and no decision has been made.
|
||||
- Accepted: finding is the current baseline or design conclusion.
|
||||
- Rejected: hypothesis was tested and should not be repeated as-is.
|
||||
- Superseded: finding was useful but has been replaced by a newer baseline.
|
||||
|
||||
## EXP-0001 - Whisper JSON interpretation
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-29
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Whisper JSON should be chunked from its segment stream, not from the aggregate
|
||||
top-level text field.
|
||||
|
||||
Setup:
|
||||
|
||||
`chunk_transcript.py` was updated to parse JSON input and prefer
|
||||
`segments[*].text` when `segments` exists. A regression test supplies JSON with
|
||||
both top-level `text` and separate segment texts.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Minimal synthetic Whisper-style JSON in `tests/test_chunking.py`.
|
||||
- Real Whisper artifacts under `samples/whisper/`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- No LLM.
|
||||
|
||||
Result:
|
||||
|
||||
The test verifies that the block stream is `["alpha", "beta", "gamma"]` and
|
||||
does not include the aggregate `"alpha beta gamma"` text. The repository history
|
||||
records this as the fix for the earlier failure where the first chunk contained
|
||||
the complete transcript.
|
||||
|
||||
Decision:
|
||||
|
||||
When `segments` exists, `segments[*].text` is the authoritative transcript
|
||||
stream. The top-level `text` field is only a fallback.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Whisper JSON is structured input. Treating it like plain text can duplicate the
|
||||
entire transcript and invalidate downstream chunking.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||
- `tests/test_chunking.py`
|
||||
- Commit `4656523` - `Fix Whisper JSON chunk extraction`
|
||||
|
||||
## EXP-0002 - Technical transcript chunking baseline
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-29 to 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Sequential technical chunks around the configured target size can preserve the
|
||||
transcript while keeping extraction calls small enough for local models.
|
||||
|
||||
Setup:
|
||||
|
||||
The chunker splits block-aligned text with configurable target, minimum,
|
||||
maximum and overlap settings. Tests verify no duplicate later blocks when
|
||||
overlap is zero. Generated meeting artifacts contain a nine-chunk manifest.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Synthetic block list in `tests/test_chunking.py`.
|
||||
- `samples/whisper/meeting_speech_cleaned.json`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- No LLM for chunking.
|
||||
- Manifest uses default chunking behavior recorded in
|
||||
`samples/whisper/meeting_speech_cleaned_chunks/manifest.json`.
|
||||
|
||||
Result:
|
||||
|
||||
With overlap set to zero, tests verify that all blocks appear exactly once. The
|
||||
real sample manifest contains nine chunks, mostly near the configured target
|
||||
size, with a smaller final chunk.
|
||||
|
||||
Decision:
|
||||
|
||||
Independent sequential chunks are the current technical baseline. One
|
||||
normalized chunk per extraction call is the preferred extraction strategy.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Chunking solves model-size constraints only. It must not perform topic
|
||||
detection or semantic merging.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||
- `tests/test_chunking.py`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/manifest.json`
|
||||
- `AGENTS.md`
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
|
||||
## EXP-0003 - Conservative transcript normalization
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Transcript cleanup should improve readability without changing meeting
|
||||
semantics.
|
||||
|
||||
Setup:
|
||||
|
||||
The normalizer removes isolated filler sounds, immediate duplicate words or
|
||||
short duplicate phrases, and redundant whitespace. It records changed blocks in
|
||||
a JSON change log and explicitly preserves semantic content categories.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Chunk text files under `samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||
- Change logs such as `chunk_01_changes.json`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- No LLM.
|
||||
|
||||
Result:
|
||||
|
||||
The implementation and generated change logs show a conservative policy:
|
||||
negations, qualifiers, dates, numbers, responsibilities, technical statements,
|
||||
deadlines, decisions and commitments are preserved.
|
||||
|
||||
Decision:
|
||||
|
||||
Normalization remains deterministic and low-risk. When uncertain, leave text
|
||||
unchanged.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Filler removal is useful only if it is tightly scoped. Broad cleanup can remove
|
||||
semantic cues needed by extraction.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/normalization/normalize_transcript.py`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_changes.json`
|
||||
- `docs/pipeline.md`
|
||||
|
||||
## EXP-0004 - Full-context topic segmentation
|
||||
|
||||
Status: Superseded
|
||||
|
||||
Date or period: 2026-07-21 to 2026-07-22
|
||||
|
||||
Hypothesis:
|
||||
|
||||
A single full-context topic segmentation call can identify topic boundaries in
|
||||
a normalized transcript chunk.
|
||||
|
||||
Setup:
|
||||
|
||||
The initial segmentation prototype asked the model for topic changes and then
|
||||
converted those boundaries into continuous, non-overlapping segments.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `samples/chunks/chunk_01_normalized.txt`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Generated artifact records `qwen3:14b`.
|
||||
|
||||
Result:
|
||||
|
||||
The generated artifact contains 85 blocks, three topic-change boundaries and
|
||||
four segments. The run metadata records a substantially longer elapsed time
|
||||
than the later windowed artifact for the same input.
|
||||
|
||||
Decision:
|
||||
|
||||
Full-context segmentation was useful as a prototype, but it was superseded by
|
||||
windowed segmentation and manual review tooling.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
The prototype established the boundary-to-segment representation, but did not
|
||||
settle segmentation quality.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/segmentation/segment_topics.py`
|
||||
- `samples/chunks/chunk_01_normalized_segments.json`
|
||||
- Commit `f234efc` - `Add initial topic segmentation prototype`
|
||||
- Commit `889a4fe` - `Detect topic boundaries as continuous segments`
|
||||
|
||||
## EXP-0005 - Windowed topic segmentation and review
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-22
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Windowed topic segmentation can reduce runtime and make boundary evaluation
|
||||
more inspectable than a single full-context call.
|
||||
|
||||
Setup:
|
||||
|
||||
`segment_topics_windowed.py` analyzes overlapping windows and reports only
|
||||
boundaries from the decision range. Python merges boundaries into continuous,
|
||||
non-overlapping segments. `review_segmentation.py` renders each boundary with
|
||||
neighboring transcript context for manual classification.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `samples/chunks/chunk_01_normalized.txt`.
|
||||
- Full meeting normalized chunks under
|
||||
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- `samples/chunks` artifact: `qwen3:8b`, window size 20, overlap 3.
|
||||
- Full-meeting chunk artifacts: `qwen3.5:9b`, window size 20, overlap 3.
|
||||
|
||||
Result:
|
||||
|
||||
The `samples/chunks` windowed artifact produced 11 boundaries and 12 segments
|
||||
for 85 blocks, with recorded elapsed time lower than the full-context artifact.
|
||||
Manual review output shows that some boundaries were assessed as subtopics
|
||||
rather than full topic changes. Full-meeting artifacts show one window per
|
||||
already-small normalized chunk and two segments per chunk.
|
||||
|
||||
Decision:
|
||||
|
||||
Windowed segmentation and review tooling are accepted as prototype tooling, not
|
||||
as a stable production segmentation stage.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Windowing improves inspectability and can reduce runtime, but it can also
|
||||
cluster boundaries and over-segment. Manual review remains necessary.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/segmentation/segment_topics_windowed.py`
|
||||
- `src/meeting_lab/segmentation/review_segmentation.py`
|
||||
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
||||
- `samples/chunks/chunk_01_normalized_windowed_segments_review.md`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
||||
- Commit `1a6d731` - `Add windowed segmentation pipeline and review tooling`
|
||||
|
||||
## EXP-0006 - Qwen model comparison
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Larger local Qwen-family models should improve meaningful extraction and
|
||||
segmentation, but model size alone will not solve prompt or pipeline problems.
|
||||
|
||||
Setup:
|
||||
|
||||
Project work used smaller models for smoke checks and larger local models for
|
||||
meaningful extraction or segmentation. Artifacts and project knowledge record
|
||||
the currently useful model roles.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Gold Standard scenarios under `tests/gold/`.
|
||||
- Generated segmentation artifacts under `samples/`.
|
||||
- Generated extraction artifacts under
|
||||
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- `qwen3:1.7b`: smoke-test model according to project knowledge.
|
||||
- `qwen3.5:9b`: current meaningful extraction and segmentation model according
|
||||
to project knowledge and generated full-meeting segmentation artifacts.
|
||||
- `qwen3:8b` and `qwen3:14b`: present in earlier segmentation artifacts.
|
||||
|
||||
Result:
|
||||
|
||||
The repository supports the conclusion that `qwen3.5:9b` is the meaningful
|
||||
current experiment model and `qwen3:1.7b` is useful for smoke tests. Larger
|
||||
models and longer contexts may increase runtime substantially, but no hardware
|
||||
benchmark suite is recorded.
|
||||
|
||||
Decision:
|
||||
|
||||
Use `qwen3:1.7b` for smoke tests and `qwen3.5:9b` for meaningful current
|
||||
experiments. Do not assume model size alone fixes prompt or pipeline design.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Evaluation must separate model capability from prompt clarity, context
|
||||
strategy, extraction schema and consolidation.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/*_windowed_segments.json`
|
||||
- `samples/chunks/chunk_01_normalized_segments.json`
|
||||
- `samples/chunks/chunk_01_normalized_windowed_segments.json`
|
||||
|
||||
## EXP-0007 - Thinking output and Ollama API behavior
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Thinking-capable Qwen models may return reasoning separately from the final
|
||||
answer, and extraction parsing should not fail merely because extra text or
|
||||
multiple JSON objects appear.
|
||||
|
||||
Setup:
|
||||
|
||||
The Ollama response reader checks `response`, chat-style `message.content` and
|
||||
then `thinking`. The JSON parser tries a full parse first, then scans JSON
|
||||
object candidates and returns the final valid object. A regression test covers
|
||||
thinking text before final JSON.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Synthetic parser test in `tests/test_extraction_protocol.py`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- No LLM run in the test.
|
||||
- Code path is used by Ollama extraction.
|
||||
|
||||
Result:
|
||||
|
||||
The parser can handle additional text and multiple JSON objects where the final
|
||||
valid object is the intended answer. Current code still falls back to `thinking`
|
||||
only if no usable response or message content is present.
|
||||
|
||||
Decision:
|
||||
|
||||
Keep parser robustness, but do not treat thinking output as the root cause of
|
||||
all extraction failures.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
API response shape and model output shape are separate concerns. Preserve raw
|
||||
responses when diagnosing failures.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/extraction/extract_chunks.py`
|
||||
- `tests/test_extraction_protocol.py`
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
|
||||
## EXP-0008 - JSON truncation and generation limits
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Some extraction failures are caused by generation limits truncating JSON rather
|
||||
than by prompt wording or parser behavior.
|
||||
|
||||
Setup:
|
||||
|
||||
A `qwen3.5:9b` extraction failure was diagnosed as truncated JSON. The
|
||||
generation limit was increased for the successful path. Exact failing limit is
|
||||
not recorded in the repository; the current extractor default is verifiably
|
||||
`--num-predict 8192`.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Local extraction runs referenced by project knowledge.
|
||||
- Current extraction CLI.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- `qwen3.5:9b`.
|
||||
- Current extractor default: `num_predict=8192`.
|
||||
|
||||
Result:
|
||||
|
||||
Increasing the generation limit fixed the technical JSON failure. This was not
|
||||
primarily a parser or prompt problem.
|
||||
|
||||
Decision:
|
||||
|
||||
When JSON is truncated, inspect raw output and generation limits before editing
|
||||
prompts.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Invalid JSON can be a runtime-budget symptom. Prompt changes are the wrong
|
||||
first response if the model simply ran out of output tokens.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `src/meeting_lab/extraction/extract_chunks.py`
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `AGENTS.md`
|
||||
|
||||
## EXP-0009 - Minimal end-to-end pipeline
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
A minimal local pipeline can transform Whisper output into chunk extractions
|
||||
and an interim protocol, proving the technical path before the final
|
||||
architecture exists.
|
||||
|
||||
Setup:
|
||||
|
||||
The repository added cleanup, normalization, chunking, extraction and protocol
|
||||
builder scripts, with sample generated artifacts.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `samples/whisper/meeting_speech.json`
|
||||
- `samples/whisper/meeting_speech_cleaned.json`
|
||||
- Generated chunks and normalized chunks under
|
||||
`samples/whisper/meeting_speech_cleaned_chunks/`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Local Ollama extraction for chunk JSON.
|
||||
- Windowed segmentation artifacts use `qwen3.5:9b`.
|
||||
|
||||
Result:
|
||||
|
||||
The repository contains cleaned input, nine chunks, nine normalized chunks, nine
|
||||
extraction JSON files, windowed segmentation artifacts and
|
||||
`meeting_protocol.md`.
|
||||
|
||||
Decision:
|
||||
|
||||
The minimal pipeline is technically validated. The first protocol builder is an
|
||||
interim validation tool, not the final architecture.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
End-to-end execution exposed the next limitation: extraction output needs
|
||||
consolidation and purpose-specific rendering.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `scripts/clean_whisper_json.py`
|
||||
- `src/meeting_lab/normalization/normalize_transcript.py`
|
||||
- `src/meeting_lab/chunking/chunk_transcript.py`
|
||||
- `src/meeting_lab/extraction/extract_chunks.py`
|
||||
- `src/meeting_lab/protocol/build_protocol.py`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/`
|
||||
- Commit `07b0d80` - `Implement first end-to-end meeting analysis pipeline`
|
||||
|
||||
## EXP-0010 - Gold Standard corpus
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Reproducible prompt engineering requires synthetic transcripts with explicit
|
||||
expected semantic outputs.
|
||||
|
||||
Setup:
|
||||
|
||||
The Gold Standard corpus defines scenario directories with `transcript.txt`,
|
||||
`expected.json` and README files describing ground truth and common model
|
||||
mistakes. The runner validates schema keys and writes `actual.json` for a
|
||||
scenario.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Gold scenarios under `tests/gold/`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Runner requires an explicit Ollama model for LLM evaluation.
|
||||
- Existing unit tests for the runner do not invoke Ollama.
|
||||
|
||||
Result:
|
||||
|
||||
The corpus gives stable semantics for decisions, facts, positions, todos,
|
||||
questions, technical details and difficult mixed cases. Initial structured
|
||||
transcripts are Phase 1 and easier than raw Whisper-style transcripts.
|
||||
|
||||
Decision:
|
||||
|
||||
Use Gold Standard tests as both regression tests and formal meeting-semantics
|
||||
specification. Raw or unlabelled transcript cases remain later-phase work.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Without expected outputs, prompt changes cannot be evaluated reproducibly.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tests/gold/`
|
||||
- `scripts/run_gold_test.py`
|
||||
- `tests/test_gold_runner.py`
|
||||
- Commit `f7ad9ba` - `Establish prompt engineering baseline with Gold Standard tests`
|
||||
|
||||
## EXP-0011 - Gold-test quality and unique ground truth
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
If a prompt produces unexpected behavior, the gold test itself may be ambiguous
|
||||
and should be reviewed before the prompt is changed.
|
||||
|
||||
Setup:
|
||||
|
||||
Decision-focused scenarios were clarified during baseline creation. The current
|
||||
methodology requires checking unique ground truth before changing prompts.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `decision_simple`
|
||||
- `decision_deferred`
|
||||
- `decision_none`
|
||||
- Gold methodology document.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Prompt Version 2 baseline work.
|
||||
|
||||
Result:
|
||||
|
||||
`decision_simple` required clarification around the explicit agreement and
|
||||
nearby non-decision wording. The earlier negative/deferral ambiguity is now
|
||||
represented by distinct `decision_none` and `decision_deferred` scenarios in
|
||||
the repository. Punctuation is not reliable evidence for Whisper transcripts;
|
||||
agreement language and wording must carry the semantics.
|
||||
|
||||
Decision:
|
||||
|
||||
Ambiguous gold tests must be reviewed before prompt changes. Do not treat
|
||||
punctuation as reliable evidence in real Whisper-style transcripts.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Bad gold tests create false prompt failures and can encourage overfitting.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
||||
- `tests/gold/decision_simple/README.md`
|
||||
- `tests/gold/decision_deferred/README.md`
|
||||
- `tests/gold/decision_none/README.md`
|
||||
- Commit `f7ad9ba`
|
||||
|
||||
## EXP-0012 - Decision taxonomy
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Decision extraction needs a formal taxonomy that distinguishes substantive
|
||||
decisions from process decisions and non-decisions.
|
||||
|
||||
Setup:
|
||||
|
||||
The decision definition document and decision prompt define included and
|
||||
excluded categories. Gold tests cover explicit decisions, true no-decision
|
||||
cases and deferrals.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `tests/gold/DECISION_DEFINITION.md`
|
||||
- `prompts/decisions.md`
|
||||
- Decision gold scenarios.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Prompt Version 2 baseline.
|
||||
|
||||
Result:
|
||||
|
||||
Accepted decision categories include substantive decisions, organizational
|
||||
decisions, process decisions, approvals, rejections, deferrals, explicit
|
||||
decisions not to decide yet and explicit agreement to gather more information
|
||||
before deciding. Opinions, preferences and proposals without agreement are not
|
||||
decisions.
|
||||
|
||||
Decision:
|
||||
|
||||
"No decision was reached" and "the decision was deferred" are distinct semantic
|
||||
outcomes.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Deferral can be a valid process decision even when the substantive topic remains
|
||||
unresolved.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tests/gold/DECISION_DEFINITION.md`
|
||||
- `tests/gold/decision_deferred/`
|
||||
- `tests/gold/decision_none/`
|
||||
- `prompts/decisions.md`
|
||||
|
||||
## EXP-0013 - Prompt engineering methodology
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Prompt iteration needs strict experimental controls to prevent regression,
|
||||
overfitting and arbitrary prompt churn.
|
||||
|
||||
Setup:
|
||||
|
||||
The methodology was documented alongside the Gold Standard corpus and later
|
||||
summarized for agents.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Gold scenarios.
|
||||
- Prompt files.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Applies to all prompt experiments.
|
||||
|
||||
Result:
|
||||
|
||||
The accepted method is one prompt change per iteration, one target test at a
|
||||
time, immediate validation, no regressions, no `expected.json` edits merely to
|
||||
force a pass, no test-specific prompt hacks, stopping after two consecutive
|
||||
non-improving iterations, and verifying unique ground truth before prompt
|
||||
changes.
|
||||
|
||||
Decision:
|
||||
|
||||
Prompt changes are controlled experiments. See `AGENTS.md` for agent operating
|
||||
rules.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Most prompt changes are not isolated unless the experiment explicitly constrains
|
||||
the target and regression set.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
||||
- `AGENTS.md`
|
||||
- Commit `f7ad9ba`
|
||||
|
||||
## EXP-0014 - Decision Prompt Version 2
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Adding explicit process-decision language to the decision prompt can preserve
|
||||
true decision detection while recognizing deferrals.
|
||||
|
||||
Setup:
|
||||
|
||||
Prompt Version 2 added explicit support for deferrals and process decisions.
|
||||
The baseline was validated on three decision scenarios.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `decision_simple`
|
||||
- `decision_deferred`
|
||||
- `decision_none`
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Prompt Version 2.
|
||||
- Model used for validation is not recorded in the committed methodology.
|
||||
|
||||
Result:
|
||||
|
||||
The committed methodology records all three baseline scenarios as passing. The
|
||||
current generated `actual.json` files also show the expected decision count for
|
||||
these decision scenarios, although some non-decision categories remain less
|
||||
complete.
|
||||
|
||||
Decision:
|
||||
|
||||
Prompt Version 2 is the current decision baseline. Explicit deferrals are
|
||||
recognized as process decisions.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Decision-count success does not imply all categories are solved. Category-level
|
||||
evaluation must continue.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tests/gold/PROMPT_ENGINEERING_METHODOLOGY.md`
|
||||
- `tests/gold/decision_simple/actual.json`
|
||||
- `tests/gold/decision_deferred/actual.json`
|
||||
- `tests/gold/decision_none/actual.json`
|
||||
- `prompts/decisions.md`
|
||||
|
||||
## EXP-0015 - Difficult synthetic meeting
|
||||
|
||||
Status: Running
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
A deliberately adversarial synthetic meeting can expose extraction failures
|
||||
that simple category tests miss.
|
||||
|
||||
Setup:
|
||||
|
||||
`evil_meeting` includes interruptions, corrections, absent referenced people,
|
||||
near-decisions, changed positions and one expected explicit decision.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `tests/gold/evil_meeting/`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Existing generated `actual.json`; exact model is not stored in the artifact.
|
||||
|
||||
Result:
|
||||
|
||||
The generated result found the expected FR-7 exclusion decision. It also
|
||||
classified "do not migrate until the mapping table is checked" as an additional
|
||||
decision. The current expected file treats that statement as a position, but
|
||||
contextual review suggests it may be a valid process instruction or decision.
|
||||
|
||||
Decision:
|
||||
|
||||
Do not classify this as a simple model failure without reviewing the gold
|
||||
standard. The scenario exposes a semantic gap in the expected output.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Difficult synthetic cases are valuable because they reveal ambiguity in the
|
||||
specification as well as model mistakes.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `tests/gold/evil_meeting/README.md`
|
||||
- `tests/gold/evil_meeting/expected.json`
|
||||
- `tests/gold/evil_meeting/actual.json`
|
||||
- See EXP-0011 and EXP-0012.
|
||||
|
||||
## EXP-0016 - Context-size extraction comparisons
|
||||
|
||||
Status: Rejected
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Increasing extraction context from one chunk to neighboring chunk groups should
|
||||
make extraction more complete and therefore should become the baseline.
|
||||
|
||||
Setup:
|
||||
|
||||
Comparisons were made across one chunk and grouped contexts: 1+2, 2 only, 2+3
|
||||
and 1+2+3.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Normalized meeting chunks.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Current project knowledge identifies `qwen3.5:9b` as the meaningful model for
|
||||
extraction experiments.
|
||||
|
||||
Result:
|
||||
|
||||
More context sometimes improved completeness, but it also shifted category
|
||||
classification, added duplicates and reduced stability. No numeric winner is
|
||||
recorded in the repository.
|
||||
|
||||
Decision:
|
||||
|
||||
Do not adopt larger extraction windows as the baseline. Independent chunk
|
||||
extraction remains current strategy. Recover global context through
|
||||
consolidation rather than continuously enlarging extraction windows.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Context size is not a monotonic quality knob. It changes the task the model is
|
||||
performing.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `AGENTS.md`
|
||||
- `ROADMAP.md`
|
||||
- See EXP-0002 and EXP-0018.
|
||||
|
||||
## EXP-0017 - Independent full-meeting chunk extraction
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Extracting every normalized chunk independently can produce enough structured
|
||||
material for a useful protocol draft.
|
||||
|
||||
Setup:
|
||||
|
||||
Nine normalized chunks were extracted into separate JSON files and then
|
||||
rendered by the interim protocol builder.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_01_normalized.txt`
|
||||
through `chunk_09_normalized.txt`.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Existing extraction artifacts do not record model metadata.
|
||||
|
||||
Result:
|
||||
|
||||
The nine extraction files contain facts, decisions, todos, questions and
|
||||
technical details. `meeting_protocol.md` aggregates them into a readable draft.
|
||||
Duplicates, category shifts and synthesis became the dominant limitations.
|
||||
|
||||
Decision:
|
||||
|
||||
Independent chunk extraction is useful enough to keep as the baseline, but it
|
||||
requires a consolidation stage.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Per-chunk extraction gives recall-oriented raw material. It does not by itself
|
||||
produce a polished or canonical meeting representation.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/chunk_*_extraction.json`
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
||||
- `src/meeting_lab/protocol/build_protocol.py`
|
||||
- See EXP-0018.
|
||||
|
||||
## EXP-0018 - Human protocol comparison and output-view split
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
One generated protocol cannot satisfy every use case; protocol output should be
|
||||
separated by purpose and audience.
|
||||
|
||||
Setup:
|
||||
|
||||
The interim machine protocol was compared against the desired human protocol
|
||||
shape and then the architecture was revised toward parallel output views.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Interim `meeting_protocol.md`.
|
||||
- Architecture and output-view documentation.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Not applicable; this is a design evaluation.
|
||||
|
||||
Result:
|
||||
|
||||
The human protocol target is denser and organized by purpose and topic rather
|
||||
than extraction categories. The machine extraction retains more context and is
|
||||
useful for recall, but it is not the right direct source for a concise
|
||||
distribution artifact or durable knowledge entry.
|
||||
|
||||
Decision:
|
||||
|
||||
Separate final outputs into Working Protocol / Arbeitsprotokoll, Distribution
|
||||
Protocol / Verteilerprotokoll and Knowledge Objects /
|
||||
Wissensdatenbankeintrag.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Rendering is a separate concern from extraction and consolidation. Output views
|
||||
must be parallel renderings of shared semantics, not transformations of one
|
||||
another.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `samples/whisper/meeting_speech_cleaned_chunks/meeting_protocol.md`
|
||||
- `docs/output-views.md`
|
||||
- `docs/architecture.md`
|
||||
- `docs/pipeline.md`
|
||||
- Commit `5c03ed7` - `Refine canonical meeting knowledge architecture`
|
||||
|
||||
## EXP-0019 - Consolidation and Canonical Meeting Knowledge
|
||||
|
||||
Status: Accepted
|
||||
|
||||
Date or period: 2026-07-30
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Extraction, consolidation and output rendering are separate problems and should
|
||||
not be collapsed into one LLM prompt or one protocol file.
|
||||
|
||||
Setup:
|
||||
|
||||
The architecture was refined after the minimal pipeline and protocol draft
|
||||
showed duplicate, synthesis and audience-specific rendering limitations.
|
||||
|
||||
Inputs:
|
||||
|
||||
- Extraction JSON artifacts.
|
||||
- Interim protocol draft.
|
||||
- Architecture and data-model documentation.
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Not applicable; this is an architectural conclusion.
|
||||
|
||||
Result:
|
||||
|
||||
The accepted design is a planned Canonical Meeting Knowledge layer as the
|
||||
semantic source of truth, with Working Protocol, Distribution Protocol and
|
||||
Knowledge Objects as parallel output views. Consolidation must merge duplicates,
|
||||
preserve evidence, reconcile category shifts and mark contradictions or
|
||||
uncertainty.
|
||||
|
||||
Decision:
|
||||
|
||||
Consolidation is the next major engineering step after stable local extraction.
|
||||
Canonical Meeting Knowledge and final output views are planned, not implemented.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
Global meeting understanding should be recovered by consolidation over
|
||||
evidence-bearing extractions, not by silently changing output views or expanding
|
||||
LLM context indefinitely.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `docs/architecture.md`
|
||||
- `docs/pipeline.md`
|
||||
- `docs/data-models.md`
|
||||
- `docs/output-views.md`
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `ROADMAP.md`
|
||||
- Commit `5c03ed7`
|
||||
|
||||
Reference in New Issue
Block a user