Document evidence and commitment model

This commit is contained in:
2026-08-11 14:19:19 +02:00
parent fd7d5e1424
commit c2b7b6b4d2
3 changed files with 525 additions and 3 deletions
+48 -1
View File
@@ -10,7 +10,54 @@ Im Mittelpunkt steht nicht die Softwarearchitektur, sondern die Frage:
> **Wie lässt sich aus einem realen Meeting möglichst zuverlässig strukturiertes Wissen extrahieren?** > **Wie lässt sich aus einem realen Meeting möglichst zuverlässig strukturiertes Wissen extrahieren?**
Neue Ideen werden zunächst hier experimentell umgesetzt. Erst wenn sich ein Ansatz bewährt hat, wird er in den eigentlichen *Meeting Assistant* übernommen. Meeting Lab ist die experimentelle R&D-Umgebung fuer den zukuenftigen
*Meeting Assistant*. Es dient gleichzeitig als Forschungsplattform,
Architektur-Spielwiese, Regressionsframework, Benchmark-Umgebung und
Prototypimplementierung. Neue Ideen werden hier untersucht und gegen reale
Meeting-Beispiele validiert, bevor sie fuer das Produkt in Betracht kommen.
Der beabsichtigte Reifeprozess ist:
```text
Research idea
->
Meeting Lab experiment
->
Regression tests
->
Stable architecture
->
Meeting Assistant implementation
```
Nur ausreichend ausgereifte und verifizierte Komponenten sollen in den
Meeting Assistant uebernommen werden. Meeting Lab darf bewusst experimentelle
Ansaetze und Entwicklungszweige enthalten, die verworfen werden oder nie den
Assistant erreichen.
## Meeting Assistant und langfristige Produktentwicklung
Der Meeting Assistant ist als Produktionsanwendung vorgesehen. Seine erste
oeffentliche Beta soll auf einem stabilen Meeting-Lab-MVP basieren und eine
polierte User Experience, Installer, Konfiguration und eine produktionsreife
Pipeline bieten. Eine GUI ist optional; experimentelle Funktionen sollen
standardmaessig nicht aktiviert sein.
Meeting Lab entwickelt sich unabhaengig weiter und bleibt der langfristige
Innovationszweig. Der erwartete Transferpfad lautet:
```text
Meeting Lab Alpha
->
Meeting Lab Beta
->
Meeting Assistant Beta
->
Meeting Assistant Release
```
Meeting-Assistant-Releases uebernehmen damit gezielt bewaehrte Meeting-Lab-
Komponenten, waehrend der Assistant den stabilen Produktzweig bildet.
--- ---
+189 -2
View File
@@ -2,13 +2,63 @@
## Purpose ## Purpose
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts. The **Meeting Lab** is the experimental R&D environment for developing and
evaluating methods to extract structured knowledge from real meeting
transcripts. It is the research platform, architecture playground, regression
framework, benchmark environment and prototype implementation for the future
Meeting Assistant.
Its purpose is not to build a complete meeting assistant, but to answer a single question: Its purpose is not to build a complete meeting assistant, but to answer a single question:
> **How can knowledge be extracted from real discussions as reliably as possible?** > **How can knowledge be extracted from real discussions as reliably as possible?**
Successful approaches will later be integrated into the Meeting Assistant project. Its purpose is to validate ideas before they are promoted into the product.
Only sufficiently mature and verified components should migrate into Meeting
Assistant. Meeting Lab may intentionally contain experiments or development
branches that are rejected, remain inconclusive or never reach the Assistant.
## Meeting Lab and Meeting Assistant lifecycle
Meeting Lab and Meeting Assistant have different long-term responsibilities:
- **Meeting Lab** is the long-term innovation branch. It favors learning,
inspectable experiments, regression evidence, benchmarks and architectural
change.
- **Meeting Assistant** is the stable product branch. It favors a polished user
experience, installation, configuration and a production pipeline.
Architectural promotion follows an evidence-based lifecycle:
```text
Research idea
->
Meeting Lab experiment
->
Regression tests
->
Stable architecture
->
Meeting Assistant implementation
```
The expected release progression is:
```text
Meeting Lab Alpha
->
Meeting Lab Beta
->
Meeting Assistant Beta
->
Meeting Assistant Release
```
The first public Meeting Assistant beta should be based on a stable Meeting
Lab MVP. It should provide a polished user experience, an installer,
configuration and a production-quality pipeline. A GUI is optional.
Experimental features should not be enabled by default. Meeting Lab continues
to evolve independently after components have migrated; promotion does not
turn the Lab itself into the product branch.
--- ---
@@ -356,6 +406,127 @@ The architecture document only describes the overall system.
--- ---
# Version 2 Accepted Architectural Direction
The following topics are accepted architectural goals for Version 2. They
record direction reached through the BUG-011 through BUG-015 investigations;
they are not descriptions of implemented behavior or authorization to change
the current pipeline.
## Speaker diarization before semantic analysis
Version 2 should determine **who is speaking before semantic analysis**.
Speaker identity contains evidence that cannot reliably be reconstructed from
text alone. It helps distinguish, for example, who answers a question, accepts
work, agrees with a proposal, or advances the discussion after another
speaker. It also preserves conversational flow that anonymous transcript text
can erase.
Diarization is therefore a semantic prerequisite in the intended Version 2
architecture, not merely a display enhancement. Its output should remain
traceable to transcript segments so later stages can preserve speaker and
source provenance.
## Persistent speaker identification
Version 2 should add a persistent speaker database and an interactive identity
workflow during import:
```text
Unknown speaker detected
->
Representative audio sample (approximately 20 seconds)
->
User selects an existing identity or creates a new identity
->
Known speaker available for future recognition
```
Automatic recognition may suggest an identity, but user confirmation governs
the persistent association. Over the long term, speaker embeddings rather
than raw meeting recordings should be the persistent recognition
representation. Representative raw audio is an import and confirmation aid,
not the intended durable identity store. Privacy, deletion and false-match
handling require separate design before implementation.
## Meeting Context as a probabilistic prior
Known meeting participants should influence semantic interpretation, but
Meeting Context is a **probabilistic prior**, not a deterministic semantic
rule. It can make one interpretation more plausible and help focus review; it
must never manufacture a commitment, decision or responsibility assignment.
For example, if Marleen is confirmed as present, “Marleen müsste sich mal
äußern” is more likely to be conversation management directed at a current
participant than future project work. Presence alone does not prove this
interpretation, and it does not establish an Action Item or responsibility.
Explicit meeting evidence remains authoritative. This extends, rather than
weakens, the responsibility attribution invariant.
The existing Meeting Context V2 entity direction is documented in
[`adr-meeting-context-v2-entity-registry.md`](adr-meeting-context-v2-entity-registry.md).
Speaker identities and meeting-specific participant confirmation should
eventually feed that context without turning registry metadata into semantic
facts.
## Conversation Management versus Meeting Content
BUG-015 reinforced that not every utterance is protocol-worthy content.
Version 2 should conceptually distinguish:
- **Conversation Management**: utterances that coordinate the meeting itself,
such as asking a present participant to speak, moderation, requesting a
slide, or asking someone to repeat something.
- **Meeting Content**: propositions that may contribute to the meeting's
durable knowledge, including facts, technical findings, decisions, action
items and open questions.
This is an architectural concept, not a currently implemented category or
filter. The distinction should prevent conversational coordination from being
promoted into project commitments while retaining sufficient provenance to
understand dialogue. Context and diarization can inform the distinction, but
neither should act as a deterministic keyword or participant rule.
## Evidence and commitment before protocol eligibility
The BUG-015 design study concludes that semantic state should be classified
before policy determines whether an item is eligible for a protocol. A binary
keep/reject verifier conflates evidence recognition with publication policy
and loses valid intermediate states.
The proposed decision progression is:
```text
idea -> option -> proposal -> preferred option -> tentative agreement -> decision
```
The proposed action progression is:
```text
possible next step -> recommendation -> requested action
-> established action -> ongoing work -> completed
```
Questions combine a communicative **kind** with an independent **resolution
state**, rather than treating every uncertainty or interrogative as an Open
Question. Responsibility remains an independent dimension and may be recorded
only when explicitly assigned, accepted or confirmed. Evidence strength is
also independent: it describes support for a semantic label, not semantic
maturity or protocol eligibility.
After semantic state classification, explicit policy should select Decisions,
Action Items and Open Questions for a particular output view. Renderers should
receive policy-selected semantic content and must not promote proposals or
conversation management into commitments. The complete taxonomy, trade-offs,
architecture interactions and migration questions are recorded in
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md).
Implementation is intentionally postponed until this architecture has been
reviewed. No Version 2 goal in this section changes current extraction,
canonicalization, consolidation or rendering behavior.
---
# Current State # Current State
Implemented: Implemented:
@@ -389,6 +560,22 @@ Action Item may have no known owner, while a named owner requires explicit
assignment, volunteering or acceptance. These are semantic LLM classifications; assignment, volunteering or acceptance. These are semantic LLM classifications;
deterministic validation must not guess intent from keywords. deterministic validation must not guess intent from keywords.
### Classification Verifier
Before extraction output is normalized for Canonicalizer input, Decision,
Action Item and Open Question candidates pass through a semantic precision
gate. Facts and technical details pass through unchanged. Each candidate is
reviewed independently with its evidence and bounded local chunk context; the
verifier may only keep or reject the existing candidate. It cannot add or
rewrite semantic items.
Verifier output contains the stable candidate ID, category, `keep|reject`
verdict, evidence-based reason and `responsibility_supported`. A kept Action
Item with an unsupported named owner is retained with its responsibility
cleared. Malformed output fails the extraction verification substage closed,
after preserving candidate input and raw response. Per-candidate results and an
aggregate audit remain traceable before Canonicalizer input is written.
## Semantic Consolidator failure handling ## Semantic Consolidator failure handling
Semantic Consolidator V0 preserves every raw model response before parsing. Semantic Consolidator V0 preserves every raw model response before parsing.
+288
View File
@@ -0,0 +1,288 @@
# Evidence / Commitment Model
Status: architecture design study for BUG-015; no implementation recommendation is made here.
## Motivation
Meeting Lab currently asks extraction and the BUG-015 verifier to cross a category boundary in one step: a candidate is either a protocol-worthy Decision, Action Item or Open Question, or it is rejected. That combines two different questions:
1. What does the transcript provide evidence for?
2. Which evidenced states should a particular protocol view publish?
Meetings do not move directly from absence to commitment. An option can become a proposal, then a preferred option, then a tentative agreement and finally a decision. Work can be suggested, requested, assigned, accepted, already underway or completed. A question can be asked and answered, or can remain explicitly unresolved. Rejecting everything below the final protocol threshold discards useful evidence; promoting it creates false commitments.
The proposed model therefore records the evidenced semantic state first. A later policy selects protocol-worthy states. It preserves the existing separation between extraction, deterministic canonicalization, semantic consolidation and rendering, and it does not make rendered output the semantic source of truth.
## Observed BUG-015 failures
The BUG-015 Gold cases expose both sides of the binary-classification problem.
| Progeo-derived evidence | Correct evidence state | Binary failure to avoid |
| --- | --- | --- |
| “Die Option steht im Raum, das Material chemisch recyceln zu lassen.” | option | False Decision |
| “Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden.” | personal preference with a conditional option | False Decision |
| “Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht. … das ist entschieden.” | explicit decision: rejection | False rejection |
| “Die Geometrie kann man vielleicht noch optimieren.” | possible next step | False Action Item |
| “Marleen müsste vielleicht mal äußern …” | suggested/requested action; no accepted or confirmed responsibility | False Action Item and false owner |
| “Wir könnten Dirk Textor vielleicht noch einmal kontaktieren.” | suggestion | False Action Item |
| “Nina, übernimmst du …? – Ja, ich übernehme …” | accepted action with explicit responsibility and deadline | False rejection |
| “Den CET-Artikel erstellen wir bereits …” | ongoing established work; owner not evidenced | Consistent false rejection by the binary verifier |
| “Ob sich das Waschen lohnt, weiß ich nicht.” | uncertainty | False Open Question |
| “Gibt es schon ein Programm? – Ja …” | explicit, resolved question | False Open Question if local resolution is ignored |
| “Welche Daten … dürfen wir veröffentlichen? … weiterhin ungeklärt.” | explicit unresolved question | Consistent false rejection by the binary verifier |
The verifier was reliable on several negative cases and on explicit commitment language, but not on valid states whose evidence did not resemble a fresh agreement: already-established work and an explicitly unresolved information need. This suggests a representation problem, not merely an insufficient keep/reject prompt.
## Model shape
The model is deliberately small and compositional. Each evidence item has:
- a **domain**: decision, action or information need;
- a **semantic state** within that domain;
- an **evidence support level** describing how directly the transcript supports that label;
- preserved transcript evidence and source references;
- domain-specific independent attributes, such as responsibility or resolution;
- no automatic claim that the item belongs in a final protocol.
Semantic maturity and evidence support are independent. An explicitly worded proposal remains a proposal; it is not a weak decision. Conversely, ongoing work can be strongly evidenced without a recorded moment of assignment.
## Semantic state diagrams
The arrows show common progressions, not mandatory workflows. Meetings may enter at any state, skip states, regress, or end without commitment.
### Decision evolution
```text
idea -> option -> proposal -> preferred_option -> tentative_agreement -> decision
| | | | |
+----------+--------------+--------------------+----------> withdrawn
superseded
reopened
```
An **idea** is an undeveloped possibility. An **option** is a candidate alternative. A **proposal** asks the group to adopt an outcome. A **preferred option** expresses comparative preference without settlement. A **tentative agreement** records provisional convergence that is explicitly conditional or awaiting confirmation. A **decision** records an outcome the meeting settled, selected, approved, rejected or committed to. The intermediate preferred and tentative states matter because they prevent likely direction from being confused with commitment.
### Action evolution
```text
possible_next_step -> recommendation -> requested_action -> established_action -> ongoing_work -> completed
| | | | |
+------------------+-----------------+--------------------+----------> cancelled
Responsibility (independent):
unset -> proposed_responsible -> assigned -> accepted/confirmed
```
An action can enter directly as **established_action** through explicit assignment, acceptance or commitment. It can enter directly as **ongoing_work** when the transcript clearly says that the work is already being performed, as in the CET example. `proposed_responsible` is evidence about a suggestion, not permission to populate the protocol's responsible field. Only `assigned`, `accepted` or `confirmed` supports recorded responsibility, and only where the meeting evidence explicitly assigns, accepts or confirms it.
### Information-need evolution
```text
uncertainty ---------> information_need(unresolved) ---------> resolved
curiosity -----------> explicit_question(unresolved) --------> resolved
request_for_clarification(unresolved) -----------------------> resolved
missing_information(unresolved) -----------------------------> resolved
rhetorical_question -> discourse only
```
Uncertainty and curiosity do not automatically create an information need. An explicit question may be resolved immediately. “Open Question” is therefore a policy result derived primarily from a qualifying kind plus `unresolved`, rather than a primitive utterance category.
## Proposed taxonomy
### Common fields
| Field | Proposed values | Purpose |
| --- | --- | --- |
| `domain` | `decision`, `action`, `information_need` | Selects the domain taxonomy. |
| `state` | Domain-specific values below | Records what the evidence says now. |
| `support` | `indirect`, `direct`, `explicit` | Rates support for the chosen state, not protocol importance. |
| `polarity` | `positive`, `negative` where applicable | Preserves approval/rejection and adoption/refusal without rewriting meaning. |
| `evidence` | transcript excerpt(s) | Proves the state and attributes. |
| `source_refs` | stable source references | Retains provenance across stages. |
`support` is intentionally not `weak / medium / strong`. Those terms mix confidence, semantic maturity and number of witnesses. The proposed meanings are narrower:
- `indirect`: the state is supported by context but not stated in a self-contained utterance; downstream commitment policy should normally be conservative.
- `direct`: an utterance directly expresses the state, such as a proposal, preference, ongoing-work statement or question.
- `explicit`: the utterance also names the decisive status, for example “das ist entschieden”, “ich übernehme”, “wir erstellen bereits” or “weiterhin ungeklärt”.
This common axis simplifies evidence auditing and threshold policy, but cannot replace domain state. An explicit suggestion is still not an Action Item, and an explicit uncertainty is still not an Open Question. Model confidence, if retained at all, must be a separate operational field and must not be presented as evidence strength.
### Decision states
| State | Meaning | Default protocol treatment |
| --- | --- | --- |
| `idea` | Undeveloped possibility or brainstorming contribution | Exclude |
| `option` | Candidate alternative under consideration | Exclude |
| `proposal` | Outcome offered for adoption | Exclude |
| `preferred_option` | Expressed preference among alternatives | Exclude |
| `tentative_agreement` | Provisional convergence with an expressed condition or pending confirmation | Exclude from Decisions; potentially expose to an editorial view |
| `decision` | Settled, selected, approved, rejected or committed outcome | Include as Decision when evidence is sufficient |
| `reopened` | Earlier decision explicitly returned to unresolved consideration | Do not render the earlier decision as currently settled without qualification |
| `superseded` | Earlier decision replaced by a later one | Retain provenance; normally render only the current decision |
| `withdrawn` | Candidate state explicitly withdrawn | Exclude as current commitment |
`idea`, `option`, `proposal`, `preferred_option`, `tentative_agreement` and `decision` are necessary distinctions for BUG-015. `reopened`, `superseded`, `withdrawn` are lifecycle states needed to avoid treating historical evidence as current policy; they need not be first-iteration extraction targets.
### Action states and attributes
| State | Meaning | Default protocol treatment |
| --- | --- | --- |
| `possible_next_step` | Hypothetical or exploratory action | Exclude |
| `recommendation` | Action advocated but not established as work | Exclude |
| `requested_action` | Someone asks that work be done, without enough evidence that it is established | Exclude by default |
| `established_action` | Concrete future work established by assignment, acceptance, commitment or confirmation | Include as Action Item |
| `ongoing_work` | Concrete work explicitly already underway | Include when still relevant; do not require a newly witnessed assignment |
| `completed` | Work explicitly reported complete | Exclude from open Action Items; retain as status/history |
| `cancelled` | Work explicitly cancelled or declined | Exclude from open Action Items; retain provenance |
`suggestion` is represented as `possible_next_step` or `recommendation`, depending on whether the speaker advocates it. `accepted_action` and `assigned_work` should not be competing lifecycle states: both establish `established_action`, while the independent commitment basis records `accepted`, `assigned`, `self_committed` or `confirmed_existing`. This avoids an artificial choice when, as with Nina, an assignment and acceptance occur together.
Responsibility is independent:
| Responsibility status | Meaning | May populate `responsible`? |
| --- | --- | --- |
| `unset` | No person/team explicitly tied to ownership | No |
| `proposed` | A person is suggested, asked speculatively or mentioned near the work | No |
| `assigned` | The meeting explicitly assigns the work | Yes |
| `accepted` | The party explicitly accepts or volunteers | Yes |
| `confirmed` | Existing ownership is explicitly confirmed | Yes |
This preserves the project invariant: discussion, expertise, adjacency, organizational role and likely ownership never establish responsibility. An action may be valid with `responsibility_status: unset`, as with established CET work.
### Information-need kinds and resolution
Question form and resolution are separate attributes.
| Kind | Meaning | Can become an Open Question? |
| --- | --- | --- |
| `uncertainty` | Speaker expresses doubt or lack of certainty without establishing a concrete need | No, by itself |
| `curiosity` | Interest without a concrete need requiring follow-up | No, by itself |
| `explicit_question` | Direct interrogative seeking an answer | Yes, if unresolved |
| `request_for_clarification` | Explicit request to clarify a concrete matter | Yes, if unresolved |
| `missing_information` | Concrete required information is stated as absent | Yes, if unresolved |
| `rhetorical_question` | Interrogative used for emphasis rather than an answer | No |
Resolution is one of `unresolved`, `resolved`, or `resolution_unclear`. `resolved_question` is therefore not a separate kind: it is, for example, `explicit_question + resolved`. The publication example is `explicit_question + unresolved`; the event-program example is `explicit_question + resolved`; the washing example is `uncertainty` unless later evidence establishes a concrete unresolved need. `resolution_unclear` preserves evidence but should not silently pass a precision-first Open Question policy.
An “unresolved question” is a derived, protocol-relevant combination rather than a fourth kind. This makes the classification testable: kind answers what communicative act occurred, and resolution answers what remained at the end of the available context.
## Progeo examples under the model
| Evidence | Proposed representation | Working Protocol policy result |
| --- | --- | --- |
| Chemical recycling “Option steht im Raum” | `decision / option / explicit` | Not a Decision |
| Technikum preference | `decision / preferred_option / direct`; personal scope preserved | Not a Decision |
| Schlummer collaboration rejected and “entschieden” | `decision / decision / explicit / negative` | Decision |
| Geometry “kann man vielleicht … optimieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item |
| Marleen “müsste vielleicht mal” | `action / requested_action / direct`, responsibility proposed only | Not an Action Item; no responsible party |
| Textor “könnten … kontaktieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item |
| Nina asks and accepts the review by Friday | `action / established_action / explicit`, basis assigned + accepted, responsible Nina, deadline Friday | Action Item |
| CET article “erstellen wir bereits” and later circulation | `action / ongoing_work / explicit`, responsibility unset | Action Item without invented owner |
| Washing “ob sich das lohnt, weiß ich nicht” | `information_need / uncertainty / direct`, resolution unclear | Not an Open Question |
| Event program asked and answered | `information_need / explicit_question / direct`, resolved | Not an Open Question |
| Publishable energy-audit data “weiterhin ungeklärt” | `information_need / explicit_question / explicit`, unresolved | Open Question |
These labels preserve every BUG-015 distinction without treating rejected protocol candidates as meaningless.
## Architectural impact
### Extraction
Extraction could emit evidence records with richer labels instead of immediately claiming `decision`, `action_item` or `open_question`. This is a bounded increase in extraction vocabulary, not a request for larger context windows, a multi-chunk strategy or a monolithic synthesis prompt. Existing Facts, Positions and Technical Details remain distinct categories; the new domains refine only commitment-sensitive content.
The immediate output would describe the observed state. Protocol category selection would happen later through an explicit policy. Some policy rules can be deterministic once semantic labels are trustworthy, for example:
```text
Decision := domain=decision AND state=decision
Action Item := domain=action AND state IN {established_action, ongoing_work}
Open Question := domain=information_need
AND kind IN {explicit_question, request_for_clarification, missing_information}
AND resolution=unresolved
Responsible := responsibility_status IN {assigned, accepted, confirmed}
```
This keeps downstream selection simple, but does not make semantic extraction deterministic. The LLM still has to distinguish proposals from decisions and uncertainty from unresolved needs.
### LLM stability
The architecture should reduce one source of instability: the model no longer has to erase a supported proposition merely because it falls below a protocol threshold, nor relocate it into another category to retain it. Labels correspond more closely to observable speech acts and lifecycle statements, and downstream thresholds become explicit.
It will not eliminate instability. More labels create adjacent-class boundaries, local context may not reveal resolution, and indirect language remains difficult. Stability depends on small schemas, evidence spans, one best supported state, conservative handling of `resolution_unclear`, and later evaluation against the Gold corpus. A generic evidence-strength score alone would likely worsen instability because it invites subjective grading.
### Canonicalizer
The Canonicalizer would benefit from carrying normalized state names, independent responsibility fields, resolution, polarity and stable source references. It could deterministically validate allowed combinations, normalize aliases, preserve evidence, group exact duplicates and reject structurally impossible combinations. It must not promote a proposal to a decision, infer resolution, infer responsibility or perform uncertain semantic merging.
Richer labels make safe comparisons easier: two `option` records can be recognized as candidates about the same subject without being merged into a `decision`; an `explicit_question + resolved` item is not confused with an unresolved instance. Provenance improves because a later decision can link back to earlier option/proposal evidence rather than overwriting it.
Semantic merging becomes better informed, but not automatically easy. Equivalence and lifecycle transitions remain semantic work. A future consolidator should preserve all source evidence, distinguish duplicate evidence from state evolution, and mark contradictions or uncertainty rather than collapsing them.
### Renderer and output policy
The renderer should not make commitment classification decisions. A policy/view-selection stage should provide protocol-worthy semantic states to the renderer. The normal Working or Distribution Protocol renderer should therefore receive Decisions, established/ongoing Action Items and unresolved information needs, not raw proposals.
Proposals may still be visible to a renderer for a view explicitly designed to show them, such as a discussion appendix or editorial drafting view. That access must be typed and intentional; a renderer must never silently format a proposal under “Decisions.” Different output views may select different states while sharing the same Canonical Meeting Knowledge.
### Future BPD-style protocol generation
Richer labels would give editorial generation better material without granting it authority to invent commitment. A BPD-style view could explain the path from options through tentative agreement to decision, separate current obligations from completed work, and describe why an information need remains open. Proposals and preferred options could be useful appendix material when the requested view values discussion history.
The editorial layer may condense or order these records, but it must preserve state, polarity, responsibility status and provenance. Appendix inclusion is a view policy, not a semantic promotion.
## Trade-offs
Benefits:
- preserves valid evidence below the final-protocol threshold;
- separates meeting semantics from publication policy;
- represents established work without inventing an assignment event or owner;
- treats question kind and resolution independently;
- makes responsibility independently auditable;
- improves provenance across lifecycle changes;
- prevents renderers from silently converting proposals into commitments;
- gives future editorial views useful, explicitly non-final material.
Costs and risks:
- expands the extraction schema and Gold specification;
- introduces adjacent semantic labels that require precise definitions;
- requires consolidation to distinguish duplicates from state transitions;
- may retain more evidence records, increasing storage and review volume;
- depends on adequate local context to determine resolution and lifecycle state;
- requires explicit view policies so consumers do not treat every evidence record as protocol-worthy;
- cannot guarantee LLM consistency merely by replacing a binary verdict with a taxonomy.
The taxonomy should remain closed and small. In particular, modality, confidence, support, commitment state and responsibility must not be collapsed into a single scalar.
## Migration strategy
This is a staged design path, not a recommendation to implement it now.
1. Specify the taxonomy and invariants against the existing BUG-015 Gold candidates and the cited Progeo excerpts, without changing current expected outputs.
2. Annotate a design-only mapping from each current Decision, Action Item and Open Question example to evidence state, support and independent attributes. Confirm objectively unique ground truth, especially for `requested_action` versus `possible_next_step` and `tentative_agreement` versus `preferred_option`.
3. Define a versioned evidence-record schema and deterministic protocol-selection policy on paper. Keep responsibility and question resolution independent.
4. Design evaluation measures separately for state accuracy, resolution accuracy, responsibility accuracy and protocol-policy output. A correct lower state must not count as a false protocol commitment.
5. Only after the design is validated, plan an experiment that compares the evidence-first approach with current extraction and the binary verifier. Any prompt or Python changes would be separate, explicitly scoped tasks following the Gold methodology and LLM safety rules.
6. If later adopted, preserve compatibility through an adapter that maps protocol-worthy evidence states to the current canonical categories while Canonical Meeting Knowledge evolves. Do not ask the renderer to interpret raw states.
No database, prompt, code, test or current output-schema migration is proposed by this document.
## Open questions
- Is `requested_action` useful as a distinct state when a direct assignment already creates `established_action`, or should it be limited to unaccepted requests?
- Does `tentative_agreement` have sufficiently unique evidence in the current corpus, or should it remain an annotation until more Gold examples exist?
- Should an ongoing-work item whose relevance to the meeting is unclear pass the Working Protocol policy, or require an explicit continuation/follow-up signal?
- How much local context is required to label a question `resolved` safely when the answer occurs across a technical chunk boundary?
- Should `resolution_unclear` be retained only in Canonical Meeting Knowledge, or also exposed to a review view?
- Should support be limited to `direct / explicit`, treating `indirect` as review-only, to reduce subjective classification?
- How should consolidation represent one subject moving from proposal to decision: linked immutable evidence records, or a current-state object with a preserved event history?
- Which output views, if any, should include withdrawn proposals, cancelled work and resolved questions?
- Can polarity and lifecycle links be normalized deterministically without introducing semantic inference?
## Recommendation
The evidence-first architecture should replace the current binary keep/reject classification approach as the target architecture. The replacement should be conceptual and staged, not implemented yet: extract a small, domain-specific semantic state plus independent support, responsibility and resolution attributes; preserve evidence and provenance; then apply explicit policy to produce protocol categories.
A single common Evidence Strength axis should complement this taxonomy but must not replace it. The decisive improvement is separation of semantic state from protocol eligibility. That separation directly explains the BUG-015 false positives and false negatives, supports the Progeo cases without invented ownership, simplifies view selection, and provides a sounder basis for Canonical Meeting Knowledge and future BPD-style rendering.