# Evidence / Commitment Model Status: architecture design study for BUG-015; no implementation recommendation is made here. ## Motivation Meeting Lab currently asks extraction and the BUG-015 verifier to cross a category boundary in one step: a candidate is either a protocol-worthy Decision, Action Item or Open Question, or it is rejected. That combines two different questions: 1. What does the transcript provide evidence for? 2. Which evidenced states should a particular protocol view publish? Meetings do not move directly from absence to commitment. An option can become a proposal, then a preferred option, then a tentative agreement and finally a decision. Work can be suggested, requested, assigned, accepted, already underway or completed. A question can be asked and answered, or can remain explicitly unresolved. Rejecting everything below the final protocol threshold discards useful evidence; promoting it creates false commitments. The proposed model therefore records the evidenced semantic state first. A later policy selects protocol-worthy states. It preserves the existing separation between extraction, deterministic canonicalization, semantic consolidation and rendering, and it does not make rendered output the semantic source of truth. ## Observed BUG-015 failures The BUG-015 Gold cases expose both sides of the binary-classification problem. | Progeo-derived evidence | Correct evidence state | Binary failure to avoid | | --- | --- | --- | | “Die Option steht im Raum, das Material chemisch recyceln zu lassen.” | option | False Decision | | “Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden.” | personal preference with a conditional option | False Decision | | “Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht. … das ist entschieden.” | explicit decision: rejection | False rejection | | “Die Geometrie kann man vielleicht noch optimieren.” | possible next step | False Action Item | | “Marleen müsste vielleicht mal äußern …” | suggested/requested action; no accepted or confirmed responsibility | False Action Item and false owner | | “Wir könnten Dirk Textor vielleicht noch einmal kontaktieren.” | suggestion | False Action Item | | “Nina, übernimmst du …? – Ja, ich übernehme …” | accepted action with explicit responsibility and deadline | False rejection | | “Den CET-Artikel erstellen wir bereits …” | ongoing established work; owner not evidenced | Consistent false rejection by the binary verifier | | “Ob sich das Waschen lohnt, weiß ich nicht.” | uncertainty | False Open Question | | “Gibt es schon ein Programm? – Ja …” | explicit, resolved question | False Open Question if local resolution is ignored | | “Welche Daten … dürfen wir veröffentlichen? … weiterhin ungeklärt.” | explicit unresolved question | Consistent false rejection by the binary verifier | The verifier was reliable on several negative cases and on explicit commitment language, but not on valid states whose evidence did not resemble a fresh agreement: already-established work and an explicitly unresolved information need. This suggests a representation problem, not merely an insufficient keep/reject prompt. ## Model shape The model is deliberately small and compositional. Each evidence item has: - a **domain**: decision, action or information need; - a **semantic state** within that domain; - an **evidence support level** describing how directly the transcript supports that label; - preserved transcript evidence and source references; - domain-specific independent attributes, such as responsibility or resolution; - no automatic claim that the item belongs in a final protocol. Semantic maturity and evidence support are independent. An explicitly worded proposal remains a proposal; it is not a weak decision. Conversely, ongoing work can be strongly evidenced without a recorded moment of assignment. ## Semantic state diagrams The arrows show common progressions, not mandatory workflows. Meetings may enter at any state, skip states, regress, or end without commitment. ### Decision evolution ```text idea -> option -> proposal -> preferred_option -> tentative_agreement -> decision | | | | | +----------+--------------+--------------------+----------> withdrawn superseded reopened ``` An **idea** is an undeveloped possibility. An **option** is a candidate alternative. A **proposal** asks the group to adopt an outcome. A **preferred option** expresses comparative preference without settlement. A **tentative agreement** records provisional convergence that is explicitly conditional or awaiting confirmation. A **decision** records an outcome the meeting settled, selected, approved, rejected or committed to. The intermediate preferred and tentative states matter because they prevent likely direction from being confused with commitment. ### Action evolution ```text possible_next_step -> recommendation -> requested_action -> established_action -> ongoing_work -> completed | | | | | +------------------+-----------------+--------------------+----------> cancelled Responsibility (independent): unset -> proposed_responsible -> assigned -> accepted/confirmed ``` An action can enter directly as **established_action** through explicit assignment, acceptance or commitment. It can enter directly as **ongoing_work** when the transcript clearly says that the work is already being performed, as in the CET example. `proposed_responsible` is evidence about a suggestion, not permission to populate the protocol's responsible field. Only `assigned`, `accepted` or `confirmed` supports recorded responsibility, and only where the meeting evidence explicitly assigns, accepts or confirms it. ### Information-need evolution ```text uncertainty ---------> information_need(unresolved) ---------> resolved curiosity -----------> explicit_question(unresolved) --------> resolved request_for_clarification(unresolved) -----------------------> resolved missing_information(unresolved) -----------------------------> resolved rhetorical_question -> discourse only ``` Uncertainty and curiosity do not automatically create an information need. An explicit question may be resolved immediately. “Open Question” is therefore a policy result derived primarily from a qualifying kind plus `unresolved`, rather than a primitive utterance category. ## Proposed taxonomy ### Common fields | Field | Proposed values | Purpose | | --- | --- | --- | | `domain` | `decision`, `action`, `information_need` | Selects the domain taxonomy. | | `state` | Domain-specific values below | Records what the evidence says now. | | `support` | `indirect`, `direct`, `explicit` | Rates support for the chosen state, not protocol importance. | | `polarity` | `positive`, `negative` where applicable | Preserves approval/rejection and adoption/refusal without rewriting meaning. | | `evidence` | transcript excerpt(s) | Proves the state and attributes. | | `source_refs` | stable source references | Retains provenance across stages. | `support` is intentionally not `weak / medium / strong`. Those terms mix confidence, semantic maturity and number of witnesses. The proposed meanings are narrower: - `indirect`: the state is supported by context but not stated in a self-contained utterance; downstream commitment policy should normally be conservative. - `direct`: an utterance directly expresses the state, such as a proposal, preference, ongoing-work statement or question. - `explicit`: the utterance also names the decisive status, for example “das ist entschieden”, “ich übernehme”, “wir erstellen bereits” or “weiterhin ungeklärt”. This common axis simplifies evidence auditing and threshold policy, but cannot replace domain state. An explicit suggestion is still not an Action Item, and an explicit uncertainty is still not an Open Question. Model confidence, if retained at all, must be a separate operational field and must not be presented as evidence strength. ### Decision states | State | Meaning | Default protocol treatment | | --- | --- | --- | | `idea` | Undeveloped possibility or brainstorming contribution | Exclude | | `option` | Candidate alternative under consideration | Exclude | | `proposal` | Outcome offered for adoption | Exclude | | `preferred_option` | Expressed preference among alternatives | Exclude | | `tentative_agreement` | Provisional convergence with an expressed condition or pending confirmation | Exclude from Decisions; potentially expose to an editorial view | | `decision` | Settled, selected, approved, rejected or committed outcome | Include as Decision when evidence is sufficient | | `reopened` | Earlier decision explicitly returned to unresolved consideration | Do not render the earlier decision as currently settled without qualification | | `superseded` | Earlier decision replaced by a later one | Retain provenance; normally render only the current decision | | `withdrawn` | Candidate state explicitly withdrawn | Exclude as current commitment | `idea`, `option`, `proposal`, `preferred_option`, `tentative_agreement` and `decision` are necessary distinctions for BUG-015. `reopened`, `superseded`, `withdrawn` are lifecycle states needed to avoid treating historical evidence as current policy; they need not be first-iteration extraction targets. ### Action states and attributes | State | Meaning | Default protocol treatment | | --- | --- | --- | | `possible_next_step` | Hypothetical or exploratory action | Exclude | | `recommendation` | Action advocated but not established as work | Exclude | | `requested_action` | Someone asks that work be done, without enough evidence that it is established | Exclude by default | | `established_action` | Concrete future work established by assignment, acceptance, commitment or confirmation | Include as Action Item | | `ongoing_work` | Concrete work explicitly already underway | Include when still relevant; do not require a newly witnessed assignment | | `completed` | Work explicitly reported complete | Exclude from open Action Items; retain as status/history | | `cancelled` | Work explicitly cancelled or declined | Exclude from open Action Items; retain provenance | `suggestion` is represented as `possible_next_step` or `recommendation`, depending on whether the speaker advocates it. `accepted_action` and `assigned_work` should not be competing lifecycle states: both establish `established_action`, while the independent commitment basis records `accepted`, `assigned`, `self_committed` or `confirmed_existing`. This avoids an artificial choice when, as with Nina, an assignment and acceptance occur together. Responsibility is independent: | Responsibility status | Meaning | May populate `responsible`? | | --- | --- | --- | | `unset` | No person/team explicitly tied to ownership | No | | `proposed` | A person is suggested, asked speculatively or mentioned near the work | No | | `assigned` | The meeting explicitly assigns the work | Yes | | `accepted` | The party explicitly accepts or volunteers | Yes | | `confirmed` | Existing ownership is explicitly confirmed | Yes | This preserves the project invariant: discussion, expertise, adjacency, organizational role and likely ownership never establish responsibility. An action may be valid with `responsibility_status: unset`, as with established CET work. ### Information-need kinds and resolution Question form and resolution are separate attributes. | Kind | Meaning | Can become an Open Question? | | --- | --- | --- | | `uncertainty` | Speaker expresses doubt or lack of certainty without establishing a concrete need | No, by itself | | `curiosity` | Interest without a concrete need requiring follow-up | No, by itself | | `explicit_question` | Direct interrogative seeking an answer | Yes, if unresolved | | `request_for_clarification` | Explicit request to clarify a concrete matter | Yes, if unresolved | | `missing_information` | Concrete required information is stated as absent | Yes, if unresolved | | `rhetorical_question` | Interrogative used for emphasis rather than an answer | No | Resolution is one of `unresolved`, `resolved`, or `resolution_unclear`. `resolved_question` is therefore not a separate kind: it is, for example, `explicit_question + resolved`. The publication example is `explicit_question + unresolved`; the event-program example is `explicit_question + resolved`; the washing example is `uncertainty` unless later evidence establishes a concrete unresolved need. `resolution_unclear` preserves evidence but should not silently pass a precision-first Open Question policy. An “unresolved question” is a derived, protocol-relevant combination rather than a fourth kind. This makes the classification testable: kind answers what communicative act occurred, and resolution answers what remained at the end of the available context. ## Progeo examples under the model | Evidence | Proposed representation | Working Protocol policy result | | --- | --- | --- | | Chemical recycling “Option steht im Raum” | `decision / option / explicit` | Not a Decision | | Technikum preference | `decision / preferred_option / direct`; personal scope preserved | Not a Decision | | Schlummer collaboration rejected and “entschieden” | `decision / decision / explicit / negative` | Decision | | Geometry “kann man vielleicht … optimieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item | | Marleen “müsste vielleicht mal” | `action / requested_action / direct`, responsibility proposed only | Not an Action Item; no responsible party | | Textor “könnten … kontaktieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item | | Nina asks and accepts the review by Friday | `action / established_action / explicit`, basis assigned + accepted, responsible Nina, deadline Friday | Action Item | | CET article “erstellen wir bereits” and later circulation | `action / ongoing_work / explicit`, responsibility unset | Action Item without invented owner | | Washing “ob sich das lohnt, weiß ich nicht” | `information_need / uncertainty / direct`, resolution unclear | Not an Open Question | | Event program asked and answered | `information_need / explicit_question / direct`, resolved | Not an Open Question | | Publishable energy-audit data “weiterhin ungeklärt” | `information_need / explicit_question / explicit`, unresolved | Open Question | These labels preserve every BUG-015 distinction without treating rejected protocol candidates as meaningless. ## Thematic protocol as the primary structure The Evidence / Commitment Model supplies semantic distinctions inside a larger meeting reconstruction. It does not imply that the primary protocol should be organized as one section per semantic category. The central Version 2 requirement is: > **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.** A primary protocol should reconstruct which topics were discussed, what relevant information emerged within each topic, how the discussion developed, which alternatives, ideas, objections or proposals mattered, what outcome or current state was reached, and which actions or unresolved questions resulted from that topic. Semantic categories remain essential, but as metadata and supporting structure attached to topics. Facts, technical findings, ideas, alternatives, proposals, objections, decisions, action items and open questions must not dictate the main document structure. For example, a human-style topic section may read: ```text ## Trial setup Several variants for the next trial were discussed. A thinner carrier material was proposed as one possible alternative. Concerns were raised regarding its mechanical suitability. The group therefore decided to continue with the existing setup for the next trial. Nina will obtain the remaining samples before the next production run. The publication question remains unresolved. ``` This preserves the relationship between the proposal, objection, decision, resulting work and unresolved question. A primary document split into separate Ideas, Objections, Decisions and Action Items sections would lose that thematic and conversational relationship. Category-oriented outputs remain valuable as derived secondary views. An Action Item table, Decision register, Open Questions list or Management summary can be generated from the same underlying meeting knowledge after thematic reconstruction. These indexes and summaries do not replace the primary topic-oriented protocol. The conceptual Version 2 flow is: ```text Transcript -> Evidence Extraction -> Topic Reconstruction -> Semantic Synthesis -> Protocol Rendering ``` Evidence items and their semantic metadata remain attached to topics and support Semantic Synthesis. The Version 2 renderer should receive topic-oriented semantic knowledge rather than a flat collection grouped by category. This direction does not define the final Topic Reconstruction schema, redesign the current pipeline or remove the existing Working Protocol V2 architecture. It is an accepted target architecture whose implementation is postponed. ## Architectural impact ### Extraction Extraction could emit evidence records with richer labels instead of immediately claiming `decision`, `action_item` or `open_question`. This is a bounded increase in extraction vocabulary, not a request for larger context windows, a multi-chunk strategy or a monolithic synthesis prompt. Existing Facts, Positions and Technical Details remain distinct categories; the new domains refine only commitment-sensitive content. The immediate output would describe the observed state. Protocol category selection would happen later through an explicit policy. Some policy rules can be deterministic once semantic labels are trustworthy, for example: ```text Decision := domain=decision AND state=decision Action Item := domain=action AND state IN {established_action, ongoing_work} Open Question := domain=information_need AND kind IN {explicit_question, request_for_clarification, missing_information} AND resolution=unresolved Responsible := responsibility_status IN {assigned, accepted, confirmed} ``` This keeps downstream selection simple, but does not make semantic extraction deterministic. The LLM still has to distinguish proposals from decisions and uncertainty from unresolved needs. ### LLM stability The architecture should reduce one source of instability: the model no longer has to erase a supported proposition merely because it falls below a protocol threshold, nor relocate it into another category to retain it. Labels correspond more closely to observable speech acts and lifecycle statements, and downstream thresholds become explicit. It will not eliminate instability. More labels create adjacent-class boundaries, local context may not reveal resolution, and indirect language remains difficult. Stability depends on small schemas, evidence spans, one best supported state, conservative handling of `resolution_unclear`, and later evaluation against the Gold corpus. A generic evidence-strength score alone would likely worsen instability because it invites subjective grading. ### Canonicalizer The Canonicalizer would benefit from carrying normalized state names, independent responsibility fields, resolution, polarity and stable source references. It could deterministically validate allowed combinations, normalize aliases, preserve evidence, group exact duplicates and reject structurally impossible combinations. It must not promote a proposal to a decision, infer resolution, infer responsibility or perform uncertain semantic merging. Richer labels make safe comparisons easier: two `option` records can be recognized as candidates about the same subject without being merged into a `decision`; an `explicit_question + resolved` item is not confused with an unresolved instance. Provenance improves because a later decision can link back to earlier option/proposal evidence rather than overwriting it. Semantic merging becomes better informed, but not automatically easy. Equivalence and lifecycle transitions remain semantic work. A future consolidator should preserve all source evidence, distinguish duplicate evidence from state evolution, and mark contradictions or uncertainty rather than collapsing them. ### Renderer and output policy The renderer should not make commitment classification decisions. A policy/view-selection stage should determine protocol-worthy semantic states while preserving their topic associations. In the Version 2 target architecture, the primary protocol renderer should receive topic-oriented semantic knowledge containing eligible evidence and state transitions, not a flat category-oriented collection. It must not silently promote raw proposals to Decisions or present conversational coordination as Meeting Content. Proposals may still be visible to a renderer for a view explicitly designed to show them, such as a discussion appendix or editorial drafting view. That access must be typed and intentional; a renderer must never silently format a proposal under “Decisions.” Different output views may select different states while sharing the same Canonical Meeting Knowledge. Category-oriented renderers remain valid for secondary views such as Action Item tables, Decision registers and Open Questions lists. This accepted direction supplements rather than removes the current Working Protocol V2 architecture; it does not change current rendering behavior. ### Future BPD-style protocol generation Richer labels would give editorial generation better material without granting it authority to invent commitment. A BPD-style view could explain the path from options through tentative agreement to decision, separate current obligations from completed work, and describe why an information need remains open. Proposals and preferred options could be useful appendix material when the requested view values discussion history. The editorial layer may condense or order these records, but it must preserve state, polarity, responsibility status and provenance. Appendix inclusion is a view policy, not a semantic promotion. ## Trade-offs Benefits: - preserves valid evidence below the final-protocol threshold; - separates meeting semantics from publication policy; - represents established work without inventing an assignment event or owner; - treats question kind and resolution independently; - makes responsibility independently auditable; - improves provenance across lifecycle changes; - prevents renderers from silently converting proposals into commitments; - gives future editorial views useful, explicitly non-final material. Costs and risks: - expands the extraction schema and Gold specification; - introduces adjacent semantic labels that require precise definitions; - requires consolidation to distinguish duplicates from state transitions; - may retain more evidence records, increasing storage and review volume; - depends on adequate local context to determine resolution and lifecycle state; - requires explicit view policies so consumers do not treat every evidence record as protocol-worthy; - cannot guarantee LLM consistency merely by replacing a binary verdict with a taxonomy. The taxonomy should remain closed and small. In particular, modality, confidence, support, commitment state and responsibility must not be collapsed into a single scalar. ## Migration strategy This is a staged design path, not a recommendation to implement it now. 1. Specify the taxonomy and invariants against the existing BUG-015 Gold candidates and the cited Progeo excerpts, without changing current expected outputs. 2. Annotate a design-only mapping from each current Decision, Action Item and Open Question example to evidence state, support and independent attributes. Confirm objectively unique ground truth, especially for `requested_action` versus `possible_next_step` and `tentative_agreement` versus `preferred_option`. 3. Define a versioned evidence-record schema and deterministic protocol-selection policy on paper. Keep responsibility and question resolution independent. 4. Design evaluation measures separately for state accuracy, resolution accuracy, responsibility accuracy and protocol-policy output. A correct lower state must not count as a false protocol commitment. 5. Only after the design is validated, plan an experiment that compares the evidence-first approach with current extraction and the binary verifier. Any prompt or Python changes would be separate, explicitly scoped tasks following the Gold methodology and LLM safety rules. 6. If later adopted, preserve compatibility through an adapter that maps protocol-worthy evidence states to the current canonical categories while Canonical Meeting Knowledge evolves. Do not ask the renderer to interpret raw states. No database, prompt, code, test or current output-schema migration is proposed by this document. ## Open questions - Is `requested_action` useful as a distinct state when a direct assignment already creates `established_action`, or should it be limited to unaccepted requests? - Does `tentative_agreement` have sufficiently unique evidence in the current corpus, or should it remain an annotation until more Gold examples exist? - Should an ongoing-work item whose relevance to the meeting is unclear pass the Working Protocol policy, or require an explicit continuation/follow-up signal? - How much local context is required to label a question `resolved` safely when the answer occurs across a technical chunk boundary? - Should `resolution_unclear` be retained only in Canonical Meeting Knowledge, or also exposed to a review view? - Should support be limited to `direct / explicit`, treating `indirect` as review-only, to reduce subjective classification? - How should consolidation represent one subject moving from proposal to decision: linked immutable evidence records, or a current-state object with a preserved event history? - Which output views, if any, should include withdrawn proposals, cancelled work and resolved questions? - Can polarity and lifecycle links be normalized deterministically without introducing semantic inference? ## Recommendation The evidence-first architecture should replace the current binary keep/reject classification approach as the target architecture. The replacement should be conceptual and staged, not implemented yet: extract a small, domain-specific semantic state plus independent support, responsibility and resolution attributes; preserve evidence and provenance; then apply explicit policy to produce protocol categories. A single common Evidence Strength axis should complement this taxonomy but must not replace it. The decisive improvement is separation of semantic state from protocol eligibility. That separation directly explains the BUG-015 false positives and false negatives, supports the Progeo cases without invented ownership, simplifies view selection, and provides a sounder basis for Canonical Meeting Knowledge and future BPD-style rendering.