Files
meeting-lab/docs/design/evidence_commitment_model.md
T

27 KiB
Raw Blame History

Evidence / Commitment Model

Status: architecture design study for BUG-015; no implementation recommendation is made here.

Motivation

Meeting Lab currently asks extraction and the BUG-015 verifier to cross a category boundary in one step: a candidate is either a protocol-worthy Decision, Action Item or Open Question, or it is rejected. That combines two different questions:

  1. What does the transcript provide evidence for?
  2. Which evidenced states should a particular protocol view publish?

Meetings do not move directly from absence to commitment. An option can become a proposal, then a preferred option, then a tentative agreement and finally a decision. Work can be suggested, requested, assigned, accepted, already underway or completed. A question can be asked and answered, or can remain explicitly unresolved. Rejecting everything below the final protocol threshold discards useful evidence; promoting it creates false commitments.

The proposed model therefore records the evidenced semantic state first. A later policy selects protocol-worthy states. It preserves the existing separation between extraction, deterministic canonicalization, semantic consolidation and rendering, and it does not make rendered output the semantic source of truth.

Observed BUG-015 failures

The BUG-015 Gold cases expose both sides of the binary-classification problem.

Progeo-derived evidence Correct evidence state Binary failure to avoid
“Die Option steht im Raum, das Material chemisch recyceln zu lassen.” option False Decision
“Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden.” personal preference with a conditional option False Decision
“Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht. … das ist entschieden.” explicit decision: rejection False rejection
“Die Geometrie kann man vielleicht noch optimieren.” possible next step False Action Item
“Marleen müsste vielleicht mal äußern …” suggested/requested action; no accepted or confirmed responsibility False Action Item and false owner
“Wir könnten Dirk Textor vielleicht noch einmal kontaktieren.” suggestion False Action Item
“Nina, übernimmst du …? – Ja, ich übernehme …” accepted action with explicit responsibility and deadline False rejection
“Den CET-Artikel erstellen wir bereits …” ongoing established work; owner not evidenced Consistent false rejection by the binary verifier
“Ob sich das Waschen lohnt, weiß ich nicht.” uncertainty False Open Question
“Gibt es schon ein Programm? – Ja …” explicit, resolved question False Open Question if local resolution is ignored
“Welche Daten … dürfen wir veröffentlichen? … weiterhin ungeklärt.” explicit unresolved question Consistent false rejection by the binary verifier

The verifier was reliable on several negative cases and on explicit commitment language, but not on valid states whose evidence did not resemble a fresh agreement: already-established work and an explicitly unresolved information need. This suggests a representation problem, not merely an insufficient keep/reject prompt.

Model shape

The model is deliberately small and compositional. Each evidence item has:

  • a domain: decision, action or information need;
  • a semantic state within that domain;
  • an evidence support level describing how directly the transcript supports that label;
  • preserved transcript evidence and source references;
  • domain-specific independent attributes, such as responsibility or resolution;
  • no automatic claim that the item belongs in a final protocol.

Semantic maturity and evidence support are independent. An explicitly worded proposal remains a proposal; it is not a weak decision. Conversely, ongoing work can be strongly evidenced without a recorded moment of assignment.

Semantic state diagrams

The arrows show common progressions, not mandatory workflows. Meetings may enter at any state, skip states, regress, or end without commitment.

Decision evolution

idea -> option -> proposal -> preferred_option -> tentative_agreement -> decision
          |          |              |                    |              |
          +----------+--------------+--------------------+----------> withdrawn
                                                                    superseded
                                                                    reopened

An idea is an undeveloped possibility. An option is a candidate alternative. A proposal asks the group to adopt an outcome. A preferred option expresses comparative preference without settlement. A tentative agreement records provisional convergence that is explicitly conditional or awaiting confirmation. A decision records an outcome the meeting settled, selected, approved, rejected or committed to. The intermediate preferred and tentative states matter because they prevent likely direction from being confused with commitment.

Action evolution

possible_next_step -> recommendation -> requested_action -> established_action -> ongoing_work -> completed
        |                  |                 |                    |               |
        +------------------+-----------------+--------------------+----------> cancelled

Responsibility (independent):
unset -> proposed_responsible -> assigned -> accepted/confirmed

An action can enter directly as established_action through explicit assignment, acceptance or commitment. It can enter directly as ongoing_work when the transcript clearly says that the work is already being performed, as in the CET example. proposed_responsible is evidence about a suggestion, not permission to populate the protocol's responsible field. Only assigned, accepted or confirmed supports recorded responsibility, and only where the meeting evidence explicitly assigns, accepts or confirms it.

Information-need evolution

uncertainty ---------> information_need(unresolved) ---------> resolved
curiosity -----------> explicit_question(unresolved) --------> resolved
request_for_clarification(unresolved) -----------------------> resolved
missing_information(unresolved) -----------------------------> resolved

rhetorical_question -> discourse only

Uncertainty and curiosity do not automatically create an information need. An explicit question may be resolved immediately. “Open Question” is therefore a policy result derived primarily from a qualifying kind plus unresolved, rather than a primitive utterance category.

Proposed taxonomy

Common fields

Field Proposed values Purpose
domain decision, action, information_need Selects the domain taxonomy.
state Domain-specific values below Records what the evidence says now.
support indirect, direct, explicit Rates support for the chosen state, not protocol importance.
polarity positive, negative where applicable Preserves approval/rejection and adoption/refusal without rewriting meaning.
evidence transcript excerpt(s) Proves the state and attributes.
source_refs stable source references Retains provenance across stages.

support is intentionally not weak / medium / strong. Those terms mix confidence, semantic maturity and number of witnesses. The proposed meanings are narrower:

  • indirect: the state is supported by context but not stated in a self-contained utterance; downstream commitment policy should normally be conservative.
  • direct: an utterance directly expresses the state, such as a proposal, preference, ongoing-work statement or question.
  • explicit: the utterance also names the decisive status, for example “das ist entschieden”, “ich übernehme”, “wir erstellen bereits” or “weiterhin ungeklärt”.

This common axis simplifies evidence auditing and threshold policy, but cannot replace domain state. An explicit suggestion is still not an Action Item, and an explicit uncertainty is still not an Open Question. Model confidence, if retained at all, must be a separate operational field and must not be presented as evidence strength.

Decision states

State Meaning Default protocol treatment
idea Undeveloped possibility or brainstorming contribution Exclude
option Candidate alternative under consideration Exclude
proposal Outcome offered for adoption Exclude
preferred_option Expressed preference among alternatives Exclude
tentative_agreement Provisional convergence with an expressed condition or pending confirmation Exclude from Decisions; potentially expose to an editorial view
decision Settled, selected, approved, rejected or committed outcome Include as Decision when evidence is sufficient
reopened Earlier decision explicitly returned to unresolved consideration Do not render the earlier decision as currently settled without qualification
superseded Earlier decision replaced by a later one Retain provenance; normally render only the current decision
withdrawn Candidate state explicitly withdrawn Exclude as current commitment

idea, option, proposal, preferred_option, tentative_agreement and decision are necessary distinctions for BUG-015. reopened, superseded, withdrawn are lifecycle states needed to avoid treating historical evidence as current policy; they need not be first-iteration extraction targets.

Action states and attributes

State Meaning Default protocol treatment
possible_next_step Hypothetical or exploratory action Exclude
recommendation Action advocated but not established as work Exclude
requested_action Someone asks that work be done, without enough evidence that it is established Exclude by default
established_action Concrete future work established by assignment, acceptance, commitment or confirmation Include as Action Item
ongoing_work Concrete work explicitly already underway Include when still relevant; do not require a newly witnessed assignment
completed Work explicitly reported complete Exclude from open Action Items; retain as status/history
cancelled Work explicitly cancelled or declined Exclude from open Action Items; retain provenance

suggestion is represented as possible_next_step or recommendation, depending on whether the speaker advocates it. accepted_action and assigned_work should not be competing lifecycle states: both establish established_action, while the independent commitment basis records accepted, assigned, self_committed or confirmed_existing. This avoids an artificial choice when, as with Nina, an assignment and acceptance occur together.

Responsibility is independent:

Responsibility status Meaning May populate responsible?
unset No person/team explicitly tied to ownership No
proposed A person is suggested, asked speculatively or mentioned near the work No
assigned The meeting explicitly assigns the work Yes
accepted The party explicitly accepts or volunteers Yes
confirmed Existing ownership is explicitly confirmed Yes

This preserves the project invariant: discussion, expertise, adjacency, organizational role and likely ownership never establish responsibility. An action may be valid with responsibility_status: unset, as with established CET work.

Information-need kinds and resolution

Question form and resolution are separate attributes.

Kind Meaning Can become an Open Question?
uncertainty Speaker expresses doubt or lack of certainty without establishing a concrete need No, by itself
curiosity Interest without a concrete need requiring follow-up No, by itself
explicit_question Direct interrogative seeking an answer Yes, if unresolved
request_for_clarification Explicit request to clarify a concrete matter Yes, if unresolved
missing_information Concrete required information is stated as absent Yes, if unresolved
rhetorical_question Interrogative used for emphasis rather than an answer No

Resolution is one of unresolved, resolved, or resolution_unclear. resolved_question is therefore not a separate kind: it is, for example, explicit_question + resolved. The publication example is explicit_question + unresolved; the event-program example is explicit_question + resolved; the washing example is uncertainty unless later evidence establishes a concrete unresolved need. resolution_unclear preserves evidence but should not silently pass a precision-first Open Question policy.

An “unresolved question” is a derived, protocol-relevant combination rather than a fourth kind. This makes the classification testable: kind answers what communicative act occurred, and resolution answers what remained at the end of the available context.

Progeo examples under the model

Evidence Proposed representation Working Protocol policy result
Chemical recycling “Option steht im Raum” decision / option / explicit Not a Decision
Technikum preference decision / preferred_option / direct; personal scope preserved Not a Decision
Schlummer collaboration rejected and “entschieden” decision / decision / explicit / negative Decision
Geometry “kann man vielleicht … optimieren” action / possible_next_step / direct, responsibility unset Not an Action Item
Marleen “müsste vielleicht mal” action / requested_action / direct, responsibility proposed only Not an Action Item; no responsible party
Textor “könnten … kontaktieren” action / possible_next_step / direct, responsibility unset Not an Action Item
Nina asks and accepts the review by Friday action / established_action / explicit, basis assigned + accepted, responsible Nina, deadline Friday Action Item
CET article “erstellen wir bereits” and later circulation action / ongoing_work / explicit, responsibility unset Action Item without invented owner
Washing “ob sich das lohnt, weiß ich nicht” information_need / uncertainty / direct, resolution unclear Not an Open Question
Event program asked and answered information_need / explicit_question / direct, resolved Not an Open Question
Publishable energy-audit data “weiterhin ungeklärt” information_need / explicit_question / explicit, unresolved Open Question

These labels preserve every BUG-015 distinction without treating rejected protocol candidates as meaningless.

Thematic protocol as the primary structure

The Evidence / Commitment Model supplies semantic distinctions inside a larger meeting reconstruction. It does not imply that the primary protocol should be organized as one section per semantic category.

The central Version 2 requirement is:

The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.

A primary protocol should reconstruct which topics were discussed, what relevant information emerged within each topic, how the discussion developed, which alternatives, ideas, objections or proposals mattered, what outcome or current state was reached, and which actions or unresolved questions resulted from that topic.

Semantic categories remain essential, but as metadata and supporting structure attached to topics. Facts, technical findings, ideas, alternatives, proposals, objections, decisions, action items and open questions must not dictate the main document structure.

For example, a human-style topic section may read:

## Trial setup

Several variants for the next trial were discussed.

A thinner carrier material was proposed as one possible alternative.
Concerns were raised regarding its mechanical suitability.
The group therefore decided to continue with the existing setup for the next trial.

Nina will obtain the remaining samples before the next production run.

The publication question remains unresolved.

This preserves the relationship between the proposal, objection, decision, resulting work and unresolved question. A primary document split into separate Ideas, Objections, Decisions and Action Items sections would lose that thematic and conversational relationship.

Category-oriented outputs remain valuable as derived secondary views. An Action Item table, Decision register, Open Questions list or Management summary can be generated from the same underlying meeting knowledge after thematic reconstruction. These indexes and summaries do not replace the primary topic-oriented protocol.

The conceptual Version 2 flow is:

Transcript
    ->
Evidence Extraction
    ->
Topic Reconstruction
    ->
Semantic Synthesis
    ->
Protocol Rendering

Evidence items and their semantic metadata remain attached to topics and support Semantic Synthesis. The Version 2 renderer should receive topic-oriented semantic knowledge rather than a flat collection grouped by category. This direction does not define the final Topic Reconstruction schema, redesign the current pipeline or remove the existing Working Protocol V2 architecture. It is an accepted target architecture whose implementation is postponed.

Architectural impact

Extraction

Extraction could emit evidence records with richer labels instead of immediately claiming decision, action_item or open_question. This is a bounded increase in extraction vocabulary, not a request for larger context windows, a multi-chunk strategy or a monolithic synthesis prompt. Existing Facts, Positions and Technical Details remain distinct categories; the new domains refine only commitment-sensitive content.

The immediate output would describe the observed state. Protocol category selection would happen later through an explicit policy. Some policy rules can be deterministic once semantic labels are trustworthy, for example:

Decision       := domain=decision AND state=decision
Action Item    := domain=action AND state IN {established_action, ongoing_work}
Open Question  := domain=information_need
                  AND kind IN {explicit_question, request_for_clarification, missing_information}
                  AND resolution=unresolved
Responsible    := responsibility_status IN {assigned, accepted, confirmed}

This keeps downstream selection simple, but does not make semantic extraction deterministic. The LLM still has to distinguish proposals from decisions and uncertainty from unresolved needs.

LLM stability

The architecture should reduce one source of instability: the model no longer has to erase a supported proposition merely because it falls below a protocol threshold, nor relocate it into another category to retain it. Labels correspond more closely to observable speech acts and lifecycle statements, and downstream thresholds become explicit.

It will not eliminate instability. More labels create adjacent-class boundaries, local context may not reveal resolution, and indirect language remains difficult. Stability depends on small schemas, evidence spans, one best supported state, conservative handling of resolution_unclear, and later evaluation against the Gold corpus. A generic evidence-strength score alone would likely worsen instability because it invites subjective grading.

Canonicalizer

The Canonicalizer would benefit from carrying normalized state names, independent responsibility fields, resolution, polarity and stable source references. It could deterministically validate allowed combinations, normalize aliases, preserve evidence, group exact duplicates and reject structurally impossible combinations. It must not promote a proposal to a decision, infer resolution, infer responsibility or perform uncertain semantic merging.

Richer labels make safe comparisons easier: two option records can be recognized as candidates about the same subject without being merged into a decision; an explicit_question + resolved item is not confused with an unresolved instance. Provenance improves because a later decision can link back to earlier option/proposal evidence rather than overwriting it.

Semantic merging becomes better informed, but not automatically easy. Equivalence and lifecycle transitions remain semantic work. A future consolidator should preserve all source evidence, distinguish duplicate evidence from state evolution, and mark contradictions or uncertainty rather than collapsing them.

Renderer and output policy

The renderer should not make commitment classification decisions. A policy/view-selection stage should determine protocol-worthy semantic states while preserving their topic associations. In the Version 2 target architecture, the primary protocol renderer should receive topic-oriented semantic knowledge containing eligible evidence and state transitions, not a flat category-oriented collection. It must not silently promote raw proposals to Decisions or present conversational coordination as Meeting Content.

Proposals may still be visible to a renderer for a view explicitly designed to show them, such as a discussion appendix or editorial drafting view. That access must be typed and intentional; a renderer must never silently format a proposal under “Decisions.” Different output views may select different states while sharing the same Canonical Meeting Knowledge.

Category-oriented renderers remain valid for secondary views such as Action Item tables, Decision registers and Open Questions lists. This accepted direction supplements rather than removes the current Working Protocol V2 architecture; it does not change current rendering behavior.

Future BPD-style protocol generation

Richer labels would give editorial generation better material without granting it authority to invent commitment. A BPD-style view could explain the path from options through tentative agreement to decision, separate current obligations from completed work, and describe why an information need remains open. Proposals and preferred options could be useful appendix material when the requested view values discussion history.

The editorial layer may condense or order these records, but it must preserve state, polarity, responsibility status and provenance. Appendix inclusion is a view policy, not a semantic promotion.

Trade-offs

Benefits:

  • preserves valid evidence below the final-protocol threshold;
  • separates meeting semantics from publication policy;
  • represents established work without inventing an assignment event or owner;
  • treats question kind and resolution independently;
  • makes responsibility independently auditable;
  • improves provenance across lifecycle changes;
  • prevents renderers from silently converting proposals into commitments;
  • gives future editorial views useful, explicitly non-final material.

Costs and risks:

  • expands the extraction schema and Gold specification;
  • introduces adjacent semantic labels that require precise definitions;
  • requires consolidation to distinguish duplicates from state transitions;
  • may retain more evidence records, increasing storage and review volume;
  • depends on adequate local context to determine resolution and lifecycle state;
  • requires explicit view policies so consumers do not treat every evidence record as protocol-worthy;
  • cannot guarantee LLM consistency merely by replacing a binary verdict with a taxonomy.

The taxonomy should remain closed and small. In particular, modality, confidence, support, commitment state and responsibility must not be collapsed into a single scalar.

Migration strategy

This is a staged design path, not a recommendation to implement it now.

  1. Specify the taxonomy and invariants against the existing BUG-015 Gold candidates and the cited Progeo excerpts, without changing current expected outputs.
  2. Annotate a design-only mapping from each current Decision, Action Item and Open Question example to evidence state, support and independent attributes. Confirm objectively unique ground truth, especially for requested_action versus possible_next_step and tentative_agreement versus preferred_option.
  3. Define a versioned evidence-record schema and deterministic protocol-selection policy on paper. Keep responsibility and question resolution independent.
  4. Design evaluation measures separately for state accuracy, resolution accuracy, responsibility accuracy and protocol-policy output. A correct lower state must not count as a false protocol commitment.
  5. Only after the design is validated, plan an experiment that compares the evidence-first approach with current extraction and the binary verifier. Any prompt or Python changes would be separate, explicitly scoped tasks following the Gold methodology and LLM safety rules.
  6. If later adopted, preserve compatibility through an adapter that maps protocol-worthy evidence states to the current canonical categories while Canonical Meeting Knowledge evolves. Do not ask the renderer to interpret raw states.

No database, prompt, code, test or current output-schema migration is proposed by this document.

Open questions

  • Is requested_action useful as a distinct state when a direct assignment already creates established_action, or should it be limited to unaccepted requests?
  • Does tentative_agreement have sufficiently unique evidence in the current corpus, or should it remain an annotation until more Gold examples exist?
  • Should an ongoing-work item whose relevance to the meeting is unclear pass the Working Protocol policy, or require an explicit continuation/follow-up signal?
  • How much local context is required to label a question resolved safely when the answer occurs across a technical chunk boundary?
  • Should resolution_unclear be retained only in Canonical Meeting Knowledge, or also exposed to a review view?
  • Should support be limited to direct / explicit, treating indirect as review-only, to reduce subjective classification?
  • How should consolidation represent one subject moving from proposal to decision: linked immutable evidence records, or a current-state object with a preserved event history?
  • Which output views, if any, should include withdrawn proposals, cancelled work and resolved questions?
  • Can polarity and lifecycle links be normalized deterministically without introducing semantic inference?

Recommendation

The evidence-first architecture should replace the current binary keep/reject classification approach as the target architecture. The replacement should be conceptual and staged, not implemented yet: extract a small, domain-specific semantic state plus independent support, responsibility and resolution attributes; preserve evidence and provenance; then apply explicit policy to produce protocol categories.

A single common Evidence Strength axis should complement this taxonomy but must not replace it. The decisive improvement is separation of semantic state from protocol eligibility. That separation directly explains the BUG-015 false positives and false negatives, supports the Progeo cases without invented ownership, simplifies view selection, and provides a sounder basis for Canonical Meeting Knowledge and future BPD-style rendering.