From 63075eaca91088beb92e646b6b0a2539cfe049ca Mon Sep 17 00:00:00 2001 From: Martin Tazl Date: Fri, 31 Jul 2026 13:01:11 +0200 Subject: [PATCH] Enforce explicit responsibility attribution - document responsibility attribution as a project-wide invariant - prevent inferred ownership in protocol rendering - add negative gold regression for false responsibility assignment - document future responsibility evidence and attribution model - record the real-life benchmark finding --- AGENTS.md | 18 +++ CHANGELOG.md | 4 + PROJECT_KNOWLEDGE.md | 17 +++ ROADMAP.md | 2 + docs/architecture.md | 14 ++ docs/data-models.md | 21 +++ docs/experiments.md | 62 ++++++++ docs/output-views.md | 7 + docs/pipeline.md | 16 +++ prompts/working_protocol.md | 133 ++++++++++++++++++ .../README.md | 15 ++ .../expected.json | 39 +++++ .../transcript.txt | 11 ++ 13 files changed, 359 insertions(+) create mode 100644 prompts/working_protocol.md create mode 100644 tests/gold/responsibility_attribution_negative/README.md create mode 100644 tests/gold/responsibility_attribution_negative/expected.json create mode 100644 tests/gold/responsibility_attribution_negative/transcript.txt diff --git a/AGENTS.md b/AGENTS.md index d2b1831..901e3c3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -70,6 +70,24 @@ Status: of the source transcript or consolidated meeting knowledge unless an explicit output language is requested. +## Responsibility Attribution Invariant + +A person, team or department may be recorded as responsible only when the +meeting evidence explicitly assigns, accepts or confirms that responsibility. + +Discussion, expertise, objection, suggestion, thematic proximity, speaker +adjacency, organizational assumptions, likely job roles or mere mention do not +establish ownership. + +When support is incomplete or ambiguous, leave the responsible person unset, +mark the item as unclear where the current schema supports it, and preserve the +supporting evidence. Never guess. + +This invariant applies to extraction, canonicalization, semantic consolidation, +Canonical Meeting Knowledge, Working Protocol / Arbeitsprotokoll, Distribution +Protocol / Verteilerprotokoll and Knowledge Objects / +Wissensdatenbankeintrag. + ## Prompt Engineering Rules The Gold Standard corpus is the reference specification. Follow Rules 1-11 from diff --git a/CHANGELOG.md b/CHANGELOG.md index fb90995..47d3127 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -27,6 +27,7 @@ - Consolidation prompt and non-LLM tests for payload construction, grouping validation and source fact coverage. - Semantic Consolidator V0 benchmark report and consolidated extraction JSON. +- Gold regression scenario for responsibility attribution integrity. ### Changed @@ -42,6 +43,9 @@ - Documented that rendered protocol language should normally match the source transcript or consolidated meeting knowledge unless explicitly requested otherwise. +- Added the Responsibility Attribution Invariant across extraction, + canonicalization, consolidation, Canonical Meeting Knowledge and output + views. - Updated decision extraction semantics to include explicit process decisions and deferrals. - Simplified extraction prompt assembly around prompt files. diff --git a/PROJECT_KNOWLEDGE.md b/PROJECT_KNOWLEDGE.md index c3a3f07..d52314f 100644 --- a/PROJECT_KNOWLEDGE.md +++ b/PROJECT_KNOWLEDGE.md @@ -92,6 +92,10 @@ validated work may specify another model explicitly. - Gold Standard tests are also a formal specification of meeting semantics. - Raw model responses should be preserved when diagnosing parser or truncation failures. +- Responsibility attribution is a critical correctness invariant: people, + teams and departments must not be assigned ownership from discussion, + expertise, objection, thematic proximity, speaker adjacency, role guesses or + world knowledge. Assignment requires explicit evidence. ## Decision Taxonomy @@ -108,6 +112,19 @@ Accepted decision semantics: - "No decision was reached" is different from "the group decided to defer the decision." +## Responsibility Attribution + +Meeting Lab distinguishes mentioned people, speakers, participants, +responsible people, departments, owners and assignees. These concepts must not +be collapsed into one field. + +A `responsible` or future `owner` / `assignee` value may be recorded only when +source evidence explicitly assigns, accepts or confirms responsibility. If the +evidence is incomplete or ambiguous, the responsible person remains `null` or +unset and the evidence is preserved. Future schema work may add +`responsibility_status` values such as `explicit`, `accepted`, `proposed` and +`unclear`, plus `attribution_evidence`. + Current Prompt Version 2 decision baseline: - `decision_simple`: passing. diff --git a/ROADMAP.md b/ROADMAP.md index 7b51c6c..7b54eb1 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -12,6 +12,8 @@ Deliverables: - Stronger Gold Standard coverage across facts, positions, decisions, todos, questions and technical details. +- Gold coverage for responsibility attribution: discussion, objection, role + proximity and department mention must not become ownership. - Improved category prompts. - Repeatable evaluation workflow. - Documented prompt experiment log. diff --git a/docs/architecture.md b/docs/architecture.md index d5e3148..8385259 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -47,6 +47,20 @@ Every processing step should be understandable. Intermediate results should remain inspectable throughout the pipeline. +## Responsibility Attribution Integrity + +Responsibility, ownership, organizational roles and action-item assignments may +be recorded only when meeting evidence explicitly assigns, accepts or confirms +them. + +The system must not infer responsibility from thematic proximity, +participation in a discussion, mentioning a task, commenting on another +department, organizational assumptions, likely job roles, speaker adjacency or +model world knowledge. + +When evidence is incomplete or ambiguous, the responsible person remains unset +or unclear and the supporting evidence is preserved. + ## Reproducible Experiments Experiments must be repeatable. diff --git a/docs/data-models.md b/docs/data-models.md index 9781be4..36db6a2 100644 --- a/docs/data-models.md +++ b/docs/data-models.md @@ -272,6 +272,27 @@ Each item contains at least: Action items also preserve deterministic fields such as `responsible` and `deadline` when present. +Responsibility attribution is stricter than mention or participation. A +`responsible` value may be kept only when source evidence explicitly assigns, +accepts or confirms responsibility. If evidence is incomplete, ambiguous or +only based on a suggestion, objection, topic expertise or department mention, +the field remains `null` or unset and the evidence is preserved. + +Do not collapse these concepts into one field: + +- `speaker`: person who uttered the evidence. +- `mentioned_person`: person named in the evidence. +- `participant`: person present in the meeting. +- `responsible_person`: person explicitly assigned to or accepting an action. +- `department`: organizational unit discussed or represented. +- `owner`: durable ownership of a process, system or knowledge object. +- `assignee`: operational person or team assigned to a concrete task. + +Future compatible fields may include: + +- `responsibility_status`: `explicit`, `accepted`, `proposed` or `unclear`. +- `attribution_evidence`: source evidence supporting the assignment status. + Example: ```json diff --git a/docs/experiments.md b/docs/experiments.md index bdb82f0..84b59ee 100644 --- a/docs/experiments.md +++ b/docs/experiments.md @@ -1169,3 +1169,65 @@ Evidence: - `samples/benchmarks/semantic_consolidator_v0/report.md` - `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json` - Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt` + +## EXP-0023 - Responsibility attribution integrity + +Status: Running + +Date or period: 2026-07-31 + +Hypothesis: + +Protocol generation is operationally unsafe if responsibility, ownership or +departmental role attribution is inferred from discussion context rather than +explicit meeting evidence. + +Setup: + +The real-life Working Protocol Renderer V2 benchmark was inspected against the +consolidated input. A false assignment connected a Marketing participant to +Business Development criteria work even though the participant's contribution +was critical or reluctant and did not establish acceptance of that task. + +Inputs: + +- `samples/benchmarks/working_protocol_renderer_v2/working_protocol.md` +- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json` +- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json` + +Model / configuration: + +- Not rerun for this finding. +- Finding is based on existing benchmark artifacts. + +Result: + +The false responsibility attribution is visible in the structured input before +rendering, so the issue is not merely stylistic renderer wording. The root +cause may originate earlier in extraction and then be preserved by +canonicalization and consolidation. Renderer guardrails are still required so +output views do not strengthen ambiguous ownership. + +Decision: + +Responsibility attribution is now treated as a critical project-wide +invariant. A person, team or department may be recorded as responsible only +when the evidence explicitly assigns, accepts or confirms that responsibility. +Discussion, expertise, objection, suggestion, thematic proximity, speaker +adjacency, organizational assumptions and likely job roles do not establish +ownership. + +Lessons learned: + +This class of error affects operational correctness, not only style. The +pipeline needs traceable attribution evidence and future schema support for +responsibility status such as explicit, accepted, proposed or unclear. + +Evidence: + +- `AGENTS.md` +- `PROJECT_KNOWLEDGE.md` +- `docs/data-models.md` +- `docs/output-views.md` +- `prompts/working_protocol.md` +- `tests/gold/responsibility_attribution_negative/` diff --git a/docs/output-views.md b/docs/output-views.md index beb7597..4db3898 100644 --- a/docs/output-views.md +++ b/docs/output-views.md @@ -78,6 +78,13 @@ Rendered protocol language should normally match the dominant language of the source transcript or consolidated meeting knowledge unless an explicit output language is requested. +Responsibility attribution is a cross-cutting invariant for every Output View. +A renderer may name a person, team or department as responsible only when the +input explicitly contains that assignment or acceptance. Discussion, +expertise, objection, suggestion, speaker adjacency, organizational +assumptions and department mentions do not establish ownership. If +responsibility is unclear, render it as open rather than guessing. + ## Working Protocol Suggested filename: `working_protocol.md` diff --git a/docs/pipeline.md b/docs/pipeline.md index d80be00..fe2d700 100644 --- a/docs/pipeline.md +++ b/docs/pipeline.md @@ -10,6 +10,22 @@ The guiding principle is simple: > **Each processing stage has exactly one responsibility.** +Cross-cutting invariant: + +> **Responsibility attribution requires explicit evidence.** + +A person, team or department may be recorded as responsible only when the +source material explicitly assigns, accepts or confirms that responsibility. +The pipeline must not infer ownership from thematic proximity, discussion +participation, mentioning a task, commenting on another department, +organizational assumptions, likely job roles, speaker adjacency or model world +knowledge. + +When support is incomplete or ambiguous, leave the responsible person unset, +mark the item as unclear where supported, and preserve the attribution +evidence. This applies to extraction, canonicalization, semantic consolidation, +Canonical Meeting Knowledge and every Output View renderer. + --- # Pipeline Overview diff --git a/prompts/working_protocol.md b/prompts/working_protocol.md new file mode 100644 index 0000000..5f3edd7 --- /dev/null +++ b/prompts/working_protocol.md @@ -0,0 +1,133 @@ +You are the Working Protocol renderer for Meeting Lab. + +Render the provided Canonical Meeting Knowledge, or the current consolidated +meeting representation, as one professional internal working protocol in +Markdown. + +Your task is presentation only. + +Use only information contained in the provided input. Do not use outside +knowledge. Do not infer missing facts. Do not merge semantically different +items. Do not resolve conflicts. Do not add conclusions that are not present in +the input. + +Output language: + +- Use the dominant language of the consolidated meeting knowledge. +- Do not translate. +- Change language only if an explicit output language is requested. + +Renderer responsibilities: + +- Organize information for human readers. +- Remove presentation-level redundancy. +- Preserve all decisions. +- Preserve all action items. +- Preserve all open questions. +- Preserve evidence-relevant context when needed to understand a decision, + action item or open question. + +The renderer is not responsible for: + +- semantic consolidation +- deduplication of meaning +- fact extraction +- conflict resolution +- inventing topics that are not supported by the supplied structure +- protocol-independent knowledge synthesis + +Target document: + +Write a professional internal working protocol for meeting participants who +need to reconstruct what was discussed, what was decided, what remains open and +who has to do what. + +This is not a management summary, a narrative meeting report, verbatim minutes +or marketing text. + +Writing style: + +- precise +- factual +- neutral +- concise +- technically accurate + +Avoid narrative framing and generic report phrases, including: + +- "The discussion focused on..." +- "It was noted that..." +- "Participants discussed..." +- "The meeting emphasized..." + +Do not restate decisions inside background paragraphs. +Do not explain action items multiple times. +Do not repeat the same information merely because it appears in multiple +categories. + +Priorities: + +1. Decisions +2. Action Items +3. Open Questions +4. Required background + +Condense everything else. + +Structure: + +Start immediately with: + +# Working Protocol + +Then organize by topic. For each topic, use only the sections that contain +information: + +## Topic title + +### Background + +Use at most one or two short paragraphs. Include only context required to +understand the decisions, action items or open questions. + +### Decisions + +Preserve every decision. Do not weaken or strengthen the decision. Do not +attribute it to a person unless the input explicitly provides that attribution. + +### Action Items + +Preserve every action item. Include responsible person and deadline only when +explicitly present. If either is missing, do not invent it. + +Never infer or strengthen responsibility. Only render an action item with a +named responsible person if the consolidated input explicitly contains that +assignment. If the input contains uncertainty, preserve it. If an item has no +explicit responsible person, render it without one or mark it as +"Verantwortung offen". Do not convert opinions into assignments, suggestions +into commitments, objections into ownership, or departmental discussion into +role attribution. + +### Open Questions + +Preserve every open question. Do not convert open questions into decisions or +action items. + +Do not create empty sections. + +Faithfulness rules: + +- Never invent facts. +- Never strengthen uncertainty. +- Never weaken decisions. +- Never attribute statements to people unless explicitly present. +- Keep distinctions between background, decisions, action items and open + questions. +- If information appears contradictory in the input, preserve the uncertainty + without resolving it. +- If two items are similar but not already consolidated, keep their meanings + distinct in the rendered document. + +Return Markdown only. +Do not include an introductory explanation. +Do not include a concluding summary. diff --git a/tests/gold/responsibility_attribution_negative/README.md b/tests/gold/responsibility_attribution_negative/README.md new file mode 100644 index 0000000..13bcce7 --- /dev/null +++ b/tests/gold/responsibility_attribution_negative/README.md @@ -0,0 +1,15 @@ +# responsibility_attribution_negative + +Tests that thematic proximity, department discussion and objection do not create +an action-item owner. + +The scenario includes a Marketing participant who comments critically on +Business Development criteria. No one assigns that participant responsibility +for defining the criteria, and the participant does not accept such a task. + +Expected behavior: + +- Do not assign Business Development criteria to the Marketing participant. +- Do not reclassify the Marketing participant as Business Development. +- Preserve the objection as a position. +- Keep responsibility open unless explicitly assigned. diff --git a/tests/gold/responsibility_attribution_negative/expected.json b/tests/gold/responsibility_attribution_negative/expected.json new file mode 100644 index 0000000..2444d98 --- /dev/null +++ b/tests/gold/responsibility_attribution_negative/expected.json @@ -0,0 +1,39 @@ +{ + "facts": [ + { + "fact": "Criteria are still needed for Business Development projects.", + "evidence": "Mara: For the new partner intake, we still need criteria for Business Development projects." + }, + { + "fact": "Noah is speaking from Marketing.", + "evidence": "Noah: From Marketing, I can tell you that generic criteria will not work." + }, + { + "fact": "Noah is not taking responsibility for defining the Business Development criteria.", + "evidence": "Noah: I am not taking that on." + } + ], + "decisions": [ + { + "decision": "Responsibility for defining the Business Development criteria remains open until Business Development confirms an owner.", + "evidence": "Mara: Agreed. Responsibility stays open for now, and we will ask Business Development for an owner." + } + ], + "todos": [ + { + "task": "Ask Business Development for an owner for defining the Business Development criteria.", + "responsible": null, + "deadline": null, + "evidence": "Mara: Agreed. Responsibility stays open for now, and we will ask Business Development for an owner." + } + ], + "questions": [], + "positions": [ + { + "speaker": "Noah", + "position": "Noah says generic criteria will not work for different Marketing and product contexts.", + "evidence": "Noah: From Marketing, I can tell you that generic criteria will not work. A digital campaign and a physical product launch need different checks." + } + ], + "technical": [] +} diff --git a/tests/gold/responsibility_attribution_negative/transcript.txt b/tests/gold/responsibility_attribution_negative/transcript.txt new file mode 100644 index 0000000..4270f5a --- /dev/null +++ b/tests/gold/responsibility_attribution_negative/transcript.txt @@ -0,0 +1,11 @@ +Mara: For the new partner intake, we still need criteria for Business Development projects. + +Noah: From Marketing, I can tell you that generic criteria will not work. A digital campaign and a physical product launch need different checks. + +Mara: Right, so the Business Development criteria remain open. + +Noah: I am not taking that on. I only want to flag that the criteria cannot be generic. + +Lea: Then we should keep the owner open until Business Development confirms who will define them. + +Mara: Agreed. Responsibility stays open for now, and we will ask Business Development for an owner.