Enforce explicit responsibility attribution
- document responsibility attribution as a project-wide invariant - prevent inferred ownership in protocol rendering - add negative gold regression for false responsibility assignment - document future responsibility evidence and attribution model - record the real-life benchmark finding
This commit is contained in:
@@ -70,6 +70,24 @@ Status:
|
||||
of the source transcript or consolidated meeting knowledge unless an explicit
|
||||
output language is requested.
|
||||
|
||||
## Responsibility Attribution Invariant
|
||||
|
||||
A person, team or department may be recorded as responsible only when the
|
||||
meeting evidence explicitly assigns, accepts or confirms that responsibility.
|
||||
|
||||
Discussion, expertise, objection, suggestion, thematic proximity, speaker
|
||||
adjacency, organizational assumptions, likely job roles or mere mention do not
|
||||
establish ownership.
|
||||
|
||||
When support is incomplete or ambiguous, leave the responsible person unset,
|
||||
mark the item as unclear where the current schema supports it, and preserve the
|
||||
supporting evidence. Never guess.
|
||||
|
||||
This invariant applies to extraction, canonicalization, semantic consolidation,
|
||||
Canonical Meeting Knowledge, Working Protocol / Arbeitsprotokoll, Distribution
|
||||
Protocol / Verteilerprotokoll and Knowledge Objects /
|
||||
Wissensdatenbankeintrag.
|
||||
|
||||
## Prompt Engineering Rules
|
||||
|
||||
The Gold Standard corpus is the reference specification. Follow Rules 1-11 from
|
||||
|
||||
@@ -27,6 +27,7 @@
|
||||
- Consolidation prompt and non-LLM tests for payload construction, grouping
|
||||
validation and source fact coverage.
|
||||
- Semantic Consolidator V0 benchmark report and consolidated extraction JSON.
|
||||
- Gold regression scenario for responsibility attribution integrity.
|
||||
|
||||
### Changed
|
||||
|
||||
@@ -42,6 +43,9 @@
|
||||
- Documented that rendered protocol language should normally match the source
|
||||
transcript or consolidated meeting knowledge unless explicitly requested
|
||||
otherwise.
|
||||
- Added the Responsibility Attribution Invariant across extraction,
|
||||
canonicalization, consolidation, Canonical Meeting Knowledge and output
|
||||
views.
|
||||
- Updated decision extraction semantics to include explicit process decisions
|
||||
and deferrals.
|
||||
- Simplified extraction prompt assembly around prompt files.
|
||||
|
||||
@@ -92,6 +92,10 @@ validated work may specify another model explicitly.
|
||||
- Gold Standard tests are also a formal specification of meeting semantics.
|
||||
- Raw model responses should be preserved when diagnosing parser or truncation
|
||||
failures.
|
||||
- Responsibility attribution is a critical correctness invariant: people,
|
||||
teams and departments must not be assigned ownership from discussion,
|
||||
expertise, objection, thematic proximity, speaker adjacency, role guesses or
|
||||
world knowledge. Assignment requires explicit evidence.
|
||||
|
||||
## Decision Taxonomy
|
||||
|
||||
@@ -108,6 +112,19 @@ Accepted decision semantics:
|
||||
- "No decision was reached" is different from "the group decided to defer the
|
||||
decision."
|
||||
|
||||
## Responsibility Attribution
|
||||
|
||||
Meeting Lab distinguishes mentioned people, speakers, participants,
|
||||
responsible people, departments, owners and assignees. These concepts must not
|
||||
be collapsed into one field.
|
||||
|
||||
A `responsible` or future `owner` / `assignee` value may be recorded only when
|
||||
source evidence explicitly assigns, accepts or confirms responsibility. If the
|
||||
evidence is incomplete or ambiguous, the responsible person remains `null` or
|
||||
unset and the evidence is preserved. Future schema work may add
|
||||
`responsibility_status` values such as `explicit`, `accepted`, `proposed` and
|
||||
`unclear`, plus `attribution_evidence`.
|
||||
|
||||
Current Prompt Version 2 decision baseline:
|
||||
|
||||
- `decision_simple`: passing.
|
||||
|
||||
@@ -12,6 +12,8 @@ Deliverables:
|
||||
|
||||
- Stronger Gold Standard coverage across facts, positions, decisions, todos,
|
||||
questions and technical details.
|
||||
- Gold coverage for responsibility attribution: discussion, objection, role
|
||||
proximity and department mention must not become ownership.
|
||||
- Improved category prompts.
|
||||
- Repeatable evaluation workflow.
|
||||
- Documented prompt experiment log.
|
||||
|
||||
@@ -47,6 +47,20 @@ Every processing step should be understandable.
|
||||
|
||||
Intermediate results should remain inspectable throughout the pipeline.
|
||||
|
||||
## Responsibility Attribution Integrity
|
||||
|
||||
Responsibility, ownership, organizational roles and action-item assignments may
|
||||
be recorded only when meeting evidence explicitly assigns, accepts or confirms
|
||||
them.
|
||||
|
||||
The system must not infer responsibility from thematic proximity,
|
||||
participation in a discussion, mentioning a task, commenting on another
|
||||
department, organizational assumptions, likely job roles, speaker adjacency or
|
||||
model world knowledge.
|
||||
|
||||
When evidence is incomplete or ambiguous, the responsible person remains unset
|
||||
or unclear and the supporting evidence is preserved.
|
||||
|
||||
## Reproducible Experiments
|
||||
|
||||
Experiments must be repeatable.
|
||||
|
||||
@@ -272,6 +272,27 @@ Each item contains at least:
|
||||
Action items also preserve deterministic fields such as `responsible` and
|
||||
`deadline` when present.
|
||||
|
||||
Responsibility attribution is stricter than mention or participation. A
|
||||
`responsible` value may be kept only when source evidence explicitly assigns,
|
||||
accepts or confirms responsibility. If evidence is incomplete, ambiguous or
|
||||
only based on a suggestion, objection, topic expertise or department mention,
|
||||
the field remains `null` or unset and the evidence is preserved.
|
||||
|
||||
Do not collapse these concepts into one field:
|
||||
|
||||
- `speaker`: person who uttered the evidence.
|
||||
- `mentioned_person`: person named in the evidence.
|
||||
- `participant`: person present in the meeting.
|
||||
- `responsible_person`: person explicitly assigned to or accepting an action.
|
||||
- `department`: organizational unit discussed or represented.
|
||||
- `owner`: durable ownership of a process, system or knowledge object.
|
||||
- `assignee`: operational person or team assigned to a concrete task.
|
||||
|
||||
Future compatible fields may include:
|
||||
|
||||
- `responsibility_status`: `explicit`, `accepted`, `proposed` or `unclear`.
|
||||
- `attribution_evidence`: source evidence supporting the assignment status.
|
||||
|
||||
Example:
|
||||
|
||||
```json
|
||||
|
||||
@@ -1169,3 +1169,65 @@ Evidence:
|
||||
- `samples/benchmarks/semantic_consolidator_v0/report.md`
|
||||
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
|
||||
- Local diagnostic only: `samples/benchmarks/semantic_consolidator_v0/raw_model_response.txt`
|
||||
|
||||
## EXP-0023 - Responsibility attribution integrity
|
||||
|
||||
Status: Running
|
||||
|
||||
Date or period: 2026-07-31
|
||||
|
||||
Hypothesis:
|
||||
|
||||
Protocol generation is operationally unsafe if responsibility, ownership or
|
||||
departmental role attribution is inferred from discussion context rather than
|
||||
explicit meeting evidence.
|
||||
|
||||
Setup:
|
||||
|
||||
The real-life Working Protocol Renderer V2 benchmark was inspected against the
|
||||
consolidated input. A false assignment connected a Marketing participant to
|
||||
Business Development criteria work even though the participant's contribution
|
||||
was critical or reluctant and did not establish acceptance of that task.
|
||||
|
||||
Inputs:
|
||||
|
||||
- `samples/benchmarks/working_protocol_renderer_v2/working_protocol.md`
|
||||
- `samples/benchmarks/semantic_consolidator_v0/consolidated_extractions.json`
|
||||
- `samples/benchmarks/canonicalizer_v1/canonicalized_extractions.json`
|
||||
|
||||
Model / configuration:
|
||||
|
||||
- Not rerun for this finding.
|
||||
- Finding is based on existing benchmark artifacts.
|
||||
|
||||
Result:
|
||||
|
||||
The false responsibility attribution is visible in the structured input before
|
||||
rendering, so the issue is not merely stylistic renderer wording. The root
|
||||
cause may originate earlier in extraction and then be preserved by
|
||||
canonicalization and consolidation. Renderer guardrails are still required so
|
||||
output views do not strengthen ambiguous ownership.
|
||||
|
||||
Decision:
|
||||
|
||||
Responsibility attribution is now treated as a critical project-wide
|
||||
invariant. A person, team or department may be recorded as responsible only
|
||||
when the evidence explicitly assigns, accepts or confirms that responsibility.
|
||||
Discussion, expertise, objection, suggestion, thematic proximity, speaker
|
||||
adjacency, organizational assumptions and likely job roles do not establish
|
||||
ownership.
|
||||
|
||||
Lessons learned:
|
||||
|
||||
This class of error affects operational correctness, not only style. The
|
||||
pipeline needs traceable attribution evidence and future schema support for
|
||||
responsibility status such as explicit, accepted, proposed or unclear.
|
||||
|
||||
Evidence:
|
||||
|
||||
- `AGENTS.md`
|
||||
- `PROJECT_KNOWLEDGE.md`
|
||||
- `docs/data-models.md`
|
||||
- `docs/output-views.md`
|
||||
- `prompts/working_protocol.md`
|
||||
- `tests/gold/responsibility_attribution_negative/`
|
||||
|
||||
@@ -78,6 +78,13 @@ Rendered protocol language should normally match the dominant language of the
|
||||
source transcript or consolidated meeting knowledge unless an explicit output
|
||||
language is requested.
|
||||
|
||||
Responsibility attribution is a cross-cutting invariant for every Output View.
|
||||
A renderer may name a person, team or department as responsible only when the
|
||||
input explicitly contains that assignment or acceptance. Discussion,
|
||||
expertise, objection, suggestion, speaker adjacency, organizational
|
||||
assumptions and department mentions do not establish ownership. If
|
||||
responsibility is unclear, render it as open rather than guessing.
|
||||
|
||||
## Working Protocol
|
||||
|
||||
Suggested filename: `working_protocol.md`
|
||||
|
||||
@@ -10,6 +10,22 @@ The guiding principle is simple:
|
||||
|
||||
> **Each processing stage has exactly one responsibility.**
|
||||
|
||||
Cross-cutting invariant:
|
||||
|
||||
> **Responsibility attribution requires explicit evidence.**
|
||||
|
||||
A person, team or department may be recorded as responsible only when the
|
||||
source material explicitly assigns, accepts or confirms that responsibility.
|
||||
The pipeline must not infer ownership from thematic proximity, discussion
|
||||
participation, mentioning a task, commenting on another department,
|
||||
organizational assumptions, likely job roles, speaker adjacency or model world
|
||||
knowledge.
|
||||
|
||||
When support is incomplete or ambiguous, leave the responsible person unset,
|
||||
mark the item as unclear where supported, and preserve the attribution
|
||||
evidence. This applies to extraction, canonicalization, semantic consolidation,
|
||||
Canonical Meeting Knowledge and every Output View renderer.
|
||||
|
||||
---
|
||||
|
||||
# Pipeline Overview
|
||||
|
||||
@@ -0,0 +1,133 @@
|
||||
You are the Working Protocol renderer for Meeting Lab.
|
||||
|
||||
Render the provided Canonical Meeting Knowledge, or the current consolidated
|
||||
meeting representation, as one professional internal working protocol in
|
||||
Markdown.
|
||||
|
||||
Your task is presentation only.
|
||||
|
||||
Use only information contained in the provided input. Do not use outside
|
||||
knowledge. Do not infer missing facts. Do not merge semantically different
|
||||
items. Do not resolve conflicts. Do not add conclusions that are not present in
|
||||
the input.
|
||||
|
||||
Output language:
|
||||
|
||||
- Use the dominant language of the consolidated meeting knowledge.
|
||||
- Do not translate.
|
||||
- Change language only if an explicit output language is requested.
|
||||
|
||||
Renderer responsibilities:
|
||||
|
||||
- Organize information for human readers.
|
||||
- Remove presentation-level redundancy.
|
||||
- Preserve all decisions.
|
||||
- Preserve all action items.
|
||||
- Preserve all open questions.
|
||||
- Preserve evidence-relevant context when needed to understand a decision,
|
||||
action item or open question.
|
||||
|
||||
The renderer is not responsible for:
|
||||
|
||||
- semantic consolidation
|
||||
- deduplication of meaning
|
||||
- fact extraction
|
||||
- conflict resolution
|
||||
- inventing topics that are not supported by the supplied structure
|
||||
- protocol-independent knowledge synthesis
|
||||
|
||||
Target document:
|
||||
|
||||
Write a professional internal working protocol for meeting participants who
|
||||
need to reconstruct what was discussed, what was decided, what remains open and
|
||||
who has to do what.
|
||||
|
||||
This is not a management summary, a narrative meeting report, verbatim minutes
|
||||
or marketing text.
|
||||
|
||||
Writing style:
|
||||
|
||||
- precise
|
||||
- factual
|
||||
- neutral
|
||||
- concise
|
||||
- technically accurate
|
||||
|
||||
Avoid narrative framing and generic report phrases, including:
|
||||
|
||||
- "The discussion focused on..."
|
||||
- "It was noted that..."
|
||||
- "Participants discussed..."
|
||||
- "The meeting emphasized..."
|
||||
|
||||
Do not restate decisions inside background paragraphs.
|
||||
Do not explain action items multiple times.
|
||||
Do not repeat the same information merely because it appears in multiple
|
||||
categories.
|
||||
|
||||
Priorities:
|
||||
|
||||
1. Decisions
|
||||
2. Action Items
|
||||
3. Open Questions
|
||||
4. Required background
|
||||
|
||||
Condense everything else.
|
||||
|
||||
Structure:
|
||||
|
||||
Start immediately with:
|
||||
|
||||
# Working Protocol
|
||||
|
||||
Then organize by topic. For each topic, use only the sections that contain
|
||||
information:
|
||||
|
||||
## Topic title
|
||||
|
||||
### Background
|
||||
|
||||
Use at most one or two short paragraphs. Include only context required to
|
||||
understand the decisions, action items or open questions.
|
||||
|
||||
### Decisions
|
||||
|
||||
Preserve every decision. Do not weaken or strengthen the decision. Do not
|
||||
attribute it to a person unless the input explicitly provides that attribution.
|
||||
|
||||
### Action Items
|
||||
|
||||
Preserve every action item. Include responsible person and deadline only when
|
||||
explicitly present. If either is missing, do not invent it.
|
||||
|
||||
Never infer or strengthen responsibility. Only render an action item with a
|
||||
named responsible person if the consolidated input explicitly contains that
|
||||
assignment. If the input contains uncertainty, preserve it. If an item has no
|
||||
explicit responsible person, render it without one or mark it as
|
||||
"Verantwortung offen". Do not convert opinions into assignments, suggestions
|
||||
into commitments, objections into ownership, or departmental discussion into
|
||||
role attribution.
|
||||
|
||||
### Open Questions
|
||||
|
||||
Preserve every open question. Do not convert open questions into decisions or
|
||||
action items.
|
||||
|
||||
Do not create empty sections.
|
||||
|
||||
Faithfulness rules:
|
||||
|
||||
- Never invent facts.
|
||||
- Never strengthen uncertainty.
|
||||
- Never weaken decisions.
|
||||
- Never attribute statements to people unless explicitly present.
|
||||
- Keep distinctions between background, decisions, action items and open
|
||||
questions.
|
||||
- If information appears contradictory in the input, preserve the uncertainty
|
||||
without resolving it.
|
||||
- If two items are similar but not already consolidated, keep their meanings
|
||||
distinct in the rendered document.
|
||||
|
||||
Return Markdown only.
|
||||
Do not include an introductory explanation.
|
||||
Do not include a concluding summary.
|
||||
@@ -0,0 +1,15 @@
|
||||
# responsibility_attribution_negative
|
||||
|
||||
Tests that thematic proximity, department discussion and objection do not create
|
||||
an action-item owner.
|
||||
|
||||
The scenario includes a Marketing participant who comments critically on
|
||||
Business Development criteria. No one assigns that participant responsibility
|
||||
for defining the criteria, and the participant does not accept such a task.
|
||||
|
||||
Expected behavior:
|
||||
|
||||
- Do not assign Business Development criteria to the Marketing participant.
|
||||
- Do not reclassify the Marketing participant as Business Development.
|
||||
- Preserve the objection as a position.
|
||||
- Keep responsibility open unless explicitly assigned.
|
||||
@@ -0,0 +1,39 @@
|
||||
{
|
||||
"facts": [
|
||||
{
|
||||
"fact": "Criteria are still needed for Business Development projects.",
|
||||
"evidence": "Mara: For the new partner intake, we still need criteria for Business Development projects."
|
||||
},
|
||||
{
|
||||
"fact": "Noah is speaking from Marketing.",
|
||||
"evidence": "Noah: From Marketing, I can tell you that generic criteria will not work."
|
||||
},
|
||||
{
|
||||
"fact": "Noah is not taking responsibility for defining the Business Development criteria.",
|
||||
"evidence": "Noah: I am not taking that on."
|
||||
}
|
||||
],
|
||||
"decisions": [
|
||||
{
|
||||
"decision": "Responsibility for defining the Business Development criteria remains open until Business Development confirms an owner.",
|
||||
"evidence": "Mara: Agreed. Responsibility stays open for now, and we will ask Business Development for an owner."
|
||||
}
|
||||
],
|
||||
"todos": [
|
||||
{
|
||||
"task": "Ask Business Development for an owner for defining the Business Development criteria.",
|
||||
"responsible": null,
|
||||
"deadline": null,
|
||||
"evidence": "Mara: Agreed. Responsibility stays open for now, and we will ask Business Development for an owner."
|
||||
}
|
||||
],
|
||||
"questions": [],
|
||||
"positions": [
|
||||
{
|
||||
"speaker": "Noah",
|
||||
"position": "Noah says generic criteria will not work for different Marketing and product contexts.",
|
||||
"evidence": "Noah: From Marketing, I can tell you that generic criteria will not work. A digital campaign and a physical product launch need different checks."
|
||||
}
|
||||
],
|
||||
"technical": []
|
||||
}
|
||||
@@ -0,0 +1,11 @@
|
||||
Mara: For the new partner intake, we still need criteria for Business Development projects.
|
||||
|
||||
Noah: From Marketing, I can tell you that generic criteria will not work. A digital campaign and a physical product launch need different checks.
|
||||
|
||||
Mara: Right, so the Business Development criteria remain open.
|
||||
|
||||
Noah: I am not taking that on. I only want to flag that the criteria cannot be generic.
|
||||
|
||||
Lea: Then we should keep the owner open until Business Development confirms who will define them.
|
||||
|
||||
Mara: Agreed. Responsibility stays open for now, and we will ask Business Development for an owner.
|
||||
Reference in New Issue
Block a user