Compare commits
21
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
8a0f38fce4 | ||
|
|
df89a38829 | ||
|
|
d77bfedb6e | ||
|
|
d94436af43 | ||
|
|
8dab928763 | ||
|
|
f2d21c1faf | ||
|
|
a9dab7c81a | ||
|
|
70007ea8f2 | ||
|
|
3918c0b1c4 | ||
|
|
7fa771a7e4 | ||
|
|
3229786b5c | ||
|
|
8ca62fbd92 | ||
|
|
0d4b426021 | ||
|
|
97a22f3ebb | ||
|
|
1c36fa76fb | ||
|
|
a1fe89de52 | ||
|
|
19672adab4 | ||
|
|
4ffd4c1c5d | ||
|
|
18beb3385f | ||
|
|
bcb197a908 | ||
|
|
c2b7b6b4d2 |
@@ -34,6 +34,7 @@ htmlcov/
|
||||
# Experiment Outputs
|
||||
experiments/**/output/
|
||||
experiments/**/results/
|
||||
artifacts/experiments/**/
|
||||
|
||||
# Pipeline runtime artifacts
|
||||
samples/raw/
|
||||
|
||||
@@ -25,7 +25,28 @@ Implemented:
|
||||
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
|
||||
- Meeting Context V1 loading, validation and optional extraction prompt
|
||||
injection with minimal extraction JSON provenance.
|
||||
- FFmpeg-backed WAV, FLAC and M4A preparation into a per-run canonical mono
|
||||
16 kHz signed PCM16 WAV artifact before transcription or diarization. Audio
|
||||
preparation always runs. Optional loudness normalization defaults to on and
|
||||
currently uses the isolated FFmpeg filter
|
||||
`loudnorm=I=-16:LRA=11:TP=-1.5`. This is a conservative speech-recording
|
||||
default and may be revisited after empirical comparison without changing the
|
||||
orchestration API.
|
||||
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
|
||||
- Direct protocol prompt input protection: diarized transcripts are rendered as
|
||||
compact adjacent-speaker blocks without per-segment timestamps. Every source
|
||||
segment remains represented in order. A deterministic heuristic enforces a
|
||||
configurable safe input budget, falls back to complete plain transcript text
|
||||
when necessary, and fails before any Ollama request if even that input is too
|
||||
large. Silent head/tail truncation is prohibited.
|
||||
- The `qwen3.8:27b` direct-protocol stage explicitly requests `num_ctx=32768`
|
||||
and `think=false`; the practical prompt target is approximately 29,000 tokens.
|
||||
A 31,038-token synthetic prompt passed, but larger prompts are not assumed safe
|
||||
from the model's advertised 262,144-token native context alone.
|
||||
- `regenerate_mvp_protocol` updates the run's validated Meeting Context and
|
||||
regenerates protocol artifacts from the existing diarized transcript when
|
||||
available. It never reruns audio preparation, Whisper or Pyannote, and it
|
||||
preserves anonymous speaker labels in the source transcript.
|
||||
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
|
||||
gold-test runner validation.
|
||||
- Meeting Context V1 scaffold and documentation for manually maintained
|
||||
@@ -139,6 +160,9 @@ departments only when they are explicitly supplied as metadata. It must not be
|
||||
used to infer responsibilities. In the current implementation this context can
|
||||
be injected into chunk extraction prompts as authoritative metadata, and only
|
||||
minimal provenance is written to extraction JSON.
|
||||
The implemented MVP statuses are exactly `present` and `mentioned_only`.
|
||||
Legacy entries without a status receive collection-appropriate defaults. Only
|
||||
present participants may be targets of explicit `SPEAKER_XX` mappings.
|
||||
|
||||
A `responsible` or future `owner` / `assignee` value may be recorded only when
|
||||
source evidence explicitly assigns, accepts or confirms responsibility. If the
|
||||
|
||||
@@ -10,7 +10,54 @@ Im Mittelpunkt steht nicht die Softwarearchitektur, sondern die Frage:
|
||||
|
||||
> **Wie lässt sich aus einem realen Meeting möglichst zuverlässig strukturiertes Wissen extrahieren?**
|
||||
|
||||
Neue Ideen werden zunächst hier experimentell umgesetzt. Erst wenn sich ein Ansatz bewährt hat, wird er in den eigentlichen *Meeting Assistant* übernommen.
|
||||
Meeting Lab ist die experimentelle R&D-Umgebung fuer den zukuenftigen
|
||||
*Meeting Assistant*. Es dient gleichzeitig als Forschungsplattform,
|
||||
Architektur-Spielwiese, Regressionsframework, Benchmark-Umgebung und
|
||||
Prototypimplementierung. Neue Ideen werden hier untersucht und gegen reale
|
||||
Meeting-Beispiele validiert, bevor sie fuer das Produkt in Betracht kommen.
|
||||
|
||||
Der beabsichtigte Reifeprozess ist:
|
||||
|
||||
```text
|
||||
Research idea
|
||||
->
|
||||
Meeting Lab experiment
|
||||
->
|
||||
Regression tests
|
||||
->
|
||||
Stable architecture
|
||||
->
|
||||
Meeting Assistant implementation
|
||||
```
|
||||
|
||||
Nur ausreichend ausgereifte und verifizierte Komponenten sollen in den
|
||||
Meeting Assistant uebernommen werden. Meeting Lab darf bewusst experimentelle
|
||||
Ansaetze und Entwicklungszweige enthalten, die verworfen werden oder nie den
|
||||
Assistant erreichen.
|
||||
|
||||
## Meeting Assistant und langfristige Produktentwicklung
|
||||
|
||||
Der Meeting Assistant ist als Produktionsanwendung vorgesehen. Seine erste
|
||||
oeffentliche Beta soll auf einem stabilen Meeting-Lab-MVP basieren und eine
|
||||
polierte User Experience, Installer, Konfiguration und eine produktionsreife
|
||||
Pipeline bieten. Eine GUI ist optional; experimentelle Funktionen sollen
|
||||
standardmaessig nicht aktiviert sein.
|
||||
|
||||
Meeting Lab entwickelt sich unabhaengig weiter und bleibt der langfristige
|
||||
Innovationszweig. Der erwartete Transferpfad lautet:
|
||||
|
||||
```text
|
||||
Meeting Lab Alpha
|
||||
->
|
||||
Meeting Lab Beta
|
||||
->
|
||||
Meeting Assistant Beta
|
||||
->
|
||||
Meeting Assistant Release
|
||||
```
|
||||
|
||||
Meeting-Assistant-Releases uebernehmen damit gezielt bewaehrte Meeting-Lab-
|
||||
Komponenten, waehrend der Assistant den stabilen Produktzweig bildet.
|
||||
|
||||
---
|
||||
|
||||
|
||||
+206
-2
@@ -2,13 +2,63 @@
|
||||
|
||||
## Purpose
|
||||
|
||||
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.
|
||||
The **Meeting Lab** is the experimental R&D environment for developing and
|
||||
evaluating methods to extract structured knowledge from real meeting
|
||||
transcripts. It is the research platform, architecture playground, regression
|
||||
framework, benchmark environment and prototype implementation for the future
|
||||
Meeting Assistant.
|
||||
|
||||
Its purpose is not to build a complete meeting assistant, but to answer a single question:
|
||||
|
||||
> **How can knowledge be extracted from real discussions as reliably as possible?**
|
||||
|
||||
Successful approaches will later be integrated into the Meeting Assistant project.
|
||||
Its purpose is to validate ideas before they are promoted into the product.
|
||||
Only sufficiently mature and verified components should migrate into Meeting
|
||||
Assistant. Meeting Lab may intentionally contain experiments or development
|
||||
branches that are rejected, remain inconclusive or never reach the Assistant.
|
||||
|
||||
## Meeting Lab and Meeting Assistant lifecycle
|
||||
|
||||
Meeting Lab and Meeting Assistant have different long-term responsibilities:
|
||||
|
||||
- **Meeting Lab** is the long-term innovation branch. It favors learning,
|
||||
inspectable experiments, regression evidence, benchmarks and architectural
|
||||
change.
|
||||
- **Meeting Assistant** is the stable product branch. It favors a polished user
|
||||
experience, installation, configuration and a production pipeline.
|
||||
|
||||
Architectural promotion follows an evidence-based lifecycle:
|
||||
|
||||
```text
|
||||
Research idea
|
||||
->
|
||||
Meeting Lab experiment
|
||||
->
|
||||
Regression tests
|
||||
->
|
||||
Stable architecture
|
||||
->
|
||||
Meeting Assistant implementation
|
||||
```
|
||||
|
||||
The expected release progression is:
|
||||
|
||||
```text
|
||||
Meeting Lab Alpha
|
||||
->
|
||||
Meeting Lab Beta
|
||||
->
|
||||
Meeting Assistant Beta
|
||||
->
|
||||
Meeting Assistant Release
|
||||
```
|
||||
|
||||
The first public Meeting Assistant beta should be based on a stable Meeting
|
||||
Lab MVP. It should provide a polished user experience, an installer,
|
||||
configuration and a production-quality pipeline. A GUI is optional.
|
||||
Experimental features should not be enabled by default. Meeting Lab continues
|
||||
to evolve independently after components have migrated; promotion does not
|
||||
turn the Lab itself into the product branch.
|
||||
|
||||
---
|
||||
|
||||
@@ -356,6 +406,144 @@ The architecture document only describes the overall system.
|
||||
|
||||
---
|
||||
|
||||
# Version 2 Accepted Architectural Direction
|
||||
|
||||
The following topics are accepted architectural goals for Version 2. They
|
||||
record direction reached through the BUG-011 through BUG-015 investigations;
|
||||
they are not descriptions of implemented behavior or authorization to change
|
||||
the current pipeline.
|
||||
|
||||
## Speaker diarization before semantic analysis
|
||||
|
||||
Version 2 should determine **who is speaking before semantic analysis**.
|
||||
Speaker identity contains evidence that cannot reliably be reconstructed from
|
||||
text alone. It helps distinguish, for example, who answers a question, accepts
|
||||
work, agrees with a proposal, or advances the discussion after another
|
||||
speaker. It also preserves conversational flow that anonymous transcript text
|
||||
can erase.
|
||||
|
||||
Diarization is therefore a semantic prerequisite in the intended Version 2
|
||||
architecture, not merely a display enhancement. Its output should remain
|
||||
traceable to transcript segments so later stages can preserve speaker and
|
||||
source provenance.
|
||||
|
||||
## Persistent speaker identification
|
||||
|
||||
Version 2 should add a persistent speaker database and an interactive identity
|
||||
workflow during import:
|
||||
|
||||
```text
|
||||
Unknown speaker detected
|
||||
->
|
||||
Representative audio sample (approximately 20 seconds)
|
||||
->
|
||||
User selects an existing identity or creates a new identity
|
||||
->
|
||||
Known speaker available for future recognition
|
||||
```
|
||||
|
||||
Automatic recognition may suggest an identity, but user confirmation governs
|
||||
the persistent association. Over the long term, speaker embeddings rather
|
||||
than raw meeting recordings should be the persistent recognition
|
||||
representation. Representative raw audio is an import and confirmation aid,
|
||||
not the intended durable identity store. Privacy, deletion and false-match
|
||||
handling require separate design before implementation.
|
||||
|
||||
## Meeting Context as a probabilistic prior
|
||||
|
||||
Known meeting participants should influence semantic interpretation, but
|
||||
Meeting Context is a **probabilistic prior**, not a deterministic semantic
|
||||
rule. It can make one interpretation more plausible and help focus review; it
|
||||
must never manufacture a commitment, decision or responsibility assignment.
|
||||
|
||||
For example, if Marleen is confirmed as present, “Marleen müsste sich mal
|
||||
äußern” is more likely to be conversation management directed at a current
|
||||
participant than future project work. Presence alone does not prove this
|
||||
interpretation, and it does not establish an Action Item or responsibility.
|
||||
Explicit meeting evidence remains authoritative. This extends, rather than
|
||||
weakens, the responsibility attribution invariant.
|
||||
|
||||
The existing Meeting Context V2 entity direction is documented in
|
||||
[`adr-meeting-context-v2-entity-registry.md`](adr-meeting-context-v2-entity-registry.md).
|
||||
Speaker identities and meeting-specific participant confirmation should
|
||||
eventually feed that context without turning registry metadata into semantic
|
||||
facts.
|
||||
|
||||
## Conversation Management versus Meeting Content
|
||||
|
||||
BUG-015 reinforced that not every utterance is protocol-worthy content.
|
||||
Version 2 should conceptually distinguish:
|
||||
|
||||
- **Conversation Management**: utterances that coordinate the meeting itself,
|
||||
such as asking a present participant to speak, moderation, requesting a
|
||||
slide, or asking someone to repeat something.
|
||||
- **Meeting Content**: propositions that may contribute to the meeting's
|
||||
durable knowledge, including facts, technical findings, decisions, action
|
||||
items and open questions.
|
||||
|
||||
This is an architectural concept, not a currently implemented category or
|
||||
filter. The distinction should prevent conversational coordination from being
|
||||
promoted into project commitments while retaining sufficient provenance to
|
||||
understand dialogue. Context and diarization can inform the distinction, but
|
||||
neither should act as a deterministic keyword or participant rule.
|
||||
|
||||
## Evidence and commitment before protocol eligibility
|
||||
|
||||
The BUG-015 design study concludes that semantic state should be classified
|
||||
before policy determines whether an item is eligible for a protocol. A binary
|
||||
keep/reject verifier conflates evidence recognition with publication policy
|
||||
and loses valid intermediate states.
|
||||
|
||||
The proposed decision progression is:
|
||||
|
||||
```text
|
||||
idea -> option -> proposal -> preferred option -> tentative agreement -> decision
|
||||
```
|
||||
|
||||
The proposed action progression is:
|
||||
|
||||
```text
|
||||
possible next step -> recommendation -> requested action
|
||||
-> established action -> ongoing work -> completed
|
||||
```
|
||||
|
||||
Questions combine a communicative **kind** with an independent **resolution
|
||||
state**, rather than treating every uncertainty or interrogative as an Open
|
||||
Question. Responsibility remains an independent dimension and may be recorded
|
||||
only when explicitly assigned, accepted or confirmed. Evidence strength is
|
||||
also independent: it describes support for a semantic label, not semantic
|
||||
maturity or protocol eligibility.
|
||||
|
||||
After semantic state classification, explicit policy should select Decisions,
|
||||
Action Items and Open Questions for a particular output view. Renderers should
|
||||
receive policy-selected semantic content and must not promote proposals or
|
||||
conversation management into commitments. The complete taxonomy, trade-offs,
|
||||
architecture interactions and migration questions are recorded in
|
||||
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md).
|
||||
|
||||
## Topic-oriented primary protocol
|
||||
|
||||
One of the highest-level Version 2 requirements is:
|
||||
|
||||
> **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.**
|
||||
|
||||
Semantic categories remain metadata and supporting structure within topics;
|
||||
they must not dictate the main document structure. The conceptual target flow
|
||||
is `Transcript -> Evidence Extraction -> Topic Reconstruction -> Semantic
|
||||
Synthesis -> Protocol Rendering`. The primary renderer should eventually
|
||||
receive topic-oriented semantic knowledge. Category-oriented Action Item,
|
||||
Decision, Open Question and management views remain useful derived outputs.
|
||||
|
||||
The detailed rationale, example, architectural implications and explicit
|
||||
deferral of a final Topic Reconstruction schema are documented in
|
||||
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md#thematic-protocol-as-the-primary-structure).
|
||||
|
||||
Implementation is intentionally postponed until this architecture has been
|
||||
reviewed. No Version 2 goal in this section changes current extraction,
|
||||
canonicalization, consolidation or rendering behavior.
|
||||
|
||||
---
|
||||
|
||||
# Current State
|
||||
|
||||
Implemented:
|
||||
@@ -389,6 +577,22 @@ Action Item may have no known owner, while a named owner requires explicit
|
||||
assignment, volunteering or acceptance. These are semantic LLM classifications;
|
||||
deterministic validation must not guess intent from keywords.
|
||||
|
||||
### Classification Verifier
|
||||
|
||||
Before extraction output is normalized for Canonicalizer input, Decision,
|
||||
Action Item and Open Question candidates pass through a semantic precision
|
||||
gate. Facts and technical details pass through unchanged. Each candidate is
|
||||
reviewed independently with its evidence and bounded local chunk context; the
|
||||
verifier may only keep or reject the existing candidate. It cannot add or
|
||||
rewrite semantic items.
|
||||
|
||||
Verifier output contains the stable candidate ID, category, `keep|reject`
|
||||
verdict, evidence-based reason and `responsibility_supported`. A kept Action
|
||||
Item with an unsupported named owner is retained with its responsibility
|
||||
cleared. Malformed output fails the extraction verification substage closed,
|
||||
after preserving candidate input and raw response. Per-candidate results and an
|
||||
aggregate audit remain traceable before Canonicalizer input is written.
|
||||
|
||||
## Semantic Consolidator failure handling
|
||||
|
||||
Semantic Consolidator V0 preserves every raw model response before parsing.
|
||||
|
||||
@@ -0,0 +1,369 @@
|
||||
# Evidence / Commitment Model
|
||||
|
||||
Status: architecture design study for BUG-015; no implementation recommendation is made here.
|
||||
|
||||
## Motivation
|
||||
|
||||
Meeting Lab currently asks extraction and the BUG-015 verifier to cross a category boundary in one step: a candidate is either a protocol-worthy Decision, Action Item or Open Question, or it is rejected. That combines two different questions:
|
||||
|
||||
1. What does the transcript provide evidence for?
|
||||
2. Which evidenced states should a particular protocol view publish?
|
||||
|
||||
Meetings do not move directly from absence to commitment. An option can become a proposal, then a preferred option, then a tentative agreement and finally a decision. Work can be suggested, requested, assigned, accepted, already underway or completed. A question can be asked and answered, or can remain explicitly unresolved. Rejecting everything below the final protocol threshold discards useful evidence; promoting it creates false commitments.
|
||||
|
||||
The proposed model therefore records the evidenced semantic state first. A later policy selects protocol-worthy states. It preserves the existing separation between extraction, deterministic canonicalization, semantic consolidation and rendering, and it does not make rendered output the semantic source of truth.
|
||||
|
||||
## Observed BUG-015 failures
|
||||
|
||||
The BUG-015 Gold cases expose both sides of the binary-classification problem.
|
||||
|
||||
| Progeo-derived evidence | Correct evidence state | Binary failure to avoid |
|
||||
| --- | --- | --- |
|
||||
| “Die Option steht im Raum, das Material chemisch recyceln zu lassen.” | option | False Decision |
|
||||
| “Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden.” | personal preference with a conditional option | False Decision |
|
||||
| “Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht. … das ist entschieden.” | explicit decision: rejection | False rejection |
|
||||
| “Die Geometrie kann man vielleicht noch optimieren.” | possible next step | False Action Item |
|
||||
| “Marleen müsste vielleicht mal äußern …” | suggested/requested action; no accepted or confirmed responsibility | False Action Item and false owner |
|
||||
| “Wir könnten Dirk Textor vielleicht noch einmal kontaktieren.” | suggestion | False Action Item |
|
||||
| “Nina, übernimmst du …? – Ja, ich übernehme …” | accepted action with explicit responsibility and deadline | False rejection |
|
||||
| “Den CET-Artikel erstellen wir bereits …” | ongoing established work; owner not evidenced | Consistent false rejection by the binary verifier |
|
||||
| “Ob sich das Waschen lohnt, weiß ich nicht.” | uncertainty | False Open Question |
|
||||
| “Gibt es schon ein Programm? – Ja …” | explicit, resolved question | False Open Question if local resolution is ignored |
|
||||
| “Welche Daten … dürfen wir veröffentlichen? … weiterhin ungeklärt.” | explicit unresolved question | Consistent false rejection by the binary verifier |
|
||||
|
||||
The verifier was reliable on several negative cases and on explicit commitment language, but not on valid states whose evidence did not resemble a fresh agreement: already-established work and an explicitly unresolved information need. This suggests a representation problem, not merely an insufficient keep/reject prompt.
|
||||
|
||||
## Model shape
|
||||
|
||||
The model is deliberately small and compositional. Each evidence item has:
|
||||
|
||||
- a **domain**: decision, action or information need;
|
||||
- a **semantic state** within that domain;
|
||||
- an **evidence support level** describing how directly the transcript supports that label;
|
||||
- preserved transcript evidence and source references;
|
||||
- domain-specific independent attributes, such as responsibility or resolution;
|
||||
- no automatic claim that the item belongs in a final protocol.
|
||||
|
||||
Semantic maturity and evidence support are independent. An explicitly worded proposal remains a proposal; it is not a weak decision. Conversely, ongoing work can be strongly evidenced without a recorded moment of assignment.
|
||||
|
||||
## Semantic state diagrams
|
||||
|
||||
The arrows show common progressions, not mandatory workflows. Meetings may enter at any state, skip states, regress, or end without commitment.
|
||||
|
||||
### Decision evolution
|
||||
|
||||
```text
|
||||
idea -> option -> proposal -> preferred_option -> tentative_agreement -> decision
|
||||
| | | | |
|
||||
+----------+--------------+--------------------+----------> withdrawn
|
||||
superseded
|
||||
reopened
|
||||
```
|
||||
|
||||
An **idea** is an undeveloped possibility. An **option** is a candidate alternative. A **proposal** asks the group to adopt an outcome. A **preferred option** expresses comparative preference without settlement. A **tentative agreement** records provisional convergence that is explicitly conditional or awaiting confirmation. A **decision** records an outcome the meeting settled, selected, approved, rejected or committed to. The intermediate preferred and tentative states matter because they prevent likely direction from being confused with commitment.
|
||||
|
||||
### Action evolution
|
||||
|
||||
```text
|
||||
possible_next_step -> recommendation -> requested_action -> established_action -> ongoing_work -> completed
|
||||
| | | | |
|
||||
+------------------+-----------------+--------------------+----------> cancelled
|
||||
|
||||
Responsibility (independent):
|
||||
unset -> proposed_responsible -> assigned -> accepted/confirmed
|
||||
```
|
||||
|
||||
An action can enter directly as **established_action** through explicit assignment, acceptance or commitment. It can enter directly as **ongoing_work** when the transcript clearly says that the work is already being performed, as in the CET example. `proposed_responsible` is evidence about a suggestion, not permission to populate the protocol's responsible field. Only `assigned`, `accepted` or `confirmed` supports recorded responsibility, and only where the meeting evidence explicitly assigns, accepts or confirms it.
|
||||
|
||||
### Information-need evolution
|
||||
|
||||
```text
|
||||
uncertainty ---------> information_need(unresolved) ---------> resolved
|
||||
curiosity -----------> explicit_question(unresolved) --------> resolved
|
||||
request_for_clarification(unresolved) -----------------------> resolved
|
||||
missing_information(unresolved) -----------------------------> resolved
|
||||
|
||||
rhetorical_question -> discourse only
|
||||
```
|
||||
|
||||
Uncertainty and curiosity do not automatically create an information need. An explicit question may be resolved immediately. “Open Question” is therefore a policy result derived primarily from a qualifying kind plus `unresolved`, rather than a primitive utterance category.
|
||||
|
||||
## Proposed taxonomy
|
||||
|
||||
### Common fields
|
||||
|
||||
| Field | Proposed values | Purpose |
|
||||
| --- | --- | --- |
|
||||
| `domain` | `decision`, `action`, `information_need` | Selects the domain taxonomy. |
|
||||
| `state` | Domain-specific values below | Records what the evidence says now. |
|
||||
| `support` | `indirect`, `direct`, `explicit` | Rates support for the chosen state, not protocol importance. |
|
||||
| `polarity` | `positive`, `negative` where applicable | Preserves approval/rejection and adoption/refusal without rewriting meaning. |
|
||||
| `evidence` | transcript excerpt(s) | Proves the state and attributes. |
|
||||
| `source_refs` | stable source references | Retains provenance across stages. |
|
||||
|
||||
`support` is intentionally not `weak / medium / strong`. Those terms mix confidence, semantic maturity and number of witnesses. The proposed meanings are narrower:
|
||||
|
||||
- `indirect`: the state is supported by context but not stated in a self-contained utterance; downstream commitment policy should normally be conservative.
|
||||
- `direct`: an utterance directly expresses the state, such as a proposal, preference, ongoing-work statement or question.
|
||||
- `explicit`: the utterance also names the decisive status, for example “das ist entschieden”, “ich übernehme”, “wir erstellen bereits” or “weiterhin ungeklärt”.
|
||||
|
||||
This common axis simplifies evidence auditing and threshold policy, but cannot replace domain state. An explicit suggestion is still not an Action Item, and an explicit uncertainty is still not an Open Question. Model confidence, if retained at all, must be a separate operational field and must not be presented as evidence strength.
|
||||
|
||||
### Decision states
|
||||
|
||||
| State | Meaning | Default protocol treatment |
|
||||
| --- | --- | --- |
|
||||
| `idea` | Undeveloped possibility or brainstorming contribution | Exclude |
|
||||
| `option` | Candidate alternative under consideration | Exclude |
|
||||
| `proposal` | Outcome offered for adoption | Exclude |
|
||||
| `preferred_option` | Expressed preference among alternatives | Exclude |
|
||||
| `tentative_agreement` | Provisional convergence with an expressed condition or pending confirmation | Exclude from Decisions; potentially expose to an editorial view |
|
||||
| `decision` | Settled, selected, approved, rejected or committed outcome | Include as Decision when evidence is sufficient |
|
||||
| `reopened` | Earlier decision explicitly returned to unresolved consideration | Do not render the earlier decision as currently settled without qualification |
|
||||
| `superseded` | Earlier decision replaced by a later one | Retain provenance; normally render only the current decision |
|
||||
| `withdrawn` | Candidate state explicitly withdrawn | Exclude as current commitment |
|
||||
|
||||
`idea`, `option`, `proposal`, `preferred_option`, `tentative_agreement` and `decision` are necessary distinctions for BUG-015. `reopened`, `superseded`, `withdrawn` are lifecycle states needed to avoid treating historical evidence as current policy; they need not be first-iteration extraction targets.
|
||||
|
||||
### Action states and attributes
|
||||
|
||||
| State | Meaning | Default protocol treatment |
|
||||
| --- | --- | --- |
|
||||
| `possible_next_step` | Hypothetical or exploratory action | Exclude |
|
||||
| `recommendation` | Action advocated but not established as work | Exclude |
|
||||
| `requested_action` | Someone asks that work be done, without enough evidence that it is established | Exclude by default |
|
||||
| `established_action` | Concrete future work established by assignment, acceptance, commitment or confirmation | Include as Action Item |
|
||||
| `ongoing_work` | Concrete work explicitly already underway | Include when still relevant; do not require a newly witnessed assignment |
|
||||
| `completed` | Work explicitly reported complete | Exclude from open Action Items; retain as status/history |
|
||||
| `cancelled` | Work explicitly cancelled or declined | Exclude from open Action Items; retain provenance |
|
||||
|
||||
`suggestion` is represented as `possible_next_step` or `recommendation`, depending on whether the speaker advocates it. `accepted_action` and `assigned_work` should not be competing lifecycle states: both establish `established_action`, while the independent commitment basis records `accepted`, `assigned`, `self_committed` or `confirmed_existing`. This avoids an artificial choice when, as with Nina, an assignment and acceptance occur together.
|
||||
|
||||
Responsibility is independent:
|
||||
|
||||
| Responsibility status | Meaning | May populate `responsible`? |
|
||||
| --- | --- | --- |
|
||||
| `unset` | No person/team explicitly tied to ownership | No |
|
||||
| `proposed` | A person is suggested, asked speculatively or mentioned near the work | No |
|
||||
| `assigned` | The meeting explicitly assigns the work | Yes |
|
||||
| `accepted` | The party explicitly accepts or volunteers | Yes |
|
||||
| `confirmed` | Existing ownership is explicitly confirmed | Yes |
|
||||
|
||||
This preserves the project invariant: discussion, expertise, adjacency, organizational role and likely ownership never establish responsibility. An action may be valid with `responsibility_status: unset`, as with established CET work.
|
||||
|
||||
### Information-need kinds and resolution
|
||||
|
||||
Question form and resolution are separate attributes.
|
||||
|
||||
| Kind | Meaning | Can become an Open Question? |
|
||||
| --- | --- | --- |
|
||||
| `uncertainty` | Speaker expresses doubt or lack of certainty without establishing a concrete need | No, by itself |
|
||||
| `curiosity` | Interest without a concrete need requiring follow-up | No, by itself |
|
||||
| `explicit_question` | Direct interrogative seeking an answer | Yes, if unresolved |
|
||||
| `request_for_clarification` | Explicit request to clarify a concrete matter | Yes, if unresolved |
|
||||
| `missing_information` | Concrete required information is stated as absent | Yes, if unresolved |
|
||||
| `rhetorical_question` | Interrogative used for emphasis rather than an answer | No |
|
||||
|
||||
Resolution is one of `unresolved`, `resolved`, or `resolution_unclear`. `resolved_question` is therefore not a separate kind: it is, for example, `explicit_question + resolved`. The publication example is `explicit_question + unresolved`; the event-program example is `explicit_question + resolved`; the washing example is `uncertainty` unless later evidence establishes a concrete unresolved need. `resolution_unclear` preserves evidence but should not silently pass a precision-first Open Question policy.
|
||||
|
||||
An “unresolved question” is a derived, protocol-relevant combination rather than a fourth kind. This makes the classification testable: kind answers what communicative act occurred, and resolution answers what remained at the end of the available context.
|
||||
|
||||
## Progeo examples under the model
|
||||
|
||||
| Evidence | Proposed representation | Working Protocol policy result |
|
||||
| --- | --- | --- |
|
||||
| Chemical recycling “Option steht im Raum” | `decision / option / explicit` | Not a Decision |
|
||||
| Technikum preference | `decision / preferred_option / direct`; personal scope preserved | Not a Decision |
|
||||
| Schlummer collaboration rejected and “entschieden” | `decision / decision / explicit / negative` | Decision |
|
||||
| Geometry “kann man vielleicht … optimieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item |
|
||||
| Marleen “müsste vielleicht mal” | `action / requested_action / direct`, responsibility proposed only | Not an Action Item; no responsible party |
|
||||
| Textor “könnten … kontaktieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item |
|
||||
| Nina asks and accepts the review by Friday | `action / established_action / explicit`, basis assigned + accepted, responsible Nina, deadline Friday | Action Item |
|
||||
| CET article “erstellen wir bereits” and later circulation | `action / ongoing_work / explicit`, responsibility unset | Action Item without invented owner |
|
||||
| Washing “ob sich das lohnt, weiß ich nicht” | `information_need / uncertainty / direct`, resolution unclear | Not an Open Question |
|
||||
| Event program asked and answered | `information_need / explicit_question / direct`, resolved | Not an Open Question |
|
||||
| Publishable energy-audit data “weiterhin ungeklärt” | `information_need / explicit_question / explicit`, unresolved | Open Question |
|
||||
|
||||
These labels preserve every BUG-015 distinction without treating rejected protocol candidates as meaningless.
|
||||
|
||||
## Thematic protocol as the primary structure
|
||||
|
||||
The Evidence / Commitment Model supplies semantic distinctions inside a larger
|
||||
meeting reconstruction. It does not imply that the primary protocol should be
|
||||
organized as one section per semantic category.
|
||||
|
||||
The central Version 2 requirement is:
|
||||
|
||||
> **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.**
|
||||
|
||||
A primary protocol should reconstruct which topics were discussed, what
|
||||
relevant information emerged within each topic, how the discussion developed,
|
||||
which alternatives, ideas, objections or proposals mattered, what outcome or
|
||||
current state was reached, and which actions or unresolved questions resulted
|
||||
from that topic.
|
||||
|
||||
Semantic categories remain essential, but as metadata and supporting structure
|
||||
attached to topics. Facts, technical findings, ideas, alternatives, proposals,
|
||||
objections, decisions, action items and open questions must not dictate the
|
||||
main document structure.
|
||||
|
||||
For example, a human-style topic section may read:
|
||||
|
||||
```text
|
||||
## Trial setup
|
||||
|
||||
Several variants for the next trial were discussed.
|
||||
|
||||
A thinner carrier material was proposed as one possible alternative.
|
||||
Concerns were raised regarding its mechanical suitability.
|
||||
The group therefore decided to continue with the existing setup for the next trial.
|
||||
|
||||
Nina will obtain the remaining samples before the next production run.
|
||||
|
||||
The publication question remains unresolved.
|
||||
```
|
||||
|
||||
This preserves the relationship between the proposal, objection, decision,
|
||||
resulting work and unresolved question. A primary document split into separate
|
||||
Ideas, Objections, Decisions and Action Items sections would lose that thematic
|
||||
and conversational relationship.
|
||||
|
||||
Category-oriented outputs remain valuable as derived secondary views. An
|
||||
Action Item table, Decision register, Open Questions list or Management summary
|
||||
can be generated from the same underlying meeting knowledge after thematic
|
||||
reconstruction. These indexes and summaries do not replace the primary
|
||||
topic-oriented protocol.
|
||||
|
||||
The conceptual Version 2 flow is:
|
||||
|
||||
```text
|
||||
Transcript
|
||||
->
|
||||
Evidence Extraction
|
||||
->
|
||||
Topic Reconstruction
|
||||
->
|
||||
Semantic Synthesis
|
||||
->
|
||||
Protocol Rendering
|
||||
```
|
||||
|
||||
Evidence items and their semantic metadata remain attached to topics and
|
||||
support Semantic Synthesis. The Version 2 renderer should receive
|
||||
topic-oriented semantic knowledge rather than a flat collection grouped by
|
||||
category. This direction does not define the final Topic Reconstruction schema,
|
||||
redesign the current pipeline or remove the existing Working Protocol V2
|
||||
architecture. It is an accepted target architecture whose implementation is
|
||||
postponed.
|
||||
|
||||
## Architectural impact
|
||||
|
||||
### Extraction
|
||||
|
||||
Extraction could emit evidence records with richer labels instead of immediately claiming `decision`, `action_item` or `open_question`. This is a bounded increase in extraction vocabulary, not a request for larger context windows, a multi-chunk strategy or a monolithic synthesis prompt. Existing Facts, Positions and Technical Details remain distinct categories; the new domains refine only commitment-sensitive content.
|
||||
|
||||
The immediate output would describe the observed state. Protocol category selection would happen later through an explicit policy. Some policy rules can be deterministic once semantic labels are trustworthy, for example:
|
||||
|
||||
```text
|
||||
Decision := domain=decision AND state=decision
|
||||
Action Item := domain=action AND state IN {established_action, ongoing_work}
|
||||
Open Question := domain=information_need
|
||||
AND kind IN {explicit_question, request_for_clarification, missing_information}
|
||||
AND resolution=unresolved
|
||||
Responsible := responsibility_status IN {assigned, accepted, confirmed}
|
||||
```
|
||||
|
||||
This keeps downstream selection simple, but does not make semantic extraction deterministic. The LLM still has to distinguish proposals from decisions and uncertainty from unresolved needs.
|
||||
|
||||
### LLM stability
|
||||
|
||||
The architecture should reduce one source of instability: the model no longer has to erase a supported proposition merely because it falls below a protocol threshold, nor relocate it into another category to retain it. Labels correspond more closely to observable speech acts and lifecycle statements, and downstream thresholds become explicit.
|
||||
|
||||
It will not eliminate instability. More labels create adjacent-class boundaries, local context may not reveal resolution, and indirect language remains difficult. Stability depends on small schemas, evidence spans, one best supported state, conservative handling of `resolution_unclear`, and later evaluation against the Gold corpus. A generic evidence-strength score alone would likely worsen instability because it invites subjective grading.
|
||||
|
||||
### Canonicalizer
|
||||
|
||||
The Canonicalizer would benefit from carrying normalized state names, independent responsibility fields, resolution, polarity and stable source references. It could deterministically validate allowed combinations, normalize aliases, preserve evidence, group exact duplicates and reject structurally impossible combinations. It must not promote a proposal to a decision, infer resolution, infer responsibility or perform uncertain semantic merging.
|
||||
|
||||
Richer labels make safe comparisons easier: two `option` records can be recognized as candidates about the same subject without being merged into a `decision`; an `explicit_question + resolved` item is not confused with an unresolved instance. Provenance improves because a later decision can link back to earlier option/proposal evidence rather than overwriting it.
|
||||
|
||||
Semantic merging becomes better informed, but not automatically easy. Equivalence and lifecycle transitions remain semantic work. A future consolidator should preserve all source evidence, distinguish duplicate evidence from state evolution, and mark contradictions or uncertainty rather than collapsing them.
|
||||
|
||||
### Renderer and output policy
|
||||
|
||||
The renderer should not make commitment classification decisions. A
|
||||
policy/view-selection stage should determine protocol-worthy semantic states
|
||||
while preserving their topic associations. In the Version 2 target
|
||||
architecture, the primary protocol renderer should receive topic-oriented
|
||||
semantic knowledge containing eligible evidence and state transitions, not a
|
||||
flat category-oriented collection. It must not silently promote raw proposals
|
||||
to Decisions or present conversational coordination as Meeting Content.
|
||||
|
||||
Proposals may still be visible to a renderer for a view explicitly designed to show them, such as a discussion appendix or editorial drafting view. That access must be typed and intentional; a renderer must never silently format a proposal under “Decisions.” Different output views may select different states while sharing the same Canonical Meeting Knowledge.
|
||||
|
||||
Category-oriented renderers remain valid for secondary views such as Action
|
||||
Item tables, Decision registers and Open Questions lists. This accepted
|
||||
direction supplements rather than removes the current Working Protocol V2
|
||||
architecture; it does not change current rendering behavior.
|
||||
|
||||
### Future BPD-style protocol generation
|
||||
|
||||
Richer labels would give editorial generation better material without granting it authority to invent commitment. A BPD-style view could explain the path from options through tentative agreement to decision, separate current obligations from completed work, and describe why an information need remains open. Proposals and preferred options could be useful appendix material when the requested view values discussion history.
|
||||
|
||||
The editorial layer may condense or order these records, but it must preserve state, polarity, responsibility status and provenance. Appendix inclusion is a view policy, not a semantic promotion.
|
||||
|
||||
## Trade-offs
|
||||
|
||||
Benefits:
|
||||
|
||||
- preserves valid evidence below the final-protocol threshold;
|
||||
- separates meeting semantics from publication policy;
|
||||
- represents established work without inventing an assignment event or owner;
|
||||
- treats question kind and resolution independently;
|
||||
- makes responsibility independently auditable;
|
||||
- improves provenance across lifecycle changes;
|
||||
- prevents renderers from silently converting proposals into commitments;
|
||||
- gives future editorial views useful, explicitly non-final material.
|
||||
|
||||
Costs and risks:
|
||||
|
||||
- expands the extraction schema and Gold specification;
|
||||
- introduces adjacent semantic labels that require precise definitions;
|
||||
- requires consolidation to distinguish duplicates from state transitions;
|
||||
- may retain more evidence records, increasing storage and review volume;
|
||||
- depends on adequate local context to determine resolution and lifecycle state;
|
||||
- requires explicit view policies so consumers do not treat every evidence record as protocol-worthy;
|
||||
- cannot guarantee LLM consistency merely by replacing a binary verdict with a taxonomy.
|
||||
|
||||
The taxonomy should remain closed and small. In particular, modality, confidence, support, commitment state and responsibility must not be collapsed into a single scalar.
|
||||
|
||||
## Migration strategy
|
||||
|
||||
This is a staged design path, not a recommendation to implement it now.
|
||||
|
||||
1. Specify the taxonomy and invariants against the existing BUG-015 Gold candidates and the cited Progeo excerpts, without changing current expected outputs.
|
||||
2. Annotate a design-only mapping from each current Decision, Action Item and Open Question example to evidence state, support and independent attributes. Confirm objectively unique ground truth, especially for `requested_action` versus `possible_next_step` and `tentative_agreement` versus `preferred_option`.
|
||||
3. Define a versioned evidence-record schema and deterministic protocol-selection policy on paper. Keep responsibility and question resolution independent.
|
||||
4. Design evaluation measures separately for state accuracy, resolution accuracy, responsibility accuracy and protocol-policy output. A correct lower state must not count as a false protocol commitment.
|
||||
5. Only after the design is validated, plan an experiment that compares the evidence-first approach with current extraction and the binary verifier. Any prompt or Python changes would be separate, explicitly scoped tasks following the Gold methodology and LLM safety rules.
|
||||
6. If later adopted, preserve compatibility through an adapter that maps protocol-worthy evidence states to the current canonical categories while Canonical Meeting Knowledge evolves. Do not ask the renderer to interpret raw states.
|
||||
|
||||
No database, prompt, code, test or current output-schema migration is proposed by this document.
|
||||
|
||||
## Open questions
|
||||
|
||||
- Is `requested_action` useful as a distinct state when a direct assignment already creates `established_action`, or should it be limited to unaccepted requests?
|
||||
- Does `tentative_agreement` have sufficiently unique evidence in the current corpus, or should it remain an annotation until more Gold examples exist?
|
||||
- Should an ongoing-work item whose relevance to the meeting is unclear pass the Working Protocol policy, or require an explicit continuation/follow-up signal?
|
||||
- How much local context is required to label a question `resolved` safely when the answer occurs across a technical chunk boundary?
|
||||
- Should `resolution_unclear` be retained only in Canonical Meeting Knowledge, or also exposed to a review view?
|
||||
- Should support be limited to `direct / explicit`, treating `indirect` as review-only, to reduce subjective classification?
|
||||
- How should consolidation represent one subject moving from proposal to decision: linked immutable evidence records, or a current-state object with a preserved event history?
|
||||
- Which output views, if any, should include withdrawn proposals, cancelled work and resolved questions?
|
||||
- Can polarity and lifecycle links be normalized deterministically without introducing semantic inference?
|
||||
|
||||
## Recommendation
|
||||
|
||||
The evidence-first architecture should replace the current binary keep/reject classification approach as the target architecture. The replacement should be conceptual and staged, not implemented yet: extract a small, domain-specific semantic state plus independent support, responsibility and resolution attributes; preserve evidence and provenance; then apply explicit policy to produce protocol categories.
|
||||
|
||||
A single common Evidence Strength axis should complement this taxonomy but must not replace it. The decisive improvement is separation of semantic state from protocol eligibility. That separation directly explains the BUG-015 false positives and false negatives, supports the Progeo cases without invented ownership, simplifies view selection, and provides a sounder basis for Canonical Meeting Knowledge and future BPD-style rendering.
|
||||
@@ -0,0 +1,62 @@
|
||||
# Optional Speaker Diarization
|
||||
|
||||
The direct-protocol MVP keeps speaker diarization disabled by default. Enable
|
||||
anonymous Community-1 speaker labels with `--diarization auto`, `gpu`, or
|
||||
`cpu`:
|
||||
|
||||
```bash
|
||||
python3 scripts/run_mvp_meeting.py meeting.wav \
|
||||
--whisper-model /path/to/ggml-model.bin \
|
||||
--diarization auto
|
||||
```
|
||||
|
||||
Native mode (the default runtime) requires a compatible local PyTorch and
|
||||
`pyannote.audio==4.0.7`. For isolated ROCm/CUDA environments, select the
|
||||
container runtime and provide its image and hardware arguments explicitly:
|
||||
|
||||
```bash
|
||||
python3 scripts/run_mvp_meeting.py meeting.wav \
|
||||
--whisper-model /path/to/ggml-model.bin \
|
||||
--diarization gpu \
|
||||
--diarization-runtime container \
|
||||
--diarization-container-image IMAGE \
|
||||
--diarization-container-arg=--device=/dev/kfd \
|
||||
--diarization-container-arg=--device=/dev/dri \
|
||||
--diarization-container-arg=--group-add \
|
||||
--diarization-container-arg=video
|
||||
```
|
||||
|
||||
The container receives `HF_TOKEN` by environment-variable name only. It mounts
|
||||
the source audio and repository read-only and writes diarization artifacts into
|
||||
the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with
|
||||
Python's `wave` module and sends an in-memory tensor to pyannote, avoiding its
|
||||
torchcodec file decoder.
|
||||
|
||||
Anonymous `SPEAKER_XX` labels are aligned to Whisper segments by maximum
|
||||
temporal overlap with Community-1 exclusive diarization. The original Whisper
|
||||
transcript is preserved; the derived transcript under `diarization/` is used as
|
||||
the direct-protocol generator's source input.
|
||||
|
||||
The full diarized JSON and timestamped text remain immutable audit artifacts,
|
||||
but their per-segment formatting is too verbose for a full-meeting LLM prompt:
|
||||
timestamps and repeated speaker labels can more than double input size. For
|
||||
protocol generation, Meeting Lab deterministically groups only adjacent
|
||||
segments assigned to the same anonymous speaker and omits timestamps. A later
|
||||
return by the same speaker starts a new block, and unassigned segments remain
|
||||
under `SPEAKER_UNASSIGNED`. `protocol/transcript_input.txt` preserves the exact
|
||||
derived representation sent to prompt construction.
|
||||
|
||||
Before contacting Ollama, Meeting Lab conservatively estimates prompt tokens
|
||||
from UTF-8 byte count without adding a model tokenizer dependency. The default safe
|
||||
budget is 29,000 estimated tokens within the explicitly configured 32,768-token
|
||||
Ollama context. The estimate is calibrated against the currently validated
|
||||
German BPD input and is configurable through
|
||||
`MvpMeetingConfig.protocol_safe_input_token_budget` or
|
||||
`--protocol-safe-input-token-budget`.
|
||||
|
||||
If compact diarized input exceeds the budget, the generator deterministically
|
||||
uses the complete plain segment transcript and records the fallback. If that
|
||||
also exceeds the budget, generation fails before model lookup or generation;
|
||||
it never truncates, chunks, summarizes, retries, or makes multiple protocol
|
||||
calls implicitly. Full diarization artifacts are never overwritten by this
|
||||
selection.
|
||||
@@ -1350,3 +1350,912 @@ accordance with the Gold Standard methodology.
|
||||
|
||||
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
|
||||
remains Open; no phrase-specific deterministic filter was introduced.
|
||||
|
||||
## EXP-0027 — Evidence-near observation extraction
|
||||
|
||||
Date: 2026-08-18
|
||||
|
||||
Hypothesis: `qwen3.5:9B` can more reliably extract evidence-near linguistic and
|
||||
semantic properties than directly synthesize protocol-level events, outcomes,
|
||||
actions and unresolved issues. This isolated experiment stops before semantic
|
||||
interpretation and does not connect to the production pipeline.
|
||||
|
||||
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
|
||||
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
|
||||
source evidence, explicit targets, a five-value relation vocabulary, modality,
|
||||
temporality, evaluation, agreement, responsibility/person, uncertainty,
|
||||
clarification need and free-text scope. It contains no protocol-level category
|
||||
field. The validator requires sequential observation IDs, known evidence IDs,
|
||||
backward-only valid observation targets, closed categorical vocabularies,
|
||||
consistent responsibility/person pairs, and one canonical absence form: JSON
|
||||
null for person and `absent` for scope.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`, no retries. All nine cases ran exactly once, for nine LLM
|
||||
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
|
||||
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
|
||||
responses, parsed observations, validation results, Ollama metadata and
|
||||
comparisons are preserved under
|
||||
`/tmp/meeting-lab-evidence-observations-v1-20260818/`.
|
||||
|
||||
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
|
||||
Five cases failed structure because the model represented a single target as a
|
||||
one-element list, usually `["discussion_subject"]`; the accepted schema permits
|
||||
a list only for two or more jointly referenced observations. Several responses
|
||||
also copied the relation label `limits_scope` into the free-text scope field.
|
||||
These were systematic model-output errors, not transport or parser failures.
|
||||
The prompt and run were not retried or tuned.
|
||||
|
||||
Human semantic review of the preserved raw responses:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
|
||||
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
|
||||
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
|
||||
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
|
||||
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
|
||||
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
|
||||
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
|
||||
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
|
||||
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
|
||||
|
||||
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override
|
||||
strict structural failures; it separates useful semantic signal from schema
|
||||
compliance.
|
||||
|
||||
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving
|
||||
closer to evidence reduced some direct promotion behavior: the geometry mention
|
||||
did not become work, the washing disadvantage did not become an unresolved
|
||||
issue, both alternatives in B remained uncommitted, and the publication query
|
||||
did not become an assignment. However, the important promotion errors did not
|
||||
disappear. C acquired unsupported acceptance, and G still promoted a personal
|
||||
preference into rejection. Positive cases were only partly preserved: explicit
|
||||
rejection was recognized but incorrectly linked; trial-only language was kept
|
||||
but responsibility was invented; Nina's request and commitment were recognized
|
||||
but responsibility states were wrong; and the publication issue was recognized
|
||||
but its uncertainty was contradicted by rejection.
|
||||
|
||||
Result: **B — evidence-near extraction is promising, but specific observation
|
||||
dimensions remain unreliable.** Target/relation selection, scope attachment,
|
||||
responsibility state/person attribution, and agreement versus uncertainty are
|
||||
not reliable enough to justify designing the later interpretation stage yet.
|
||||
No production integration or later interpretation stage was implemented.
|
||||
|
||||
## EXP-0028 — Evidence-Near Observation Extraction V2
|
||||
|
||||
Date: 2026-08-19
|
||||
|
||||
V2 tested whether `qwen3.5:9B` preserves the evidence needed by a later
|
||||
controlled interpretation stage when direct responsibility, agreement and
|
||||
semantic graph relations are removed. Responsibility was replaced by explicit
|
||||
participant/discourse facts (`speaker`, `named_person`, `addressee`, singular
|
||||
self-reference, collective `we`, and impersonal person reference). Agreement
|
||||
was replaced by explicit affirmation, explicit negation and determination
|
||||
statement signals. Graph relations were reduced to nullable scalar
|
||||
`refers_to`; scope became free-text `qualifier` plus nullable scalar
|
||||
`limits_target`. No later derivation stage was implemented.
|
||||
|
||||
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
|
||||
human interpretations. It contains no responsibility, agreement, action,
|
||||
decision, open-question, accepted-trial, rejected-alternative or protocol
|
||||
eligibility fields. Validation enforces known evidence IDs, sequential unique
|
||||
observation IDs, backward-only scalar references, closed vocabularies, boolean
|
||||
participant flags, JSON-nullable participant/qualifier/reference fields and no
|
||||
string `"null"`.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`, no retries or voting. A launch-path defect was corrected
|
||||
before the live run; the failed launch made zero model calls. A sandbox-blocked
|
||||
localhost attempt also made zero model calls. The completed run called the
|
||||
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
|
||||
Persistent prompts, Gold input and expectations, raw and parsed model output,
|
||||
validation, automatic comparison, Ollama metadata and human evaluation are in
|
||||
`artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/`.
|
||||
|
||||
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
|
||||
were schema-invalid. The dominant serialization pattern was use of `present`
|
||||
instead of the specified `explicit` for affirmation/negation; E additionally
|
||||
used `none` instead of `absent` for a determination signal, while D emitted the
|
||||
separate uncertainty concept as an invalid modality. These errors are
|
||||
contract violations, although most `present`/`explicit` differences are
|
||||
deterministically normalizable without changing meaning. A and C were valid
|
||||
JSON/schema outputs but had critical semantic mismatches.
|
||||
|
||||
Human semantic review:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
|
||||
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; `present` caused schema failure. |
|
||||
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
|
||||
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
|
||||
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
|
||||
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
|
||||
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
|
||||
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
|
||||
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
|
||||
|
||||
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced
|
||||
V1 promotion errors: collective speech and speaker identity no longer became
|
||||
individual responsibility; personal preference no longer became a group-level
|
||||
rejection field; an information question did not become work; and explicit
|
||||
negation/affirmation survived as separate evidence. Useful participant evidence
|
||||
also survived strongly in H and collective-speech evidence in F.
|
||||
|
||||
Simplification did not make all evidence-near dimensions reliable. Scalar
|
||||
references and `limits_target` were almost entirely omitted, qualifiers were
|
||||
usually omitted, committed modality was missed in E, F and H, and C/G repeated
|
||||
important modality, uncertainty and participant-feature errors. Some positive
|
||||
semantic information therefore survived only in free-text `content`, not in
|
||||
the structural signals a controlled derivation stage would need.
|
||||
|
||||
Result: **B — V2 is materially better, but specific evidence-near dimensions
|
||||
still require refinement.** Direct responsibility, agreement and graph-relation
|
||||
classification should remain excluded. Before designing the derivation stage,
|
||||
the next work should examine the minimal reliable representation of explicit
|
||||
reference/scope limitation, commitment modality and participant deixis. No
|
||||
production integration, Progeo run or derivation implementation was performed.
|
||||
|
||||
## EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
|
||||
|
||||
Date: 2026-08-19
|
||||
|
||||
Hypothesis: `qwen3.5:9B` is substantially more reliable when the first semantic
|
||||
stage preserves meeting meaning as atomic natural-language observations with
|
||||
provenance and only simple participant information, without classifying or
|
||||
deriving higher-level meeting semantics.
|
||||
|
||||
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
|
||||
observation contains exactly `observation_id`, `evidence_id`, `content`,
|
||||
`speaker`, nullable `named_person`, and nullable `addressee`. It contains no
|
||||
modality, temporality, evaluation, affirmation, negation, determination,
|
||||
uncertainty, clarification, responsibility, agreement, relation, reference,
|
||||
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
|
||||
Instead, the prompt asks for conservative atomic content that retains hedges,
|
||||
conditions, personal/collective/impersonal language, requests, acceptances,
|
||||
rejections, quantities, deadlines and boundaries in natural language.
|
||||
|
||||
Structural validation is intentionally small: exact schema keys, non-empty
|
||||
observations/content, unique `obs_N` identifiers, known evidence IDs, speaker
|
||||
matching its evidence, explicit named people/addressees, and no string
|
||||
`"null"`. Human semantic preservation against per-case requirements is the
|
||||
primary evaluation; wording differences do not fail a case.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`, no retries, voting or per-case tuning. One sandbox-blocked
|
||||
localhost launch made zero model calls. The completed run made exactly nine
|
||||
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
|
||||
structural validation. Persistent source evidence, semantic requirements,
|
||||
exact prompts, raw and parsed responses, validation, Ollama metadata and human
|
||||
evaluation are stored under
|
||||
`artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/`.
|
||||
|
||||
Human semantic preservation results:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PASS | Preserved `kann`, `vielleicht`, tentative follow-up, and explicit `Dann` sequence without commitment. |
|
||||
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
|
||||
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
|
||||
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
|
||||
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
|
||||
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
|
||||
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, `wenn überhaupt`, impersonal checking need and no decision, but dropped collective `wir` from who would receive contaminated material. |
|
||||
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
|
||||
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
|
||||
|
||||
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
|
||||
“that we receive contaminated material back” into an impersonal passive phrase;
|
||||
the risk itself remained hypothetical. H translated `Freitag` to `Friday`, a
|
||||
harmless wording difference. I retained two compound observations rather than
|
||||
splitting every proposition, but all required semantic boundaries and
|
||||
dependencies remained explicit.
|
||||
|
||||
Compared with V2, categorical-field removal improved content preservation in
|
||||
A, G and I: A retained `Dann`; G retained `wenn überhaupt`, personal `Ich` and
|
||||
impersonal `Man`; I retained publication purpose and all boundaries. It also
|
||||
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
|
||||
strengthening into responsibility, group rejection, established work or
|
||||
assigned clarification. F and H remain sufficiently complete in natural
|
||||
language for a later interpretation experiment. No useful meaning was shown to
|
||||
depend on the removed fields; the V3 content retained the useful signals that
|
||||
V2's fields had attempted to encode.
|
||||
|
||||
Result: **A — MINIMAL FIRST STAGE ACCEPTED.** On A-I, minimal atomic content
|
||||
with evidence provenance and simple participants is sufficiently reliable to
|
||||
be the candidate first semantic stage. A later bounded experiment may examine
|
||||
controlled semantic interpretation, but no derivation stage, production
|
||||
integration or Progeo run was implemented here.
|
||||
|
||||
## EXP-0030 — V3 model comparison: qwen3.5:9B vs qwen3.6:35B-A3B
|
||||
|
||||
Date: 2026-08-19
|
||||
|
||||
This controlled comparison reran the accepted V3 minimal semantic-preservation
|
||||
experiment unchanged with the locally installed `qwen3.6:35B-A3B`. It used the
|
||||
same implementation, A-I fixture, evidence, prompt, minimal schema, temperature
|
||||
0, `think=false`, `num_ctx=16384`, `num_predict=4096`, no retries, no voting and
|
||||
one call per case. The completed run made exactly nine calls in 96.392 seconds.
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/evidence_observations_v3/20260819_v3_qwen36_35b_a3b_single_run/`.
|
||||
|
||||
Structural validation passed 8/9 cases. D was semantically faithful but invalid
|
||||
because the model copied transcript speakers Antonius and Martin into
|
||||
`named_person`, although those names were not explicitly named within their
|
||||
utterances. Human semantic preservation produced 7 PASS, 1 PARTIAL and 1 FAIL:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PASS | Preserved `can`, `perhaps`, tentative `would`, explicit `then` and no commitment, but changed the content language to English. |
|
||||
| B | PARTIAL | Preserved both approaches overall, but removed `Oder` from the 20-20 observation and locally strengthened it into collective planned conduct. |
|
||||
| C | PASS | Preserved Tim's tentative personal Textor contact and possible follow-up without established work. |
|
||||
| D | PASS | Preserved possibility, uncertainty, process and energy consequence; structural failure was confined to invalid speaker-as-named-person values. |
|
||||
| E | PASS | Preserved cost, collective rejection/non-pursuit and later determination, though the confirmation dropped explicit `Ja`. |
|
||||
| F | FAIL | Preserved quantity, timing and trial/series boundaries, but changed collective `wir` into “Martin suggests” and “Tim agrees,” inventing individual proposal/agreement meaning. |
|
||||
| G | PASS | Preserved hypothetical risk, collective recipient `wir`, Martin's personal stance, `wenn überhaupt`, conditional Technikum, impersonal `man müsste` and no decision/owner. |
|
||||
| H | PASS | Preserved request, addressee, Friday, explicit acceptance and future personal commitment without a responsibility field; content was English. |
|
||||
| I | PASS | Preserved local/pure-production and comparison boundaries, five-degree difference, upstream effort, publication purpose, unresolved permission and clarification without assignment. |
|
||||
|
||||
Direct comparison:
|
||||
|
||||
| Measure | `qwen3.5:9B` | `qwen3.6:35B-A3B` |
|
||||
| --- | ---: | ---: |
|
||||
| Structurally valid | 9/9 | 8/9 |
|
||||
| Human PASS | 8 | 7 |
|
||||
| Human PARTIAL | 1 | 1 |
|
||||
| Human FAIL | 0 | 1 |
|
||||
| Observations | 28 | 30 |
|
||||
| Runtime | 40.074 s | 96.392 s |
|
||||
| LLM calls | 9 | 9 |
|
||||
| Prompt-evaluation tokens | 6,268 | 6,268 |
|
||||
| Evaluation tokens | 2,590 | 2,718 |
|
||||
|
||||
The larger model fixed the 9B weakness in G by preserving collective `wir`, and
|
||||
it split I's compound production/effort observations more cleanly. Those gains
|
||||
did not offset regressions: B was locally strengthened, F materially converted
|
||||
collective conduct into individual agreement, D violated the participant
|
||||
schema, observation count increased, and runtime was 2.4 times higher. Both
|
||||
models preserved German consistently in six of nine cases, but in different
|
||||
cases; the 35B-A3B model changed A, F and H to English, while 9B changed C, F
|
||||
and H wholly or partly to English.
|
||||
|
||||
Result: **D — REGRESSION.** `qwen3.6:35B-A3B` does not materially improve the
|
||||
accepted minimal V3 first-stage preservation over `qwen3.5:9B`; it is worse on
|
||||
the A-I comparison because of the F ownership-adjacent strengthening and lower
|
||||
structural validity. This conclusion applies only to the minimal V3 first
|
||||
stage and does not determine model choice for any later semantic derivation.
|
||||
No production integration, derivation implementation or Progeo run occurred.
|
||||
|
||||
## EXP-0031 — Controlled Semantic Derivation H V0
|
||||
|
||||
Date: 2026-08-19
|
||||
|
||||
This isolated experiment tested the first controlled second-stage derivation
|
||||
using only the accepted `qwen3.5:9B` V3 observations for case H. The derivation
|
||||
LLM received the two V3 observations, not the transcript or Gold expectation.
|
||||
Its deliberately narrow task was limited to recognizing whether `obs_1` is a
|
||||
concrete request and whether `obs_2` explicitly commits its speaker to
|
||||
substantially the same work. Its strict output schema forbids responsibility,
|
||||
requested actor, establishment/status, Action Item, protocol, confidence and
|
||||
generic relation/graph fields.
|
||||
|
||||
Deterministic code validates observation/evidence provenance, obtains the
|
||||
requested actor only from the request observation's addressee, requires the
|
||||
acceptance to follow the request, requires the accepting speaker to equal that
|
||||
addressee, and establishes responsibility only after all semantic and
|
||||
structural gates pass. A bounded weekday normalizer reconciles `Friday` and
|
||||
`Freitag`, rejects conflicting weekdays, and separates the supported due date
|
||||
from the normalized action text. No general temporal or action ontology was
|
||||
introduced.
|
||||
|
||||
Twenty focused deterministic tests cover the positive H path and the required
|
||||
negative invariants: request alone, acknowledgement/non-commitment, tentative
|
||||
acceptance, different response speaker, different work, reversed order,
|
||||
speaker/name/addressee alone, conflicting deadlines, unknown observation IDs,
|
||||
inconsistent evidence provenance, forbidden semantic fields, malformed JSON
|
||||
and persistent artifacts. The complete non-LLM suite passed 192/192.
|
||||
|
||||
Configuration: one `qwen3.5:9B` call, temperature 0, `think=false`,
|
||||
`num_ctx=16384`, `num_predict=1024`, no retries or voting. The call took 11.765
|
||||
seconds, with 469 prompt-evaluation and 124 evaluation tokens. The model
|
||||
returned a valid recognition object: `obs_1` is a concrete request, `obs_2` is
|
||||
an explicit commitment, and both concern substantially the same work. It
|
||||
returned no responsibility or establishment judgment.
|
||||
|
||||
All deterministic gates passed. The final derived result is an established
|
||||
action `Prüfung der Messdaten`, requested from and assigned to Nina, due
|
||||
`Freitag`, supported by request `obs_1/e1` and acceptance `obs_2/e2`. The model
|
||||
included `bis Friday` in its normalized request text; after the single call, a
|
||||
deterministic-only bounded correction separated that already-recognized due
|
||||
phrase from action content without changing the prompt, recognition schema,
|
||||
semantic result or call count. Focused and complete non-LLM suites still
|
||||
passed after this correction.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/controlled_semantic_derivation_h/20260819_h_qwen35_9b_single_run/`.
|
||||
Result: the H mechanism succeeded. This establishes only that the narrow
|
||||
request-plus-explicit-acceptance pattern can be recognized and gated for H; it
|
||||
does not generalize the derivation architecture to other cases or semantic
|
||||
categories. No production integration, other case run, semantic graph,
|
||||
protocol derivation or Progeo run occurred.
|
||||
|
||||
## EXP-0032 — Request / Acceptance Gold V0
|
||||
|
||||
Status: Experimental; promising with semantic precision gaps
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
This isolated regression experiment tested whether the EXP-0031 mechanism
|
||||
generalizes beyond H. It used ten short synthetic cases containing only
|
||||
V3-style observations. Evidence Observation V3 was neither called nor changed,
|
||||
and the model received no raw transcript or expected result. The fixed
|
||||
recognition schema permits only a nullable concrete request and nullable later
|
||||
explicit personal commitment, plus the same-requested-work judgment and
|
||||
normalized action text. Responsibility, requested actor, established status,
|
||||
Action Item, protocol, confidence and generic graph fields remain forbidden.
|
||||
|
||||
Cases:
|
||||
|
||||
- RA-01 explicit positive acceptance: PASS.
|
||||
- RA-02 paraphrased positive acceptance: PASS.
|
||||
- RA-03 acknowledgement only: PASS.
|
||||
- RA-04 tentative response: PASS.
|
||||
- RA-05 different responder without personal acceptance: PASS.
|
||||
- RA-06 explicit commitment to different work: PARTIAL. The model returned no
|
||||
acceptance instead of recognizing a commitment with `same_requested_work`
|
||||
false. The requested action correctly remained unestablished.
|
||||
- RA-07 request without response: PASS.
|
||||
- RA-08 collective commitment: PARTIAL. The model over-recognized the
|
||||
collective `wir` statement as an explicit commitment, but no request existed
|
||||
and deterministic gates prevented individual responsibility.
|
||||
- RA-09 impersonal necessity: PARTIAL. The model over-recognized the impersonal
|
||||
necessity as a concrete request, but the observation had no addressee and
|
||||
deterministic gates prevented establishment.
|
||||
- RA-10 tentative personal suggestion: PASS.
|
||||
|
||||
Configuration: exactly ten sequential `qwen3.5:9B` calls, one per case,
|
||||
temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no retries,
|
||||
no voting and no prompt change between cases. Summed call time was 23.754
|
||||
seconds, with 4,826 prompt-evaluation tokens and 793 evaluation tokens. The
|
||||
strict schema validated every response and no responsibility or establishment
|
||||
field leaked into model output.
|
||||
|
||||
Both positive cases recognized the request, explicit commitment and same-work
|
||||
relationship, including the paraphrased acceptance, and deterministically
|
||||
established Clara as responsible with due date `Dienstag`. The model rendered
|
||||
the normalized action in semantically equivalent English; evaluation therefore
|
||||
checks the structural deterministic result exactly while treating normalized
|
||||
action wording as evidence-near semantic text rather than requiring lexical
|
||||
identity. Acknowledgement and tentative response were not promoted. Every
|
||||
negative case remained unestablished, and no individual responsibility was
|
||||
invented.
|
||||
|
||||
Recognition-level errors were two false positives (RA-08 commitment and RA-09
|
||||
request) and one false negative (RA-06 different-work commitment). Final
|
||||
established-action false positives and false negatives were both zero. The
|
||||
overall result was seven PASS, three PARTIAL and zero FAIL.
|
||||
|
||||
Conclusion: the narrow request-plus-acceptance architecture remains promising
|
||||
for established individual actions because deterministic addressee, ordering,
|
||||
speaker, same-work, provenance and deadline gates contained all recognition
|
||||
errors. The recognition layer is not yet precise enough to generalize: its
|
||||
handling of collective commitment, impersonal necessity and commitments to
|
||||
different work needs further isolated study. No production integration or
|
||||
additional semantic category is justified by this result.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/request_acceptance_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0033 — Collective Commitment Gold V0
|
||||
|
||||
Status: Experimental; architecturally successful with one contained
|
||||
recognition false positive
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
This isolated second-stage experiment tested whether an explicit collective
|
||||
first-person commitment can establish an action without inventing an individual
|
||||
owner. It used ten synthetic cases containing one minimal V3-style observation
|
||||
each. Evidence Observation V3 was neither called nor changed, and the accepted
|
||||
Request/Acceptance mechanism remained unchanged and independent.
|
||||
|
||||
The strict semantic schema contains exactly `observation_id`,
|
||||
`commitment_form` and `normalized_action_text`. `commitment_form` is closed to
|
||||
`individual_first_person`, `collective_first_person` and `none`. The model
|
||||
cannot output responsibility, ownership, requested actor, establishment,
|
||||
Action Item, protocol, confidence, relations, graphs, decisions or unresolved
|
||||
issues. Deterministic code validates schema and provenance, requires collective
|
||||
commitment plus non-empty action text, applies bounded deadline consistency and
|
||||
explicit-negation gates, and only then sets `status: established`,
|
||||
`commitment_scope: collective` and `responsible_person: null`.
|
||||
|
||||
Gold results:
|
||||
|
||||
- CC-01 explicit collective commitment: PASS; established, due `nächste
|
||||
Woche`, no person.
|
||||
- CC-02 individual commitment: PASS; correctly routed out of the collective
|
||||
path.
|
||||
- CC-03 tentative collective possibility: PASS; unestablished.
|
||||
- CC-04 collective suggestion: PASS; unestablished.
|
||||
- CC-05 impersonal necessity: PASS; unestablished.
|
||||
- CC-06 passive future statement: PASS; unestablished.
|
||||
- CC-07 collective rejection: PARTIAL. The model incorrectly returned
|
||||
`collective_first_person`, but the deterministic negation gate detected
|
||||
`nicht` and prevented establishment.
|
||||
- CC-08 qualified collective commitment: PASS; established with `nur im
|
||||
Technikum` preserved, null due and no person.
|
||||
- CC-09 collective commitment without deadline: PASS; established with null
|
||||
due and no person.
|
||||
- CC-10 speaker ownership trap: PASS; established collectively while Martin
|
||||
remained only the speaker and was not assigned ownership.
|
||||
|
||||
Configuration: exactly ten successful sequential `qwen3.5:9B` calls, one per
|
||||
case, temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no
|
||||
retries, no voting and no prompt change. There were zero technical failed
|
||||
calls. Aggregate runner time was 10.504 seconds; summed per-call time was 10.500
|
||||
seconds, with 4,267 prompt-evaluation tokens and 415 evaluation tokens.
|
||||
|
||||
The outcome was nine PASS, one PARTIAL and zero FAIL. There was one recognition
|
||||
false positive and no recognition false negatives. No qualifier was lost, no
|
||||
individual owner was invented, and no responsibility or status field leaked
|
||||
into recognition. Bounded due handling preserved `nächste Woche` verbatim and
|
||||
returned null when no deadline was present.
|
||||
|
||||
Conclusion: the collective-commitment path is architecturally successful for
|
||||
this narrow Gold set. The deterministic negation gate contained the only model
|
||||
error, and every successful collective result necessarily retained
|
||||
`responsible_person: null`. This does not justify a generic commitment system,
|
||||
production integration, group identity inference or another semantic category.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0034 — Explicit Rejection Gold V0
|
||||
|
||||
Status: Failed architecturally
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
This isolated Stage-2 experiment tested the narrow evidence fact that a
|
||||
concrete action, option, proposal or future course was explicitly rejected,
|
||||
abandoned, discontinued or ruled out. It used twelve synthetic cases containing
|
||||
one self-contained observation or one local target/rejection pair. Evidence
|
||||
Observation V3 was not called or changed. The accepted Request/Acceptance and
|
||||
Collective Commitment paths remained unchanged and were not invoked.
|
||||
|
||||
The strict semantic schema contains exactly `rejection_observation_id`,
|
||||
`target_observation_id`, `rejection_form` and
|
||||
`normalized_rejected_action_text`. `rejection_form` is closed to
|
||||
`explicit_action_rejection` and `none`. A positive recognition requires a
|
||||
known local target and non-empty normalized target; `none` requires both target
|
||||
and normalized text to be null. Decision, outcome, topic-closure,
|
||||
responsibility, ownership, protocol, confidence and graph fields are forbidden.
|
||||
Target resolution is limited to the same observation or one earlier supplied
|
||||
observation. Deterministic code validates schema, IDs, ordering and complete
|
||||
provenance before emitting the narrow status `explicitly_rejected`.
|
||||
|
||||
`explicitly_rejected` means rejected by the cited evidence only. It is not yet
|
||||
a final meeting decision or final topic outcome, does not close a topic, and
|
||||
does not supersede an earlier commitment.
|
||||
|
||||
Gold results:
|
||||
|
||||
- RJ-01 explicit collective rejection with local target: PASS.
|
||||
- RJ-02 explicit non-pursuit with paired target: PASS.
|
||||
- RJ-03 self-contained collaboration rejection: FAIL. The model returned
|
||||
`none`, producing one recognition false negative.
|
||||
- RJ-04 personal preference: FAIL. The model promoted the preference to an
|
||||
explicit rejection and derived an unsupported rejection.
|
||||
- RJ-05 concern: PASS; remained a non-rejection.
|
||||
- RJ-06 uncertainty: PASS; remained a non-rejection.
|
||||
- RJ-07 negative recommendation: FAIL. The model promoted advice to an
|
||||
explicit rejection and derived an unsupported rejection.
|
||||
- RJ-08 deferral: PASS; remained a non-rejection.
|
||||
- RJ-09 factual negation: PASS; remained a non-rejection.
|
||||
- RJ-10 temporary non-action: FAIL. The model treated `erstmal noch nicht` as
|
||||
abandonment and derived an unsupported rejection.
|
||||
- RJ-11 explicit rejection with material scope: PASS. Real-plant and
|
||||
Druckversuch scope were preserved.
|
||||
- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option
|
||||
was rejected; the Technikum alternative was not absorbed.
|
||||
|
||||
Configuration: exactly twelve successful sequential `qwen3.5:9B` calls, one
|
||||
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
||||
technical failed calls. Aggregate runner time was 15.518 seconds; summed
|
||||
per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681
|
||||
evaluation tokens.
|
||||
|
||||
The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced
|
||||
three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03).
|
||||
There were four strict target-field expectation mismatches: three were
|
||||
consequences of false-positive rejection objects populating otherwise locally
|
||||
correct antecedents, and one was the missing self-contained RJ-03 target. No
|
||||
derived positive selected the wrong concrete antecedent. Qualifier-loss count
|
||||
was zero, positive-alternative absorption count was zero, and no responsibility,
|
||||
decision, outcome or topic-closure field leaked into model output.
|
||||
|
||||
Conclusion: the experiment is not architecturally successful. Deterministic
|
||||
structural gates cannot contain a semantically well-formed false-positive
|
||||
rejection with valid local target and provenance. The model did distinguish
|
||||
concern, uncertainty, deferral and factual negation, and it handled scoped and
|
||||
alternative-bearing positives correctly, but it did not reliably separate
|
||||
explicit rejection from personal preference, advice or temporary non-action.
|
||||
The current binary recognition `explicit_action_rejection | none` is
|
||||
insufficient for reliable generalization.
|
||||
No production integration, generic rejection system, prompt tuning or
|
||||
cross-pattern reconciliation is justified.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0035 — Negative Act Form V0
|
||||
|
||||
Status: Experimental; successful for form classification with normalization
|
||||
limitations
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
EXP-0034 failed because the binary `explicit_action_rejection | none` question
|
||||
collapsed materially different negative acts. It missed self-contained
|
||||
non-pursuit and promoted personal preference, recommendation and temporary
|
||||
non-action to rejection. This isolated follow-up tested only whether those
|
||||
evidence-near forms can be distinguished before any normative derivation. It
|
||||
does not derive rejection, decision, outcome, topic closure, responsibility or
|
||||
protocol status, and EXP-0034 remained unchanged.
|
||||
|
||||
The strict output schema contains exactly `observation_id`,
|
||||
`negative_act_form` and `normalized_action_text`. The closed form vocabulary is
|
||||
`explicit_non_pursuit`, `personal_preference`, `recommendation`,
|
||||
`temporary_non_action` and `none`. Non-`none` forms require non-empty normalized
|
||||
action text; `none` requires null. Rejection, status, decision, outcome,
|
||||
responsibility and other normative fields are forbidden recursively. Local
|
||||
context may resolve a candidate observation's pronoun, but the schema contains
|
||||
no target relation and the experiment exposes no derivation function.
|
||||
|
||||
Gold results:
|
||||
|
||||
- NA-01 explicit non-pursuit: PARTIAL. The form was correct; `working with Dr.
|
||||
Schlummer` omitted the continuation aspect from normalization.
|
||||
- NA-02 paraphrased explicit non-pursuit: PASS.
|
||||
- NA-03 personal preference: PARTIAL. The form was correct, but normalization
|
||||
repeated `Ich würde das nicht machen` instead of resolving the real-plant
|
||||
trial target.
|
||||
- NA-04 negative recommendation: PARTIAL. The form was correct; the normalized
|
||||
English action used the loose rendering `real asset` for `reale Anlage`.
|
||||
- NA-05 temporary non-action: PASS.
|
||||
- NA-06 concern only: PASS with `none` and null action text.
|
||||
- NA-07 uncertainty: PASS with `none` and null action text.
|
||||
- NA-08 factual negation: PASS with `none` and null action text.
|
||||
|
||||
Expected-versus-actual form confusion was entirely diagonal:
|
||||
|
||||
| Expected form | Actual form | Count |
|
||||
| --- | --- | ---: |
|
||||
| `explicit_non_pursuit` | `explicit_non_pursuit` | 2 |
|
||||
| `personal_preference` | `personal_preference` | 1 |
|
||||
| `recommendation` | `recommendation` | 1 |
|
||||
| `temporary_non_action` | `temporary_non_action` | 1 |
|
||||
| `none` | `none` | 3 |
|
||||
|
||||
Configuration: exactly eight successful sequential `qwen3.5:9B` calls, one
|
||||
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
||||
technical failures. Aggregate runner time was 8.688 seconds; summed per-call
|
||||
time was 8.686 seconds, with 4,183 prompt-evaluation tokens and 310 evaluation
|
||||
tokens.
|
||||
|
||||
The result was five PASS, three PARTIAL and zero FAIL. All eight
|
||||
`negative_act_form` classifications matched Gold. There was no unsupported
|
||||
semantic strengthening and no rejection, status, decision, outcome,
|
||||
responsibility or topic-closure leakage. Normalized action meaning was fully
|
||||
acceptable in five cases and imperfect in three.
|
||||
|
||||
Conclusion: the finer evidence-near form vocabulary successfully distinguished
|
||||
the four semantic boundaries that defeated the binary rejection experiment in
|
||||
this small Gold set. The result supports separating negative-act-form
|
||||
recognition from later normative derivation, but local target normalization is
|
||||
not yet uniformly reliable. It does not justify modifying EXP-0034, deriving
|
||||
rejection, production integration or beginning cross-pattern reconciliation.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0036 — Controlled Rejection Derivation V1
|
||||
|
||||
Status: Experimental; architecturally unsuccessful
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
This isolated experiment followed the failed binary rejection baseline
|
||||
(EXP-0034) and successful Negative Act Form classification (EXP-0035). Its V1
|
||||
hypothesis was to classify the negative act first, resolve its local target in
|
||||
a separate semantic call, and only then derive `explicitly_rejected`
|
||||
deterministically. It did not modify either predecessor or any accepted Stage-2
|
||||
pattern, and it has no production integration.
|
||||
|
||||
The target recognizer emitted exactly `candidate_observation_id`,
|
||||
`target_observation_id`, and `normalized_target_text`. Only
|
||||
`explicit_non_pursuit` was deterministically eligible. Personal preference,
|
||||
recommendation, temporary non-action, and `none` could never derive rejection,
|
||||
even with a valid target. Provenance, local membership, ordering, non-empty
|
||||
target text, and strict non-normative output were additional gates.
|
||||
`explicitly_rejected` means rejected by the cited evidence only, not a final
|
||||
decision, topic outcome, permanent state, or closure.
|
||||
|
||||
The run reused five exact accepted Negative Act Form outputs and made three new
|
||||
Negative Act calls plus eight target-resolution calls. All calls used
|
||||
`qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=1024`, no retries, voting, or prompt changes.
|
||||
|
||||
| Case | Negative Act expected / actual | Target result | Verdict |
|
||||
| --- | --- | --- | --- |
|
||||
| CR-01 | `explicit_non_pursuit` / same | Model returned the string `"null"` as an unknown ID; self-contained target was not linked | FAIL |
|
||||
| CR-02 | `explicit_non_pursuit` / same | `obs_1`, external solution and continuation preserved | PASS |
|
||||
| CR-03 | `personal_preference` / same | `obs_1`; eligibility gate prevented rejection | PASS |
|
||||
| CR-04 | `recommendation` / same | `obs_1`; eligibility gate prevented rejection | PASS |
|
||||
| CR-05 | `temporary_non_action` / same | `obs_1`; eligibility gate prevented rejection | PASS |
|
||||
| CR-06 | `none` / same | Null target; final non-rejection was correct, but expected local target was unresolved | FAIL |
|
||||
| CR-07 | `explicit_non_pursuit` / same | `obs_1`; real-plant and pressure-test scope survived, but normalization remained proposition-like | PARTIAL |
|
||||
| CR-08 | `explicit_non_pursuit` / same | `obs_1`; real-plant scope preserved and Technikum alternative excluded | PASS |
|
||||
|
||||
Result: five PASS, one PARTIAL, two FAIL. All eight Negative Act forms were
|
||||
correct. There were no false-positive rejections: the valid targets in CR-03,
|
||||
CR-04, and CR-05 could not override their ineligible forms. There was one
|
||||
false-negative rejection, CR-01, caused by invalid target output. Target
|
||||
resolution missed two expected links (invalid CR-01 and null CR-06), so the
|
||||
wrong/unresolved-target count was two. CR-06 exposed a strategy flaw: a
|
||||
non-eligible Negative Act form should not be required to pass target resolution
|
||||
when it cannot derive rejection. Qualifier-loss count was zero. CR-08 isolated
|
||||
the positive alternative successfully. No individual owner, responsibility,
|
||||
decision, outcome, topic-closure, or LLM-emitted rejection status appeared.
|
||||
|
||||
There were 3 new Negative Act calls, 8 target calls, 5 accepted classification
|
||||
reuses, zero technical call failures, and one structural target-validation
|
||||
failure. Aggregate runner time was 11.951 seconds.
|
||||
|
||||
Conclusion: negative-act-form gating is promising and successfully contains
|
||||
the semantic false positives that defeated EXP-0034, but the experiment is not
|
||||
architecturally successful. The current target-resolution strategy failed the
|
||||
required self-contained positive CR-01 and unnecessarily evaluated the
|
||||
ineligible CR-06 path; it is not reliable enough for rejection derivation.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/controlled_rejection_v1/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0037 — Target Resolution V0
|
||||
|
||||
Status: Experimental; FAILED for target resolution
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
Controlled Rejection V1 showed that fine-grained Negative Act Form eligibility
|
||||
contained false-positive rejection, but its target strategy failed a
|
||||
self-contained positive and unnecessarily resolved a target for an ineligible
|
||||
`none` form. This isolated experiment tested target resolution only. It
|
||||
contains no rejection derivation, status, decision, outcome, responsibility,
|
||||
topic closure, or production integration.
|
||||
|
||||
Eligibility was deterministic: only `explicit_non_pursuit` could reach the
|
||||
resolver. TR-05 personal preference, TR-06 recommendation, TR-07 temporary
|
||||
non-action, and TR-08 `none` stopped before prompt construction and recorded an
|
||||
explicit skipped-call artifact. This hard gate worked in all four cases.
|
||||
|
||||
The target schema contained exactly `candidate_observation_id`,
|
||||
`target_observation_id`, and `normalized_target_text`, with local IDs,
|
||||
same-or-earlier ordering, unique evidence provenance, null consistency, and
|
||||
recursive normative-field exclusion. TR-01 used the self-contained strategy:
|
||||
the prompt stated that linkage was deterministically fixed to the candidate and
|
||||
requested semantic normalization only. TR-02 through TR-04 used paired local
|
||||
resolution. No original transcript or new Negative Act classification call was
|
||||
used.
|
||||
|
||||
| Case | Form / eligible | Call | Target result | Verdict |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| TR-01 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; required same-observation target unresolved | FAIL |
|
||||
| TR-02 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; `obs_1` unresolved | FAIL |
|
||||
| TR-03 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; scoped `obs_1` unresolved | FAIL |
|
||||
| TR-04 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; real-plant target unresolved | FAIL |
|
||||
| TR-05 | `personal_preference` / no | no | Deterministically skipped | PASS |
|
||||
| TR-06 | `recommendation` / no | no | Deterministically skipped | PASS |
|
||||
| TR-07 | `temporary_non_action` / no | no | Deterministically skipped | PASS |
|
||||
| TR-08 | `none` / no | no | Deterministically skipped | PASS |
|
||||
|
||||
Result: four PASS, zero PARTIAL, four FAIL. Exactly four successful Ollama
|
||||
calls were made, all for eligible cases; there were zero technical call
|
||||
failures and four structural validation failures. All four raw responses used
|
||||
the JSON string `"null"` as target ID rather than a supplied observation ID or
|
||||
JSON null. Wrong-target count and unresolved-target count were therefore four.
|
||||
No qualifier-preservation claim can be made because no eligible positive target
|
||||
passed validation. TR-04 alternative isolation likewise could not be
|
||||
established. No rejection, status, decision, outcome, responsibility, or other
|
||||
normative leakage occurred, and no rejection derivation was performed.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`,
|
||||
`num_ctx=16384`, `num_predict=1024`, no retries, voting, or prompt changes.
|
||||
Aggregate runner time was 4.034 seconds.
|
||||
|
||||
Conclusion: eligibility gating is successful and should be retained; it fully
|
||||
prevents unnecessary target calls for ineligible Negative Act forms. Target
|
||||
Resolution V0 itself failed structurally across all eligible cases. Neither the
|
||||
self-contained nor paired strategy produced a valid target, and merely
|
||||
instructing deterministic self-linkage in the semantic prompt did not make the
|
||||
linkage structurally deterministic. The repeated `"null"` string pattern
|
||||
requires diagnosis before changing the architecture or prompt. No rejection
|
||||
derivation is justified by this result.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/target_resolution_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0038 — Target Resolution V1 Diagnostic
|
||||
|
||||
Status: Experimental; linkage boundary successful, normalization incomplete
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
Forensics on failed Target Resolution V0 found a definite prompt defect: its
|
||||
illustrative value `"observation ID or null"` placed both alternatives inside
|
||||
a JSON string. V0 also sent only `format: "json"`, which enforced JSON syntax
|
||||
but not field types. This isolated diagnostic changed only the linkage/output
|
||||
boundary. It contains no rejection derivation or normative semantics.
|
||||
|
||||
Ollama 0.32.6 accepted a true JSON Schema object in `format`. TR1-V1 removed
|
||||
target selection from the model output entirely and deterministically linked
|
||||
the self-contained candidate to itself. TR2-V1 through TR4-V1 used a closed
|
||||
allowed-ID list, an enum of those IDs plus JSON null, typed positive and null
|
||||
examples, recursive strict validation, and one fixed paired prompt. Linkage and
|
||||
normalization were persisted separately.
|
||||
|
||||
| Case | Strategy / ID source | Target | Normalized target | Verdict |
|
||||
| --- | --- | --- | --- | --- |
|
||||
| TR1-V1 | self-contained / deterministic | `obs_1` | `Mit Dr. Schlummer arbeiten wir nicht weiter.` retained negation instead of a positive action meaning | FAIL |
|
||||
| TR2-V1 | paired / LLM | `obs_1` | `externe Lösung weiterverfolgen` | PASS |
|
||||
| TR3-V1 | paired / LLM | `obs_1` | `reale Anlage zur Diskussion` lost `Druckversuch` purpose and the `nutzen` action | FAIL |
|
||||
| TR4-V1 | paired / LLM | `obs_1` | `Versuch in der realen Anlage durchführen`; Technikum excluded | PASS |
|
||||
|
||||
Result: two PASS, zero PARTIAL, two FAIL. All four responses passed their true
|
||||
JSON Schemas. Every resulting target was `obs_1`; wrong-target and
|
||||
unresolved-target counts were zero. The string `"null"` recurrence count was
|
||||
zero, and there were zero structural validation failures. TR1 preserved the
|
||||
collaboration, person, and continuation wording but failed positive-action
|
||||
normalization by retaining negation. TR3 had one material scope loss. TR4
|
||||
preserved real-plant scope and isolated the Technikum alternative. No
|
||||
normative leakage occurred.
|
||||
|
||||
Configuration: exactly four `qwen3.5:9B` calls, temperature 0,
|
||||
`think=false`, `num_ctx=16384`, `num_predict=1024`, no retries, voting, or
|
||||
prompt tuning. Aggregate runner time was 4.557 seconds.
|
||||
|
||||
Conclusion: the V0 string-null failure was primarily a linkage/output-boundary
|
||||
failure rather than evidence that observation-ID linkage is semantically
|
||||
impossible. True typed schemas, closed ID lists, and deterministic self-linkage
|
||||
eliminated every structural and target-ID failure. The experiment still fails
|
||||
its complete acceptance criterion because target normalization is not reliably
|
||||
positive or scope-preserving. These results justify separating linkage from
|
||||
normalization, but not deriving rejection or integrating a new pipeline.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/target_resolution_v1_diagnostic/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0039 — Target Normalization V0
|
||||
|
||||
Status: Experimental; normalization improved but incomplete
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
Target Resolution V1 established correct linkage for all four narrow cases and
|
||||
eliminated structural ID failures with deterministic self-linkage, closed ID
|
||||
lists, and true JSON Schemas. Its remaining failures were normalization-only.
|
||||
This isolated follow-up therefore accepted candidate and target IDs as fixed
|
||||
input and tested only reconstruction of the positive German action meaning. It
|
||||
contains no target selection, Negative Act classification, eligibility logic,
|
||||
rejection derivation, or production integration.
|
||||
|
||||
The strict output schema contained exactly `candidate_observation_id`,
|
||||
`target_observation_id`, and `normalized_target_text`. Both IDs were constrained
|
||||
to their supplied values with JSON Schema `const`; normalized text was a
|
||||
non-empty string and null was disallowed. The one fixed prompt required removal
|
||||
of negative polarity, preservation of action, continuation, material scope and
|
||||
source language, and exclusion of separate alternatives.
|
||||
|
||||
| Case | Actual normalized target | Verdict |
|
||||
| --- | --- | --- |
|
||||
| TN-01 | `Mit Dr. Schlummer zusammenarbeiten` | FAIL: positive polarity and collaboration survived, but continuation was lost |
|
||||
| TN-02 | `externe Lösung weiterverfolgen` | PASS |
|
||||
| TN-03 | `reale Anlage für den Druckversuch nutzen` | PASS |
|
||||
| TN-04 | `Versuch in der realen Anlage durchführen` | PASS; Technikum alternative excluded |
|
||||
|
||||
Result: three PASS, zero PARTIAL, one FAIL. All four outputs passed strict
|
||||
schema validation and copied both fixed IDs exactly, so changed-ID count was
|
||||
zero. Polarity-error count was zero: even TN-01 removed rejection and negation.
|
||||
Action/continuation-loss count was one (TN-01); material purpose/location
|
||||
scope-loss count was zero; alternative-absorption count was zero. There was no
|
||||
unsupported strengthening or normative leakage.
|
||||
|
||||
Configuration: exactly four `qwen3.5:9B` calls, temperature 0,
|
||||
`think=false`, true JSON Schema, `num_ctx=16384`, `num_predict=1024`, no
|
||||
retries, voting, or prompt tuning. Aggregate runner time was 5.066 seconds.
|
||||
|
||||
Conclusion: isolating normalization solved the polarity and scoped-action
|
||||
failures seen in Target Resolution V1 for three of four cases, including exact
|
||||
pressure-test scope and alternative isolation. Continuation semantics remain
|
||||
unreliable in the self-contained collaboration case, so Target Normalization
|
||||
V0 does not meet its full acceptance criterion. The result does not justify
|
||||
rejection derivation or production integration.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/target_normalization_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
|
||||
|
||||
Date: 2026-08-11
|
||||
|
||||
Hypothesis: the primary protocol should be a topic-oriented reconstruction of
|
||||
the meeting rather than a category-oriented list of extracted information.
|
||||
|
||||
This first isolated prototype does not replace or connect to the production
|
||||
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
|
||||
transcript excerpts to `qwen3.5:9B` and requests Discussion Subjects. Each
|
||||
subject may contain supported discourse events, an outcome with mandatory
|
||||
scope, resulting actions and unresolved issues. Optional structures must be
|
||||
omitted when absent. Every semantic object must reference known evidence IDs.
|
||||
|
||||
The strict experimental schema validates:
|
||||
|
||||
- non-empty subjects and globally unique semantic identifiers;
|
||||
- a closed discourse-event vocabulary;
|
||||
- non-empty, known and non-duplicated evidence references;
|
||||
- outcome text, scope, certainty and evidence;
|
||||
- action text, JSON-nullable responsibility/deadline and evidence;
|
||||
- unresolved-issue text and evidence;
|
||||
- omission rather than null or empty optional structures.
|
||||
|
||||
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only,
|
||||
multiple options, unaccepted proposal, proposal with objection, rejected
|
||||
alternative, trial-scoped acceptance, no-decision discussion, resulting Action
|
||||
Item, and outcome plus unresolved issue. Evaluation targets semantic identity,
|
||||
development, outcome scope, actions, unresolved issues, traceability and
|
||||
absence of invented commitments rather than exact wording.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`. Each case received exactly one model call; there were no
|
||||
model retries or prompt iterations. The nine completed calls took 59.251
|
||||
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
|
||||
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
|
||||
failure artifacts were preserved under
|
||||
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/` and
|
||||
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/`. Two earlier launch
|
||||
attempts made zero LLM calls: one failed on the script import path and one was
|
||||
blocked by sandbox networking.
|
||||
|
||||
Human-reviewed results after correcting two objectively wrong Gold assumptions
|
||||
without another model call:
|
||||
|
||||
| Case | Verdict | Reason |
|
||||
| --- | --- | --- |
|
||||
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled `considered_option` rather than `introduced_idea`. |
|
||||
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
|
||||
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
|
||||
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
|
||||
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
|
||||
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
|
||||
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
|
||||
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
|
||||
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
|
||||
|
||||
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case
|
||||
F: the model distinguished acceptance for a bounded trial from acceptance as a
|
||||
final solution. It also handled the explicit action in case H well. However,
|
||||
the experiment failed systematically on sparse structured output, subject
|
||||
grouping and restraint around absent outcomes/actions/unresolved issues. The
|
||||
model frequently mirrored optional schema fields as empty/null values, treated
|
||||
alternatives as outcomes, split one discussion into multiple subjects, or
|
||||
invented open issues from mere non-selection.
|
||||
|
||||
The focused experiment is not promising enough to justify a real Progeo chunk
|
||||
sanity check. No such run was performed, and no architecture is accepted on
|
||||
the basis of this prototype. Further work should first analyze whether the
|
||||
failure comes from the schema/prompt representation, the model's sparse-output
|
||||
reliability, or the boundary between subject grouping and semantic synthesis.
|
||||
It should not proceed through repeated prompt tuning against these nine cases.
|
||||
|
||||
@@ -107,8 +107,10 @@ or aliases are corrected.
|
||||
|
||||
`department`: Organizational unit. Optional and nullable.
|
||||
|
||||
`attendance_status`: `present` for participants. This distinguishes attendees
|
||||
from mentioned people.
|
||||
`attendance_status`: exactly `present` for participants or `mentioned_only` for
|
||||
people who are relevant but did not attend. For backward compatibility, a
|
||||
missing status defaults to `present` in `participants` and `mentioned_only` in
|
||||
`mentioned_people`.
|
||||
|
||||
`mentioned_people`: People discussed or referenced but not present. They are
|
||||
not participants and must not be treated as speakers.
|
||||
@@ -153,7 +155,7 @@ mentioned_people:
|
||||
aliases: []
|
||||
role: null
|
||||
department: null
|
||||
attendance_status: "not_present"
|
||||
attendance_status: "mentioned_only"
|
||||
notes: "Wurde erwaehnt, war aber nicht anwesend."
|
||||
|
||||
organization:
|
||||
@@ -282,6 +284,8 @@ The current validator checks that:
|
||||
- participant ids and mentioned-person ids do not collide
|
||||
- referenced departments exist in `organization.departments`
|
||||
- `attendance_status` values are valid
|
||||
- speaker mappings reference present participants only; mentioned-only people
|
||||
cannot be diarized speakers
|
||||
- participants are marked `present`
|
||||
- mentioned people are not marked `present`
|
||||
|
||||
|
||||
@@ -0,0 +1,229 @@
|
||||
# Protocol Generation Decision
|
||||
|
||||
## Executive Summary
|
||||
|
||||
Meeting Lab tested ten protocol-generation and runtime variants against the
|
||||
same 93.5-minute reference meeting, `project_process_meeting`. The current
|
||||
evidence does not support a fully automatic protocol. The best practical local
|
||||
baseline remains one direct call to `qwen3.6:35B-A3B`, followed by informed
|
||||
human review. It is fast and produces readable, broadly useful Markdown, but it
|
||||
still overstates consensus, compresses unresolved process boundaries, misses
|
||||
some challenge and follow-up paths, and can infer unsafe ownership. Its current
|
||||
quality verdict is **C — promising but insufficient**.
|
||||
|
||||
Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model
|
||||
and 70B-scale variants did not produce a reliable step change. Some improved
|
||||
individual dimensions, but none reached a stable B result or removed the need
|
||||
for substantive review. This is a current evidence-based product choice, not a
|
||||
permanent architecture decision.
|
||||
|
||||
## Experiments Compared
|
||||
|
||||
All protocol rows used the complete cleaned transcript and meeting context
|
||||
unless stated otherwise. A dash means that the artifact did not record the
|
||||
number; it is not an estimate. Verdicts marked “assessment” are comparative
|
||||
assessments of saved outputs because those older artifact directories contain
|
||||
no formal `quality_review.json`.
|
||||
|
||||
| Experiment | Model | Architecture | LLM calls | Prompt tokens | Runtime | Human editing | Verdict | Main strength | Main failure | MVP | Research |
|
||||
| --- | --- | --- | ---: | ---: | ---: | --- | --- | --- | --- | --- | --- |
|
||||
| Direct one-shot baseline | Qwen3.6 35B-A3B | Direct full-context protocol | 1 | 18,385 | 72.5 s cold; about 38 s inference | Not recorded | C | Fast, readable, broad topic outline | Consensus and ownership overpromotion; missing boundaries and follow-up | **Yes, with review** | Baseline |
|
||||
| Conservative one-shot | Qwen3.6 35B-A3B | Direct with stronger safety instructions | 1 | 19,120 | 39.5 s warm | Not recorded | C (assessment) | Better uncertainty and pending-feedback language | Still invents or upgrades named follow-up actions | No | Limited |
|
||||
| Draft → review | Qwen3.6 35B-A3B | Conservative draft plus review call | 2 | 39,071 total | About 110.4 s summed | Not recorded | C (assessment) | Removes some unsafe named attribution | Does not reliably restore omitted content; empty/weak action sections remain | No | Limited |
|
||||
| Meeting Map → protocol | Qwen3.6 35B-A3B | Semantic map followed by rendering | 2 | 39,129 total | 134.8 s | Not recorded | C (assessment) | Explicit intermediate structure | Map errors propagate: false consensus and named ownership remain | No | Yes |
|
||||
| Hierarchical notes → protocol | Qwen3.6 35B-A3B | Five chunk-note calls plus synthesis | 6 | 32,145 total | 251.1 s | Not recorded | C (assessment) | Highest recall in several detailed/open topics | Amplifies unsupported speaker/name interpretations and confirmed actions | No | Yes |
|
||||
| Segment-level anonymous diarization | Qwen3.6 35B-A3B | One-shot over 1,154 labeled segments | 1 | 39,308 | 125.4 s | 25–35 min | C | Preserves some filtered-idea challenge structure | Token count more than doubled; actions and deadlines became less safe | No | No further protocol tests |
|
||||
| Turn-merged anonymous diarization | Qwen3.6 35B-A3B | One-shot over 435 merged turns | 1 | 22,490 | 101.2 s | 25–35 min | C | Corrected token inflation; recovered some topic and feedback detail | Still did not beat raw input; unsafe actions/deadlines persisted | No | UI/search only |
|
||||
| Dense one-shot | Qwen3.5 27B | Direct full-context protocol | 1 | 18,385 | 181.2 s | Not recorded | C (assessment) | Somewhat better recall of process details | Much slower; no material overall quality gain | No | No |
|
||||
| 70B scale one-shot | Llama 3.3 70B Q3_K_S | Dense, 64% CPU / 36% GPU | 1 | 21,156 | 470.1 s warm | 45–60 min | D | Technically proved a 70B hybrid load can run | Severe coverage loss, invented governance, internal contradiction | No | Negative scale result |
|
||||
| Ollama vs native llama.cpp | Qwen3.6 35B-A3B | Same Q4_K_M GGUF; ROCm/Vulkan servers | 1 per backend | 18,385 | ROCm 38.4 s; Vulkan 42.3 s | N/A | Runtime only | Native ROCm reached 51.38 generated tok/s | No meaningful end-to-end advantage; more operational complexity | Ollama | Runtime reference |
|
||||
|
||||
The draft-review total combines the saved conservative draft call and the
|
||||
saved review call. Its review metadata itself reports only the one new review
|
||||
call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second
|
||||
wall time includes a 34.5-second cold load; its measured prompt evaluation plus
|
||||
generation was 37.8 seconds. These distinctions explain apparent runtime
|
||||
differences between otherwise similar Qwen3.6 calls.
|
||||
|
||||
### Recurring quality patterns
|
||||
|
||||
- **Topic coverage and factual accuracy:** Direct Qwen3.6 captures the main
|
||||
process but misses the second review after enrichment, project reporting and
|
||||
parts of the filtered-idea challenge path. Hierarchical processing recalls
|
||||
more detail but introduces too many unsupported interpretations. Llama 3.3
|
||||
loses most of the meeting and invents a governance role for the
|
||||
Geschäftsführung.
|
||||
- **Consensus and unresolved boundaries:** Every broad one-shot family remains
|
||||
vulnerable to turning discussion or a working direction into agreement. The
|
||||
unresolved boundary between central coordination and autonomous department
|
||||
work, and the uncertainty around universal filter criteria, are especially
|
||||
fragile.
|
||||
- **Visibility, veto and reconsideration:** No approach consistently preserves
|
||||
initial filtering, later cross-functional challenge, reconsideration after
|
||||
enrichment and the return through the project cycle together.
|
||||
- **Stakeholder feedback:** Pending Jovana and Björn feedback is an important
|
||||
quality probe. Some variants preserve both; segment-level diarization drops
|
||||
Björn, while Llama 3.3 drops both.
|
||||
- **Actions and attribution:** Added structure does not guarantee safety.
|
||||
Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still
|
||||
promote proposals or expected work into confirmed actions, infer owners from
|
||||
roles or conversational context, or invent deadlines. Human review remains
|
||||
mandatory.
|
||||
|
||||
## Model Findings
|
||||
|
||||
### Qwen3.6:35B-A3B
|
||||
|
||||
Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is
|
||||
operationally fast on the RX 9070/CPU hybrid setup, follows the requested
|
||||
Markdown form and usually provides a useful first draft. It remains verdict C:
|
||||
larger context and fluent synthesis do not reliably protect evidence strength,
|
||||
responsibility attribution or unresolved process boundaries.
|
||||
|
||||
### Qwen3.5:27b dense
|
||||
|
||||
The dense 27B run recalled some process details better than the MoE baseline,
|
||||
but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount
|
||||
to a material overall quality improvement. This result does not prove that
|
||||
dense models are generally inferior; it shows that this dense model is not a
|
||||
better product choice on this hardware and meeting.
|
||||
|
||||
### Llama 3.3 70B Q3_K_S
|
||||
|
||||
Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded
|
||||
footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a
|
||||
very short, materially worse protocol: one critical invented governance claim,
|
||||
five new major errors and an estimated 45–60 minutes of editing. Raw parameter
|
||||
count alone is therefore insufficient. The older model generation and
|
||||
aggressive Q3 quantization are plausible contributors, but this experiment
|
||||
does not isolate or prove either cause.
|
||||
|
||||
## Runtime Findings
|
||||
|
||||
The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF
|
||||
that Ollama uses. The tested `llama-server` binary was the llama.cpp runtime
|
||||
shipped with the installed Ollama distribution, not an independent source
|
||||
build.
|
||||
|
||||
Native ROCm processed the reference request in 38.4 seconds and generated at
|
||||
51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The
|
||||
comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds
|
||||
of prompt evaluation plus generation when load time is excluded. Output token
|
||||
counts differed, so generation throughput alone is not an end-to-end quality or
|
||||
latency comparison.
|
||||
|
||||
Native ROCm gained some generation throughput, but did not provide a meaningful
|
||||
end-to-end advantage for this workload. Vulkan required more host spill and was
|
||||
not preferable. Ollama already provides the relevant llama.cpp runtime
|
||||
components, model lifecycle and API integration; it remains the preferred
|
||||
routine Meeting Lab runtime.
|
||||
|
||||
## Diarization Findings
|
||||
|
||||
### Technical feasibility
|
||||
|
||||
Pyannote `speaker-diarization-community-1` successfully processed the
|
||||
93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It
|
||||
detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments
|
||||
(99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it
|
||||
does not establish speaker identity or diarization accuracy against labeled
|
||||
ground truth.
|
||||
|
||||
### Protocol-quality impact
|
||||
|
||||
Annotating every Whisper segment increased the Qwen prompt from 18,385 to
|
||||
39,308 tokens, confounding speaker structure with fragmentation and token
|
||||
inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and
|
||||
the complete prompt to 22,490 tokens. That controlled the main representation
|
||||
confound, but the resulting protocol still did not materially outperform the
|
||||
raw transcript and remained verdict C.
|
||||
|
||||
Anonymous diarization is therefore not justified as a mandatory MVP
|
||||
protocol-quality feature. This does **not** mean diarization is generally
|
||||
useless. It may remain valuable for speaker-aware UI, navigation and search,
|
||||
participation statistics, traceability, or later carefully validated real-name
|
||||
mapping.
|
||||
|
||||
## Semantic Research Findings
|
||||
|
||||
The semantic experiments provide architectural evidence, but should not
|
||||
dominate the product decision:
|
||||
|
||||
- **Evidence Observation V3** is a strong evidence-near candidate stage. With
|
||||
Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges,
|
||||
alternatives, requests, commitments and boundaries in natural language.
|
||||
- **Request/Acceptance** and **Collective Commitment** show that narrow semantic
|
||||
recognition followed by deterministic provenance, ordering, addressee,
|
||||
negation and deadline gates can safely derive limited consequences. Model
|
||||
recognition errors were contained without inventing individual ownership.
|
||||
- **Explicit Rejection** failed when reduced to a coarse binary recognition
|
||||
problem: semantically valid false positives passed structural gates.
|
||||
- **Negative Act Form** worked better by distinguishing non-pursuit, personal
|
||||
preference, recommendation and temporary non-action before any normative
|
||||
derivation. All eight form classifications matched Gold, although normalized
|
||||
action text was imperfect in three cases.
|
||||
- **Target Resolution V0** failed because prompt examples and a weak JSON
|
||||
boundary encouraged the string `"null"` instead of typed linkage.
|
||||
**Target Resolution V1** fixed all linkage/ID failures with deterministic
|
||||
self-linkage, closed ID lists and true JSON Schema, but normalization remained
|
||||
incomplete.
|
||||
- **Target Normalization V0** improved polarity and scope preservation to 3/4
|
||||
PASS, but still lost continuation meaning in the collaboration case.
|
||||
|
||||
These findings support Meeting Lab as a research and validation track. They do
|
||||
not yet justify placing a multi-stage semantic pipeline on the MVP critical
|
||||
path.
|
||||
|
||||
## Current MVP Decision
|
||||
|
||||
The current product path is:
|
||||
|
||||
```text
|
||||
Audio
|
||||
-> transcription
|
||||
-> direct qwen3.6:35B-A3B protocol generation through Ollama
|
||||
-> informed human review
|
||||
-> final protocol
|
||||
```
|
||||
|
||||
The first MVP should treat the generated protocol as an editable draft, not an
|
||||
authoritative semantic record. Human review must specifically check consensus,
|
||||
unresolved boundaries, competing positions, action status, owners, deadlines
|
||||
and pending stakeholder feedback.
|
||||
|
||||
Diarization is optional and deferred. The semantic research pipeline remains
|
||||
in Meeting Lab, outside the MVP critical path. Ollama remains the default local
|
||||
runtime.
|
||||
|
||||
## Rejected / Deferred Directions
|
||||
|
||||
- Do not continue Qwen3.6 prompt variants as the main quality strategy.
|
||||
- Do not add draft-review, Meeting Map or hierarchical generation to the MVP;
|
||||
their added calls and complexity did not deliver reliable quality gains.
|
||||
- Do not continue anonymous-diarization protocol experiments. Revisit
|
||||
diarization for UI, search, statistics or traceability instead.
|
||||
- Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model.
|
||||
- Do not replace Ollama with a manually managed native llama.cpp service for
|
||||
this workload.
|
||||
- Retain semantic experiments, but defer production integration and broad
|
||||
semantic consolidation.
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. How does a genuinely newer, materially stronger model perform when a useful
|
||||
quantization fits the available RAM/VRAM without severe swap?
|
||||
2. If project policy permits, what quality ceiling does the unchanged reference
|
||||
prompt achieve with a commercial frontier model?
|
||||
3. What is the measured reviewer time and correction distribution once the
|
||||
direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow?
|
||||
4. Which non-protocol product benefits justify revisiting diarization later?
|
||||
|
||||
No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old
|
||||
70B Q3 scale tests are recommended.
|
||||
|
||||
## Recommended Next Product Step
|
||||
|
||||
Build the practical end-to-end MVP around direct Qwen3.6 generation and an
|
||||
explicit human review handoff. Measure reviewer time and correction categories
|
||||
in real use. Keep the experiment artifacts and semantic Gold work as validation
|
||||
evidence, but do not block the first product loop on broader research stages.
|
||||
@@ -75,7 +75,7 @@ mentioned_people:
|
||||
aliases: []
|
||||
role: null
|
||||
department: null
|
||||
attendance_status: "not_present"
|
||||
attendance_status: "mentioned_only"
|
||||
notes: null
|
||||
|
||||
organization:
|
||||
@@ -127,4 +127,4 @@ context_rules:
|
||||
do_not_infer_departments: true
|
||||
do_not_infer_responsibilities: true
|
||||
do_not_infer_attendance: true
|
||||
mentioned_people_are_not_participants: true
|
||||
mentioned_people_are_not_participants: true
|
||||
|
||||
@@ -64,7 +64,7 @@ mentioned_people:
|
||||
- "Giovana"
|
||||
role: "Leiterin Business Development"
|
||||
department_id: "bd"
|
||||
attendance_status: "not_present"
|
||||
attendance_status: "mentioned_only"
|
||||
notes: null
|
||||
|
||||
organization:
|
||||
|
||||
@@ -4,6 +4,8 @@
|
||||
schema_version: "1"
|
||||
|
||||
meeting:
|
||||
# Stable identifier used for provenance across corrections and later runs.
|
||||
meeting_id: ""
|
||||
# Human-readable title for the meeting.
|
||||
title: ""
|
||||
# Dominant meeting language, for example "de" or "en".
|
||||
@@ -28,6 +30,12 @@ participants:
|
||||
attendance_status: "present"
|
||||
notes: null
|
||||
|
||||
# Optional authoritative mapping from diarization labels to actual participants.
|
||||
# Add entries only after a human or trusted external process confirms identity.
|
||||
# Never infer mappings from conversational context. Unmapped labels stay anonymous.
|
||||
speaker_mappings: {}
|
||||
# SPEAKER_00: "participant-id"
|
||||
|
||||
mentioned_people:
|
||||
# People discussed or referenced but not present in the meeting.
|
||||
# Mentioned people are not speakers and must not become responsible persons
|
||||
@@ -37,7 +45,7 @@ mentioned_people:
|
||||
aliases: []
|
||||
role: null
|
||||
department: null
|
||||
attendance_status: "not_present"
|
||||
attendance_status: "mentioned_only"
|
||||
notes: null
|
||||
|
||||
organization:
|
||||
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the collective-commitment Gold experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_collective import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path: sys.path.insert(0, str(ROOT))
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection_v1 import main
|
||||
if __name__ == "__main__": raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the H-only controlled derivation experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_h import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,153 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run the one-call direct protocol MVP from compact Whisper JSON."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import shutil
|
||||
import sys
|
||||
import time
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT # noqa: E402
|
||||
from src.meeting_lab.protocol.generate_direct_protocol import ( # noqa: E402
|
||||
DEFAULT_MODEL,
|
||||
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
DirectProtocolResult,
|
||||
generate_direct_protocol,
|
||||
)
|
||||
|
||||
|
||||
DEFAULT_OUTPUT_ROOT = Path("meeting_data/runs")
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Generate one direct protocol from compact Whisper JSON.")
|
||||
parser.add_argument("transcript", type=Path)
|
||||
parser.add_argument("--context", type=Path)
|
||||
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument(
|
||||
"--safe-input-token-budget",
|
||||
type=int,
|
||||
default=DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
)
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def create_unique_run_dir(
|
||||
output_root: Path,
|
||||
transcript_stem: str,
|
||||
now: Callable[[], datetime] = datetime.now,
|
||||
) -> Path:
|
||||
safe_stem = re.sub(r"[^A-Za-z0-9_.-]+", "_", transcript_stem).strip("._-") or "meeting"
|
||||
base = output_root / f"{safe_stem}_{now().strftime('%Y%m%d_%H%M%S')}"
|
||||
candidate = base
|
||||
suffix = 1
|
||||
while candidate.exists():
|
||||
candidate = output_root / f"{base.name}_{suffix:02d}"
|
||||
suffix += 1
|
||||
candidate.mkdir(parents=True)
|
||||
return candidate
|
||||
|
||||
|
||||
def write_json(path: Path, data: Any) -> None:
|
||||
path.write_text(json.dumps(data, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def persist_result(run_dir: Path, result: DirectProtocolResult) -> Path:
|
||||
protocol_dir = run_dir / "protocol"
|
||||
protocol_dir.mkdir()
|
||||
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
|
||||
write_json(protocol_dir / "raw_response.json", result.raw_response)
|
||||
write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
|
||||
transcript_input = getattr(result, "transcript_input", None)
|
||||
if transcript_input is not None:
|
||||
(protocol_dir / "transcript_input.txt").write_text(
|
||||
transcript_input, encoding="utf-8"
|
||||
)
|
||||
protocol_path = run_dir / "protocol.md"
|
||||
protocol_path.write_text(result.protocol_text, encoding="utf-8")
|
||||
return protocol_path
|
||||
|
||||
|
||||
def run(args: argparse.Namespace) -> tuple[int, Path, Path | None]:
|
||||
run_dir = create_unique_run_dir(args.output_root, args.transcript.stem)
|
||||
timestamp = datetime.now().astimezone().isoformat(timespec="seconds")
|
||||
started = time.perf_counter()
|
||||
protocol_path: Path | None = None
|
||||
metadata: dict[str, Any] = {
|
||||
"run_id": run_dir.name,
|
||||
"timestamp": timestamp,
|
||||
"transcript_path": str(args.transcript.resolve()),
|
||||
"context_path": str(args.context.resolve()) if args.context else None,
|
||||
"model": args.model,
|
||||
"ollama_endpoint": args.ollama_endpoint,
|
||||
"status": "running",
|
||||
"total_runtime_seconds": None,
|
||||
"final_protocol_path": None,
|
||||
}
|
||||
try:
|
||||
transcript_dir = run_dir / "transcript"
|
||||
transcript_dir.mkdir()
|
||||
if not args.transcript.is_file():
|
||||
raise FileNotFoundError(f"Transcript file does not exist: {args.transcript}")
|
||||
preserved_transcript = transcript_dir / "transcript.json"
|
||||
shutil.copy2(args.transcript, preserved_transcript)
|
||||
|
||||
preserved_context: Path | None = None
|
||||
if args.context is not None:
|
||||
if not args.context.is_file():
|
||||
raise FileNotFoundError(f"Meeting Context file does not exist: {args.context}")
|
||||
context_dir = run_dir / "context"
|
||||
context_dir.mkdir()
|
||||
preserved_context = context_dir / "meeting_context.yaml"
|
||||
shutil.copy2(args.context, preserved_context)
|
||||
|
||||
write_json(
|
||||
run_dir / "input_manifest.json",
|
||||
{
|
||||
"transcript_source": str(args.transcript.resolve()),
|
||||
"transcript_copy": str(preserved_transcript.resolve()),
|
||||
"context_source": str(args.context.resolve()) if args.context else None,
|
||||
"context_copy": str(preserved_context.resolve()) if preserved_context else None,
|
||||
},
|
||||
)
|
||||
result = generate_direct_protocol(
|
||||
preserved_transcript,
|
||||
preserved_context,
|
||||
model=args.model,
|
||||
endpoint=args.ollama_endpoint,
|
||||
safe_input_token_budget=args.safe_input_token_budget,
|
||||
)
|
||||
protocol_path = persist_result(run_dir, result)
|
||||
metadata["status"] = "completed"
|
||||
metadata["final_protocol_path"] = str(protocol_path.resolve())
|
||||
except Exception as exc:
|
||||
metadata["status"] = "failed"
|
||||
metadata["failure"] = f"{type(exc).__name__}: {exc}"
|
||||
print(f"Error: {metadata['failure']}", file=sys.stderr)
|
||||
finally:
|
||||
metadata["total_runtime_seconds"] = round(time.perf_counter() - started, 3)
|
||||
write_json(run_dir / "run_metadata.json", metadata)
|
||||
return (0 if metadata["status"] == "completed" else 2), run_dir, protocol_path
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
code, _run_dir, protocol_path = run(parse_args(argv))
|
||||
if protocol_path is not None:
|
||||
print(protocol_path)
|
||||
return code
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the evidence-near observation experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.evidence_observations.experiment import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for evidence-near observation experiment V2."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.evidence_observations_v2.experiment import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for evidence-near observation experiment V3."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.evidence_observations_v3.experiment import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the explicit-rejection Gold experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,151 @@
|
||||
#!/usr/bin/env python3
|
||||
"""CLI adapter for the reusable Meeting Lab MVP orchestration API."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT # noqa: E402
|
||||
from src.meeting_lab.models.meeting_context import MeetingContext # noqa: E402
|
||||
from src.meeting_lab.orchestration.mvp import ( # noqa: E402
|
||||
DEFAULT_DIARIZATION_MODEL,
|
||||
DEFAULT_MODEL,
|
||||
DEFAULT_OUTPUT_ROOT,
|
||||
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
MvpMeetingConfig,
|
||||
create_unique_run_dir,
|
||||
run_mvp_meeting,
|
||||
)
|
||||
from src.meeting_lab.progress import ProgressSink # noqa: E402
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Transcribe one meeting and generate one direct protocol."
|
||||
)
|
||||
parser.add_argument("audio_file", type=Path)
|
||||
parser.add_argument("--whisper-model", type=Path, required=True)
|
||||
parser.add_argument("--whisper-executable", default="whisper-cli")
|
||||
parser.add_argument("--ffmpeg-executable", default="ffmpeg")
|
||||
parser.add_argument(
|
||||
"--audio-normalization",
|
||||
action=argparse.BooleanOptionalAction,
|
||||
default=True,
|
||||
help=(
|
||||
"Enable FFmpeg loudness normalization during canonical audio preparation "
|
||||
"(default: enabled)."
|
||||
),
|
||||
)
|
||||
parser.add_argument("--context", type=Path)
|
||||
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
|
||||
parser.add_argument("--language", default="de")
|
||||
parser.add_argument(
|
||||
"--threads",
|
||||
default="auto",
|
||||
help="Thread count or 'auto' for physical CPU cores (default: auto).",
|
||||
)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument(
|
||||
"--protocol-safe-input-token-budget",
|
||||
type=int,
|
||||
default=DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
help="Conservative estimated prompt-token limit before any Ollama request.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--diarization",
|
||||
choices=("auto", "gpu", "cpu", "off"),
|
||||
default="off",
|
||||
help="Optional Community-1 diarization device mode (default: off).",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--diarization-runtime",
|
||||
choices=("native", "container"),
|
||||
default="native",
|
||||
help="Run pyannote in this Python environment or an explicit container.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--diarization-container-image",
|
||||
help="Container image required with --diarization-runtime container.",
|
||||
)
|
||||
parser.add_argument(
|
||||
"--diarization-container-arg",
|
||||
action="append",
|
||||
default=[],
|
||||
help="Additional docker argument; repeat and use = for values beginning with --.",
|
||||
)
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def config_from_args(args: argparse.Namespace) -> MvpMeetingConfig:
|
||||
return MvpMeetingConfig(
|
||||
audio_file=args.audio_file,
|
||||
whisper_model=args.whisper_model,
|
||||
whisper_executable=args.whisper_executable,
|
||||
ffmpeg_executable=args.ffmpeg_executable,
|
||||
audio_normalization=args.audio_normalization,
|
||||
context_file=args.context,
|
||||
output_root=args.output_root,
|
||||
language=args.language,
|
||||
threads=args.threads,
|
||||
model=args.model,
|
||||
ollama_endpoint=args.ollama_endpoint,
|
||||
protocol_safe_input_token_budget=args.protocol_safe_input_token_budget,
|
||||
diarization=args.diarization,
|
||||
diarization_runtime=args.diarization_runtime,
|
||||
diarization_container_image=args.diarization_container_image,
|
||||
diarization_container_args=tuple(args.diarization_container_arg),
|
||||
)
|
||||
|
||||
|
||||
def run(
|
||||
args: argparse.Namespace,
|
||||
*,
|
||||
context_override: MeetingContext | dict[str, Any] | None = None,
|
||||
progress_sink: ProgressSink | None = None,
|
||||
) -> tuple[int, Path | None, Path | None]:
|
||||
"""Compatibility wrapper for existing Python callers of the CLI module."""
|
||||
result = run_mvp_meeting(
|
||||
config_from_args(args),
|
||||
meeting_context=context_override,
|
||||
progress_sink=progress_sink,
|
||||
)
|
||||
return result.exit_code, result.run_dir, result.protocol_path
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
args = parse_args(argv)
|
||||
if args.diarization == "off":
|
||||
print("Diarization: disabled")
|
||||
else:
|
||||
print(
|
||||
f"Diarization: enabled; backend=pyannote.audio; "
|
||||
f"model={DEFAULT_DIARIZATION_MODEL}; requested_device={args.diarization}; "
|
||||
f"runtime={args.diarization_runtime}"
|
||||
)
|
||||
code, run_dir, protocol_path = run(args)
|
||||
if run_dir is not None and args.diarization != "off":
|
||||
metadata_path = run_dir / "diarization" / "metadata.json"
|
||||
if metadata_path.is_file():
|
||||
details = json.loads(metadata_path.read_text(encoding="utf-8"))
|
||||
print(
|
||||
f"Diarization result: device={details.get('actual_device')}; "
|
||||
f"device_name={details.get('device_name') or 'n/a'}; "
|
||||
f"runtime={details.get('runtime_seconds'):.3f}s; "
|
||||
f"speakers={details.get('speaker_count')}; artifacts={metadata_path.parent}"
|
||||
)
|
||||
if protocol_path is not None:
|
||||
print(protocol_path)
|
||||
return code
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the Negative Act Form experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_negative_act import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the request/acceptance Gold experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_gold import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the isolated semantic synthesis experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.semantic_synthesis.experiment import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
ROOT=Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path: sys.path.insert(0,str(ROOT))
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_target_normalization import main
|
||||
if __name__=="__main__": raise SystemExit(main())
|
||||
@@ -0,0 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path: sys.path.insert(0, str(ROOT))
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution import main
|
||||
if __name__ == "__main__": raise SystemExit(main())
|
||||
@@ -0,0 +1,7 @@
|
||||
#!/usr/bin/env python3
|
||||
import sys
|
||||
from pathlib import Path
|
||||
ROOT=Path(__file__).resolve().parents[1]
|
||||
if str(ROOT) not in sys.path: sys.path.insert(0,str(ROOT))
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution_v1 import main
|
||||
if __name__=="__main__": raise SystemExit(main())
|
||||
@@ -0,0 +1,16 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repository entry point for the isolated topic reconstruction experiment."""
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.topic_reconstruction.experiment import main # noqa: E402
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,51 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Transcribe one audio file with whisper.cpp; do not generate a protocol."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||
if str(REPO_ROOT) not in sys.path:
|
||||
sys.path.insert(0, str(REPO_ROOT))
|
||||
|
||||
from src.meeting_lab.transcription.whisper import ( # noqa: E402
|
||||
TranscriptionError,
|
||||
transcribe_audio,
|
||||
)
|
||||
|
||||
|
||||
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Create a compact Meeting Lab transcript with whisper.cpp.")
|
||||
parser.add_argument("audio_file", type=Path)
|
||||
parser.add_argument("--model", type=Path, required=True, help="Path to a whisper.cpp GGML model.")
|
||||
parser.add_argument("--output-dir", type=Path, required=True)
|
||||
parser.add_argument("--language", default="auto", help="Language code or 'auto' (default: auto).")
|
||||
parser.add_argument("--threads", default="auto", help="Thread count or 'auto' for physical CPU cores (default: auto).")
|
||||
parser.add_argument("--whisper-executable", default="whisper-cli", help="whisper.cpp CLI executable (default: whisper-cli).")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
args = parse_args(argv)
|
||||
try:
|
||||
result = transcribe_audio(
|
||||
args.audio_file,
|
||||
args.model,
|
||||
args.output_dir,
|
||||
args.language,
|
||||
executable=args.whisper_executable,
|
||||
threads=args.threads,
|
||||
)
|
||||
except TranscriptionError as exc:
|
||||
print(f"Error: {exc}")
|
||||
return 1
|
||||
print(f"Transcript: {result.transcript_json}")
|
||||
print(f"Runtime: {result.runtime_seconds:.3f} seconds")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Canonical audio preparation boundary."""
|
||||
|
||||
from .preparation import AudioPreparationError, PreparedAudio, prepare_audio
|
||||
|
||||
__all__ = ["AudioPreparationError", "PreparedAudio", "prepare_audio"]
|
||||
@@ -0,0 +1,190 @@
|
||||
"""Prepare supported recordings for deterministic downstream processing."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import wave
|
||||
from collections.abc import Callable, Sequence
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
|
||||
SUPPORTED_EXTENSIONS = {".wav", ".flac", ".m4a"}
|
||||
CANONICAL_SAMPLE_RATE = 16_000
|
||||
CANONICAL_CHANNELS = 1
|
||||
CANONICAL_SAMPLE_WIDTH_BYTES = 2
|
||||
CANONICAL_CODEC = "pcm_s16le"
|
||||
DEFAULT_NORMALIZATION_FILTER = "loudnorm=I=-16:LRA=11:TP=-1.5"
|
||||
DEFAULT_NORMALIZATION_METHOD = "ffmpeg_loudnorm"
|
||||
|
||||
|
||||
class AudioPreparationError(RuntimeError):
|
||||
"""Raised when source audio cannot be prepared as canonical WAV."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PreparedAudio:
|
||||
source_path: Path
|
||||
source_format: str
|
||||
prepared_path: Path
|
||||
method: str
|
||||
ffmpeg_executable: str
|
||||
normalization_enabled: bool = True
|
||||
normalization_method: str | None = DEFAULT_NORMALIZATION_METHOD
|
||||
normalization_filter: str | None = DEFAULT_NORMALIZATION_FILTER
|
||||
|
||||
def metadata(self) -> dict[str, object]:
|
||||
return {
|
||||
"original_source_path": str(self.source_path.resolve()),
|
||||
"original_source_name": self.source_path.name,
|
||||
"original_format": self.source_format,
|
||||
"prepared_audio_path": str(self.prepared_path.resolve()),
|
||||
"preparation_method": self.method,
|
||||
"ffmpeg_executable": self.ffmpeg_executable,
|
||||
"normalization_enabled": self.normalization_enabled,
|
||||
"normalization_method": self.normalization_method,
|
||||
"normalization_filter": self.normalization_filter,
|
||||
"canonical_output": {
|
||||
"container": "wav",
|
||||
"codec": CANONICAL_CODEC,
|
||||
"channels": CANONICAL_CHANNELS,
|
||||
"sample_rate_hz": CANONICAL_SAMPLE_RATE,
|
||||
"bits_per_sample": CANONICAL_SAMPLE_WIDTH_BYTES * 8,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
Runner = Callable[..., subprocess.CompletedProcess[str]]
|
||||
|
||||
|
||||
def prepare_audio(
|
||||
source_path: Path,
|
||||
prepared_path: Path,
|
||||
*,
|
||||
ffmpeg_executable: str = "ffmpeg",
|
||||
normalization_enabled: bool = True,
|
||||
runner: Runner = subprocess.run,
|
||||
) -> PreparedAudio:
|
||||
"""Create and validate a canonical mono 16 kHz signed PCM16 WAV artifact."""
|
||||
source_path = Path(source_path)
|
||||
prepared_path = Path(prepared_path)
|
||||
source_format = source_path.suffix.lower()
|
||||
if not source_path.is_file():
|
||||
raise AudioPreparationError(f"Source audio does not exist: {source_path}")
|
||||
if source_format not in SUPPORTED_EXTENSIONS:
|
||||
supported = ", ".join(sorted(SUPPORTED_EXTENSIONS))
|
||||
raise AudioPreparationError(
|
||||
f"Unsupported audio format {source_format or '<none>'!r}; supported: {supported}."
|
||||
)
|
||||
|
||||
resolved_executable = _resolve_executable(ffmpeg_executable)
|
||||
prepared_path.parent.mkdir(parents=True, exist_ok=True)
|
||||
temporary_path = prepared_path.with_name(f".{prepared_path.name}.tmp.wav")
|
||||
command_parts = [
|
||||
resolved_executable,
|
||||
"-nostdin",
|
||||
"-hide_banner",
|
||||
"-loglevel",
|
||||
"error",
|
||||
"-y",
|
||||
"-i",
|
||||
str(source_path),
|
||||
"-map_metadata",
|
||||
"-1",
|
||||
"-vn",
|
||||
]
|
||||
if normalization_enabled:
|
||||
command_parts.extend(("-af", DEFAULT_NORMALIZATION_FILTER))
|
||||
command_parts.extend(
|
||||
(
|
||||
"-ac",
|
||||
str(CANONICAL_CHANNELS),
|
||||
"-ar",
|
||||
str(CANONICAL_SAMPLE_RATE),
|
||||
"-c:a",
|
||||
CANONICAL_CODEC,
|
||||
"-fflags",
|
||||
"+bitexact",
|
||||
str(temporary_path),
|
||||
)
|
||||
)
|
||||
command: Sequence[str] = tuple(command_parts)
|
||||
try:
|
||||
completed = runner(command, capture_output=True, text=True, check=False)
|
||||
except OSError as exc:
|
||||
raise AudioPreparationError(f"Could not run FFmpeg: {exc}") from exc
|
||||
if completed.returncode != 0:
|
||||
detail = (
|
||||
completed.stderr or completed.stdout or "no diagnostic output"
|
||||
).strip()
|
||||
raise AudioPreparationError(
|
||||
f"FFmpeg failed to prepare {source_path.name} (exit {completed.returncode}): "
|
||||
f"{detail}"
|
||||
)
|
||||
try:
|
||||
_validate_canonical_wav(temporary_path)
|
||||
os.replace(temporary_path, prepared_path)
|
||||
except Exception:
|
||||
temporary_path.unlink(missing_ok=True)
|
||||
raise
|
||||
|
||||
return PreparedAudio(
|
||||
source_path=source_path,
|
||||
source_format=source_format.removeprefix("."),
|
||||
prepared_path=prepared_path,
|
||||
method="ffmpeg",
|
||||
ffmpeg_executable=resolved_executable,
|
||||
normalization_enabled=normalization_enabled,
|
||||
normalization_method=(
|
||||
DEFAULT_NORMALIZATION_METHOD if normalization_enabled else None
|
||||
),
|
||||
normalization_filter=(
|
||||
DEFAULT_NORMALIZATION_FILTER if normalization_enabled else None
|
||||
),
|
||||
)
|
||||
|
||||
|
||||
def _resolve_executable(executable: str) -> str:
|
||||
value = executable.strip()
|
||||
if not value:
|
||||
raise AudioPreparationError("FFmpeg executable must not be empty.")
|
||||
if Path(value).parent != Path("."):
|
||||
path = Path(value)
|
||||
if path.is_file() and os.access(path, os.X_OK):
|
||||
return str(path)
|
||||
raise AudioPreparationError(f"FFmpeg executable is not available: {value}")
|
||||
resolved = shutil.which(value)
|
||||
if resolved is None:
|
||||
raise AudioPreparationError(
|
||||
f"FFmpeg executable {value!r} was not found on PATH. Install FFmpeg or "
|
||||
"configure its executable path."
|
||||
)
|
||||
return resolved
|
||||
|
||||
|
||||
def _validate_canonical_wav(path: Path) -> None:
|
||||
try:
|
||||
with wave.open(str(path), "rb") as recording:
|
||||
properties = (
|
||||
recording.getnchannels(),
|
||||
recording.getframerate(),
|
||||
recording.getsampwidth(),
|
||||
recording.getcomptype(),
|
||||
)
|
||||
except (OSError, EOFError, wave.Error) as exc:
|
||||
raise AudioPreparationError(
|
||||
f"FFmpeg did not produce a readable WAV file: {path}: {exc}"
|
||||
) from exc
|
||||
expected = (
|
||||
CANONICAL_CHANNELS,
|
||||
CANONICAL_SAMPLE_RATE,
|
||||
CANONICAL_SAMPLE_WIDTH_BYTES,
|
||||
"NONE",
|
||||
)
|
||||
if properties != expected:
|
||||
raise AudioPreparationError(
|
||||
"Prepared audio is not canonical mono 16 kHz PCM16 WAV: "
|
||||
f"channels={properties[0]}, sample_rate={properties[1]}, "
|
||||
f"sample_width={properties[2]}, compression={properties[3]}."
|
||||
)
|
||||
@@ -0,0 +1 @@
|
||||
"""Isolated controlled semantic derivation experiments."""
|
||||
@@ -0,0 +1,349 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Isolated collective-commitment Gold reliability experiment."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from .experiment_h import (
|
||||
DEFAULT_ENDPOINT,
|
||||
DEFAULT_MODEL,
|
||||
DerivationValidationError,
|
||||
OBSERVATION_KEYS,
|
||||
call_ollama,
|
||||
)
|
||||
|
||||
|
||||
GOLD_SCHEMA_VERSION = "experimental-collective-commitment-gold-v0"
|
||||
RECOGNITION_KEYS = {"observation_id", "commitment_form", "normalized_action_text"}
|
||||
COMMITMENT_FORMS = {"individual_first_person", "collective_first_person", "none"}
|
||||
FORBIDDEN_LLM_KEYS = {
|
||||
"responsible_person", "responsibility", "responsibility_scope",
|
||||
"requested_actor", "owner", "ownership", "assignee", "status",
|
||||
"established", "action_item", "protocol", "protocol_category", "decision",
|
||||
"unresolved_issue", "confidence", "relation", "relations", "graph",
|
||||
}
|
||||
WEEKDAYS = {
|
||||
"monday": "Montag", "montag": "Montag", "tuesday": "Dienstag",
|
||||
"dienstag": "Dienstag", "wednesday": "Mittwoch", "mittwoch": "Mittwoch",
|
||||
"thursday": "Donnerstag", "donnerstag": "Donnerstag", "friday": "Freitag",
|
||||
"freitag": "Freitag", "saturday": "Samstag", "samstag": "Samstag",
|
||||
"sunday": "Sonntag", "sonntag": "Sonntag",
|
||||
}
|
||||
|
||||
PROMPT_TEMPLATE = """Recognize only the explicit first-person commitment form and concise action meaning in the supplied single V3-style observation.
|
||||
|
||||
Answer only:
|
||||
1. What explicit first-person commitment form is present?
|
||||
- individual_first_person: the speaker explicitly commits themself personally.
|
||||
- collective_first_person: the speaker explicitly commits a "we" group.
|
||||
- none: there is no explicit first-person commitment.
|
||||
2. What is the concise normalized action meaning?
|
||||
|
||||
Tentative possibility is not commitment. Suggestion or recommendation is not commitment. Impersonal necessity is not commitment. Passive future wording is not commitment. Rejection or negation is not positive commitment. Speaker identity does not convert collective "we" into individual commitment.
|
||||
|
||||
Preserve material limitations such as "nur im Technikum", "nur als Versuch", or "nur 20 Meter" in normalized_action_text. Keep normalized action text in the observation language. When commitment_form is not "none", normalized_action_text must be a non-empty string. When commitment_form is "none", normalized_action_text may be a non-empty action meaning or null.
|
||||
|
||||
Do not infer who is responsible. Do not decide whether an action is established. Do not output responsibility, responsibility scope, requested actor, owner, assignee, status, established, Action Item, protocol, confidence, semantic relations, graphs, decisions, or unresolved issues.
|
||||
|
||||
Return exactly this JSON shape and no additional fields:
|
||||
{{
|
||||
"observation_id": "observation ID",
|
||||
"commitment_form": "individual_first_person | collective_first_person | none",
|
||||
"normalized_action_text": "concise action meaning" | null
|
||||
}}
|
||||
|
||||
V3-style observation:
|
||||
{observation_json}
|
||||
"""
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required
|
||||
if missing:
|
||||
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _nonempty_text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def _validate_observations(observations: Any) -> None:
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise DerivationValidationError("observations must be a non-empty list")
|
||||
seen_observations: set[str] = set()
|
||||
seen_evidence: set[str] = set()
|
||||
for index, observation in enumerate(observations):
|
||||
location = f"observations[{index}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise DerivationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
|
||||
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if observation_id in seen_observations or evidence_id in seen_evidence:
|
||||
raise DerivationValidationError("observation and evidence provenance must be unique")
|
||||
seen_observations.add(observation_id)
|
||||
seen_evidence.add(evidence_id)
|
||||
_nonempty_text(observation["content"], f"{location}.content")
|
||||
_nonempty_text(observation["speaker"], f"{location}.speaker")
|
||||
for field in ("named_person", "addressee"):
|
||||
if observation[field] is not None:
|
||||
_nonempty_text(observation[field], f"{location}.{field}")
|
||||
|
||||
|
||||
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("Gold fixture must be an object")
|
||||
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
|
||||
if data["schema_version"] != GOLD_SCHEMA_VERSION:
|
||||
raise DerivationValidationError("unexpected Gold fixture schema_version")
|
||||
cases = data["cases"]
|
||||
if not isinstance(cases, list) or not cases:
|
||||
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for case in cases:
|
||||
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
|
||||
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
|
||||
if case_id in seen:
|
||||
raise DerivationValidationError(f"duplicate case ID: {case_id}")
|
||||
seen.add(case_id)
|
||||
_validate_observations(case["observations"])
|
||||
if len(case["observations"]) != 1:
|
||||
raise DerivationValidationError("collective Gold cases require exactly one observation")
|
||||
return cases
|
||||
|
||||
|
||||
def build_prompt(observations: list[dict[str, Any]]) -> str:
|
||||
_validate_observations(observations)
|
||||
if len(observations) != 1:
|
||||
raise DerivationValidationError("collective recognition requires exactly one observation")
|
||||
return PROMPT_TEMPLATE.format(
|
||||
observation_json=json.dumps(observations[0], ensure_ascii=False, indent=2)
|
||||
)
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||
if isinstance(value, dict):
|
||||
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||
if forbidden:
|
||||
raise DerivationValidationError(
|
||||
f"{location} contains forbidden semantic keys: {sorted(forbidden)}"
|
||||
)
|
||||
for key, item in value.items():
|
||||
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||
elif isinstance(value, list):
|
||||
for index, item in enumerate(value):
|
||||
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||
|
||||
|
||||
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
_validate_observations(observations)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
_reject_forbidden_keys(data)
|
||||
_exact_keys(data, RECOGNITION_KEYS, "output")
|
||||
observation_id = _nonempty_text(data["observation_id"], "output.observation_id")
|
||||
if observation_id not in {item["observation_id"] for item in observations}:
|
||||
raise DerivationValidationError("recognition references unknown observation")
|
||||
if data["commitment_form"] not in COMMITMENT_FORMS:
|
||||
raise DerivationValidationError("commitment_form has an unsupported value")
|
||||
action_text = data["normalized_action_text"]
|
||||
if action_text is not None:
|
||||
_nonempty_text(action_text, "output.normalized_action_text")
|
||||
if data["commitment_form"] != "none" and action_text is None:
|
||||
raise DerivationValidationError("non-none commitment requires normalized_action_text")
|
||||
return data
|
||||
|
||||
|
||||
def _bounded_due(observations: list[dict[str, Any]]) -> tuple[str | None, bool]:
|
||||
due_forms: set[str] = set()
|
||||
for observation in observations:
|
||||
content = observation["content"].casefold()
|
||||
for token in re.findall(r"\b[A-Za-zÄÖÜäöü]+\b", content):
|
||||
if token in WEEKDAYS:
|
||||
due_forms.add(WEEKDAYS[token])
|
||||
if re.search(r"\bnächste\s+woche\b", content):
|
||||
due_forms.add("nächste Woche")
|
||||
return (next(iter(due_forms)) if len(due_forms) == 1 else None, len(due_forms) <= 1)
|
||||
|
||||
|
||||
def _has_explicit_negation(observations: list[dict[str, Any]]) -> bool:
|
||||
return any(
|
||||
re.search(r"\b(?:nicht|kein(?:e|en|er|es)?|nein|not|no)\b", item["content"], re.IGNORECASE)
|
||||
for item in observations
|
||||
)
|
||||
|
||||
|
||||
def derive_collective_action(
|
||||
observations: list[dict[str, Any]], recognition: dict[str, Any]
|
||||
) -> tuple[dict[str, bool], dict[str, Any] | None]:
|
||||
validate_recognition(recognition, observations)
|
||||
by_id = {item["observation_id"]: item for item in observations}
|
||||
observation = by_id.get(recognition["observation_id"])
|
||||
due, deadline_consistent = _bounded_due(observations)
|
||||
gates = {
|
||||
"recognition_schema_valid": True,
|
||||
"observation_exists": observation is not None,
|
||||
"provenance_valid_and_unique": observation is not None and len({item["evidence_id"] for item in observations}) == len(observations),
|
||||
"collective_commitment_form": recognition["commitment_form"] == "collective_first_person",
|
||||
"normalized_action_present": isinstance(recognition["normalized_action_text"], str) and bool(recognition["normalized_action_text"].strip()),
|
||||
"deadline_supported_and_consistent": deadline_consistent,
|
||||
"no_explicit_negation": not _has_explicit_negation(observations),
|
||||
}
|
||||
if not all(gates.values()):
|
||||
return gates, None
|
||||
return gates, {
|
||||
"action_id": "action_1",
|
||||
"content": recognition["normalized_action_text"].strip(),
|
||||
"status": "established",
|
||||
"commitment_scope": "collective",
|
||||
"responsible_person": None,
|
||||
"due": due,
|
||||
"support": {
|
||||
"commitment": {
|
||||
"observation_id": observation["observation_id"],
|
||||
"evidence_id": observation["evidence_id"],
|
||||
}
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
|
||||
if not concepts:
|
||||
return True
|
||||
if not isinstance(text, str):
|
||||
return False
|
||||
folded = text.casefold()
|
||||
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
|
||||
|
||||
|
||||
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_recognition(recognition, case["observations"])
|
||||
gates, result = derive_collective_action(case["observations"], recognition)
|
||||
expected_recognition = case["expected_recognition"]
|
||||
expected_result = case["expected_result"]
|
||||
form_correct = recognition["commitment_form"] == expected_recognition["commitment_form"]
|
||||
action_correct = _concepts_present(recognition["normalized_action_text"], expected_recognition["action_concepts"])
|
||||
qualifier_preserved = _concepts_present(recognition["normalized_action_text"], expected_recognition["qualifier_concepts"])
|
||||
established = result is not None
|
||||
final_correct = established == expected_result["established"]
|
||||
if result is not None:
|
||||
final_correct = final_correct and result["due"] == expected_result["due"] and result["responsible_person"] is None and result["commitment_scope"] == "collective"
|
||||
owner_correct = result is None or result["responsible_person"] is None
|
||||
automatic_failure = (established and not expected_result["established"]) or not owner_correct or (established and not qualifier_preserved)
|
||||
semantic_correct = form_correct and action_correct and qualifier_preserved
|
||||
classification = "FAIL" if automatic_failure or not final_correct else ("PASS" if semantic_correct else "PARTIAL")
|
||||
return {
|
||||
"case_id": case["case_id"], "classification": classification,
|
||||
"commitment_form_correct": form_correct,
|
||||
"normalized_action_meaning_correct": action_correct,
|
||||
"material_qualifier_preserved": qualifier_preserved,
|
||||
"deterministic_gates_correct": final_correct,
|
||||
"final_result_correct": final_correct,
|
||||
"responsible_person_correctly_null": owner_correct,
|
||||
"due_correct": result is None or result["due"] == expected_result["due"],
|
||||
"unsupported_semantic_strengthening": recognition["commitment_form"] == "collective_first_person" and expected_recognition["commitment_form"] != "collective_first_person",
|
||||
"responsibility_status_leakage": False,
|
||||
"gates": gates, "result": result,
|
||||
}
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_gold_cases(args.cases)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
|
||||
evaluations: list[dict[str, Any]] = []
|
||||
successful_calls = 0
|
||||
technical_failures = 0
|
||||
started = time.perf_counter()
|
||||
for case in cases:
|
||||
case_dir = args.output / case["case_id"].lower()
|
||||
case_dir.mkdir()
|
||||
observations = case["observations"]
|
||||
_write_json(case_dir / "v3_style_input_observations.json", observations)
|
||||
prompt = build_prompt(observations)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
try:
|
||||
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||
successful_calls += 1
|
||||
except Exception as exc: # one recorded attempt; never retry
|
||||
technical_failures += 1
|
||||
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
|
||||
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
|
||||
_write_json(case_dir / "deterministic_gate_results.json", {})
|
||||
_write_json(case_dir / "final_derived_result.json", None)
|
||||
_write_json(case_dir / "evaluation.json", failure)
|
||||
evaluations.append(failure)
|
||||
continue
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
try:
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
|
||||
evaluation = evaluate_case(case, parsed)
|
||||
validation = {"valid": True, "error": None}
|
||||
gates, result = derive_collective_action(observations, parsed)
|
||||
except (DerivationValidationError, json.JSONDecodeError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "responsibility_status_leakage": "forbidden" in str(exc)}
|
||||
gates, result = {}, None
|
||||
_write_json(case_dir / "structural_validation.json", validation)
|
||||
_write_json(case_dir / "deterministic_gate_results.json", gates)
|
||||
_write_json(case_dir / "final_derived_result.json", result)
|
||||
_write_json(case_dir / "evaluation.json", evaluation)
|
||||
evaluations.append(evaluation)
|
||||
summary = {
|
||||
"experiment": "collective_commitment_gold_v0", "model": args.model,
|
||||
"successful_llm_call_count": successful_calls,
|
||||
"technical_failed_call_count": technical_failures,
|
||||
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||
"evaluations": evaluations,
|
||||
}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run isolated collective-commitment Gold experiment")
|
||||
parser.add_argument("cases", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=1024)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
summary = run_gold(parse_args())
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,367 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Isolated request/acceptance Gold reliability experiment."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from .experiment_h import (
|
||||
DEFAULT_ENDPOINT,
|
||||
DEFAULT_MODEL,
|
||||
DerivationValidationError,
|
||||
FORBIDDEN_LLM_KEYS,
|
||||
OBSERVATION_KEYS,
|
||||
call_ollama,
|
||||
)
|
||||
|
||||
|
||||
GOLD_SCHEMA_VERSION = "experimental-request-acceptance-gold-v0"
|
||||
RECOGNITION_SCHEMA_VERSION = "experimental-request-acceptance-recognition-v0"
|
||||
REQUEST_KEYS = {"observation_id", "is_concrete_request", "normalized_action_text"}
|
||||
ACCEPTANCE_KEYS = {
|
||||
"observation_id", "is_explicit_commitment", "same_requested_work",
|
||||
"normalized_action_text",
|
||||
}
|
||||
WEEKDAYS = {
|
||||
"monday": "Montag", "montag": "Montag", "tuesday": "Dienstag",
|
||||
"dienstag": "Dienstag", "wednesday": "Mittwoch", "mittwoch": "Mittwoch",
|
||||
"thursday": "Donnerstag", "donnerstag": "Donnerstag", "friday": "Freitag",
|
||||
"freitag": "Freitag", "saturday": "Samstag", "samstag": "Samstag",
|
||||
"sunday": "Sonntag", "sonntag": "Sonntag",
|
||||
}
|
||||
|
||||
PROMPT_TEMPLATE = """Recognize only a concrete directed request and a later explicit personal commitment in the supplied V3-style observations.
|
||||
|
||||
The input is observations only, not a transcript. Identify:
|
||||
1. A concrete request directed to the observation's explicit addressee, if one exists.
|
||||
2. A later response that explicitly commits its speaker to work, if one exists.
|
||||
3. Whether that explicit commitment concerns substantially the same requested work.
|
||||
|
||||
Lexical identity is not required: a contextual paraphrase may denote the same work. Mere acknowledgement, tentative or conditional language, collective "we" statements, impersonal necessity, suggestions, and statements that work should be done are not explicit personal commitments. A commitment to different work is an explicit commitment but not the same requested work.
|
||||
|
||||
Do not decide or output responsibility, requested actor, established status, Action Item status, protocol eligibility, confidence, semantic relations, or graphs. Do not answer who is responsible. Deterministic code applies those gates later.
|
||||
|
||||
Return exactly this JSON shape and no other fields. Use null for request or acceptance when no qualifying observation exists:
|
||||
{{
|
||||
"schema_version": "experimental-request-acceptance-recognition-v0",
|
||||
"request": null | {{
|
||||
"observation_id": "observation ID",
|
||||
"is_concrete_request": true,
|
||||
"normalized_action_text": "concise requested work"
|
||||
}},
|
||||
"acceptance": null | {{
|
||||
"observation_id": "observation ID",
|
||||
"is_explicit_commitment": true,
|
||||
"same_requested_work": true,
|
||||
"normalized_action_text": "concise committed work"
|
||||
}}
|
||||
}}
|
||||
|
||||
V3-style observations:
|
||||
{observations_json}
|
||||
"""
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required
|
||||
if missing:
|
||||
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _nonempty_text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("Gold fixture must be an object")
|
||||
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
|
||||
if data["schema_version"] != GOLD_SCHEMA_VERSION:
|
||||
raise DerivationValidationError("unexpected Gold fixture schema_version")
|
||||
cases = data["cases"]
|
||||
if not isinstance(cases, list) or not cases:
|
||||
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
|
||||
seen_cases: set[str] = set()
|
||||
for case in cases:
|
||||
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
|
||||
case_id = _nonempty_text(case["case_id"], "case_id")
|
||||
if case_id in seen_cases:
|
||||
raise DerivationValidationError(f"duplicate case ID: {case_id}")
|
||||
seen_cases.add(case_id)
|
||||
_validate_observations(case["observations"])
|
||||
return cases
|
||||
|
||||
|
||||
def _validate_observations(observations: Any) -> None:
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise DerivationValidationError("observations must be a non-empty list")
|
||||
seen_ids: set[str] = set()
|
||||
seen_evidence: set[str] = set()
|
||||
for index, observation in enumerate(observations):
|
||||
location = f"observations[{index}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise DerivationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
|
||||
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if observation_id in seen_ids or evidence_id in seen_evidence:
|
||||
raise DerivationValidationError("observation and evidence IDs must be unique")
|
||||
seen_ids.add(observation_id)
|
||||
seen_evidence.add(evidence_id)
|
||||
_nonempty_text(observation["content"], f"{location}.content")
|
||||
_nonempty_text(observation["speaker"], f"{location}.speaker")
|
||||
for field in ("named_person", "addressee"):
|
||||
if observation[field] is not None:
|
||||
_nonempty_text(observation[field], f"{location}.{field}")
|
||||
|
||||
|
||||
def build_prompt(observations: list[dict[str, Any]]) -> str:
|
||||
_validate_observations(observations)
|
||||
return PROMPT_TEMPLATE.format(
|
||||
observations_json=json.dumps(observations, ensure_ascii=False, indent=2)
|
||||
)
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||
if isinstance(value, dict):
|
||||
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||
if forbidden:
|
||||
raise DerivationValidationError(
|
||||
f"{location} contains forbidden semantic keys: {sorted(forbidden)}"
|
||||
)
|
||||
for key, item in value.items():
|
||||
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||
elif isinstance(value, list):
|
||||
for index, item in enumerate(value):
|
||||
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||
|
||||
|
||||
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
_reject_forbidden_keys(data)
|
||||
_exact_keys(data, {"schema_version", "request", "acceptance"}, "output")
|
||||
if data["schema_version"] != RECOGNITION_SCHEMA_VERSION:
|
||||
raise DerivationValidationError("unexpected recognition schema_version")
|
||||
known_ids = {item["observation_id"] for item in observations}
|
||||
request = data["request"]
|
||||
acceptance = data["acceptance"]
|
||||
if request is not None:
|
||||
if not isinstance(request, dict):
|
||||
raise DerivationValidationError("output.request must be an object or null")
|
||||
_exact_keys(request, REQUEST_KEYS, "output.request")
|
||||
if request["observation_id"] not in known_ids:
|
||||
raise DerivationValidationError("request references unknown observation")
|
||||
if not isinstance(request["is_concrete_request"], bool):
|
||||
raise DerivationValidationError("is_concrete_request must be boolean")
|
||||
_nonempty_text(request["normalized_action_text"], "request.normalized_action_text")
|
||||
if acceptance is not None:
|
||||
if not isinstance(acceptance, dict):
|
||||
raise DerivationValidationError("output.acceptance must be an object or null")
|
||||
_exact_keys(acceptance, ACCEPTANCE_KEYS, "output.acceptance")
|
||||
if acceptance["observation_id"] not in known_ids:
|
||||
raise DerivationValidationError("acceptance references unknown observation")
|
||||
for field in ("is_explicit_commitment", "same_requested_work"):
|
||||
if not isinstance(acceptance[field], bool):
|
||||
raise DerivationValidationError(f"{field} must be boolean")
|
||||
_nonempty_text(acceptance["normalized_action_text"], "acceptance.normalized_action_text")
|
||||
if request is not None and acceptance is not None and request["observation_id"] == acceptance["observation_id"]:
|
||||
raise DerivationValidationError("request and acceptance must reference different observations")
|
||||
return data
|
||||
|
||||
|
||||
def _bounded_due(observations: list[dict[str, Any]]) -> tuple[str | None, bool]:
|
||||
forms: set[str] = set()
|
||||
for observation in observations:
|
||||
for token in re.findall(r"\b[A-Za-zÄÖÜäöü]+\b", observation["content"].casefold()):
|
||||
if token in WEEKDAYS:
|
||||
forms.add(WEEKDAYS[token])
|
||||
return (next(iter(forms)) if len(forms) == 1 else None, len(forms) <= 1)
|
||||
|
||||
|
||||
def _strip_due(action_text: str) -> str:
|
||||
weekday = "|".join(re.escape(value) for value in WEEKDAYS)
|
||||
result = re.sub(rf"\s+(?:bis|by)\s+(?:{weekday})\b", "", action_text, flags=re.IGNORECASE)
|
||||
return result.strip(" .,:;-") or action_text.strip()
|
||||
|
||||
|
||||
def derive_action(
|
||||
observations: list[dict[str, Any]], recognition: dict[str, Any]
|
||||
) -> tuple[dict[str, bool], dict[str, Any] | None]:
|
||||
_validate_observations(observations)
|
||||
validate_recognition(recognition, observations)
|
||||
by_id = {item["observation_id"]: item for item in observations}
|
||||
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
|
||||
request_semantic = recognition["request"]
|
||||
acceptance_semantic = recognition["acceptance"]
|
||||
request = by_id.get(request_semantic["observation_id"]) if request_semantic else None
|
||||
acceptance = by_id.get(acceptance_semantic["observation_id"]) if acceptance_semantic else None
|
||||
due, deadline_consistent = _bounded_due(observations)
|
||||
gates = {
|
||||
"request_semantic_positive": request_semantic is not None and request_semantic["is_concrete_request"] is True,
|
||||
"request_observation_exists": request is not None,
|
||||
"request_has_addressee": request is not None and isinstance(request["addressee"], str) and bool(request["addressee"].strip()),
|
||||
"acceptance_semantic_positive": acceptance_semantic is not None and acceptance_semantic["is_explicit_commitment"] is True,
|
||||
"same_requested_work": acceptance_semantic is not None and acceptance_semantic["same_requested_work"] is True,
|
||||
"acceptance_observation_exists": acceptance is not None,
|
||||
"acceptance_after_request": request is not None and acceptance is not None and positions[acceptance["observation_id"]] > positions[request["observation_id"]],
|
||||
"acceptance_speaker_matches_addressee": request is not None and acceptance is not None and acceptance["speaker"] == request["addressee"],
|
||||
"provenance_valid_and_consistent": request is not None and acceptance is not None and request["evidence_id"] in {item["evidence_id"] for item in observations} and acceptance["evidence_id"] in {item["evidence_id"] for item in observations},
|
||||
"deadline_consistent": deadline_consistent,
|
||||
}
|
||||
if not all(gates.values()):
|
||||
return gates, None
|
||||
return gates, {
|
||||
"action_id": "action_1",
|
||||
"content": _strip_due(request_semantic["normalized_action_text"]),
|
||||
"status": "established",
|
||||
"requested_actor": request["addressee"],
|
||||
"responsible_person": acceptance["speaker"],
|
||||
"due": due,
|
||||
"support": {
|
||||
"request": {"observation_id": request["observation_id"], "evidence_id": request["evidence_id"]},
|
||||
"acceptance": {"observation_id": acceptance["observation_id"], "evidence_id": acceptance["evidence_id"]},
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
|
||||
observations = case["observations"]
|
||||
validate_recognition(recognition, observations)
|
||||
gates, result = derive_action(observations, recognition)
|
||||
expected_recognition = case["expected_recognition"]
|
||||
request_correct = (recognition["request"] is not None and recognition["request"]["is_concrete_request"]) == expected_recognition["request"]
|
||||
commitment_correct = (recognition["acceptance"] is not None and recognition["acceptance"]["is_explicit_commitment"]) == expected_recognition["commitment"]
|
||||
same_work_correct = (recognition["acceptance"] is not None and recognition["acceptance"]["same_requested_work"]) == expected_recognition["same_work"]
|
||||
expected = case["expected_result"]
|
||||
actual_established = result is not None
|
||||
final_correct = actual_established == expected["established"]
|
||||
if result is not None:
|
||||
final_correct = final_correct and all(
|
||||
result[key] == expected[key]
|
||||
for key in ("requested_actor", "responsible_person", "due")
|
||||
)
|
||||
else:
|
||||
final_correct = final_correct and expected["responsible_person"] is None
|
||||
semantic_correct = request_correct and commitment_correct and same_work_correct
|
||||
classification = "PASS" if final_correct and semantic_correct else ("PARTIAL" if final_correct else "FAIL")
|
||||
return {
|
||||
"case_id": case["case_id"], "classification": classification,
|
||||
"request_correct": request_correct, "commitment_correct": commitment_correct,
|
||||
"same_work_correct": same_work_correct, "deterministic_gates_correct": final_correct,
|
||||
"final_result_correct": final_correct, "result": result, "gates": gates,
|
||||
"unsupported_semantic_strengthening": not semantic_correct and actual_established,
|
||||
"responsibility_status_leakage": False,
|
||||
}
|
||||
|
||||
|
||||
def reevaluate_existing(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_gold_cases(args.cases)
|
||||
evaluations = []
|
||||
for case in cases:
|
||||
case_dir = args.output / case["case_id"].lower()
|
||||
parsed = json.loads(
|
||||
(case_dir / "parsed_semantic_recognition.json").read_text(encoding="utf-8")
|
||||
)
|
||||
evaluation = evaluate_case(case, parsed)
|
||||
_write_json(case_dir / "evaluation.json", evaluation)
|
||||
evaluations.append(evaluation)
|
||||
metadata = [
|
||||
json.loads(
|
||||
(args.output / case["case_id"].lower() / "ollama_metadata.json").read_text(
|
||||
encoding="utf-8"
|
||||
)
|
||||
)
|
||||
for case in cases
|
||||
]
|
||||
summary = {
|
||||
"experiment": "request_acceptance_gold_v0", "model": args.model,
|
||||
"llm_call_count": len(cases),
|
||||
"runtime_seconds": round(sum(item["elapsed_seconds"] for item in metadata), 3),
|
||||
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||
"evaluations": evaluations,
|
||||
}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_gold_cases(args.cases)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
|
||||
evaluations = []
|
||||
total_started = time.perf_counter()
|
||||
for case in cases:
|
||||
case_dir = args.output / case["case_id"].lower()
|
||||
case_dir.mkdir()
|
||||
_write_json(case_dir / "v3_style_input_observations.json", case["observations"])
|
||||
prompt = build_prompt(case["observations"])
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
|
||||
try:
|
||||
evaluation = evaluate_case(case, parsed)
|
||||
validation = {"valid": True, "error": None}
|
||||
except (DerivationValidationError, json.JSONDecodeError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "responsibility_status_leakage": "forbidden" in str(exc)}
|
||||
_write_json(case_dir / "structural_validation.json", validation)
|
||||
_write_json(case_dir / "evaluation.json", evaluation)
|
||||
evaluations.append(evaluation)
|
||||
summary = {
|
||||
"experiment": "request_acceptance_gold_v0", "model": args.model,
|
||||
"llm_call_count": len(cases), "runtime_seconds": round(time.perf_counter() - total_started, 3),
|
||||
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||
"evaluations": evaluations,
|
||||
}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run isolated request/acceptance Gold experiment")
|
||||
parser.add_argument("cases", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=1024)
|
||||
parser.add_argument("--reevaluate-existing", action="store_true")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
summary = reevaluate_existing(args) if args.reevaluate_existing else run_gold(args)
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["counts"]["FAIL"] == 0 else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,320 @@
|
||||
#!/usr/bin/env python3
|
||||
"""H-only request/acceptance recognition and deterministic action derivation."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
SCHEMA_VERSION = "experimental-controlled-semantic-recognition-h-v0"
|
||||
DEFAULT_MODEL = "qwen3.5:9B"
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
EXPECTED_PROVENANCE = {"obs_1": "e1", "obs_2": "e2"}
|
||||
OBSERVATION_KEYS = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
|
||||
SEMANTIC_KEYS = {"schema_version", "request", "acceptance"}
|
||||
REQUEST_KEYS = {"observation_id", "is_concrete_request", "normalized_action_text"}
|
||||
ACCEPTANCE_KEYS = {"observation_id", "is_explicit_commitment", "same_requested_work", "normalized_action_text"}
|
||||
FORBIDDEN_LLM_KEYS = {
|
||||
"responsible_person", "responsibility", "requested_actor", "status", "established",
|
||||
"action_item", "protocol_section", "protocol_category", "confidence", "relation",
|
||||
"relations", "graph", "decision", "open_question", "unresolved_issue",
|
||||
}
|
||||
|
||||
|
||||
class DerivationValidationError(ValueError):
|
||||
"""Raised when experiment input or LLM recognition violates the contract."""
|
||||
|
||||
|
||||
PROMPT_TEMPLATE = """Recognize only two narrow semantic facts in the supplied V3 observations.
|
||||
|
||||
The input contains V3 observations only, not a transcript. Answer only:
|
||||
1. Is obs_1 a concrete request directed to its recorded addressee?
|
||||
2. Does obs_2 explicitly commit its speaker to substantially the same requested work?
|
||||
|
||||
Lexical identity is not required. Conversational paraphrases such as "Prüfung der
|
||||
Messdaten" and "die Prüfung" may denote the same work when the supplied observation
|
||||
sequence clearly supports that reading.
|
||||
|
||||
Do not decide or output responsibility, requested actor, established status, Action
|
||||
Item status, protocol eligibility, confidence, semantic relations, or graphs. Do not
|
||||
answer who is responsible. Deterministic code will apply those gates later.
|
||||
|
||||
Return exactly this JSON shape and no other fields:
|
||||
{{
|
||||
"schema_version": "experimental-controlled-semantic-recognition-h-v0",
|
||||
"request": {{
|
||||
"observation_id": "obs_1",
|
||||
"is_concrete_request": true,
|
||||
"normalized_action_text": "concise requested work in the observation language"
|
||||
}},
|
||||
"acceptance": {{
|
||||
"observation_id": "obs_2",
|
||||
"is_explicit_commitment": true,
|
||||
"same_requested_work": true,
|
||||
"normalized_action_text": "concise accepted work in the observation language"
|
||||
}}
|
||||
}}
|
||||
|
||||
V3 observations:
|
||||
{observations_json}
|
||||
"""
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run the H-only controlled semantic derivation experiment.")
|
||||
parser.add_argument("observations", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=1024)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing, unknown = required - value.keys(), value.keys() - required
|
||||
if missing:
|
||||
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||
result = value.strip()
|
||||
if result.casefold() == "null":
|
||||
raise DerivationValidationError(f"{location} must not be the string 'null'")
|
||||
return result
|
||||
|
||||
|
||||
def load_v3_observations(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("V3 input must be an object")
|
||||
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "V3 input")
|
||||
if data["schema_version"] != "experimental-evidence-observations-v3":
|
||||
raise DerivationValidationError("V3 input has an unexpected schema_version")
|
||||
observations = data["observations"]
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise DerivationValidationError("V3 observations must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for index, observation in enumerate(observations):
|
||||
location = f"V3 observations[{index}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise DerivationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
|
||||
if observation_id in seen:
|
||||
raise DerivationValidationError(f"duplicate observation ID: {observation_id}")
|
||||
seen.add(observation_id)
|
||||
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if observation_id not in EXPECTED_PROVENANCE:
|
||||
raise DerivationValidationError(f"unknown H observation ID: {observation_id}")
|
||||
if EXPECTED_PROVENANCE[observation_id] != evidence_id:
|
||||
raise DerivationValidationError(f"inconsistent evidence provenance for {observation_id}")
|
||||
_text(observation["content"], f"{location}.content")
|
||||
_text(observation["speaker"], f"{location}.speaker")
|
||||
for field in ("named_person", "addressee"):
|
||||
if observation[field] is not None:
|
||||
_text(observation[field], f"{location}.{field}")
|
||||
if seen != set(EXPECTED_PROVENANCE):
|
||||
raise DerivationValidationError("H input must contain exactly obs_1/e1 and obs_2/e2")
|
||||
return observations
|
||||
|
||||
|
||||
def build_prompt(observations: list[dict[str, Any]]) -> str:
|
||||
validate_observation_sequence(observations)
|
||||
return PROMPT_TEMPLATE.format(observations_json=json.dumps(observations, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
def validate_observation_sequence(observations: list[dict[str, Any]]) -> None:
|
||||
if [item.get("observation_id") for item in observations] != ["obs_1", "obs_2"]:
|
||||
raise DerivationValidationError("H observations must be ordered obs_1, obs_2")
|
||||
for observation in observations:
|
||||
if EXPECTED_PROVENANCE.get(observation.get("observation_id")) != observation.get("evidence_id"):
|
||||
raise DerivationValidationError("H observation provenance is inconsistent")
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||
if isinstance(value, dict):
|
||||
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||
if forbidden:
|
||||
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
|
||||
for key, item in value.items():
|
||||
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||
elif isinstance(value, list):
|
||||
for index, item in enumerate(value):
|
||||
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||
|
||||
|
||||
def validate_semantic_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
_reject_forbidden_keys(data)
|
||||
_exact_keys(data, SEMANTIC_KEYS, "output")
|
||||
if data["schema_version"] != SCHEMA_VERSION:
|
||||
raise DerivationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
|
||||
request, acceptance = data["request"], data["acceptance"]
|
||||
if not isinstance(request, dict) or not isinstance(acceptance, dict):
|
||||
raise DerivationValidationError("request and acceptance must be objects")
|
||||
_exact_keys(request, REQUEST_KEYS, "output.request")
|
||||
_exact_keys(acceptance, ACCEPTANCE_KEYS, "output.acceptance")
|
||||
known_ids = {item["observation_id"] for item in observations}
|
||||
for location, item in (("output.request", request), ("output.acceptance", acceptance)):
|
||||
observation_id = _text(item["observation_id"], f"{location}.observation_id")
|
||||
if observation_id not in known_ids:
|
||||
raise DerivationValidationError(f"{location} references unknown observation: {observation_id}")
|
||||
_text(item["normalized_action_text"], f"{location}.normalized_action_text")
|
||||
for field, value in (
|
||||
("output.request.is_concrete_request", request["is_concrete_request"]),
|
||||
("output.acceptance.is_explicit_commitment", acceptance["is_explicit_commitment"]),
|
||||
("output.acceptance.same_requested_work", acceptance["same_requested_work"]),
|
||||
):
|
||||
if not isinstance(value, bool):
|
||||
raise DerivationValidationError(f"{field} must be boolean")
|
||||
if request["observation_id"] == acceptance["observation_id"]:
|
||||
raise DerivationValidationError("request and acceptance must reference different observations")
|
||||
return data
|
||||
|
||||
|
||||
def _bounded_due(observations: list[dict[str, Any]]) -> tuple[str | None, bool]:
|
||||
weekday_forms = {
|
||||
"monday": "Montag", "montag": "Montag",
|
||||
"tuesday": "Dienstag", "dienstag": "Dienstag",
|
||||
"wednesday": "Mittwoch", "mittwoch": "Mittwoch",
|
||||
"thursday": "Donnerstag", "donnerstag": "Donnerstag",
|
||||
"friday": "Freitag", "freitag": "Freitag",
|
||||
"saturday": "Samstag", "samstag": "Samstag",
|
||||
"sunday": "Sonntag", "sonntag": "Sonntag",
|
||||
}
|
||||
forms: set[str] = set()
|
||||
for observation in observations:
|
||||
for token in re.findall(r"\b[A-Za-zÄÖÜäöü]+\b", observation["content"].casefold()):
|
||||
if token in weekday_forms:
|
||||
forms.add(weekday_forms[token])
|
||||
return (next(iter(forms)) if len(forms) == 1 else None, len(forms) <= 1)
|
||||
|
||||
|
||||
def _remove_bounded_due_from_action(action_text: str) -> str:
|
||||
result = re.sub(
|
||||
r"\s+(?:bis|by)\s+(?:Friday|Freitag)\b", "", action_text,
|
||||
flags=re.IGNORECASE,
|
||||
).strip(" .,:;-")
|
||||
return result or action_text.strip()
|
||||
|
||||
|
||||
def derive_action(
|
||||
observations: list[dict[str, Any]], recognition: dict[str, Any]
|
||||
) -> tuple[dict[str, bool], dict[str, Any] | None]:
|
||||
by_id = {item["observation_id"]: item for item in observations}
|
||||
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
|
||||
request_semantic = recognition["request"]
|
||||
acceptance_semantic = recognition["acceptance"]
|
||||
request = by_id.get(request_semantic["observation_id"])
|
||||
acceptance = by_id.get(acceptance_semantic["observation_id"])
|
||||
due, deadline_consistent = _bounded_due(observations)
|
||||
gates = {
|
||||
"request_semantic_positive": request_semantic["is_concrete_request"] is True,
|
||||
"request_observation_exists": request is not None,
|
||||
"request_has_addressee": request is not None and isinstance(request.get("addressee"), str) and bool(request["addressee"].strip()),
|
||||
"acceptance_semantic_positive": acceptance_semantic["is_explicit_commitment"] is True,
|
||||
"same_requested_work": acceptance_semantic["same_requested_work"] is True,
|
||||
"acceptance_observation_exists": acceptance is not None,
|
||||
"acceptance_after_request": request is not None and acceptance is not None and positions[acceptance["observation_id"]] > positions[request["observation_id"]],
|
||||
"acceptance_speaker_matches_addressee": request is not None and acceptance is not None and acceptance["speaker"] == request["addressee"],
|
||||
"provenance_valid_and_consistent": request is not None and acceptance is not None and EXPECTED_PROVENANCE.get(request["observation_id"]) == request["evidence_id"] and EXPECTED_PROVENANCE.get(acceptance["observation_id"]) == acceptance["evidence_id"],
|
||||
"deadline_consistent": deadline_consistent,
|
||||
}
|
||||
if not all(gates.values()):
|
||||
return gates, None
|
||||
action_text = _remove_bounded_due_from_action(
|
||||
request_semantic["normalized_action_text"]
|
||||
)
|
||||
result = {
|
||||
"action_id": "action_1",
|
||||
"content": action_text,
|
||||
"status": "established",
|
||||
"requested_actor": request["addressee"],
|
||||
"responsible_person": acceptance["speaker"],
|
||||
"due": due,
|
||||
"support": {
|
||||
"request": {"observation_id": request["observation_id"], "evidence_id": request["evidence_id"]},
|
||||
"acceptance": {"observation_id": acceptance["observation_id"], "evidence_id": acceptance["evidence_id"]},
|
||||
},
|
||||
}
|
||||
return gates, result
|
||||
|
||||
|
||||
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
|
||||
|
||||
|
||||
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
|
||||
started = time.perf_counter()
|
||||
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
|
||||
elapsed = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
raw = body.get("response") if isinstance(body, dict) else None
|
||||
if not isinstance(raw, str) or not raw.strip():
|
||||
raise ValueError("Ollama returned no usable response text")
|
||||
metadata = {"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3), "total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"), "prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"), "eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"), "configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0}}
|
||||
return raw.strip(), metadata
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
observations = load_v3_observations(args.observations)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(args.output / "v3_input_observations.json", observations)
|
||||
prompt = build_prompt(observations)
|
||||
(args.output / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
started = time.perf_counter()
|
||||
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||
(args.output / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(args.output / "ollama_metadata.json", metadata)
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(args.output / "parsed_semantic_recognition.json", parsed)
|
||||
try:
|
||||
validate_semantic_recognition(parsed, observations)
|
||||
validation = {"valid": True, "error": None}
|
||||
gates, result = derive_action(observations, parsed)
|
||||
except DerivationValidationError as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
gates, result = {}, None
|
||||
_write_json(args.output / "structural_validation.json", validation)
|
||||
_write_json(args.output / "deterministic_gate_results.json", gates)
|
||||
_write_json(args.output / "final_derived_result.json", result)
|
||||
summary = {"experiment": "controlled_semantic_derivation_h_v0", "model": args.model, "llm_call_count": 1, "runtime_seconds": round(time.perf_counter() - started, 3), "semantic_recognition_valid": validation["valid"], "all_gates_passed": bool(gates) and all(gates.values()), "action_established": result is not None}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
summary = run_experiment(args)
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["action_established"] else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,277 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Isolated evidence-near Negative Act Form classification experiment."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from .experiment_h import (
|
||||
DEFAULT_ENDPOINT,
|
||||
DEFAULT_MODEL,
|
||||
DerivationValidationError,
|
||||
OBSERVATION_KEYS,
|
||||
build_ollama_payload,
|
||||
call_ollama,
|
||||
)
|
||||
|
||||
|
||||
GOLD_SCHEMA_VERSION = "experimental-negative-act-form-gold-v0"
|
||||
RECOGNITION_KEYS = {"observation_id", "negative_act_form", "normalized_action_text"}
|
||||
NEGATIVE_ACT_FORMS = {
|
||||
"explicit_non_pursuit", "personal_preference", "recommendation",
|
||||
"temporary_non_action", "none",
|
||||
}
|
||||
FORBIDDEN_LLM_KEYS = {
|
||||
"rejection_form", "explicitly_rejected", "status", "decision", "outcome",
|
||||
"topic_status", "responsible_person", "responsibility", "owner",
|
||||
"requested_actor", "action_item", "protocol", "protocol_category",
|
||||
"confidence", "relation", "relations", "graph", "unresolved_issue",
|
||||
}
|
||||
|
||||
PROMPT_TEMPLATE = """Classify only the negative semantic form expressed by the candidate observation, using earlier supplied V3-style observations only as local context for pronouns or shortened references.
|
||||
|
||||
The candidate observation is {candidate_observation_id}.
|
||||
|
||||
Choose exactly one negative_act_form:
|
||||
- explicit_non_pursuit: explicitly states that an action, option, collaboration, or course will not be continued or pursued. This is stronger than preference, advice, or temporary delay.
|
||||
- personal_preference: the speaker states what they personally would or would not do, without establishing collective non-pursuit.
|
||||
- recommendation: the speaker advises for or against an action without establishing abandonment.
|
||||
- temporary_non_action: the action is postponed, deferred, or explicitly not done for now without abandonment.
|
||||
- none: none of those four forms is present, including mere concern, uncertainty, negative sentiment, or factual negation.
|
||||
|
||||
Do not collapse non-pursuit into temporary non-action. Do not convert a personal conditional preference into collective non-pursuit. Do not convert advice into non-pursuit. Speaker identity does not change personal preference into collective non-pursuit.
|
||||
|
||||
When the form is not none, return concise normalized action meaning. Resolve a pronoun only from the supplied local context. If its target is genuinely ambiguous, return none rather than guessing. When the form is none, normalized_action_text must be null. Keep normalized action text in the observation language.
|
||||
|
||||
Do not derive or output rejection, status, decision, outcome, topic closure, responsibility, ownership, Action Item, protocol category, confidence, relations, graphs, or unresolved issues.
|
||||
|
||||
Return exactly this JSON shape and no additional fields:
|
||||
{{
|
||||
"observation_id": "{candidate_observation_id}",
|
||||
"negative_act_form": "explicit_non_pursuit | personal_preference | recommendation | temporary_non_action | none",
|
||||
"normalized_action_text": "concise action meaning" | null
|
||||
}}
|
||||
|
||||
V3-style observations:
|
||||
{observations_json}
|
||||
"""
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required
|
||||
if missing:
|
||||
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _nonempty_text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def _validate_observations(observations: Any) -> None:
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise DerivationValidationError("observations must be a non-empty list")
|
||||
seen_observations: set[str] = set()
|
||||
seen_evidence: set[str] = set()
|
||||
for index, observation in enumerate(observations):
|
||||
location = f"observations[{index}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise DerivationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
|
||||
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if observation_id in seen_observations or evidence_id in seen_evidence:
|
||||
raise DerivationValidationError("observation and evidence provenance must be unique")
|
||||
seen_observations.add(observation_id)
|
||||
seen_evidence.add(evidence_id)
|
||||
_nonempty_text(observation["content"], f"{location}.content")
|
||||
_nonempty_text(observation["speaker"], f"{location}.speaker")
|
||||
for field in ("named_person", "addressee"):
|
||||
if observation[field] is not None:
|
||||
_nonempty_text(observation[field], f"{location}.{field}")
|
||||
|
||||
|
||||
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("Gold fixture must be an object")
|
||||
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
|
||||
if data["schema_version"] != GOLD_SCHEMA_VERSION:
|
||||
raise DerivationValidationError("unexpected Gold fixture schema_version")
|
||||
cases = data["cases"]
|
||||
if not isinstance(cases, list) or not cases:
|
||||
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for case in cases:
|
||||
_exact_keys(case, {"case_id", "description", "observations", "expected"}, "Gold case")
|
||||
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
|
||||
if case_id in seen:
|
||||
raise DerivationValidationError(f"duplicate case ID: {case_id}")
|
||||
seen.add(case_id)
|
||||
_validate_observations(case["observations"])
|
||||
if len(case["observations"]) not in (1, 2):
|
||||
raise DerivationValidationError("Negative Act cases require one or two observations")
|
||||
return cases
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
observations = case["observations"]
|
||||
_validate_observations(observations)
|
||||
candidate_id = observations[-1]["observation_id"]
|
||||
return PROMPT_TEMPLATE.format(
|
||||
candidate_observation_id=candidate_id,
|
||||
observations_json=json.dumps(observations, ensure_ascii=False, indent=2),
|
||||
)
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic classification must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||
if isinstance(value, dict):
|
||||
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||
if forbidden:
|
||||
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
|
||||
for key, item in value.items():
|
||||
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||
elif isinstance(value, list):
|
||||
for index, item in enumerate(value):
|
||||
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||
|
||||
|
||||
def validate_classification(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
_validate_observations(observations)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic classification must be an object")
|
||||
_reject_forbidden_keys(data)
|
||||
_exact_keys(data, RECOGNITION_KEYS, "output")
|
||||
observation_id = _nonempty_text(data["observation_id"], "output.observation_id")
|
||||
if observation_id not in {item["observation_id"] for item in observations}:
|
||||
raise DerivationValidationError("classification references unknown observation")
|
||||
form = data["negative_act_form"]
|
||||
if form not in NEGATIVE_ACT_FORMS:
|
||||
raise DerivationValidationError("negative_act_form has an unsupported value")
|
||||
action_text = data["normalized_action_text"]
|
||||
if form == "none":
|
||||
if action_text is not None:
|
||||
raise DerivationValidationError("none form requires null normalized_action_text")
|
||||
else:
|
||||
_nonempty_text(action_text, "output.normalized_action_text")
|
||||
return data
|
||||
|
||||
|
||||
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
|
||||
if not concepts:
|
||||
return text is None
|
||||
if not isinstance(text, str):
|
||||
return False
|
||||
folded = text.casefold()
|
||||
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
|
||||
|
||||
|
||||
def evaluate_case(case: dict[str, Any], classification: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_classification(classification, case["observations"])
|
||||
expected = case["expected"]
|
||||
observation_correct = classification["observation_id"] == expected["observation_id"]
|
||||
form_correct = classification["negative_act_form"] == expected["negative_act_form"]
|
||||
action_correct = _concepts_present(classification["normalized_action_text"], expected["action_concepts"])
|
||||
unsupported_strengthening = expected["negative_act_form"] == "none" and classification["negative_act_form"] != "none"
|
||||
classification_label = "PASS" if observation_correct and form_correct and action_correct else ("PARTIAL" if observation_correct and form_correct else "FAIL")
|
||||
return {
|
||||
"case_id": case["case_id"], "classification": classification_label,
|
||||
"expected_negative_act_form": expected["negative_act_form"],
|
||||
"actual_negative_act_form": classification["negative_act_form"],
|
||||
"observation_id_correct": observation_correct,
|
||||
"normalized_action_meaning_correct": action_correct,
|
||||
"unsupported_semantic_strengthening": unsupported_strengthening,
|
||||
"normative_leakage": False,
|
||||
}
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_gold_cases(args.cases)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
|
||||
evaluations: list[dict[str, Any]] = []
|
||||
successful_calls = 0
|
||||
technical_failures = 0
|
||||
started = time.perf_counter()
|
||||
for case in cases:
|
||||
case_dir = args.output / case["case_id"].lower()
|
||||
case_dir.mkdir()
|
||||
observations = case["observations"]
|
||||
_write_json(case_dir / "v3_style_input_observations.json", observations)
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
try:
|
||||
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||
successful_calls += 1
|
||||
except Exception as exc: # one recorded attempt; never retry
|
||||
technical_failures += 1
|
||||
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
|
||||
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
|
||||
_write_json(case_dir / "evaluation.json", failure)
|
||||
evaluations.append(failure)
|
||||
continue
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
try:
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_semantic_classification.json", parsed)
|
||||
evaluation = evaluate_case(case, parsed)
|
||||
validation = {"valid": True, "error": None}
|
||||
except (DerivationValidationError, json.JSONDecodeError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)}
|
||||
_write_json(case_dir / "structural_validation.json", validation)
|
||||
_write_json(case_dir / "evaluation.json", evaluation)
|
||||
evaluations.append(evaluation)
|
||||
summary = {
|
||||
"experiment": "negative_act_form_v0", "model": args.model,
|
||||
"successful_llm_call_count": successful_calls,
|
||||
"technical_failed_call_count": technical_failures,
|
||||
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||
"evaluations": evaluations,
|
||||
}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run isolated Negative Act Form experiment")
|
||||
parser.add_argument("cases", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=1024)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
summary = run_experiment(parse_args())
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,347 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Isolated explicit-action-rejection Gold reliability experiment."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from .experiment_h import (
|
||||
DEFAULT_ENDPOINT,
|
||||
DEFAULT_MODEL,
|
||||
DerivationValidationError,
|
||||
OBSERVATION_KEYS,
|
||||
call_ollama,
|
||||
)
|
||||
|
||||
|
||||
GOLD_SCHEMA_VERSION = "experimental-explicit-rejection-gold-v0"
|
||||
RECOGNITION_KEYS = {
|
||||
"rejection_observation_id", "target_observation_id", "rejection_form",
|
||||
"normalized_rejected_action_text",
|
||||
}
|
||||
REJECTION_FORMS = {"explicit_action_rejection", "none"}
|
||||
FORBIDDEN_LLM_KEYS = {
|
||||
"decision", "decision_status", "outcome", "topic_status", "closed",
|
||||
"agreement", "responsible_person", "responsibility", "responsibility_scope",
|
||||
"owner", "ownership", "assignee", "requested_actor", "status",
|
||||
"explicitly_rejected", "action_item", "protocol", "protocol_category",
|
||||
"confidence", "relation", "relations", "graph", "unresolved_issue",
|
||||
}
|
||||
|
||||
PROMPT_TEMPLATE = """Recognize only whether the candidate rejection observation explicitly rejects a concrete action, option, proposal, or future course of action in this small local set of V3-style observations.
|
||||
|
||||
Answer only:
|
||||
1. Does the candidate rejection observation explicitly reject, abandon, discontinue, or rule out a concrete action, option, proposal, or future course of action?
|
||||
2. If yes, which supplied observation identifies the rejected target?
|
||||
3. What is the concise normalized meaning of the rejected action or option?
|
||||
|
||||
The candidate rejection observation is {rejection_observation_id}.
|
||||
|
||||
Use explicit_action_rejection only for an asserted rejection, abandonment, discontinuation, or non-pursuit with a concrete locally resolvable target. Personal preference is not meeting-level explicit rejection. Concern or objection without refusal is not rejection. Uncertainty is not rejection. Negative recommendation or advice is not established rejection. Deferral is not rejection. "Not yet" or temporary non-action is not abandonment. Factual negation is not action rejection. Lack of commitment is not rejection.
|
||||
|
||||
The rejected target may be self-contained in the candidate observation or introduced by one earlier supplied observation. Choose only among supplied observation IDs. If the target is ambiguous or unresolved, return rejection_form none. Preserve material scope limitations in normalized_rejected_action_text. Ignore a separate positive alternative when describing the rejected target. Keep normalized text in the observation language.
|
||||
|
||||
Do not infer responsibility, ownership, decision status, final outcome, topic closure, protocol status, confidence, relations, graphs, or unresolved issues. Do not answer whether this was finally decided, what the meeting outcome was, who is responsible, or whether the topic is closed.
|
||||
|
||||
Return exactly this JSON shape and no additional fields:
|
||||
{{
|
||||
"rejection_observation_id": "{rejection_observation_id}",
|
||||
"target_observation_id": "supplied observation ID" | null,
|
||||
"rejection_form": "explicit_action_rejection | none",
|
||||
"normalized_rejected_action_text": "concise rejected target" | null
|
||||
}}
|
||||
|
||||
For rejection_form none, target_observation_id and normalized_rejected_action_text must both be null.
|
||||
|
||||
V3-style observations:
|
||||
{observations_json}
|
||||
"""
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required
|
||||
if missing:
|
||||
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _nonempty_text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def _validate_observations(observations: Any) -> None:
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise DerivationValidationError("observations must be a non-empty list")
|
||||
seen_observations: set[str] = set()
|
||||
seen_evidence: set[str] = set()
|
||||
for index, observation in enumerate(observations):
|
||||
location = f"observations[{index}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise DerivationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
|
||||
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if observation_id in seen_observations:
|
||||
raise DerivationValidationError("observation IDs must be unique")
|
||||
if evidence_id in seen_evidence:
|
||||
raise DerivationValidationError("evidence provenance must be unique and consistent")
|
||||
seen_observations.add(observation_id)
|
||||
seen_evidence.add(evidence_id)
|
||||
_nonempty_text(observation["content"], f"{location}.content")
|
||||
_nonempty_text(observation["speaker"], f"{location}.speaker")
|
||||
for field in ("named_person", "addressee"):
|
||||
if observation[field] is not None:
|
||||
_nonempty_text(observation[field], f"{location}.{field}")
|
||||
|
||||
|
||||
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("Gold fixture must be an object")
|
||||
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
|
||||
if data["schema_version"] != GOLD_SCHEMA_VERSION:
|
||||
raise DerivationValidationError("unexpected Gold fixture schema_version")
|
||||
cases = data["cases"]
|
||||
if not isinstance(cases, list) or not cases:
|
||||
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for case in cases:
|
||||
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
|
||||
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
|
||||
if case_id in seen:
|
||||
raise DerivationValidationError(f"duplicate case ID: {case_id}")
|
||||
seen.add(case_id)
|
||||
_validate_observations(case["observations"])
|
||||
if len(case["observations"]) not in (1, 2):
|
||||
raise DerivationValidationError("rejection Gold cases require one or two observations")
|
||||
return cases
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
observations = case["observations"]
|
||||
_validate_observations(observations)
|
||||
rejection_observation_id = observations[-1]["observation_id"]
|
||||
return PROMPT_TEMPLATE.format(
|
||||
rejection_observation_id=rejection_observation_id,
|
||||
observations_json=json.dumps(observations, ensure_ascii=False, indent=2),
|
||||
)
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||
if isinstance(value, dict):
|
||||
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||
if forbidden:
|
||||
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
|
||||
for key, item in value.items():
|
||||
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||
elif isinstance(value, list):
|
||||
for index, item in enumerate(value):
|
||||
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||
|
||||
|
||||
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
_validate_observations(observations)
|
||||
if not isinstance(data, dict):
|
||||
raise DerivationValidationError("semantic recognition must be an object")
|
||||
_reject_forbidden_keys(data)
|
||||
_exact_keys(data, RECOGNITION_KEYS, "output")
|
||||
rejection_id = _nonempty_text(data["rejection_observation_id"], "output.rejection_observation_id")
|
||||
known_ids = {item["observation_id"] for item in observations}
|
||||
if rejection_id not in known_ids:
|
||||
raise DerivationValidationError("unknown rejection observation ID")
|
||||
form = data["rejection_form"]
|
||||
if form not in REJECTION_FORMS:
|
||||
raise DerivationValidationError("rejection_form has an unsupported value")
|
||||
target_id = data["target_observation_id"]
|
||||
action_text = data["normalized_rejected_action_text"]
|
||||
if form == "none":
|
||||
if target_id is not None:
|
||||
raise DerivationValidationError("none rejection must have null target_observation_id")
|
||||
if action_text is not None:
|
||||
raise DerivationValidationError("none rejection must have null normalized_rejected_action_text")
|
||||
else:
|
||||
target_id = _nonempty_text(target_id, "output.target_observation_id")
|
||||
if target_id not in known_ids:
|
||||
raise DerivationValidationError("unknown target observation ID")
|
||||
_nonempty_text(action_text, "output.normalized_rejected_action_text")
|
||||
return data
|
||||
|
||||
|
||||
def derive_rejection(
|
||||
observations: list[dict[str, Any]], recognition: dict[str, Any]
|
||||
) -> tuple[dict[str, bool], dict[str, Any] | None]:
|
||||
validate_recognition(recognition, observations)
|
||||
by_id = {item["observation_id"]: item for item in observations}
|
||||
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
|
||||
rejection = by_id.get(recognition["rejection_observation_id"])
|
||||
target_id = recognition["target_observation_id"]
|
||||
target = by_id.get(target_id) if target_id is not None else None
|
||||
gates = {
|
||||
"recognition_schema_valid": True,
|
||||
"explicit_action_rejection": recognition["rejection_form"] == "explicit_action_rejection",
|
||||
"rejection_observation_exists": rejection is not None,
|
||||
"target_observation_exists": target is not None,
|
||||
"observation_ids_valid_and_unique": len(by_id) == len(observations),
|
||||
"evidence_provenance_valid_unique_consistent": len({item["evidence_id"] for item in observations}) == len(observations),
|
||||
"target_same_or_before_rejection": target is not None and rejection is not None and positions[target["observation_id"]] <= positions[rejection["observation_id"]],
|
||||
"normalized_rejected_action_present": isinstance(recognition["normalized_rejected_action_text"], str) and bool(recognition["normalized_rejected_action_text"].strip()),
|
||||
"target_local_to_case": target_id in by_id if target_id is not None else False,
|
||||
"schema_state_consistent": recognition["rejection_form"] == "explicit_action_rejection" and target_id is not None,
|
||||
"referenced_provenance_available": target is not None and rejection is not None and bool(target["evidence_id"]) and bool(rejection["evidence_id"]),
|
||||
}
|
||||
if not all(gates.values()):
|
||||
return gates, None
|
||||
return gates, {
|
||||
"rejection_id": "rejection_1",
|
||||
"content": recognition["normalized_rejected_action_text"].strip(),
|
||||
"status": "explicitly_rejected",
|
||||
"support": {
|
||||
"target": {"observation_id": target["observation_id"], "evidence_id": target["evidence_id"]},
|
||||
"rejection": {"observation_id": rejection["observation_id"], "evidence_id": rejection["evidence_id"]},
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
|
||||
if not concepts:
|
||||
return True
|
||||
if not isinstance(text, str):
|
||||
return False
|
||||
folded = text.casefold()
|
||||
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
|
||||
|
||||
|
||||
def _contains_forbidden_concept(text: str | None, concepts: list[str]) -> bool:
|
||||
return isinstance(text, str) and any(concept.casefold() in text.casefold() for concept in concepts)
|
||||
|
||||
|
||||
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_recognition(recognition, case["observations"])
|
||||
gates, result = derive_rejection(case["observations"], recognition)
|
||||
expected = case["expected_recognition"]
|
||||
expected_result = case["expected_result"]
|
||||
form_correct = recognition["rejection_form"] == expected["rejection_form"]
|
||||
rejection_observation_correct = recognition["rejection_observation_id"] == expected["rejection_observation_id"]
|
||||
target_correct = recognition["target_observation_id"] == expected["target_observation_id"]
|
||||
action_correct = _concepts_present(recognition["normalized_rejected_action_text"], expected["action_concepts"])
|
||||
qualifier_preserved = _concepts_present(recognition["normalized_rejected_action_text"], expected["qualifier_concepts"])
|
||||
alternative_absorbed = _contains_forbidden_concept(recognition["normalized_rejected_action_text"], expected["forbidden_action_concepts"])
|
||||
derived = result is not None
|
||||
final_correct = derived == expected_result["explicitly_rejected"]
|
||||
if result is not None:
|
||||
final_correct = final_correct and result["status"] == "explicitly_rejected"
|
||||
semantic_correct = form_correct and rejection_observation_correct and target_correct and action_correct and qualifier_preserved and not alternative_absorbed
|
||||
automatic_failure = (derived and not expected_result["explicitly_rejected"]) or (derived and not target_correct) or (derived and not qualifier_preserved) or alternative_absorbed
|
||||
classification = "FAIL" if automatic_failure or not final_correct else ("PASS" if semantic_correct else "PARTIAL")
|
||||
return {
|
||||
"case_id": case["case_id"], "classification": classification,
|
||||
"rejection_form_correct": form_correct,
|
||||
"rejection_observation_correct": rejection_observation_correct,
|
||||
"target_observation_correct": target_correct,
|
||||
"normalized_rejected_action_correct": action_correct,
|
||||
"material_qualifiers_preserved": qualifier_preserved,
|
||||
"positive_alternative_absorbed": alternative_absorbed,
|
||||
"deterministic_gates_correct": final_correct,
|
||||
"final_result_correct": final_correct,
|
||||
"unsupported_semantic_strengthening": recognition["rejection_form"] == "explicit_action_rejection" and expected["rejection_form"] == "none",
|
||||
"normative_leakage": False,
|
||||
"gates": gates, "result": result,
|
||||
}
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_gold_cases(args.cases)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
|
||||
evaluations: list[dict[str, Any]] = []
|
||||
successful_calls = 0
|
||||
technical_failures = 0
|
||||
started = time.perf_counter()
|
||||
for case in cases:
|
||||
case_dir = args.output / case["case_id"].lower()
|
||||
case_dir.mkdir()
|
||||
observations = case["observations"]
|
||||
_write_json(case_dir / "v3_style_input_observations.json", observations)
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
try:
|
||||
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||
successful_calls += 1
|
||||
except Exception as exc: # one recorded attempt; never retry
|
||||
technical_failures += 1
|
||||
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
|
||||
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
|
||||
_write_json(case_dir / "deterministic_gate_results.json", {})
|
||||
_write_json(case_dir / "final_derived_result.json", None)
|
||||
_write_json(case_dir / "evaluation.json", failure)
|
||||
evaluations.append(failure)
|
||||
continue
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
try:
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
|
||||
evaluation = evaluate_case(case, parsed)
|
||||
validation = {"valid": True, "error": None}
|
||||
gates, result = derive_rejection(observations, parsed)
|
||||
except (DerivationValidationError, json.JSONDecodeError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)}
|
||||
gates, result = {}, None
|
||||
_write_json(case_dir / "structural_validation.json", validation)
|
||||
_write_json(case_dir / "deterministic_gate_results.json", gates)
|
||||
_write_json(case_dir / "final_derived_result.json", result)
|
||||
_write_json(case_dir / "evaluation.json", evaluation)
|
||||
evaluations.append(evaluation)
|
||||
summary = {
|
||||
"experiment": "explicit_rejection_gold_v0", "model": args.model,
|
||||
"successful_llm_call_count": successful_calls,
|
||||
"technical_failed_call_count": technical_failures,
|
||||
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||
"evaluations": evaluations,
|
||||
}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run isolated explicit-rejection Gold experiment")
|
||||
parser.add_argument("cases", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=1024)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
summary = run_gold(parse_args())
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,100 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Isolated controlled rejection derivation V1 experiment."""
|
||||
from __future__ import annotations
|
||||
import argparse, json, time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
from .experiment_h import DEFAULT_ENDPOINT, DEFAULT_MODEL, DerivationValidationError, OBSERVATION_KEYS, call_ollama
|
||||
from .experiment_negative_act import build_prompt as build_negative_prompt, parse_model_json, validate_classification
|
||||
|
||||
SCHEMA_VERSION="experimental-controlled-rejection-v1"
|
||||
TARGET_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||
FORBIDDEN={"rejection_form","negative_act_form","explicitly_rejected","status","decision","outcome","topic_status","closed","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
|
||||
PROMPT="""Resolve only the concrete local action or option referred to by the candidate negative act. The candidate is {candidate}. Choose only a supplied observation ID. Use the same observation for a self-contained target. If no unique local target exists, return null for both target fields. Preserve source-language meaning and material scope such as purpose and location. Preserve continuation when non-pursuit concerns continuing something. Ignore any separate positive alternative. Do not classify the negative act and do not output rejection, status, decision, outcome, responsibility, protocol concepts, confidence, relations, or graphs. Return exactly JSON with candidate_observation_id, target_observation_id, normalized_target_text and no other fields.\nObservations:\n{observations}"""
|
||||
|
||||
def _keys(v,r,loc):
|
||||
if not isinstance(v,dict): raise DerivationValidationError(f"{loc} must be an object")
|
||||
if set(v)!=r: raise DerivationValidationError(f"{loc} keys invalid: missing={sorted(r-set(v))}, unknown={sorted(set(v)-r)}")
|
||||
def _text(v,loc):
|
||||
if not isinstance(v,str) or not v.strip(): raise DerivationValidationError(f"{loc} must be non-empty")
|
||||
return v.strip()
|
||||
def _forbidden(v,loc="output"):
|
||||
if isinstance(v,dict):
|
||||
bad=FORBIDDEN & set(v)
|
||||
if bad: raise DerivationValidationError(f"{loc} contains forbidden fields: {sorted(bad)}")
|
||||
for k,x in v.items(): _forbidden(x,f"{loc}.{k}")
|
||||
elif isinstance(v,list):
|
||||
for i,x in enumerate(v): _forbidden(x,f"{loc}[{i}]")
|
||||
def validate_observations(obs):
|
||||
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
|
||||
ids=set(); evid=set()
|
||||
for i,o in enumerate(obs):
|
||||
_keys(o,OBSERVATION_KEYS,f"observations[{i}]"); oid=_text(o["observation_id"],"observation_id"); eid=_text(o["evidence_id"],"evidence_id")
|
||||
if oid in ids or eid in evid: raise DerivationValidationError("provenance must be unique")
|
||||
ids.add(oid); evid.add(eid); _text(o["content"],"content"); _text(o["speaker"],"speaker")
|
||||
return ids
|
||||
def validate_target(data,obs):
|
||||
ids=validate_observations(obs); _forbidden(data); _keys(data,TARGET_KEYS,"target output")
|
||||
candidate=_text(data["candidate_observation_id"],"candidate_observation_id")
|
||||
if candidate not in ids: raise DerivationValidationError("unknown candidate observation")
|
||||
target=data["target_observation_id"]; normalized=data["normalized_target_text"]
|
||||
if target is None:
|
||||
if normalized is not None: raise DerivationValidationError("null target requires null text")
|
||||
else:
|
||||
target=_text(target,"target_observation_id")
|
||||
if target not in ids: raise DerivationValidationError("unknown target observation")
|
||||
_text(normalized,"normalized_target_text")
|
||||
return data
|
||||
def build_target_prompt(case):
|
||||
validate_observations(case["observations"])
|
||||
return PROMPT.format(candidate=case["expected"]["candidate_observation_id"],observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||
def derive(obs,negative,target):
|
||||
ids=validate_observations(obs); validate_classification(negative,obs); validate_target(target,obs)
|
||||
candidate=negative["observation_id"]
|
||||
if target["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate outputs disagree")
|
||||
positions={o["observation_id"]:i for i,o in enumerate(obs)}; tid=target["target_observation_id"]
|
||||
gates={"negative_act_valid":True,"eligible_explicit_non_pursuit":negative["negative_act_form"]=="explicit_non_pursuit","candidate_exists":candidate in ids,"target_valid":True,"target_present":tid is not None,"target_exists":tid in ids if tid else False,"target_not_after_candidate":positions[tid]<=positions[candidate] if tid else False,"provenance_valid_unique":True,"normalized_target_nonempty":bool(target["normalized_target_text"] and target["normalized_target_text"].strip()),"same_isolated_case":tid in ids if tid else False,"no_forbidden_fields":True}
|
||||
established=all(gates.values())
|
||||
result=None
|
||||
if established:
|
||||
byid={o["observation_id"]:o for o in obs}
|
||||
result={"rejection_id":"rejection_1","content":target["normalized_target_text"].strip(),"status":"explicitly_rejected","support":{"target":{"observation_id":tid,"evidence_id":byid[tid]["evidence_id"]},"negative_act":{"observation_id":candidate,"evidence_id":byid[candidate]["evidence_id"]}}}
|
||||
return {"gates":gates,"derived_result":result}
|
||||
def load_cases(path):
|
||||
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
|
||||
return data["cases"]
|
||||
def _concepts(text,groups):
|
||||
folded=(text or "").casefold(); return all(any(x.casefold() in folded for x in g) for g in groups)
|
||||
def evaluate(case,negative,target,derivation):
|
||||
e=case["expected"]; text=target["normalized_target_text"]
|
||||
form=negative["negative_act_form"]==e["negative_act_form"]; target_ok=target["target_observation_id"]==e["target_observation_id"]
|
||||
action=_concepts(text,e["action_concepts"]); material=_concepts(text,e["material_concepts"]); forbidden=any(x.casefold() in (text or "").casefold() for x in e["forbidden_concepts"])
|
||||
final=(derivation["derived_result"] is not None)==e["explicitly_rejected"]
|
||||
label="PASS" if form and target_ok and action and material and not forbidden and final else ("PARTIAL" if form and target_ok and material and not forbidden and final else "FAIL")
|
||||
return {"case_id":case["case_id"],"classification":label,"expected_negative_act_form":e["negative_act_form"],"actual_negative_act_form":negative["negative_act_form"],"expected_target_observation_id":e["target_observation_id"],"actual_target_observation_id":target["target_observation_id"],"normalized_target_text":text,"normalized_action_correct":action,"material_scope_preserved":material,"alternative_absorbed":forbidden,"final_rejection_correct":final}
|
||||
def _write(p,v): p.write_text(json.dumps(v,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||
def run(args):
|
||||
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases})
|
||||
evals=[]; naf_calls=target_calls=technical_failures=structural_failures=0; start=time.perf_counter()
|
||||
for case in cases:
|
||||
d=args.output/case["case_id"].lower(); d.mkdir(); obs=case["observations"]; _write(d/"v3_style_input_observations.json",obs)
|
||||
try:
|
||||
source=case["negative_act_source"]
|
||||
if source=="live":
|
||||
np=build_negative_prompt(case); (d/"negative_act_prompt.txt").write_text(np,encoding="utf-8"); raw,nmeta=call_ollama(args.endpoint,args.model,np,args.timeout,args.num_ctx,args.num_predict); naf_calls+=1; (d/"negative_act_raw_response.txt").write_text(raw+"\n",encoding="utf-8"); negative=parse_model_json(raw)
|
||||
_write(d/"negative_act_ollama_metadata.json",nmeta); _write(d/"negative_act_source.json",{"kind":"live_call"})
|
||||
else:
|
||||
sd=args.negative_act_artifacts/source.lower(); accepted=json.loads((sd/"v3_style_input_observations.json").read_text());
|
||||
if accepted!=obs: raise DerivationValidationError(f"{source} observations do not exactly match")
|
||||
negative=json.loads((sd/"parsed_semantic_classification.json").read_text()); _write(d/"negative_act_source.json",{"kind":"accepted_artifact_reuse","case_id":source,"path":str(sd)})
|
||||
validate_classification(negative,obs); _write(d/"negative_act_classification.json",negative)
|
||||
tp=build_target_prompt(case); (d/"target_prompt.txt").write_text(tp,encoding="utf-8"); traw,tmeta=call_ollama(args.endpoint,args.model,tp,args.timeout,args.num_ctx,args.num_predict); target_calls+=1; (d/"target_raw_response.txt").write_text(traw+"\n",encoding="utf-8"); _write(d/"target_ollama_metadata.json",tmeta); target=parse_model_json(traw); _write(d/"target_recognition.json",target); validate_target(target,obs)
|
||||
derivation=derive(obs,negative,target); _write(d/"deterministic_gate_results.json",derivation["gates"]); _write(d/"final_derived_result.json",derivation["derived_result"]); ev=evaluate(case,negative,target,derivation)
|
||||
_write(d/"structural_validation.json",{"valid":True})
|
||||
except Exception as exc:
|
||||
structural_failures+=1; ev={"case_id":case["case_id"],"classification":"FAIL","technical_or_validation_failure":str(exc)}; _write(d/"structural_validation.json",{"valid":False,"error":str(exc)})
|
||||
_write(d/"evaluation.json",ev); evals.append(ev)
|
||||
summary={"experiment":"controlled_rejection_v1","model":args.model,"negative_act_llm_call_count":naf_calls,"reused_negative_act_count":len(cases)-naf_calls,"target_resolution_llm_call_count":target_calls,"technical_failed_call_count":technical_failures,"structural_validation_failure_count":structural_failures,"runtime_seconds":round(time.perf_counter()-start,3),"counts":{x:sum(e["classification"]==x for e in evals) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evals}; _write(args.output/"summary.json",summary); return summary
|
||||
def main():
|
||||
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--negative-act-artifacts",type=Path,default=Path("artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run")); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); args=p.parse_args(); print(json.dumps(run(args),ensure_ascii=False,indent=2)); return 0
|
||||
@@ -0,0 +1,81 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Target Normalization V0: reconstruct target text with fixed linkage."""
|
||||
from __future__ import annotations
|
||||
import argparse,json,time
|
||||
from pathlib import Path
|
||||
from typing import Any,Callable
|
||||
import requests
|
||||
from .experiment_h import DEFAULT_ENDPOINT,DEFAULT_MODEL,DerivationValidationError,OBSERVATION_KEYS
|
||||
|
||||
SCHEMA_VERSION="experimental-target-normalization-v0"
|
||||
OUTPUT_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||
LINK_KEYS={"candidate_observation_id","target_observation_id"}
|
||||
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","topic_status","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
|
||||
PROMPT="""The candidate and target observation IDs below are already resolved. Copy both IDs exactly; do not perform target selection. Reconstruct only the concrete POSITIVE action or option meaning targeted by the negative act. Remove rejection and negation polarity while preserving the underlying positive action. Preserve German source language, material qualifiers, purpose, location, named people, and continuation. Exclude separate positive alternatives. Do not summarize the discussion or infer rejection, decision, outcome, status, responsibility, ownership, protocol relevance, confidence, relations, graphs, or topic state. Return only the JSON-Schema-conforming object; null is not permitted.\n\nExample A observations: [{{"observation_id":"obs_a","content":"Mit Frau Beispiel arbeiten wir nicht weiter."}}]\nFixed IDs: candidate=obs_a, target=obs_a\nOutput: {{"candidate_observation_id":"obs_a","target_observation_id":"obs_a","normalized_target_text":"Zusammenarbeit mit Frau Beispiel fortsetzen"}}\n\nExample B observations: [{{"observation_id":"obs_a","content":"Für den Druckversuch steht die reale Anlage zur Diskussion."}},{{"observation_id":"obs_b","content":"Die reale Anlage nutzen wir dafür nicht."}}]\nFixed IDs: candidate=obs_b, target=obs_a\nOutput: {{"candidate_observation_id":"obs_b","target_observation_id":"obs_a","normalized_target_text":"reale Anlage für den Druckversuch nutzen"}}\n\nFixed candidate_observation_id: {candidate}\nFixed target_observation_id: {target}\nV3-style observations:\n{observations}"""
|
||||
|
||||
def _keys(value,required,where):
|
||||
if not isinstance(value,dict): raise DerivationValidationError(f"{where} must be an object")
|
||||
if set(value)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(value))}, unknown={sorted(set(value)-required)}")
|
||||
def _text(value,where):
|
||||
if not isinstance(value,str) or not value.strip(): raise DerivationValidationError(f"{where} must be non-empty")
|
||||
return value.strip()
|
||||
def _reject_forbidden(value,where="output"):
|
||||
if isinstance(value,dict):
|
||||
bad=FORBIDDEN & set(value)
|
||||
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
|
||||
for key,item in value.items(): _reject_forbidden(item,f"{where}.{key}")
|
||||
elif isinstance(value,list):
|
||||
for index,item in enumerate(value): _reject_forbidden(item,f"{where}[{index}]")
|
||||
def validate_observations(observations):
|
||||
if not isinstance(observations,list) or not observations: raise DerivationValidationError("observations must be non-empty")
|
||||
ids=set(); evidence=set()
|
||||
for index,item in enumerate(observations):
|
||||
_keys(item,OBSERVATION_KEYS,f"observations[{index}]"); oid=_text(item["observation_id"],"observation_id"); eid=_text(item["evidence_id"],"evidence_id")
|
||||
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
|
||||
ids.add(oid); evidence.add(eid); _text(item["content"],"content"); _text(item["speaker"],"speaker")
|
||||
return ids
|
||||
def validate_linkage(case):
|
||||
ids=validate_observations(case["observations"]); linkage=case["fixed_linkage"]; _keys(linkage,LINK_KEYS,"fixed_linkage")
|
||||
for field in LINK_KEYS:
|
||||
if _text(linkage[field],field) not in ids: raise DerivationValidationError(f"{field} is unknown")
|
||||
return linkage
|
||||
def output_schema(case):
|
||||
link=validate_linkage(case)
|
||||
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","target_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":link["candidate_observation_id"]},"target_observation_id":{"const":link["target_observation_id"]},"normalized_target_text":{"type":"string","minLength":1}}}
|
||||
def build_prompt(case):
|
||||
link=validate_linkage(case)
|
||||
return PROMPT.format(candidate=link["candidate_observation_id"],target=link["target_observation_id"],observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||
def validate_output(data,case):
|
||||
_reject_forbidden(data); _keys(data,OUTPUT_KEYS,"output"); link=validate_linkage(case)
|
||||
if data["candidate_observation_id"]!=link["candidate_observation_id"]: raise DerivationValidationError("candidate ID changed")
|
||||
if data["target_observation_id"]!=link["target_observation_id"]: raise DerivationValidationError("target ID changed")
|
||||
_text(data["normalized_target_text"],"normalized_target_text"); return data
|
||||
def build_payload(model,prompt,schema,num_ctx,num_predict): return {"model":model,"prompt":prompt,"think":False,"stream":False,"format":schema,"options":{"temperature":0,"num_ctx":num_ctx,"num_predict":num_predict}}
|
||||
def call_schema(endpoint,model,prompt,schema,timeout,num_ctx,num_predict):
|
||||
started=time.perf_counter(); response=requests.post(endpoint,json=build_payload(model,prompt,schema,num_ctx,num_predict),timeout=timeout); elapsed=time.perf_counter()-started; response.raise_for_status(); body=response.json(); raw=body.get("response")
|
||||
if not isinstance(raw,str) or not raw.strip(): raise ValueError("Ollama returned no usable response")
|
||||
return raw.strip(),{"model":body.get("model",model),"elapsed_seconds":round(elapsed,3),"total_duration_ns":body.get("total_duration"),"prompt_eval_count":body.get("prompt_eval_count"),"eval_count":body.get("eval_count"),"configuration":{"temperature":0,"think":False,"format":"json_schema_object","num_ctx":num_ctx,"num_predict":num_predict,"retries":0}}
|
||||
def _concepts(text,groups):
|
||||
folded=text.casefold(); return all(any(alias.casefold() in folded for alias in group) for group in groups)
|
||||
def evaluate(case,output):
|
||||
validate_output(output,case); expected=case["expected"]; text=output["normalized_target_text"]; folded=text.casefold(); action=_concepts(text,expected["action_concepts"]); scope=_concepts(text,expected["material_concepts"]); forbidden=[x for x in expected["forbidden_concepts"] if x.casefold() in folded]; positive=not any(x in forbidden for x in ("nicht","beenden")); german=any(x.casefold() in folded for x in expected["german_markers"]); alternative=not any(x.casefold() in folded for x in ("technikum","stattdessen")); strengthening=False
|
||||
label="PASS" if positive and action and scope and german and alternative and not forbidden and not strengthening else "FAIL"
|
||||
return {"case_id":case["case_id"],"classification":label,"expected_normalized_target_text":expected["normalized_target_text"],"actual_normalized_target_text":text,"positive_polarity_correct":positive,"action_semantics_preserved":action,"material_scope_preserved":scope,"source_language_preserved":german,"separate_alternative_excluded":alternative,"forbidden_semantics_present":forbidden,"unsupported_strengthening":strengthening,"normative_leakage":False,"candidate_id_unchanged":True,"target_id_unchanged":True}
|
||||
def load_cases(path):
|
||||
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
|
||||
for case in data["cases"]: validate_linkage(case)
|
||||
return data["cases"]
|
||||
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||
def run(args,caller:Callable=call_schema):
|
||||
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases}); evaluations=[]; calls=failures=0; started=time.perf_counter()
|
||||
for case in cases:
|
||||
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"fixed_linkage.json",case["fixed_linkage"]); schema=output_schema(case); _write(folder/"ollama_json_schema.json",schema); prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
|
||||
try:
|
||||
raw,metadata=caller(args.endpoint,args.model,prompt,schema,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",metadata); parsed=json.loads(raw); _write(folder/"parsed_response.json",parsed); validate_output(parsed,case); _write(folder/"structural_validation.json",{"valid":True}); _write(folder/"normalized_target_result.json",{"normalized_target_text":parsed["normalized_target_text"]}); evaluation=evaluate(case,parsed)
|
||||
except Exception as exc:
|
||||
failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); evaluation={"case_id":case["case_id"],"classification":"FAIL","error":str(exc),"normative_leakage":False}
|
||||
_write(folder/"evaluation.json",evaluation); evaluations.append(evaluation)
|
||||
summary={"experiment":"target_normalization_v0","model":args.model,"llm_call_count":calls,"structural_validation_failure_count":failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
|
||||
def main():
|
||||
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
|
||||
@@ -0,0 +1,92 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Isolated local target-resolution experiment; performs no rejection derivation."""
|
||||
from __future__ import annotations
|
||||
import argparse, json, time
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable
|
||||
from .experiment_h import DEFAULT_ENDPOINT, DEFAULT_MODEL, DerivationValidationError, OBSERVATION_KEYS, call_ollama
|
||||
from .experiment_negative_act import validate_classification
|
||||
|
||||
SCHEMA_VERSION="experimental-target-resolution-v0"
|
||||
TARGET_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","topic_status","closed","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
|
||||
PROMPT="""Resolve and normalize only the concrete action or option referred to by the candidate negative act. Candidate: {candidate}. Strategy: {instruction} Choose only a supplied observation ID. If no unique local target exists, use JSON null for both target fields. Preserve German source meaning, continuation, purpose, location, and other material scope. Do not translate Anlage as asset. Ignore separate positive alternatives. Do not output negative-act form, rejection, status, decision, outcome, responsibility, topic closure, protocol concepts, confidence, relations, or graphs. Return exactly this JSON object with no additional fields: {{"candidate_observation_id":"{candidate}","target_observation_id":"observation ID or null","normalized_target_text":"concise positive action in German or null"}}\nObservations:\n{observations}"""
|
||||
|
||||
def _keys(v:Any, required:set[str], where:str):
|
||||
if not isinstance(v,dict): raise DerivationValidationError(f"{where} must be an object")
|
||||
if set(v)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(v))}, unknown={sorted(set(v)-required)}")
|
||||
def _text(v:Any, where:str):
|
||||
if not isinstance(v,str) or not v.strip(): raise DerivationValidationError(f"{where} must be non-empty")
|
||||
return v.strip()
|
||||
def _reject_forbidden(v:Any, where="output"):
|
||||
if isinstance(v,dict):
|
||||
bad=FORBIDDEN & set(v)
|
||||
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
|
||||
for k,x in v.items(): _reject_forbidden(x,f"{where}.{k}")
|
||||
elif isinstance(v,list):
|
||||
for i,x in enumerate(v): _reject_forbidden(x,f"{where}[{i}]")
|
||||
def validate_observations(obs):
|
||||
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
|
||||
ids=set(); evidence=set()
|
||||
for i,o in enumerate(obs):
|
||||
_keys(o,OBSERVATION_KEYS,f"observations[{i}]"); oid=_text(o["observation_id"],"observation_id"); eid=_text(o["evidence_id"],"evidence_id")
|
||||
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
|
||||
ids.add(oid); evidence.add(eid); _text(o["content"],"content"); _text(o["speaker"],"speaker")
|
||||
return ids
|
||||
def eligibility(negative,obs):
|
||||
validate_classification(negative,obs)
|
||||
eligible=negative["negative_act_form"]=="explicit_non_pursuit"
|
||||
return {"eligible_for_target_resolution":eligible,"reason":None if eligible else "negative_act_form_not_explicit_non_pursuit"}
|
||||
def validate_target(data,obs,candidate):
|
||||
ids=validate_observations(obs); _reject_forbidden(data); _keys(data,TARGET_KEYS,"target output")
|
||||
if _text(data["candidate_observation_id"],"candidate_observation_id")!=candidate: raise DerivationValidationError("candidate observation mismatch")
|
||||
if candidate not in ids: raise DerivationValidationError("unknown candidate observation")
|
||||
target=data["target_observation_id"]; normalized=data["normalized_target_text"]
|
||||
if target is None:
|
||||
if normalized is not None: raise DerivationValidationError("null target requires null text")
|
||||
else:
|
||||
target=_text(target,"target_observation_id")
|
||||
if target not in ids: raise DerivationValidationError("unknown target observation")
|
||||
if [o["observation_id"] for o in obs].index(target)>[o["observation_id"] for o in obs].index(candidate): raise DerivationValidationError("target must not occur after candidate")
|
||||
_text(normalized,"normalized_target_text")
|
||||
return data
|
||||
def build_prompt(case):
|
||||
gate=eligibility(case["negative_act"],case["observations"])
|
||||
if not gate["eligible_for_target_resolution"]: raise DerivationValidationError("ineligible case must not build a target prompt")
|
||||
candidate=case["negative_act"]["observation_id"]
|
||||
instruction=(f"The target linkage is deterministically fixed to {candidate}; output that exact target ID and only normalize its positive action meaning." if case["strategy"]=="self_contained" else "Resolve the unique preceding local observation that supplies the referenced action.")
|
||||
return PROMPT.format(candidate=candidate,instruction=instruction,observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||
def _concepts(text,groups):
|
||||
folded=(text or "").casefold(); return all(any(alias.casefold() in folded for alias in group) for group in groups)
|
||||
def evaluate(case,gate,called,target):
|
||||
e=case["expected"]; eligible=gate["eligible_for_target_resolution"]==e["eligible"]; call_ok=called==e["eligible"]
|
||||
if not e["eligible"]:
|
||||
label="PASS" if eligible and call_ok and target is None else "FAIL"
|
||||
return {"case_id":case["case_id"],"classification":label,"negative_act_form":case["negative_act"]["negative_act_form"],"eligible":gate["eligible_for_target_resolution"],"target_resolution_call_made":called,"expected_target_observation_id":None,"actual_target_observation_id":None,"normalized_target_text":None,"material_scope_preserved":True,"alternative_isolation":True,"normative_leakage":False}
|
||||
text=target["normalized_target_text"]; target_ok=target["target_observation_id"]==e["target_observation_id"]; concepts=_concepts(text,e["concepts"]); material=_concepts(text,e["material_concepts"]); isolated=not any(x.casefold() in (text or "").casefold() for x in e["forbidden_concepts"])
|
||||
label="PASS" if eligible and call_ok and target_ok and concepts and material and isolated else ("PARTIAL" if eligible and call_ok and target_ok and material and isolated else "FAIL")
|
||||
return {"case_id":case["case_id"],"classification":label,"negative_act_form":case["negative_act"]["negative_act_form"],"eligible":gate["eligible_for_target_resolution"],"target_resolution_call_made":called,"expected_target_observation_id":e["target_observation_id"],"actual_target_observation_id":target["target_observation_id"],"normalized_target_text":text,"normalized_action_correct":concepts,"material_scope_preserved":material,"alternative_isolation":isolated,"normative_leakage":False}
|
||||
def load_cases(path):
|
||||
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("unexpected schema version")
|
||||
for case in data["cases"]: validate_observations(case["observations"]); validate_classification(case["negative_act"],case["observations"])
|
||||
return data["cases"]
|
||||
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||
def run(args, resolver:Callable=call_ollama):
|
||||
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases})
|
||||
evaluations=[]; calls=technical_failures=structural_failures=0; started=time.perf_counter()
|
||||
for case in cases:
|
||||
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"negative_act_form.json",case["negative_act"])
|
||||
gate=eligibility(case["negative_act"],case["observations"]); _write(folder/"eligibility.json",gate)
|
||||
if not gate["eligible_for_target_resolution"]:
|
||||
skipped={"call_made":False,"reason":gate["reason"]}; _write(folder/"target_resolution_skipped.json",skipped); ev=evaluate(case,gate,False,None)
|
||||
else:
|
||||
prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
|
||||
try:
|
||||
raw,meta=resolver(args.endpoint,args.model,prompt,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",meta); parsed=json.loads(raw); _write(folder/"parsed_target_resolution.json",parsed); validate_target(parsed,case["observations"],case["negative_act"]["observation_id"]); _write(folder/"structural_validation.json",{"valid":True}); ev=evaluate(case,gate,True,parsed)
|
||||
except Exception as exc:
|
||||
structural_failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); ev={"case_id":case["case_id"],"classification":"FAIL","negative_act_form":case["negative_act"]["negative_act_form"],"eligible":True,"target_resolution_call_made":True,"error":str(exc)}
|
||||
_write(folder/"evaluation.json",ev); evaluations.append(ev)
|
||||
summary={"experiment":"target_resolution_v0","model":args.model,"target_resolution_llm_call_count":calls,"technical_failed_call_count":technical_failures,"structural_validation_failure_count":structural_failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
|
||||
def main():
|
||||
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
|
||||
@@ -0,0 +1,107 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Target Resolution V1 diagnostic: linkage and normalization only."""
|
||||
from __future__ import annotations
|
||||
import argparse,json,time
|
||||
from pathlib import Path
|
||||
from typing import Any,Callable
|
||||
import requests
|
||||
from .experiment_h import DEFAULT_ENDPOINT,DEFAULT_MODEL,DerivationValidationError,OBSERVATION_KEYS
|
||||
from .experiment_negative_act import validate_classification
|
||||
|
||||
SCHEMA_VERSION="experimental-target-resolution-v1-diagnostic"
|
||||
SELF_KEYS={"candidate_observation_id","normalized_target_text"}; PAIRED_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","topic_status","closed","unresolved_issue"}
|
||||
SELF_PROMPT="""Normalize only the concrete positive action meaning in the self-contained candidate observation. The target linkage is already deterministic and is not your task. Preserve German, collaboration, named people, and continuation meaning. Do not output a target ID, rejection, status, decision, outcome, responsibility, ownership, protocol concepts, confidence, relations, graphs, or topic closure. Return only the schema-conforming object.\nCandidate observation ID: {candidate}\nObservation:\n{observations}"""
|
||||
PAIRED_PROMPT="""Resolve and normalize only the concrete local action or option referred to by the candidate negative act. Choose exactly one listed allowed target observation ID, or use JSON null only when no unique local target exists. Never return the string \"null\". Preserve German source language and all material purpose/location scope. Do not absorb a separate positive alternative. Do not output negative-act form, rejection, status, decision, outcome, responsibility, ownership, protocol concepts, confidence, relations, graphs, or topic closure.\nAllowed target observation IDs:\n{allowed}\nConcrete positive typed example:\n{{"candidate_observation_id":"obs_2","target_observation_id":"obs_1","normalized_target_text":"externe Lösung weiterverfolgen"}}\nActual JSON-null example:\n{{"candidate_observation_id":"obs_2","target_observation_id":null,"normalized_target_text":null}}\nReturn only the schema-conforming object.\nCandidate observation ID: {candidate}\nObservations:\n{observations}"""
|
||||
|
||||
def _keys(value,required,where):
|
||||
if not isinstance(value,dict): raise DerivationValidationError(f"{where} must be an object")
|
||||
if set(value)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(value))}, unknown={sorted(set(value)-required)}")
|
||||
def _text(value,where):
|
||||
if not isinstance(value,str) or not value.strip(): raise DerivationValidationError(f"{where} must be non-empty")
|
||||
return value.strip()
|
||||
def _forbidden(value,where="output"):
|
||||
if isinstance(value,dict):
|
||||
bad=FORBIDDEN & set(value)
|
||||
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
|
||||
for key,item in value.items(): _forbidden(item,f"{where}.{key}")
|
||||
elif isinstance(value,list):
|
||||
for index,item in enumerate(value): _forbidden(item,f"{where}[{index}]")
|
||||
def validate_observations(obs):
|
||||
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
|
||||
ids=[]; evidence=set()
|
||||
for index,item in enumerate(obs):
|
||||
_keys(item,OBSERVATION_KEYS,f"observations[{index}]"); oid=_text(item["observation_id"],"observation_id"); eid=_text(item["evidence_id"],"evidence_id")
|
||||
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
|
||||
ids.append(oid); evidence.add(eid); _text(item["content"],"content"); _text(item["speaker"],"speaker")
|
||||
return ids
|
||||
def allowed_ids(case): return validate_observations(case["observations"])
|
||||
def deterministic_self_link(case):
|
||||
if case["strategy"]!="self_contained": raise DerivationValidationError("self-linkage requires self-contained strategy")
|
||||
ids=validate_observations(case["observations"]); candidate=case["negative_act"]["observation_id"]
|
||||
if candidate not in ids: raise DerivationValidationError("unknown candidate")
|
||||
return {"linkage_source":"deterministic","candidate_observation_id":candidate,"target_observation_id":candidate}
|
||||
def output_schema(case):
|
||||
candidate=case["negative_act"]["observation_id"]
|
||||
if case["strategy"]=="self_contained":
|
||||
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":candidate},"normalized_target_text":{"type":"string","minLength":1}}}
|
||||
ids=allowed_ids(case)
|
||||
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","target_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":candidate},"target_observation_id":{"enum":ids+[None]},"normalized_target_text":{"type":["string","null"]}},"allOf":[{"if":{"properties":{"target_observation_id":{"type":"null"}}},"then":{"properties":{"normalized_target_text":{"type":"null"}}},"else":{"properties":{"normalized_target_text":{"type":"string","minLength":1}}}}]}
|
||||
def build_prompt(case):
|
||||
validate_classification(case["negative_act"],case["observations"]); candidate=case["negative_act"]["observation_id"]
|
||||
if case["strategy"]=="self_contained": return SELF_PROMPT.format(candidate=candidate,observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||
return PAIRED_PROMPT.format(candidate=candidate,allowed=json.dumps(allowed_ids(case),ensure_ascii=False),observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||
def validate_semantic_output(data,case):
|
||||
_forbidden(data); candidate=case["negative_act"]["observation_id"]
|
||||
if case["strategy"]=="self_contained":
|
||||
_keys(data,SELF_KEYS,"self output")
|
||||
if data["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate mismatch")
|
||||
_text(data["normalized_target_text"],"normalized_target_text")
|
||||
else:
|
||||
_keys(data,PAIRED_KEYS,"paired output")
|
||||
if data["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate mismatch")
|
||||
target=data["target_observation_id"]
|
||||
if target=="null": raise DerivationValidationError('string "null" is forbidden')
|
||||
if target is None:
|
||||
if data["normalized_target_text"] is not None: raise DerivationValidationError("null target requires null text")
|
||||
else:
|
||||
if target not in allowed_ids(case): raise DerivationValidationError("target is not an allowed ID")
|
||||
ids=allowed_ids(case)
|
||||
if ids.index(target)>ids.index(candidate): raise DerivationValidationError("target must not occur after candidate")
|
||||
_text(data["normalized_target_text"],"normalized_target_text")
|
||||
return data
|
||||
def combine(case,semantic):
|
||||
validate_semantic_output(semantic,case)
|
||||
if case["strategy"]=="self_contained":
|
||||
link=deterministic_self_link(case); return {**link,"normalized_target_text":semantic["normalized_target_text"]}
|
||||
return {"linkage_source":"llm","candidate_observation_id":semantic["candidate_observation_id"],"target_observation_id":semantic["target_observation_id"],"normalized_target_text":semantic["normalized_target_text"]}
|
||||
def build_payload(model,prompt,schema,num_ctx,num_predict):
|
||||
return {"model":model,"prompt":prompt,"think":False,"stream":False,"format":schema,"options":{"temperature":0,"num_ctx":num_ctx,"num_predict":num_predict}}
|
||||
def call_schema(endpoint,model,prompt,schema,timeout,num_ctx,num_predict):
|
||||
started=time.perf_counter(); response=requests.post(endpoint,json=build_payload(model,prompt,schema,num_ctx,num_predict),timeout=timeout); elapsed=time.perf_counter()-started; response.raise_for_status(); body=response.json(); raw=body.get("response")
|
||||
if not isinstance(raw,str) or not raw.strip(): raise ValueError("Ollama returned no usable response")
|
||||
meta={"model":body.get("model",model),"elapsed_seconds":round(elapsed,3),"total_duration_ns":body.get("total_duration"),"prompt_eval_count":body.get("prompt_eval_count"),"eval_count":body.get("eval_count"),"configuration":{"temperature":0,"think":False,"format":"json_schema_object","num_ctx":num_ctx,"num_predict":num_predict,"retries":0}}
|
||||
return raw.strip(),meta
|
||||
def _concepts(text,groups):
|
||||
folded=(text or "").casefold(); return all(any(x.casefold() in folded for x in group) for group in groups)
|
||||
def evaluate(case,semantic,combined):
|
||||
expected=case["expected"]; text=combined["normalized_target_text"]; target=combined["target_observation_id"]; concepts=_concepts(text,expected["concepts"]); material=_concepts(text,expected["material_concepts"]); isolated=not any(x.casefold() in (text or "").casefold() for x in expected["forbidden_concepts"]); recurrence=semantic.get("target_observation_id")=="null"; correct=target==expected["target_observation_id"]
|
||||
label="PASS" if correct and concepts and material and isolated and not recurrence else ("PARTIAL" if correct and material and isolated and not recurrence else "FAIL")
|
||||
return {"case_id":case["case_id"],"classification":label,"strategy":case["strategy"],"target_id_decision_source":combined["linkage_source"],"expected_target_observation_id":expected["target_observation_id"],"actual_target_observation_id":target,"normalized_target_text":text,"continuation_or_action_preserved":concepts,"material_scope_preserved":material,"alternative_isolated":isolated,"schema_valid":True,"string_null_recurrence":recurrence,"normative_leakage":False}
|
||||
def load_cases(path):
|
||||
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
|
||||
return data["cases"]
|
||||
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||
def run(args,caller:Callable=call_schema):
|
||||
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases}); evaluations=[]; calls=failures=0; started=time.perf_counter()
|
||||
for case in cases:
|
||||
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"negative_act_form.json",case["negative_act"]); _write(folder/"eligibility.json",{"eligible_for_target_resolution":True,"reason":None}); _write(folder/"deterministic_strategy.json",{"strategy":case["strategy"],"target_id_decision_source":"deterministic" if case["strategy"]=="self_contained" else "llm"}); _write(folder/"allowed_target_ids.json",allowed_ids(case)); schema=output_schema(case); _write(folder/"ollama_json_schema.json",schema); prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
|
||||
try:
|
||||
raw,meta=caller(args.endpoint,args.model,prompt,schema,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",meta); semantic=json.loads(raw); _write(folder/"parsed_semantic_output.json",semantic); validate_semantic_output(semantic,case); combined=combine(case,semantic); _write(folder/"structural_validation.json",{"valid":True}); _write(folder/"deterministic_linkage_result.json",{k:combined[k] for k in ("linkage_source","candidate_observation_id","target_observation_id")}); _write(folder/"normalized_target_result.json",{"normalized_target_text":combined["normalized_target_text"]}); evaluation=evaluate(case,semantic,combined)
|
||||
except Exception as exc:
|
||||
failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); evaluation={"case_id":case["case_id"],"classification":"FAIL","strategy":case["strategy"],"schema_valid":False,"error":str(exc)}
|
||||
_write(folder/"evaluation.json",evaluation); evaluations.append(evaluation)
|
||||
summary={"experiment":"target_resolution_v1_diagnostic","model":args.model,"llm_call_count":calls,"structural_validation_failure_count":failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
|
||||
def main():
|
||||
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
|
||||
@@ -0,0 +1,20 @@
|
||||
"""Optional speaker diarization and transcript alignment."""
|
||||
|
||||
from src.meeting_lab.diarization.alignment import align_transcript, write_diarized_transcript
|
||||
from src.meeting_lab.diarization.backend import (
|
||||
DEFAULT_MODEL,
|
||||
DiarizationError,
|
||||
DiarizationResult,
|
||||
diarize_audio,
|
||||
select_device,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"DEFAULT_MODEL",
|
||||
"DiarizationError",
|
||||
"DiarizationResult",
|
||||
"align_transcript",
|
||||
"diarize_audio",
|
||||
"select_device",
|
||||
"write_diarized_transcript",
|
||||
]
|
||||
@@ -0,0 +1,129 @@
|
||||
"""Deterministic Whisper-segment alignment to anonymous diarization turns."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
class AlignmentError(ValueError):
|
||||
"""Raised when transcript or diarization inputs are malformed."""
|
||||
|
||||
|
||||
def _number(value: Any, description: str) -> float:
|
||||
if not isinstance(value, (int, float)) or isinstance(value, bool):
|
||||
raise AlignmentError(f"{description} must be a number.")
|
||||
return float(value)
|
||||
|
||||
|
||||
def _validated_turns(turns: list[dict[str, Any]]) -> list[dict[str, Any]]:
|
||||
validated = []
|
||||
for index, turn in enumerate(turns):
|
||||
if not isinstance(turn, dict):
|
||||
raise AlignmentError(f"Diarization turn {index} must be an object.")
|
||||
start = _number(turn.get("start"), f"Diarization turn {index} start")
|
||||
end = _number(turn.get("end"), f"Diarization turn {index} end")
|
||||
speaker = turn.get("speaker_id", turn.get("speaker"))
|
||||
if end < start:
|
||||
raise AlignmentError(f"Diarization turn {index} ends before it starts.")
|
||||
if not isinstance(speaker, str) or not speaker.startswith("SPEAKER_"):
|
||||
raise AlignmentError(
|
||||
f"Diarization turn {index} must have an anonymous SPEAKER_ label."
|
||||
)
|
||||
validated.append({"start": start, "end": end, "speaker_id": speaker})
|
||||
return validated
|
||||
|
||||
|
||||
def align_transcript(
|
||||
transcript: dict[str, Any], exclusive_turns: list[dict[str, Any]]
|
||||
) -> dict[str, Any]:
|
||||
"""Return a derived transcript using maximum exclusive-turn overlap per segment."""
|
||||
if not isinstance(transcript, dict) or not isinstance(transcript.get("segments"), list):
|
||||
raise AlignmentError("Whisper transcript must contain a 'segments' list.")
|
||||
turns = _validated_turns(exclusive_turns)
|
||||
aligned_segments: list[dict[str, Any]] = []
|
||||
|
||||
for index, source in enumerate(transcript["segments"]):
|
||||
if not isinstance(source, dict):
|
||||
raise AlignmentError(f"Transcript segment {index} must be an object.")
|
||||
start = _number(source.get("start"), f"Transcript segment {index} start")
|
||||
end = _number(source.get("end"), f"Transcript segment {index} end")
|
||||
if end < start:
|
||||
raise AlignmentError(f"Transcript segment {index} ends before it starts.")
|
||||
overlap_by_speaker: dict[str, float] = {}
|
||||
for turn in turns:
|
||||
overlap = max(0.0, min(end, turn["end"]) - max(start, turn["start"]))
|
||||
if overlap:
|
||||
speaker = turn["speaker_id"]
|
||||
overlap_by_speaker[speaker] = overlap_by_speaker.get(speaker, 0.0) + overlap
|
||||
speaker_id = None
|
||||
overlap_seconds = 0.0
|
||||
if overlap_by_speaker:
|
||||
speaker_id, overlap_seconds = min(
|
||||
overlap_by_speaker.items(), key=lambda item: (-item[1], item[0])
|
||||
)
|
||||
duration = end - start
|
||||
aligned = dict(source)
|
||||
aligned.update(
|
||||
{
|
||||
"speaker_id": speaker_id,
|
||||
"speaker_overlap_seconds": round(overlap_seconds, 6),
|
||||
"speaker_overlap_ratio": round(
|
||||
overlap_seconds / duration if duration > 0 else 0.0, 6
|
||||
),
|
||||
}
|
||||
)
|
||||
aligned_segments.append(aligned)
|
||||
|
||||
return {
|
||||
"text": diarized_transcript_text(aligned_segments, include_end=True),
|
||||
"segments": aligned_segments,
|
||||
"speaker_labels_anonymous": True,
|
||||
"alignment_source": "exclusive_diarization",
|
||||
}
|
||||
|
||||
|
||||
def _timestamp(seconds: float) -> str:
|
||||
milliseconds = int(round(seconds * 1000))
|
||||
hours, remainder = divmod(milliseconds, 3_600_000)
|
||||
minutes, remainder = divmod(remainder, 60_000)
|
||||
secs, millis = divmod(remainder, 1000)
|
||||
return f"{hours:02d}:{minutes:02d}:{secs:02d}.{millis:03d}"
|
||||
|
||||
|
||||
def diarized_transcript_text(
|
||||
segments: list[dict[str, Any]], *, include_end: bool = True
|
||||
) -> str:
|
||||
lines = []
|
||||
for segment in segments:
|
||||
start = _timestamp(float(segment["start"]))
|
||||
end = _timestamp(float(segment["end"]))
|
||||
speaker = segment.get("speaker_id") or "SPEAKER_UNASSIGNED"
|
||||
timestamp = f"[{start} - {end}]" if include_end else f"[{start}]"
|
||||
lines.append(f"{timestamp} {speaker}: {str(segment.get('text', '')).strip()}")
|
||||
return "\n".join(lines) + ("\n" if lines else "")
|
||||
|
||||
|
||||
def write_diarized_transcript(
|
||||
transcript_path: Path,
|
||||
exclusive_turns_path: Path,
|
||||
output_dir: Path,
|
||||
) -> tuple[Path, Path]:
|
||||
"""Read source artifacts and write a separate speaker-aware transcript pair."""
|
||||
transcript = json.loads(Path(transcript_path).read_text(encoding="utf-8-sig"))
|
||||
turns = json.loads(Path(exclusive_turns_path).read_text(encoding="utf-8"))
|
||||
if not isinstance(turns, list):
|
||||
raise AlignmentError("Exclusive diarization turns must contain a JSON list.")
|
||||
derived = align_transcript(transcript, turns)
|
||||
output_dir = Path(output_dir)
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
json_path = output_dir / "transcript_diarized.json"
|
||||
text_path = output_dir / "transcript_diarized.txt"
|
||||
json_path.write_text(
|
||||
json.dumps(derived, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
text_path.write_text(
|
||||
diarized_transcript_text(derived["segments"], include_end=True), encoding="utf-8"
|
||||
)
|
||||
return json_path, text_path
|
||||
@@ -0,0 +1,314 @@
|
||||
"""pyannote Community-1 backend with native and isolated-container runtimes."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import importlib.metadata
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import time
|
||||
import wave
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable, Literal, Sequence
|
||||
|
||||
|
||||
DEFAULT_MODEL = "pyannote/speaker-diarization-community-1"
|
||||
PYANNOTE_VERSION = "4.0.7"
|
||||
DeviceMode = Literal["auto", "gpu", "cpu"]
|
||||
RuntimeMode = Literal["native", "container"]
|
||||
|
||||
|
||||
class DiarizationError(RuntimeError):
|
||||
"""Raised when diarization configuration or execution fails."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DiarizationResult:
|
||||
output_dir: Path
|
||||
metadata_path: Path
|
||||
ordinary_rttm: Path
|
||||
exclusive_rttm: Path
|
||||
turns_json: Path
|
||||
exclusive_turns_json: Path
|
||||
metadata: dict[str, Any]
|
||||
|
||||
|
||||
def _host_uid() -> int:
|
||||
getter = getattr(os, "getuid", None)
|
||||
if getter is None:
|
||||
raise DiarizationError("Container diarization requires host UID discovery.")
|
||||
return int(getter())
|
||||
|
||||
|
||||
def _host_gid() -> int:
|
||||
getter = getattr(os, "getgid", None)
|
||||
if getter is None:
|
||||
raise DiarizationError("Container diarization requires host GID discovery.")
|
||||
return int(getter())
|
||||
|
||||
|
||||
def select_device(mode: DeviceMode, torch_module: Any) -> tuple[Any, str | None]:
|
||||
"""Resolve CPU/GPU without depending on the GPU vendor."""
|
||||
if mode == "cpu":
|
||||
return torch_module.device("cpu"), None
|
||||
if mode not in ("auto", "gpu"):
|
||||
raise DiarizationError(f"Unsupported diarization device mode: {mode}")
|
||||
try:
|
||||
available = bool(torch_module.cuda.is_available())
|
||||
if available:
|
||||
name = str(torch_module.cuda.get_device_name(0))
|
||||
probe = torch_module.zeros(1, device="cuda")
|
||||
del probe
|
||||
return torch_module.device("cuda"), name
|
||||
except Exception as exc:
|
||||
if mode == "gpu":
|
||||
raise DiarizationError(f"Requested PyTorch GPU is not usable: {exc}") from exc
|
||||
if mode == "gpu":
|
||||
raise DiarizationError("Requested PyTorch GPU is unavailable.")
|
||||
return torch_module.device("cpu"), None
|
||||
|
||||
|
||||
def _load_pcm_wave(audio_path: Path, torch_module: Any) -> tuple[Any, int, float, dict[str, Any]]:
|
||||
try:
|
||||
with wave.open(str(audio_path), "rb") as source:
|
||||
channels = source.getnchannels()
|
||||
sample_rate = source.getframerate()
|
||||
sample_width = source.getsampwidth()
|
||||
frame_count = source.getnframes()
|
||||
pcm = bytearray(source.readframes(frame_count))
|
||||
except (OSError, wave.Error) as exc:
|
||||
raise DiarizationError(f"Cannot read PCM WAV input {audio_path}: {exc}") from exc
|
||||
if channels != 1 or sample_rate != 16000 or sample_width != 2:
|
||||
raise DiarizationError(
|
||||
"Diarization currently requires mono 16 kHz signed 16-bit PCM WAV; "
|
||||
f"got channels={channels}, sample_rate={sample_rate}, sample_width={sample_width}."
|
||||
)
|
||||
waveform = torch_module.frombuffer(pcm, dtype=torch_module.int16).to(
|
||||
torch_module.float32
|
||||
)
|
||||
waveform = (waveform / 32768.0).reshape(channels, frame_count)
|
||||
duration = frame_count / sample_rate
|
||||
validation = {
|
||||
"waveform_dtype": str(waveform.dtype),
|
||||
"waveform_shape": list(waveform.shape),
|
||||
"sample_rate": sample_rate,
|
||||
"sample_count": frame_count,
|
||||
"duration_seconds": duration,
|
||||
"min_sample_value": waveform.min().item(),
|
||||
"max_sample_value": waveform.max().item(),
|
||||
"audio_loading": "python_wave_pcm16",
|
||||
}
|
||||
return waveform, sample_rate, duration, validation
|
||||
|
||||
|
||||
def _turns(annotation: Any) -> list[dict[str, Any]]:
|
||||
return [
|
||||
{
|
||||
"start": segment.start,
|
||||
"end": segment.end,
|
||||
"speaker_id": speaker,
|
||||
}
|
||||
for segment, _track, speaker in annotation.itertracks(yield_label=True)
|
||||
]
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _result_from_output(output_dir: Path) -> DiarizationResult:
|
||||
metadata_path = output_dir / "metadata.json"
|
||||
try:
|
||||
metadata = json.loads(metadata_path.read_text(encoding="utf-8"))
|
||||
except (OSError, json.JSONDecodeError) as exc:
|
||||
raise DiarizationError(f"Cannot read diarization metadata: {exc}") from exc
|
||||
return DiarizationResult(
|
||||
output_dir=output_dir,
|
||||
metadata_path=metadata_path,
|
||||
ordinary_rttm=output_dir / "diarization.rttm",
|
||||
exclusive_rttm=output_dir / "exclusive_diarization.rttm",
|
||||
turns_json=output_dir / "turns.json",
|
||||
exclusive_turns_json=output_dir / "exclusive_turns.json",
|
||||
metadata=metadata,
|
||||
)
|
||||
|
||||
|
||||
def _require_writable_output(output_dir: Path) -> None:
|
||||
unwritable = [
|
||||
path
|
||||
for path in (output_dir, *output_dir.rglob("*"))
|
||||
if not os.access(path, os.W_OK)
|
||||
]
|
||||
if unwritable:
|
||||
rendered = ", ".join(str(path) for path in unwritable[:3])
|
||||
if len(unwritable) > 3:
|
||||
rendered += f", and {len(unwritable) - 3} more"
|
||||
raise DiarizationError(
|
||||
f"Container diarization artifacts are not writable by the host user: {rendered}"
|
||||
)
|
||||
|
||||
|
||||
def run_native_pyannote(
|
||||
audio_path: Path,
|
||||
output_dir: Path,
|
||||
device_mode: DeviceMode,
|
||||
*,
|
||||
model: str = DEFAULT_MODEL,
|
||||
) -> DiarizationResult:
|
||||
"""Run one local pyannote inference using an in-memory waveform mapping."""
|
||||
try:
|
||||
import torch
|
||||
from pyannote.audio import Pipeline
|
||||
except ImportError as exc:
|
||||
raise DiarizationError(
|
||||
f"Native diarization requires pyannote.audio=={PYANNOTE_VERSION} and PyTorch."
|
||||
) from exc
|
||||
|
||||
token = os.environ.get("HF_TOKEN")
|
||||
if not token:
|
||||
raise DiarizationError("HF_TOKEN is required for the pyannote model.")
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
waveform, sample_rate, duration, audio_metadata = _load_pcm_wave(audio_path, torch)
|
||||
device, device_name = select_device(device_mode, torch)
|
||||
|
||||
try:
|
||||
pipeline = Pipeline.from_pretrained(model, token=token)
|
||||
pipeline.to(device)
|
||||
started = time.perf_counter()
|
||||
output = pipeline(
|
||||
{
|
||||
"waveform": waveform,
|
||||
"sample_rate": sample_rate,
|
||||
"uri": audio_path.stem,
|
||||
}
|
||||
)
|
||||
runtime = time.perf_counter() - started
|
||||
except Exception as exc:
|
||||
raise DiarizationError(f"pyannote diarization failed: {type(exc).__name__}: {exc}") from exc
|
||||
|
||||
ordinary = getattr(output, "speaker_diarization", output)
|
||||
exclusive = getattr(output, "exclusive_speaker_diarization", None)
|
||||
if exclusive is None:
|
||||
raise DiarizationError("Community-1 did not return exclusive diarization.")
|
||||
ordinary_turns = _turns(ordinary)
|
||||
exclusive_turns = _turns(exclusive)
|
||||
with (output_dir / "diarization.rttm").open("w", encoding="utf-8") as handle:
|
||||
ordinary.write_rttm(handle)
|
||||
with (output_dir / "exclusive_diarization.rttm").open(
|
||||
"w", encoding="utf-8"
|
||||
) as handle:
|
||||
exclusive.write_rttm(handle)
|
||||
_write_json(output_dir / "turns.json", ordinary_turns)
|
||||
_write_json(output_dir / "exclusive_turns.json", exclusive_turns)
|
||||
speakers = sorted({turn["speaker_id"] for turn in ordinary_turns})
|
||||
actual_device = str(device)
|
||||
metadata = {
|
||||
"backend": "pyannote.audio",
|
||||
"model": model,
|
||||
"pyannote_version": importlib.metadata.version("pyannote.audio"),
|
||||
"torch_version": torch.__version__,
|
||||
"hip_version": getattr(torch.version, "hip", None),
|
||||
"cuda_version": getattr(torch.version, "cuda", None),
|
||||
"runtime_adapter": "native",
|
||||
"requested_device_mode": device_mode,
|
||||
"actual_device": actual_device,
|
||||
"device_name": device_name if actual_device == "cuda" else None,
|
||||
"audio_duration_seconds": duration,
|
||||
"runtime_seconds": runtime,
|
||||
"rtf": runtime / duration,
|
||||
"speaker_count": len(speakers),
|
||||
"speaker_labels": speakers,
|
||||
"turn_count": len(ordinary_turns),
|
||||
"exclusive_turn_count": len(exclusive_turns),
|
||||
"audio": audio_metadata,
|
||||
"credentials_persisted": False,
|
||||
"output_files": {
|
||||
"ordinary_rttm": "diarization.rttm",
|
||||
"exclusive_rttm": "exclusive_diarization.rttm",
|
||||
"turns": "turns.json",
|
||||
"exclusive_turns": "exclusive_turns.json",
|
||||
},
|
||||
}
|
||||
_write_json(output_dir / "metadata.json", metadata)
|
||||
return _result_from_output(output_dir)
|
||||
|
||||
|
||||
def run_container_pyannote(
|
||||
audio_path: Path,
|
||||
output_dir: Path,
|
||||
device_mode: DeviceMode,
|
||||
*,
|
||||
image: str,
|
||||
container_args: Sequence[str] = (),
|
||||
runner: Callable[..., subprocess.CompletedProcess[str]] = subprocess.run,
|
||||
uid_getter: Callable[[], int] = _host_uid,
|
||||
gid_getter: Callable[[], int] = _host_gid,
|
||||
) -> DiarizationResult:
|
||||
"""Run the same backend in an explicitly configured disposable container."""
|
||||
if not image.strip():
|
||||
raise DiarizationError("A diarization container image is required.")
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
host_uid = uid_getter()
|
||||
host_gid = gid_getter()
|
||||
if host_uid < 0 or host_gid < 0:
|
||||
raise DiarizationError("Host UID and GID must be non-negative integers.")
|
||||
command = [
|
||||
"docker", "run", "--rm", "--ipc=host", "--shm-size=8g", "-e", "HF_TOKEN",
|
||||
*container_args,
|
||||
"-v", f"{Path(audio_path).resolve()}:/input/audio.wav:ro",
|
||||
"-v", f"{output_dir.resolve()}:/output:rw",
|
||||
"-v", f"{Path(__file__).resolve().parents[3]}:/work/meeting-lab:ro",
|
||||
"-w", "/work/meeting-lab",
|
||||
image,
|
||||
"/bin/bash", "-lc",
|
||||
(
|
||||
"inference_status=0; "
|
||||
f"python -m pip install --disable-pip-version-check pyannote.audio=={PYANNOTE_VERSION} "
|
||||
"> /output/pip-install.log 2>&1 && "
|
||||
"python -m src.meeting_lab.diarization.container_entry "
|
||||
f"/input/audio.wav /output --device {device_mode} || inference_status=$?; "
|
||||
f"chown -R {host_uid}:{host_gid} /output || exit $?; "
|
||||
"chmod -R u+rwX /output || exit $?; "
|
||||
'exit "$inference_status"'
|
||||
),
|
||||
]
|
||||
try:
|
||||
completed = runner(command, check=False, capture_output=True, text=True)
|
||||
except OSError as exc:
|
||||
raise DiarizationError(f"Could not start diarization container: {exc}") from exc
|
||||
(output_dir / "container_stdout.log").write_text(completed.stdout, encoding="utf-8")
|
||||
(output_dir / "container_stderr.log").write_text(completed.stderr, encoding="utf-8")
|
||||
if completed.returncode != 0:
|
||||
detail = completed.stderr.strip() or completed.stdout.strip() or "no diagnostic output"
|
||||
raise DiarizationError(
|
||||
f"Diarization container failed with exit code {completed.returncode}: {detail}"
|
||||
)
|
||||
_require_writable_output(output_dir)
|
||||
result = _result_from_output(output_dir)
|
||||
metadata = dict(result.metadata)
|
||||
metadata["runtime_adapter"] = "container"
|
||||
_write_json(result.metadata_path, metadata)
|
||||
return _result_from_output(output_dir)
|
||||
|
||||
|
||||
def diarize_audio(
|
||||
audio_path: Path,
|
||||
output_dir: Path,
|
||||
device_mode: DeviceMode,
|
||||
*,
|
||||
runtime: RuntimeMode = "native",
|
||||
container_image: str | None = None,
|
||||
container_args: Sequence[str] = (),
|
||||
) -> DiarizationResult:
|
||||
if runtime == "native":
|
||||
return run_native_pyannote(audio_path, output_dir, device_mode)
|
||||
if runtime == "container":
|
||||
return run_container_pyannote(
|
||||
audio_path,
|
||||
output_dir,
|
||||
device_mode,
|
||||
image=container_image or "",
|
||||
container_args=container_args,
|
||||
)
|
||||
raise DiarizationError(f"Unsupported diarization runtime: {runtime}")
|
||||
@@ -0,0 +1,22 @@
|
||||
"""Internal entry point for the isolated pyannote container adapter."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
from pathlib import Path
|
||||
|
||||
from src.meeting_lab.diarization.backend import run_native_pyannote
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("audio", type=Path)
|
||||
parser.add_argument("output", type=Path)
|
||||
parser.add_argument("--device", choices=("auto", "gpu", "cpu"), required=True)
|
||||
args = parser.parse_args()
|
||||
run_native_pyannote(args.audio, args.output, args.device)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1 @@
|
||||
"""Isolated evidence-near observation experiment."""
|
||||
@@ -0,0 +1,395 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Extract evidence-near observations for a fixed Discussion Subject."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
SCHEMA_VERSION = "experimental-evidence-observations-v1"
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
DEFAULT_MODEL = "qwen3.5:9B"
|
||||
DEFAULT_TIMEOUT = 300
|
||||
DEFAULT_NUM_CTX = 16384
|
||||
DEFAULT_NUM_PREDICT = 4096
|
||||
|
||||
RELATIONS = {"none", "supports", "opposes", "qualifies", "limits_scope"}
|
||||
MODALITIES = {
|
||||
"factual",
|
||||
"possible",
|
||||
"suggested",
|
||||
"interpersonal_request",
|
||||
"impersonal_necessity",
|
||||
"information_question",
|
||||
"committed",
|
||||
}
|
||||
TEMPORALITIES = {"existing", "future", "completed", "unspecified"}
|
||||
EVALUATIONS = {"positive", "negative", "none"}
|
||||
AGREEMENTS = {"accepted", "rejected", "unclear", "none"}
|
||||
RESPONSIBILITIES = {"none", "named", "accepted"}
|
||||
UNCERTAINTIES = {"present", "absent"}
|
||||
CLARIFICATION_NEEDS = {"explicit", "implicit", "none"}
|
||||
OBSERVATION_ID_RE = re.compile(r"^obs_[1-9][0-9]*$")
|
||||
|
||||
|
||||
class ObservationValidationError(ValueError):
|
||||
"""Raised when an experimental fixture or model output is invalid."""
|
||||
|
||||
|
||||
PROMPT_TEMPLATE = """You extract atomic, evidence-near observations for one fixed Discussion Subject.
|
||||
|
||||
Stop before protocol interpretation. Never classify anything as an idea, proposal,
|
||||
objection, decision, action item, or open question. Do not determine protocol
|
||||
eligibility, reconstruct topics, generate a protocol, or invent missing stages.
|
||||
|
||||
Split an evidence unit into multiple observations when it directly contains multiple
|
||||
propositions. Preserve every observation's source evidence ID. Use concise content in
|
||||
the evidence language.
|
||||
|
||||
Return exactly one JSON object with this shape:
|
||||
{{
|
||||
"schema_version": "experimental-evidence-observations-v1",
|
||||
"subject_id": "copy exactly",
|
||||
"subject": "copy exactly",
|
||||
"observations": [
|
||||
{{
|
||||
"observation_id": "obs_1",
|
||||
"evidence_id": "e1",
|
||||
"content": "directly supported atomic observation",
|
||||
"target": "discussion_subject",
|
||||
"relation": "none",
|
||||
"modality": "factual",
|
||||
"temporality": "existing",
|
||||
"evaluation": "none",
|
||||
"agreement": "none",
|
||||
"responsibility": "none",
|
||||
"person": null,
|
||||
"uncertainty": "absent",
|
||||
"clarification_need": "none",
|
||||
"scope": "absent"
|
||||
}}
|
||||
]
|
||||
}}
|
||||
|
||||
Rules:
|
||||
- Number observation_id sequentially as obs_1, obs_2, ... in evidence order.
|
||||
- target is "discussion_subject", one earlier observation_id, or a non-empty list of
|
||||
earlier observation_ids only when the evidence jointly refers to them.
|
||||
- relation is only none, supports, opposes, qualifies, or limits_scope.
|
||||
- modality is only factual, possible, suggested, interpersonal_request,
|
||||
impersonal_necessity, information_question, or committed.
|
||||
- interpersonal_request is a direct request to another person.
|
||||
- impersonal_necessity says something needs to happen without assigning it.
|
||||
- information_question expresses missing information without assigning work.
|
||||
- temporality is only existing, future, completed, or unspecified.
|
||||
- evaluation is positive, negative, or none. Do not infer evaluation from world
|
||||
knowledge. A bare cost or technical fact normally has evaluation none.
|
||||
- agreement is only accepted, rejected, unclear, or none and applies to target.
|
||||
- responsibility is none, named, or accepted. Use named only for an explicitly
|
||||
addressed candidate and accepted only for explicit acceptance/commitment.
|
||||
- person is the explicit person's name for named/accepted responsibility; otherwise
|
||||
use JSON null. Mentioning or speaking in first person does not establish ownership.
|
||||
- uncertainty is present or absent.
|
||||
- clarification_need is explicit, implicit, or none.
|
||||
- scope is an evidence-grounded qualifier, or exactly "absent". Never use null or the
|
||||
string "null" anywhere.
|
||||
- Confirmation of a rejection targets the rejection observation, not the option.
|
||||
- A trial-only qualification targets and limits the accepted trial.
|
||||
- A negative consequence can oppose another observation without requiring
|
||||
clarification.
|
||||
- Personal preference is not group rejection.
|
||||
- Collective "we" does not name an individual owner.
|
||||
|
||||
Fixed Gold input:
|
||||
{input_json}
|
||||
"""
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description="Run the evidence-observation Gold experiment.")
|
||||
parser.add_argument("fixture", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT)
|
||||
parser.add_argument("--num-ctx", type=int, default=DEFAULT_NUM_CTX)
|
||||
parser.add_argument("--num-predict", type=int, default=DEFAULT_NUM_PREDICT)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required
|
||||
if missing:
|
||||
raise ObservationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise ObservationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ObservationValidationError(f"{location} must be a non-empty string")
|
||||
result = value.strip()
|
||||
if result.casefold() == "null":
|
||||
raise ObservationValidationError(f"{location} must not be the string 'null'")
|
||||
return result
|
||||
|
||||
|
||||
OBSERVATION_KEYS = {
|
||||
"observation_id", "evidence_id", "content", "target", "relation", "modality",
|
||||
"temporality", "evaluation", "agreement", "responsibility", "person",
|
||||
"uncertainty", "clarification_need", "scope",
|
||||
}
|
||||
|
||||
|
||||
def _validate_target(value: Any, location: str, earlier: set[str]) -> None:
|
||||
if isinstance(value, str):
|
||||
target = _text(value, location)
|
||||
if target != "discussion_subject" and target not in earlier:
|
||||
raise ObservationValidationError(f"{location} references unknown or later observation: {target}")
|
||||
return
|
||||
if not isinstance(value, list) or not value:
|
||||
raise ObservationValidationError(f"{location} must be discussion_subject, an earlier observation ID, or a non-empty list")
|
||||
if len(value) < 2:
|
||||
raise ObservationValidationError(f"{location} list must contain at least two jointly referenced observations")
|
||||
seen: set[str] = set()
|
||||
for index, item in enumerate(value):
|
||||
target = _text(item, f"{location}[{index}]")
|
||||
if target not in earlier:
|
||||
raise ObservationValidationError(f"{location}[{index}] references unknown or later observation: {target}")
|
||||
if target in seen:
|
||||
raise ObservationValidationError(f"{location} contains duplicate target: {target}")
|
||||
seen.add(target)
|
||||
|
||||
|
||||
def validate_observations(data: Any, case: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_case(case)
|
||||
if not isinstance(data, dict):
|
||||
raise ObservationValidationError("output must be an object")
|
||||
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "output")
|
||||
if data["schema_version"] != SCHEMA_VERSION:
|
||||
raise ObservationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
|
||||
if data["subject_id"] != case["subject_id"] or data["subject"] != case["subject"]:
|
||||
raise ObservationValidationError("model changed the fixed Discussion Subject")
|
||||
observations = data["observations"]
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise ObservationValidationError("output.observations must be a non-empty array")
|
||||
known_evidence = {item["evidence_id"] for item in case["evidence"]}
|
||||
earlier: set[str] = set()
|
||||
for index, observation in enumerate(observations, start=1):
|
||||
location = f"output.observations[{index - 1}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise ObservationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
|
||||
if not OBSERVATION_ID_RE.fullmatch(observation_id) or observation_id != f"obs_{index}":
|
||||
raise ObservationValidationError(f"{location}.observation_id must be obs_{index}")
|
||||
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if evidence_id not in known_evidence:
|
||||
raise ObservationValidationError(f"{location}.evidence_id references unknown evidence: {evidence_id}")
|
||||
_text(observation["content"], f"{location}.content")
|
||||
_validate_target(observation["target"], f"{location}.target", earlier)
|
||||
for field, values in (
|
||||
("relation", RELATIONS), ("modality", MODALITIES),
|
||||
("temporality", TEMPORALITIES), ("evaluation", EVALUATIONS),
|
||||
("agreement", AGREEMENTS), ("responsibility", RESPONSIBILITIES),
|
||||
("uncertainty", UNCERTAINTIES), ("clarification_need", CLARIFICATION_NEEDS),
|
||||
):
|
||||
if observation[field] not in values:
|
||||
raise ObservationValidationError(f"{location}.{field} is invalid: {observation[field]!r}")
|
||||
person = observation["person"]
|
||||
if observation["responsibility"] == "none":
|
||||
if person is not None:
|
||||
raise ObservationValidationError(f"{location}.person must be JSON null when responsibility is none")
|
||||
else:
|
||||
_text(person, f"{location}.person")
|
||||
scope = _text(observation["scope"], f"{location}.scope")
|
||||
if scope.casefold() == "null":
|
||||
raise ObservationValidationError(f"{location}.scope must use 'absent', not 'null'")
|
||||
earlier.add(observation_id)
|
||||
return data
|
||||
|
||||
|
||||
def validate_case(case: Any) -> dict[str, Any]:
|
||||
if not isinstance(case, dict):
|
||||
raise ObservationValidationError("case must be an object")
|
||||
_exact_keys(case, {"case_id", "description", "subject_id", "subject", "evidence", "expected_observations"}, "case")
|
||||
for field in ("case_id", "description", "subject_id", "subject"):
|
||||
_text(case[field], f"case.{field}")
|
||||
evidence = case["evidence"]
|
||||
if not isinstance(evidence, list) or not evidence:
|
||||
raise ObservationValidationError("case.evidence must be a non-empty array")
|
||||
seen: set[str] = set()
|
||||
for index, unit in enumerate(evidence):
|
||||
location = f"case.evidence[{index}]"
|
||||
if not isinstance(unit, dict):
|
||||
raise ObservationValidationError(f"{location} must be an object")
|
||||
_exact_keys(unit, {"evidence_id", "text"}, location)
|
||||
evidence_id = _text(unit["evidence_id"], f"{location}.evidence_id")
|
||||
if evidence_id in seen:
|
||||
raise ObservationValidationError(f"duplicate evidence ID: {evidence_id}")
|
||||
seen.add(evidence_id)
|
||||
_text(unit["text"], f"{location}.text")
|
||||
expected = case["expected_observations"]
|
||||
if not isinstance(expected, list) or not expected:
|
||||
raise ObservationValidationError("case.expected_observations must be a non-empty array")
|
||||
return case
|
||||
|
||||
|
||||
def validate_fixture_case(case: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_case(case)
|
||||
data = {"schema_version": SCHEMA_VERSION, "subject_id": case["subject_id"], "subject": case["subject"], "observations": case["expected_observations"]}
|
||||
validate_observations(data, case)
|
||||
return case
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
validate_fixture_case(case)
|
||||
model_input = {"subject_id": case["subject_id"], "subject": case["subject"], "evidence": case["evidence"]}
|
||||
return PROMPT_TEMPLATE.format(input_json=json.dumps(model_input, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise ObservationValidationError("model response JSON must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
|
||||
|
||||
|
||||
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
|
||||
started = time.perf_counter()
|
||||
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
|
||||
elapsed = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
raw_text = body.get("response") if isinstance(body, dict) else None
|
||||
if not isinstance(raw_text, str) or not raw_text.strip():
|
||||
raise ValueError("Ollama returned no usable response text")
|
||||
metadata = {
|
||||
"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3),
|
||||
"total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"),
|
||||
"prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"),
|
||||
"eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"),
|
||||
"configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0},
|
||||
}
|
||||
return raw_text.strip(), metadata
|
||||
|
||||
|
||||
COMPARE_FIELDS = ("evidence_id", "target", "relation", "modality", "temporality", "evaluation", "agreement", "responsibility", "person", "uncertainty", "clarification_need")
|
||||
|
||||
|
||||
def _scope_matches(actual: str, expected: str) -> bool:
|
||||
if expected == "absent":
|
||||
return actual == "absent"
|
||||
expected_terms = [term.strip().casefold() for term in expected.split("|")]
|
||||
folded = actual.casefold()
|
||||
return any(term in folded for term in expected_terms)
|
||||
|
||||
|
||||
def evaluate_observations(data: dict[str, Any], expected: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
actual = data["observations"]
|
||||
checks: list[dict[str, Any]] = []
|
||||
pair_count = min(len(actual), len(expected))
|
||||
checks.append({"name": "observation_count", "passed": len(actual) == len(expected), "critical": False})
|
||||
categories = {"missing_observations": max(0, len(expected) - len(actual)), "invented_observations": max(0, len(actual) - len(expected)), "stronger_commitment": 0, "weaker_commitment": 0, "incorrect_targets_relations": 0, "incorrect_responsibility": 0, "incorrect_uncertainty_clarification": 0}
|
||||
commitment_rank = {"factual": 0, "possible": 1, "suggested": 1, "information_question": 1, "impersonal_necessity": 2, "interpersonal_request": 2, "committed": 3}
|
||||
for index in range(pair_count):
|
||||
got, want = actual[index], expected[index]
|
||||
for field in COMPARE_FIELDS:
|
||||
passed = got[field] == want[field]
|
||||
checks.append({"name": f"obs_{index + 1}:{field}", "passed": passed, "critical": field in {"evidence_id", "target", "relation", "modality", "agreement", "responsibility", "person"}})
|
||||
if not passed:
|
||||
if field in {"target", "relation"}: categories["incorrect_targets_relations"] += 1
|
||||
if field in {"responsibility", "person"}: categories["incorrect_responsibility"] += 1
|
||||
if field in {"uncertainty", "clarification_need"}: categories["incorrect_uncertainty_clarification"] += 1
|
||||
scope_ok = _scope_matches(got["scope"], want["scope"])
|
||||
checks.append({"name": f"obs_{index + 1}:scope", "passed": scope_ok, "critical": False})
|
||||
got_rank, want_rank = commitment_rank[got["modality"]], commitment_rank[want["modality"]]
|
||||
if got_rank > want_rank or (want["agreement"] == "none" and got["agreement"] in {"accepted", "rejected"}): categories["stronger_commitment"] += 1
|
||||
if got_rank < want_rank or (want["agreement"] in {"accepted", "rejected"} and got["agreement"] == "none"): categories["weaker_commitment"] += 1
|
||||
passed_count = sum(check["passed"] for check in checks)
|
||||
critical_failures = [check["name"] for check in checks if check["critical"] and not check["passed"]]
|
||||
ratio = passed_count / len(checks)
|
||||
if ratio == 1:
|
||||
verdict = "PASS"
|
||||
elif ratio >= 0.7 and categories["stronger_commitment"] == 0 and categories["incorrect_responsibility"] == 0:
|
||||
verdict = "PARTIAL"
|
||||
else:
|
||||
verdict = "FAIL"
|
||||
return {"verdict": verdict, "matched_checks": passed_count, "check_count": len(checks), "match_ratio": round(ratio, 3), "critical_failures": critical_failures, "error_categories": categories, "checks": checks}
|
||||
|
||||
|
||||
def load_fixture(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict) or set(data) != {"cases"} or not isinstance(data["cases"], list) or not data["cases"]:
|
||||
raise ObservationValidationError("fixture must contain exactly one non-empty cases list")
|
||||
seen: set[str] = set()
|
||||
for case in data["cases"]:
|
||||
validate_fixture_case(case)
|
||||
if case["case_id"] in seen:
|
||||
raise ObservationValidationError(f"duplicate case ID: {case['case_id']}")
|
||||
seen.add(case["case_id"])
|
||||
return data["cases"]
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_case(case: dict[str, Any], output_root: Path, endpoint: str, model: str, timeout: int, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
case_dir = output_root / case["case_id"]
|
||||
case_dir.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(case_dir / "gold_input.json", {key: case[key] for key in ("case_id", "description", "subject_id", "subject", "evidence")})
|
||||
_write_json(case_dir / "gold_expected_observations.json", case["expected_observations"])
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
raw, metadata = call_ollama(endpoint, model, prompt, timeout, num_ctx, num_predict)
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_observations.json", parsed)
|
||||
validate_observations(parsed, case)
|
||||
validation = {"valid": True, "error": None}
|
||||
evaluation = evaluate_observations(parsed, case["expected_observations"])
|
||||
except requests.RequestException:
|
||||
raise
|
||||
except (json.JSONDecodeError, ObservationValidationError, ValueError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
evaluation = {"verdict": "FAIL", "matched_checks": 0, "check_count": 0, "match_ratio": 0, "critical_failures": ["schema_validation"], "error_categories": {}, "checks": []}
|
||||
_write_json(case_dir / "validation_result.json", validation)
|
||||
result = {"case_id": case["case_id"], **evaluation, "elapsed_seconds": round(time.perf_counter() - started, 3)}
|
||||
_write_json(case_dir / "evaluation_result.json", result)
|
||||
return result
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_fixture(args.fixture)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
started = time.perf_counter()
|
||||
results = []
|
||||
for index, case in enumerate(cases, start=1):
|
||||
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
|
||||
results.append(run_case(case, args.output, args.endpoint, args.model, args.timeout, args.num_ctx, args.num_predict))
|
||||
summary = {"experiment": "evidence_near_observation_extraction", "schema_version": SCHEMA_VERSION, "model": args.model, "temperature": 0, "think": False, "retries": 0, "case_count": len(cases), "llm_call_count": len(results), "runtime_seconds": round(time.perf_counter() - started, 3), "verdict_counts": {v: sum(r["verdict"] == v for r in results) for v in ("PASS", "PARTIAL", "FAIL")}, "results": results}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
summary = run_experiment(args)
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["verdict_counts"]["FAIL"] == 0 else 1
|
||||
@@ -0,0 +1 @@
|
||||
"""Reduced-semantic-load evidence observation experiment."""
|
||||
@@ -0,0 +1,344 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Extract reduced-semantic-load evidence-near observations."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
SCHEMA_VERSION = "experimental-evidence-observations-v2"
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
DEFAULT_MODEL = "qwen3.5:9B"
|
||||
MODALITIES = {"factual", "possible", "suggested", "interpersonal_request", "impersonal_necessity", "information_question", "committed"}
|
||||
TEMPORALITIES = {"existing", "future", "completed", "unspecified"}
|
||||
EVALUATIONS = {"positive", "negative", "none"}
|
||||
BINARY_SIGNALS = {"explicit", "absent"}
|
||||
PRESENCE_SIGNALS = {"present", "absent"}
|
||||
CLARIFICATION_NEEDS = {"explicit", "implicit", "none"}
|
||||
OBSERVATION_ID_RE = re.compile(r"^obs_[1-9][0-9]*$")
|
||||
|
||||
|
||||
class ObservationValidationError(ValueError):
|
||||
"""Raised for invalid fixtures or model output."""
|
||||
|
||||
|
||||
PROMPT_TEMPLATE = """You extract atomic linguistic and discourse observations for one fixed Discussion Subject.
|
||||
|
||||
Preserve only facts directly expressed by the evidence. Do not derive responsibility,
|
||||
agreement, decisions, action items, open questions, accepted trials, rejected
|
||||
alternatives, established actions, or protocol eligibility. Speaker identity, a name,
|
||||
an addressee, first-person language, collective "we", and impersonal "man" never by
|
||||
themselves establish responsibility.
|
||||
|
||||
Return exactly one JSON object with this shape:
|
||||
{{
|
||||
"schema_version": "experimental-evidence-observations-v2",
|
||||
"subject_id": "copy exactly",
|
||||
"subject": "copy exactly",
|
||||
"observations": [
|
||||
{{
|
||||
"observation_id": "obs_1",
|
||||
"evidence_id": "e1",
|
||||
"content": "directly supported atomic observation",
|
||||
"refers_to": null,
|
||||
"speaker": "name copied from evidence or null",
|
||||
"named_person": null,
|
||||
"addressee": null,
|
||||
"self_reference": false,
|
||||
"collective_we": false,
|
||||
"impersonal_person_reference": false,
|
||||
"modality": "factual",
|
||||
"temporality": "existing",
|
||||
"evaluation": "none",
|
||||
"affirmation": "absent",
|
||||
"negation": "absent",
|
||||
"determination_statement": "absent",
|
||||
"uncertainty": "absent",
|
||||
"clarification_need": "none",
|
||||
"qualifier": null,
|
||||
"limits_target": null
|
||||
}}
|
||||
]
|
||||
}}
|
||||
|
||||
Rules:
|
||||
- Produce multiple observations for distinct propositions in one evidence unit, but do
|
||||
not fragment a single proposition unnecessarily.
|
||||
- observation_id is sequential in evidence order. evidence_id must be copied exactly.
|
||||
- refers_to is null or one earlier observation_id when the utterance explicitly refers
|
||||
to it. Never use arrays. Preserve joint-reference utterances without inventing a
|
||||
multi-target graph.
|
||||
- speaker is the explicit transcript speaker. named_person is a person explicitly
|
||||
named in the proposition. addressee is a person explicitly addressed.
|
||||
- self_reference marks singular first-person self-reference. collective_we marks
|
||||
collective first-person language. impersonal_person_reference marks impersonal
|
||||
person expressions such as German "man".
|
||||
- modality is factual, possible, suggested, interpersonal_request,
|
||||
impersonal_necessity, information_question, or committed.
|
||||
- temporality is existing, future, completed, or unspecified.
|
||||
- evaluation is positive, negative, or none, only when linguistically supported.
|
||||
- affirmation is explicit only for an explicit affirmative discourse signal such as
|
||||
"ja". negation is explicit only for directly expressed negation/rejection.
|
||||
- determination_statement is present only when the utterance explicitly says a
|
||||
determination has been made.
|
||||
- uncertainty is present or absent. clarification_need is explicit, implicit, or none.
|
||||
- qualifier is null or concise evidence-grounded qualifying text.
|
||||
- limits_target is null or one earlier observation explicitly limited in validity or
|
||||
scope by this observation.
|
||||
- Use JSON null, never the string "null". Output no fields beyond the schema.
|
||||
|
||||
Fixed Gold input:
|
||||
{input_json}
|
||||
"""
|
||||
|
||||
|
||||
OBSERVATION_KEYS = {
|
||||
"observation_id", "evidence_id", "content", "refers_to", "speaker",
|
||||
"named_person", "addressee", "self_reference", "collective_we",
|
||||
"impersonal_person_reference", "modality", "temporality", "evaluation",
|
||||
"affirmation", "negation", "determination_statement", "uncertainty",
|
||||
"clarification_need", "qualifier", "limits_target",
|
||||
}
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("fixture", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=4096)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing, unknown = required - value.keys(), value.keys() - required
|
||||
if missing:
|
||||
raise ObservationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise ObservationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ObservationValidationError(f"{location} must be a non-empty string")
|
||||
result = value.strip()
|
||||
if result.casefold() == "null":
|
||||
raise ObservationValidationError(f"{location} must not be the string 'null'")
|
||||
return result
|
||||
|
||||
|
||||
def _nullable_text(value: Any, location: str) -> None:
|
||||
if value is not None:
|
||||
_text(value, location)
|
||||
|
||||
|
||||
def _prior_reference(value: Any, location: str, earlier: set[str]) -> None:
|
||||
if value is None:
|
||||
return
|
||||
reference = _text(value, location)
|
||||
if reference not in earlier:
|
||||
raise ObservationValidationError(f"{location} references unknown or later observation: {reference}")
|
||||
|
||||
|
||||
def validate_observations(data: Any, case: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_case(case)
|
||||
if not isinstance(data, dict):
|
||||
raise ObservationValidationError("output must be an object")
|
||||
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "output")
|
||||
if data["schema_version"] != SCHEMA_VERSION:
|
||||
raise ObservationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
|
||||
if data["subject_id"] != case["subject_id"] or data["subject"] != case["subject"]:
|
||||
raise ObservationValidationError("model changed the fixed Discussion Subject")
|
||||
observations = data["observations"]
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise ObservationValidationError("output.observations must be a non-empty array")
|
||||
known_evidence = {item["evidence_id"] for item in case["evidence"]}
|
||||
earlier: set[str] = set()
|
||||
for index, observation in enumerate(observations, 1):
|
||||
location = f"output.observations[{index - 1}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise ObservationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
|
||||
if not OBSERVATION_ID_RE.fullmatch(observation_id) or observation_id != f"obs_{index}":
|
||||
raise ObservationValidationError(f"{location}.observation_id must be obs_{index}")
|
||||
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if evidence_id not in known_evidence:
|
||||
raise ObservationValidationError(f"{location}.evidence_id references unknown evidence: {evidence_id}")
|
||||
_text(observation["content"], f"{location}.content")
|
||||
_prior_reference(observation["refers_to"], f"{location}.refers_to", earlier)
|
||||
_prior_reference(observation["limits_target"], f"{location}.limits_target", earlier)
|
||||
for field in ("speaker", "named_person", "addressee", "qualifier"):
|
||||
_nullable_text(observation[field], f"{location}.{field}")
|
||||
for field in ("self_reference", "collective_we", "impersonal_person_reference"):
|
||||
if not isinstance(observation[field], bool):
|
||||
raise ObservationValidationError(f"{location}.{field} must be boolean")
|
||||
for field, values in (
|
||||
("modality", MODALITIES), ("temporality", TEMPORALITIES),
|
||||
("evaluation", EVALUATIONS), ("affirmation", BINARY_SIGNALS),
|
||||
("negation", BINARY_SIGNALS), ("determination_statement", PRESENCE_SIGNALS),
|
||||
("uncertainty", PRESENCE_SIGNALS), ("clarification_need", CLARIFICATION_NEEDS),
|
||||
):
|
||||
if observation[field] not in values:
|
||||
raise ObservationValidationError(f"{location}.{field} is invalid: {observation[field]!r}")
|
||||
earlier.add(observation_id)
|
||||
return data
|
||||
|
||||
|
||||
def validate_case(case: Any) -> dict[str, Any]:
|
||||
if not isinstance(case, dict):
|
||||
raise ObservationValidationError("case must be an object")
|
||||
_exact_keys(case, {"case_id", "description", "subject_id", "subject", "evidence", "expected_observations"}, "case")
|
||||
for field in ("case_id", "description", "subject_id", "subject"):
|
||||
_text(case[field], f"case.{field}")
|
||||
if not isinstance(case["evidence"], list) or not case["evidence"]:
|
||||
raise ObservationValidationError("case.evidence must be a non-empty array")
|
||||
seen: set[str] = set()
|
||||
for index, unit in enumerate(case["evidence"]):
|
||||
_exact_keys(unit, {"evidence_id", "text"}, f"case.evidence[{index}]")
|
||||
evidence_id = _text(unit["evidence_id"], f"case.evidence[{index}].evidence_id")
|
||||
if evidence_id in seen:
|
||||
raise ObservationValidationError(f"duplicate evidence ID: {evidence_id}")
|
||||
seen.add(evidence_id)
|
||||
_text(unit["text"], f"case.evidence[{index}].text")
|
||||
if not isinstance(case["expected_observations"], list) or not case["expected_observations"]:
|
||||
raise ObservationValidationError("case.expected_observations must be a non-empty array")
|
||||
return case
|
||||
|
||||
|
||||
def validate_fixture_case(case: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_case(case)
|
||||
validate_observations({"schema_version": SCHEMA_VERSION, "subject_id": case["subject_id"], "subject": case["subject"], "observations": case["expected_observations"]}, case)
|
||||
return case
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
validate_fixture_case(case)
|
||||
model_input = {key: case[key] for key in ("subject_id", "subject", "evidence")}
|
||||
return PROMPT_TEMPLATE.format(input_json=json.dumps(model_input, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise ObservationValidationError("model response JSON must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
|
||||
|
||||
|
||||
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
|
||||
started = time.perf_counter()
|
||||
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
|
||||
elapsed = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
raw = body.get("response") if isinstance(body, dict) else None
|
||||
if not isinstance(raw, str) or not raw.strip():
|
||||
raise ValueError("Ollama returned no usable response text")
|
||||
metadata = {"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3), "total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"), "prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"), "eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"), "configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0}}
|
||||
return raw.strip(), metadata
|
||||
|
||||
|
||||
COMPARE_FIELDS = tuple(sorted(OBSERVATION_KEYS - {"observation_id", "content", "qualifier"}))
|
||||
|
||||
|
||||
def _qualifier_matches(actual: str | None, expected: str | None) -> bool:
|
||||
if expected is None:
|
||||
return actual is None
|
||||
if actual is None:
|
||||
return False
|
||||
return any(term.strip().casefold() in actual.casefold() for term in expected.split("|"))
|
||||
|
||||
|
||||
def evaluate_observations(data: dict[str, Any], expected: list[dict[str, Any]]) -> dict[str, Any]:
|
||||
actual = data["observations"]
|
||||
checks = [{"name": "observation_count", "passed": len(actual) == len(expected), "critical": False}]
|
||||
for index, (got, want) in enumerate(zip(actual, expected), 1):
|
||||
for field in COMPARE_FIELDS:
|
||||
checks.append({"name": f"obs_{index}:{field}", "passed": got[field] == want[field], "critical": field in {"evidence_id", "refers_to", "limits_target", "modality", "affirmation", "negation", "determination_statement"}})
|
||||
checks.append({"name": f"obs_{index}:qualifier", "passed": _qualifier_matches(got["qualifier"], want["qualifier"]), "critical": False})
|
||||
passed = sum(check["passed"] for check in checks)
|
||||
ratio = passed / len(checks)
|
||||
critical = [check["name"] for check in checks if check["critical"] and not check["passed"]]
|
||||
verdict = "PASS" if ratio == 1 else "PARTIAL" if ratio >= 0.75 and not critical else "FAIL"
|
||||
return {"verdict": verdict, "matched_checks": passed, "check_count": len(checks), "match_ratio": round(ratio, 3), "critical_failures": critical, "checks": checks}
|
||||
|
||||
|
||||
def load_fixture(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict) or set(data) != {"cases"} or not isinstance(data["cases"], list) or not data["cases"]:
|
||||
raise ObservationValidationError("fixture must contain exactly one non-empty cases list")
|
||||
seen: set[str] = set()
|
||||
for case in data["cases"]:
|
||||
validate_fixture_case(case)
|
||||
if case["case_id"] in seen:
|
||||
raise ObservationValidationError(f"duplicate case ID: {case['case_id']}")
|
||||
seen.add(case["case_id"])
|
||||
return data["cases"]
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_case(case: dict[str, Any], output_root: Path, endpoint: str, model: str, timeout: int, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
case_dir = output_root / case["case_id"]
|
||||
case_dir.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(case_dir / "gold_input.json", {key: case[key] for key in ("case_id", "description", "subject_id", "subject", "evidence")})
|
||||
_write_json(case_dir / "gold_expected_observations.json", case["expected_observations"])
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
started = time.perf_counter()
|
||||
raw, metadata = call_ollama(endpoint, model, prompt, timeout, num_ctx, num_predict)
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
try:
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_observations.json", parsed)
|
||||
validate_observations(parsed, case)
|
||||
validation = {"valid": True, "error": None}
|
||||
evaluation = evaluate_observations(parsed, case["expected_observations"])
|
||||
except (json.JSONDecodeError, ObservationValidationError, ValueError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
evaluation = {"verdict": "FAIL", "matched_checks": 0, "check_count": 0, "match_ratio": 0, "critical_failures": ["schema_validation"], "checks": []}
|
||||
_write_json(case_dir / "validation_result.json", validation)
|
||||
result = {"case_id": case["case_id"], **evaluation, "elapsed_seconds": round(time.perf_counter() - started, 3)}
|
||||
_write_json(case_dir / "evaluation_result.json", result)
|
||||
return result
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_fixture(args.fixture)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
started = time.perf_counter()
|
||||
results = []
|
||||
for index, case in enumerate(cases, 1):
|
||||
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
|
||||
results.append(run_case(case, args.output, args.endpoint, args.model, args.timeout, args.num_ctx, args.num_predict))
|
||||
summary = {"experiment": "evidence_near_observation_extraction_v2", "schema_version": SCHEMA_VERSION, "model": args.model, "temperature": 0, "think": False, "retries": 0, "case_count": len(cases), "llm_call_count": len(results), "runtime_seconds": round(time.perf_counter() - started, 3), "verdict_counts": {verdict: sum(result["verdict"] == verdict for result in results) for verdict in ("PASS", "PARTIAL", "FAIL")}, "results": results}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
summary = run_experiment(args)
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["verdict_counts"]["FAIL"] == 0 else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1 @@
|
||||
"""Minimal semantic-preservation observation experiment."""
|
||||
@@ -0,0 +1,271 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Preserve meeting meaning as minimal atomic natural-language observations."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
SCHEMA_VERSION = "experimental-evidence-observations-v3"
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
DEFAULT_MODEL = "qwen3.5:9B"
|
||||
OBSERVATION_ID_RE = re.compile(r"^obs_[1-9][0-9]*$")
|
||||
OBSERVATION_KEYS = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
|
||||
|
||||
|
||||
class ObservationValidationError(ValueError):
|
||||
"""Raised for invalid fixtures or model output."""
|
||||
|
||||
|
||||
PROMPT_TEMPLATE = """Preserve the meeting meaning in atomic natural-language observations.
|
||||
|
||||
This is semantic preservation, not classification or summarization. Return only facts
|
||||
faithfully contributed by the evidence. Conservative wording is more important than
|
||||
elegant prose. When in doubt, preserve the source wording closely.
|
||||
|
||||
Return exactly one JSON object:
|
||||
{{
|
||||
"schema_version": "experimental-evidence-observations-v3",
|
||||
"subject_id": "copy exactly",
|
||||
"subject": "copy exactly",
|
||||
"observations": [
|
||||
{{
|
||||
"observation_id": "obs_1",
|
||||
"evidence_id": "e1",
|
||||
"content": "atomic, semantically faithful observation",
|
||||
"speaker": "speaker copied from evidence",
|
||||
"named_person": null,
|
||||
"addressee": null
|
||||
}}
|
||||
]
|
||||
}}
|
||||
|
||||
Rules:
|
||||
- Use only the six observation fields shown. Do not output classifications, labels,
|
||||
relations, scope fields, responsibility, agreement, decisions, actions, questions,
|
||||
eligibility, or any other field.
|
||||
- observation_id is sequential in evidence order. Copy evidence_id and speaker.
|
||||
- named_person is null or a person explicitly named in that observation's evidence.
|
||||
- addressee is null or a person explicitly addressed in that observation's evidence.
|
||||
- A name, speaker, or addressee never implies responsibility, acceptance, ownership,
|
||||
or assignment.
|
||||
- content is not a summary. Preserve distinctions needed for later interpretation:
|
||||
maybe/perhaps; can/could; should/must; personal, collective, or impersonal wording;
|
||||
explicit requests, acceptances, and rejections; uncertainty and unresolved status;
|
||||
conditions such as "if at all"; quantities; deadlines; trial/process/comparison
|
||||
boundaries; "not yet"; and sequence such as "then".
|
||||
- Never strengthen modality, weaken uncertainty, turn possibility into fact, turn a
|
||||
preference into group rejection, turn a request into established work, turn "we"
|
||||
into individual ownership, remove conditions/limits, generalize, or invent relations.
|
||||
- Split one evidence unit only when it contributes propositions that may later require
|
||||
different interpretations. Do not split merely because it has several clauses.
|
||||
- Do not emit observation-ID relations. When evidence clearly makes an observation
|
||||
depend on the immediately preceding proposition, state that dependency naturally in
|
||||
content, without inventing an antecedent.
|
||||
- Preserve content in the evidence language. Use JSON null, never the string "null".
|
||||
|
||||
Fixed input:
|
||||
{input_json}
|
||||
"""
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("fixture", type=Path)
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=300)
|
||||
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||
parser.add_argument("--num-predict", type=int, default=4096)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||
missing, unknown = required - value.keys(), value.keys() - required
|
||||
if missing:
|
||||
raise ObservationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||
if unknown:
|
||||
raise ObservationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||
|
||||
|
||||
def _text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ObservationValidationError(f"{location} must be a non-empty string")
|
||||
result = value.strip()
|
||||
if result.casefold() == "null":
|
||||
raise ObservationValidationError(f"{location} must not be the string 'null'")
|
||||
return result
|
||||
|
||||
|
||||
def _explicit_people(text: str) -> set[str]:
|
||||
prefix = text.split(":", 1)[0].strip() if ":" in text else ""
|
||||
candidates = set(re.findall(r"\b(?:Dr\.\s+)?[A-ZÄÖÜ][A-Za-zÄÖÜäöüß-]+(?:\s+[A-ZÄÖÜ][A-Za-zÄÖÜäöüß-]+)*", text))
|
||||
candidates.discard(prefix)
|
||||
return candidates
|
||||
|
||||
|
||||
def validate_observations(data: Any, case: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_case(case)
|
||||
if not isinstance(data, dict):
|
||||
raise ObservationValidationError("output must be an object")
|
||||
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "output")
|
||||
if data["schema_version"] != SCHEMA_VERSION:
|
||||
raise ObservationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
|
||||
if data["subject_id"] != case["subject_id"] or data["subject"] != case["subject"]:
|
||||
raise ObservationValidationError("model changed the fixed Discussion Subject")
|
||||
observations = data["observations"]
|
||||
if not isinstance(observations, list) or not observations:
|
||||
raise ObservationValidationError("output.observations must be a non-empty list")
|
||||
evidence = {item["evidence_id"]: item["text"] for item in case["evidence"]}
|
||||
seen: set[str] = set()
|
||||
for index, observation in enumerate(observations):
|
||||
location = f"output.observations[{index}]"
|
||||
if not isinstance(observation, dict):
|
||||
raise ObservationValidationError(f"{location} must be an object")
|
||||
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
|
||||
if not OBSERVATION_ID_RE.fullmatch(observation_id) or observation_id in seen:
|
||||
raise ObservationValidationError(f"{location}.observation_id must be unique and match obs_N")
|
||||
seen.add(observation_id)
|
||||
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
|
||||
if evidence_id not in evidence:
|
||||
raise ObservationValidationError(f"{location}.evidence_id references unknown evidence: {evidence_id}")
|
||||
source = evidence[evidence_id]
|
||||
source_speaker = source.split(":", 1)[0].strip()
|
||||
speaker = _text(observation["speaker"], f"{location}.speaker")
|
||||
if speaker != source_speaker:
|
||||
raise ObservationValidationError(f"{location}.speaker must match evidence speaker {source_speaker!r}")
|
||||
_text(observation["content"], f"{location}.content")
|
||||
explicit_people = _explicit_people(source)
|
||||
for field in ("named_person", "addressee"):
|
||||
person = observation[field]
|
||||
if person is not None:
|
||||
person = _text(person, f"{location}.{field}")
|
||||
if person not in explicit_people:
|
||||
raise ObservationValidationError(f"{location}.{field} is not an explicit person in evidence: {person!r}")
|
||||
return data
|
||||
|
||||
|
||||
def validate_case(case: Any) -> dict[str, Any]:
|
||||
required = {"case_id", "description", "subject_id", "subject", "evidence", "semantic_requirements"}
|
||||
if not isinstance(case, dict):
|
||||
raise ObservationValidationError("case must be an object")
|
||||
_exact_keys(case, required, "case")
|
||||
for field in ("case_id", "description", "subject_id", "subject"):
|
||||
_text(case[field], f"case.{field}")
|
||||
if not isinstance(case["evidence"], list) or not case["evidence"]:
|
||||
raise ObservationValidationError("case.evidence must be a non-empty list")
|
||||
evidence_ids: set[str] = set()
|
||||
for index, unit in enumerate(case["evidence"]):
|
||||
_exact_keys(unit, {"evidence_id", "text"}, f"case.evidence[{index}]")
|
||||
evidence_id = _text(unit["evidence_id"], f"case.evidence[{index}].evidence_id")
|
||||
if evidence_id in evidence_ids:
|
||||
raise ObservationValidationError(f"duplicate evidence ID: {evidence_id}")
|
||||
evidence_ids.add(evidence_id)
|
||||
_text(unit["text"], f"case.evidence[{index}].text")
|
||||
if not isinstance(case["semantic_requirements"], list) or not case["semantic_requirements"]:
|
||||
raise ObservationValidationError("case.semantic_requirements must be a non-empty list")
|
||||
for index, requirement in enumerate(case["semantic_requirements"]):
|
||||
_text(requirement, f"case.semantic_requirements[{index}]")
|
||||
return case
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
validate_case(case)
|
||||
model_input = {key: case[key] for key in ("subject_id", "subject", "evidence")}
|
||||
return PROMPT_TEMPLATE.format(input_json=json.dumps(model_input, ensure_ascii=False, indent=2))
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise ObservationValidationError("model response JSON must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
|
||||
|
||||
|
||||
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
|
||||
started = time.perf_counter()
|
||||
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
|
||||
elapsed = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
raw = body.get("response") if isinstance(body, dict) else None
|
||||
if not isinstance(raw, str) or not raw.strip():
|
||||
raise ValueError("Ollama returned no usable response text")
|
||||
metadata = {"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3), "total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"), "prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"), "eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"), "configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0}}
|
||||
return raw.strip(), metadata
|
||||
|
||||
|
||||
def load_fixture(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict) or set(data) != {"cases"} or not isinstance(data["cases"], list) or not data["cases"]:
|
||||
raise ObservationValidationError("fixture must contain exactly one non-empty cases list")
|
||||
seen: set[str] = set()
|
||||
for case in data["cases"]:
|
||||
validate_case(case)
|
||||
if case["case_id"] in seen:
|
||||
raise ObservationValidationError(f"duplicate case ID: {case['case_id']}")
|
||||
seen.add(case["case_id"])
|
||||
return data["cases"]
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def run_case(case: dict[str, Any], output_root: Path, endpoint: str, model: str, timeout: int, num_ctx: int, num_predict: int) -> dict[str, Any]:
|
||||
case_dir = output_root / case["case_id"]
|
||||
case_dir.mkdir(parents=True, exist_ok=False)
|
||||
_write_json(case_dir / "source_evidence.json", {key: case[key] for key in ("case_id", "description", "subject_id", "subject", "evidence")})
|
||||
_write_json(case_dir / "gold_semantic_requirements.json", case["semantic_requirements"])
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
started = time.perf_counter()
|
||||
raw, metadata = call_ollama(endpoint, model, prompt, timeout, num_ctx, num_predict)
|
||||
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||
try:
|
||||
parsed = parse_model_json(raw)
|
||||
_write_json(case_dir / "parsed_observations.json", parsed)
|
||||
validate_observations(parsed, case)
|
||||
validation = {"valid": True, "error": None}
|
||||
except (json.JSONDecodeError, ObservationValidationError, ValueError) as exc:
|
||||
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||
_write_json(case_dir / "structural_validation.json", validation)
|
||||
return {"case_id": case["case_id"], "structurally_valid": validation["valid"], "elapsed_seconds": round(time.perf_counter() - started, 3)}
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_fixture(args.fixture)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
started = time.perf_counter()
|
||||
results = []
|
||||
for index, case in enumerate(cases, 1):
|
||||
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
|
||||
results.append(run_case(case, args.output, args.endpoint, args.model, args.timeout, args.num_ctx, args.num_predict))
|
||||
summary = {"experiment": "evidence_near_observation_extraction_v3", "schema_version": SCHEMA_VERSION, "model": args.model, "temperature": 0, "think": False, "retries": 0, "case_count": len(cases), "llm_call_count": len(results), "runtime_seconds": round(time.perf_counter() - started, 3), "structurally_valid_count": sum(result["structurally_valid"] for result in results), "results": results}
|
||||
_write_json(args.output / "summary.json", summary)
|
||||
return summary
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
summary = run_experiment(args)
|
||||
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||
return 0 if summary["structurally_valid_count"] == summary["case_count"] else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,95 @@
|
||||
"""Minimal Ollama client behavior used by the direct protocol MVP."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434"
|
||||
|
||||
|
||||
class OllamaError(RuntimeError):
|
||||
"""Raised when Ollama cannot safely complete the requested operation."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class OllamaGeneration:
|
||||
raw_response: dict[str, Any]
|
||||
text: str
|
||||
client_wall_time_seconds: float
|
||||
|
||||
|
||||
def ollama_base_url(endpoint: str) -> str:
|
||||
endpoint = endpoint.rstrip("/")
|
||||
return endpoint.rsplit("/api/", 1)[0] if "/api/" in endpoint else endpoint
|
||||
|
||||
|
||||
def generate_url(endpoint: str) -> str:
|
||||
return f"{ollama_base_url(endpoint)}/api/generate"
|
||||
|
||||
|
||||
def require_model(endpoint: str, model: str, timeout: int = 10) -> dict[str, Any]:
|
||||
base_url = ollama_base_url(endpoint)
|
||||
try:
|
||||
response = requests.get(f"{base_url}/api/tags", timeout=timeout)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
except (requests.RequestException, ValueError) as exc:
|
||||
raise OllamaError(f"Ollama endpoint is not reachable at {base_url}: {exc}") from exc
|
||||
|
||||
models = data.get("models") if isinstance(data, dict) else None
|
||||
if not isinstance(models, list):
|
||||
raise OllamaError("Ollama /api/tags returned a malformed response.")
|
||||
installed = {
|
||||
item.get("name")
|
||||
for item in models
|
||||
if isinstance(item, dict) and isinstance(item.get("name"), str)
|
||||
}
|
||||
if model not in installed:
|
||||
raise OllamaError(f"Requested model is not installed in Ollama: {model}")
|
||||
return {"base_url": base_url, "model": model, "installed": True}
|
||||
|
||||
|
||||
def generate_once(
|
||||
endpoint: str,
|
||||
model: str,
|
||||
prompt: str,
|
||||
*,
|
||||
timeout: int,
|
||||
num_ctx: int,
|
||||
num_predict: int,
|
||||
) -> OllamaGeneration:
|
||||
payload = {
|
||||
"model": model,
|
||||
"prompt": prompt,
|
||||
"think": False,
|
||||
"stream": False,
|
||||
"options": {
|
||||
"temperature": 0.0,
|
||||
"num_ctx": num_ctx,
|
||||
"num_predict": num_predict,
|
||||
},
|
||||
}
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
response = requests.post(generate_url(endpoint), json=payload, timeout=timeout)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
except requests.RequestException as exc:
|
||||
raise OllamaError(f"Ollama generation request failed: {exc}") from exc
|
||||
except ValueError as exc:
|
||||
raise OllamaError("Ollama generation response is not valid JSON.") from exc
|
||||
wall_time = time.perf_counter() - started
|
||||
|
||||
if not isinstance(data, dict):
|
||||
raise OllamaError("Ollama generation response must be a JSON object.")
|
||||
text = data.get("response")
|
||||
if not isinstance(text, str):
|
||||
raise OllamaError("Ollama generation response has no string 'response' field.")
|
||||
if not text.strip():
|
||||
raise OllamaError("Ollama returned an empty protocol.")
|
||||
return OllamaGeneration(data, text, wall_time)
|
||||
|
||||
@@ -3,13 +3,15 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import ast
|
||||
import copy
|
||||
import re
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
SUPPORTED_SCHEMA_VERSIONS = {"1"}
|
||||
VALID_ATTENDANCE_STATUSES = {"present", "not_present", "absent"}
|
||||
VALID_ATTENDANCE_STATUSES = {"present", "mentioned_only"}
|
||||
|
||||
|
||||
class MeetingContextValidationError(ValueError):
|
||||
@@ -29,6 +31,21 @@ class MeetingContext:
|
||||
def meeting_id(self) -> str:
|
||||
return str(self.data["meeting"]["meeting_id"])
|
||||
|
||||
@property
|
||||
def speaker_mappings(self) -> dict[str, str]:
|
||||
mappings = self.data.get("speaker_mappings")
|
||||
return dict(mappings) if isinstance(mappings, dict) else {}
|
||||
|
||||
def participant_for_speaker(self, speaker_label: str) -> dict[str, Any] | None:
|
||||
"""Resolve only an explicit authoritative mapping; never infer identity."""
|
||||
participant_id = self.speaker_mappings.get(speaker_label)
|
||||
if participant_id is None:
|
||||
return None
|
||||
for participant in self.data.get("participants", []):
|
||||
if participant.get("participant_id") == participant_id:
|
||||
return participant
|
||||
return None
|
||||
|
||||
def provenance(self) -> dict[str, str]:
|
||||
return {
|
||||
"meeting_id": self.meeting_id,
|
||||
@@ -42,8 +59,43 @@ def load_meeting_context(path: Path) -> MeetingContext:
|
||||
if not isinstance(loaded, dict):
|
||||
raise MeetingContextValidationError("Meeting Context must be a YAML object.")
|
||||
|
||||
validate_meeting_context(loaded)
|
||||
return MeetingContext(data=loaded, source_file=path)
|
||||
normalized = _with_attendance_defaults(loaded)
|
||||
validate_meeting_context(normalized)
|
||||
return MeetingContext(data=normalized, source_file=path)
|
||||
|
||||
|
||||
def create_meeting_context(
|
||||
data: dict[str, Any], *, source_file: Path = Path("<generated>")
|
||||
) -> MeetingContext:
|
||||
"""Validate structured data and return an immutable context boundary."""
|
||||
validated = _with_attendance_defaults(data)
|
||||
validate_meeting_context(validated)
|
||||
return MeetingContext(data=validated, source_file=source_file)
|
||||
|
||||
|
||||
def serialize_meeting_context_yaml(context: MeetingContext) -> str:
|
||||
"""Serialize validated Meeting Context data deterministically as YAML."""
|
||||
validate_meeting_context(context.data)
|
||||
try:
|
||||
import yaml # type: ignore[import-not-found]
|
||||
except ModuleNotFoundError as exc:
|
||||
raise MeetingContextValidationError(
|
||||
"PyYAML is required to write Meeting Context YAML."
|
||||
) from exc
|
||||
return yaml.safe_dump(
|
||||
context.data,
|
||||
allow_unicode=True,
|
||||
sort_keys=False,
|
||||
default_flow_style=False,
|
||||
)
|
||||
|
||||
|
||||
def write_meeting_context(context: MeetingContext, path: Path) -> Path:
|
||||
"""Persist a validated context without changing its schema or semantics."""
|
||||
path = Path(path)
|
||||
path.parent.mkdir(parents=True, exist_ok=True)
|
||||
path.write_text(serialize_meeting_context_yaml(context), encoding="utf-8")
|
||||
return path
|
||||
|
||||
|
||||
def validate_meeting_context(data: dict[str, Any]) -> None:
|
||||
@@ -74,10 +126,26 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
|
||||
+ ", ".join(collisions)
|
||||
)
|
||||
|
||||
speaker_mappings = _optional_mapping(
|
||||
data.get("speaker_mappings"), "speaker_mappings"
|
||||
)
|
||||
for speaker_label, participant_id in speaker_mappings.items():
|
||||
if not isinstance(speaker_label, str) or re.fullmatch(
|
||||
r"SPEAKER_\d+", speaker_label
|
||||
) is None:
|
||||
raise MeetingContextValidationError(
|
||||
f"Invalid diarization speaker label: {speaker_label!r}."
|
||||
)
|
||||
if not isinstance(participant_id, str) or participant_id not in participant_ids:
|
||||
raise MeetingContextValidationError(
|
||||
f"speaker_mappings.{speaker_label} references unknown participant: "
|
||||
f"{participant_id!r}."
|
||||
)
|
||||
|
||||
for index, participant in enumerate(participants):
|
||||
item_path = f"participants[{index}]"
|
||||
_validate_attendance(participant, item_path)
|
||||
if participant.get("attendance_status") != "present":
|
||||
status = _validate_attendance(participant, item_path, default="present")
|
||||
if status != "present":
|
||||
raise MeetingContextValidationError(
|
||||
f"{item_path}.attendance_status must be 'present'."
|
||||
)
|
||||
@@ -85,10 +153,10 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
|
||||
|
||||
for index, person in enumerate(mentioned_people):
|
||||
item_path = f"mentioned_people[{index}]"
|
||||
_validate_attendance(person, item_path)
|
||||
if person.get("attendance_status") == "present":
|
||||
status = _validate_attendance(person, item_path, default="mentioned_only")
|
||||
if status != "mentioned_only":
|
||||
raise MeetingContextValidationError(
|
||||
f"{item_path}.attendance_status must not be 'present'."
|
||||
f"{item_path}.attendance_status must be 'mentioned_only'."
|
||||
)
|
||||
_validate_department_reference(person, item_path, department_ids)
|
||||
|
||||
@@ -131,6 +199,28 @@ def render_meeting_context_for_prompt(context: MeetingContext) -> str:
|
||||
for participant in participants:
|
||||
lines.append(_render_person_line(participant, "participant_id", departments_by_id))
|
||||
|
||||
speaker_mappings = context.speaker_mappings
|
||||
if speaker_mappings:
|
||||
participants_by_id = {
|
||||
participant["participant_id"]: participant
|
||||
for participant in participants
|
||||
if isinstance(participant, dict) and participant.get("participant_id")
|
||||
}
|
||||
lines.extend(
|
||||
[
|
||||
"",
|
||||
"Confirmed diarization speaker mappings (authoritative):",
|
||||
"- Use only these explicit mappings. Never infer identities for other speaker labels.",
|
||||
"- Unmapped SPEAKER_XX labels must remain anonymous.",
|
||||
]
|
||||
)
|
||||
for speaker_label, participant_id in sorted(speaker_mappings.items()):
|
||||
participant = participants_by_id[participant_id]
|
||||
lines.append(
|
||||
f"- {speaker_label}: {_text(participant.get('display_name'))} "
|
||||
f"(participant_id: {participant_id})"
|
||||
)
|
||||
|
||||
mentioned_people = _optional_list(data.get("mentioned_people"), "mentioned_people")
|
||||
if mentioned_people:
|
||||
lines.extend(["", "Mentioned but absent people:"])
|
||||
@@ -245,12 +335,28 @@ def _collect_unique_ids(items: list[Any], key: str, path: str) -> set[str]:
|
||||
return ids
|
||||
|
||||
|
||||
def _validate_attendance(item: dict[str, Any], path: str) -> None:
|
||||
status = item.get("attendance_status")
|
||||
def _validate_attendance(item: dict[str, Any], path: str, *, default: str) -> str:
|
||||
status = item.get("attendance_status", default)
|
||||
if status not in VALID_ATTENDANCE_STATUSES:
|
||||
raise MeetingContextValidationError(
|
||||
f"{path}.attendance_status has invalid value: {status!r}."
|
||||
)
|
||||
return status
|
||||
|
||||
|
||||
def _with_attendance_defaults(data: dict[str, Any]) -> dict[str, Any]:
|
||||
normalized = copy.deepcopy(data)
|
||||
participants = normalized.get("participants")
|
||||
if isinstance(participants, list):
|
||||
for participant in participants:
|
||||
if isinstance(participant, dict):
|
||||
participant.setdefault("attendance_status", "present")
|
||||
mentioned_people = normalized.get("mentioned_people")
|
||||
if isinstance(mentioned_people, list):
|
||||
for person in mentioned_people:
|
||||
if isinstance(person, dict):
|
||||
person.setdefault("attendance_status", "mentioned_only")
|
||||
return normalized
|
||||
|
||||
|
||||
def _validate_department_reference(
|
||||
|
||||
@@ -0,0 +1,17 @@
|
||||
"""Reusable orchestration APIs for Meeting Lab applications and CLIs."""
|
||||
|
||||
from src.meeting_lab.orchestration.mvp import (
|
||||
DEFAULT_OUTPUT_ROOT,
|
||||
MvpMeetingConfig,
|
||||
MvpRunResult,
|
||||
create_unique_run_dir,
|
||||
run_mvp_meeting,
|
||||
)
|
||||
|
||||
__all__ = [
|
||||
"DEFAULT_OUTPUT_ROOT",
|
||||
"MvpMeetingConfig",
|
||||
"MvpRunResult",
|
||||
"create_unique_run_dir",
|
||||
"run_mvp_meeting",
|
||||
]
|
||||
@@ -0,0 +1,467 @@
|
||||
"""Reusable audio-to-direct-protocol MVP orchestration."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import re
|
||||
import shutil
|
||||
import sys
|
||||
import time
|
||||
from collections.abc import Callable, Mapping, Sequence
|
||||
from dataclasses import dataclass
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
from src.meeting_lab.audio import prepare_audio
|
||||
from src.meeting_lab.diarization import (
|
||||
DEFAULT_MODEL as DEFAULT_DIARIZATION_MODEL,
|
||||
diarize_audio,
|
||||
write_diarized_transcript,
|
||||
)
|
||||
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT
|
||||
from src.meeting_lab.models.meeting_context import (
|
||||
MeetingContext,
|
||||
create_meeting_context,
|
||||
load_meeting_context,
|
||||
validate_meeting_context,
|
||||
write_meeting_context,
|
||||
)
|
||||
from src.meeting_lab.progress import ProgressEvent, ProgressSink, ProgressStatus
|
||||
from src.meeting_lab.protocol.generate_direct_protocol import (
|
||||
DEFAULT_MODEL,
|
||||
DEFAULT_NUM_CTX,
|
||||
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
DirectProtocolResult,
|
||||
generate_direct_protocol,
|
||||
load_compact_transcript,
|
||||
)
|
||||
from src.meeting_lab.transcription.whisper import transcribe_audio
|
||||
|
||||
|
||||
DEFAULT_OUTPUT_ROOT = Path("meeting_data/runs")
|
||||
ContextInput = MeetingContext | Mapping[str, Any]
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MvpMeetingConfig:
|
||||
audio_file: Path
|
||||
whisper_model: Path
|
||||
whisper_executable: str = "whisper-cli"
|
||||
ffmpeg_executable: str = "ffmpeg"
|
||||
audio_normalization: bool = True
|
||||
context_file: Path | None = None
|
||||
output_root: Path = DEFAULT_OUTPUT_ROOT
|
||||
language: str = "de"
|
||||
threads: str | int = "auto"
|
||||
model: str = DEFAULT_MODEL
|
||||
ollama_endpoint: str = DEFAULT_ENDPOINT
|
||||
protocol_num_ctx: int = DEFAULT_NUM_CTX
|
||||
protocol_safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET
|
||||
diarization: str = "off"
|
||||
diarization_runtime: str = "native"
|
||||
diarization_container_image: str | None = None
|
||||
diarization_container_args: Sequence[str] = ()
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class MvpRunResult:
|
||||
exit_code: int
|
||||
run_dir: Path | None
|
||||
protocol_path: Path | None
|
||||
|
||||
|
||||
def regenerate_mvp_protocol(
|
||||
run_dir: Path,
|
||||
*,
|
||||
meeting_context: ContextInput,
|
||||
model: str = DEFAULT_MODEL,
|
||||
ollama_endpoint: str = DEFAULT_ENDPOINT,
|
||||
protocol_num_ctx: int = DEFAULT_NUM_CTX,
|
||||
protocol_safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
progress_sink: ProgressSink | None = None,
|
||||
) -> MvpRunResult:
|
||||
"""Regenerate only protocol artifacts from an existing completed run."""
|
||||
started = time.perf_counter()
|
||||
run_dir = Path(run_dir)
|
||||
context = _effective_context(meeting_context)
|
||||
if context is None:
|
||||
raise ValueError("Meeting Context is required for protocol regeneration.")
|
||||
if protocol_num_ctx <= 0:
|
||||
raise ValueError("Protocol Ollama context size must be positive.")
|
||||
if protocol_safe_input_token_budget <= 0:
|
||||
raise ValueError("Protocol safe input token budget must be positive.")
|
||||
|
||||
diarized_transcript = run_dir / "diarization" / "transcript_diarized.json"
|
||||
plain_transcript = run_dir / "transcript" / "transcript.json"
|
||||
transcript_path = (
|
||||
diarized_transcript if diarized_transcript.is_file() else plain_transcript
|
||||
)
|
||||
if not transcript_path.is_file():
|
||||
raise FileNotFoundError(
|
||||
f"Existing run has no protocol transcript artifact: {run_dir}"
|
||||
)
|
||||
|
||||
context_path = run_dir / "context" / "meeting_context.yaml"
|
||||
write_meeting_context(context, context_path)
|
||||
_emit(progress_sink, "protocol_generation", "started", started)
|
||||
try:
|
||||
result = generate_direct_protocol(
|
||||
transcript_path,
|
||||
context_path,
|
||||
model=model,
|
||||
endpoint=ollama_endpoint,
|
||||
num_ctx=protocol_num_ctx,
|
||||
safe_input_token_budget=protocol_safe_input_token_budget,
|
||||
)
|
||||
protocol_path = _persist_protocol(run_dir, result)
|
||||
except Exception as exc:
|
||||
_emit(
|
||||
progress_sink,
|
||||
"failed",
|
||||
"failed",
|
||||
started,
|
||||
message=f"protocol_generation: {type(exc).__name__}: {exc}",
|
||||
)
|
||||
raise
|
||||
_emit(progress_sink, "protocol_generation", "completed", started)
|
||||
_emit(progress_sink, "completed", "completed", started)
|
||||
return MvpRunResult(0, run_dir, protocol_path)
|
||||
|
||||
|
||||
def create_unique_run_dir(
|
||||
output_root: Path,
|
||||
meeting_name: str,
|
||||
now: Callable[[], datetime] = datetime.now,
|
||||
) -> Path:
|
||||
safe_name = re.sub(r"[^A-Za-z0-9_.-]+", "_", meeting_name).strip("._-") or "meeting"
|
||||
base = output_root / f"{safe_name}_{now().strftime('%Y%m%d_%H%M%S')}"
|
||||
candidate = base
|
||||
suffix = 1
|
||||
while candidate.exists():
|
||||
candidate = output_root / f"{base.name}_{suffix:02d}"
|
||||
suffix += 1
|
||||
candidate.mkdir(parents=True)
|
||||
return candidate
|
||||
|
||||
|
||||
def _write_json(path: Path, value: Any) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def _effective_context(value: ContextInput | None) -> MeetingContext | None:
|
||||
if value is None:
|
||||
return None
|
||||
if isinstance(value, MeetingContext):
|
||||
validate_meeting_context(value.data)
|
||||
return value
|
||||
if isinstance(value, Mapping):
|
||||
return create_meeting_context(dict(value), source_file=Path("<programmatic>"))
|
||||
raise TypeError("meeting_context must be MeetingContext, mapping, or None.")
|
||||
|
||||
|
||||
def _validate_inputs(
|
||||
config: MvpMeetingConfig, meeting_context: MeetingContext | None
|
||||
) -> None:
|
||||
if not config.audio_file.is_file():
|
||||
raise FileNotFoundError(f"Audio file does not exist: {config.audio_file}")
|
||||
if not config.whisper_model.is_file():
|
||||
raise FileNotFoundError(f"Whisper model does not exist: {config.whisper_model}")
|
||||
if meeting_context is not None and config.context_file is not None:
|
||||
raise ValueError("Use either a context file or a programmatic Meeting Context, not both.")
|
||||
if meeting_context is not None:
|
||||
validate_meeting_context(meeting_context.data)
|
||||
elif config.context_file is not None:
|
||||
if not config.context_file.is_file():
|
||||
raise FileNotFoundError(
|
||||
f"Meeting Context file does not exist: {config.context_file}"
|
||||
)
|
||||
load_meeting_context(config.context_file)
|
||||
if config.diarization not in ("off", "auto", "gpu", "cpu"):
|
||||
raise ValueError(f"Unsupported diarization mode: {config.diarization}")
|
||||
if config.diarization_runtime not in ("native", "container"):
|
||||
raise ValueError(
|
||||
f"Unsupported diarization runtime: {config.diarization_runtime}"
|
||||
)
|
||||
if (
|
||||
config.diarization != "off"
|
||||
and config.diarization_runtime == "container"
|
||||
and not config.diarization_container_image
|
||||
):
|
||||
raise ValueError("A diarization container image is required.")
|
||||
if config.protocol_safe_input_token_budget <= 0:
|
||||
raise ValueError("Protocol safe input token budget must be positive.")
|
||||
if config.protocol_num_ctx <= 0:
|
||||
raise ValueError("Protocol Ollama context size must be positive.")
|
||||
|
||||
|
||||
def _emit(
|
||||
sink: ProgressSink | None,
|
||||
stage: str,
|
||||
status: ProgressStatus,
|
||||
overall_started: float,
|
||||
*,
|
||||
message: str | None = None,
|
||||
) -> None:
|
||||
if sink is not None:
|
||||
sink(
|
||||
ProgressEvent(
|
||||
stage=stage,
|
||||
status=status,
|
||||
elapsed_seconds=time.perf_counter() - overall_started,
|
||||
message=message,
|
||||
)
|
||||
)
|
||||
|
||||
|
||||
def _persist_protocol(run_dir: Path, result: DirectProtocolResult) -> Path:
|
||||
protocol_dir = run_dir / "protocol"
|
||||
protocol_dir.mkdir(exist_ok=True)
|
||||
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
|
||||
_write_json(protocol_dir / "raw_response.json", result.raw_response)
|
||||
_write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
|
||||
transcript_input = getattr(result, "transcript_input", None)
|
||||
if transcript_input is not None:
|
||||
(protocol_dir / "transcript_input.txt").write_text(
|
||||
transcript_input, encoding="utf-8"
|
||||
)
|
||||
protocol_path = run_dir / "protocol.md"
|
||||
protocol_path.write_text(result.protocol_text, encoding="utf-8")
|
||||
return protocol_path
|
||||
|
||||
|
||||
def run_mvp_meeting(
|
||||
config: MvpMeetingConfig,
|
||||
*,
|
||||
meeting_context: ContextInput | None = None,
|
||||
progress_sink: ProgressSink | None = None,
|
||||
) -> MvpRunResult:
|
||||
"""Run the existing MVP directly, without subprocess or GUI dependencies."""
|
||||
overall_started = time.perf_counter()
|
||||
validation_started = time.perf_counter()
|
||||
_emit(progress_sink, "preparing", "started", overall_started)
|
||||
try:
|
||||
effective_context = _effective_context(meeting_context)
|
||||
_validate_inputs(config, effective_context)
|
||||
except Exception as exc:
|
||||
_emit(
|
||||
progress_sink,
|
||||
"failed",
|
||||
"failed",
|
||||
overall_started,
|
||||
message=f"preparing: {type(exc).__name__}: {exc}",
|
||||
)
|
||||
print(f"Error: {type(exc).__name__}: {exc}", file=sys.stderr)
|
||||
return MvpRunResult(2, None, None)
|
||||
|
||||
validation_runtime = time.perf_counter() - validation_started
|
||||
run_dir = create_unique_run_dir(config.output_root, config.audio_file.stem)
|
||||
timestamp = datetime.now().astimezone().isoformat(timespec="seconds")
|
||||
transcript_path = run_dir / "transcript" / "transcript.json"
|
||||
protocol_path = run_dir / "protocol.md"
|
||||
stage_runtimes: dict[str, float | None] = {
|
||||
"validation": round(validation_runtime, 3),
|
||||
"setup": None,
|
||||
"audio_preparation": None,
|
||||
"whisper": None,
|
||||
"transcript_validation": None,
|
||||
"protocol": None,
|
||||
}
|
||||
if config.diarization != "off":
|
||||
stage_runtimes["diarization"] = None
|
||||
stage_runtimes["diarization_alignment"] = None
|
||||
metadata: dict[str, Any] = {
|
||||
"run_id": run_dir.name,
|
||||
"timestamp": timestamp,
|
||||
"input_audio": str(config.audio_file.resolve()),
|
||||
"audio_preparation": None,
|
||||
"transcript_output": str(transcript_path.resolve()),
|
||||
"protocol_output": str(protocol_path.resolve()),
|
||||
"whisper_model": str(config.whisper_model.resolve()),
|
||||
"model": config.model,
|
||||
"ollama_endpoint": config.ollama_endpoint,
|
||||
"status": "running",
|
||||
"stage_runtimes_seconds": stage_runtimes,
|
||||
"total_runtime_seconds": None,
|
||||
"failure": None,
|
||||
"diarization": {
|
||||
"enabled": config.diarization != "off",
|
||||
"backend": "pyannote.audio" if config.diarization != "off" else None,
|
||||
"model": DEFAULT_DIARIZATION_MODEL if config.diarization != "off" else None,
|
||||
"requested_device_mode": config.diarization,
|
||||
"runtime": config.diarization_runtime if config.diarization != "off" else None,
|
||||
"metadata_path": None,
|
||||
"transcript_diarized": None,
|
||||
},
|
||||
}
|
||||
current_stage = "preparing"
|
||||
stage_started = time.perf_counter()
|
||||
|
||||
try:
|
||||
audio_dir = run_dir / "audio"
|
||||
transcript_dir = run_dir / "transcript"
|
||||
context_dir = run_dir / "context"
|
||||
protocol_dir = run_dir / "protocol"
|
||||
audio_dir.mkdir()
|
||||
transcript_dir.mkdir()
|
||||
context_dir.mkdir()
|
||||
protocol_dir.mkdir()
|
||||
_write_json(
|
||||
audio_dir / "input_manifest.json",
|
||||
{
|
||||
"source_file": str(config.audio_file.resolve()),
|
||||
"filename": config.audio_file.name,
|
||||
"size_bytes": config.audio_file.stat().st_size,
|
||||
},
|
||||
)
|
||||
|
||||
preparation_started = time.perf_counter()
|
||||
current_stage = "audio_preparation"
|
||||
stage_started = preparation_started
|
||||
prepared_audio = prepare_audio(
|
||||
config.audio_file,
|
||||
audio_dir / "prepared.wav",
|
||||
ffmpeg_executable=config.ffmpeg_executable,
|
||||
normalization_enabled=config.audio_normalization,
|
||||
)
|
||||
stage_runtimes["audio_preparation"] = round(
|
||||
time.perf_counter() - preparation_started, 3
|
||||
)
|
||||
current_stage = "preparing"
|
||||
preparation_metadata = prepared_audio.metadata()
|
||||
metadata["audio_preparation"] = preparation_metadata
|
||||
_write_json(audio_dir / "preparation_metadata.json", preparation_metadata)
|
||||
_write_json(
|
||||
audio_dir / "input_manifest.json",
|
||||
{
|
||||
"source_file": str(config.audio_file.resolve()),
|
||||
"filename": config.audio_file.name,
|
||||
"size_bytes": config.audio_file.stat().st_size,
|
||||
"format": config.audio_file.suffix.lower().removeprefix("."),
|
||||
"prepared_audio": preparation_metadata,
|
||||
},
|
||||
)
|
||||
|
||||
preserved_context: Path | None = None
|
||||
if effective_context is not None:
|
||||
preserved_context = context_dir / "meeting_context.yaml"
|
||||
write_meeting_context(effective_context, preserved_context)
|
||||
elif config.context_file is not None:
|
||||
preserved_context = context_dir / "meeting_context.yaml"
|
||||
shutil.copy2(config.context_file, preserved_context)
|
||||
stage_runtimes["setup"] = round(time.perf_counter() - stage_started, 3)
|
||||
_emit(progress_sink, "preparing", "completed", overall_started)
|
||||
|
||||
current_stage = "transcription"
|
||||
stage_started = time.perf_counter()
|
||||
_emit(progress_sink, "transcription", "started", overall_started)
|
||||
transcription = transcribe_audio(
|
||||
prepared_audio.prepared_path,
|
||||
config.whisper_model,
|
||||
transcript_dir,
|
||||
config.language,
|
||||
executable=config.whisper_executable,
|
||||
threads=config.threads,
|
||||
)
|
||||
stage_runtimes["whisper"] = round(time.perf_counter() - stage_started, 3)
|
||||
_emit(progress_sink, "transcription", "completed", overall_started)
|
||||
|
||||
stage_started = time.perf_counter()
|
||||
load_compact_transcript(transcription.transcript_json)
|
||||
stage_runtimes["transcript_validation"] = round(
|
||||
time.perf_counter() - stage_started, 3
|
||||
)
|
||||
|
||||
protocol_transcript = transcription.transcript_json
|
||||
if config.diarization != "off":
|
||||
current_stage = "diarization"
|
||||
stage_started = time.perf_counter()
|
||||
_emit(progress_sink, "diarization", "started", overall_started)
|
||||
diarization_dir = run_dir / "diarization"
|
||||
diarization = diarize_audio(
|
||||
prepared_audio.prepared_path,
|
||||
diarization_dir,
|
||||
config.diarization,
|
||||
runtime=config.diarization_runtime,
|
||||
container_image=config.diarization_container_image,
|
||||
container_args=config.diarization_container_args,
|
||||
)
|
||||
stage_runtimes["diarization"] = round(
|
||||
time.perf_counter() - stage_started, 3
|
||||
)
|
||||
metadata["diarization"].update(
|
||||
{
|
||||
"actual_device": diarization.metadata.get("actual_device"),
|
||||
"device_name": diarization.metadata.get("device_name"),
|
||||
"runtime_seconds": diarization.metadata.get("runtime_seconds"),
|
||||
"speaker_count": diarization.metadata.get("speaker_count"),
|
||||
"metadata_path": str(diarization.metadata_path.resolve()),
|
||||
}
|
||||
)
|
||||
|
||||
stage_started = time.perf_counter()
|
||||
protocol_transcript, diarized_text = write_diarized_transcript(
|
||||
transcription.transcript_json,
|
||||
diarization.exclusive_turns_json,
|
||||
diarization_dir,
|
||||
)
|
||||
load_compact_transcript(protocol_transcript)
|
||||
stage_runtimes["diarization_alignment"] = round(
|
||||
time.perf_counter() - stage_started, 3
|
||||
)
|
||||
metadata["diarization"].update(
|
||||
{
|
||||
"transcript_diarized": str(protocol_transcript.resolve()),
|
||||
"transcript_diarized_text": str(diarized_text.resolve()),
|
||||
}
|
||||
)
|
||||
_emit(progress_sink, "diarization", "completed", overall_started)
|
||||
|
||||
current_stage = "protocol_generation"
|
||||
stage_started = time.perf_counter()
|
||||
_emit(progress_sink, "protocol_generation", "started", overall_started)
|
||||
result = generate_direct_protocol(
|
||||
protocol_transcript,
|
||||
preserved_context,
|
||||
model=config.model,
|
||||
endpoint=config.ollama_endpoint,
|
||||
num_ctx=config.protocol_num_ctx,
|
||||
safe_input_token_budget=config.protocol_safe_input_token_budget,
|
||||
)
|
||||
stage_runtimes["protocol"] = round(time.perf_counter() - stage_started, 3)
|
||||
protocol_path = _persist_protocol(run_dir, result)
|
||||
_emit(progress_sink, "protocol_generation", "completed", overall_started)
|
||||
metadata["status"] = "completed"
|
||||
_emit(progress_sink, "completed", "completed", overall_started)
|
||||
except Exception as exc:
|
||||
metadata_stage = {
|
||||
"preparing": "setup",
|
||||
"audio_preparation": "audio_preparation",
|
||||
"transcription": "whisper",
|
||||
"diarization": "diarization",
|
||||
"protocol_generation": "protocol",
|
||||
}.get(current_stage, current_stage)
|
||||
runtime_key = metadata_stage
|
||||
if runtime_key in stage_runtimes and stage_runtimes[runtime_key] is None:
|
||||
stage_runtimes[runtime_key] = round(time.perf_counter() - stage_started, 3)
|
||||
metadata["status"] = "failed"
|
||||
metadata["failure"] = {
|
||||
"stage": metadata_stage,
|
||||
"type": type(exc).__name__,
|
||||
"message": str(exc),
|
||||
}
|
||||
protocol_path = None
|
||||
_emit(
|
||||
progress_sink,
|
||||
"failed",
|
||||
"failed",
|
||||
overall_started,
|
||||
message=f"{current_stage}: {type(exc).__name__}: {exc}",
|
||||
)
|
||||
print(f"Error: {type(exc).__name__}: {exc}", file=sys.stderr)
|
||||
finally:
|
||||
metadata["total_runtime_seconds"] = round(time.perf_counter() - overall_started, 3)
|
||||
_write_json(run_dir / "run_metadata.json", metadata)
|
||||
|
||||
exit_code = 0 if metadata["status"] == "completed" else 2
|
||||
return MvpRunResult(exit_code, run_dir, protocol_path)
|
||||
@@ -0,0 +1,21 @@
|
||||
"""Small observer boundary for long-running Meeting Lab operations."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Callable, Literal
|
||||
|
||||
|
||||
ProgressStatus = Literal["started", "completed", "failed"]
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ProgressEvent:
|
||||
stage: str
|
||||
status: ProgressStatus
|
||||
elapsed_seconds: float
|
||||
progress: float | None = None
|
||||
message: str | None = None
|
||||
|
||||
|
||||
ProgressSink = Callable[[ProgressEvent], None]
|
||||
@@ -0,0 +1,32 @@
|
||||
"""Prompt construction for the direct transcript-to-protocol MVP."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
|
||||
DIRECT_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und dem Meeting-Kontext ein vollständiges, strukturiertes und professionelles internes Besprechungsprotokoll in deutscher Sprache.
|
||||
|
||||
Das Protokoll muss themenorientiert sein, nicht chronologisch und nicht nach technischen Kategorien gegliedert. Beginne mit # Meeting Protocol. Verwende für jedes kohärente Thema eine Überschrift ## <Thema> und darunter eine strukturierte Synthese der Diskussion. Bewahre relevante Diskussionsverläufe, unterschiedliche Positionen, offene Punkte und Entscheidungsgrundlagen. Dokumentiere die wesentlichen Inhalte nachvollziehbar und fasse Themenblöcke so zusammen, dass auch Personen, die nicht am Meeting teilgenommen haben, den Kontext und die Entwicklung der Diskussion verstehen können. Nenne Entscheidungen oder abgestimmte Positionen nur, wenn sie tatsächlich belegt sind. Führe Maßnahmen nur auf, wenn eine konkrete zukünftige Handlung gestützt ist; nenne verantwortliche Personen und Fristen ausschließlich bei expliziter Zuweisung, Annahme oder Bestätigung im Transkript. Vorschläge, Einwände, Möglichkeiten und vorläufige Ideen sind keine Entscheidungen oder Verpflichtungen. Bewahre relevante Einschränkungen und ungelöste Meinungsverschiedenheiten. Nenne offene Punkte nur, wenn sie wirklich offen bleiben. Nicht jedes Thema benötigt Entscheidungen, Maßnahmen oder offene Punkte.
|
||||
|
||||
Erzeuge keine reine Wiedergabe des Transkripts und verlängere das Protokoll nicht unnötig durch Wiederholungen. Synthetisiere zusammengehörige Aussagen, entferne Füllwörter und Gesprächsrauschen und erfinde keine Fakten, Entscheidungen, Zustimmungen, Verantwortlichen oder Fristen. Gib kein JSON, keine internen Labels und keine Analyse oder Denkprotokolle aus. Das Ergebnis soll als Markdown-Protokoll nach geringfügiger menschlicher Redaktion intern versendbar sein. Eine kompakte themenübergreifende Maßnahmenliste am Ende ist optional, wenn sie nützlich und vollständig belegt ist."""
|
||||
|
||||
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und Meeting-Kontext ein vollständiges, professionelles internes Besprechungsprotokoll auf Deutsch. Das Transkript ist in aufeinanderfolgende anonyme Sprecherblöcke gegliedert.
|
||||
|
||||
Beginne mit # Meeting Protocol. Gliedere themenorientiert mit ## <Thema> und synthetisiere je Thema den relevanten Diskussionsverlauf, Kontext, unterschiedliche Positionen, Entscheidungsgrundlagen, Einschränkungen und ungelöste Meinungsverschiedenheiten so, dass Dritte ihn nachvollziehen können. Nenne Entscheidungen nur bei Beleg. Nenne Maßnahmen, Verantwortliche und Fristen nur bei expliziter Zuweisung, Annahme oder Bestätigung; Vorschläge sind keine Verpflichtungen.
|
||||
|
||||
Entferne nur Wiederholungen, Füllwörter und Gesprächsrauschen. Erfinde keine Fakten oder Identitäten. Gib kein JSON, keine Sprecherlabels und kein Denkprotokoll aus. Eine belegte themenübergreifende Maßnahmenliste am Ende ist optional."""
|
||||
|
||||
MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION = """Nutze die autoritativen SPEAKER_XX-zu-Teilnehmer-Zuordnungen im Meeting-Kontext, um ausdrücklich belegte Aussagen, Positionen, Entscheidungen, Zuweisungen und angenommene persönliche Verpflichtungen namentlich zuzuordnen. Eine ausdrückliche Ich-Zusage eines zugeordneten Sprechers belegt persönliche Verantwortung. Unterscheide stets den Sprecher einer Aussage von darin nur erwähnten Personen. Leite für nicht zugeordnete Sprecher keine Identität ab und erfinde keine persönliche Verantwortung. Gib die technischen SPEAKER_XX-Bezeichnungen nicht im nutzerseitigen Protokoll aus."""
|
||||
|
||||
|
||||
def build_direct_protocol_prompt(
|
||||
transcript: str,
|
||||
meeting_context: str | None = None,
|
||||
*,
|
||||
instruction: str = DIRECT_PROTOCOL_INSTRUCTION,
|
||||
) -> str:
|
||||
context = meeting_context.strip() if meeting_context else "Kein Meeting-Kontext bereitgestellt."
|
||||
return (
|
||||
f"{instruction}\n\n"
|
||||
f"MEETING-KONTEXT:\n{context}\n\n"
|
||||
f"VOLLSTAENDIGES TRANSKRIPT:\n{transcript.strip()}\n"
|
||||
)
|
||||
@@ -0,0 +1,236 @@
|
||||
"""One-call direct protocol generation from a compact Whisper transcript."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from dataclasses import dataclass
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable
|
||||
|
||||
from src.meeting_lab.llm.ollama import (
|
||||
DEFAULT_ENDPOINT,
|
||||
OllamaGeneration,
|
||||
generate_once,
|
||||
require_model,
|
||||
)
|
||||
from src.meeting_lab.models.meeting_context import (
|
||||
MeetingContext,
|
||||
load_meeting_context,
|
||||
render_meeting_context_for_prompt,
|
||||
)
|
||||
from src.meeting_lab.protocol.direct_protocol_prompt import (
|
||||
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||
MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION,
|
||||
build_direct_protocol_prompt,
|
||||
)
|
||||
from src.meeting_lab.protocol.transcript_input import (
|
||||
TranscriptInputError,
|
||||
compact_diarized_transcript,
|
||||
plain_segment_transcript,
|
||||
)
|
||||
|
||||
|
||||
DEFAULT_MODEL = "qwen3.6:35B-A3B"
|
||||
DEFAULT_NUM_CTX = 32768
|
||||
DEFAULT_NUM_PREDICT = 8192
|
||||
DEFAULT_TIMEOUT = 1800
|
||||
DEFAULT_SAFE_INPUT_TOKEN_BUDGET = 29_000
|
||||
ESTIMATED_UTF8_BYTES_PER_TOKEN = 4.4
|
||||
|
||||
|
||||
class DirectProtocolError(ValueError):
|
||||
"""Raised for invalid direct-protocol inputs or model output."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class DirectProtocolResult:
|
||||
protocol_text: str
|
||||
exact_prompt: str
|
||||
model_metadata: dict[str, Any]
|
||||
runtime_metadata: dict[str, Any]
|
||||
raw_response: dict[str, Any]
|
||||
transcript_input: str | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SelectedTranscriptInput:
|
||||
text: str
|
||||
prompt: str
|
||||
representation: str
|
||||
estimated_input_tokens: int
|
||||
safe_input_token_budget: int
|
||||
fallback_used: bool
|
||||
diarization_enabled: bool
|
||||
|
||||
|
||||
def load_compact_transcript(path: Path) -> str:
|
||||
data = _load_transcript_document(path)
|
||||
text = data.get("text")
|
||||
if not isinstance(text, str) or not text.strip():
|
||||
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
|
||||
return text
|
||||
|
||||
|
||||
def _load_transcript_document(path: Path) -> dict[str, Any]:
|
||||
if not path.is_file():
|
||||
raise DirectProtocolError(f"Transcript file does not exist: {path}")
|
||||
try:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
except json.JSONDecodeError as exc:
|
||||
raise DirectProtocolError(f"Transcript is not valid JSON: {path}: {exc}") from exc
|
||||
if not isinstance(data, dict):
|
||||
raise DirectProtocolError("Transcript JSON must contain a top-level object.")
|
||||
if "text" not in data:
|
||||
raise DirectProtocolError("Transcript JSON must contain top-level 'text'.")
|
||||
return data
|
||||
|
||||
|
||||
def estimate_input_tokens(prompt: str) -> int:
|
||||
"""Estimate tokens without adding a model-specific tokenizer dependency."""
|
||||
byte_count = len(prompt.encode("utf-8"))
|
||||
return max(1, int(byte_count / ESTIMATED_UTF8_BYTES_PER_TOKEN + 0.999999))
|
||||
|
||||
|
||||
def select_transcript_input(
|
||||
transcript: dict[str, Any],
|
||||
rendered_context: str | None,
|
||||
*,
|
||||
safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
) -> SelectedTranscriptInput:
|
||||
"""Select complete prompt input without allowing silent tail truncation."""
|
||||
if safe_input_token_budget <= 0:
|
||||
raise DirectProtocolError("Safe protocol input token budget must be positive.")
|
||||
diarization_enabled = transcript.get("speaker_labels_anonymous") is True
|
||||
if diarization_enabled:
|
||||
try:
|
||||
compact = compact_diarized_transcript(transcript.get("segments"))
|
||||
plain_text = plain_segment_transcript(transcript.get("segments"))
|
||||
except TranscriptInputError as exc:
|
||||
raise DirectProtocolError(str(exc)) from exc
|
||||
instruction = COMPACT_DIARIZED_PROTOCOL_INSTRUCTION
|
||||
if (
|
||||
rendered_context
|
||||
and "Confirmed diarization speaker mappings" in rendered_context
|
||||
):
|
||||
instruction = f"{instruction}\n\n{MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION}"
|
||||
compact_prompt = build_direct_protocol_prompt(
|
||||
compact.text,
|
||||
rendered_context,
|
||||
instruction=instruction,
|
||||
)
|
||||
compact_estimate = estimate_input_tokens(compact_prompt)
|
||||
if compact_estimate <= safe_input_token_budget:
|
||||
return SelectedTranscriptInput(
|
||||
text=compact.text,
|
||||
prompt=compact_prompt,
|
||||
representation="diarized_compact",
|
||||
estimated_input_tokens=compact_estimate,
|
||||
safe_input_token_budget=safe_input_token_budget,
|
||||
fallback_used=False,
|
||||
diarization_enabled=True,
|
||||
)
|
||||
representation = "plain_transcript_fallback"
|
||||
fallback_used = True
|
||||
else:
|
||||
plain_text = transcript.get("text")
|
||||
if not isinstance(plain_text, str) or not plain_text.strip():
|
||||
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
|
||||
representation = "plain_transcript"
|
||||
fallback_used = False
|
||||
|
||||
plain_prompt = build_direct_protocol_prompt(plain_text, rendered_context)
|
||||
plain_estimate = estimate_input_tokens(plain_prompt)
|
||||
if plain_estimate > safe_input_token_budget:
|
||||
raise DirectProtocolError(
|
||||
"Protocol prompt/input is too large for the configured safe input budget "
|
||||
f"({plain_estimate} estimated tokens > {safe_input_token_budget}). "
|
||||
"No LLM request was made; silent truncation is not allowed."
|
||||
)
|
||||
return SelectedTranscriptInput(
|
||||
text=plain_text,
|
||||
prompt=plain_prompt,
|
||||
representation=representation,
|
||||
estimated_input_tokens=plain_estimate,
|
||||
safe_input_token_budget=safe_input_token_budget,
|
||||
fallback_used=fallback_used,
|
||||
diarization_enabled=diarization_enabled,
|
||||
)
|
||||
|
||||
|
||||
def generate_direct_protocol(
|
||||
transcript_path: Path,
|
||||
context_path: Path | None = None,
|
||||
*,
|
||||
model: str = DEFAULT_MODEL,
|
||||
endpoint: str = DEFAULT_ENDPOINT,
|
||||
timeout: int = DEFAULT_TIMEOUT,
|
||||
num_ctx: int = DEFAULT_NUM_CTX,
|
||||
num_predict: int = DEFAULT_NUM_PREDICT,
|
||||
safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
|
||||
model_check: Callable[[str, str, int], dict[str, Any]] = require_model,
|
||||
generation_call: Callable[..., OllamaGeneration] = generate_once,
|
||||
) -> DirectProtocolResult:
|
||||
transcript = _load_transcript_document(transcript_path)
|
||||
context: MeetingContext | None = (
|
||||
load_meeting_context(context_path) if context_path is not None else None
|
||||
)
|
||||
rendered_context = render_meeting_context_for_prompt(context) if context else None
|
||||
selected = select_transcript_input(
|
||||
transcript,
|
||||
rendered_context,
|
||||
safe_input_token_budget=safe_input_token_budget,
|
||||
)
|
||||
|
||||
model_metadata = model_check(endpoint, model, 10)
|
||||
generation = generation_call(
|
||||
endpoint,
|
||||
model,
|
||||
selected.prompt,
|
||||
timeout=timeout,
|
||||
num_ctx=num_ctx,
|
||||
num_predict=num_predict,
|
||||
)
|
||||
data = generation.raw_response
|
||||
runtime_metadata = {
|
||||
"model": model,
|
||||
"prompt_token_count": data.get("prompt_eval_count"),
|
||||
"output_token_count": data.get("eval_count"),
|
||||
"prompt_evaluation_duration_ns": data.get("prompt_eval_duration"),
|
||||
"generation_duration_ns": data.get("eval_duration"),
|
||||
"total_ollama_duration_ns": data.get("total_duration"),
|
||||
"client_wall_time_seconds": generation.client_wall_time_seconds,
|
||||
"completion_reason": data.get("done_reason"),
|
||||
"done": data.get("done"),
|
||||
"request_count": 1,
|
||||
"temperature": 0.0,
|
||||
"think": False,
|
||||
"num_ctx": num_ctx,
|
||||
"num_predict": num_predict,
|
||||
"selected_transcript_representation": selected.representation,
|
||||
"estimated_input_tokens": selected.estimated_input_tokens,
|
||||
"safe_input_token_budget": selected.safe_input_token_budget,
|
||||
"input_token_estimation_method": "utf8_bytes_divided_by_4.4",
|
||||
"fallback_used": selected.fallback_used,
|
||||
"diarization_enabled": selected.diarization_enabled,
|
||||
"speaker_attribution_available": (
|
||||
True
|
||||
if selected.representation == "diarized_compact"
|
||||
else False
|
||||
if selected.representation == "plain_transcript_fallback"
|
||||
else None
|
||||
),
|
||||
"speaker_attribution_loss_reason": (
|
||||
"plain_transcript_fallback"
|
||||
if selected.representation == "plain_transcript_fallback"
|
||||
else None
|
||||
),
|
||||
"speaker_mapping_count": len(context.speaker_mappings) if context else 0,
|
||||
}
|
||||
return DirectProtocolResult(
|
||||
protocol_text=generation.text,
|
||||
exact_prompt=selected.prompt,
|
||||
model_metadata=model_metadata,
|
||||
runtime_metadata=runtime_metadata,
|
||||
raw_response=data,
|
||||
transcript_input=selected.text,
|
||||
)
|
||||
@@ -0,0 +1,97 @@
|
||||
"""Deterministic transcript representations for one-call protocol prompts."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
from dataclasses import dataclass
|
||||
from typing import Any
|
||||
|
||||
|
||||
class TranscriptInputError(ValueError):
|
||||
"""Raised when a transcript cannot be represented without content loss."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SpeakerBlock:
|
||||
"""One contiguous run of transcript segments assigned to one speaker."""
|
||||
|
||||
speaker_id: str
|
||||
segment_texts: tuple[str, ...]
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class CompactDiarizedTranscript:
|
||||
"""Compact prompt text plus structural evidence of segment preservation."""
|
||||
|
||||
text: str
|
||||
blocks: tuple[SpeakerBlock, ...]
|
||||
source_segment_count: int
|
||||
|
||||
@property
|
||||
def represented_segment_count(self) -> int:
|
||||
return sum(len(block.segment_texts) for block in self.blocks)
|
||||
|
||||
@property
|
||||
def segment_texts(self) -> tuple[str, ...]:
|
||||
return tuple(text for block in self.blocks for text in block.segment_texts)
|
||||
|
||||
|
||||
def normalize_segment_text(value: Any, index: int) -> str:
|
||||
"""Normalize formatting whitespace while retaining all semantic text."""
|
||||
if not isinstance(value, str):
|
||||
raise TranscriptInputError(f"Transcript segment {index} text must be a string.")
|
||||
return " ".join(value.split())
|
||||
|
||||
|
||||
def compact_diarized_transcript(segments: Any) -> CompactDiarizedTranscript:
|
||||
"""Group only adjacent same-speaker segments and omit repeated timestamps."""
|
||||
if not isinstance(segments, list) or not segments:
|
||||
raise TranscriptInputError(
|
||||
"Diarized transcript must contain a non-empty 'segments' list."
|
||||
)
|
||||
|
||||
mutable_blocks: list[tuple[str, list[str]]] = []
|
||||
source_texts: list[str] = []
|
||||
for index, segment in enumerate(segments):
|
||||
if not isinstance(segment, dict):
|
||||
raise TranscriptInputError(f"Transcript segment {index} must be an object.")
|
||||
speaker = segment.get("speaker_id") or "SPEAKER_UNASSIGNED"
|
||||
if not isinstance(speaker, str) or not speaker.startswith("SPEAKER_"):
|
||||
raise TranscriptInputError(
|
||||
f"Transcript segment {index} must use an anonymous SPEAKER_ label."
|
||||
)
|
||||
text = normalize_segment_text(segment.get("text"), index)
|
||||
source_texts.append(text)
|
||||
if mutable_blocks and mutable_blocks[-1][0] == speaker:
|
||||
mutable_blocks[-1][1].append(text)
|
||||
else:
|
||||
mutable_blocks.append((speaker, [text]))
|
||||
|
||||
blocks = tuple(
|
||||
SpeakerBlock(speaker_id=speaker, segment_texts=tuple(texts))
|
||||
for speaker, texts in mutable_blocks
|
||||
)
|
||||
rendered = "\n".join(
|
||||
f"{block.speaker_id}: {' '.join(block.segment_texts)}" for block in blocks
|
||||
)
|
||||
result = CompactDiarizedTranscript(
|
||||
text=rendered + "\n",
|
||||
blocks=blocks,
|
||||
source_segment_count=len(segments),
|
||||
)
|
||||
if result.represented_segment_count != len(segments):
|
||||
raise TranscriptInputError("Compact diarized transcript lost source segments.")
|
||||
if result.segment_texts != tuple(source_texts):
|
||||
raise TranscriptInputError("Compact diarized transcript changed segment order or text.")
|
||||
return result
|
||||
|
||||
|
||||
def plain_segment_transcript(segments: Any) -> str:
|
||||
"""Reconstruct plain transcript text from every segment in source order."""
|
||||
if not isinstance(segments, list) or not segments:
|
||||
raise TranscriptInputError("Transcript must contain a non-empty 'segments' list.")
|
||||
texts = []
|
||||
for index, segment in enumerate(segments):
|
||||
if not isinstance(segment, dict):
|
||||
raise TranscriptInputError(f"Transcript segment {index} must be an object.")
|
||||
texts.append(normalize_segment_text(segment.get("text"), index))
|
||||
return " ".join(texts)
|
||||
@@ -0,0 +1 @@
|
||||
"""Isolated experimental semantic synthesis for known discussion subjects."""
|
||||
@@ -0,0 +1,640 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run semantic synthesis with subject detection and evidence assignment fixed."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
SCHEMA_VERSION = "experimental-semantic-synthesis-v1"
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
DEFAULT_MODEL = "qwen3.5:9B"
|
||||
DEFAULT_TIMEOUT = 300
|
||||
DEFAULT_NUM_CTX = 8192
|
||||
DEFAULT_NUM_PREDICT = 2048
|
||||
|
||||
EVENT_TYPES = {
|
||||
"idea",
|
||||
"option",
|
||||
"proposal",
|
||||
"objection",
|
||||
"supporting_argument",
|
||||
"clarification",
|
||||
"rejection",
|
||||
"scoped_acceptance",
|
||||
"fact",
|
||||
"technical_finding",
|
||||
}
|
||||
OUTCOME_STATUSES = {"established", "rejected", "scoped_acceptance", "tentative"}
|
||||
|
||||
|
||||
class SynthesisValidationError(ValueError):
|
||||
"""Raised when isolated semantic synthesis output is structurally invalid."""
|
||||
|
||||
|
||||
PROMPT_TEMPLATE = """You perform semantic synthesis for one already known discussion subject.
|
||||
|
||||
The subject boundary and evidence assignment are fixed and complete. Do not discover,
|
||||
split, merge, rename, or omit the subject. Do not assign evidence to another subject.
|
||||
Interpret only what the supplied evidence semantically establishes.
|
||||
|
||||
Semantic distinctions:
|
||||
- idea: mentioned possibility without stronger commitment
|
||||
- option: alternative considered without commitment
|
||||
- proposal: suggested course of action not yet established as work
|
||||
- objection: argument or concern against something; not automatically unresolved
|
||||
- rejection: an alternative is explicitly rejected
|
||||
- scoped_acceptance: accepted only for the stated test, trial, condition, or scope
|
||||
- proposal is not an action
|
||||
- no decision is not a tentative decision
|
||||
- mention is not an unresolved issue
|
||||
- an action requires explicit assignment, acceptance, commitment, or established work
|
||||
- an unresolved issue requires a concrete need explicitly left unresolved
|
||||
|
||||
Preserve explicit rejection, explicit accepted work, explicit unresolved questions,
|
||||
and all limits on an outcome. Never generalize trial acceptance into final acceptance.
|
||||
Use only supplied evidence IDs. Keep concise semantic text in the evidence language.
|
||||
|
||||
Return exactly one JSON object. Always include these fields:
|
||||
{{
|
||||
"schema_version": "experimental-semantic-synthesis-v1",
|
||||
"subject_id": "copy the supplied subject_id exactly",
|
||||
"subject": "copy the supplied subject exactly",
|
||||
"events": [
|
||||
{{
|
||||
"type": "idea|option|proposal|objection|supporting_argument|clarification|rejection|scoped_acceptance|fact|technical_finding",
|
||||
"text": "supported semantic event",
|
||||
"evidence_ids": ["e1"]
|
||||
}}
|
||||
],
|
||||
"actions": [
|
||||
{{
|
||||
"text": "established action",
|
||||
"responsible": null,
|
||||
"due": null,
|
||||
"evidence_ids": ["e2"]
|
||||
}}
|
||||
],
|
||||
"unresolved_issues": [
|
||||
{{
|
||||
"text": "explicitly unresolved issue",
|
||||
"evidence_ids": ["e3"]
|
||||
}}
|
||||
]
|
||||
}}
|
||||
|
||||
The three arrays are structurally required; use [] when none exist.
|
||||
Add "outcome" only when an outcome was actually established:
|
||||
{{
|
||||
"status": "established|rejected|scoped_acceptance|tentative",
|
||||
"text": "what was actually established",
|
||||
"scope": "the exact scope, condition, or limit",
|
||||
"evidence_ids": ["e2"]
|
||||
}}
|
||||
Omit outcome completely when there is none. Never use null for outcome. Never use the
|
||||
string "null"; use JSON null only for unknown responsible or due values.
|
||||
|
||||
Fixed Gold input:
|
||||
{input_json}
|
||||
"""
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Run the isolated semantic-synthesis Gold experiment."
|
||||
)
|
||||
parser.add_argument("fixture", type=Path, help="Fixed-subject Gold bundle JSON.")
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT)
|
||||
parser.add_argument("--num-ctx", type=int, default=DEFAULT_NUM_CTX)
|
||||
parser.add_argument("--num-predict", type=int, default=DEFAULT_NUM_PREDICT)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def _exact_keys(
|
||||
value: dict[str, Any], required: set[str], optional: set[str], location: str
|
||||
) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required - optional
|
||||
if missing:
|
||||
raise SynthesisValidationError(
|
||||
f"{location} missing required keys: {sorted(missing)}"
|
||||
)
|
||||
if unknown:
|
||||
raise SynthesisValidationError(
|
||||
f"{location} has unknown keys: {sorted(unknown)}"
|
||||
)
|
||||
|
||||
|
||||
def _text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise SynthesisValidationError(f"{location} must be a non-empty string")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def validate_bundle(case: Any) -> dict[str, Any]:
|
||||
if not isinstance(case, dict):
|
||||
raise SynthesisValidationError("case must be an object")
|
||||
_exact_keys(
|
||||
case,
|
||||
{
|
||||
"case_id",
|
||||
"description",
|
||||
"subject_id",
|
||||
"subject",
|
||||
"evidence",
|
||||
"allowed_responsible",
|
||||
"expected",
|
||||
},
|
||||
set(),
|
||||
"case",
|
||||
)
|
||||
_text(case["case_id"], "case.case_id")
|
||||
_text(case["description"], "case.description")
|
||||
_text(case["subject_id"], "case.subject_id")
|
||||
_text(case["subject"], "case.subject")
|
||||
evidence = case["evidence"]
|
||||
if not isinstance(evidence, list) or not evidence:
|
||||
raise SynthesisValidationError("case.evidence must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for index, item in enumerate(evidence):
|
||||
location = f"case.evidence[{index}]"
|
||||
if not isinstance(item, dict):
|
||||
raise SynthesisValidationError(f"{location} must be an object")
|
||||
_exact_keys(item, {"evidence_id", "text"}, set(), location)
|
||||
evidence_id = _text(item["evidence_id"], f"{location}.evidence_id")
|
||||
if evidence_id in seen:
|
||||
raise SynthesisValidationError(f"duplicate evidence ID: {evidence_id}")
|
||||
seen.add(evidence_id)
|
||||
_text(item["text"], f"{location}.text")
|
||||
allowed = case["allowed_responsible"]
|
||||
if not isinstance(allowed, list) or any(
|
||||
not isinstance(value, str) or not value.strip() for value in allowed
|
||||
):
|
||||
raise SynthesisValidationError(
|
||||
"case.allowed_responsible must be a list of non-empty strings"
|
||||
)
|
||||
if len(set(allowed)) != len(allowed):
|
||||
raise SynthesisValidationError("case.allowed_responsible contains duplicates")
|
||||
if not isinstance(case["expected"], dict):
|
||||
raise SynthesisValidationError("case.expected must be an object")
|
||||
return case
|
||||
|
||||
|
||||
def _evidence_ids(value: Any, location: str, known: set[str]) -> list[str]:
|
||||
if not isinstance(value, list) or not value:
|
||||
raise SynthesisValidationError(f"{location} must be a non-empty list")
|
||||
result: list[str] = []
|
||||
for index, evidence_id in enumerate(value):
|
||||
evidence_id = _text(evidence_id, f"{location}[{index}]")
|
||||
if evidence_id not in known:
|
||||
raise SynthesisValidationError(
|
||||
f"{location}[{index}] references unknown evidence ID: {evidence_id}"
|
||||
)
|
||||
if evidence_id in result:
|
||||
raise SynthesisValidationError(
|
||||
f"{location} contains duplicate evidence ID: {evidence_id}"
|
||||
)
|
||||
result.append(evidence_id)
|
||||
return result
|
||||
|
||||
|
||||
def _nullable_text(value: Any, location: str) -> str | None:
|
||||
if value is None:
|
||||
return None
|
||||
result = _text(value, location)
|
||||
if result.casefold() == "null":
|
||||
raise SynthesisValidationError(
|
||||
f"{location} must use JSON null, not the string 'null'"
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def validate_synthesis(data: Any, case: dict[str, Any]) -> dict[str, Any]:
|
||||
validate_bundle(case)
|
||||
if not isinstance(data, dict):
|
||||
raise SynthesisValidationError("output must be an object")
|
||||
_exact_keys(
|
||||
data,
|
||||
{
|
||||
"schema_version",
|
||||
"subject_id",
|
||||
"subject",
|
||||
"events",
|
||||
"actions",
|
||||
"unresolved_issues",
|
||||
},
|
||||
{"outcome"},
|
||||
"output",
|
||||
)
|
||||
if data["schema_version"] != SCHEMA_VERSION:
|
||||
raise SynthesisValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
|
||||
if data["subject_id"] != case["subject_id"]:
|
||||
raise SynthesisValidationError("model changed fixed subject_id")
|
||||
if data["subject"] != case["subject"]:
|
||||
raise SynthesisValidationError("model changed fixed subject")
|
||||
|
||||
known = {item["evidence_id"] for item in case["evidence"]}
|
||||
events = data["events"]
|
||||
if not isinstance(events, list):
|
||||
raise SynthesisValidationError("output.events must be an array")
|
||||
for index, event in enumerate(events):
|
||||
location = f"output.events[{index}]"
|
||||
if not isinstance(event, dict):
|
||||
raise SynthesisValidationError(f"{location} must be an object")
|
||||
_exact_keys(event, {"type", "text", "evidence_ids"}, set(), location)
|
||||
if event["type"] not in EVENT_TYPES:
|
||||
raise SynthesisValidationError(f"{location}.type is invalid")
|
||||
_text(event["text"], f"{location}.text")
|
||||
_evidence_ids(event["evidence_ids"], f"{location}.evidence_ids", known)
|
||||
|
||||
if "outcome" in data:
|
||||
outcome = data["outcome"]
|
||||
if not isinstance(outcome, dict):
|
||||
raise SynthesisValidationError(
|
||||
"output.outcome must be an object when present; omit it when absent"
|
||||
)
|
||||
_exact_keys(
|
||||
outcome, {"status", "text", "scope", "evidence_ids"}, set(), "output.outcome"
|
||||
)
|
||||
if outcome["status"] not in OUTCOME_STATUSES:
|
||||
raise SynthesisValidationError("output.outcome.status is invalid")
|
||||
_text(outcome["text"], "output.outcome.text")
|
||||
_text(outcome["scope"], "output.outcome.scope")
|
||||
_evidence_ids(outcome["evidence_ids"], "output.outcome.evidence_ids", known)
|
||||
|
||||
actions = data["actions"]
|
||||
if not isinstance(actions, list):
|
||||
raise SynthesisValidationError("output.actions must be an array")
|
||||
allowed = set(case["allowed_responsible"])
|
||||
for index, action in enumerate(actions):
|
||||
location = f"output.actions[{index}]"
|
||||
if not isinstance(action, dict):
|
||||
raise SynthesisValidationError(f"{location} must be an object")
|
||||
_exact_keys(
|
||||
action,
|
||||
{"text", "responsible", "due", "evidence_ids"},
|
||||
set(),
|
||||
location,
|
||||
)
|
||||
_text(action["text"], f"{location}.text")
|
||||
responsible = _nullable_text(action["responsible"], f"{location}.responsible")
|
||||
if responsible is not None and responsible not in allowed:
|
||||
raise SynthesisValidationError(
|
||||
f"{location}.responsible is not allowed: {responsible}"
|
||||
)
|
||||
_nullable_text(action["due"], f"{location}.due")
|
||||
_evidence_ids(action["evidence_ids"], f"{location}.evidence_ids", known)
|
||||
|
||||
issues = data["unresolved_issues"]
|
||||
if not isinstance(issues, list):
|
||||
raise SynthesisValidationError("output.unresolved_issues must be an array")
|
||||
for index, issue in enumerate(issues):
|
||||
location = f"output.unresolved_issues[{index}]"
|
||||
if not isinstance(issue, dict):
|
||||
raise SynthesisValidationError(f"{location} must be an object")
|
||||
_exact_keys(issue, {"text", "evidence_ids"}, set(), location)
|
||||
_text(issue["text"], f"{location}.text")
|
||||
_evidence_ids(issue["evidence_ids"], f"{location}.evidence_ids", known)
|
||||
return data
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
validate_bundle(case)
|
||||
model_input = {
|
||||
"subject_id": case["subject_id"],
|
||||
"subject": case["subject"],
|
||||
"evidence": case["evidence"],
|
||||
}
|
||||
return PROMPT_TEMPLATE.format(
|
||||
input_json=json.dumps(model_input, ensure_ascii=False, indent=2)
|
||||
)
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise SynthesisValidationError("model response JSON must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def build_ollama_payload(
|
||||
model: str, prompt: str, num_ctx: int, num_predict: int
|
||||
) -> dict[str, Any]:
|
||||
return {
|
||||
"model": model,
|
||||
"prompt": prompt,
|
||||
"think": False,
|
||||
"stream": False,
|
||||
"format": "json",
|
||||
"options": {
|
||||
"temperature": 0,
|
||||
"num_ctx": num_ctx,
|
||||
"num_predict": num_predict,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def call_ollama(
|
||||
endpoint: str,
|
||||
model: str,
|
||||
prompt: str,
|
||||
timeout: int,
|
||||
num_ctx: int,
|
||||
num_predict: int,
|
||||
) -> tuple[str, dict[str, Any]]:
|
||||
payload = build_ollama_payload(model, prompt, num_ctx, num_predict)
|
||||
started = time.perf_counter()
|
||||
response = requests.post(endpoint, json=payload, timeout=timeout)
|
||||
elapsed = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
body = response.json()
|
||||
if not isinstance(body, dict):
|
||||
raise ValueError("Ollama response must be an object")
|
||||
raw_text = body.get("response")
|
||||
if not isinstance(raw_text, str) or not raw_text.strip():
|
||||
raise ValueError("Ollama returned no usable response text")
|
||||
metadata = {
|
||||
"model": body.get("model", model),
|
||||
"elapsed_seconds": round(elapsed, 3),
|
||||
"total_duration_ns": body.get("total_duration"),
|
||||
"load_duration_ns": body.get("load_duration"),
|
||||
"prompt_eval_count": body.get("prompt_eval_count"),
|
||||
"prompt_eval_duration_ns": body.get("prompt_eval_duration"),
|
||||
"eval_count": body.get("eval_count"),
|
||||
"eval_duration_ns": body.get("eval_duration"),
|
||||
"configuration": {
|
||||
"temperature": 0,
|
||||
"think": False,
|
||||
"num_ctx": num_ctx,
|
||||
"num_predict": num_predict,
|
||||
},
|
||||
}
|
||||
return raw_text.strip(), metadata
|
||||
|
||||
|
||||
def _contains(text: str, terms: list[str]) -> bool:
|
||||
folded = text.casefold()
|
||||
return any(term.casefold() in folded for term in terms)
|
||||
|
||||
|
||||
def _refs_cover(items: list[dict[str, Any]], expected: list[str]) -> bool:
|
||||
actual = {
|
||||
evidence_id
|
||||
for item in items
|
||||
for evidence_id in item.get("evidence_ids", [])
|
||||
}
|
||||
return set(expected).issubset(actual)
|
||||
|
||||
|
||||
def evaluate_synthesis(data: dict[str, Any], expected: dict[str, Any]) -> dict[str, Any]:
|
||||
checks: list[dict[str, Any]] = []
|
||||
|
||||
def add(name: str, passed: bool, critical: bool = False) -> None:
|
||||
checks.append({"name": name, "passed": passed, "critical": critical})
|
||||
|
||||
events = data["events"]
|
||||
event_types = [item["type"] for item in events]
|
||||
for event_type, minimum in expected.get("event_type_minimums", {}).items():
|
||||
add(f"event:{event_type}", event_types.count(event_type) >= minimum)
|
||||
allowed_types = set(expected.get("allowed_event_types", EVENT_TYPES))
|
||||
add("no_unexpected_event_types", set(event_types).issubset(allowed_types))
|
||||
add(
|
||||
"event_evidence",
|
||||
_refs_cover(events, expected.get("event_evidence_ids", [])),
|
||||
)
|
||||
|
||||
outcome_expected = expected["outcome"]
|
||||
outcome = data.get("outcome")
|
||||
add(
|
||||
"outcome_presence",
|
||||
(outcome is not None) == outcome_expected["required"],
|
||||
critical=True,
|
||||
)
|
||||
if outcome_expected["required"] and outcome is not None:
|
||||
add("outcome_status", outcome["status"] in outcome_expected["statuses"])
|
||||
combined = f"{outcome['text']} {outcome['scope']}"
|
||||
add("outcome_meaning", _contains(combined, outcome_expected["terms"]))
|
||||
add(
|
||||
"outcome_scope",
|
||||
_contains(combined, outcome_expected["scope_terms"]),
|
||||
critical=True,
|
||||
)
|
||||
add(
|
||||
"outcome_evidence",
|
||||
set(outcome_expected["evidence_ids"]).issubset(outcome["evidence_ids"]),
|
||||
critical=True,
|
||||
)
|
||||
|
||||
actions = data["actions"]
|
||||
expected_actions = expected["actions"]
|
||||
add(
|
||||
"action_count",
|
||||
len(actions) == expected_actions["count"],
|
||||
critical=True,
|
||||
)
|
||||
if expected_actions["count"] and actions:
|
||||
action_text = " ".join(item["text"] for item in actions)
|
||||
add("action_meaning", _contains(action_text, expected_actions["terms"]))
|
||||
if "responsible" in expected_actions:
|
||||
add(
|
||||
"action_responsible",
|
||||
any(item["responsible"] == expected_actions["responsible"] for item in actions),
|
||||
critical=True,
|
||||
)
|
||||
if expected_actions.get("due_terms"):
|
||||
due_text = " ".join(str(item["due"] or "") for item in actions)
|
||||
add("action_due", _contains(due_text, expected_actions["due_terms"]))
|
||||
add(
|
||||
"action_evidence",
|
||||
_refs_cover(actions, expected_actions["evidence_ids"]),
|
||||
critical=True,
|
||||
)
|
||||
|
||||
issues = data["unresolved_issues"]
|
||||
expected_issues = expected["unresolved_issues"]
|
||||
add(
|
||||
"unresolved_count",
|
||||
len(issues) == expected_issues["count"],
|
||||
critical=True,
|
||||
)
|
||||
if expected_issues["count"] and issues:
|
||||
issue_text = " ".join(item["text"] for item in issues)
|
||||
add("unresolved_meaning", _contains(issue_text, expected_issues["terms"]))
|
||||
add(
|
||||
"unresolved_evidence",
|
||||
_refs_cover(issues, expected_issues["evidence_ids"]),
|
||||
critical=True,
|
||||
)
|
||||
|
||||
passed = sum(item["passed"] for item in checks)
|
||||
critical_failures = [
|
||||
item["name"] for item in checks if item["critical"] and not item["passed"]
|
||||
]
|
||||
ratio = passed / len(checks)
|
||||
if ratio == 1:
|
||||
verdict = "PASS"
|
||||
elif ratio >= 0.7 and not critical_failures:
|
||||
verdict = "PARTIAL"
|
||||
else:
|
||||
verdict = "FAIL"
|
||||
failed = [item["name"] for item in checks if not item["passed"]]
|
||||
return {
|
||||
"verdict": verdict,
|
||||
"reason": "All semantic checks passed." if not failed else "Failed: " + ", ".join(failed),
|
||||
"passed_checks": passed,
|
||||
"check_count": len(checks),
|
||||
"critical_failures": critical_failures,
|
||||
"checks": checks,
|
||||
}
|
||||
|
||||
|
||||
def load_fixture(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict) or set(data) != {"cases"}:
|
||||
raise SynthesisValidationError("fixture must contain exactly a cases list")
|
||||
cases = data["cases"]
|
||||
if not isinstance(cases, list) or not cases:
|
||||
raise SynthesisValidationError("fixture cases must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for case in cases:
|
||||
validate_bundle(case)
|
||||
if case["case_id"] in seen:
|
||||
raise SynthesisValidationError(f"duplicate case ID: {case['case_id']}")
|
||||
seen.add(case["case_id"])
|
||||
return cases
|
||||
|
||||
|
||||
def run_case(
|
||||
case: dict[str, Any],
|
||||
output_root: Path,
|
||||
endpoint: str,
|
||||
model: str,
|
||||
timeout: int,
|
||||
num_ctx: int,
|
||||
num_predict: int,
|
||||
) -> dict[str, Any]:
|
||||
case_dir = output_root / case["case_id"]
|
||||
case_dir.mkdir(parents=True, exist_ok=False)
|
||||
gold_input = {
|
||||
"case_id": case["case_id"],
|
||||
"description": case["description"],
|
||||
"subject_id": case["subject_id"],
|
||||
"subject": case["subject"],
|
||||
"evidence": case["evidence"],
|
||||
}
|
||||
(case_dir / "gold_input.json").write_text(
|
||||
json.dumps(gold_input, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
raw_text, metadata = call_ollama(
|
||||
endpoint, model, prompt, timeout, num_ctx, num_predict
|
||||
)
|
||||
(case_dir / "raw_model_response.txt").write_text(raw_text + "\n", encoding="utf-8")
|
||||
(case_dir / "ollama_metadata.json").write_text(
|
||||
json.dumps(metadata, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
parsed = parse_model_json(raw_text)
|
||||
(case_dir / "parsed_response.json").write_text(
|
||||
json.dumps(parsed, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
validated = validate_synthesis(parsed, case)
|
||||
validation = {"valid": True, "error": None}
|
||||
evaluation = evaluate_synthesis(validated, case["expected"])
|
||||
except requests.RequestException:
|
||||
raise
|
||||
except (json.JSONDecodeError, SynthesisValidationError, ValueError) as exc:
|
||||
validation = {
|
||||
"valid": False,
|
||||
"error_type": type(exc).__name__,
|
||||
"error": str(exc),
|
||||
}
|
||||
evaluation = {
|
||||
"verdict": "FAIL",
|
||||
"reason": f"Schema validation failed: {exc}",
|
||||
"passed_checks": 0,
|
||||
"check_count": 0,
|
||||
"critical_failures": ["schema_validation"],
|
||||
"checks": [],
|
||||
}
|
||||
(case_dir / "validation_result.json").write_text(
|
||||
json.dumps(validation, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
result = {
|
||||
"case_id": case["case_id"],
|
||||
"description": case["description"],
|
||||
**evaluation,
|
||||
"elapsed_seconds": round(time.perf_counter() - started, 3),
|
||||
}
|
||||
(case_dir / "evaluation_result.json").write_text(
|
||||
json.dumps(result, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_fixture(args.fixture)
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
started = time.perf_counter()
|
||||
results: list[dict[str, Any]] = []
|
||||
for index, case in enumerate(cases, start=1):
|
||||
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
|
||||
results.append(
|
||||
run_case(
|
||||
case,
|
||||
args.output,
|
||||
args.endpoint,
|
||||
args.model,
|
||||
args.timeout,
|
||||
args.num_ctx,
|
||||
args.num_predict,
|
||||
)
|
||||
)
|
||||
summary = {
|
||||
"experiment": "semantic_synthesis_isolation",
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"model": args.model,
|
||||
"temperature": 0,
|
||||
"think": False,
|
||||
"num_ctx": args.num_ctx,
|
||||
"num_predict": args.num_predict,
|
||||
"case_count": len(cases),
|
||||
"llm_call_count": len(results),
|
||||
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||
"verdict_counts": {
|
||||
verdict: sum(result["verdict"] == verdict for result in results)
|
||||
for verdict in ("PASS", "PARTIAL", "FAIL")
|
||||
},
|
||||
"results": results,
|
||||
}
|
||||
(args.output / "summary.json").write_text(
|
||||
json.dumps(summary, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
return summary
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
try:
|
||||
summary = run_experiment(args)
|
||||
except (OSError, ValueError, requests.RequestException) as exc:
|
||||
print(f"Error: {exc}")
|
||||
return 1
|
||||
print(json.dumps(summary["verdict_counts"], sort_keys=True))
|
||||
print(f"Artifacts: {args.output.resolve()}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1 @@
|
||||
"""Experimental topic-oriented discussion reconstruction."""
|
||||
@@ -0,0 +1,721 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run an isolated Discussion Subject reconstruction experiment with Ollama."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import re
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
SCHEMA_VERSION = "experimental-discussion-subjects-v1"
|
||||
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
DEFAULT_MODEL = "qwen3.5:9B"
|
||||
DEFAULT_TIMEOUT = 300
|
||||
DEFAULT_NUM_CTX = 16384
|
||||
DEFAULT_NUM_PREDICT = 4096
|
||||
|
||||
EVENT_TYPES = {
|
||||
"introduced_idea",
|
||||
"considered_option",
|
||||
"proposal",
|
||||
"supporting_argument",
|
||||
"objection",
|
||||
"clarification",
|
||||
"modification",
|
||||
"fact",
|
||||
"technical_finding",
|
||||
}
|
||||
OUTCOME_CERTAINTIES = {"established", "tentative", "conditional", "rejected"}
|
||||
IDENTIFIER_RE = re.compile(r"^[a-z][a-z0-9_]*$")
|
||||
|
||||
|
||||
class ReconstructionValidationError(ValueError):
|
||||
"""Raised when experimental reconstruction output violates the schema."""
|
||||
|
||||
|
||||
PROMPT_TEMPLATE = """You reconstruct discussion subjects from meeting evidence.
|
||||
|
||||
This is semantic reconstruction, not protocol writing and not flat category extraction.
|
||||
Group evidence by what participants are actually discussing. For each subject, record
|
||||
only supported discourse events and, when present, the actual outcome, resulting
|
||||
actions, and genuinely unresolved issues.
|
||||
|
||||
Important distinctions:
|
||||
- discussed is not necessarily proposed
|
||||
- proposed is not necessarily preferred or accepted
|
||||
- preferred is not accepted
|
||||
- accepted for a trial is not accepted as a final solution
|
||||
- mentioned is not an unresolved question
|
||||
- an outcome must preserve its scope, conditions, polarity, and uncertainty
|
||||
- do not infer responsibility from mention, expertise, adjacency, or likely role
|
||||
- do not invent missing stages or emit empty optional structures
|
||||
|
||||
Evidence discipline:
|
||||
- Use only the supplied evidence IDs in evidence_refs.
|
||||
- Every subject, event, outcome, action, and unresolved issue needs at least one
|
||||
evidence reference.
|
||||
- Keep statements concise; do not copy long evidence passages.
|
||||
- A subject may consist only of one introduced idea.
|
||||
|
||||
Return one JSON object with exactly:
|
||||
{{
|
||||
"schema_version": "experimental-discussion-subjects-v1",
|
||||
"subjects": [
|
||||
{{
|
||||
"subject_id": "subject_1",
|
||||
"title": "concise discussion subject",
|
||||
"evidence_refs": ["e1"],
|
||||
"development": [
|
||||
{{
|
||||
"event_id": "event_1",
|
||||
"type": "introduced_idea|considered_option|proposal|supporting_argument|objection|clarification|modification|fact|technical_finding",
|
||||
"text": "what happened in the discussion",
|
||||
"evidence_refs": ["e1"]
|
||||
}}
|
||||
],
|
||||
"outcome": {{
|
||||
"text": "only what was established",
|
||||
"scope": "explicit limit or full scope of the outcome",
|
||||
"certainty": "established|tentative|conditional|rejected",
|
||||
"evidence_refs": ["e2"]
|
||||
}},
|
||||
"actions": [
|
||||
{{
|
||||
"action_id": "action_1",
|
||||
"text": "established work only",
|
||||
"responsible": "explicitly supported name or null",
|
||||
"deadline": "explicitly supported deadline or null",
|
||||
"evidence_refs": ["e3"]
|
||||
}}
|
||||
],
|
||||
"unresolved_issues": [
|
||||
{{
|
||||
"issue_id": "issue_1",
|
||||
"text": "concrete unresolved issue",
|
||||
"evidence_refs": ["e4"]
|
||||
}}
|
||||
]
|
||||
}}
|
||||
]
|
||||
}}
|
||||
|
||||
Only subject_id, title, evidence_refs are required for each subject. Omit
|
||||
development, outcome, actions, or unresolved_issues when absent. Never emit null
|
||||
or an empty optional list/object.
|
||||
|
||||
Case ID: {case_id}
|
||||
Evidence units:
|
||||
{evidence_json}
|
||||
"""
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(
|
||||
description="Run the isolated topic-reconstruction Gold experiment."
|
||||
)
|
||||
parser.add_argument("fixture", type=Path, help="Focused Gold cases JSON.")
|
||||
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT)
|
||||
parser.add_argument("--num-ctx", type=int, default=DEFAULT_NUM_CTX)
|
||||
parser.add_argument("--num-predict", type=int, default=DEFAULT_NUM_PREDICT)
|
||||
parser.add_argument(
|
||||
"--case", action="append", dest="case_ids", help="Run only this case ID."
|
||||
)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def _expect_exact_keys(
|
||||
value: dict[str, Any], required: set[str], optional: set[str], location: str
|
||||
) -> None:
|
||||
missing = required - value.keys()
|
||||
unknown = value.keys() - required - optional
|
||||
if missing:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location} missing required keys: {sorted(missing)}"
|
||||
)
|
||||
if unknown:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location} has unknown keys: {sorted(unknown)}"
|
||||
)
|
||||
|
||||
|
||||
def _nonempty_text(value: Any, location: str) -> str:
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
raise ReconstructionValidationError(f"{location} must be a non-empty string")
|
||||
return value.strip()
|
||||
|
||||
|
||||
def _identifier(value: Any, location: str, seen: set[str]) -> str:
|
||||
text = _nonempty_text(value, location)
|
||||
if not IDENTIFIER_RE.fullmatch(text):
|
||||
raise ReconstructionValidationError(f"{location} is not a valid identifier")
|
||||
if text in seen:
|
||||
raise ReconstructionValidationError(f"duplicate identifier: {text}")
|
||||
seen.add(text)
|
||||
return text
|
||||
|
||||
|
||||
def _nullable_text(value: Any, location: str) -> str | None:
|
||||
if value is None:
|
||||
return None
|
||||
text = _nonempty_text(value, location)
|
||||
if text.casefold() == "null":
|
||||
raise ReconstructionValidationError(
|
||||
f"{location} must use JSON null, not the string 'null'"
|
||||
)
|
||||
return text
|
||||
|
||||
|
||||
def _evidence_refs(value: Any, location: str, known: set[str]) -> list[str]:
|
||||
if not isinstance(value, list) or not value:
|
||||
raise ReconstructionValidationError(f"{location} must be a non-empty list")
|
||||
refs: list[str] = []
|
||||
for index, ref in enumerate(value):
|
||||
ref = _nonempty_text(ref, f"{location}[{index}]")
|
||||
if ref not in known:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location}[{index}] references unknown evidence ID: {ref}"
|
||||
)
|
||||
if ref in refs:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location} contains duplicate evidence reference: {ref}"
|
||||
)
|
||||
refs.append(ref)
|
||||
return refs
|
||||
|
||||
|
||||
def validate_evidence_units(evidence_units: Any) -> set[str]:
|
||||
if not isinstance(evidence_units, list) or not evidence_units:
|
||||
raise ReconstructionValidationError("evidence_units must be a non-empty list")
|
||||
known: set[str] = set()
|
||||
for index, unit in enumerate(evidence_units):
|
||||
location = f"evidence_units[{index}]"
|
||||
if not isinstance(unit, dict):
|
||||
raise ReconstructionValidationError(f"{location} must be an object")
|
||||
_expect_exact_keys(unit, {"evidence_id", "text"}, set(), location)
|
||||
evidence_id = _nonempty_text(unit["evidence_id"], f"{location}.evidence_id")
|
||||
if evidence_id in known:
|
||||
raise ReconstructionValidationError(
|
||||
f"duplicate input evidence identifier: {evidence_id}"
|
||||
)
|
||||
known.add(evidence_id)
|
||||
_nonempty_text(unit["text"], f"{location}.text")
|
||||
return known
|
||||
|
||||
|
||||
def validate_reconstruction(data: Any, evidence_units: Any) -> dict[str, Any]:
|
||||
known = validate_evidence_units(evidence_units)
|
||||
if not isinstance(data, dict):
|
||||
raise ReconstructionValidationError("model output must be an object")
|
||||
_expect_exact_keys(data, {"schema_version", "subjects"}, set(), "output")
|
||||
if data["schema_version"] != SCHEMA_VERSION:
|
||||
raise ReconstructionValidationError(
|
||||
f"schema_version must be {SCHEMA_VERSION!r}"
|
||||
)
|
||||
subjects = data["subjects"]
|
||||
if not isinstance(subjects, list) or not subjects:
|
||||
raise ReconstructionValidationError("subjects must be a non-empty list")
|
||||
|
||||
seen: set[str] = set()
|
||||
for subject_index, subject in enumerate(subjects):
|
||||
location = f"subjects[{subject_index}]"
|
||||
if not isinstance(subject, dict):
|
||||
raise ReconstructionValidationError(f"{location} must be an object")
|
||||
_expect_exact_keys(
|
||||
subject,
|
||||
{"subject_id", "title", "evidence_refs"},
|
||||
{"development", "outcome", "actions", "unresolved_issues"},
|
||||
location,
|
||||
)
|
||||
_identifier(subject["subject_id"], f"{location}.subject_id", seen)
|
||||
_nonempty_text(subject["title"], f"{location}.title")
|
||||
_evidence_refs(subject["evidence_refs"], f"{location}.evidence_refs", known)
|
||||
|
||||
if "development" in subject:
|
||||
events = subject["development"]
|
||||
if not isinstance(events, list) or not events:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location}.development must be a non-empty list when present"
|
||||
)
|
||||
for event_index, event in enumerate(events):
|
||||
event_location = f"{location}.development[{event_index}]"
|
||||
if not isinstance(event, dict):
|
||||
raise ReconstructionValidationError(
|
||||
f"{event_location} must be an object"
|
||||
)
|
||||
_expect_exact_keys(
|
||||
event,
|
||||
{"event_id", "type", "text", "evidence_refs"},
|
||||
set(),
|
||||
event_location,
|
||||
)
|
||||
_identifier(event["event_id"], f"{event_location}.event_id", seen)
|
||||
if event["type"] not in EVENT_TYPES:
|
||||
raise ReconstructionValidationError(
|
||||
f"{event_location}.type is invalid: {event['type']!r}"
|
||||
)
|
||||
_nonempty_text(event["text"], f"{event_location}.text")
|
||||
_evidence_refs(
|
||||
event["evidence_refs"], f"{event_location}.evidence_refs", known
|
||||
)
|
||||
|
||||
if "outcome" in subject:
|
||||
outcome = subject["outcome"]
|
||||
outcome_location = f"{location}.outcome"
|
||||
if not isinstance(outcome, dict):
|
||||
raise ReconstructionValidationError(
|
||||
f"{outcome_location} must be a non-empty object when present"
|
||||
)
|
||||
_expect_exact_keys(
|
||||
outcome,
|
||||
{"text", "scope", "certainty", "evidence_refs"},
|
||||
set(),
|
||||
outcome_location,
|
||||
)
|
||||
_nonempty_text(outcome["text"], f"{outcome_location}.text")
|
||||
_nonempty_text(outcome["scope"], f"{outcome_location}.scope")
|
||||
if outcome["certainty"] not in OUTCOME_CERTAINTIES:
|
||||
raise ReconstructionValidationError(
|
||||
f"{outcome_location}.certainty is invalid: {outcome['certainty']!r}"
|
||||
)
|
||||
_evidence_refs(
|
||||
outcome["evidence_refs"], f"{outcome_location}.evidence_refs", known
|
||||
)
|
||||
|
||||
if "actions" in subject:
|
||||
actions = subject["actions"]
|
||||
if not isinstance(actions, list) or not actions:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location}.actions must be a non-empty list when present"
|
||||
)
|
||||
for action_index, action in enumerate(actions):
|
||||
action_location = f"{location}.actions[{action_index}]"
|
||||
if not isinstance(action, dict):
|
||||
raise ReconstructionValidationError(
|
||||
f"{action_location} must be an object"
|
||||
)
|
||||
_expect_exact_keys(
|
||||
action,
|
||||
{"action_id", "text", "responsible", "deadline", "evidence_refs"},
|
||||
set(),
|
||||
action_location,
|
||||
)
|
||||
_identifier(action["action_id"], f"{action_location}.action_id", seen)
|
||||
_nonempty_text(action["text"], f"{action_location}.text")
|
||||
for field in ("responsible", "deadline"):
|
||||
_nullable_text(action[field], f"{action_location}.{field}")
|
||||
_evidence_refs(
|
||||
action["evidence_refs"], f"{action_location}.evidence_refs", known
|
||||
)
|
||||
|
||||
if "unresolved_issues" in subject:
|
||||
issues = subject["unresolved_issues"]
|
||||
if not isinstance(issues, list) or not issues:
|
||||
raise ReconstructionValidationError(
|
||||
f"{location}.unresolved_issues must be a non-empty list when present"
|
||||
)
|
||||
for issue_index, issue in enumerate(issues):
|
||||
issue_location = f"{location}.unresolved_issues[{issue_index}]"
|
||||
if not isinstance(issue, dict):
|
||||
raise ReconstructionValidationError(
|
||||
f"{issue_location} must be an object"
|
||||
)
|
||||
_expect_exact_keys(
|
||||
issue,
|
||||
{"issue_id", "text", "evidence_refs"},
|
||||
set(),
|
||||
issue_location,
|
||||
)
|
||||
_identifier(issue["issue_id"], f"{issue_location}.issue_id", seen)
|
||||
_nonempty_text(issue["text"], f"{issue_location}.text")
|
||||
_evidence_refs(
|
||||
issue["evidence_refs"], f"{issue_location}.evidence_refs", known
|
||||
)
|
||||
|
||||
return data
|
||||
|
||||
|
||||
def build_prompt(case: dict[str, Any]) -> str:
|
||||
evidence_units = case["evidence_units"]
|
||||
validate_evidence_units(evidence_units)
|
||||
return PROMPT_TEMPLATE.format(
|
||||
case_id=case["case_id"],
|
||||
evidence_json=json.dumps(evidence_units, ensure_ascii=False, indent=2),
|
||||
)
|
||||
|
||||
|
||||
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||
data = json.loads(raw_text)
|
||||
if not isinstance(data, dict):
|
||||
raise ReconstructionValidationError("model response JSON must be an object")
|
||||
return data
|
||||
|
||||
|
||||
def build_ollama_payload(
|
||||
model: str, prompt: str, num_ctx: int, num_predict: int
|
||||
) -> dict[str, Any]:
|
||||
return {
|
||||
"model": model,
|
||||
"prompt": prompt,
|
||||
"think": False,
|
||||
"stream": False,
|
||||
"format": "json",
|
||||
"options": {
|
||||
"temperature": 0,
|
||||
"num_ctx": num_ctx,
|
||||
"num_predict": num_predict,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def call_ollama(
|
||||
endpoint: str,
|
||||
model: str,
|
||||
prompt: str,
|
||||
timeout: int,
|
||||
num_ctx: int,
|
||||
num_predict: int,
|
||||
) -> tuple[str, dict[str, Any]]:
|
||||
payload = build_ollama_payload(model, prompt, num_ctx, num_predict)
|
||||
started = time.perf_counter()
|
||||
response = requests.post(endpoint, json=payload, timeout=timeout)
|
||||
elapsed = time.perf_counter() - started
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
if not isinstance(data, dict):
|
||||
raise ValueError("Ollama response must be a JSON object")
|
||||
raw_text = data.get("response")
|
||||
if not isinstance(raw_text, str) or not raw_text.strip():
|
||||
raise ValueError("Ollama returned no usable response text")
|
||||
metadata = {
|
||||
"model": data.get("model", model),
|
||||
"elapsed_seconds": round(elapsed, 3),
|
||||
"total_duration_ns": data.get("total_duration"),
|
||||
"load_duration_ns": data.get("load_duration"),
|
||||
"prompt_eval_count": data.get("prompt_eval_count"),
|
||||
"prompt_eval_duration_ns": data.get("prompt_eval_duration"),
|
||||
"eval_count": data.get("eval_count"),
|
||||
"eval_duration_ns": data.get("eval_duration"),
|
||||
"configuration": {
|
||||
"temperature": 0,
|
||||
"think": False,
|
||||
"num_ctx": num_ctx,
|
||||
"num_predict": num_predict,
|
||||
},
|
||||
}
|
||||
return raw_text.strip(), metadata
|
||||
|
||||
|
||||
def _all_text(subjects: list[dict[str, Any]]) -> str:
|
||||
parts: list[str] = []
|
||||
for subject in subjects:
|
||||
parts.append(subject["title"])
|
||||
for event in subject.get("development", []):
|
||||
parts.append(event["text"])
|
||||
outcome = subject.get("outcome")
|
||||
if outcome:
|
||||
parts.extend((outcome["text"], outcome["scope"]))
|
||||
for action in subject.get("actions", []):
|
||||
parts.append(action["text"])
|
||||
for issue in subject.get("unresolved_issues", []):
|
||||
parts.append(issue["text"])
|
||||
return " ".join(parts).casefold()
|
||||
|
||||
|
||||
def _contains_any(text: str, terms: list[str]) -> bool:
|
||||
return any(term.casefold() in text for term in terms)
|
||||
|
||||
|
||||
def evaluate_reconstruction(
|
||||
reconstruction: dict[str, Any], expected: dict[str, Any]
|
||||
) -> dict[str, Any]:
|
||||
subjects = reconstruction["subjects"]
|
||||
combined = _all_text(subjects)
|
||||
events = [event for subject in subjects for event in subject.get("development", [])]
|
||||
outcomes = [subject["outcome"] for subject in subjects if "outcome" in subject]
|
||||
actions = [action for subject in subjects for action in subject.get("actions", [])]
|
||||
issues = [issue for subject in subjects for issue in subject.get("unresolved_issues", [])]
|
||||
checks: list[dict[str, Any]] = []
|
||||
|
||||
def add(name: str, passed: bool, critical: bool = False) -> None:
|
||||
checks.append({"name": name, "passed": passed, "critical": critical})
|
||||
|
||||
add("subject_count", len(subjects) == expected.get("subject_count", 1))
|
||||
add("subject_identity", _contains_any(combined, expected["subject_terms"]))
|
||||
|
||||
event_types = {event["type"] for event in events}
|
||||
for event_type in expected.get("required_event_types", []):
|
||||
add(f"event_type:{event_type}", event_type in event_types)
|
||||
|
||||
expected_outcome = expected.get("outcome", {})
|
||||
outcome_required = expected_outcome.get("required", False)
|
||||
add(
|
||||
"outcome_presence",
|
||||
bool(outcomes) is outcome_required,
|
||||
critical=not outcome_required and bool(outcomes),
|
||||
)
|
||||
if outcome_required and outcomes:
|
||||
outcome_text = " ".join(
|
||||
f"{item['text']} {item['scope']}" for item in outcomes
|
||||
).casefold()
|
||||
add("outcome_meaning", _contains_any(outcome_text, expected_outcome["terms"]))
|
||||
add(
|
||||
"outcome_scope",
|
||||
_contains_any(outcome_text, expected_outcome.get("scope_terms", [])),
|
||||
critical=True,
|
||||
)
|
||||
add(
|
||||
"outcome_certainty",
|
||||
any(
|
||||
item["certainty"] in expected_outcome.get("certainties", [])
|
||||
for item in outcomes
|
||||
),
|
||||
)
|
||||
|
||||
expected_actions = expected.get("actions", {})
|
||||
minimum_actions = expected_actions.get("minimum", 0)
|
||||
add(
|
||||
"action_count",
|
||||
len(actions) >= minimum_actions if minimum_actions else not actions,
|
||||
critical=minimum_actions == 0 and bool(actions),
|
||||
)
|
||||
if minimum_actions and actions:
|
||||
action_text = " ".join(item["text"] for item in actions).casefold()
|
||||
add("action_meaning", _contains_any(action_text, expected_actions["terms"]))
|
||||
if "responsible" in expected_actions:
|
||||
add(
|
||||
"action_responsibility",
|
||||
any(
|
||||
item["responsible"] == expected_actions["responsible"]
|
||||
for item in actions
|
||||
),
|
||||
critical=True,
|
||||
)
|
||||
|
||||
expected_issues = expected.get("unresolved", {})
|
||||
minimum_issues = expected_issues.get("minimum", 0)
|
||||
add(
|
||||
"unresolved_count",
|
||||
len(issues) >= minimum_issues if minimum_issues else not issues,
|
||||
critical=minimum_issues == 0 and bool(issues),
|
||||
)
|
||||
if minimum_issues and issues:
|
||||
issue_text = " ".join(item["text"] for item in issues).casefold()
|
||||
add("unresolved_meaning", _contains_any(issue_text, expected_issues["terms"]))
|
||||
|
||||
passed = sum(check["passed"] for check in checks)
|
||||
critical_failures = [
|
||||
check["name"] for check in checks if check["critical"] and not check["passed"]
|
||||
]
|
||||
ratio = passed / len(checks)
|
||||
if ratio == 1:
|
||||
verdict = "PASS"
|
||||
elif ratio >= 0.6 and not critical_failures:
|
||||
verdict = "PARTIAL"
|
||||
else:
|
||||
verdict = "FAIL"
|
||||
failed = [check["name"] for check in checks if not check["passed"]]
|
||||
reason = "All semantic checks passed." if not failed else "Failed: " + ", ".join(failed)
|
||||
return {
|
||||
"verdict": verdict,
|
||||
"reason": reason,
|
||||
"passed_checks": passed,
|
||||
"check_count": len(checks),
|
||||
"critical_failures": critical_failures,
|
||||
"checks": checks,
|
||||
}
|
||||
|
||||
|
||||
def load_fixture(path: Path) -> list[dict[str, Any]]:
|
||||
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||
if not isinstance(data, dict) or set(data) != {"cases"}:
|
||||
raise ValueError("fixture must contain exactly one 'cases' list")
|
||||
cases = data["cases"]
|
||||
if not isinstance(cases, list) or not cases:
|
||||
raise ValueError("fixture cases must be a non-empty list")
|
||||
seen: set[str] = set()
|
||||
for index, case in enumerate(cases):
|
||||
if not isinstance(case, dict):
|
||||
raise ValueError(f"cases[{index}] must be an object")
|
||||
required = {"case_id", "description", "evidence_units", "expected"}
|
||||
if set(case) != required:
|
||||
raise ValueError(f"cases[{index}] must contain exactly {sorted(required)}")
|
||||
case_id = _nonempty_text(case["case_id"], f"cases[{index}].case_id")
|
||||
if case_id in seen:
|
||||
raise ValueError(f"duplicate case_id: {case_id}")
|
||||
seen.add(case_id)
|
||||
_nonempty_text(case["description"], f"cases[{index}].description")
|
||||
validate_evidence_units(case["evidence_units"])
|
||||
if not isinstance(case["expected"], dict):
|
||||
raise ValueError(f"cases[{index}].expected must be an object")
|
||||
return cases
|
||||
|
||||
|
||||
def run_case(
|
||||
case: dict[str, Any],
|
||||
output_root: Path,
|
||||
endpoint: str,
|
||||
model: str,
|
||||
timeout: int,
|
||||
num_ctx: int,
|
||||
num_predict: int,
|
||||
) -> dict[str, Any]:
|
||||
case_dir = output_root / case["case_id"]
|
||||
case_dir.mkdir(parents=True, exist_ok=False)
|
||||
input_payload = {
|
||||
"case_id": case["case_id"],
|
||||
"description": case["description"],
|
||||
"evidence_units": case["evidence_units"],
|
||||
}
|
||||
(case_dir / "input.json").write_text(
|
||||
json.dumps(input_payload, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
prompt = build_prompt(case)
|
||||
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||
|
||||
started = time.perf_counter()
|
||||
try:
|
||||
raw_text, metadata = call_ollama(
|
||||
endpoint, model, prompt, timeout, num_ctx, num_predict
|
||||
)
|
||||
(case_dir / "raw_model_response.txt").write_text(
|
||||
raw_text + "\n", encoding="utf-8"
|
||||
)
|
||||
(case_dir / "ollama_metadata.json").write_text(
|
||||
json.dumps(metadata, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
parsed = parse_model_json(raw_text)
|
||||
(case_dir / "parsed_output.json").write_text(
|
||||
json.dumps(parsed, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
validated = validate_reconstruction(parsed, case["evidence_units"])
|
||||
evaluation = evaluate_reconstruction(validated, case["expected"])
|
||||
except requests.RequestException as exc:
|
||||
failure = {
|
||||
"case_id": case["case_id"],
|
||||
"error_type": type(exc).__name__,
|
||||
"error": str(exc),
|
||||
"elapsed_seconds": round(time.perf_counter() - started, 3),
|
||||
}
|
||||
(case_dir / "validation_failure.json").write_text(
|
||||
json.dumps(failure, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
raise
|
||||
except (json.JSONDecodeError, ReconstructionValidationError, ValueError) as exc:
|
||||
elapsed = round(time.perf_counter() - started, 3)
|
||||
failure = {
|
||||
"case_id": case["case_id"],
|
||||
"error_type": type(exc).__name__,
|
||||
"error": str(exc),
|
||||
"elapsed_seconds": elapsed,
|
||||
}
|
||||
(case_dir / "validation_failure.json").write_text(
|
||||
json.dumps(failure, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
result = {
|
||||
"case_id": case["case_id"],
|
||||
"description": case["description"],
|
||||
"verdict": "FAIL",
|
||||
"reason": f"{type(exc).__name__}: {exc}",
|
||||
"passed_checks": 0,
|
||||
"check_count": 0,
|
||||
"critical_failures": ["schema_validation"],
|
||||
"checks": [],
|
||||
"elapsed_seconds": elapsed,
|
||||
"subject_titles": [],
|
||||
}
|
||||
(case_dir / "evaluation.json").write_text(
|
||||
json.dumps(result, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return result
|
||||
|
||||
result = {
|
||||
"case_id": case["case_id"],
|
||||
"description": case["description"],
|
||||
**evaluation,
|
||||
"elapsed_seconds": metadata["elapsed_seconds"],
|
||||
"subject_titles": [item["title"] for item in validated["subjects"]],
|
||||
}
|
||||
(case_dir / "evaluation.json").write_text(
|
||||
json.dumps(result, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||
cases = load_fixture(args.fixture)
|
||||
selected = set(args.case_ids or [])
|
||||
if selected:
|
||||
known = {case["case_id"] for case in cases}
|
||||
unknown = selected - known
|
||||
if unknown:
|
||||
raise ValueError(f"unknown requested case IDs: {sorted(unknown)}")
|
||||
cases = [case for case in cases if case["case_id"] in selected]
|
||||
|
||||
args.output.mkdir(parents=True, exist_ok=False)
|
||||
results: list[dict[str, Any]] = []
|
||||
started = time.perf_counter()
|
||||
for index, case in enumerate(cases, start=1):
|
||||
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
|
||||
results.append(
|
||||
run_case(
|
||||
case,
|
||||
args.output,
|
||||
args.endpoint,
|
||||
args.model,
|
||||
args.timeout,
|
||||
args.num_ctx,
|
||||
args.num_predict,
|
||||
)
|
||||
)
|
||||
summary = {
|
||||
"experiment": "topic_reconstruction_v2",
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"model": args.model,
|
||||
"temperature": 0,
|
||||
"think": False,
|
||||
"case_count": len(cases),
|
||||
"llm_call_count": len(results),
|
||||
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||
"verdict_counts": {
|
||||
verdict: sum(item["verdict"] == verdict for item in results)
|
||||
for verdict in ("PASS", "PARTIAL", "FAIL")
|
||||
},
|
||||
"results": results,
|
||||
}
|
||||
(args.output / "summary.json").write_text(
|
||||
json.dumps(summary, ensure_ascii=False, indent=2) + "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
return summary
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
try:
|
||||
summary = run_experiment(args)
|
||||
except (OSError, ValueError, requests.RequestException) as exc:
|
||||
print(f"Error: {exc}")
|
||||
return 1
|
||||
print(json.dumps(summary["verdict_counts"], sort_keys=True))
|
||||
print(f"Artifacts: {args.output.resolve()}")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,5 @@
|
||||
"""Audio transcription support for the direct-protocol MVP."""
|
||||
|
||||
from .whisper import TranscriptionError, TranscriptionResult, transcribe_audio
|
||||
|
||||
__all__ = ["TranscriptionError", "TranscriptionResult", "transcribe_audio"]
|
||||
@@ -0,0 +1,290 @@
|
||||
"""Isolated whisper.cpp wrapper producing Meeting Lab compact transcripts."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import platform
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import tempfile
|
||||
import time
|
||||
from dataclasses import dataclass
|
||||
from datetime import datetime, timezone
|
||||
from pathlib import Path
|
||||
from typing import Any, Callable, Sequence
|
||||
|
||||
|
||||
BACKEND = "whisper.cpp"
|
||||
RAW_FILENAME = "whisper_raw.json"
|
||||
TRANSCRIPT_FILENAME = "transcript.json"
|
||||
TEXT_FILENAME = "transcript.txt"
|
||||
METADATA_FILENAME = "runtime_metadata.json"
|
||||
DEFAULT_THREADS = "auto"
|
||||
|
||||
|
||||
class TranscriptionError(RuntimeError):
|
||||
"""Raised when parameters, Whisper execution, or output are invalid."""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class TranscriptionResult:
|
||||
output_dir: Path
|
||||
raw_output: Path
|
||||
transcript_json: Path
|
||||
transcript_text: Path
|
||||
runtime_metadata: Path
|
||||
runtime_seconds: float
|
||||
|
||||
|
||||
def _logical_cpu_count() -> int:
|
||||
"""Return the available logical CPU count as a last-resort fallback."""
|
||||
if hasattr(os, "sched_getaffinity"):
|
||||
try:
|
||||
count = len(os.sched_getaffinity(0))
|
||||
if count > 0:
|
||||
return count
|
||||
except OSError:
|
||||
pass
|
||||
return os.cpu_count() or 1
|
||||
|
||||
|
||||
def _linux_physical_core_count() -> int | None:
|
||||
affinity = None
|
||||
if hasattr(os, "sched_getaffinity"):
|
||||
try:
|
||||
affinity = os.sched_getaffinity(0)
|
||||
except OSError:
|
||||
pass
|
||||
cores: set[tuple[str, str]] = set()
|
||||
for cpu_dir in Path("/sys/devices/system/cpu").glob("cpu[0-9]*"):
|
||||
try:
|
||||
cpu_number = int(cpu_dir.name[3:])
|
||||
if affinity is not None and cpu_number not in affinity:
|
||||
continue
|
||||
topology = cpu_dir / "topology"
|
||||
package = (topology / "physical_package_id").read_text().strip()
|
||||
core = (topology / "core_id").read_text().strip()
|
||||
cores.add((package, core))
|
||||
except (OSError, ValueError):
|
||||
continue
|
||||
return len(cores) or None
|
||||
|
||||
|
||||
def _darwin_physical_core_count() -> int | None:
|
||||
try:
|
||||
completed = subprocess.run(
|
||||
("sysctl", "-n", "hw.physicalcpu"),
|
||||
check=False,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
)
|
||||
count = int(completed.stdout.strip())
|
||||
return count if completed.returncode == 0 and count > 0 else None
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
def physical_core_count() -> int:
|
||||
"""Detect physical cores where supported, falling back to available threads."""
|
||||
system = platform.system()
|
||||
detected = _linux_physical_core_count() if system == "Linux" else None
|
||||
if system == "Darwin":
|
||||
detected = _darwin_physical_core_count()
|
||||
return detected or _logical_cpu_count()
|
||||
|
||||
|
||||
def resolve_threads(
|
||||
threads: str | int,
|
||||
detector: Callable[[], int] = physical_core_count,
|
||||
) -> int:
|
||||
if isinstance(threads, bool):
|
||||
raise TranscriptionError("Threads must be 'auto' or a positive integer.")
|
||||
if threads == "auto":
|
||||
count = detector()
|
||||
else:
|
||||
try:
|
||||
count = int(threads)
|
||||
except (TypeError, ValueError) as exc:
|
||||
raise TranscriptionError("Threads must be 'auto' or a positive integer.") from exc
|
||||
if count <= 0:
|
||||
raise TranscriptionError("Threads must be 'auto' or a positive integer.")
|
||||
return count
|
||||
|
||||
|
||||
def _vulkan_support(raw: dict[str, Any]) -> bool | None:
|
||||
system_info = raw.get("systeminfo")
|
||||
if not isinstance(system_info, str) or "VULKAN" not in system_info.upper():
|
||||
return None
|
||||
return re.search(r"VULKAN\s*=\s*1", system_info, re.IGNORECASE) is not None
|
||||
|
||||
|
||||
def _json_object(path: Path) -> dict[str, Any]:
|
||||
try:
|
||||
data = json.loads(path.read_text(encoding="utf-8"))
|
||||
except (OSError, UnicodeError, json.JSONDecodeError) as exc:
|
||||
raise TranscriptionError(f"Cannot read Whisper JSON output {path}: {exc}") from exc
|
||||
if not isinstance(data, dict):
|
||||
raise TranscriptionError("Whisper JSON output must contain a top-level object.")
|
||||
return data
|
||||
|
||||
|
||||
def compact_transcript(raw: dict[str, Any]) -> dict[str, Any]:
|
||||
"""Convert whisper.cpp JSON without linguistic cleanup or reordering."""
|
||||
entries = raw.get("transcription")
|
||||
if not isinstance(entries, list):
|
||||
raise TranscriptionError("Whisper JSON output must contain a 'transcription' list.")
|
||||
|
||||
segments: list[dict[str, Any]] = []
|
||||
for index, entry in enumerate(entries):
|
||||
if not isinstance(entry, dict):
|
||||
raise TranscriptionError(f"transcription[{index}] must be an object.")
|
||||
offsets = entry.get("offsets")
|
||||
if not isinstance(offsets, dict):
|
||||
raise TranscriptionError(f"transcription[{index}].offsets must be an object.")
|
||||
start_ms = offsets.get("from")
|
||||
end_ms = offsets.get("to")
|
||||
if not isinstance(start_ms, (int, float)) or isinstance(start_ms, bool):
|
||||
raise TranscriptionError(f"transcription[{index}].offsets.from must be a number.")
|
||||
if not isinstance(end_ms, (int, float)) or isinstance(end_ms, bool):
|
||||
raise TranscriptionError(f"transcription[{index}].offsets.to must be a number.")
|
||||
if end_ms < start_ms:
|
||||
raise TranscriptionError(
|
||||
f"transcription[{index}].offsets.to must be greater than or equal to offsets.from."
|
||||
)
|
||||
text_value = entry.get("text", "")
|
||||
if not isinstance(text_value, str):
|
||||
raise TranscriptionError(f"transcription[{index}].text must be a string.")
|
||||
text = text_value.strip()
|
||||
if text:
|
||||
segments.append(
|
||||
{
|
||||
"id": len(segments),
|
||||
"start": float(start_ms) / 1000.0,
|
||||
"end": float(end_ms) / 1000.0,
|
||||
"text": text,
|
||||
}
|
||||
)
|
||||
return {"text": " ".join(item["text"] for item in segments), "segments": segments}
|
||||
|
||||
|
||||
def _timestamp(seconds: float) -> str:
|
||||
milliseconds = int(round(seconds * 1000))
|
||||
hours, remainder = divmod(milliseconds, 3_600_000)
|
||||
minutes, remainder = divmod(remainder, 60_000)
|
||||
secs, millis = divmod(remainder, 1000)
|
||||
return f"{hours:02d}:{minutes:02d}:{secs:02d}.{millis:03d}"
|
||||
|
||||
|
||||
def transcript_text(transcript: dict[str, Any]) -> str:
|
||||
lines = [
|
||||
f"[{_timestamp(item['start'])} - {_timestamp(item['end'])}] {item['text']}"
|
||||
for item in transcript["segments"]
|
||||
]
|
||||
return "\n".join(lines) + ("\n" if lines else "")
|
||||
|
||||
|
||||
def _write_json(path: Path, value: dict[str, Any]) -> None:
|
||||
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def transcribe_audio(
|
||||
audio_path: Path,
|
||||
model_path: Path,
|
||||
output_dir: Path,
|
||||
language: str = "auto",
|
||||
*,
|
||||
executable: str = "whisper-cli",
|
||||
threads: str | int = DEFAULT_THREADS,
|
||||
thread_detector: Callable[[], int] = physical_core_count,
|
||||
runner: Callable[..., subprocess.CompletedProcess[str]] = subprocess.run,
|
||||
monotonic: Callable[[], float] = time.monotonic,
|
||||
now: Callable[[], datetime] = lambda: datetime.now(timezone.utc),
|
||||
) -> TranscriptionResult:
|
||||
"""Run one whisper.cpp call and write raw, compact, text, and metadata outputs."""
|
||||
audio_path = Path(audio_path)
|
||||
model_path = Path(model_path)
|
||||
output_dir = Path(output_dir)
|
||||
if not audio_path.is_file():
|
||||
raise TranscriptionError(f"Audio file does not exist: {audio_path}")
|
||||
if not model_path.is_file():
|
||||
raise TranscriptionError(f"Whisper model does not exist: {model_path}")
|
||||
if not isinstance(language, str) or not language.strip():
|
||||
raise TranscriptionError("Language must be a non-empty string.")
|
||||
if not executable.strip():
|
||||
raise TranscriptionError("Whisper executable must be a non-empty string.")
|
||||
if output_dir.exists() and not output_dir.is_dir():
|
||||
raise TranscriptionError(f"Output directory path is not a directory: {output_dir}")
|
||||
thread_count = resolve_threads(threads, thread_detector)
|
||||
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
raw_output = output_dir / RAW_FILENAME
|
||||
started_at = now().astimezone(timezone.utc)
|
||||
started = monotonic()
|
||||
|
||||
with tempfile.TemporaryDirectory(prefix=".whisper-", dir=output_dir) as temp_name:
|
||||
temporary_prefix = Path(temp_name) / "whisper_raw"
|
||||
command: Sequence[str] = (
|
||||
executable,
|
||||
"-m", str(model_path),
|
||||
"-f", str(audio_path),
|
||||
"-l", language.strip(),
|
||||
"-t", str(thread_count),
|
||||
"-fa",
|
||||
"-oj",
|
||||
"-of", str(temporary_prefix),
|
||||
)
|
||||
try:
|
||||
completed = runner(command, check=False, capture_output=True, text=True)
|
||||
except OSError as exc:
|
||||
raise TranscriptionError(f"Could not start {BACKEND}: {exc}") from exc
|
||||
runtime_seconds = monotonic() - started
|
||||
temporary_raw = temporary_prefix.with_suffix(".json")
|
||||
if temporary_raw.is_file():
|
||||
shutil.copyfile(temporary_raw, raw_output)
|
||||
if completed.returncode != 0:
|
||||
detail = completed.stderr.strip() or completed.stdout.strip() or "no diagnostic output"
|
||||
raise TranscriptionError(
|
||||
f"{BACKEND} failed with exit code {completed.returncode}: {detail}"
|
||||
)
|
||||
if not raw_output.is_file():
|
||||
raise TranscriptionError(f"{BACKEND} completed without producing JSON output.")
|
||||
|
||||
raw_data = _json_object(raw_output)
|
||||
transcript = compact_transcript(raw_data)
|
||||
transcript_json_path = output_dir / TRANSCRIPT_FILENAME
|
||||
transcript_text_path = output_dir / TEXT_FILENAME
|
||||
metadata_path = output_dir / METADATA_FILENAME
|
||||
_write_json(transcript_json_path, transcript)
|
||||
transcript_text_path.write_text(transcript_text(transcript), encoding="utf-8")
|
||||
duration = max((item["end"] for item in transcript["segments"]), default=None)
|
||||
metadata = {
|
||||
"input_file": str(audio_path.resolve()),
|
||||
"model": str(model_path.resolve()),
|
||||
"backend": BACKEND,
|
||||
"whisper_executable": executable,
|
||||
"language": language.strip(),
|
||||
"threads": thread_count,
|
||||
"threads_option": str(threads),
|
||||
"flash_attention": True,
|
||||
"vulkan_support_detected": _vulkan_support(raw_data),
|
||||
"duration_seconds": duration,
|
||||
"runtime_seconds": runtime_seconds,
|
||||
"timestamp": started_at.isoformat(),
|
||||
"output_files": {
|
||||
"whisper_raw": RAW_FILENAME,
|
||||
"transcript_json": TRANSCRIPT_FILENAME,
|
||||
"transcript_text": TEXT_FILENAME,
|
||||
"runtime_metadata": METADATA_FILENAME,
|
||||
},
|
||||
}
|
||||
_write_json(metadata_path, metadata)
|
||||
return TranscriptionResult(
|
||||
output_dir=output_dir,
|
||||
raw_output=raw_output,
|
||||
transcript_json=transcript_json_path,
|
||||
transcript_text=transcript_text_path,
|
||||
runtime_metadata=metadata_path,
|
||||
runtime_seconds=runtime_seconds,
|
||||
)
|
||||
@@ -0,0 +1,95 @@
|
||||
{
|
||||
"schema_version": "experimental-collective-commitment-gold-v0",
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "CC-01",
|
||||
"description": "Explicit collective commitment",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": true, "due": "nächste Woche"}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-02",
|
||||
"description": "Individual commitment",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, ich teste nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "individual_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": false, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-03",
|
||||
"description": "Tentative collective possibility",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten nächste Woche 20 Meter testen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": false, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-04",
|
||||
"description": "Collective suggestion",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Vielleicht sollten wir nächste Woche 20 Meter testen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": false, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-05",
|
||||
"description": "Impersonal necessity",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Man müsste nächste Woche 20 Meter testen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": false, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-06",
|
||||
"description": "Passive future statement",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Nächste Woche werden 20 Meter getestet.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": false, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-07",
|
||||
"description": "Collective rejection",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Nein, das testen wir nächste Woche nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "none", "action_concepts": [], "qualifier_concepts": []},
|
||||
"expected_result": {"established": false, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-08",
|
||||
"description": "Collective commitment with qualifier",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen 20 Meter, aber nur im Technikum.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": [["nur", "only"], ["technikum", "technical facility", "technical center", "technical centre"]]},
|
||||
"expected_result": {"established": true, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-09",
|
||||
"description": "Collective commitment without deadline",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": true, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "CC-10",
|
||||
"description": "Speaker ownership trap",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
|
||||
"expected_result": {"established": true, "due": "nächste Woche"}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"schema_version": "experimental-controlled-rejection-v1",
|
||||
"cases": [
|
||||
{"case_id":"CR-01","description":"self-contained non-pursuit","negative_act_source":"NA-01","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_1","target_observation_id":"obs_1","action_concepts":[["Schlummer"],["Zusammenarbeit","arbeiten"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":true}},
|
||||
{"case_id":"CR-02","description":"paired non-pursuit","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":true}},
|
||||
{"case_id":"CR-03","description":"personal preference","negative_act_source":"NA-03","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage für den Versuch nutzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde das nicht machen.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"personal_preference","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["reale Anlage"],["Versuch"],["nutzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||
{"case_id":"CR-04","description":"recommendation","negative_act_source":"NA-04","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage verwenden.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde eher davon abraten.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"recommendation","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["reale Anlage"],["verwenden","nutzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||
{"case_id":"CR-05","description":"temporary non-action","negative_act_source":"NA-05","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die Waschstufe einbauen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir erstmal noch nicht.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"temporary_non_action","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Waschstufe"],["einbauen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||
{"case_id":"CR-06","description":"concern","negative_act_source":"NA-06","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten das neue Material einsetzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das wäre kritisch.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"none","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["neue Material","neues Material"],["einsetzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||
{"case_id":"CR-07","description":"scoped explicit rejection","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[],"explicitly_rejected":true}},
|
||||
{"case_id":"CR-08","description":"rejection plus alternative","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"],"explicitly_rejected":true}}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,147 @@
|
||||
{
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "a_idea_only",
|
||||
"description": "Possible geometry optimization without commitment.",
|
||||
"subject_id": "subject_a",
|
||||
"subject": "Optimierung der Geometrie",
|
||||
"evidence": [{"evidence_id": "e1", "text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Geometrie kann vielleicht optimiert werden.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"positive","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"absent"},
|
||||
{"observation_id":"obs_2","evidence_id":"e1","content":"Danach könnte betrachtet werden, was herauskommt.","target":"obs_1","relation":"qualifies","modality":"suggested","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"implicit","scope":"nach der Optimierung|danach"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "b_multiple_options",
|
||||
"description": "Two alternatives for insufficient grid strength.",
|
||||
"subject_id": "subject_b",
|
||||
"subject": "Umgang mit unzureichender Festigkeit des 40-40-Gitters",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},
|
||||
{"evidence_id":"e2","text":"Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},
|
||||
{"evidence_id":"e3","text":"Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Festigkeit reicht noch nicht aus.","target":"discussion_subject","relation":"none","modality":"factual","temporality":"existing","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"40-40-Gitter|40-40"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Mehr Masse könnte für die gleiche Festigkeit eingesetzt werden.","target":"obs_1","relation":"qualifies","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"mehr Masse|gleiche Festigkeit"},
|
||||
{"observation_id":"obs_3","evidence_id":"e3","content":"Das Produkt könnte als 20-20 statt 40-40 ausgeführt werden.","target":"obs_1","relation":"qualifies","modality":"suggested","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"20-20|statt 40-40"},
|
||||
{"observation_id":"obs_4","evidence_id":"e3","content":"Die vorherigen Möglichkeiten sind die zwei Ansätze.","target":["obs_2","obs_3"],"relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"absent"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "c_unaccepted_proposal",
|
||||
"description": "Suggested Textor contact without established work.",
|
||||
"subject_id": "subject_c",
|
||||
"subject": "Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},
|
||||
{"evidence_id":"e2","text":"Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Tim erwägt, Dirk Textor erneut zu kontaktieren und nach seiner Einschätzung zu fragen.","target":"discussion_subject","relation":"none","modality":"suggested","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"Dirk Textors Einschätzung|erneut kontaktieren"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Eine erneute Rückkopplung mit Dirk Textor ist möglich.","target":"obs_1","relation":"supports","modality":"possible","temporality":"future","evaluation":"positive","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Rückkopplung mit Dirk Textor|mit ihm"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "d_proposal_with_objection",
|
||||
"description": "Washing possibility and explicit energy disadvantage.",
|
||||
"subject_id": "subject_d",
|
||||
"subject": "Waschen des Materials vor der weiteren Verarbeitung",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},
|
||||
{"evidence_id":"e2","text":"Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Material könnte gewaschen werden.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"vor der weiteren Verarbeitung"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin weiß nicht, ob sich das Waschen lohnt.","target":"obs_1","relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"Nutzen des Waschens|ob es sich lohnt"},
|
||||
{"observation_id":"obs_3","evidence_id":"e2","content":"Waschen umfasst Nassmachen und erneutes Trocknen.","target":"obs_1","relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Waschprozess|Nassmachen und Trocknen"},
|
||||
{"observation_id":"obs_4","evidence_id":"e2","content":"Waschen und Trocknen verursachen einen sehr hohen Energieaufwand.","target":"obs_1","relation":"opposes","modality":"factual","temporality":"existing","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Energieaufwand des Waschens|Waschen und Trocknen"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "e_rejected_alternative",
|
||||
"description": "Explicit rejection followed by confirmation of that rejection.",
|
||||
"subject_id": "subject_e",
|
||||
"subject": "Zusammenarbeit mit Dr. Schlummer für Versuche",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},
|
||||
{"evidence_id":"e2","text":"Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},
|
||||
{"evidence_id":"e3","text":"Antonius: Ja, das ist entschieden."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Angebot kostet 30.000 Euro.","target":"discussion_subject","relation":"none","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Angebot für die Versuche|30.000 Euro"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Die Zusammenarbeit mit Dr. Schlummer wird nicht durchgeführt.","target":"discussion_subject","relation":"none","modality":"committed","temporality":"future","evaluation":"none","agreement":"rejected","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Zusammenarbeit für die Versuche|Dr. Schlummer"},
|
||||
{"observation_id":"obs_3","evidence_id":"e3","content":"Die vorherige Ablehnung ist entschieden.","target":"obs_2","relation":"supports","modality":"factual","temporality":"completed","evaluation":"none","agreement":"accepted","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"absent"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "f_trial_only_acceptance",
|
||||
"description": "Acceptance limited to a 20-metre trial.",
|
||||
"subject_id": "subject_f",
|
||||
"subject": "20-Prozent-Variante im Versuch am kleinen Extruder",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},
|
||||
{"evidence_id":"e2","text":"Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},
|
||||
{"evidence_id":"e3","text":"Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Die 20-Prozent-Variante könnte am kleinen Extruder nachgestellt werden.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"kleiner Extruder"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"20 Meter der Variante werden beim nächsten Versuch getestet.","target":"obs_1","relation":"supports","modality":"committed","temporality":"future","evaluation":"none","agreement":"accepted","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"20 Meter beim nächsten Versuch|20 Meter"},
|
||||
{"observation_id":"obs_3","evidence_id":"e3","content":"Die Zusage gilt nur für einen Versuch.","target":"obs_2","relation":"limits_scope","modality":"factual","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"nur ein Versuch|Versuch"},
|
||||
{"observation_id":"obs_4","evidence_id":"e3","content":"Die Variante ist noch nicht als Serienlösung festgelegt.","target":"discussion_subject","relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"implicit","scope":"Serienlösung|finale Produktion"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "g_no_decision",
|
||||
"description": "Preference, alternative, and impersonal checking need without decision.",
|
||||
"subject_id": "subject_g",
|
||||
"subject": "Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},
|
||||
{"evidence_id":"e2","text":"Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},
|
||||
{"evidence_id":"e3","text":"Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Eine reale Recyclinganlage birgt das Risiko kontaminierten Rückmaterials.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"reale Recyclinganlage|kontaminiertes Material"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin würde nicht in eine reale Anlage gehen.","target":"discussion_subject","relation":"opposes","modality":"suggested","temporality":"future","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Martins persönliche Präferenz|reale Anlage"},
|
||||
{"observation_id":"obs_3","evidence_id":"e2","content":"Ein Technikum bleibt als bedingte Möglichkeit im Gespräch.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"wenn überhaupt|Technikum"},
|
||||
{"observation_id":"obs_4","evidence_id":"e3","content":"Zunächst muss geprüft werden, welcher Reinigungsansatz verfügbar ist.","target":"discussion_subject","relation":"qualifies","modality":"impersonal_necessity","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"zunächst|verfügbarer Reinigungsansatz"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "h_resulting_action",
|
||||
"description": "Interpersonal request followed by accepted responsibility.",
|
||||
"subject_id": "subject_h",
|
||||
"subject": "Prüfung der Messdaten bis Freitag",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},
|
||||
{"evidence_id":"e2","text":"Nina: Ja, ich übernehme die Prüfung bis Freitag."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Antonius bittet Nina um die Prüfung der Messdaten.","target":"discussion_subject","relation":"none","modality":"interpersonal_request","temporality":"future","evaluation":"none","agreement":"none","responsibility":"named","person":"Nina","uncertainty":"absent","clarification_need":"none","scope":"bis Freitag|Freitag"},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Nina übernimmt die Prüfung.","target":"obs_1","relation":"supports","modality":"committed","temporality":"future","evaluation":"none","agreement":"accepted","responsibility":"accepted","person":"Nina","uncertainty":"absent","clarification_need":"none","scope":"bis Freitag|Freitag"}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "i_outcome_and_unresolved",
|
||||
"description": "Bounded production finding and unresolved publication information.",
|
||||
"subject_id": "subject_i",
|
||||
"subject": "Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
|
||||
"evidence": [
|
||||
{"evidence_id":"e1","text":"Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},
|
||||
{"evidence_id":"e2","text":"Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},
|
||||
{"evidence_id":"e3","text":"Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},
|
||||
{"evidence_id":"e4","text":"Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}
|
||||
],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Bei der reinen Produktion gab es praktisch keine Änderung gegenüber dem Standardprodukt.","target":"discussion_subject","relation":"none","modality":"factual","temporality":"completed","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"eigene Anlage, reine Produktion, Standardprodukt|reine Produktion"},
|
||||
{"observation_id":"obs_2","evidence_id":"e1","content":"Die Produktion erfolgte fünf Grad kälter.","target":"obs_1","relation":"qualifies","modality":"factual","temporality":"completed","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"fünf Grad kälter|5 Grad"},
|
||||
{"observation_id":"obs_3","evidence_id":"e2","content":"Gegenüber Virgin Material ist bei reiner Produktion kein zusätzlicher Aufwand notwendig.","target":"obs_1","relation":"supports","modality":"factual","temporality":"existing","evaluation":"none","agreement":"accepted","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"reine Produktion gegenüber Virgin Material|Virgin Material"},
|
||||
{"observation_id":"obs_4","evidence_id":"e2","content":"Vor der reinen Produktion entsteht Aufwand.","target":"obs_3","relation":"limits_scope","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"vor der reinen Produktion|davor"},
|
||||
{"observation_id":"obs_5","evidence_id":"e3","content":"Es wird gefragt, welche Energieaudit-Daten veröffentlicht werden dürfen.","target":"discussion_subject","relation":"none","modality":"information_question","temporality":"unspecified","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"Veröffentlichung von Energieaudit-Daten|Energieaudit"},
|
||||
{"observation_id":"obs_6","evidence_id":"e4","content":"Die Veröffentlichungserlaubnis ist weiterhin ungeklärt.","target":"obs_5","relation":"supports","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"Veröffentlichungserlaubnis|Freigabe"},
|
||||
{"observation_id":"obs_7","evidence_id":"e4","content":"Die Freigabe muss noch geklärt werden.","target":"obs_5","relation":"supports","modality":"impersonal_necessity","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"Freigabe zur Veröffentlichung|Freigabe"}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,99 @@
|
||||
{
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "a_idea_only", "description": "Possible geometry optimization without commitment.",
|
||||
"subject_id": "subject_a", "subject": "Optimierung der Geometrie",
|
||||
"evidence": [{"evidence_id": "e1", "text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Geometrie kann vielleicht optimiert werden.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"positive","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":null,"limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e1","content":"Danach könnte betrachtet werden, was herauskommt.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"suggested","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"implicit","qualifier":"danach","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "b_multiple_options", "description": "Two alternatives for insufficient grid strength.",
|
||||
"subject_id": "subject_b", "subject": "Umgang mit unzureichender Festigkeit des 40-40-Gitters",
|
||||
"evidence": [{"evidence_id":"e1","text":"Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},{"evidence_id":"e2","text":"Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},{"evidence_id":"e3","text":"Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Festigkeit des 40-40-Gitters reicht noch nicht aus.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"negative","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"40-40-Gitter","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Mehr Masse könnte für die gleiche Festigkeit eingesetzt werden.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"mehr Masse für die gleiche Festigkeit","limits_target":null},
|
||||
{"observation_id":"obs_3","evidence_id":"e3","content":"Das Produkt könnte als 20-20 statt 40-40 ausgeführt werden.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"suggested","temporality":"future","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"20-20 statt 40-40","limits_target":null},
|
||||
{"observation_id":"obs_4","evidence_id":"e3","content":"Die vorherigen Möglichkeiten sind die zwei Ansätze.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"zwei Ansätze","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "c_unaccepted_proposal", "description": "Suggested Textor contact without established work.",
|
||||
"subject_id": "subject_c", "subject": "Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
|
||||
"evidence": [{"evidence_id":"e1","text":"Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},{"evidence_id":"e2","text":"Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Tim erwägt, Dirk Textor erneut zu kontaktieren und nach seiner Einschätzung zu fragen.","refers_to":null,"speaker":"Tim","named_person":"Dirk Textor","addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"suggested","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"erneut; Dirk Textors Einschätzung","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Eine erneute Rückkopplung mit Dirk Textor ist möglich.","refers_to":"obs_1","speaker":"Tim","named_person":"Dirk Textor","addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"positive","affirmation":"explicit","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"noch einmal mit ihm","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "d_proposal_with_objection", "description": "Washing possibility and explicit energy disadvantage.",
|
||||
"subject_id": "subject_d", "subject": "Waschen des Materials vor der weiteren Verarbeitung",
|
||||
"evidence": [{"evidence_id":"e1","text":"Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},{"evidence_id":"e2","text":"Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Material könnte gewaschen werden.","refers_to":null,"speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"vor der weiteren Verarbeitung","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin weiß nicht, ob sich das Waschen lohnt.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"ob es sich lohnt","limits_target":null},
|
||||
{"observation_id":"obs_3","evidence_id":"e2","content":"Waschen umfasst Nassmachen und erneutes Trocknen.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"nass machen und wieder trocknen","limits_target":null},
|
||||
{"observation_id":"obs_4","evidence_id":"e2","content":"Waschen und Trocknen verursachen einen sehr hohen Energieaufwand.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"negative","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"Waschen und Trocknen","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "e_rejected_alternative", "description": "Explicit negation followed by confirmation of that determination.",
|
||||
"subject_id": "subject_e", "subject": "Zusammenarbeit mit Dr. Schlummer für Versuche",
|
||||
"evidence": [{"evidence_id":"e1","text":"Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},{"evidence_id":"e2","text":"Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},{"evidence_id":"e3","text":"Antonius: Ja, das ist entschieden."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Angebot von Dr. Schlummer kostet 30.000 Euro.","refers_to":null,"speaker":"Antonius","named_person":"Dr. Schlummer","addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"für die Versuche; 30.000 Euro","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Die Zusammenarbeit mit Dr. Schlummer wird nicht durchgeführt.","refers_to":null,"speaker":"Tim","named_person":"Dr. Schlummer","addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"committed","temporality":"future","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"present","uncertainty":"absent","clarification_need":"none","qualifier":"Zusammenarbeit für die Versuche","limits_target":null},
|
||||
{"observation_id":"obs_3","evidence_id":"e3","content":"Die vorherige Festlegung ist entschieden.","refers_to":"obs_2","speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"completed","evaluation":"none","affirmation":"explicit","negation":"absent","determination_statement":"present","uncertainty":"absent","clarification_need":"none","qualifier":null,"limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "f_trial_only_acceptance", "description": "Affirmed commitment limited to a 20-metre trial.",
|
||||
"subject_id": "subject_f", "subject": "20-Prozent-Variante im Versuch am kleinen Extruder",
|
||||
"evidence": [{"evidence_id":"e1","text":"Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},{"evidence_id":"e2","text":"Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},{"evidence_id":"e3","text":"Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Die 20-Prozent-Variante könnte am kleinen Extruder nachgestellt werden.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"am kleinen Extruder","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"20 Meter der Variante werden beim nächsten Versuch getestet.","refers_to":"obs_1","speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"committed","temporality":"future","evaluation":"none","affirmation":"explicit","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"20 Meter beim nächsten Versuch","limits_target":null},
|
||||
{"observation_id":"obs_3","evidence_id":"e3","content":"Die Zusage gilt nur für einen Versuch.","refers_to":"obs_2","speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"nur ein Versuch","limits_target":"obs_2"},
|
||||
{"observation_id":"obs_4","evidence_id":"e3","content":"Die Variante ist noch nicht als Serienlösung festgelegt.","refers_to":"obs_2","speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"present","uncertainty":"present","clarification_need":"implicit","qualifier":"als Serienlösung","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "g_no_decision", "description": "Preference, alternative, and impersonal checking need without decision.",
|
||||
"subject_id": "subject_g", "subject": "Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
|
||||
"evidence": [{"evidence_id":"e1","text":"Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},{"evidence_id":"e2","text":"Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},{"evidence_id":"e3","text":"Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Eine reale Recyclinganlage birgt das Risiko kontaminierten Rückmaterials.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"possible","temporality":"future","evaluation":"negative","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"reale Recyclinganlage; kontaminiertes Material","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin würde persönlich nicht in eine reale Anlage gehen.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"suggested","temporality":"future","evaluation":"negative","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"reale Anlage","limits_target":null},
|
||||
{"observation_id":"obs_3","evidence_id":"e2","content":"Ein Technikum bleibt als bedingte Möglichkeit im Gespräch.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"wenn überhaupt; Technikum","limits_target":null},
|
||||
{"observation_id":"obs_4","evidence_id":"e3","content":"Zunächst muss geprüft werden, welcher Reinigungsansatz verfügbar ist.","refers_to":null,"speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"impersonal_necessity","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"zunächst; verfügbarer Reinigungsansatz","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "h_resulting_action", "description": "Interpersonal request followed by explicit personal acceptance.",
|
||||
"subject_id": "subject_h", "subject": "Prüfung der Messdaten bis Freitag",
|
||||
"evidence": [{"evidence_id":"e1","text":"Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},{"evidence_id":"e2","text":"Nina: Ja, ich übernehme die Prüfung bis Freitag."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Antonius richtet an Nina die Bitte, die Messdaten zu prüfen.","refers_to":null,"speaker":"Antonius","named_person":"Nina","addressee":"Nina","self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"interpersonal_request","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"bis Freitag","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e2","content":"Nina sagt zu, die Prüfung zu übernehmen.","refers_to":"obs_1","speaker":"Nina","named_person":null,"addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"committed","temporality":"future","evaluation":"none","affirmation":"explicit","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"bis Freitag","limits_target":null}
|
||||
]
|
||||
},
|
||||
{
|
||||
"case_id": "i_outcome_and_unresolved", "description": "Bounded production finding and unresolved publication information.",
|
||||
"subject_id": "subject_i", "subject": "Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
|
||||
"evidence": [{"evidence_id":"e1","text":"Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},{"evidence_id":"e2","text":"Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},{"evidence_id":"e3","text":"Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},{"evidence_id":"e4","text":"Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}],
|
||||
"expected_observations": [
|
||||
{"observation_id":"obs_1","evidence_id":"e1","content":"Bei der reinen Produktion gab es praktisch keine Änderung gegenüber dem Standardprodukt.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"completed","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"an unserer Anlage; reine Produktion; gegenüber dem Standardprodukt","limits_target":null},
|
||||
{"observation_id":"obs_2","evidence_id":"e1","content":"Die Produktion erfolgte fünf Grad kälter.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"factual","temporality":"completed","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"fünf Grad kälter","limits_target":null},
|
||||
{"observation_id":"obs_3","evidence_id":"e2","content":"Gegenüber Virgin Material ist bei reiner Produktion kein zusätzlicher Aufwand notwendig.","refers_to":"obs_1","speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"present","uncertainty":"absent","clarification_need":"none","qualifier":"bei reiner Produktion; gegenüber Virgin Material","limits_target":null},
|
||||
{"observation_id":"obs_4","evidence_id":"e2","content":"Vor der reinen Produktion entsteht Aufwand.","refers_to":"obs_3","speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"davor","limits_target":"obs_3"},
|
||||
{"observation_id":"obs_5","evidence_id":"e3","content":"Es wird gefragt, welche Energieaudit-Daten veröffentlicht werden dürfen.","refers_to":null,"speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"information_question","temporality":"unspecified","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"Veröffentlichung von Energieaudit-Daten","limits_target":null},
|
||||
{"observation_id":"obs_6","evidence_id":"e4","content":"Die Veröffentlichungserlaubnis ist weiterhin ungeklärt.","refers_to":"obs_5","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"weiterhin","limits_target":null},
|
||||
{"observation_id":"obs_7","evidence_id":"e4","content":"Die Freigabe muss noch geklärt werden.","refers_to":"obs_5","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"impersonal_necessity","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"noch; Freigabe zur Veröffentlichung","limits_target":null}
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,49 @@
|
||||
{
|
||||
"cases": [
|
||||
{
|
||||
"case_id":"a_idea_only","description":"Possible geometry optimization without commitment.","subject_id":"subject_a","subject":"Optimierung der Geometrie",
|
||||
"evidence":[{"evidence_id":"e1","text":"Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}],
|
||||
"semantic_requirements":["Geometry optimization remains possible and tentative.","Subsequent checking remains conditional and tentative.","The then/sequential dependency survives.","No commitment or owner is introduced."]
|
||||
},
|
||||
{
|
||||
"case_id":"b_multiple_options","description":"Two alternatives for insufficient grid strength.","subject_id":"subject_b","subject":"Umgang mit unzureichender Festigkeit des 40-40-Gitters",
|
||||
"evidence":[{"evidence_id":"e1","text":"Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},{"evidence_id":"e2","text":"Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},{"evidence_id":"e3","text":"Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}],
|
||||
"semantic_requirements":["Insufficient 40-40 strength survives.","Additional mass remains one alternative.","20-20 remains another alternative.","Both remain alternatives and neither is selected."]
|
||||
},
|
||||
{
|
||||
"case_id":"c_unaccepted_proposal","description":"Suggested Textor contact without established work.","subject_id":"subject_c","subject":"Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
|
||||
"evidence":[{"evidence_id":"e1","text":"Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},{"evidence_id":"e2","text":"Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}],
|
||||
"semantic_requirements":["Contacting Dirk Textor remains Tim's tentative personal suggestion.","The follow-up remains possible and relates to that contact.","No established work or responsibility is introduced."]
|
||||
},
|
||||
{
|
||||
"case_id":"d_proposal_with_objection","description":"Washing possibility and explicit energy disadvantage.","subject_id":"subject_d","subject":"Waschen des Materials vor der weiteren Verarbeitung",
|
||||
"evidence":[{"evidence_id":"e1","text":"Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},{"evidence_id":"e2","text":"Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}],
|
||||
"semantic_requirements":["Washing before further processing remains possible.","Martin's uncertainty whether washing is worthwhile survives.","The washing and drying process survives.","The high energy consequence survives.","No unresolved task is invented."]
|
||||
},
|
||||
{
|
||||
"case_id":"e_rejected_alternative","description":"Explicit rejection followed by confirmation of that determination.","subject_id":"subject_e","subject":"Zusammenarbeit mit Dr. Schlummer für Versuche",
|
||||
"evidence":[{"evidence_id":"e1","text":"Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},{"evidence_id":"e2","text":"Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},{"evidence_id":"e3","text":"Antonius: Ja, das ist entschieden."}],
|
||||
"semantic_requirements":["The offer cost survives without inferred evaluation.","Collaboration is explicitly not to be pursued.","The explicit rejection survives.","The later statement confirms that the preceding determination has been made."]
|
||||
},
|
||||
{
|
||||
"case_id":"f_trial_only_acceptance","description":"Collective commitment limited to a 20-metre trial.","subject_id":"subject_f","subject":"20-Prozent-Variante im Versuch am kleinen Extruder",
|
||||
"evidence":[{"evidence_id":"e1","text":"Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},{"evidence_id":"e2","text":"Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},{"evidence_id":"e3","text":"Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}],
|
||||
"semantic_requirements":["The 20-percent variant at the small extruder remains initially possible.","The later statement collectively commits to a test.","Twenty metres and next-trial timing survive.","The test remains limited to a trial.","Series adoption remains explicitly not yet established.","No individual owner is invented."]
|
||||
},
|
||||
{
|
||||
"case_id":"g_no_decision","description":"Preference, conditional alternative, and impersonal checking need.","subject_id":"subject_g","subject":"Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
|
||||
"evidence":[{"evidence_id":"e1","text":"Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},{"evidence_id":"e2","text":"Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},{"evidence_id":"e3","text":"Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}],
|
||||
"semantic_requirements":["Contamination remains a risk rather than a fact.","Martin's negative stance remains personal.","The Technikum remains conditional and if-at-all survives.","Cleaning-method availability still needs to be checked.","The need remains impersonal.","No group decision or owner is invented."]
|
||||
},
|
||||
{
|
||||
"case_id":"h_resulting_action","description":"Interpersonal request followed by explicit personal acceptance.","subject_id":"subject_h","subject":"Prüfung der Messdaten bis Freitag",
|
||||
"evidence":[{"evidence_id":"e1","text":"Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},{"evidence_id":"e2","text":"Nina: Ja, ich übernehme die Prüfung bis Freitag."}],
|
||||
"semantic_requirements":["Antonius requests measurement-data review from Nina.","The Friday deadline survives.","Nina explicitly accepts the preceding request.","Nina's response expresses future personal commitment.","No responsibility field or unsupported inference is introduced."]
|
||||
},
|
||||
{
|
||||
"case_id":"i_outcome_and_unresolved","description":"Bounded production finding and unresolved publication information.","subject_id":"subject_i","subject":"Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
|
||||
"evidence":[{"evidence_id":"e1","text":"Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},{"evidence_id":"e2","text":"Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},{"evidence_id":"e3","text":"Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},{"evidence_id":"e4","text":"Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}],
|
||||
"semantic_requirements":["The pure-production finding remains bounded to the local plant and standard-product comparison.","The five-degree difference survives.","No-extra-effort remains bounded to pure production compared with Virgin material.","Upstream effort before that production stage survives.","The publication purpose of the energy-audit question remains explicit.","Publication permission remains unresolved and clarification remains necessary.","No assigned work is invented."]
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,111 @@
|
||||
{
|
||||
"schema_version": "experimental-explicit-rejection-gold-v0",
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "RJ-01", "description": "Explicit collective rejection with local target",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Nein, das machen wir nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": true}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-02", "description": "Explicit non-pursuit with paired target",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": true}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-03", "description": "Self-contained collaboration rejection",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_1", "target_observation_id": "obs_1", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": true}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-04", "description": "Personal preference",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-05", "description": "Concern",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-06", "description": "Uncertainty",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-07", "description": "Negative recommendation",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-08", "description": "Deferral",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die externe Lösung einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das entscheiden wir nächste Woche.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-09", "description": "Factual negation",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_1", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-10", "description": "Temporary non-action",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": false}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-11", "description": "Explicit rejection with material scope",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Die reale Anlage nutzen wir dafür nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"]], "qualifier_concepts": [["druckversuch", "dafür", "pressure test"]], "forbidden_action_concepts": []},
|
||||
"expected_result": {"explicitly_rejected": true}
|
||||
},
|
||||
{
|
||||
"case_id": "RJ-12", "description": "Rejection plus positive alternative",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten den Versuch in der realen Anlage durchführen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir nicht; wir testen stattdessen im Technikum.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["versuch", "trial", "test"], ["real"], ["anlage", "plant"]], "qualifier_concepts": [["real"], ["anlage", "plant"]], "forbidden_action_concepts": ["technikum", "technical facility", "technical center", "technical centre"]},
|
||||
"expected_result": {"explicitly_rejected": true}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
{
|
||||
"schema_version": "experimental-negative-act-form-gold-v0",
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "NA-01", "description": "Explicit non-pursuit",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_1", "negative_act_form": "explicit_non_pursuit", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]]}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-02", "description": "Explicit non-pursuit paraphrase",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Die externe Lösung verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_1", "negative_act_form": "explicit_non_pursuit", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]]}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-03", "description": "Personal preference with local context",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_2", "negative_act_form": "personal_preference", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]]}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-04", "description": "Negative recommendation",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_2", "negative_act_form": "recommendation", "action_concepts": [["real"], ["anlage", "plant"], ["verwend", "use"]]}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-05", "description": "Temporary non-action",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_2", "negative_act_form": "temporary_non_action", "action_concepts": [["waschstufe", "washing stage"], ["einbau", "install"]]}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-06", "description": "Concern only",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_2", "negative_act_form": "none", "action_concepts": []}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-07", "description": "Uncertainty",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_2", "negative_act_form": "none", "action_concepts": []}
|
||||
},
|
||||
{
|
||||
"case_id": "NA-08", "description": "Factual negation",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected": {"observation_id": "obs_1", "negative_act_form": "none", "action_concepts": []}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,101 @@
|
||||
{
|
||||
"schema_version": "experimental-request-acceptance-gold-v0",
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "RA-01",
|
||||
"description": "Explicit positive acceptance",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja, ich übernehme die Auswertung bis Dienstag.", "speaker": "Clara", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": true, "same_work": true},
|
||||
"expected_result": {"established": true, "content": "Auswertung der Messwerte", "requested_actor": "Clara", "responsible_person": "Clara", "due": "Dienstag"}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-02",
|
||||
"description": "Paraphrased positive acceptance",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja. Ich kümmere mich darum und habe die Auswertung bis Dienstag fertig.", "speaker": "Clara", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": true, "same_work": true},
|
||||
"expected_result": {"established": true, "content": "Auswertung der Messwerte", "requested_actor": "Clara", "responsible_person": "Clara", "due": "Dienstag"}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-03",
|
||||
"description": "Acknowledgement only",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja, ich habe verstanden, worum es geht.", "speaker": "Clara", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-04",
|
||||
"description": "Tentative response",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ich schaue mal, ob ich das schaffe.", "speaker": "Clara", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-05",
|
||||
"description": "Different responder without personal acceptance",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ja, das sollte gemacht werden.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-06",
|
||||
"description": "Explicit commitment to different work",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
|
||||
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja, ich kümmere mich um die Präsentation.", "speaker": "Clara", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": true, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-07",
|
||||
"description": "Request without response",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"}
|
||||
],
|
||||
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-08",
|
||||
"description": "Collective commitment",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": false, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-09",
|
||||
"description": "Impersonal necessity",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Man müsste zuerst prüfen, welches Reinigungsverfahren verfügbar ist.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": false, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
},
|
||||
{
|
||||
"case_id": "RA-10",
|
||||
"description": "Personal suggestion",
|
||||
"observations": [
|
||||
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ich würde vielleicht Dirk Textor kontaktieren.", "speaker": "Martin", "named_person": "Dirk Textor", "addressee": null}
|
||||
],
|
||||
"expected_recognition": {"request": false, "commitment": false, "same_work": false},
|
||||
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,11 @@
|
||||
# semantic_synthesis_isolation
|
||||
|
||||
Isolation Gold set derived from the existing Topic Reconstruction V2 A-I
|
||||
cases. Every case supplies one manually fixed Discussion Subject and the
|
||||
complete original evidence bundle. The model performs Semantic Synthesis only;
|
||||
subject detection, subject grouping, and evidence assignment are outside the
|
||||
experiment.
|
||||
|
||||
Expected criteria evaluate semantic event distinctions, outcomes and scope,
|
||||
actions, unresolved issues, and supporting evidence. They do not evaluate
|
||||
subject discovery or exact wording.
|
||||
@@ -0,0 +1,227 @@
|
||||
{
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "a_idea_only",
|
||||
"description": "Idea mentioned without stronger commitment.",
|
||||
"subject_id": "subject_a",
|
||||
"subject": "Optimierung der Geometrie",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"idea": 1},
|
||||
"allowed_event_types": ["idea"],
|
||||
"event_evidence_ids": ["e1"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {"count": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "b_multiple_options",
|
||||
"description": "Two alternatives for insufficient 40-40 grid strength.",
|
||||
"subject_id": "subject_b",
|
||||
"subject": "Umgang mit unzureichender Festigkeit des 40-40-Gitters",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},
|
||||
{"evidence_id": "e2", "text": "Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},
|
||||
{"evidence_id": "e3", "text": "Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"option": 2},
|
||||
"allowed_event_types": ["technical_finding", "fact", "option"],
|
||||
"event_evidence_ids": ["e1", "e2", "e3"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {"count": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "c_unaccepted_proposal",
|
||||
"description": "Possible Textor contact remains a proposal only.",
|
||||
"subject_id": "subject_c",
|
||||
"subject": "Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},
|
||||
{"evidence_id": "e2", "text": "Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"proposal": 1},
|
||||
"allowed_event_types": ["proposal"],
|
||||
"event_evidence_ids": ["e1", "e2"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {"count": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "d_proposal_with_objection",
|
||||
"description": "Washing proposal with energy objection but no unresolved issue.",
|
||||
"subject_id": "subject_d",
|
||||
"subject": "Waschen des Materials vor der weiteren Verarbeitung",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},
|
||||
{"evidence_id": "e2", "text": "Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"proposal": 1, "objection": 1},
|
||||
"allowed_event_types": ["proposal", "objection"],
|
||||
"event_evidence_ids": ["e1", "e2"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {"count": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "e_rejected_alternative",
|
||||
"description": "Explicit rejection of Schlummer collaboration.",
|
||||
"subject_id": "subject_e",
|
||||
"subject": "Zusammenarbeit mit Dr. Schlummer für Versuche",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},
|
||||
{"evidence_id": "e2", "text": "Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},
|
||||
{"evidence_id": "e3", "text": "Antonius: Ja, das ist entschieden."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"fact": 1, "rejection": 1},
|
||||
"allowed_event_types": ["fact", "rejection", "clarification"],
|
||||
"event_evidence_ids": ["e1", "e2", "e3"],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"statuses": ["rejected"],
|
||||
"terms": ["nicht", "abgelehnt", "keine"],
|
||||
"scope_terms": ["zusammenarbeit", "versuch", "schlummer"],
|
||||
"evidence_ids": ["e2", "e3"]
|
||||
},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {"count": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "f_trial_only_acceptance",
|
||||
"description": "Acceptance limited to a 20-metre trial.",
|
||||
"subject_id": "subject_f",
|
||||
"subject": "20-Prozent-Variante im Versuch am kleinen Extruder",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},
|
||||
{"evidence_id": "e2", "text": "Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},
|
||||
{"evidence_id": "e3", "text": "Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"proposal": 1, "scoped_acceptance": 1, "clarification": 1},
|
||||
"allowed_event_types": ["proposal", "scoped_acceptance", "clarification"],
|
||||
"event_evidence_ids": ["e1", "e2", "e3"],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"statuses": ["scoped_acceptance"],
|
||||
"terms": ["test", "versuch"],
|
||||
"scope_terms": ["20 meter", "20 m", "nur", "begrenzt"],
|
||||
"evidence_ids": ["e2", "e3"]
|
||||
},
|
||||
"actions": {
|
||||
"count": 1,
|
||||
"terms": ["test", "versuch", "20 meter"],
|
||||
"evidence_ids": ["e2"],
|
||||
"due_terms": ["nächsten versuch", "next trial"]
|
||||
},
|
||||
"unresolved_issues": {
|
||||
"count": 1,
|
||||
"terms": ["serienlösung", "final", "serie", "festgelegt"],
|
||||
"evidence_ids": ["e3"]
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "g_no_decision",
|
||||
"description": "Plant versus Technikum discussion ending without a decision.",
|
||||
"subject_id": "subject_g",
|
||||
"subject": "Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},
|
||||
{"evidence_id": "e2", "text": "Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},
|
||||
{"evidence_id": "e3", "text": "Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"objection": 1, "option": 1},
|
||||
"allowed_event_types": ["objection", "option", "proposal", "clarification"],
|
||||
"event_evidence_ids": ["e1", "e2", "e3"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {
|
||||
"count": 1,
|
||||
"terms": ["reinigungsansatz", "verfügbar", "prüfen", "reinigung"],
|
||||
"evidence_ids": ["e3"]
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "h_resulting_action",
|
||||
"description": "Explicitly accepted action with owner and deadline.",
|
||||
"subject_id": "subject_h",
|
||||
"subject": "Prüfung der Messdaten bis Freitag",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},
|
||||
{"evidence_id": "e2", "text": "Nina: Ja, ich übernehme die Prüfung bis Freitag."}
|
||||
],
|
||||
"allowed_responsible": ["Nina"],
|
||||
"expected": {
|
||||
"event_type_minimums": {},
|
||||
"allowed_event_types": ["proposal", "clarification", "scoped_acceptance"],
|
||||
"event_evidence_ids": [],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"statuses": ["established"],
|
||||
"terms": ["übernimmt", "prüf", "accepted", "review", "assigned"],
|
||||
"scope_terms": ["messdaten", "prüfung", "measurement", "review"],
|
||||
"evidence_ids": ["e2"]
|
||||
},
|
||||
"actions": {
|
||||
"count": 1,
|
||||
"terms": ["messdaten", "prüf", "measurement", "review"],
|
||||
"responsible": "Nina",
|
||||
"due_terms": ["freitag", "friday"],
|
||||
"evidence_ids": ["e2"]
|
||||
},
|
||||
"unresolved_issues": {"count": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "i_outcome_and_unresolved",
|
||||
"description": "Bounded production outcome and unresolved publication question.",
|
||||
"subject_id": "subject_i",
|
||||
"subject": "Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},
|
||||
{"evidence_id": "e2", "text": "Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},
|
||||
{"evidence_id": "e3", "text": "Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},
|
||||
{"evidence_id": "e4", "text": "Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}
|
||||
],
|
||||
"allowed_responsible": [],
|
||||
"expected": {
|
||||
"event_type_minimums": {"fact": 1},
|
||||
"allowed_event_types": ["technical_finding", "fact", "clarification"],
|
||||
"event_evidence_ids": ["e1", "e2"],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"statuses": ["established"],
|
||||
"terms": ["kein zusätzlicher", "keine zusätzliche", "unverändert"],
|
||||
"scope_terms": ["reine produktion", "produktion", "virgin"],
|
||||
"evidence_ids": ["e1", "e2"]
|
||||
},
|
||||
"actions": {"count": 0},
|
||||
"unresolved_issues": {
|
||||
"count": 1,
|
||||
"terms": ["veröffentlich", "freigabe", "energieaudit", "daten"],
|
||||
"evidence_ids": ["e3", "e4"]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"schema_version":"experimental-target-normalization-v0",
|
||||
"cases":[
|
||||
{"case_id":"TN-01","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_1","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"Zusammenarbeit mit Dr. Schlummer fortsetzen","action_concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":["nicht","beenden"],"german_markers":["Zusammenarbeit","arbeiten","fortsetzen"]}},
|
||||
{"case_id":"TN-02","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"externe Lösung weiterverfolgen","action_concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":["nicht"],"german_markers":["Lösung","verfolgen"]}},
|
||||
{"case_id":"TN-03","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"reale Anlage für den Druckversuch nutzen","action_concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":["nicht","zur Diskussion"],"german_markers":["Anlage","Druckversuch","nutzen"]}},
|
||||
{"case_id":"TN-04","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"Versuch in der realen Anlage durchführen","action_concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["nicht","Technikum","stattdessen"],"german_markers":["Versuch","Anlage","durchführen"]}}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"schema_version": "experimental-target-resolution-v0",
|
||||
"cases": [
|
||||
{"case_id":"TR-01","description":"self-contained continuation target","strategy":"self_contained","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"negative_act":{"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"working with Dr. Schlummer"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR-02","description":"paired pronoun target","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"verfolgen wir nicht weiter"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR-03","description":"scoped location and purpose target","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"reale Anlage dafür nicht nutzen"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR-04","description":"rejection plus alternative","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"Versuch in der realen Anlage nicht durchführen"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"]}},
|
||||
{"case_id":"TR-05","description":"personal preference","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage für den Versuch nutzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde das nicht machen.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"personal_preference","normalized_action_text":"Ich würde das nicht machen"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR-06","description":"recommendation","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage verwenden.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde eher davon abraten.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"recommendation","normalized_action_text":"advise against using the real asset"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR-07","description":"temporary non-action","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die Waschstufe einbauen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir erstmal noch nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"temporary_non_action","normalized_action_text":"install the washing stage"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR-08","description":"concern only","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten das neue Material einsetzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das wäre kritisch.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"none","normalized_action_text":null},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"schema_version":"experimental-target-resolution-v1-diagnostic",
|
||||
"cases":[
|
||||
{"case_id":"TR1-V1","strategy":"self_contained","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"negative_act":{"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"working with Dr. Schlummer"},"expected":{"target_observation_id":"obs_1","concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":["nicht"]}},
|
||||
{"case_id":"TR2-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"verfolgen wir nicht weiter"},"expected":{"target_observation_id":"obs_1","concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR3-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"reale Anlage dafür nicht nutzen"},"expected":{"target_observation_id":"obs_1","concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[]}},
|
||||
{"case_id":"TR4-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"Versuch in der realen Anlage nicht durchführen"},"expected":{"target_observation_id":"obs_1","concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"]}}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
# topic_reconstruction_v2
|
||||
|
||||
Focused experimental Gold material derived from BUG-015 and the Progeo
|
||||
discussion. These cases evaluate topic-oriented reconstruction rather than
|
||||
exact protocol wording or flat category extraction.
|
||||
|
||||
The nine cases cover:
|
||||
|
||||
- an idea mentioned without further development;
|
||||
- multiple alternatives;
|
||||
- an unaccepted proposal;
|
||||
- a proposal with an objection;
|
||||
- an explicitly rejected alternative;
|
||||
- acceptance limited to a bounded trial;
|
||||
- discussion ending without a decision;
|
||||
- a resulting Action Item;
|
||||
- an outcome accompanied by an unresolved issue.
|
||||
|
||||
Evidence units carry stable local IDs. Expected criteria describe semantic
|
||||
features and prohibited promotions rather than exact generated sentences.
|
||||
@@ -0,0 +1,258 @@
|
||||
{
|
||||
"cases": [
|
||||
{
|
||||
"case_id": "a_idea_only",
|
||||
"description": "A geometry optimization idea is mentioned but not developed.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["geometr"],
|
||||
"required_event_types": ["introduced_idea"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {"minimum": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "b_multiple_options",
|
||||
"description": "Two alternatives for compensating insufficient specimen strength are discussed.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e3",
|
||||
"text": "Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["festigkeit", "gitter", "geometr"],
|
||||
"required_event_types": ["considered_option"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {"minimum": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "c_unaccepted_proposal",
|
||||
"description": "Contacting Dirk Textor is proposed but not accepted as work.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["textor", "einschätzung", "kontakt"],
|
||||
"required_event_types": ["proposal"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {"minimum": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "d_proposal_with_objection",
|
||||
"description": "Washing is considered and an energy-cost objection is raised.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["wasch", "reinig"],
|
||||
"required_event_types": ["proposal", "objection"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {"minimum": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "e_rejected_alternative",
|
||||
"description": "The collaboration with Dr. Schlummer is explicitly rejected after its cost is discussed.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e3",
|
||||
"text": "Antonius: Ja, das ist entschieden."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["schlummer", "zusammenarbeit"],
|
||||
"required_event_types": ["fact"],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"terms": ["nicht", "abgelehnt", "keine"],
|
||||
"scope_terms": ["zusammenarbeit", "versuch"],
|
||||
"certainties": ["rejected", "established"]
|
||||
},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {"minimum": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "f_trial_only_acceptance",
|
||||
"description": "A 20 percent variant is accepted only for a bounded trial, not as the final production solution.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e3",
|
||||
"text": "Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["20-prozent", "variante", "extruder"],
|
||||
"required_event_types": ["proposal", "clarification"],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"terms": ["test", "versuch"],
|
||||
"scope_terms": ["20 meter", "20 m", "nur", "begrenzt"],
|
||||
"certainties": ["established", "conditional"]
|
||||
},
|
||||
"actions": {
|
||||
"minimum": 1,
|
||||
"terms": ["test", "versuch", "20 meters", "20 meter"]
|
||||
},
|
||||
"unresolved": {
|
||||
"minimum": 1,
|
||||
"terms": ["final", "series", "serie", "adopt", "festgelegt"]
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "g_no_decision",
|
||||
"description": "Real recycling plant and Technikum alternatives are discussed without a group decision.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e3",
|
||||
"text": "Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["technikum", "reinig", "anlage"],
|
||||
"required_event_types": ["considered_option", "objection"],
|
||||
"outcome": {"required": false},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {
|
||||
"minimum": 1,
|
||||
"terms": ["reinigungsansatz", "verfügbar", "anlage", "prüfen"]
|
||||
}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "h_resulting_action",
|
||||
"description": "The discussion establishes an accepted review action with owner and deadline.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Nina: Ja, ich übernehme die Prüfung bis Freitag."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["messdaten", "prüfung", "prüfen", "verification", "data"],
|
||||
"required_event_types": [],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"terms": ["agrees", "übernimmt", "accepted", "verify"],
|
||||
"scope_terms": ["measurement", "messdaten", "verification"],
|
||||
"certainties": ["established"]
|
||||
},
|
||||
"actions": {
|
||||
"minimum": 1,
|
||||
"terms": ["messdaten", "prüf", "verify", "measurement"],
|
||||
"responsible": "Nina"
|
||||
},
|
||||
"unresolved": {"minimum": 0}
|
||||
}
|
||||
},
|
||||
{
|
||||
"case_id": "i_outcome_and_unresolved",
|
||||
"description": "The production-energy discussion establishes one bounded finding while publication remains unresolved.",
|
||||
"evidence_units": [
|
||||
{
|
||||
"evidence_id": "e1",
|
||||
"text": "Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e2",
|
||||
"text": "Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."
|
||||
},
|
||||
{
|
||||
"evidence_id": "e3",
|
||||
"text": "Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"
|
||||
},
|
||||
{
|
||||
"evidence_id": "e4",
|
||||
"text": "Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."
|
||||
}
|
||||
],
|
||||
"expected": {
|
||||
"subject_count": 1,
|
||||
"subject_terms": ["energie", "aufwand", "produktion"],
|
||||
"required_event_types": ["technical_finding"],
|
||||
"outcome": {
|
||||
"required": true,
|
||||
"terms": ["kein zusätzlicher", "keine zusätzliche", "unverändert"],
|
||||
"scope_terms": ["reine produktion", "produktion", "gegenüber virgin"],
|
||||
"certainties": ["established"]
|
||||
},
|
||||
"actions": {"minimum": 0},
|
||||
"unresolved": {
|
||||
"minimum": 1,
|
||||
"terms": ["veröffentlich", "freigabe", "energieaudit", "daten"]
|
||||
}
|
||||
}
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,230 @@
|
||||
import subprocess
|
||||
import tempfile
|
||||
import unittest
|
||||
import wave
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.audio.preparation import (
|
||||
DEFAULT_NORMALIZATION_FILTER,
|
||||
DEFAULT_NORMALIZATION_METHOD,
|
||||
AudioPreparationError,
|
||||
prepare_audio,
|
||||
)
|
||||
|
||||
|
||||
def _write_wav(
|
||||
path: Path, *, channels: int = 1, sample_rate: int = 16_000, sample_width: int = 2
|
||||
) -> None:
|
||||
with wave.open(str(path), "wb") as recording:
|
||||
recording.setnchannels(channels)
|
||||
recording.setsampwidth(sample_width)
|
||||
recording.setframerate(sample_rate)
|
||||
recording.writeframes(b"\x00" * channels * sample_width * 32)
|
||||
|
||||
|
||||
def _successful_runner(commands: list[list[str]]):
|
||||
def run(command, **kwargs):
|
||||
commands.append(list(command))
|
||||
_write_wav(Path(command[-1]))
|
||||
return subprocess.CompletedProcess(command, 0, "", "")
|
||||
|
||||
return run
|
||||
|
||||
|
||||
class AudioPreparationTests(unittest.TestCase):
|
||||
def test_supported_inputs_are_prepared_with_normalization_on_and_off(self) -> None:
|
||||
for suffix in (".wav", ".flac", ".m4a"):
|
||||
for normalization_enabled in (True, False):
|
||||
with (
|
||||
self.subTest(
|
||||
suffix=suffix, normalization_enabled=normalization_enabled
|
||||
),
|
||||
tempfile.TemporaryDirectory() as directory,
|
||||
):
|
||||
root = Path(directory)
|
||||
source = root / f"meeting{suffix}"
|
||||
if suffix == ".wav":
|
||||
_write_wav(source)
|
||||
else:
|
||||
source.write_bytes(b"original encoded audio")
|
||||
original = source.read_bytes()
|
||||
destination = root / "run" / "audio" / "prepared.wav"
|
||||
commands: list[list[str]] = []
|
||||
|
||||
with patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which",
|
||||
return_value="/usr/bin/ffmpeg",
|
||||
):
|
||||
result = prepare_audio(
|
||||
source,
|
||||
destination,
|
||||
normalization_enabled=normalization_enabled,
|
||||
runner=_successful_runner(commands),
|
||||
)
|
||||
|
||||
self.assertEqual(source.read_bytes(), original)
|
||||
self.assertEqual(result.prepared_path, destination)
|
||||
with wave.open(str(destination), "rb") as recording:
|
||||
self.assertEqual(recording.getnchannels(), 1)
|
||||
self.assertEqual(recording.getframerate(), 16_000)
|
||||
self.assertEqual(recording.getsampwidth(), 2)
|
||||
self.assertEqual(recording.getcomptype(), "NONE")
|
||||
self.assertEqual(commands[0][commands[0].index("-ac") + 1], "1")
|
||||
self.assertEqual(commands[0][commands[0].index("-ar") + 1], "16000")
|
||||
self.assertEqual(
|
||||
commands[0][commands[0].index("-c:a") + 1], "pcm_s16le"
|
||||
)
|
||||
self.assertEqual("-af" in commands[0], normalization_enabled)
|
||||
if normalization_enabled:
|
||||
self.assertEqual(
|
||||
commands[0][commands[0].index("-af") + 1],
|
||||
DEFAULT_NORMALIZATION_FILTER,
|
||||
)
|
||||
self.assertEqual(
|
||||
result.normalization_enabled, normalization_enabled
|
||||
)
|
||||
|
||||
def test_normalization_defaults_to_on_and_explicit_on_matches(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "meeting.wav"
|
||||
_write_wav(source)
|
||||
commands: list[list[str]] = []
|
||||
with patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which",
|
||||
return_value="/usr/bin/ffmpeg",
|
||||
):
|
||||
default = prepare_audio(
|
||||
source, root / "default.wav", runner=_successful_runner(commands)
|
||||
)
|
||||
explicit = prepare_audio(
|
||||
source,
|
||||
root / "explicit.wav",
|
||||
normalization_enabled=True,
|
||||
runner=_successful_runner(commands),
|
||||
)
|
||||
|
||||
self.assertTrue(default.normalization_enabled)
|
||||
self.assertTrue(explicit.normalization_enabled)
|
||||
self.assertEqual(
|
||||
commands[0][commands[0].index("-af") + 1],
|
||||
commands[1][commands[1].index("-af") + 1],
|
||||
)
|
||||
|
||||
def test_noncanonical_wav_is_normalized(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "stereo-48k.wav"
|
||||
_write_wav(source, channels=2, sample_rate=48_000)
|
||||
destination = root / "prepared.wav"
|
||||
|
||||
with patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which",
|
||||
return_value="/usr/bin/ffmpeg",
|
||||
):
|
||||
prepare_audio(source, destination, runner=_successful_runner([]))
|
||||
|
||||
with wave.open(str(destination), "rb") as recording:
|
||||
self.assertEqual(
|
||||
(recording.getnchannels(), recording.getframerate()), (1, 16_000)
|
||||
)
|
||||
|
||||
def test_ffmpeg_missing_has_actionable_error(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "meeting.flac"
|
||||
source.write_bytes(b"audio")
|
||||
|
||||
with (
|
||||
patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which", return_value=None
|
||||
),
|
||||
self.assertRaisesRegex(AudioPreparationError, "not found on PATH"),
|
||||
):
|
||||
prepare_audio(source, root / "prepared.wav")
|
||||
|
||||
def test_ffmpeg_failure_includes_diagnostic_and_preserves_source(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "meeting.m4a"
|
||||
source.write_bytes(b"original")
|
||||
|
||||
def fail(command, **kwargs):
|
||||
return subprocess.CompletedProcess(command, 1, "", "decoder exploded")
|
||||
|
||||
with (
|
||||
patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which",
|
||||
return_value="/usr/bin/ffmpeg",
|
||||
),
|
||||
self.assertRaisesRegex(AudioPreparationError, "decoder exploded"),
|
||||
):
|
||||
prepare_audio(source, root / "prepared.wav", runner=fail)
|
||||
|
||||
self.assertEqual(source.read_bytes(), b"original")
|
||||
self.assertFalse((root / "prepared.wav").exists())
|
||||
|
||||
def test_prepared_audio_metadata_is_traceable(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "unknown_meeting.flac"
|
||||
source.write_bytes(b"source")
|
||||
destination = root / "audio" / "prepared.wav"
|
||||
with patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which",
|
||||
return_value="/usr/bin/ffmpeg",
|
||||
):
|
||||
result = prepare_audio(
|
||||
source, destination, runner=_successful_runner([])
|
||||
)
|
||||
|
||||
metadata = result.metadata()
|
||||
self.assertEqual(metadata["original_source_name"], "unknown_meeting.flac")
|
||||
self.assertEqual(metadata["original_format"], "flac")
|
||||
self.assertEqual(
|
||||
metadata["prepared_audio_path"], str(destination.resolve())
|
||||
)
|
||||
self.assertEqual(metadata["preparation_method"], "ffmpeg")
|
||||
self.assertTrue(metadata["normalization_enabled"])
|
||||
self.assertEqual(
|
||||
metadata["normalization_method"], DEFAULT_NORMALIZATION_METHOD
|
||||
)
|
||||
self.assertEqual(
|
||||
metadata["normalization_filter"], DEFAULT_NORMALIZATION_FILTER
|
||||
)
|
||||
self.assertEqual(
|
||||
metadata["canonical_output"],
|
||||
{
|
||||
"container": "wav",
|
||||
"codec": "pcm_s16le",
|
||||
"channels": 1,
|
||||
"sample_rate_hz": 16_000,
|
||||
"bits_per_sample": 16,
|
||||
},
|
||||
)
|
||||
|
||||
def test_disabled_normalization_metadata_has_no_method_or_filter(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "meeting.m4a"
|
||||
source.write_bytes(b"source")
|
||||
with patch(
|
||||
"src.meeting_lab.audio.preparation.shutil.which",
|
||||
return_value="/usr/bin/ffmpeg",
|
||||
):
|
||||
result = prepare_audio(
|
||||
source,
|
||||
root / "prepared.wav",
|
||||
normalization_enabled=False,
|
||||
runner=_successful_runner([]),
|
||||
)
|
||||
|
||||
metadata = result.metadata()
|
||||
self.assertFalse(metadata["normalization_enabled"])
|
||||
self.assertIsNone(metadata["normalization_method"])
|
||||
self.assertIsNone(metadata["normalization_filter"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,200 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_collective import (
|
||||
DerivationValidationError,
|
||||
build_prompt,
|
||||
derive_collective_action,
|
||||
evaluate_case,
|
||||
load_gold_cases,
|
||||
validate_recognition,
|
||||
)
|
||||
|
||||
|
||||
GOLD_PATH = Path("tests/gold/collective_commitment_v0/cases.json")
|
||||
|
||||
|
||||
def recognition_for(case):
|
||||
form = case["expected_recognition"]["commitment_form"]
|
||||
action = "20 Meter testen"
|
||||
if case["case_id"] == "CC-08":
|
||||
action = "20 Meter testen, aber nur im Technikum"
|
||||
return {
|
||||
"observation_id": "obs_1",
|
||||
"commitment_form": form,
|
||||
"normalized_action_text": action,
|
||||
}
|
||||
|
||||
|
||||
class CollectiveCommitmentGoldExperimentTests(unittest.TestCase):
|
||||
@classmethod
|
||||
def setUpClass(cls):
|
||||
cls.cases = load_gold_cases(GOLD_PATH)
|
||||
cls.by_id = {case["case_id"]: case for case in cls.cases}
|
||||
|
||||
def test_fixture_contains_exactly_cc_01_through_cc_10(self):
|
||||
self.assertEqual(list(self.by_id), [f"CC-{number:02d}" for number in range(1, 11)])
|
||||
|
||||
def test_all_cases_are_single_minimal_v3_style_observations(self):
|
||||
keys = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
|
||||
for case in self.cases:
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertEqual(len(case["observations"]), 1)
|
||||
self.assertEqual(set(case["observations"][0]), keys)
|
||||
|
||||
def test_cc_01_establishes_collective_action_without_person_and_with_due(self):
|
||||
case = self.by_id["CC-01"]
|
||||
gates, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertTrue(all(gates.values()))
|
||||
self.assertEqual(result["status"], "established")
|
||||
self.assertEqual(result["commitment_scope"], "collective")
|
||||
self.assertIsNone(result["responsible_person"])
|
||||
self.assertEqual(result["due"], "nächste Woche")
|
||||
|
||||
def test_individual_commitment_routes_out_of_collective_path(self):
|
||||
self._assert_unestablished("CC-02", "collective_commitment_form")
|
||||
|
||||
def test_tentative_suggestion_impersonal_and_passive_remain_unestablished(self):
|
||||
for case_id in ("CC-03", "CC-04", "CC-05", "CC-06"):
|
||||
with self.subTest(case=case_id):
|
||||
self._assert_unestablished(case_id, "collective_commitment_form")
|
||||
|
||||
def test_rejection_remains_unestablished_and_negation_gate_is_negative(self):
|
||||
case = self.by_id["CC-07"]
|
||||
gates, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertFalse(gates["collective_commitment_form"])
|
||||
self.assertFalse(gates["no_explicit_negation"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_qualifier_case_establishes_preserves_limit_and_has_no_due(self):
|
||||
case = self.by_id["CC-08"]
|
||||
_, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertIsNotNone(result)
|
||||
self.assertIn("nur im Technikum", result["content"])
|
||||
self.assertIsNone(result["due"])
|
||||
self.assertIsNone(result["responsible_person"])
|
||||
|
||||
def test_collective_without_deadline_establishes_with_null_due(self):
|
||||
case = self.by_id["CC-09"]
|
||||
_, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertIsNotNone(result)
|
||||
self.assertIsNone(result["due"])
|
||||
|
||||
def test_speaker_ownership_trap_never_assigns_martin(self):
|
||||
case = self.by_id["CC-10"]
|
||||
_, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertIsNotNone(result)
|
||||
self.assertEqual(case["observations"][0]["speaker"], "Martin")
|
||||
self.assertIsNone(result["responsible_person"])
|
||||
|
||||
def test_changing_only_speaker_cannot_create_individual_owner(self):
|
||||
case = deepcopy(self.by_id["CC-01"])
|
||||
for speaker in ("Martin", "Clara", "Antonius"):
|
||||
case["observations"][0]["speaker"] = speaker
|
||||
_, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
with self.subTest(speaker=speaker):
|
||||
self.assertIsNotNone(result)
|
||||
self.assertIsNone(result["responsible_person"])
|
||||
|
||||
def test_none_and_individual_forms_never_establish(self):
|
||||
case = self.by_id["CC-01"]
|
||||
for form in ("none", "individual_first_person"):
|
||||
recognition = recognition_for(case)
|
||||
recognition["commitment_form"] = form
|
||||
_, result = derive_collective_action(case["observations"], recognition)
|
||||
with self.subTest(form=form):
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_non_none_commitment_requires_action_text(self):
|
||||
case = self.by_id["CC-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["normalized_action_text"] = None
|
||||
with self.assertRaisesRegex(DerivationValidationError, "requires normalized_action_text"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_none_commitment_allows_null_action_text_but_never_establishes(self):
|
||||
case = self.by_id["CC-03"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["normalized_action_text"] = None
|
||||
_, result = derive_collective_action(case["observations"], recognition)
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_unknown_observation_id_is_rejected(self):
|
||||
case = self.by_id["CC-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["observation_id"] = "obs_99"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown observation"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_inconsistent_duplicate_evidence_provenance_is_rejected(self):
|
||||
fixture = json.loads(GOLD_PATH.read_text())
|
||||
fixture["cases"][0]["observations"].append({
|
||||
"observation_id": "obs_2", "evidence_id": "e1", "content": "Martin: Zusatz.",
|
||||
"speaker": "Martin", "named_person": None, "addressee": None,
|
||||
})
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
path = Path(temporary) / "cases.json"
|
||||
path.write_text(json.dumps(fixture), encoding="utf-8")
|
||||
with self.assertRaisesRegex(DerivationValidationError, "provenance must be unique"):
|
||||
load_gold_cases(path)
|
||||
|
||||
def test_conflicting_deadlines_prevent_establishment(self):
|
||||
case = deepcopy(self.by_id["CC-01"])
|
||||
case["observations"][0]["content"] += " Bis Mittwoch."
|
||||
gates, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertFalse(gates["deadline_supported_and_consistent"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_forbidden_fields_are_rejected_recursively(self):
|
||||
case = self.by_id["CC-01"]
|
||||
forbidden = (
|
||||
"responsible_person", "responsibility", "responsibility_scope",
|
||||
"requested_actor", "owner", "ownership", "assignee", "status",
|
||||
"established", "action_item", "protocol_category", "decision",
|
||||
"unresolved_issue", "confidence", "relation", "relations", "graph",
|
||||
)
|
||||
for field in forbidden:
|
||||
recognition = recognition_for(case)
|
||||
recognition["wrapper"] = {field: "forbidden"}
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_unknown_schema_field_is_rejected(self):
|
||||
case = self.by_id["CC-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["explanation"] = "extra"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown keys"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_provenance_survives_and_successes_always_have_null_person(self):
|
||||
for case_id in ("CC-01", "CC-08", "CC-09", "CC-10"):
|
||||
case = self.by_id[case_id]
|
||||
_, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
with self.subTest(case=case_id):
|
||||
self.assertEqual(result["support"]["commitment"], {"observation_id": "obs_1", "evidence_id": "e1"})
|
||||
self.assertIsNone(result["responsible_person"])
|
||||
|
||||
def test_all_expected_recognitions_have_correct_final_outcome(self):
|
||||
for case in self.cases:
|
||||
evaluation = evaluate_case(case, recognition_for(case))
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertEqual(evaluation["classification"], "PASS")
|
||||
|
||||
def test_prompt_is_fixed_narrow_and_contains_no_gold_expectation(self):
|
||||
prompt = build_prompt(self.by_id["CC-01"]["observations"])
|
||||
self.assertIn("commitment_form", prompt)
|
||||
self.assertNotIn("expected_result", prompt)
|
||||
self.assertNotIn("Who is responsible", prompt)
|
||||
|
||||
def _assert_unestablished(self, case_id, failed_gate):
|
||||
case = self.by_id[case_id]
|
||||
gates, result = derive_collective_action(case["observations"], recognition_for(case))
|
||||
self.assertFalse(gates[failed_gate])
|
||||
self.assertIsNone(result)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,82 @@
|
||||
import copy, json, unittest
|
||||
from pathlib import Path
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection_v1 import (
|
||||
build_target_prompt, derive, load_cases, validate_target,
|
||||
)
|
||||
|
||||
CASES=load_cases(Path("tests/gold/controlled_rejection_v1/cases.json"))
|
||||
BY_ID={c["case_id"]:c for c in CASES}
|
||||
|
||||
def negative(case, form=None):
|
||||
return {"observation_id":case["expected"]["candidate_observation_id"],"negative_act_form":form or case["expected"]["negative_act_form"],"normalized_action_text":None if (form or case["expected"]["negative_act_form"])=="none" else "semantische Aktion"}
|
||||
def target(case, oid=None, text="konkrete Zielhandlung"):
|
||||
return {"candidate_observation_id":case["expected"]["candidate_observation_id"],"target_observation_id":oid or case["expected"]["target_observation_id"],"normalized_target_text":text}
|
||||
|
||||
class ControlledRejectionV1Tests(unittest.TestCase):
|
||||
def test_positive_forms_derive_and_provenance_survives(self):
|
||||
for cid in ("CR-01","CR-02","CR-07","CR-08"):
|
||||
c=BY_ID[cid]; out=derive(c["observations"],negative(c),target(c))
|
||||
self.assertEqual(out["derived_result"]["status"],"explicitly_rejected")
|
||||
self.assertEqual(out["derived_result"]["support"]["target"]["evidence_id"],"e1")
|
||||
self.assertNotIn("responsible_person",json.dumps(out["derived_result"]))
|
||||
|
||||
def test_noneligible_form_never_derives_even_with_target(self):
|
||||
for form in ["personal_preference","recommendation","temporary_non_action","none"]:
|
||||
c=BY_ID["CR-03"]; self.assertIsNone(derive(c["observations"],negative(c,form),target(c))["derived_result"])
|
||||
|
||||
def test_missing_target_prevents_derivation(self):
|
||||
c=BY_ID["CR-02"]; t=target(c); t.update(target_observation_id=None,normalized_target_text=None)
|
||||
self.assertIsNone(derive(c["observations"],negative(c),t)["derived_result"])
|
||||
|
||||
def test_unknown_ids_rejected(self):
|
||||
for field in ["candidate_observation_id","target_observation_id"]:
|
||||
c=BY_ID["CR-02"]; t=target(c); t[field]="obs_unknown"
|
||||
with self.assertRaises(DerivationValidationError): validate_target(t,c["observations"])
|
||||
|
||||
def test_target_after_candidate_cannot_derive(self):
|
||||
c=copy.deepcopy(BY_ID["CR-02"]); n={"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"x"}; t={"candidate_observation_id":"obs_1","target_observation_id":"obs_2","normalized_target_text":"x"}
|
||||
self.assertIsNone(derive(c["observations"],n,t)["derived_result"])
|
||||
|
||||
def test_same_observation_target_allowed(self):
|
||||
c=BY_ID["CR-01"]; self.assertIsNotNone(derive(c["observations"],negative(c),target(c))["derived_result"])
|
||||
|
||||
def test_duplicate_provenance_rejected(self):
|
||||
for field in ["observation_id","evidence_id"]:
|
||||
c=copy.deepcopy(BY_ID["CR-02"]); c["observations"][1][field]=c["observations"][0][field]
|
||||
with self.assertRaises(DerivationValidationError): derive(c["observations"],negative(BY_ID["CR-02"]),target(BY_ID["CR-02"]))
|
||||
|
||||
def test_target_text_constraints(self):
|
||||
c=BY_ID["CR-02"]
|
||||
with self.assertRaises(DerivationValidationError): validate_target(target(c,text=""),c["observations"])
|
||||
t=target(c); t["target_observation_id"]=None
|
||||
with self.assertRaises(DerivationValidationError): validate_target(t,c["observations"])
|
||||
|
||||
def test_null_target_accepts_only_null_text(self):
|
||||
c=BY_ID["CR-02"]; t=target(c); t.update(target_observation_id=None,normalized_target_text=None)
|
||||
self.assertEqual(validate_target(t,c["observations"]),t)
|
||||
|
||||
def test_forbidden_and_unknown_fields_rejected(self):
|
||||
for extra in [{"status":"rejected"},{"nested":{"decision":True}},{"extra":1}]:
|
||||
c=BY_ID["CR-02"]; t=target(c); t.update(extra)
|
||||
with self.assertRaises(DerivationValidationError): validate_target(t,c["observations"])
|
||||
|
||||
def test_candidate_outputs_must_agree(self):
|
||||
c=BY_ID["CR-02"]; t=target(c); t["candidate_observation_id"]="obs_1"
|
||||
with self.assertRaises(DerivationValidationError): derive(c["observations"],negative(c),t)
|
||||
|
||||
def test_scope_and_alternative_fixture_contract(self):
|
||||
self.assertLessEqual({"real","Druckversuch"},{x for group in BY_ID["CR-07"]["expected"]["material_concepts"] for x in group})
|
||||
self.assertEqual(BY_ID["CR-08"]["expected"]["forbidden_concepts"],["Technikum"])
|
||||
|
||||
def test_target_prompt_is_semantic_only_and_fixed(self):
|
||||
prompt=build_target_prompt(BY_ID["CR-08"])
|
||||
self.assertIn("Ignore any separate positive alternative",prompt)
|
||||
self.assertIn("Do not classify the negative act",prompt)
|
||||
|
||||
def test_accepted_negative_act_inputs_are_exactly_reused(self):
|
||||
root=Path("artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run")
|
||||
for cid,nid in (("CR-01","NA-01"),("CR-03","NA-03"),("CR-04","NA-04"),("CR-05","NA-05"),("CR-06","NA-06")):
|
||||
accepted=json.loads((root/nid.lower()/"v3_style_input_observations.json").read_text())
|
||||
self.assertEqual(BY_ID[cid]["observations"],accepted)
|
||||
@@ -0,0 +1,215 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_h import (
|
||||
SCHEMA_VERSION,
|
||||
DerivationValidationError,
|
||||
build_ollama_payload,
|
||||
derive_action,
|
||||
load_v3_observations,
|
||||
parse_model_json,
|
||||
run_experiment,
|
||||
validate_semantic_recognition,
|
||||
)
|
||||
|
||||
|
||||
ACCEPTED_H_PATH = Path(
|
||||
"artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/"
|
||||
"h_resulting_action/parsed_observations.json"
|
||||
)
|
||||
|
||||
|
||||
class ControlledSemanticDerivationHTests(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
self.observations = [
|
||||
{
|
||||
"observation_id": "obs_1", "evidence_id": "e1",
|
||||
"content": "Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Friday?",
|
||||
"speaker": "Antonius", "named_person": "Nina", "addressee": "Nina",
|
||||
},
|
||||
{
|
||||
"observation_id": "obs_2", "evidence_id": "e2",
|
||||
"content": "Nina: Ja, ich übernehme die Prüfung bis Freitag.",
|
||||
"speaker": "Nina", "named_person": None, "addressee": None,
|
||||
},
|
||||
]
|
||||
self.recognition = {
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"request": {
|
||||
"observation_id": "obs_1", "is_concrete_request": True,
|
||||
"normalized_action_text": "Prüfung der Messdaten",
|
||||
},
|
||||
"acceptance": {
|
||||
"observation_id": "obs_2", "is_explicit_commitment": True,
|
||||
"same_requested_work": True,
|
||||
"normalized_action_text": "die Prüfung",
|
||||
},
|
||||
}
|
||||
|
||||
def derive(self, observations=None, recognition=None):
|
||||
return derive_action(
|
||||
observations if observations is not None else self.observations,
|
||||
recognition if recognition is not None else self.recognition,
|
||||
)
|
||||
|
||||
def test_actual_accepted_v3_artifact_is_the_experiment_input(self):
|
||||
actual = load_v3_observations(ACCEPTED_H_PATH)
|
||||
self.assertEqual(actual, self.observations)
|
||||
|
||||
def test_valid_semantic_recognition_has_no_derivation_fields(self):
|
||||
validate_semantic_recognition(self.recognition, self.observations)
|
||||
serialized = json.dumps(self.recognition)
|
||||
for forbidden in ("responsible_person", "requested_actor", "status", "established", "action_item"):
|
||||
self.assertNotIn(forbidden, serialized)
|
||||
|
||||
def test_request_and_acceptance_provenance_survive(self):
|
||||
gates, result = self.derive()
|
||||
self.assertTrue(all(gates.values()))
|
||||
self.assertEqual(result["support"]["request"], {"observation_id": "obs_1", "evidence_id": "e1"})
|
||||
self.assertEqual(result["support"]["acceptance"], {"observation_id": "obs_2", "evidence_id": "e2"})
|
||||
|
||||
def test_valid_sequence_establishes_expected_action(self):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition["request"]["normalized_action_text"] = "Prüfung der Messdaten bis Friday"
|
||||
_, result = self.derive(recognition=recognition)
|
||||
self.assertEqual(result["content"], "Prüfung der Messdaten")
|
||||
self.assertEqual(result["status"], "established")
|
||||
self.assertEqual(result["requested_actor"], "Nina")
|
||||
self.assertEqual(result["responsible_person"], "Nina")
|
||||
self.assertEqual(result["due"], "Freitag")
|
||||
|
||||
def test_lexical_identity_is_not_required(self):
|
||||
self.assertNotEqual(
|
||||
self.recognition["request"]["normalized_action_text"],
|
||||
self.recognition["acceptance"]["normalized_action_text"],
|
||||
)
|
||||
gates, result = self.derive()
|
||||
self.assertTrue(gates["same_requested_work"])
|
||||
self.assertIsNotNone(result)
|
||||
|
||||
def test_request_alone_does_not_establish(self):
|
||||
gates, result = self.derive(observations=self.observations[:1])
|
||||
self.assertFalse(gates["acceptance_observation_exists"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_noncommitting_or_acknowledging_response_does_not_establish(self):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition["acceptance"]["is_explicit_commitment"] = False
|
||||
gates, result = self.derive(recognition=recognition)
|
||||
self.assertFalse(gates["acceptance_semantic_positive"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_different_response_speaker_does_not_establish(self):
|
||||
observations = deepcopy(self.observations)
|
||||
observations[1]["speaker"] = "Martin"
|
||||
gates, result = self.derive(observations=observations)
|
||||
self.assertFalse(gates["acceptance_speaker_matches_addressee"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_different_accepted_work_does_not_establish(self):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition["acceptance"]["same_requested_work"] = False
|
||||
recognition["acceptance"]["normalized_action_text"] = "Angebot prüfen"
|
||||
gates, result = self.derive(recognition=recognition)
|
||||
self.assertFalse(gates["same_requested_work"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_tentative_acceptance_does_not_establish(self):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition["acceptance"]["is_explicit_commitment"] = False
|
||||
_, result = self.derive(recognition=recognition)
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_named_person_speaker_or_addressee_alone_cannot_establish(self):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition["acceptance"]["is_explicit_commitment"] = False
|
||||
gates, result = self.derive(recognition=recognition)
|
||||
self.assertEqual(self.observations[0]["named_person"], "Nina")
|
||||
self.assertEqual(self.observations[0]["addressee"], "Nina")
|
||||
self.assertEqual(self.observations[1]["speaker"], "Nina")
|
||||
self.assertFalse(gates["acceptance_semantic_positive"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_acceptance_must_follow_request(self):
|
||||
observations = list(reversed(deepcopy(self.observations)))
|
||||
gates, result = self.derive(observations=observations)
|
||||
self.assertFalse(gates["acceptance_after_request"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_conflicting_deadlines_do_not_establish(self):
|
||||
observations = deepcopy(self.observations)
|
||||
observations[1]["content"] = "Nina: Ja, ich übernehme die Prüfung bis Donnerstag."
|
||||
gates, result = self.derive(observations=observations)
|
||||
self.assertFalse(gates["deadline_consistent"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_unknown_observation_reference_is_rejected(self):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition["acceptance"]["observation_id"] = "obs_9"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown observation"):
|
||||
validate_semantic_recognition(recognition, self.observations)
|
||||
|
||||
def test_inconsistent_evidence_provenance_is_rejected(self):
|
||||
data = json.loads(ACCEPTED_H_PATH.read_text())
|
||||
data["observations"][1]["evidence_id"] = "e1"
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
path = Path(temporary) / "observations.json"
|
||||
path.write_text(json.dumps(data), encoding="utf-8")
|
||||
with self.assertRaisesRegex(DerivationValidationError, "inconsistent evidence provenance"):
|
||||
load_v3_observations(path)
|
||||
|
||||
def test_responsibility_or_status_in_llm_output_is_rejected(self):
|
||||
for field in ("responsibility", "responsible_person", "status", "established"):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition[field] = "forbidden"
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||
validate_semantic_recognition(recognition, self.observations)
|
||||
|
||||
def test_protocol_or_unrelated_semantic_concepts_are_rejected(self):
|
||||
for field in ("protocol_category", "decision", "unresolved_issue", "graph", "confidence"):
|
||||
recognition = deepcopy(self.recognition)
|
||||
recognition[field] = "forbidden"
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||
validate_semantic_recognition(recognition, self.observations)
|
||||
|
||||
def test_malformed_json_is_rejected(self):
|
||||
with self.assertRaises(json.JSONDecodeError):
|
||||
parse_model_json("{bad json")
|
||||
|
||||
def test_payload_has_one_call_controls(self):
|
||||
payload = build_ollama_payload("qwen3.5:9B", "prompt", 16384, 1024)
|
||||
self.assertFalse(payload["think"])
|
||||
self.assertFalse(payload["stream"])
|
||||
self.assertEqual(payload["options"]["temperature"], 0)
|
||||
|
||||
def test_run_preserves_all_artifacts_without_real_ollama(self):
|
||||
raw = json.dumps(self.recognition, ensure_ascii=False)
|
||||
from argparse import Namespace
|
||||
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
output = Path(temporary) / "run"
|
||||
args = Namespace(
|
||||
observations=ACCEPTED_H_PATH, output=output, model="qwen3.5:9B",
|
||||
endpoint="http://unused", timeout=1, num_ctx=16384, num_predict=1024,
|
||||
)
|
||||
with patch(
|
||||
"src.meeting_lab.controlled_semantic_derivation.experiment_h.call_ollama",
|
||||
return_value=(raw, {"model": "qwen3.5:9B"}),
|
||||
):
|
||||
summary = run_experiment(args)
|
||||
self.assertTrue(summary["action_established"])
|
||||
for filename in (
|
||||
"v3_input_observations.json", "prompt.txt", "raw_model_response.txt",
|
||||
"parsed_semantic_recognition.json", "structural_validation.json",
|
||||
"deterministic_gate_results.json", "final_derived_result.json",
|
||||
"ollama_metadata.json", "summary.json",
|
||||
):
|
||||
self.assertTrue((output / filename).is_file(), filename)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,222 @@
|
||||
import json
|
||||
import os
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from types import SimpleNamespace
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.diarization.alignment import align_transcript, write_diarized_transcript
|
||||
from src.meeting_lab.diarization.backend import (
|
||||
DiarizationError,
|
||||
diarize_audio,
|
||||
run_container_pyannote,
|
||||
select_device,
|
||||
)
|
||||
|
||||
|
||||
class FakeCuda:
|
||||
def __init__(self, available: bool, *, name: str = "Test GPU", failure=None):
|
||||
self.available = available
|
||||
self.name = name
|
||||
self.failure = failure
|
||||
|
||||
def is_available(self):
|
||||
return self.available
|
||||
|
||||
def get_device_name(self, index):
|
||||
if self.failure:
|
||||
raise self.failure
|
||||
return self.name
|
||||
|
||||
|
||||
class FakeTorch:
|
||||
def __init__(self, available: bool, *, failure=None):
|
||||
self.cuda = FakeCuda(available, failure=failure)
|
||||
self.probes = []
|
||||
|
||||
def device(self, name):
|
||||
return name
|
||||
|
||||
def zeros(self, size, *, device):
|
||||
self.probes.append(device)
|
||||
if self.cuda.failure:
|
||||
raise self.cuda.failure
|
||||
return [0]
|
||||
|
||||
|
||||
class DeviceSelectionTests(unittest.TestCase):
|
||||
def test_auto_selects_usable_gpu(self):
|
||||
torch = FakeTorch(True)
|
||||
self.assertEqual(select_device("auto", torch), ("cuda", "Test GPU"))
|
||||
self.assertEqual(torch.probes, ["cuda"])
|
||||
|
||||
def test_auto_falls_back_to_cpu(self):
|
||||
self.assertEqual(select_device("auto", FakeTorch(False)), ("cpu", None))
|
||||
self.assertEqual(
|
||||
select_device("auto", FakeTorch(True, failure=RuntimeError("probe"))),
|
||||
("cpu", None),
|
||||
)
|
||||
|
||||
def test_explicit_cpu_does_not_probe_gpu(self):
|
||||
torch = FakeTorch(True)
|
||||
self.assertEqual(select_device("cpu", torch), ("cpu", None))
|
||||
self.assertEqual(torch.probes, [])
|
||||
|
||||
def test_explicit_gpu_fails_when_unavailable(self):
|
||||
with self.assertRaisesRegex(DiarizationError, "GPU is unavailable"):
|
||||
select_device("gpu", FakeTorch(False))
|
||||
|
||||
|
||||
class AlignmentTests(unittest.TestCase):
|
||||
def test_exclusive_overlap_assigns_anonymous_speakers(self):
|
||||
transcript = {
|
||||
"text": "Original unchanged text.",
|
||||
"segments": [
|
||||
{"id": 0, "start": 0.0, "end": 4.0, "text": "Hallo"},
|
||||
{"id": 1, "start": 4.0, "end": 6.0, "text": "Antwort"},
|
||||
],
|
||||
}
|
||||
original = json.loads(json.dumps(transcript))
|
||||
turns = [
|
||||
{"start": 0.0, "end": 3.0, "speaker_id": "SPEAKER_00"},
|
||||
{"start": 3.0, "end": 6.0, "speaker_id": "SPEAKER_01"},
|
||||
]
|
||||
|
||||
derived = align_transcript(transcript, turns)
|
||||
|
||||
self.assertEqual(transcript, original)
|
||||
self.assertEqual(derived["segments"][0]["speaker_id"], "SPEAKER_00")
|
||||
self.assertEqual(derived["segments"][0]["speaker_overlap_seconds"], 3.0)
|
||||
self.assertEqual(derived["segments"][1]["speaker_id"], "SPEAKER_01")
|
||||
self.assertIn("SPEAKER_00: Hallo", derived["text"])
|
||||
self.assertTrue(derived["speaker_labels_anonymous"])
|
||||
self.assertEqual(derived["alignment_source"], "exclusive_diarization")
|
||||
|
||||
def test_speaker_aware_transcript_is_a_separate_artifact(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
source = root / "transcript.json"
|
||||
turns = root / "exclusive_turns.json"
|
||||
source_text = json.dumps(
|
||||
{"text": "Original", "segments": [{"start": 0, "end": 1, "text": "Hi"}]}
|
||||
)
|
||||
source.write_text(source_text, encoding="utf-8")
|
||||
turns.write_text(
|
||||
json.dumps([{"start": 0, "end": 1, "speaker_id": "SPEAKER_07"}]),
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
json_path, text_path = write_diarized_transcript(source, turns, root / "derived")
|
||||
|
||||
self.assertEqual(source.read_text(encoding="utf-8"), source_text)
|
||||
self.assertNotEqual(json_path, source)
|
||||
self.assertIn("SPEAKER_07", json_path.read_text(encoding="utf-8"))
|
||||
self.assertIn("SPEAKER_07", text_path.read_text(encoding="utf-8"))
|
||||
|
||||
|
||||
class ContainerAdapterTests(unittest.TestCase):
|
||||
def test_container_configuration_and_metadata_do_not_persist_token(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
audio = root / "audio.wav"
|
||||
output = root / "output"
|
||||
audio.write_bytes(b"audio")
|
||||
observed = {}
|
||||
|
||||
def fake_runner(command, **kwargs):
|
||||
observed["command"] = command
|
||||
output.mkdir(exist_ok=True)
|
||||
metadata = {
|
||||
"backend": "pyannote.audio",
|
||||
"model": "pyannote/speaker-diarization-community-1",
|
||||
"requested_device_mode": "gpu",
|
||||
"actual_device": "cuda",
|
||||
"credentials_persisted": False,
|
||||
}
|
||||
(output / "metadata.json").write_text(json.dumps(metadata))
|
||||
return SimpleNamespace(returncode=0, stdout="ok", stderr="")
|
||||
|
||||
with patch.dict("os.environ", {"HF_TOKEN": "secret-token"}):
|
||||
result = run_container_pyannote(
|
||||
audio,
|
||||
output,
|
||||
"gpu",
|
||||
image="test/image",
|
||||
container_args=("--device=/dev/test",),
|
||||
runner=fake_runner,
|
||||
uid_getter=lambda: 2345,
|
||||
gid_getter=lambda: 3456,
|
||||
)
|
||||
|
||||
command = observed["command"]
|
||||
shell_command = command[-1]
|
||||
persisted = "".join(
|
||||
path.read_text(encoding="utf-8")
|
||||
for path in output.iterdir()
|
||||
if path.is_file()
|
||||
)
|
||||
self.assertNotIn("secret-token", persisted)
|
||||
self.assertNotIn("secret-token", command)
|
||||
self.assertIn("HF_TOKEN", command)
|
||||
self.assertIn("chown -R 2345:3456 /output", shell_command)
|
||||
self.assertIn("chmod -R u+rwX /output", shell_command)
|
||||
self.assertNotIn("1000:1000", shell_command)
|
||||
device_index = command.index("--device=/dev/test")
|
||||
self.assertLess(device_index, command.index("test/image"))
|
||||
self.assertFalse(result.metadata["credentials_persisted"])
|
||||
self.assertEqual(result.metadata["runtime_adapter"], "container")
|
||||
self.assertTrue(
|
||||
all(os.access(path, os.W_OK) for path in (output, *output.rglob("*")))
|
||||
)
|
||||
|
||||
def test_unwritable_container_artifact_is_rejected_before_metadata_update(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
audio = root / "audio.wav"
|
||||
output = root / "output"
|
||||
audio.write_bytes(b"audio")
|
||||
|
||||
def fake_runner(command, **kwargs):
|
||||
output.mkdir(exist_ok=True)
|
||||
metadata = output / "metadata.json"
|
||||
metadata.write_text("{}", encoding="utf-8")
|
||||
metadata.chmod(0o444)
|
||||
return SimpleNamespace(returncode=0, stdout="", stderr="")
|
||||
|
||||
with patch(
|
||||
"src.meeting_lab.diarization.backend.os.access",
|
||||
side_effect=lambda path, mode: Path(path).name != "metadata.json",
|
||||
):
|
||||
with self.assertRaisesRegex(DiarizationError, "not writable"):
|
||||
run_container_pyannote(
|
||||
audio,
|
||||
output,
|
||||
"cpu",
|
||||
image="test/image",
|
||||
runner=fake_runner,
|
||||
)
|
||||
|
||||
def test_orchestrator_dispatches_runtime(self):
|
||||
with patch(
|
||||
"src.meeting_lab.diarization.backend.run_container_pyannote"
|
||||
) as container:
|
||||
diarize_audio(
|
||||
Path("audio.wav"),
|
||||
Path("out"),
|
||||
"cpu",
|
||||
runtime="container",
|
||||
container_image="image",
|
||||
container_args=("--arg",),
|
||||
)
|
||||
container.assert_called_once_with(
|
||||
Path("audio.wav"),
|
||||
Path("out"),
|
||||
"cpu",
|
||||
image="image",
|
||||
container_args=("--arg",),
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,321 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
import requests
|
||||
|
||||
from scripts import run_direct_protocol
|
||||
from src.meeting_lab.llm import ollama
|
||||
from src.meeting_lab.llm.ollama import OllamaError, OllamaGeneration
|
||||
from src.meeting_lab.protocol.generate_direct_protocol import (
|
||||
DirectProtocolError,
|
||||
generate_direct_protocol,
|
||||
load_compact_transcript,
|
||||
)
|
||||
|
||||
|
||||
VALID_CONTEXT = """schema_version: "1"
|
||||
meeting:
|
||||
meeting_id: "test-meeting"
|
||||
title: "Test Meeting"
|
||||
language: "de"
|
||||
participants: []
|
||||
mentioned_people: []
|
||||
organization:
|
||||
departments: []
|
||||
known_entities: {}
|
||||
"""
|
||||
|
||||
|
||||
def write_transcript(path: Path, text: str = "Wir besprechen den Projektstatus.") -> None:
|
||||
path.write_text(json.dumps({"text": text, "segments": []}), encoding="utf-8")
|
||||
|
||||
|
||||
def generation(text: str = "# Meeting Protocol\n\n## Status\nUnveraendert.") -> OllamaGeneration:
|
||||
return OllamaGeneration(
|
||||
raw_response={
|
||||
"response": text,
|
||||
"done": True,
|
||||
"done_reason": "stop",
|
||||
"prompt_eval_count": 123,
|
||||
"eval_count": 17,
|
||||
"prompt_eval_duration": 1000,
|
||||
"eval_duration": 2000,
|
||||
"total_duration": 4000,
|
||||
},
|
||||
text=text,
|
||||
client_wall_time_seconds=0.25,
|
||||
)
|
||||
|
||||
|
||||
class TranscriptLoadingTests(unittest.TestCase):
|
||||
def test_valid_transcript_is_accepted(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
path = Path(directory) / "transcript.json"
|
||||
write_transcript(path)
|
||||
self.assertEqual(load_compact_transcript(path), "Wir besprechen den Projektstatus.")
|
||||
|
||||
def test_missing_top_level_text_is_rejected(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
path = Path(directory) / "transcript.json"
|
||||
path.write_text('{"segments": []}', encoding="utf-8")
|
||||
with self.assertRaisesRegex(DirectProtocolError, "top-level 'text'"):
|
||||
load_compact_transcript(path)
|
||||
|
||||
def test_empty_transcript_is_rejected(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
path = Path(directory) / "transcript.json"
|
||||
write_transcript(path, " \n")
|
||||
with self.assertRaisesRegex(DirectProtocolError, "non-empty string"):
|
||||
load_compact_transcript(path)
|
||||
|
||||
def test_malformed_json_is_rejected(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
path = Path(directory) / "transcript.json"
|
||||
path.write_text("{", encoding="utf-8")
|
||||
with self.assertRaisesRegex(DirectProtocolError, "not valid JSON"):
|
||||
load_compact_transcript(path)
|
||||
|
||||
|
||||
class GeneratorTests(unittest.TestCase):
|
||||
def test_prompt_requires_contextual_discussion_density_without_transcript_replay(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
transcript = Path(directory) / "transcript.json"
|
||||
write_transcript(transcript)
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
model_check=Mock(return_value={}),
|
||||
generation_call=Mock(return_value=generation()),
|
||||
)
|
||||
|
||||
self.assertIn("vollständiges, strukturiertes", result.exact_prompt)
|
||||
self.assertIn("relevante Diskussionsverläufe", result.exact_prompt)
|
||||
self.assertIn("unterschiedliche Positionen", result.exact_prompt)
|
||||
self.assertIn("Entscheidungsgrundlagen", result.exact_prompt)
|
||||
self.assertIn("nicht am Meeting teilgenommen haben", result.exact_prompt)
|
||||
self.assertIn("keine reine Wiedergabe des Transkripts", result.exact_prompt)
|
||||
self.assertIn("nicht unnötig durch Wiederholungen", result.exact_prompt)
|
||||
|
||||
def test_optional_context_absent_and_generation_called_once(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
transcript = Path(directory) / "transcript.json"
|
||||
write_transcript(transcript)
|
||||
check = Mock(return_value={"model": "qwen3.6:35B-A3B"})
|
||||
call = Mock(return_value=generation())
|
||||
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
model_check=check,
|
||||
generation_call=call,
|
||||
)
|
||||
|
||||
self.assertIn("Kein Meeting-Kontext", result.exact_prompt)
|
||||
self.assertEqual(check.call_count, 1)
|
||||
self.assertEqual(call.call_count, 1)
|
||||
self.assertEqual(call.call_args.args[1], "qwen3.6:35B-A3B")
|
||||
self.assertEqual(call.call_args.kwargs["num_ctx"], 32768)
|
||||
self.assertEqual(result.runtime_metadata["request_count"], 1)
|
||||
self.assertEqual(result.runtime_metadata["prompt_token_count"], 123)
|
||||
self.assertFalse(result.runtime_metadata["think"])
|
||||
self.assertEqual(result.runtime_metadata["temperature"], 0.0)
|
||||
|
||||
def test_valid_context_is_loaded_and_rendered(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = root / "transcript.json"
|
||||
context = root / "context.yaml"
|
||||
write_transcript(transcript)
|
||||
context.write_text(VALID_CONTEXT, encoding="utf-8")
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
context,
|
||||
model_check=Mock(return_value={}),
|
||||
generation_call=Mock(return_value=generation()),
|
||||
)
|
||||
|
||||
self.assertIn("MEETING CONTEXT V1", result.exact_prompt)
|
||||
self.assertIn("Test Meeting", result.exact_prompt)
|
||||
|
||||
def test_invalid_context_is_rejected_before_network_calls(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = root / "transcript.json"
|
||||
context = root / "context.yaml"
|
||||
write_transcript(transcript)
|
||||
context.write_text("schema_version: wrong", encoding="utf-8")
|
||||
check = Mock()
|
||||
call = Mock()
|
||||
with self.assertRaisesRegex(ValueError, "schema_version"):
|
||||
generate_direct_protocol(
|
||||
transcript,
|
||||
context,
|
||||
model_check=check,
|
||||
generation_call=call,
|
||||
)
|
||||
|
||||
check.assert_not_called()
|
||||
call.assert_not_called()
|
||||
|
||||
|
||||
class OllamaTests(unittest.TestCase):
|
||||
def test_unavailable_endpoint_failure(self) -> None:
|
||||
with patch.object(ollama.requests, "get", side_effect=requests.ConnectionError("down")):
|
||||
with self.assertRaisesRegex(OllamaError, "not reachable"):
|
||||
ollama.require_model("http://127.0.0.1:11434", "model")
|
||||
|
||||
def test_missing_model_failure(self) -> None:
|
||||
response = Mock()
|
||||
response.raise_for_status.return_value = None
|
||||
response.json.return_value = {"models": [{"name": "other:model"}]}
|
||||
with patch.object(ollama.requests, "get", return_value=response):
|
||||
with self.assertRaisesRegex(OllamaError, "not installed"):
|
||||
ollama.require_model("http://127.0.0.1:11434", "model")
|
||||
|
||||
def test_request_settings_and_raw_response(self) -> None:
|
||||
raw = {"response": "# Meeting Protocol", "done": True}
|
||||
response = Mock()
|
||||
response.raise_for_status.return_value = None
|
||||
response.json.return_value = raw
|
||||
with patch.object(ollama.requests, "post", return_value=response) as post:
|
||||
result = ollama.generate_once(
|
||||
"http://localhost:11434",
|
||||
"qwen3.8:27b",
|
||||
"prompt",
|
||||
timeout=30,
|
||||
num_ctx=32768,
|
||||
num_predict=8192,
|
||||
)
|
||||
|
||||
self.assertEqual(post.call_count, 1)
|
||||
payload = post.call_args.kwargs["json"]
|
||||
self.assertEqual(payload["model"], "qwen3.8:27b")
|
||||
self.assertEqual(payload["options"]["temperature"], 0.0)
|
||||
self.assertEqual(payload["options"]["num_ctx"], 32768)
|
||||
self.assertFalse(payload["think"])
|
||||
self.assertFalse(payload["stream"])
|
||||
self.assertEqual(result.raw_response, raw)
|
||||
|
||||
def test_malformed_response_failure_without_retry(self) -> None:
|
||||
response = Mock()
|
||||
response.raise_for_status.return_value = None
|
||||
response.json.return_value = {"message": "missing response"}
|
||||
with patch.object(ollama.requests, "post", return_value=response) as post:
|
||||
with self.assertRaisesRegex(OllamaError, "no string 'response'"):
|
||||
ollama.generate_once("url", "model", "prompt", timeout=1, num_ctx=1, num_predict=1)
|
||||
self.assertEqual(post.call_count, 1)
|
||||
|
||||
def test_empty_response_failure(self) -> None:
|
||||
response = Mock()
|
||||
response.raise_for_status.return_value = None
|
||||
response.json.return_value = {"response": " "}
|
||||
with patch.object(ollama.requests, "post", return_value=response):
|
||||
with self.assertRaisesRegex(OllamaError, "empty protocol"):
|
||||
ollama.generate_once("url", "model", "prompt", timeout=1, num_ctx=1, num_predict=1)
|
||||
|
||||
|
||||
class DirectProtocolCliTests(unittest.TestCase):
|
||||
def test_artifacts_are_preserved_and_protocol_is_untouched(self) -> None:
|
||||
protocol_text = "# Meeting Protocol\n\nExact output. \n"
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = root / "source.json"
|
||||
context = root / "source.yaml"
|
||||
write_transcript(transcript)
|
||||
context.write_text(VALID_CONTEXT, encoding="utf-8")
|
||||
args = run_direct_protocol.parse_args(
|
||||
[str(transcript), "--context", str(context), "--output-root", str(root / "runs")]
|
||||
)
|
||||
with patch.object(
|
||||
run_direct_protocol,
|
||||
"generate_direct_protocol",
|
||||
return_value=type("Result", (), {
|
||||
"protocol_text": protocol_text,
|
||||
"exact_prompt": "exact prompt\n",
|
||||
"transcript_input": "selected transcript\n",
|
||||
"raw_response": {"response": protocol_text},
|
||||
"runtime_metadata": {"request_count": 1},
|
||||
})(),
|
||||
) as generator:
|
||||
code, run_dir, protocol_path = run_direct_protocol.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
self.assertEqual(generator.call_count, 1)
|
||||
self.assertEqual(protocol_path.read_text(encoding="utf-8"), protocol_text)
|
||||
self.assertEqual(
|
||||
(run_dir / "protocol/exact_prompt.txt").read_text(encoding="utf-8"),
|
||||
"exact prompt\n",
|
||||
)
|
||||
self.assertEqual(
|
||||
(run_dir / "protocol/transcript_input.txt").read_text(encoding="utf-8"),
|
||||
"selected transcript\n",
|
||||
)
|
||||
self.assertEqual(
|
||||
json.loads((run_dir / "protocol/raw_response.json").read_text())["response"],
|
||||
protocol_text,
|
||||
)
|
||||
self.assertEqual(
|
||||
json.loads((run_dir / "protocol/runtime_metadata.json").read_text())["request_count"],
|
||||
1,
|
||||
)
|
||||
self.assertTrue((run_dir / "transcript/transcript.json").is_file())
|
||||
self.assertTrue((run_dir / "context/meeting_context.yaml").is_file())
|
||||
self.assertTrue((run_dir / "input_manifest.json").is_file())
|
||||
self.assertEqual(json.loads((run_dir / "run_metadata.json").read_text())["status"], "completed")
|
||||
|
||||
def test_unique_run_directories_do_not_overwrite(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
fixed = datetime(2026, 8, 20, 12, 0, 0)
|
||||
first = run_direct_protocol.create_unique_run_dir(root, "meeting", lambda: fixed)
|
||||
marker = first / "keep.txt"
|
||||
marker.write_text("keep", encoding="utf-8")
|
||||
second = run_direct_protocol.create_unique_run_dir(root, "meeting", lambda: fixed)
|
||||
self.assertEqual(first.name, "meeting_20260820_120000")
|
||||
self.assertEqual(second.name, "meeting_20260820_120000_01")
|
||||
self.assertEqual(marker.read_text(encoding="utf-8"), "keep")
|
||||
|
||||
def test_failure_after_directory_creation_preserves_metadata(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = run_direct_protocol.parse_args(
|
||||
[str(root / "missing.json"), "--output-root", str(root / "runs")]
|
||||
)
|
||||
code, run_dir, protocol_path = run_direct_protocol.run(args)
|
||||
|
||||
metadata = json.loads((run_dir / "run_metadata.json").read_text())
|
||||
self.assertEqual(code, 2)
|
||||
self.assertIsNone(protocol_path)
|
||||
self.assertEqual(metadata["status"], "failed")
|
||||
self.assertIn("does not exist", metadata["failure"])
|
||||
|
||||
def test_semantic_pipeline_functions_are_never_invoked(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = root / "source.json"
|
||||
write_transcript(transcript)
|
||||
args = run_direct_protocol.parse_args(
|
||||
[str(transcript), "--output-root", str(root / "runs")]
|
||||
)
|
||||
fake_result = type("Result", (), {
|
||||
"protocol_text": "# Meeting Protocol",
|
||||
"exact_prompt": "prompt",
|
||||
"raw_response": {"response": "# Meeting Protocol"},
|
||||
"runtime_metadata": {},
|
||||
})()
|
||||
with (
|
||||
patch("src.meeting_lab.extraction.extract_chunks.extract_input") as extraction,
|
||||
patch("src.meeting_lab.consolidation.consolidate_facts.call_ollama") as consolidation,
|
||||
patch.object(run_direct_protocol, "generate_direct_protocol", return_value=fake_result),
|
||||
):
|
||||
code, _run_dir, _protocol_path = run_direct_protocol.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
extraction.assert_not_called()
|
||||
consolidation.assert_not_called()
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,168 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.evidence_observations.experiment import (
|
||||
SCHEMA_VERSION,
|
||||
ObservationValidationError,
|
||||
build_ollama_payload,
|
||||
load_fixture,
|
||||
parse_model_json,
|
||||
run_case,
|
||||
validate_observations,
|
||||
)
|
||||
|
||||
|
||||
class EvidenceObservationExperimentTests(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
self.case = {
|
||||
"case_id": "test_case",
|
||||
"description": "Validator fixture.",
|
||||
"subject_id": "subject_test",
|
||||
"subject": "Prüfung der Messdaten",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Nina, prüfst du die Daten?"},
|
||||
{"evidence_id": "e2", "text": "Ja, ich prüfe sie."},
|
||||
],
|
||||
"expected_observations": [],
|
||||
}
|
||||
self.output = {
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"subject_id": "subject_test",
|
||||
"subject": "Prüfung der Messdaten",
|
||||
"observations": [self.observation()],
|
||||
}
|
||||
self.case["expected_observations"] = deepcopy(self.output["observations"])
|
||||
|
||||
def observation(self, **updates):
|
||||
value = {
|
||||
"observation_id": "obs_1",
|
||||
"evidence_id": "e1",
|
||||
"content": "Nina wird um Prüfung gebeten.",
|
||||
"target": "discussion_subject",
|
||||
"relation": "none",
|
||||
"modality": "interpersonal_request",
|
||||
"temporality": "future",
|
||||
"evaluation": "none",
|
||||
"agreement": "none",
|
||||
"responsibility": "named",
|
||||
"person": "Nina",
|
||||
"uncertainty": "absent",
|
||||
"clarification_need": "none",
|
||||
"scope": "absent",
|
||||
}
|
||||
value.update(updates)
|
||||
return value
|
||||
|
||||
def test_valid_observation_and_discussion_subject_target(self):
|
||||
self.assertIs(validate_observations(self.output, self.case), self.output)
|
||||
|
||||
def test_multiple_observations_from_one_evidence_unit_and_observation_target(self):
|
||||
second = self.observation(
|
||||
observation_id="obs_2", target="obs_1", relation="supports"
|
||||
)
|
||||
self.output["observations"].append(second)
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_plural_target_is_allowed_for_joint_reference(self):
|
||||
self.output["observations"].extend(
|
||||
[
|
||||
self.observation(observation_id="obs_2"),
|
||||
self.observation(
|
||||
observation_id="obs_3",
|
||||
target=["obs_1", "obs_2"],
|
||||
relation="qualifies",
|
||||
),
|
||||
]
|
||||
)
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_unknown_evidence_reference_is_rejected(self):
|
||||
self.output["observations"][0]["evidence_id"] = "missing"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "unknown evidence"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_unknown_observation_target_is_rejected(self):
|
||||
self.output["observations"][0]["target"] = "obs_9"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "unknown or later"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_invalid_relation_is_rejected(self):
|
||||
self.output["observations"][0]["relation"] = "causes"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "relation is invalid"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_invalid_modality_is_rejected(self):
|
||||
self.output["observations"][0]["modality"] = "proposal"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "modality is invalid"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_invalid_responsibility_person_combinations_are_rejected(self):
|
||||
self.output["observations"][0].update(responsibility="none", person="Nina")
|
||||
with self.assertRaisesRegex(ObservationValidationError, "person must be JSON null"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
self.output["observations"][0].update(responsibility="accepted", person=None)
|
||||
with self.assertRaisesRegex(ObservationValidationError, "person must be a non-empty"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_scope_uses_absent_or_nonempty_evidence_grounded_text(self):
|
||||
validate_observations(self.output, self.case)
|
||||
self.output["observations"][0]["scope"] = "bis Freitag"
|
||||
validate_observations(self.output, self.case)
|
||||
self.output["observations"][0]["scope"] = None
|
||||
with self.assertRaisesRegex(ObservationValidationError, "non-empty string"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_string_null_is_rejected_in_text_fields(self):
|
||||
self.output["observations"][0]["scope"] = "null"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "string 'null'"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_malformed_model_json_is_rejected(self):
|
||||
with self.assertRaises(json.JSONDecodeError):
|
||||
parse_model_json("{not json")
|
||||
|
||||
def test_payload_has_exact_live_controls(self):
|
||||
payload = build_ollama_payload("qwen3.5:9B", "prompt", 16384, 4096)
|
||||
self.assertIs(payload["think"], False)
|
||||
self.assertIs(payload["stream"], False)
|
||||
self.assertEqual(payload["format"], "json")
|
||||
self.assertEqual(payload["options"]["temperature"], 0)
|
||||
|
||||
def test_fixture_contains_all_nine_cases(self):
|
||||
cases = load_fixture(Path("tests/gold/evidence_observations_v1/cases.json"))
|
||||
self.assertEqual(len(cases), 9)
|
||||
self.assertEqual(cases[0]["case_id"], "a_idea_only")
|
||||
self.assertEqual(cases[-1]["case_id"], "i_outcome_and_unresolved")
|
||||
|
||||
def test_case_run_preserves_all_artifacts(self):
|
||||
raw = json.dumps(self.output, ensure_ascii=False)
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
root = Path(temporary)
|
||||
with patch(
|
||||
"src.meeting_lab.evidence_observations.experiment.call_ollama",
|
||||
return_value=(raw, {"model": "qwen3.5:9B"}),
|
||||
):
|
||||
result = run_case(
|
||||
self.case, root, "http://unused", "qwen3.5:9B", 1, 16384, 4096
|
||||
)
|
||||
self.assertEqual(result["verdict"], "PASS")
|
||||
for filename in (
|
||||
"gold_input.json",
|
||||
"gold_expected_observations.json",
|
||||
"prompt.txt",
|
||||
"raw_model_response.txt",
|
||||
"parsed_observations.json",
|
||||
"validation_result.json",
|
||||
"ollama_metadata.json",
|
||||
"evaluation_result.json",
|
||||
):
|
||||
self.assertTrue((root / "test_case" / filename).is_file(), filename)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,160 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.evidence_observations_v2.experiment import (
|
||||
SCHEMA_VERSION,
|
||||
ObservationValidationError,
|
||||
build_ollama_payload,
|
||||
load_fixture,
|
||||
parse_model_json,
|
||||
run_case,
|
||||
validate_observations,
|
||||
)
|
||||
|
||||
|
||||
class EvidenceObservationV2ExperimentTests(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
self.case = {
|
||||
"case_id": "test_case", "description": "Validator fixture.",
|
||||
"subject_id": "subject_test", "subject": "Prüfung der Messdaten",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Antonius: Nina, prüfst du die Daten?"},
|
||||
{"evidence_id": "e2", "text": "Nina: Ja, ich prüfe sie."},
|
||||
],
|
||||
"expected_observations": [],
|
||||
}
|
||||
self.output = {
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"subject_id": self.case["subject_id"], "subject": self.case["subject"],
|
||||
"observations": [self.observation()],
|
||||
}
|
||||
self.case["expected_observations"] = deepcopy(self.output["observations"])
|
||||
|
||||
def observation(self, **updates):
|
||||
value = {
|
||||
"observation_id": "obs_1", "evidence_id": "e1",
|
||||
"content": "Antonius bittet Nina um eine Prüfung.", "refers_to": None,
|
||||
"speaker": "Antonius", "named_person": "Nina", "addressee": "Nina",
|
||||
"self_reference": False, "collective_we": False,
|
||||
"impersonal_person_reference": False,
|
||||
"modality": "interpersonal_request", "temporality": "future",
|
||||
"evaluation": "none", "affirmation": "absent", "negation": "absent",
|
||||
"determination_statement": "absent", "uncertainty": "absent",
|
||||
"clarification_need": "none", "qualifier": "bis Freitag",
|
||||
"limits_target": None,
|
||||
}
|
||||
value.update(updates)
|
||||
return value
|
||||
|
||||
def test_participant_facts_do_not_include_responsibility(self):
|
||||
validate_observations(self.output, self.case)
|
||||
observation = self.output["observations"][0]
|
||||
self.assertEqual(observation["speaker"], "Antonius")
|
||||
self.assertEqual(observation["named_person"], "Nina")
|
||||
self.assertEqual(observation["addressee"], "Nina")
|
||||
self.assertNotIn("responsibility", observation)
|
||||
|
||||
def test_named_person_and_speaker_do_not_imply_any_extra_field(self):
|
||||
keys = self.output["observations"][0].keys()
|
||||
self.assertNotIn("person", keys)
|
||||
self.assertNotIn("agreement", keys)
|
||||
|
||||
def test_self_reference_collective_we_and_impersonal_reference_are_boolean(self):
|
||||
self.output["observations"][0].update(
|
||||
self_reference=True, collective_we=True, impersonal_person_reference=True
|
||||
)
|
||||
validate_observations(self.output, self.case)
|
||||
self.output["observations"][0]["collective_we"] = "true"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "must be boolean"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_explicit_affirmation_negation_and_determination(self):
|
||||
self.output["observations"][0].update(
|
||||
affirmation="explicit", negation="explicit", determination_statement="present"
|
||||
)
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_scalar_reference_to_prior_observation(self):
|
||||
self.output["observations"].append(self.observation(
|
||||
observation_id="obs_2", evidence_id="e2", refers_to="obs_1",
|
||||
speaker="Nina", named_person=None, addressee=None,
|
||||
))
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_array_and_invalid_reference_are_rejected(self):
|
||||
self.output["observations"][0]["refers_to"] = ["obs_1"]
|
||||
with self.assertRaisesRegex(ObservationValidationError, "non-empty string"):
|
||||
validate_observations(self.output, self.case)
|
||||
self.output["observations"][0]["refers_to"] = "obs_9"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "unknown or later"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_qualifier_is_null_or_nonempty_text(self):
|
||||
self.output["observations"][0]["qualifier"] = None
|
||||
validate_observations(self.output, self.case)
|
||||
self.output["observations"][0]["qualifier"] = ""
|
||||
with self.assertRaisesRegex(ObservationValidationError, "non-empty string"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_limits_target_must_reference_prior_observation(self):
|
||||
self.output["observations"].append(self.observation(
|
||||
observation_id="obs_2", evidence_id="e2", refers_to="obs_1",
|
||||
limits_target="obs_1", speaker="Nina", named_person=None, addressee=None,
|
||||
))
|
||||
validate_observations(self.output, self.case)
|
||||
self.output["observations"][1]["limits_target"] = "obs_7"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "unknown or later"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_multiple_atomic_observations_may_share_evidence(self):
|
||||
self.output["observations"].append(self.observation(observation_id="obs_2"))
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_string_null_is_rejected(self):
|
||||
self.output["observations"][0]["named_person"] = "null"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "string 'null'"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_malformed_json_is_rejected(self):
|
||||
with self.assertRaises(json.JSONDecodeError):
|
||||
parse_model_json("{not json")
|
||||
|
||||
def test_payload_has_exact_live_controls(self):
|
||||
payload = build_ollama_payload("qwen3.5:9B", "prompt", 16384, 4096)
|
||||
self.assertFalse(payload["think"])
|
||||
self.assertFalse(payload["stream"])
|
||||
self.assertEqual(payload["options"]["temperature"], 0)
|
||||
|
||||
def test_fixture_contains_unchanged_a_i_source_evidence(self):
|
||||
v1 = load_fixture(Path("tests/gold/evidence_observations_v2/cases.json"))
|
||||
original = json.loads(Path("tests/gold/evidence_observations_v1/cases.json").read_text())["cases"]
|
||||
self.assertEqual(len(v1), 9)
|
||||
self.assertEqual(
|
||||
[[item["text"] for item in case["evidence"]] for case in v1],
|
||||
[[item["text"] for item in case["evidence"]] for case in original],
|
||||
)
|
||||
|
||||
def test_case_run_preserves_all_artifacts(self):
|
||||
raw = json.dumps(self.output, ensure_ascii=False)
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
root = Path(temporary)
|
||||
with patch(
|
||||
"src.meeting_lab.evidence_observations_v2.experiment.call_ollama",
|
||||
return_value=(raw, {"model": "qwen3.5:9B"}),
|
||||
):
|
||||
result = run_case(self.case, root, "http://unused", "qwen3.5:9B", 1, 16384, 4096)
|
||||
self.assertEqual(result["verdict"], "PASS")
|
||||
for filename in (
|
||||
"gold_input.json", "gold_expected_observations.json", "prompt.txt",
|
||||
"raw_model_response.txt", "parsed_observations.json",
|
||||
"validation_result.json", "ollama_metadata.json", "evaluation_result.json",
|
||||
):
|
||||
self.assertTrue((root / "test_case" / filename).is_file(), filename)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,114 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.evidence_observations_v3.experiment import (
|
||||
SCHEMA_VERSION,
|
||||
ObservationValidationError,
|
||||
build_ollama_payload,
|
||||
load_fixture,
|
||||
parse_model_json,
|
||||
run_case,
|
||||
validate_observations,
|
||||
)
|
||||
|
||||
|
||||
class EvidenceObservationV3ExperimentTests(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
self.case = {
|
||||
"case_id": "test", "description": "Minimal fixture.",
|
||||
"subject_id": "subject_test", "subject": "Messdatenprüfung",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Antonius: Nina, prüfst du die Messdaten?"},
|
||||
{"evidence_id": "e2", "text": "Nina: Ja, ich prüfe sie bis Freitag."},
|
||||
],
|
||||
"semantic_requirements": ["Request and response survive."],
|
||||
}
|
||||
self.output = {
|
||||
"schema_version": SCHEMA_VERSION, "subject_id": "subject_test",
|
||||
"subject": "Messdatenprüfung", "observations": [self.observation()],
|
||||
}
|
||||
|
||||
def observation(self, **updates):
|
||||
value = {
|
||||
"observation_id": "obs_1", "evidence_id": "e1",
|
||||
"content": "Antonius fragt Nina, ob sie die Messdaten prüft.",
|
||||
"speaker": "Antonius", "named_person": "Nina", "addressee": "Nina",
|
||||
}
|
||||
value.update(updates)
|
||||
return value
|
||||
|
||||
def test_minimal_schema_is_valid(self):
|
||||
self.assertIs(validate_observations(self.output, self.case), self.output)
|
||||
self.assertEqual(set(self.output["observations"][0]), {
|
||||
"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"
|
||||
})
|
||||
|
||||
def test_unknown_semantic_field_is_rejected(self):
|
||||
self.output["observations"][0]["modality"] = "factual"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "unknown keys"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_unknown_evidence_is_rejected(self):
|
||||
self.output["observations"][0]["evidence_id"] = "e9"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "unknown evidence"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_speaker_must_match_evidence(self):
|
||||
self.output["observations"][0]["speaker"] = "Nina"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "match evidence speaker"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_named_person_does_not_add_responsibility(self):
|
||||
validate_observations(self.output, self.case)
|
||||
self.assertNotIn("responsibility", self.output["observations"][0])
|
||||
|
||||
def test_addressee_does_not_add_assignment(self):
|
||||
validate_observations(self.output, self.case)
|
||||
self.assertNotIn("action_item", self.output["observations"][0])
|
||||
|
||||
def test_nonexplicit_person_is_rejected(self):
|
||||
self.output["observations"][0]["named_person"] = "Martin"
|
||||
with self.assertRaisesRegex(ObservationValidationError, "not an explicit person"):
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_multiple_atomic_observations_can_share_evidence(self):
|
||||
self.output["observations"].append(self.observation(observation_id="obs_2"))
|
||||
validate_observations(self.output, self.case)
|
||||
|
||||
def test_malformed_json_is_rejected(self):
|
||||
with self.assertRaises(json.JSONDecodeError):
|
||||
parse_model_json("{bad json")
|
||||
|
||||
def test_payload_controls_are_fixed(self):
|
||||
payload = build_ollama_payload("qwen3.5:9B", "prompt", 16384, 4096)
|
||||
self.assertFalse(payload["think"])
|
||||
self.assertEqual(payload["options"]["temperature"], 0)
|
||||
|
||||
def test_fixture_reuses_exact_v2_evidence(self):
|
||||
v3 = load_fixture(Path("tests/gold/evidence_observations_v3/cases.json"))
|
||||
v2 = json.loads(Path("tests/gold/evidence_observations_v2/cases.json").read_text())["cases"]
|
||||
self.assertEqual([case["evidence"] for case in v3], [case["evidence"] for case in v2])
|
||||
|
||||
def test_case_run_preserves_persistent_artifact_set(self):
|
||||
raw = json.dumps(self.output, ensure_ascii=False)
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
root = Path(temporary)
|
||||
with patch(
|
||||
"src.meeting_lab.evidence_observations_v3.experiment.call_ollama",
|
||||
return_value=(raw, {"model": "qwen3.5:9B"}),
|
||||
):
|
||||
result = run_case(self.case, root, "http://unused", "qwen3.5:9B", 1, 16384, 4096)
|
||||
self.assertTrue(result["structurally_valid"])
|
||||
for filename in (
|
||||
"source_evidence.json", "gold_semantic_requirements.json", "prompt.txt",
|
||||
"raw_model_response.txt", "parsed_observations.json",
|
||||
"structural_validation.json", "ollama_metadata.json",
|
||||
):
|
||||
self.assertTrue((root / "test" / filename).is_file(), filename)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,210 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import (
|
||||
DerivationValidationError,
|
||||
build_prompt,
|
||||
derive_rejection,
|
||||
evaluate_case,
|
||||
load_gold_cases,
|
||||
validate_recognition,
|
||||
)
|
||||
|
||||
|
||||
GOLD_PATH = Path("tests/gold/explicit_rejection_v0/cases.json")
|
||||
|
||||
|
||||
POSITIVE_TEXT = {
|
||||
"RJ-01": "reale Anlage für den Versuch nutzen",
|
||||
"RJ-02": "externe Lösung weiterverfolgen",
|
||||
"RJ-03": "Zusammenarbeit mit Dr. Schlummer fortsetzen",
|
||||
"RJ-11": "reale Anlage für den Druckversuch nutzen",
|
||||
"RJ-12": "Versuch in der realen Anlage durchführen",
|
||||
}
|
||||
|
||||
|
||||
def recognition_for(case):
|
||||
expected = case["expected_recognition"]
|
||||
positive = expected["rejection_form"] == "explicit_action_rejection"
|
||||
return {
|
||||
"rejection_observation_id": expected["rejection_observation_id"],
|
||||
"target_observation_id": expected["target_observation_id"] if positive else None,
|
||||
"rejection_form": expected["rejection_form"],
|
||||
"normalized_rejected_action_text": POSITIVE_TEXT.get(case["case_id"]) if positive else None,
|
||||
}
|
||||
|
||||
|
||||
class ExplicitRejectionGoldExperimentTests(unittest.TestCase):
|
||||
@classmethod
|
||||
def setUpClass(cls):
|
||||
cls.cases = load_gold_cases(GOLD_PATH)
|
||||
cls.by_id = {case["case_id"]: case for case in cls.cases}
|
||||
|
||||
def test_fixture_contains_exactly_rj_01_through_rj_12(self):
|
||||
self.assertEqual(list(self.by_id), [f"RJ-{number:02d}" for number in range(1, 13)])
|
||||
|
||||
def test_cases_use_only_minimal_v3_style_observations(self):
|
||||
keys = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
|
||||
for case in self.cases:
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertIn(len(case["observations"]), (1, 2))
|
||||
self.assertTrue(all(set(item) == keys for item in case["observations"]))
|
||||
|
||||
def test_rj_01_derives_target_and_both_provenance_paths(self):
|
||||
case = self.by_id["RJ-01"]
|
||||
gates, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
self.assertTrue(all(gates.values()))
|
||||
self.assertEqual(result["status"], "explicitly_rejected")
|
||||
self.assertEqual(result["support"]["target"], {"observation_id": "obs_1", "evidence_id": "e1"})
|
||||
self.assertEqual(result["support"]["rejection"], {"observation_id": "obs_2", "evidence_id": "e2"})
|
||||
|
||||
def test_rj_02_requires_paired_target_and_derives_abandonment(self):
|
||||
case = self.by_id["RJ-02"]
|
||||
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
self.assertIn("externe Lösung", result["content"])
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown target"):
|
||||
derive_rejection(case["observations"][1:], recognition_for(case))
|
||||
|
||||
def test_rj_03_supports_same_observation_target_and_rejection(self):
|
||||
case = self.by_id["RJ-03"]
|
||||
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
self.assertEqual(result["support"]["target"], result["support"]["rejection"])
|
||||
self.assertIn("Dr. Schlummer", result["content"])
|
||||
|
||||
def test_all_required_negative_cases_remain_non_rejections(self):
|
||||
for case_id in ("RJ-04", "RJ-05", "RJ-06", "RJ-07", "RJ-08", "RJ-09", "RJ-10"):
|
||||
case = self.by_id[case_id]
|
||||
gates, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
with self.subTest(case=case_id):
|
||||
self.assertFalse(gates["explicit_action_rejection"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_rj_11_derives_and_preserves_location_and_purpose_scope(self):
|
||||
case = self.by_id["RJ-11"]
|
||||
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
self.assertIsNotNone(result)
|
||||
self.assertIn("reale Anlage", result["content"])
|
||||
self.assertIn("Druckversuch", result["content"])
|
||||
|
||||
def test_rj_12_rejects_only_real_plant_action_and_not_alternative(self):
|
||||
case = self.by_id["RJ-12"]
|
||||
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
self.assertIn("realen Anlage", result["content"])
|
||||
self.assertNotIn("Technikum", result["content"])
|
||||
|
||||
def test_separate_target_cannot_follow_rejection(self):
|
||||
case = deepcopy(self.by_id["RJ-01"])
|
||||
case["observations"].reverse()
|
||||
gates, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
self.assertFalse(gates["target_same_or_before_rejection"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_unknown_target_observation_id_is_rejected(self):
|
||||
case = self.by_id["RJ-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["target_observation_id"] = "obs_99"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown target"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_unknown_rejection_observation_id_is_rejected(self):
|
||||
case = self.by_id["RJ-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["rejection_observation_id"] = "obs_99"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown rejection"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_duplicate_observation_ids_are_rejected(self):
|
||||
fixture = json.loads(GOLD_PATH.read_text())
|
||||
fixture["cases"][0]["observations"][1]["observation_id"] = "obs_1"
|
||||
self._assert_bad_fixture(fixture, "observation IDs must be unique")
|
||||
|
||||
def test_inconsistent_evidence_provenance_is_rejected(self):
|
||||
fixture = json.loads(GOLD_PATH.read_text())
|
||||
fixture["cases"][0]["observations"][1]["evidence_id"] = "e1"
|
||||
self._assert_bad_fixture(fixture, "evidence provenance must be unique")
|
||||
|
||||
def test_none_rejects_populated_target_or_action(self):
|
||||
case = self.by_id["RJ-04"]
|
||||
for field, value, message in (
|
||||
("target_observation_id", "obs_1", "null target"),
|
||||
("normalized_rejected_action_text", "Anlage nutzen", "null normalized"),
|
||||
):
|
||||
recognition = recognition_for(case)
|
||||
recognition[field] = value
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, message):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_explicit_rejection_requires_normalized_target_text(self):
|
||||
case = self.by_id["RJ-01"]
|
||||
for value in (None, ""):
|
||||
recognition = recognition_for(case)
|
||||
recognition["normalized_rejected_action_text"] = value
|
||||
with self.subTest(value=value), self.assertRaises(DerivationValidationError):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_unknown_schema_fields_are_rejected(self):
|
||||
case = self.by_id["RJ-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["explanation"] = "extra"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown keys"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_forbidden_normative_fields_are_rejected_recursively(self):
|
||||
case = self.by_id["RJ-01"]
|
||||
fields = (
|
||||
"decision", "decision_status", "outcome", "topic_status", "closed",
|
||||
"agreement", "responsible_person", "responsibility", "responsibility_scope",
|
||||
"owner", "ownership", "assignee", "requested_actor", "status",
|
||||
"explicitly_rejected", "action_item", "protocol_category", "confidence",
|
||||
"relation", "relations", "graph", "unresolved_issue",
|
||||
)
|
||||
for field in fields:
|
||||
recognition = recognition_for(case)
|
||||
recognition["wrapper"] = {field: "forbidden"}
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_speaker_identity_creates_no_ownership_or_responsibility(self):
|
||||
case = deepcopy(self.by_id["RJ-03"])
|
||||
for speaker in ("Martin", "Clara", "Antonius"):
|
||||
case["observations"][0]["speaker"] = speaker
|
||||
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||
with self.subTest(speaker=speaker):
|
||||
self.assertNotIn("responsible_person", result)
|
||||
self.assertNotIn("owner", result)
|
||||
|
||||
def test_all_expected_recognitions_evaluate_as_pass(self):
|
||||
for case in self.cases:
|
||||
evaluation = evaluate_case(case, recognition_for(case))
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertEqual(evaluation["classification"], "PASS")
|
||||
|
||||
def test_rj_11_qualifier_loss_and_rj_12_alternative_absorption_fail(self):
|
||||
rj11 = self.by_id["RJ-11"]
|
||||
recognition = recognition_for(rj11)
|
||||
recognition["normalized_rejected_action_text"] = "reale Anlage nutzen"
|
||||
self.assertEqual(evaluate_case(rj11, recognition)["classification"], "FAIL")
|
||||
rj12 = self.by_id["RJ-12"]
|
||||
recognition = recognition_for(rj12)
|
||||
recognition["normalized_rejected_action_text"] += "; stattdessen im Technikum testen"
|
||||
self.assertEqual(evaluate_case(rj12, recognition)["classification"], "FAIL")
|
||||
|
||||
def test_prompt_is_fixed_narrow_and_does_not_expose_gold_expectation(self):
|
||||
prompt = build_prompt(self.by_id["RJ-01"])
|
||||
self.assertIn("candidate rejection observation is obs_2", prompt)
|
||||
self.assertNotIn("expected_result", prompt)
|
||||
self.assertNotIn("Who is responsible", prompt)
|
||||
|
||||
def _assert_bad_fixture(self, fixture, message):
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
path = Path(temporary) / "cases.json"
|
||||
path.write_text(json.dumps(fixture), encoding="utf-8")
|
||||
with self.assertRaisesRegex(DerivationValidationError, message):
|
||||
load_gold_cases(path)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -11,8 +11,10 @@ from src.meeting_lab.extraction.extract_chunks import (
|
||||
)
|
||||
from src.meeting_lab.models.meeting_context import (
|
||||
MeetingContextValidationError,
|
||||
create_meeting_context,
|
||||
load_meeting_context,
|
||||
render_meeting_context_for_prompt,
|
||||
serialize_meeting_context_yaml,
|
||||
validate_meeting_context,
|
||||
)
|
||||
|
||||
@@ -70,6 +72,38 @@ class MeetingContextTests(unittest.TestCase):
|
||||
with self.assertRaisesRegex(MeetingContextValidationError, "invalid value"):
|
||||
validate_meeting_context(data)
|
||||
|
||||
def test_missing_participant_attendance_defaults_to_present(self) -> None:
|
||||
data = copy.deepcopy(self.context.data)
|
||||
del data["participants"][0]["attendance_status"]
|
||||
|
||||
context = create_meeting_context(data)
|
||||
|
||||
self.assertEqual(
|
||||
context.data["participants"][0]["attendance_status"], "present"
|
||||
)
|
||||
|
||||
def test_explicit_mentioned_only_is_preserved(self) -> None:
|
||||
data = copy.deepcopy(self.context.data)
|
||||
data["mentioned_people"][0]["attendance_status"] = "mentioned_only"
|
||||
|
||||
context = create_meeting_context(data)
|
||||
|
||||
self.assertEqual(
|
||||
context.data["mentioned_people"][0]["attendance_status"],
|
||||
"mentioned_only",
|
||||
)
|
||||
self.assertIn(
|
||||
"Mentioned but absent people:", render_meeting_context_for_prompt(context)
|
||||
)
|
||||
|
||||
def test_mentioned_only_person_cannot_be_a_diarized_speaker(self) -> None:
|
||||
data = copy.deepcopy(self.context.data)
|
||||
mentioned_id = data["mentioned_people"][0]["person_id"]
|
||||
data["speaker_mappings"] = {"SPEAKER_00": mentioned_id}
|
||||
|
||||
with self.assertRaisesRegex(MeetingContextValidationError, "unknown participant"):
|
||||
validate_meeting_context(data)
|
||||
|
||||
def test_prompt_representation_is_deterministic(self) -> None:
|
||||
first = render_meeting_context_for_prompt(self.context)
|
||||
second = render_meeting_context_for_prompt(self.context)
|
||||
@@ -192,6 +226,51 @@ class MeetingContextTests(unittest.TestCase):
|
||||
self.assertNotIn("responsible: Björn", prompt_context)
|
||||
self.assertNotIn("responsible: Jovana", prompt_context)
|
||||
|
||||
def test_existing_context_without_speaker_mappings_remains_valid(self) -> None:
|
||||
self.assertEqual(self.context.speaker_mappings, {})
|
||||
self.assertIsNone(self.context.participant_for_speaker("SPEAKER_00"))
|
||||
|
||||
def test_explicit_speaker_mapping_is_authoritative(self) -> None:
|
||||
data = copy.deepcopy(self.context.data)
|
||||
participant = data["participants"][0]
|
||||
data["speaker_mappings"] = {"SPEAKER_03": participant["participant_id"]}
|
||||
|
||||
context = create_meeting_context(data)
|
||||
rendered = render_meeting_context_for_prompt(context)
|
||||
|
||||
self.assertEqual(
|
||||
context.participant_for_speaker("SPEAKER_03")["participant_id"],
|
||||
participant["participant_id"],
|
||||
)
|
||||
self.assertIsNone(context.participant_for_speaker("SPEAKER_04"))
|
||||
self.assertIn("Confirmed diarization speaker mappings (authoritative)", rendered)
|
||||
self.assertIn("Unmapped SPEAKER_XX labels must remain anonymous", rendered)
|
||||
|
||||
def test_speaker_mapping_must_reference_existing_participant(self) -> None:
|
||||
data = copy.deepcopy(self.context.data)
|
||||
data["speaker_mappings"] = {"SPEAKER_00": "unknown-person"}
|
||||
with self.assertRaisesRegex(MeetingContextValidationError, "unknown participant"):
|
||||
validate_meeting_context(data)
|
||||
|
||||
def test_speaker_mapping_label_must_use_pyannote_shape(self) -> None:
|
||||
data = copy.deepcopy(self.context.data)
|
||||
data["speaker_mappings"] = {
|
||||
"Martin": data["participants"][0]["participant_id"]
|
||||
}
|
||||
with self.assertRaisesRegex(MeetingContextValidationError, "speaker label"):
|
||||
validate_meeting_context(data)
|
||||
|
||||
def test_generated_context_yaml_is_deterministic_and_round_trips(self) -> None:
|
||||
first = serialize_meeting_context_yaml(self.context)
|
||||
second = serialize_meeting_context_yaml(self.context)
|
||||
self.assertEqual(first, second)
|
||||
|
||||
SCRATCH_DIR.mkdir(exist_ok=True)
|
||||
path = SCRATCH_DIR / "generated_context.yaml"
|
||||
path.write_text(first, encoding="utf-8")
|
||||
loaded = load_meeting_context(path)
|
||||
self.assertEqual(loaded.data, self.context.data)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
|
||||
@@ -0,0 +1,310 @@
|
||||
import json
|
||||
import shutil
|
||||
import subprocess
|
||||
import tempfile
|
||||
import unittest
|
||||
from dataclasses import replace
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from scripts import run_mvp_meeting as cli
|
||||
from src.meeting_lab.audio import PreparedAudio
|
||||
from src.meeting_lab.models.meeting_context import load_meeting_context
|
||||
from src.meeting_lab.orchestration import mvp as mvp_api
|
||||
from src.meeting_lab.orchestration.mvp import MvpMeetingConfig, MvpRunResult
|
||||
from src.meeting_lab.protocol.generate_direct_protocol import DirectProtocolResult
|
||||
from src.meeting_lab.transcription.whisper import TranscriptionError, TranscriptionResult
|
||||
|
||||
|
||||
def context_data():
|
||||
return {
|
||||
"schema_version": "1",
|
||||
"meeting": {
|
||||
"meeting_id": "programmatic-test",
|
||||
"title": "Programmatic Test",
|
||||
"language": "de",
|
||||
"date": None,
|
||||
"objective": "API prüfen",
|
||||
"notes": "",
|
||||
},
|
||||
"participants": [
|
||||
{
|
||||
"participant_id": "person-1",
|
||||
"display_name": "Test Person",
|
||||
"aliases": [],
|
||||
"role": "Projektleitung",
|
||||
"department": None,
|
||||
"attendance_status": "present",
|
||||
"notes": None,
|
||||
}
|
||||
],
|
||||
"speaker_mappings": {"SPEAKER_00": "person-1"},
|
||||
"mentioned_people": [],
|
||||
"organization": {"name": "Example", "departments": []},
|
||||
"known_entities": {},
|
||||
"context_rules": {"do_not_infer_responsibilities": True},
|
||||
}
|
||||
|
||||
|
||||
def fake_transcribe(audio, model, output, language, **kwargs):
|
||||
output.mkdir(parents=True, exist_ok=True)
|
||||
raw = output / "whisper_raw.json"
|
||||
transcript = output / "transcript.json"
|
||||
text = output / "transcript.txt"
|
||||
metadata = output / "runtime_metadata.json"
|
||||
raw.write_text('{"transcription": []}\n', encoding="utf-8")
|
||||
transcript.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"text": "Ein kurzer Besprechungstext.",
|
||||
"segments": [
|
||||
{
|
||||
"id": 0,
|
||||
"start": 0.0,
|
||||
"end": 1.0,
|
||||
"text": "Ein kurzer Besprechungstext.",
|
||||
}
|
||||
],
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
text.write_text("Ein kurzer Besprechungstext.\n", encoding="utf-8")
|
||||
metadata.write_text('{"runtime_seconds": 0.1}\n', encoding="utf-8")
|
||||
return TranscriptionResult(output, raw, transcript, text, metadata, 0.1)
|
||||
|
||||
|
||||
def fake_prepare(source, destination, **kwargs):
|
||||
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||
shutil.copyfile(source, destination)
|
||||
return PreparedAudio(source, source.suffix.removeprefix("."), destination, "ffmpeg", "ffmpeg")
|
||||
|
||||
|
||||
def fake_protocol(transcript, context, **kwargs):
|
||||
rendered_context = load_meeting_context(context)
|
||||
assert rendered_context.meeting_id == "programmatic-test"
|
||||
return DirectProtocolResult(
|
||||
protocol_text="# Meeting Protocol\n",
|
||||
exact_prompt="prompt",
|
||||
model_metadata={"model": kwargs["model"]},
|
||||
runtime_metadata={"request_count": 1},
|
||||
raw_response={"response": "# Meeting Protocol\n"},
|
||||
)
|
||||
|
||||
|
||||
class MvpApiTests(unittest.TestCase):
|
||||
def config(self, root: Path):
|
||||
audio = root / "meeting.wav"
|
||||
model = root / "model.bin"
|
||||
audio.write_bytes(b"audio")
|
||||
model.write_bytes(b"model")
|
||||
return MvpMeetingConfig(
|
||||
audio_file=audio,
|
||||
whisper_model=model,
|
||||
output_root=root / "runs",
|
||||
model="test:model",
|
||||
)
|
||||
|
||||
def test_programmatic_context_is_persisted_without_source_yaml(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
config = self.config(root)
|
||||
events = []
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe),
|
||||
patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare),
|
||||
patch.object(
|
||||
mvp_api, "generate_direct_protocol", side_effect=fake_protocol
|
||||
) as protocol_generator,
|
||||
patch.object(subprocess, "run") as subprocess_run,
|
||||
):
|
||||
result = mvp_api.run_mvp_meeting(
|
||||
config, meeting_context=context_data(), progress_sink=events.append
|
||||
)
|
||||
|
||||
self.assertEqual(result.exit_code, 0)
|
||||
context_path = result.run_dir / "context/meeting_context.yaml"
|
||||
self.assertTrue(context_path.is_file())
|
||||
persisted = load_meeting_context(context_path)
|
||||
self.assertEqual(persisted.meeting_id, "programmatic-test")
|
||||
self.assertEqual(persisted.speaker_mappings, {"SPEAKER_00": "person-1"})
|
||||
subprocess_run.assert_not_called()
|
||||
self.assertEqual(
|
||||
[(event.stage, event.status) for event in events],
|
||||
[
|
||||
("preparing", "started"),
|
||||
("preparing", "completed"),
|
||||
("transcription", "started"),
|
||||
("transcription", "completed"),
|
||||
("protocol_generation", "started"),
|
||||
("protocol_generation", "completed"),
|
||||
("completed", "completed"),
|
||||
],
|
||||
)
|
||||
self.assertTrue(all(event.progress is None for event in events))
|
||||
self.assertEqual(protocol_generator.call_args.kwargs["num_ctx"], 32_768)
|
||||
self.assertEqual(
|
||||
protocol_generator.call_args.kwargs["safe_input_token_budget"],
|
||||
29_000,
|
||||
)
|
||||
|
||||
def test_protocol_only_regeneration_reuses_diarized_artifacts(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
run_dir = root / "existing-run"
|
||||
diarization_dir = run_dir / "diarization"
|
||||
diarization_dir.mkdir(parents=True)
|
||||
source = diarization_dir / "transcript_diarized.json"
|
||||
source.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"text": "SPEAKER_00: Existing statement.\n",
|
||||
"segments": [
|
||||
{
|
||||
"start": 0.0,
|
||||
"end": 1.0,
|
||||
"speaker_id": "SPEAKER_00",
|
||||
"text": "Existing statement.",
|
||||
}
|
||||
],
|
||||
"speaker_labels_anonymous": True,
|
||||
}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
source_before = source.read_bytes()
|
||||
mapped_context = context_data()
|
||||
mapped_context["speaker_mappings"] = {"SPEAKER_00": "person-1"}
|
||||
with (
|
||||
patch.object(mvp_api, "prepare_audio") as preparation,
|
||||
patch.object(mvp_api, "transcribe_audio") as transcription,
|
||||
patch.object(mvp_api, "diarize_audio") as diarization,
|
||||
patch.object(
|
||||
mvp_api, "generate_direct_protocol", side_effect=fake_protocol
|
||||
) as protocol,
|
||||
):
|
||||
result = mvp_api.regenerate_mvp_protocol(
|
||||
run_dir,
|
||||
meeting_context=mapped_context,
|
||||
model="qwen3.8:27b",
|
||||
protocol_num_ctx=32_768,
|
||||
protocol_safe_input_token_budget=29_000,
|
||||
)
|
||||
|
||||
self.assertEqual(result.exit_code, 0)
|
||||
self.assertEqual(result.protocol_path, run_dir / "protocol.md")
|
||||
preparation.assert_not_called()
|
||||
transcription.assert_not_called()
|
||||
diarization.assert_not_called()
|
||||
self.assertEqual(protocol.call_args.args[0], source)
|
||||
self.assertEqual(protocol.call_args.kwargs["num_ctx"], 32_768)
|
||||
self.assertEqual(source.read_bytes(), source_before)
|
||||
persisted = load_meeting_context(run_dir / "context/meeting_context.yaml")
|
||||
self.assertEqual(persisted.speaker_mappings, {"SPEAKER_00": "person-1"})
|
||||
|
||||
def test_failure_emits_terminal_failure_event(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
config = self.config(root)
|
||||
events = []
|
||||
with patch.object(
|
||||
mvp_api,
|
||||
"transcribe_audio",
|
||||
side_effect=TranscriptionError("stopped"),
|
||||
), patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare):
|
||||
result = mvp_api.run_mvp_meeting(config, progress_sink=events.append)
|
||||
|
||||
self.assertEqual(result.exit_code, 2)
|
||||
self.assertEqual(events[-1].stage, "failed")
|
||||
self.assertEqual(events[-1].status, "failed")
|
||||
self.assertIn("transcription", events[-1].message)
|
||||
metadata = json.loads((result.run_dir / "run_metadata.json").read_text())
|
||||
self.assertEqual(metadata["status"], "failed")
|
||||
|
||||
def test_cli_defaults_and_wrapper_delegate_without_subprocess(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
config = self.config(root)
|
||||
args = cli.parse_args(
|
||||
[str(config.audio_file), "--whisper-model", str(config.whisper_model)]
|
||||
)
|
||||
expected = MvpRunResult(0, root / "run", root / "run/protocol.md")
|
||||
with patch.object(cli, "run_mvp_meeting", return_value=expected) as api:
|
||||
actual = cli.run(args, context_override=context_data())
|
||||
|
||||
self.assertEqual(actual, (0, expected.run_dir, expected.protocol_path))
|
||||
delegated = api.call_args.args[0]
|
||||
self.assertEqual(delegated.diarization, "off")
|
||||
self.assertEqual(delegated.language, "de")
|
||||
self.assertEqual(delegated.whisper_executable, "whisper-cli")
|
||||
self.assertEqual(delegated.ffmpeg_executable, "ffmpeg")
|
||||
self.assertTrue(delegated.audio_normalization)
|
||||
self.assertEqual(delegated.protocol_num_ctx, 32_768)
|
||||
self.assertEqual(delegated.protocol_safe_input_token_budget, 29_000)
|
||||
self.assertEqual(api.call_args.kwargs["meeting_context"], context_data())
|
||||
|
||||
def test_cli_explicit_audio_normalization_values_are_propagated(self):
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
config = self.config(root)
|
||||
enabled = cli.config_from_args(
|
||||
cli.parse_args(
|
||||
[
|
||||
str(config.audio_file),
|
||||
"--whisper-model",
|
||||
str(config.whisper_model),
|
||||
"--audio-normalization",
|
||||
]
|
||||
)
|
||||
)
|
||||
disabled = cli.config_from_args(
|
||||
cli.parse_args(
|
||||
[
|
||||
str(config.audio_file),
|
||||
"--whisper-model",
|
||||
str(config.whisper_model),
|
||||
"--no-audio-normalization",
|
||||
]
|
||||
)
|
||||
)
|
||||
|
||||
self.assertTrue(enabled.audio_normalization)
|
||||
self.assertFalse(disabled.audio_normalization)
|
||||
|
||||
def test_transcription_receives_prepared_wav_for_encoded_inputs(self):
|
||||
for suffix in (".flac", ".m4a"):
|
||||
with self.subTest(suffix=suffix), tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
config = self.config(root)
|
||||
encoded = config.audio_file.with_suffix(suffix)
|
||||
config.audio_file.rename(encoded)
|
||||
config = replace(config, audio_file=encoded)
|
||||
received = []
|
||||
|
||||
def capture_transcribe(audio, *args, received_paths=received, **kwargs):
|
||||
received_paths.append(audio)
|
||||
return fake_transcribe(audio, *args, **kwargs)
|
||||
|
||||
with (
|
||||
patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare),
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=capture_transcribe),
|
||||
patch.object(mvp_api, "generate_direct_protocol", side_effect=fake_protocol),
|
||||
):
|
||||
result = mvp_api.run_mvp_meeting(
|
||||
config, meeting_context=context_data()
|
||||
)
|
||||
|
||||
self.assertEqual(result.exit_code, 0)
|
||||
self.assertEqual(received, [result.run_dir / "audio" / "prepared.wav"])
|
||||
manifest = json.loads(
|
||||
(result.run_dir / "audio" / "input_manifest.json").read_text()
|
||||
)
|
||||
self.assertEqual(manifest["format"], suffix.removeprefix("."))
|
||||
self.assertEqual(
|
||||
manifest["prepared_audio"]["prepared_audio_path"],
|
||||
str((result.run_dir / "audio" / "prepared.wav").resolve()),
|
||||
)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,444 @@
|
||||
import json
|
||||
import shutil
|
||||
import tempfile
|
||||
import unittest
|
||||
from datetime import datetime
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
from scripts import run_mvp_meeting
|
||||
from src.meeting_lab.audio import PreparedAudio
|
||||
from src.meeting_lab.audio.preparation import (
|
||||
DEFAULT_NORMALIZATION_FILTER,
|
||||
DEFAULT_NORMALIZATION_METHOD,
|
||||
)
|
||||
from src.meeting_lab.orchestration import mvp as mvp_api
|
||||
from src.meeting_lab.protocol.generate_direct_protocol import DirectProtocolResult
|
||||
from src.meeting_lab.diarization.backend import DiarizationResult
|
||||
from src.meeting_lab.transcription.whisper import TranscriptionError, TranscriptionResult
|
||||
|
||||
|
||||
VALID_CONTEXT = """schema_version: "1"
|
||||
meeting:
|
||||
meeting_id: "mvp-test"
|
||||
title: "MVP Test"
|
||||
language: "de"
|
||||
participants: []
|
||||
mentioned_people: []
|
||||
organization:
|
||||
departments: []
|
||||
known_entities: {}
|
||||
"""
|
||||
|
||||
|
||||
def protocol_result(model: str = "chosen:model") -> DirectProtocolResult:
|
||||
text = "# Protokoll\n\nUnverändert. \n"
|
||||
return DirectProtocolResult(
|
||||
protocol_text=text,
|
||||
exact_prompt="exact prompt\n",
|
||||
model_metadata={"model": model},
|
||||
runtime_metadata={"model": model, "request_count": 1, "client_wall_time_seconds": 0.5},
|
||||
raw_response={"response": text, "done": True},
|
||||
transcript_input="selected transcript\n",
|
||||
)
|
||||
|
||||
|
||||
def fake_transcribe(
|
||||
audio_path: Path,
|
||||
model_path: Path,
|
||||
output_dir: Path,
|
||||
language: str,
|
||||
*,
|
||||
executable: str,
|
||||
threads: str | int,
|
||||
) -> TranscriptionResult:
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
raw = output_dir / "whisper_raw.json"
|
||||
transcript = output_dir / "transcript.json"
|
||||
text = output_dir / "transcript.txt"
|
||||
metadata = output_dir / "runtime_metadata.json"
|
||||
raw.write_text('{"transcription": []}\n', encoding="utf-8")
|
||||
transcript.write_text(
|
||||
json.dumps(
|
||||
{
|
||||
"text": "Besprechungstext.",
|
||||
"segments": [
|
||||
{"id": 0, "start": 0.0, "end": 1.0, "text": "Besprechungstext."}
|
||||
],
|
||||
}
|
||||
)
|
||||
+ "\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
text.write_text("Besprechungstext.\n", encoding="utf-8")
|
||||
metadata.write_text('{"runtime_seconds": 1.25}\n', encoding="utf-8")
|
||||
return TranscriptionResult(output_dir, raw, transcript, text, metadata, 1.25)
|
||||
|
||||
|
||||
def fake_diarize(audio_path, output_dir, device_mode, **kwargs):
|
||||
output_dir.mkdir(parents=True, exist_ok=True)
|
||||
metadata = {
|
||||
"backend": "pyannote.audio",
|
||||
"model": "pyannote/speaker-diarization-community-1",
|
||||
"requested_device_mode": device_mode,
|
||||
"actual_device": "cuda",
|
||||
"device_name": "Fake GPU",
|
||||
"runtime_seconds": 2.5,
|
||||
"speaker_count": 1,
|
||||
"credentials_persisted": False,
|
||||
}
|
||||
paths = {
|
||||
"metadata": output_dir / "metadata.json",
|
||||
"ordinary": output_dir / "diarization.rttm",
|
||||
"exclusive": output_dir / "exclusive_diarization.rttm",
|
||||
"turns": output_dir / "turns.json",
|
||||
"exclusive_turns": output_dir / "exclusive_turns.json",
|
||||
}
|
||||
paths["metadata"].write_text(json.dumps(metadata), encoding="utf-8")
|
||||
paths["ordinary"].write_text("", encoding="utf-8")
|
||||
paths["exclusive"].write_text("", encoding="utf-8")
|
||||
paths["turns"].write_text("[]", encoding="utf-8")
|
||||
paths["exclusive_turns"].write_text(
|
||||
json.dumps([{"start": 0, "end": 10, "speaker_id": "SPEAKER_00"}]),
|
||||
encoding="utf-8",
|
||||
)
|
||||
return DiarizationResult(
|
||||
output_dir,
|
||||
paths["metadata"],
|
||||
paths["ordinary"],
|
||||
paths["exclusive"],
|
||||
paths["turns"],
|
||||
paths["exclusive_turns"],
|
||||
metadata,
|
||||
)
|
||||
|
||||
|
||||
def fake_prepare(source, destination, **kwargs):
|
||||
destination.parent.mkdir(parents=True, exist_ok=True)
|
||||
shutil.copyfile(source, destination)
|
||||
normalization_enabled = kwargs.get("normalization_enabled", True)
|
||||
return PreparedAudio(
|
||||
source,
|
||||
source.suffix.removeprefix("."),
|
||||
destination,
|
||||
"ffmpeg",
|
||||
kwargs.get("ffmpeg_executable", "ffmpeg"),
|
||||
normalization_enabled,
|
||||
DEFAULT_NORMALIZATION_METHOD if normalization_enabled else None,
|
||||
DEFAULT_NORMALIZATION_FILTER if normalization_enabled else None,
|
||||
)
|
||||
|
||||
|
||||
class MvpOrchestratorTests(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
patcher = patch.object(mvp_api, "prepare_audio", side_effect=fake_prepare)
|
||||
self.prepare_audio = patcher.start()
|
||||
self.addCleanup(patcher.stop)
|
||||
|
||||
def create_inputs(self, root: Path) -> tuple[Path, Path, Path]:
|
||||
audio = root / "team meeting.wav"
|
||||
whisper_model = root / "ggml-model.bin"
|
||||
context = root / "source-context.yaml"
|
||||
audio.write_bytes(b"audio")
|
||||
whisper_model.write_bytes(b"model")
|
||||
context.write_text(VALID_CONTEXT, encoding="utf-8")
|
||||
return audio, whisper_model, context
|
||||
|
||||
def args(self, root: Path, extra: list[str] | None = None):
|
||||
audio, whisper_model, context = self.create_inputs(root)
|
||||
values = [
|
||||
str(audio),
|
||||
"--whisper-model", str(whisper_model),
|
||||
"--context", str(context),
|
||||
"--output-root", str(root / "runs"),
|
||||
"--language", "de",
|
||||
"--model", "chosen:model",
|
||||
"--ollama-endpoint", "http://ollama.test:11434",
|
||||
]
|
||||
if extra:
|
||||
values.extend(extra)
|
||||
return run_mvp_meeting.parse_args(values)
|
||||
|
||||
def test_successful_full_orchestration_and_artifact_layout(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root)
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe) as whisper,
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
return_value=protocol_result(),
|
||||
) as protocol,
|
||||
):
|
||||
code, run_dir, protocol_path = run_mvp_meeting.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
self.assertIsNotNone(run_dir)
|
||||
self.assertEqual(protocol_path, run_dir / "protocol.md")
|
||||
expected = {
|
||||
"run_metadata.json",
|
||||
"audio/input_manifest.json",
|
||||
"audio/prepared.wav",
|
||||
"audio/preparation_metadata.json",
|
||||
"transcript/whisper_raw.json",
|
||||
"transcript/transcript.json",
|
||||
"transcript/transcript.txt",
|
||||
"transcript/runtime_metadata.json",
|
||||
"context/meeting_context.yaml",
|
||||
"protocol/exact_prompt.txt",
|
||||
"protocol/transcript_input.txt",
|
||||
"protocol/raw_response.json",
|
||||
"protocol/runtime_metadata.json",
|
||||
"protocol.md",
|
||||
}
|
||||
self.assertTrue(all((run_dir / item).is_file() for item in expected))
|
||||
metadata = json.loads((run_dir / "run_metadata.json").read_text())
|
||||
self.assertEqual(metadata["status"], "completed")
|
||||
self.assertEqual(metadata["model"], "chosen:model")
|
||||
self.assertIsNone(metadata["failure"])
|
||||
self.assertEqual(whisper.call_count, 1)
|
||||
self.assertEqual(
|
||||
whisper.call_args.args[0], run_dir / "audio" / "prepared.wav"
|
||||
)
|
||||
self.assertEqual(protocol.call_count, 1)
|
||||
self.assertTrue(
|
||||
self.prepare_audio.call_args.kwargs["normalization_enabled"]
|
||||
)
|
||||
preparation = json.loads(
|
||||
(run_dir / "audio" / "preparation_metadata.json").read_text()
|
||||
)
|
||||
self.assertTrue(preparation["normalization_enabled"])
|
||||
|
||||
def test_cli_can_disable_audio_normalization_without_bypassing_preparation(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root, ["--no-audio-normalization"])
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe) as whisper,
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
return_value=protocol_result(),
|
||||
),
|
||||
):
|
||||
code, run_dir, _ = run_mvp_meeting.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
self.prepare_audio.assert_called_once()
|
||||
self.assertFalse(
|
||||
self.prepare_audio.call_args.kwargs["normalization_enabled"]
|
||||
)
|
||||
self.assertEqual(
|
||||
whisper.call_args.args[0], run_dir / "audio" / "prepared.wav"
|
||||
)
|
||||
preparation = json.loads(
|
||||
(run_dir / "audio" / "preparation_metadata.json").read_text()
|
||||
)
|
||||
self.assertFalse(preparation["normalization_enabled"])
|
||||
self.assertIsNone(preparation["normalization_method"])
|
||||
self.assertIsNone(preparation["normalization_filter"])
|
||||
|
||||
def test_context_model_endpoint_and_whisper_options_are_forwarded(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(
|
||||
root,
|
||||
[
|
||||
"--whisper-executable",
|
||||
"/tools/whisper-cli",
|
||||
"--ffmpeg-executable",
|
||||
"/tools/ffmpeg",
|
||||
"--threads",
|
||||
"4",
|
||||
],
|
||||
)
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe) as whisper,
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
return_value=protocol_result(),
|
||||
) as protocol,
|
||||
):
|
||||
code, run_dir, _ = run_mvp_meeting.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
self.assertEqual(whisper.call_args.args[3], "de")
|
||||
self.assertEqual(whisper.call_args.kwargs["executable"], "/tools/whisper-cli")
|
||||
self.assertEqual(whisper.call_args.kwargs["threads"], "4")
|
||||
self.assertEqual(
|
||||
self.prepare_audio.call_args.kwargs["ffmpeg_executable"],
|
||||
"/tools/ffmpeg",
|
||||
)
|
||||
self.assertEqual(protocol.call_args.args[1], run_dir / "context/meeting_context.yaml")
|
||||
self.assertEqual(protocol.call_args.kwargs["model"], "chosen:model")
|
||||
self.assertEqual(
|
||||
protocol.call_args.kwargs["endpoint"], "http://ollama.test:11434"
|
||||
)
|
||||
|
||||
def test_diarization_is_off_by_default_and_preserves_protocol_input(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root)
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe),
|
||||
patch.object(mvp_api, "diarize_audio") as diarization,
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
return_value=protocol_result(),
|
||||
) as protocol,
|
||||
):
|
||||
code, run_dir, _ = run_mvp_meeting.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
diarization.assert_not_called()
|
||||
self.assertEqual(protocol.call_args.args[0], run_dir / "transcript/transcript.json")
|
||||
self.assertFalse((run_dir / "diarization").exists())
|
||||
|
||||
def test_diarization_cli_propagates_and_uses_derived_protocol_input(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(
|
||||
root,
|
||||
[
|
||||
"--diarization", "gpu",
|
||||
"--diarization-runtime", "container",
|
||||
"--diarization-container-image", "rocm/test",
|
||||
"--diarization-container-arg=--device=/dev/kfd",
|
||||
],
|
||||
)
|
||||
with (
|
||||
patch.object(
|
||||
mvp_api, "transcribe_audio", side_effect=fake_transcribe
|
||||
),
|
||||
patch.object(
|
||||
mvp_api, "diarize_audio", side_effect=fake_diarize
|
||||
) as diarization,
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
return_value=protocol_result(),
|
||||
) as protocol,
|
||||
):
|
||||
code, run_dir, _ = run_mvp_meeting.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
self.assertEqual(diarization.call_args.args[2], "gpu")
|
||||
self.assertEqual(
|
||||
diarization.call_args.args[0], run_dir / "audio" / "prepared.wav"
|
||||
)
|
||||
self.assertEqual(diarization.call_args.kwargs["runtime"], "container")
|
||||
self.assertEqual(
|
||||
diarization.call_args.kwargs["container_args"], ("--device=/dev/kfd",)
|
||||
)
|
||||
derived = run_dir / "diarization/transcript_diarized.json"
|
||||
self.assertEqual(protocol.call_args.args[0], derived)
|
||||
self.assertIn("SPEAKER_00", derived.read_text(encoding="utf-8"))
|
||||
self.assertEqual(
|
||||
json.loads((run_dir / "transcript/transcript.json").read_text())["text"],
|
||||
"Besprechungstext.",
|
||||
)
|
||||
run_metadata = json.loads((run_dir / "run_metadata.json").read_text())
|
||||
self.assertTrue(run_metadata["diarization"]["enabled"])
|
||||
self.assertNotIn("HF_TOKEN", json.dumps(run_metadata))
|
||||
|
||||
def test_whisper_failure_is_recorded_and_protocol_is_not_called(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root)
|
||||
with (
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"transcribe_audio",
|
||||
side_effect=TranscriptionError("whisper stopped"),
|
||||
),
|
||||
patch.object(mvp_api, "generate_direct_protocol") as protocol,
|
||||
):
|
||||
code, run_dir, protocol_path = run_mvp_meeting.run(args)
|
||||
|
||||
metadata = json.loads((run_dir / "run_metadata.json").read_text())
|
||||
self.assertEqual(code, 2)
|
||||
self.assertIsNone(protocol_path)
|
||||
self.assertEqual(metadata["status"], "failed")
|
||||
self.assertEqual(metadata["failure"]["stage"], "whisper")
|
||||
self.assertIn("whisper stopped", metadata["failure"]["message"])
|
||||
protocol.assert_not_called()
|
||||
self.assertTrue((run_dir / "audio/input_manifest.json").is_file())
|
||||
|
||||
def test_protocol_failure_preserves_transcript_and_failure_metadata(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root)
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe),
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
side_effect=ValueError("generation stopped"),
|
||||
),
|
||||
):
|
||||
code, run_dir, protocol_path = run_mvp_meeting.run(args)
|
||||
|
||||
metadata = json.loads((run_dir / "run_metadata.json").read_text())
|
||||
self.assertEqual(code, 2)
|
||||
self.assertIsNone(protocol_path)
|
||||
self.assertEqual(metadata["failure"]["stage"], "protocol")
|
||||
self.assertTrue((run_dir / "transcript/transcript.json").is_file())
|
||||
self.assertFalse((run_dir / "protocol.md").exists())
|
||||
|
||||
def test_unique_run_directories_do_not_overwrite(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
fixed = datetime(2026, 8, 21, 9, 15, 30)
|
||||
first = run_mvp_meeting.create_unique_run_dir(root, "team meeting", lambda: fixed)
|
||||
marker = first / "keep.txt"
|
||||
marker.write_text("keep", encoding="utf-8")
|
||||
second = run_mvp_meeting.create_unique_run_dir(root, "team meeting", lambda: fixed)
|
||||
self.assertEqual(first.name, "team_meeting_20260821_091530")
|
||||
self.assertEqual(second.name, "team_meeting_20260821_091530_01")
|
||||
self.assertEqual(marker.read_text(), "keep")
|
||||
|
||||
def test_semantic_pipeline_functions_are_never_invoked(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root)
|
||||
with (
|
||||
patch.object(mvp_api, "transcribe_audio", side_effect=fake_transcribe),
|
||||
patch.object(
|
||||
mvp_api,
|
||||
"generate_direct_protocol",
|
||||
return_value=protocol_result(),
|
||||
),
|
||||
patch("src.meeting_lab.chunking.chunk_transcript.build_chunks") as chunking,
|
||||
patch("src.meeting_lab.extraction.extract_chunks.extract_input") as extraction,
|
||||
patch("src.meeting_lab.consolidation.consolidate_facts.call_ollama") as consolidation,
|
||||
):
|
||||
code, _, _ = run_mvp_meeting.run(args)
|
||||
|
||||
self.assertEqual(code, 0)
|
||||
chunking.assert_not_called()
|
||||
extraction.assert_not_called()
|
||||
consolidation.assert_not_called()
|
||||
|
||||
def test_main_returns_nonzero_for_whisper_failure(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
args = self.args(root)
|
||||
argv = [
|
||||
str(args.audio_file),
|
||||
"--whisper-model", str(args.whisper_model),
|
||||
"--context", str(args.context),
|
||||
"--output-root", str(args.output_root),
|
||||
]
|
||||
with patch.object(
|
||||
mvp_api,
|
||||
"transcribe_audio",
|
||||
side_effect=TranscriptionError("failed"),
|
||||
):
|
||||
self.assertEqual(run_mvp_meeting.main(argv), 2)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,162 @@
|
||||
import argparse
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
import src.meeting_lab.controlled_semantic_derivation.experiment_negative_act as module
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_negative_act import (
|
||||
DerivationValidationError,
|
||||
build_ollama_payload,
|
||||
build_prompt,
|
||||
evaluate_case,
|
||||
load_gold_cases,
|
||||
parse_model_json,
|
||||
run_experiment,
|
||||
validate_classification,
|
||||
)
|
||||
|
||||
|
||||
GOLD_PATH = Path("tests/gold/negative_act_form_v0/cases.json")
|
||||
|
||||
|
||||
FORM_TEXT = {
|
||||
"NA-01": "Zusammenarbeit mit Dr. Schlummer fortsetzen",
|
||||
"NA-02": "externe Lösung weiterverfolgen",
|
||||
"NA-03": "reale Anlage für den Versuch nutzen",
|
||||
"NA-04": "reale Anlage verwenden",
|
||||
"NA-05": "Waschstufe einbauen",
|
||||
}
|
||||
|
||||
|
||||
def classification_for(case):
|
||||
expected = case["expected"]
|
||||
return {
|
||||
"observation_id": expected["observation_id"],
|
||||
"negative_act_form": expected["negative_act_form"],
|
||||
"normalized_action_text": FORM_TEXT.get(case["case_id"]),
|
||||
}
|
||||
|
||||
|
||||
class NegativeActFormExperimentTests(unittest.TestCase):
|
||||
@classmethod
|
||||
def setUpClass(cls):
|
||||
cls.cases = load_gold_cases(GOLD_PATH)
|
||||
cls.by_id = {case["case_id"]: case for case in cls.cases}
|
||||
|
||||
def test_fixture_contains_exactly_na_01_through_na_08(self):
|
||||
self.assertEqual(list(self.by_id), [f"NA-{number:02d}" for number in range(1, 9)])
|
||||
|
||||
def test_exact_schema_is_accepted(self):
|
||||
case = self.by_id["NA-01"]
|
||||
self.assertEqual(validate_classification(classification_for(case), case["observations"]), classification_for(case))
|
||||
|
||||
def test_unknown_field_is_rejected(self):
|
||||
case = self.by_id["NA-01"]
|
||||
classification = classification_for(case)
|
||||
classification["explanation"] = "extra"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown keys"):
|
||||
validate_classification(classification, case["observations"])
|
||||
|
||||
def test_invalid_enum_is_rejected(self):
|
||||
case = self.by_id["NA-01"]
|
||||
classification = classification_for(case)
|
||||
classification["negative_act_form"] = "rejection"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unsupported value"):
|
||||
validate_classification(classification, case["observations"])
|
||||
|
||||
def test_non_none_requires_normalized_action_text(self):
|
||||
case = self.by_id["NA-01"]
|
||||
for value in (None, ""):
|
||||
classification = classification_for(case)
|
||||
classification["normalized_action_text"] = value
|
||||
with self.subTest(value=value), self.assertRaises(DerivationValidationError):
|
||||
validate_classification(classification, case["observations"])
|
||||
|
||||
def test_none_requires_null_normalized_action_text(self):
|
||||
case = self.by_id["NA-06"]
|
||||
classification = classification_for(case)
|
||||
self.assertIsNone(classification["normalized_action_text"])
|
||||
classification["normalized_action_text"] = "Material einsetzen"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "requires null"):
|
||||
validate_classification(classification, case["observations"])
|
||||
|
||||
def test_forbidden_normative_fields_are_rejected_recursively(self):
|
||||
case = self.by_id["NA-01"]
|
||||
fields = (
|
||||
"rejection_form", "explicitly_rejected", "status", "decision", "outcome",
|
||||
"topic_status", "responsible_person", "responsibility", "owner",
|
||||
"requested_actor", "action_item", "protocol_category", "confidence",
|
||||
"relation", "relations", "graph", "unresolved_issue",
|
||||
)
|
||||
for field in fields:
|
||||
classification = classification_for(case)
|
||||
classification["wrapper"] = {field: "forbidden"}
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||
validate_classification(classification, case["observations"])
|
||||
|
||||
def test_unknown_observation_id_is_rejected(self):
|
||||
case = self.by_id["NA-01"]
|
||||
classification = classification_for(case)
|
||||
classification["observation_id"] = "obs_99"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown observation"):
|
||||
validate_classification(classification, case["observations"])
|
||||
|
||||
def test_malformed_json_is_rejected(self):
|
||||
with self.assertRaises(json.JSONDecodeError):
|
||||
parse_model_json("{bad json")
|
||||
|
||||
def test_all_expected_classifications_evaluate_as_pass(self):
|
||||
for case in self.cases:
|
||||
evaluation = evaluate_case(case, classification_for(case))
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertEqual(evaluation["classification"], "PASS")
|
||||
|
||||
def test_fixed_prompt_contains_candidate_and_no_gold_expectation(self):
|
||||
prompt = build_prompt(self.by_id["NA-03"])
|
||||
self.assertIn("candidate observation is obs_2", prompt)
|
||||
self.assertNotIn("expected", prompt)
|
||||
self.assertNotIn("Who is responsible", prompt)
|
||||
|
||||
def test_fixed_model_configuration(self):
|
||||
payload = build_ollama_payload("qwen3.5:9B", "prompt", 16384, 1024)
|
||||
self.assertFalse(payload["think"])
|
||||
self.assertFalse(payload["stream"])
|
||||
self.assertEqual(payload["options"]["temperature"], 0)
|
||||
|
||||
def test_no_rejection_or_status_derivation_function_exists(self):
|
||||
public_names = {name for name in dir(module) if not name.startswith("_")}
|
||||
self.assertNotIn("derive_rejection", public_names)
|
||||
self.assertFalse(any(name.startswith("derive_") for name in public_names))
|
||||
|
||||
def test_artifacts_preserve_semantic_classification_only(self):
|
||||
case = self.by_id["NA-01"]
|
||||
raw = json.dumps(classification_for(case), ensure_ascii=False)
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
output = Path(temporary) / "run"
|
||||
args = argparse.Namespace(
|
||||
cases=GOLD_PATH, output=output, model="qwen3.5:9B",
|
||||
endpoint="http://unused", timeout=1, num_ctx=16384, num_predict=1024,
|
||||
)
|
||||
with patch.object(module, "load_gold_cases", return_value=[deepcopy(case)]), patch.object(
|
||||
module, "call_ollama", return_value=(raw, {"model": "qwen3.5:9B"})
|
||||
):
|
||||
summary = run_experiment(args)
|
||||
self.assertEqual(summary["successful_llm_call_count"], 1)
|
||||
case_dir = output / "na-01"
|
||||
for filename in (
|
||||
"v3_style_input_observations.json", "prompt.txt", "raw_model_response.txt",
|
||||
"parsed_semantic_classification.json", "structural_validation.json",
|
||||
"evaluation.json", "ollama_metadata.json",
|
||||
):
|
||||
self.assertTrue((case_dir / filename).is_file(), filename)
|
||||
self.assertFalse((case_dir / "final_derived_result.json").exists())
|
||||
self.assertFalse((case_dir / "deterministic_gate_results.json").exists())
|
||||
parsed = json.loads((case_dir / "parsed_semantic_classification.json").read_text())
|
||||
self.assertEqual(set(parsed), {"observation_id", "negative_act_form", "normalized_action_text"})
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,315 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock, patch
|
||||
|
||||
from src.meeting_lab.diarization.alignment import diarized_transcript_text
|
||||
from src.meeting_lab.llm import ollama
|
||||
from src.meeting_lab.llm.ollama import OllamaGeneration
|
||||
from src.meeting_lab.protocol.direct_protocol_prompt import (
|
||||
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||
build_direct_protocol_prompt,
|
||||
)
|
||||
from src.meeting_lab.protocol.generate_direct_protocol import (
|
||||
DirectProtocolError,
|
||||
estimate_input_tokens,
|
||||
generate_direct_protocol,
|
||||
)
|
||||
from src.meeting_lab.protocol.transcript_input import compact_diarized_transcript
|
||||
|
||||
|
||||
def segments() -> list[dict[str, object]]:
|
||||
return [
|
||||
{"id": 0, "start": 0.0, "end": 1.0, "text": "First.", "speaker_id": "SPEAKER_01"},
|
||||
{"id": 1, "start": 1.0, "end": 2.0, "text": "Second.", "speaker_id": "SPEAKER_01"},
|
||||
{"id": 2, "start": 2.0, "end": 3.0, "text": "Third.", "speaker_id": "SPEAKER_04"},
|
||||
{"id": 3, "start": 3.0, "end": 4.0, "text": "Unassigned.", "speaker_id": None},
|
||||
{"id": 4, "start": 4.0, "end": 5.0, "text": "Last.", "speaker_id": "SPEAKER_01"},
|
||||
]
|
||||
|
||||
|
||||
def diarized_document(repetitions: int = 1) -> dict[str, object]:
|
||||
source = segments() * repetitions
|
||||
return {
|
||||
"text": diarized_transcript_text(source),
|
||||
"segments": source,
|
||||
"speaker_labels_anonymous": True,
|
||||
"alignment_source": "exclusive_diarization",
|
||||
}
|
||||
|
||||
|
||||
def completed_generation() -> OllamaGeneration:
|
||||
text = "# Meeting Protocol\n\nComplete."
|
||||
return OllamaGeneration(
|
||||
raw_response={
|
||||
"response": text,
|
||||
"done": True,
|
||||
"done_reason": "stop",
|
||||
"prompt_eval_count": 100,
|
||||
"eval_count": 10,
|
||||
},
|
||||
text=text,
|
||||
client_wall_time_seconds=0.1,
|
||||
)
|
||||
|
||||
|
||||
class CompactDiarizedTranscriptTests(unittest.TestCase):
|
||||
def test_adjacent_segments_group_and_transitions_remain_separate(self) -> None:
|
||||
compact = compact_diarized_transcript(segments())
|
||||
|
||||
self.assertEqual(
|
||||
[block.speaker_id for block in compact.blocks],
|
||||
["SPEAKER_01", "SPEAKER_04", "SPEAKER_UNASSIGNED", "SPEAKER_01"],
|
||||
)
|
||||
self.assertEqual(compact.blocks[0].segment_texts, ("First.", "Second."))
|
||||
self.assertEqual(compact.blocks[-1].segment_texts, ("Last.",))
|
||||
self.assertEqual(compact.text.count("SPEAKER_01:"), 2)
|
||||
|
||||
def test_every_segment_text_and_order_are_preserved(self) -> None:
|
||||
source = segments()
|
||||
compact = compact_diarized_transcript(source)
|
||||
|
||||
self.assertEqual(compact.source_segment_count, len(source))
|
||||
self.assertEqual(compact.represented_segment_count, len(source))
|
||||
self.assertEqual(
|
||||
compact.segment_texts,
|
||||
tuple(str(segment["text"]) for segment in source),
|
||||
)
|
||||
self.assertEqual(compact.segment_texts[0], "First.")
|
||||
self.assertEqual(compact.segment_texts[-1], "Last.")
|
||||
self.assertIn("SPEAKER_UNASSIGNED: Unassigned.", compact.text)
|
||||
|
||||
def test_compact_form_is_materially_smaller_than_per_segment_format(self) -> None:
|
||||
source = [
|
||||
{
|
||||
"start": index,
|
||||
"end": index + 1,
|
||||
"text": "Repeated transcript content.",
|
||||
"speaker_id": "SPEAKER_01",
|
||||
}
|
||||
for index in range(100)
|
||||
]
|
||||
|
||||
compact = compact_diarized_transcript(source).text
|
||||
verbose = diarized_transcript_text(source)
|
||||
|
||||
self.assertLess(len(compact), len(verbose) * 0.6)
|
||||
|
||||
|
||||
class ProtocolInputBudgetTests(unittest.TestCase):
|
||||
def _write(self, root: Path, document: dict[str, object]) -> Path:
|
||||
path = root / "transcript.json"
|
||||
path.write_text(json.dumps(document), encoding="utf-8")
|
||||
return path
|
||||
|
||||
def test_token_estimate_uses_utf8_bytes_for_non_ascii_safety(self) -> None:
|
||||
self.assertEqual(estimate_input_tokens("ä" * 44), 20)
|
||||
|
||||
def test_compact_diarized_representation_selected_within_budget(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
document = diarized_document()
|
||||
transcript = self._write(root, document)
|
||||
compact = compact_diarized_transcript(document["segments"]).text
|
||||
budget = estimate_input_tokens(
|
||||
build_direct_protocol_prompt(
|
||||
compact,
|
||||
instruction=COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||
)
|
||||
)
|
||||
call = Mock(return_value=completed_generation())
|
||||
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
safe_input_token_budget=budget,
|
||||
model_check=Mock(return_value={}),
|
||||
generation_call=call,
|
||||
)
|
||||
|
||||
self.assertEqual(
|
||||
result.runtime_metadata["selected_transcript_representation"],
|
||||
"diarized_compact",
|
||||
)
|
||||
self.assertFalse(result.runtime_metadata["fallback_used"])
|
||||
self.assertTrue(result.runtime_metadata["diarization_enabled"])
|
||||
self.assertEqual(result.runtime_metadata["safe_input_token_budget"], budget)
|
||||
self.assertEqual(result.transcript_input, compact)
|
||||
self.assertEqual(call.call_count, 1)
|
||||
|
||||
def test_mapped_speakers_and_statements_reach_final_ollama_payload(self) -> None:
|
||||
diarized = {
|
||||
"text": "",
|
||||
"segments": [
|
||||
{
|
||||
"start": 0.0,
|
||||
"end": 1.0,
|
||||
"speaker_id": "SPEAKER_00",
|
||||
"text": "We will run the trial on Wednesday.",
|
||||
},
|
||||
{
|
||||
"start": 1.0,
|
||||
"end": 2.0,
|
||||
"speaker_id": "SPEAKER_01",
|
||||
"text": "I will prepare the raw materials before then.",
|
||||
},
|
||||
{
|
||||
"start": 2.0,
|
||||
"end": 3.0,
|
||||
"speaker_id": "SPEAKER_00",
|
||||
"text": "Good. Anna owns the material preparation.",
|
||||
},
|
||||
],
|
||||
"speaker_labels_anonymous": True,
|
||||
"alignment_source": "exclusive_diarization",
|
||||
}
|
||||
context = {
|
||||
"schema_version": "1",
|
||||
"meeting": {
|
||||
"meeting_id": "speaker-test",
|
||||
"title": "Speaker test",
|
||||
"language": "en",
|
||||
},
|
||||
"participants": [
|
||||
{"participant_id": "martin", "display_name": "Martin"},
|
||||
{"participant_id": "anna", "display_name": "Anna"},
|
||||
],
|
||||
"speaker_mappings": {"SPEAKER_00": "martin", "SPEAKER_01": "anna"},
|
||||
"mentioned_people": [],
|
||||
"organization": {"departments": []},
|
||||
"known_entities": {},
|
||||
}
|
||||
response = Mock()
|
||||
response.raise_for_status.return_value = None
|
||||
response.json.return_value = {"response": "# Meeting Protocol\n", "done": True}
|
||||
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = self._write(root, diarized)
|
||||
context_path = root / "context.yaml"
|
||||
context_path.write_text(json.dumps(context), encoding="utf-8")
|
||||
with patch.object(ollama.requests, "post", return_value=response) as post:
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
context_path,
|
||||
model="qwen3.8:27b",
|
||||
model_check=Mock(return_value={}),
|
||||
)
|
||||
|
||||
prompt = post.call_args.kwargs["json"]["prompt"]
|
||||
self.assertEqual(result.exact_prompt, prompt)
|
||||
self.assertIn("- SPEAKER_00: Martin (participant_id: martin)", prompt)
|
||||
self.assertIn("- SPEAKER_01: Anna (participant_id: anna)", prompt)
|
||||
self.assertIn("SPEAKER_00: We will run the trial on Wednesday.", prompt)
|
||||
self.assertIn(
|
||||
"SPEAKER_01: I will prepare the raw materials before then.", prompt
|
||||
)
|
||||
self.assertIn("SPEAKER_00: Good. Anna owns the material preparation.", prompt)
|
||||
self.assertNotIn("Martin: We will run the trial on Wednesday.", prompt)
|
||||
self.assertNotIn("Anna: I will prepare the raw materials before then.", prompt)
|
||||
self.assertIn("autoritativen SPEAKER_XX-zu-Teilnehmer-Zuordnungen", prompt)
|
||||
self.assertIn("Ich-Zusage", prompt)
|
||||
self.assertIn("nur erwähnten Personen", prompt)
|
||||
self.assertIn("keine persönliche Verantwortung", prompt)
|
||||
|
||||
def test_plain_fallback_selected_when_diarized_compact_exceeds_budget(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
alternating_segments = [
|
||||
{
|
||||
"id": index,
|
||||
"start": float(index),
|
||||
"end": float(index + 1),
|
||||
"text": "Word.",
|
||||
"speaker_id": f"SPEAKER_{index % 2:02d}",
|
||||
}
|
||||
for index in range(200)
|
||||
]
|
||||
document = {
|
||||
"text": diarized_transcript_text(alternating_segments),
|
||||
"segments": alternating_segments,
|
||||
"speaker_labels_anonymous": True,
|
||||
"alignment_source": "exclusive_diarization",
|
||||
}
|
||||
transcript = self._write(root, document)
|
||||
compact = compact_diarized_transcript(document["segments"]).text
|
||||
plain = " ".join(str(segment["text"]) for segment in document["segments"])
|
||||
compact_estimate = estimate_input_tokens(
|
||||
build_direct_protocol_prompt(
|
||||
compact,
|
||||
instruction=COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
|
||||
)
|
||||
)
|
||||
plain_estimate = estimate_input_tokens(build_direct_protocol_prompt(plain))
|
||||
self.assertLess(plain_estimate, compact_estimate)
|
||||
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
safe_input_token_budget=plain_estimate,
|
||||
model_check=Mock(return_value={}),
|
||||
generation_call=Mock(return_value=completed_generation()),
|
||||
)
|
||||
|
||||
self.assertEqual(
|
||||
result.runtime_metadata["selected_transcript_representation"],
|
||||
"plain_transcript_fallback",
|
||||
)
|
||||
self.assertTrue(result.runtime_metadata["fallback_used"])
|
||||
self.assertEqual(result.transcript_input, plain)
|
||||
self.assertNotIn("SPEAKER_00", result.exact_prompt)
|
||||
self.assertNotIn("SPEAKER_01", result.exact_prompt)
|
||||
self.assertFalse(result.runtime_metadata["speaker_attribution_available"])
|
||||
self.assertEqual(
|
||||
result.runtime_metadata["speaker_attribution_loss_reason"],
|
||||
"plain_transcript_fallback",
|
||||
)
|
||||
|
||||
def test_oversized_plain_transcript_fails_before_any_network_call(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = self._write(
|
||||
root,
|
||||
{"text": "large input " * 1000, "segments": []},
|
||||
)
|
||||
model_check = Mock()
|
||||
generation_call = Mock()
|
||||
|
||||
with self.assertRaisesRegex(
|
||||
DirectProtocolError,
|
||||
"No LLM request was made; silent truncation is not allowed",
|
||||
):
|
||||
generate_direct_protocol(
|
||||
transcript,
|
||||
safe_input_token_budget=1,
|
||||
model_check=model_check,
|
||||
generation_call=generation_call,
|
||||
)
|
||||
|
||||
model_check.assert_not_called()
|
||||
generation_call.assert_not_called()
|
||||
|
||||
def test_existing_plain_path_and_metadata_remain_direct(self) -> None:
|
||||
with tempfile.TemporaryDirectory() as directory:
|
||||
root = Path(directory)
|
||||
transcript = self._write(
|
||||
root,
|
||||
{"text": "Plain original transcript.", "segments": []},
|
||||
)
|
||||
result = generate_direct_protocol(
|
||||
transcript,
|
||||
model_check=Mock(return_value={}),
|
||||
generation_call=Mock(return_value=completed_generation()),
|
||||
)
|
||||
|
||||
self.assertEqual(result.transcript_input, "Plain original transcript.")
|
||||
self.assertEqual(
|
||||
result.runtime_metadata["selected_transcript_representation"],
|
||||
"plain_transcript",
|
||||
)
|
||||
self.assertFalse(result.runtime_metadata["fallback_used"])
|
||||
self.assertFalse(result.runtime_metadata["diarization_enabled"])
|
||||
self.assertIsNone(result.runtime_metadata["speaker_attribution_available"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,178 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from copy import deepcopy
|
||||
from pathlib import Path
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_gold import (
|
||||
RECOGNITION_SCHEMA_VERSION,
|
||||
DerivationValidationError,
|
||||
build_prompt,
|
||||
derive_action,
|
||||
evaluate_case,
|
||||
load_gold_cases,
|
||||
validate_recognition,
|
||||
)
|
||||
|
||||
|
||||
GOLD_PATH = Path("tests/gold/request_acceptance_v0/cases.json")
|
||||
|
||||
|
||||
def recognition_for(case):
|
||||
expected = case["expected_recognition"]
|
||||
request = None
|
||||
acceptance = None
|
||||
if expected["request"]:
|
||||
request = {
|
||||
"observation_id": "obs_1",
|
||||
"is_concrete_request": True,
|
||||
"normalized_action_text": "Auswertung der Messwerte bis Dienstag",
|
||||
}
|
||||
if expected["commitment"]:
|
||||
acceptance = {
|
||||
"observation_id": "obs_2",
|
||||
"is_explicit_commitment": True,
|
||||
"same_requested_work": expected["same_work"],
|
||||
"normalized_action_text": (
|
||||
"Auswertung der Messwerte" if expected["same_work"] else "Präsentation"
|
||||
),
|
||||
}
|
||||
return {
|
||||
"schema_version": RECOGNITION_SCHEMA_VERSION,
|
||||
"request": request,
|
||||
"acceptance": acceptance,
|
||||
}
|
||||
|
||||
|
||||
class RequestAcceptanceGoldExperimentTests(unittest.TestCase):
|
||||
@classmethod
|
||||
def setUpClass(cls):
|
||||
cls.cases = load_gold_cases(GOLD_PATH)
|
||||
cls.by_id = {case["case_id"]: case for case in cls.cases}
|
||||
|
||||
def test_fixture_has_exactly_required_ten_cases(self):
|
||||
self.assertEqual(list(self.by_id), [f"RA-{number:02d}" for number in range(1, 11)])
|
||||
|
||||
def test_all_cases_use_minimal_v3_style_observations(self):
|
||||
required = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
|
||||
for case in self.cases:
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertGreaterEqual(len(case["observations"]), 1)
|
||||
self.assertLessEqual(len(case["observations"]), 2)
|
||||
self.assertTrue(all(set(observation) == required for observation in case["observations"]))
|
||||
|
||||
def test_positive_cases_establish_exact_action_person_due_and_provenance(self):
|
||||
for case_id in ("RA-01", "RA-02"):
|
||||
case = self.by_id[case_id]
|
||||
gates, result = derive_action(case["observations"], recognition_for(case))
|
||||
with self.subTest(case=case_id):
|
||||
self.assertTrue(all(gates.values()))
|
||||
self.assertEqual(result["content"], "Auswertung der Messwerte")
|
||||
self.assertEqual(result["requested_actor"], "Clara")
|
||||
self.assertEqual(result["responsible_person"], "Clara")
|
||||
self.assertEqual(result["due"], "Dienstag")
|
||||
self.assertEqual(result["support"]["request"], {"observation_id": "obs_1", "evidence_id": "e1"})
|
||||
self.assertEqual(result["support"]["acceptance"], {"observation_id": "obs_2", "evidence_id": "e2"})
|
||||
|
||||
def test_paraphrase_does_not_require_lexical_identity(self):
|
||||
case = self.by_id["RA-02"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["acceptance"]["normalized_action_text"] = "darum kümmern und fertigstellen"
|
||||
gates, result = derive_action(case["observations"], recognition)
|
||||
self.assertTrue(gates["same_requested_work"])
|
||||
self.assertIsNotNone(result)
|
||||
|
||||
def test_acknowledgement_is_unestablished(self):
|
||||
self._assert_unestablished("RA-03", "acceptance_semantic_positive")
|
||||
|
||||
def test_tentative_response_is_unestablished(self):
|
||||
self._assert_unestablished("RA-04", "acceptance_semantic_positive")
|
||||
|
||||
def test_different_responder_is_not_personal_acceptance(self):
|
||||
self._assert_unestablished("RA-05", "acceptance_semantic_positive")
|
||||
case = self.by_id["RA-05"]
|
||||
recognition = recognition_for(self.by_id["RA-01"])
|
||||
gates, result = derive_action(case["observations"], recognition)
|
||||
self.assertFalse(gates["acceptance_speaker_matches_addressee"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_different_work_fails_same_work_gate(self):
|
||||
self._assert_unestablished("RA-06", "same_requested_work")
|
||||
|
||||
def test_request_without_response_is_unestablished(self):
|
||||
self._assert_unestablished("RA-07", "acceptance_observation_exists")
|
||||
|
||||
def test_collective_impersonal_and_suggestion_controls_are_unestablished(self):
|
||||
for case_id in ("RA-08", "RA-09", "RA-10"):
|
||||
with self.subTest(case=case_id):
|
||||
self._assert_unestablished(case_id, "request_semantic_positive")
|
||||
|
||||
def test_named_person_without_request_cannot_create_responsibility(self):
|
||||
case = self.by_id["RA-10"]
|
||||
self.assertEqual(case["observations"][0]["named_person"], "Dirk Textor")
|
||||
_, result = derive_action(case["observations"], recognition_for(case))
|
||||
self.assertIsNone(result)
|
||||
|
||||
def test_all_expected_recognitions_evaluate_as_pass(self):
|
||||
for case in self.cases:
|
||||
evaluation = evaluate_case(case, recognition_for(case))
|
||||
with self.subTest(case=case["case_id"]):
|
||||
self.assertEqual(evaluation["classification"], "PASS")
|
||||
|
||||
def test_equivalent_translated_normalized_action_does_not_fail_structure(self):
|
||||
case = self.by_id["RA-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["request"]["normalized_action_text"] = "evaluate measurement values"
|
||||
evaluation = evaluate_case(case, recognition)
|
||||
self.assertEqual(evaluation["classification"], "PASS")
|
||||
self.assertEqual(evaluation["result"]["content"], "evaluate measurement values")
|
||||
|
||||
def test_forbidden_llm_fields_are_rejected_recursively(self):
|
||||
case = self.by_id["RA-01"]
|
||||
for field in (
|
||||
"responsible_person", "responsibility", "requested_actor", "status",
|
||||
"established", "action_item", "protocol_category", "confidence", "graph",
|
||||
):
|
||||
recognition = recognition_for(case)
|
||||
recognition["request"][field] = "forbidden"
|
||||
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_unknown_observation_reference_is_rejected(self):
|
||||
case = self.by_id["RA-01"]
|
||||
recognition = recognition_for(case)
|
||||
recognition["acceptance"]["observation_id"] = "obs_99"
|
||||
with self.assertRaisesRegex(DerivationValidationError, "unknown observation"):
|
||||
validate_recognition(recognition, case["observations"])
|
||||
|
||||
def test_duplicate_evidence_provenance_is_rejected(self):
|
||||
fixture = json.loads(GOLD_PATH.read_text())
|
||||
fixture["cases"][0]["observations"][1]["evidence_id"] = "e1"
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
path = Path(temporary) / "cases.json"
|
||||
path.write_text(json.dumps(fixture), encoding="utf-8")
|
||||
with self.assertRaisesRegex(DerivationValidationError, "must be unique"):
|
||||
load_gold_cases(path)
|
||||
|
||||
def test_prompt_is_fixed_narrow_and_contains_observations_only(self):
|
||||
prompt = build_prompt(self.by_id["RA-01"]["observations"])
|
||||
self.assertIn("V3-style observations", prompt)
|
||||
self.assertNotIn("expected_result", prompt)
|
||||
self.assertNotIn("Who is responsible", prompt)
|
||||
|
||||
def test_conflicting_weekdays_fail_deadline_gate(self):
|
||||
case = deepcopy(self.by_id["RA-01"])
|
||||
case["observations"][1]["content"] = "Clara: Ja, ich übernehme die Auswertung bis Mittwoch."
|
||||
gates, result = derive_action(case["observations"], recognition_for(case))
|
||||
self.assertFalse(gates["deadline_consistent"])
|
||||
self.assertIsNone(result)
|
||||
|
||||
def _assert_unestablished(self, case_id, failed_gate):
|
||||
case = self.by_id[case_id]
|
||||
gates, result = derive_action(case["observations"], recognition_for(case))
|
||||
self.assertFalse(gates[failed_gate])
|
||||
self.assertIsNone(result)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,250 @@
|
||||
import json
|
||||
import tempfile
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
from src.meeting_lab.semantic_synthesis.experiment import (
|
||||
SCHEMA_VERSION,
|
||||
SynthesisValidationError,
|
||||
build_ollama_payload,
|
||||
evaluate_synthesis,
|
||||
load_fixture,
|
||||
run_case,
|
||||
validate_bundle,
|
||||
validate_synthesis,
|
||||
)
|
||||
|
||||
|
||||
class SemanticSynthesisExperimentTests(unittest.TestCase):
|
||||
def setUp(self) -> None:
|
||||
self.case = {
|
||||
"case_id": "case_1",
|
||||
"description": "Known subject test.",
|
||||
"subject_id": "subject_1",
|
||||
"subject": "Prüfung der Messdaten",
|
||||
"evidence": [
|
||||
{"evidence_id": "e1", "text": "Nina übernimmt die Prüfung."},
|
||||
{"evidence_id": "e2", "text": "Die Freigabe bleibt offen."},
|
||||
],
|
||||
"allowed_responsible": ["Nina"],
|
||||
"expected": {
|
||||
"event_type_minimums": {"proposal": 1},
|
||||
"allowed_event_types": ["proposal"],
|
||||
"event_evidence_ids": ["e1"],
|
||||
"outcome": {
|
||||
"required": True,
|
||||
"statuses": ["established"],
|
||||
"terms": ["prüfung"],
|
||||
"scope_terms": ["messdaten"],
|
||||
"evidence_ids": ["e1"],
|
||||
},
|
||||
"actions": {
|
||||
"count": 1,
|
||||
"terms": ["prüfung"],
|
||||
"responsible": "Nina",
|
||||
"due_terms": [],
|
||||
"evidence_ids": ["e1"],
|
||||
},
|
||||
"unresolved_issues": {
|
||||
"count": 1,
|
||||
"terms": ["freigabe"],
|
||||
"evidence_ids": ["e2"],
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
def valid_output(self):
|
||||
return {
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"subject_id": "subject_1",
|
||||
"subject": "Prüfung der Messdaten",
|
||||
"events": [
|
||||
{
|
||||
"type": "proposal",
|
||||
"text": "Die Prüfung wird vorgeschlagen.",
|
||||
"evidence_ids": ["e1"],
|
||||
}
|
||||
],
|
||||
"outcome": {
|
||||
"status": "established",
|
||||
"text": "Die Prüfung wird übernommen.",
|
||||
"scope": "Prüfung der Messdaten",
|
||||
"evidence_ids": ["e1"],
|
||||
},
|
||||
"actions": [
|
||||
{
|
||||
"text": "Prüfung der Messdaten durchführen.",
|
||||
"responsible": "Nina",
|
||||
"due": None,
|
||||
"evidence_ids": ["e1"],
|
||||
}
|
||||
],
|
||||
"unresolved_issues": [
|
||||
{
|
||||
"text": "Die Freigabe bleibt offen.",
|
||||
"evidence_ids": ["e2"],
|
||||
}
|
||||
],
|
||||
}
|
||||
|
||||
def test_bundle_validation_accepts_fixed_subject_and_complete_evidence(self):
|
||||
self.assertIs(validate_bundle(self.case), self.case)
|
||||
|
||||
def test_bundle_validation_rejects_duplicate_evidence_ids(self):
|
||||
case = dict(self.case)
|
||||
case["evidence"] = self.case["evidence"] * 2
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "duplicate evidence ID"):
|
||||
validate_bundle(case)
|
||||
|
||||
def test_sparse_absence_uses_empty_arrays_and_omitted_outcome(self):
|
||||
output = {
|
||||
"schema_version": SCHEMA_VERSION,
|
||||
"subject_id": "subject_1",
|
||||
"subject": "Prüfung der Messdaten",
|
||||
"events": [],
|
||||
"actions": [],
|
||||
"unresolved_issues": [],
|
||||
}
|
||||
|
||||
self.assertIs(validate_synthesis(output, self.case), output)
|
||||
|
||||
def test_outcome_null_is_rejected_but_omission_is_allowed(self):
|
||||
output = self.valid_output()
|
||||
output["outcome"] = None
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "omit it when absent"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_required_arrays_must_exist(self):
|
||||
for field in ("events", "actions", "unresolved_issues"):
|
||||
with self.subTest(field=field):
|
||||
output = self.valid_output()
|
||||
del output[field]
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "missing required"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_fixed_subject_identity_cannot_change(self):
|
||||
output = self.valid_output()
|
||||
output["subject"] = "Different subject"
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "changed fixed subject"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_unknown_evidence_id_is_rejected_in_every_structure(self):
|
||||
mutations = (
|
||||
lambda output: output["events"][0].update(evidence_ids=["unknown"]),
|
||||
lambda output: output["outcome"].update(evidence_ids=["unknown"]),
|
||||
lambda output: output["actions"][0].update(evidence_ids=["unknown"]),
|
||||
lambda output: output["unresolved_issues"][0].update(
|
||||
evidence_ids=["unknown"]
|
||||
),
|
||||
)
|
||||
for mutate in mutations:
|
||||
output = self.valid_output()
|
||||
mutate(output)
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "unknown evidence ID"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_duplicate_evidence_reference_is_rejected(self):
|
||||
output = self.valid_output()
|
||||
output["events"][0]["evidence_ids"] = ["e1", "e1"]
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "duplicate evidence ID"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_responsibility_must_be_allowed_or_json_null(self):
|
||||
output = self.valid_output()
|
||||
output["actions"][0]["responsible"] = None
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
output["actions"][0]["responsible"] = "Martin"
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "not allowed"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_string_null_is_rejected(self):
|
||||
output = self.valid_output()
|
||||
output["actions"][0]["due"] = "null"
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "JSON null"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_outcome_action_and_unresolved_structures_are_strict(self):
|
||||
for field, target in (
|
||||
("extra", lambda output: output["outcome"]),
|
||||
("extra", lambda output: output["actions"][0]),
|
||||
("extra", lambda output: output["unresolved_issues"][0]),
|
||||
):
|
||||
output = self.valid_output()
|
||||
target(output)[field] = "not allowed"
|
||||
with self.assertRaisesRegex(SynthesisValidationError, "unknown keys"):
|
||||
validate_synthesis(output, self.case)
|
||||
|
||||
def test_evaluator_passes_complete_semantics(self):
|
||||
result = evaluate_synthesis(self.valid_output(), self.case["expected"])
|
||||
self.assertEqual(result["verdict"], "PASS")
|
||||
|
||||
def test_evaluator_treats_invented_action_as_critical(self):
|
||||
output = self.valid_output()
|
||||
expected = dict(self.case["expected"])
|
||||
expected["actions"] = {"count": 0}
|
||||
result = evaluate_synthesis(output, expected)
|
||||
self.assertEqual(result["verdict"], "FAIL")
|
||||
self.assertIn("action_count", result["critical_failures"])
|
||||
|
||||
def test_ollama_payload_is_bounded_and_has_required_controls(self):
|
||||
payload = build_ollama_payload("qwen3.5:9B", "prompt", 8192, 2048)
|
||||
self.assertEqual(payload["format"], "json")
|
||||
self.assertIs(payload["think"], False)
|
||||
self.assertIs(payload["stream"], False)
|
||||
self.assertEqual(payload["options"]["temperature"], 0)
|
||||
self.assertEqual(payload["options"]["num_ctx"], 8192)
|
||||
self.assertEqual(payload["options"]["num_predict"], 2048)
|
||||
|
||||
def test_fixture_contains_all_nine_isolation_cases(self):
|
||||
cases = load_fixture(Path("tests/gold/semantic_synthesis_isolation/cases.json"))
|
||||
self.assertEqual(
|
||||
[case["case_id"] for case in cases],
|
||||
[
|
||||
"a_idea_only",
|
||||
"b_multiple_options",
|
||||
"c_unaccepted_proposal",
|
||||
"d_proposal_with_objection",
|
||||
"e_rejected_alternative",
|
||||
"f_trial_only_acceptance",
|
||||
"g_no_decision",
|
||||
"h_resulting_action",
|
||||
"i_outcome_and_unresolved",
|
||||
],
|
||||
)
|
||||
|
||||
def test_case_run_preserves_all_inspection_artifacts(self):
|
||||
raw = json.dumps(self.valid_output(), ensure_ascii=False)
|
||||
metadata = {"elapsed_seconds": 0.01}
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
root = Path(temporary)
|
||||
with patch(
|
||||
"src.meeting_lab.semantic_synthesis.experiment.call_ollama",
|
||||
return_value=(raw, metadata),
|
||||
):
|
||||
result = run_case(
|
||||
self.case,
|
||||
root,
|
||||
"http://unused",
|
||||
"qwen3.5:9B",
|
||||
1,
|
||||
8192,
|
||||
2048,
|
||||
)
|
||||
self.assertEqual(result["verdict"], "PASS")
|
||||
case_dir = root / "case_1"
|
||||
for filename in (
|
||||
"gold_input.json",
|
||||
"prompt.txt",
|
||||
"raw_model_response.txt",
|
||||
"parsed_response.json",
|
||||
"ollama_metadata.json",
|
||||
"validation_result.json",
|
||||
"evaluation_result.json",
|
||||
):
|
||||
self.assertTrue((case_dir / filename).exists(), filename)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,51 @@
|
||||
import argparse,copy,json,tempfile,unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock
|
||||
import src.meeting_lab.controlled_semantic_derivation.experiment_target_normalization as module
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||
|
||||
CASES=module.load_cases(Path("tests/gold/target_normalization_v0/cases.json")); BY={c["case_id"]:c for c in CASES}
|
||||
def output(case,text=None):
|
||||
link=case["fixed_linkage"]; return {"candidate_observation_id":link["candidate_observation_id"],"target_observation_id":link["target_observation_id"],"normalized_target_text":text or case["expected"]["normalized_target_text"]}
|
||||
|
||||
class TargetNormalizationTests(unittest.TestCase):
|
||||
def test_exact_id_copying_accepted(self):
|
||||
for case in CASES: self.assertEqual(module.validate_output(output(case),case),output(case))
|
||||
def test_changed_candidate_rejected(self):
|
||||
case=BY["TN-02"]; data=output(case); data["candidate_observation_id"]="obs_1"
|
||||
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||
def test_changed_target_rejected(self):
|
||||
case=BY["TN-02"]; data=output(case); data["target_observation_id"]="obs_2"
|
||||
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||
def test_empty_and_null_text_rejected(self):
|
||||
case=BY["TN-01"]
|
||||
for value in ("",None):
|
||||
data=output(case); data["normalized_target_text"]=value
|
||||
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||
def test_unknown_and_forbidden_fields_rejected(self):
|
||||
case=BY["TN-01"]
|
||||
for extra in ({"extra":1},{"status":"x"},{"nested":{"decision":True}}):
|
||||
data=output(case); data.update(extra)
|
||||
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||
def test_true_schema_fixes_both_ids_and_disallows_null(self):
|
||||
case=BY["TN-02"]; schema=module.output_schema(case); self.assertEqual(schema["properties"]["candidate_observation_id"]["const"],"obs_2"); self.assertEqual(schema["properties"]["target_observation_id"]["const"],"obs_1"); self.assertEqual(schema["properties"]["normalized_target_text"]["type"],"string"); self.assertFalse(schema["additionalProperties"])
|
||||
def test_payload_uses_schema_object(self):
|
||||
schema=module.output_schema(BY["TN-01"]); payload=module.build_payload("qwen3.5:9B","p",schema,16384,1024); self.assertIs(payload["format"],schema); self.assertIsInstance(payload["format"],dict)
|
||||
def test_duplicate_ids_and_evidence_rejected(self):
|
||||
for field in ("observation_id","evidence_id"):
|
||||
case=copy.deepcopy(BY["TN-02"]); case["observations"][1][field]=case["observations"][0][field]
|
||||
with self.assertRaises(DerivationValidationError): module.validate_linkage(case)
|
||||
def test_prompt_is_fixed_normalization_only(self):
|
||||
first=module.build_prompt(BY["TN-01"]); second=module.build_prompt(BY["TN-02"]); self.assertIn("do not perform target selection",first); self.assertIn("concrete POSITIVE action",first); self.assertEqual(first.split("Fixed candidate_observation_id:")[0],second.split("Fixed candidate_observation_id:")[0])
|
||||
def test_no_target_selection_or_rejection_derivation_exists(self):
|
||||
self.assertFalse(hasattr(module,"select_target")); self.assertFalse(hasattr(module,"derive")); self.assertNotIn("explicitly_rejected",module.OUTPUT_KEYS); self.assertNotIn("status",module.OUTPUT_KEYS)
|
||||
def test_artifacts_preserve_fixed_linkage(self):
|
||||
case=BY["TN-01"]; fixture={"schema_version":module.SCHEMA_VERSION,"cases":[case]}; caller=Mock(return_value=(json.dumps(output(case)),{"model":"qwen3.5:9B"}))
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
root=Path(tmp); path=root/"cases.json"; path.write_text(json.dumps(fixture)); out=root/"out"; args=argparse.Namespace(cases=path,output=out,endpoint="x",model="qwen3.5:9B",timeout=1,num_ctx=16384,num_predict=1024); summary=module.run(args,caller); self.assertEqual(summary["llm_call_count"],1); self.assertEqual(json.loads((out/"tn-01"/"fixed_linkage.json").read_text()),case["fixed_linkage"]); self.assertTrue((out/"tn-01"/"normalized_target_result.json").exists())
|
||||
def test_negative_polarity_fails_evaluation(self):
|
||||
case=BY["TN-01"]; self.assertEqual(module.evaluate(case,output(case,"Mit Dr. Schlummer arbeiten wir nicht weiter"))["classification"],"FAIL")
|
||||
def test_material_scope_and_alternative_contract(self):
|
||||
self.assertEqual(module.evaluate(BY["TN-03"],output(BY["TN-03"]))["classification"],"PASS"); self.assertEqual(module.evaluate(BY["TN-04"],output(BY["TN-04"]))["classification"],"PASS")
|
||||
|
||||
if __name__=="__main__": unittest.main()
|
||||
@@ -0,0 +1,75 @@
|
||||
import argparse, copy, json, tempfile, unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock
|
||||
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||
import src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution as module
|
||||
|
||||
GOLD=Path("tests/gold/target_resolution_v0/cases.json")
|
||||
CASES=module.load_cases(GOLD); BY_ID={c["case_id"]:c for c in CASES}
|
||||
|
||||
def target(case, target_id=None, text="konkrete Zielhandlung"):
|
||||
return {"candidate_observation_id":case["negative_act"]["observation_id"],"target_observation_id":target_id if target_id is not None else case["expected"]["target_observation_id"],"normalized_target_text":text}
|
||||
|
||||
class TargetResolutionTests(unittest.TestCase):
|
||||
def test_eligibility_enum_boundary(self):
|
||||
for cid in ("TR-01","TR-02","TR-03","TR-04"):
|
||||
self.assertTrue(module.eligibility(BY_ID[cid]["negative_act"],BY_ID[cid]["observations"])["eligible_for_target_resolution"])
|
||||
for cid in ("TR-05","TR-06","TR-07","TR-08"):
|
||||
gate=module.eligibility(BY_ID[cid]["negative_act"],BY_ID[cid]["observations"])
|
||||
self.assertFalse(gate["eligible_for_target_resolution"]); self.assertEqual(gate["reason"],"negative_act_form_not_explicit_non_pursuit")
|
||||
|
||||
def test_ineligible_cases_never_call_resolver_and_record_skip(self):
|
||||
fixture={"schema_version":module.SCHEMA_VERSION,"cases":[BY_ID[x] for x in ("TR-05","TR-06","TR-07","TR-08")]}; resolver=Mock()
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
root=Path(tmp); cases=root/"cases.json"; cases.write_text(json.dumps(fixture)); out=root/"out"
|
||||
summary=module.run(argparse.Namespace(cases=cases,output=out,endpoint="x",model="qwen3.5:9B",timeout=1,num_ctx=16384,num_predict=1024),resolver)
|
||||
self.assertEqual(summary["target_resolution_llm_call_count"],0); resolver.assert_not_called()
|
||||
for cid in ("tr-05","tr-06","tr-07","tr-08"):
|
||||
skipped=json.loads((out/cid/"target_resolution_skipped.json").read_text()); self.assertFalse(skipped["call_made"])
|
||||
|
||||
def test_self_contained_target_equals_candidate(self):
|
||||
c=BY_ID["TR-01"]; self.assertEqual(module.validate_target(target(c,text="Zusammenarbeit mit Dr. Schlummer fortsetzen"),c["observations"],"obs_1")["target_observation_id"],"obs_1")
|
||||
|
||||
def test_paired_target_precedes_candidate(self):
|
||||
c=BY_ID["TR-02"]; self.assertEqual(module.validate_target(target(c,text="externe Lösung weiterverfolgen"),c["observations"],"obs_2")["target_observation_id"],"obs_1")
|
||||
|
||||
def test_target_after_candidate_rejected(self):
|
||||
c=copy.deepcopy(BY_ID["TR-02"]); data={"candidate_observation_id":"obs_1","target_observation_id":"obs_2","normalized_target_text":"x"}
|
||||
with self.assertRaises(DerivationValidationError): module.validate_target(data,c["observations"],"obs_1")
|
||||
|
||||
def test_unknown_candidate_and_target_rejected(self):
|
||||
c=BY_ID["TR-02"]
|
||||
with self.assertRaises(DerivationValidationError): module.validate_target({"candidate_observation_id":"missing","target_observation_id":"obs_1","normalized_target_text":"x"},c["observations"],"missing")
|
||||
with self.assertRaises(DerivationValidationError): module.validate_target({"candidate_observation_id":"obs_2","target_observation_id":"missing","normalized_target_text":"x"},c["observations"],"obs_2")
|
||||
|
||||
def test_duplicate_observation_and_evidence_ids_rejected(self):
|
||||
for field in ("observation_id","evidence_id"):
|
||||
obs=copy.deepcopy(BY_ID["TR-02"]["observations"]); obs[1][field]=obs[0][field]
|
||||
with self.assertRaises(DerivationValidationError): module.validate_observations(obs)
|
||||
|
||||
def test_null_and_non_null_text_constraints(self):
|
||||
c=BY_ID["TR-02"]
|
||||
valid={"candidate_observation_id":"obs_2","target_observation_id":None,"normalized_target_text":None}; self.assertEqual(module.validate_target(valid,c["observations"],"obs_2"),valid)
|
||||
for bad in ({"candidate_observation_id":"obs_2","target_observation_id":None,"normalized_target_text":"x"},{"candidate_observation_id":"obs_2","target_observation_id":"obs_1","normalized_target_text":""}):
|
||||
with self.assertRaises(DerivationValidationError): module.validate_target(bad,c["observations"],"obs_2")
|
||||
|
||||
def test_unknown_and_recursive_forbidden_fields_rejected(self):
|
||||
c=BY_ID["TR-02"]
|
||||
for extra in ({"extra":1},{"nested":{"status":"rejected"}}):
|
||||
data=target(c,text="x"); data.update(extra)
|
||||
with self.assertRaises(DerivationValidationError): module.validate_target(data,c["observations"],"obs_2")
|
||||
|
||||
def test_self_contained_prompt_fixes_linkage_deterministically(self):
|
||||
prompt=module.build_prompt(BY_ID["TR-01"]); self.assertIn("deterministically fixed to obs_1",prompt); self.assertIn("only normalize",prompt)
|
||||
|
||||
def test_ineligible_prompt_is_impossible(self):
|
||||
with self.assertRaises(DerivationValidationError): module.build_prompt(BY_ID["TR-05"])
|
||||
|
||||
def test_experiment_has_no_rejection_derivation(self):
|
||||
self.assertFalse(hasattr(module,"derive")); self.assertNotIn("explicitly_rejected",module.TARGET_KEYS); self.assertNotIn("status",module.TARGET_KEYS)
|
||||
|
||||
def test_evaluation_preserves_scope_and_alternative_contract(self):
|
||||
c=BY_ID["TR-04"]; gate=module.eligibility(c["negative_act"],c["observations"]); result=module.evaluate(c,gate,True,target(c,text="Versuch in der realen Anlage durchführen")); self.assertEqual(result["classification"],"PASS"); self.assertTrue(result["alternative_isolation"])
|
||||
|
||||
if __name__=="__main__": unittest.main()
|
||||
@@ -0,0 +1,59 @@
|
||||
import argparse,copy,json,tempfile,unittest
|
||||
from pathlib import Path
|
||||
from unittest.mock import Mock
|
||||
import src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution_v1 as module
|
||||
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||
|
||||
CASES=module.load_cases(Path("tests/gold/target_resolution_v1/cases.json")); BY={c["case_id"]:c for c in CASES}
|
||||
def self_output(text="Zusammenarbeit mit Dr. Schlummer fortsetzen"): return {"candidate_observation_id":"obs_1","normalized_target_text":text}
|
||||
def paired(case,target="obs_1",text="externe Lösung weiterverfolgen"): return {"candidate_observation_id":case["negative_act"]["observation_id"],"target_observation_id":target,"normalized_target_text":text}
|
||||
|
||||
class TargetResolutionV1Tests(unittest.TestCase):
|
||||
def test_self_linkage_is_deterministic_and_equals_candidate(self):
|
||||
result=module.deterministic_self_link(BY["TR1-V1"]); self.assertEqual(result["linkage_source"],"deterministic"); self.assertEqual(result["target_observation_id"],result["candidate_observation_id"])
|
||||
def test_self_llm_output_has_no_target_id(self):
|
||||
self.assertEqual(module.SELF_KEYS,{"candidate_observation_id","normalized_target_text"}); self.assertNotIn("target_observation_id",module.output_schema(BY["TR1-V1"])["properties"])
|
||||
def test_self_combination_records_link_and_normalization_separately(self):
|
||||
combined=module.combine(BY["TR1-V1"],self_output()); self.assertEqual(combined["target_observation_id"],"obs_1"); self.assertIn("fortsetzen",combined["normalized_target_text"])
|
||||
def test_self_link_rejects_paired_strategy(self):
|
||||
with self.assertRaises(DerivationValidationError): module.deterministic_self_link(BY["TR2-V1"])
|
||||
def test_paired_schema_enumerates_allowed_ids_and_null(self):
|
||||
schema=module.output_schema(BY["TR2-V1"]); self.assertEqual(schema["properties"]["target_observation_id"]["enum"],["obs_1","obs_2",None]); self.assertFalse(schema["additionalProperties"])
|
||||
def test_allowed_obs1_and_obs2_are_structurally_accepted(self):
|
||||
case=BY["TR2-V1"]
|
||||
module.validate_semantic_output(paired(case,"obs_1"),case); module.validate_semantic_output(paired(case,"obs_2"),case)
|
||||
def test_unknown_and_string_null_targets_rejected(self):
|
||||
case=BY["TR2-V1"]
|
||||
for value in ("obs_9","null"):
|
||||
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(paired(case,value),case)
|
||||
def test_json_null_accepted_and_requires_null_text(self):
|
||||
case=BY["TR2-V1"]; valid=paired(case,None,None); self.assertEqual(module.validate_semantic_output(valid,case),valid)
|
||||
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(paired(case,None,"x"),case)
|
||||
def test_non_null_requires_nonempty_text(self):
|
||||
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(paired(BY["TR2-V1"],"obs_1",""),BY["TR2-V1"])
|
||||
def test_target_after_candidate_rejected(self):
|
||||
case=copy.deepcopy(BY["TR2-V1"]); case["negative_act"]["observation_id"]="obs_1"
|
||||
with self.assertRaises(DerivationValidationError): module.validate_semantic_output({"candidate_observation_id":"obs_1","target_observation_id":"obs_2","normalized_target_text":"x"},case)
|
||||
def test_duplicate_ids_and_evidence_rejected(self):
|
||||
for field in ("observation_id","evidence_id"):
|
||||
obs=copy.deepcopy(BY["TR2-V1"]["observations"]); obs[1][field]=obs[0][field]
|
||||
with self.assertRaises(DerivationValidationError): module.validate_observations(obs)
|
||||
def test_forbidden_and_unknown_fields_rejected(self):
|
||||
case=BY["TR2-V1"]
|
||||
for extra in ({"status":"x"},{"nested":{"decision":True}},{"extra":1}):
|
||||
data=paired(case); data.update(extra)
|
||||
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(data,case)
|
||||
def test_payload_uses_true_schema_object(self):
|
||||
schema=module.output_schema(BY["TR2-V1"]); payload=module.build_payload("qwen3.5:9B","p",schema,16384,1024); self.assertIs(payload["format"],schema); self.assertIsInstance(payload["format"],dict); self.assertEqual(payload["options"]["temperature"],0)
|
||||
def test_prompt_has_typed_examples_and_allowed_ids(self):
|
||||
prompt=module.build_prompt(BY["TR2-V1"]); self.assertIn('["obs_1", "obs_2"]',prompt); self.assertIn('"target_observation_id":null',prompt); self.assertIn('Never return the string "null"',prompt); self.assertNotIn('observation ID or null',prompt)
|
||||
def test_self_normalization_must_be_positive(self):
|
||||
case=BY["TR1-V1"]; semantic=self_output("Mit Dr. Schlummer arbeiten wir nicht weiter."); combined=module.combine(case,semantic); self.assertEqual(module.evaluate(case,semantic,combined)["classification"],"FAIL")
|
||||
def test_no_rejection_derivation_exists(self):
|
||||
self.assertFalse(hasattr(module,"derive")); self.assertNotIn("status",module.PAIRED_KEYS); self.assertNotIn("explicitly_rejected",module.PAIRED_KEYS)
|
||||
def test_runner_artifacts_distinguish_linkage_and_normalization(self):
|
||||
case=BY["TR1-V1"]; fixture={"schema_version":module.SCHEMA_VERSION,"cases":[case]}; caller=Mock(return_value=(json.dumps(self_output()),{"model":"qwen3.5:9B"}))
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
root=Path(tmp); path=root/"cases.json"; path.write_text(json.dumps(fixture)); out=root/"out"; args=argparse.Namespace(cases=path,output=out,endpoint="x",model="qwen3.5:9B",timeout=1,num_ctx=16384,num_predict=1024); summary=module.run(args,caller); self.assertEqual(summary["llm_call_count"],1); self.assertTrue((out/"tr1-v1"/"deterministic_linkage_result.json").exists()); self.assertTrue((out/"tr1-v1"/"normalized_target_result.json").exists())
|
||||
|
||||
if __name__=="__main__": unittest.main()
|
||||
Some files were not shown because too many files have changed in this diff Show More
Reference in New Issue
Block a user