Author SHA1 Message Date
admin 0cc86deb67 Add Vanheede real-world regression case 2026-09-15 16:43:45 +02:00
admin 21082e66b3 Add GTM Hub real-world regression case 2026-09-12 11:23:18 +02:00
admin 6fc07690d9 Make protocol language follow meeting language 2026-09-11 10:25:45 +02:00
admin 8a0f38fce4 feat: add post-diarization speaker mapping workflow 2026-08-25 15:30:16 +02:00
admin df89a38829 feat: configure 32k context for protocol generation 2026-08-25 15:29:40 +02:00
admin d77bfedb6e Guard protocol generation against context truncation 2026-08-24 16:48:06 +02:00
admin d94436af43 Add canonical audio preparation and meeting context support 2026-08-24 10:22:16 +02:00
admin 8dab928763 Add diarization and reusable MVP meeting pipeline 2026-08-23 19:29:47 +02:00
admin f2d21c1faf Use physical CPU cores for Whisper runtime defaults 2026-08-21 11:11:24 +02:00
admin a9dab7c81a Add direct protocol MVP core 2026-08-20 21:42:36 +02:00
admin 70007ea8f2 Document protocol generation decision 2026-08-20 21:26:09 +02:00
admin 3918c0b1c4 Document target normalization V0 experiment 2026-08-20 14:29:53 +02:00
admin 7fa771a7e4 Document target resolution V1 diagnostic 2026-08-20 14:21:39 +02:00
admin 3229786b5c Document failed target resolution V0 experiment 2026-08-20 13:39:17 +02:00
admin 8ca62fbd92 Document controlled rejection V1 baseline 2026-08-20 13:19:54 +02:00
admin 0d4b426021 Add negative act form experiment 2026-08-20 12:19:55 +02:00
admin 97a22f3ebb Document failed explicit rejection experiment 2026-08-20 11:57:57 +02:00
admin 1c36fa76fb Add collective commitment gold experiment 2026-08-20 09:53:59 +02:00
admin a1fe89de52 Add request-acceptance gold experiment 2026-08-20 09:13:30 +02:00
admin 19672adab4 Add controlled request-acceptance derivation experiment 2026-08-20 08:20:10 +02:00
admin 4ffd4c1c5d Document V3 model comparison 2026-08-19 15:55:01 +02:00
admin 18beb3385f Add evidence-near semantic architecture experiments
Record the V1-V3 experiments and accept the minimal semantic-preservation first stage.
2026-08-19 15:46:22 +02:00
admin bcb197a908 Define topic-oriented protocol architecture 2026-08-11 14:42:18 +02:00
admin c2b7b6b4d2 Document evidence and commitment model 2026-08-11 14:19:19 +02:00
125 changed files with 17371 additions and 21 deletions
+1
View File
@@ -34,6 +34,7 @@ htmlcov/
# Experiment Outputs
experiments/**/output/
experiments/**/results/
artifacts/experiments/**/
# Pipeline runtime artifacts
samples/raw/
+37
View File
@@ -1,5 +1,18 @@
# Project Knowledge
## Meeting language in direct protocols
Direct protocol prompts derive their explicit output language from persisted
Meeting Context `meeting.language`, including regeneration and diarized fallback.
Missing context/language explicitly defaults to `de`; omitted language is accepted
without modifying context data. Protocol runtime `output_language` is derived
provenance, not a separate setting. No transcript/context translation is performed.
Prompt evolution (2026-09-11): replaced unconditional German output with the
meeting-language instruction in the shared prompt builder. Names remain verbatim.
Validation uses mocked generation for German/English, plain/diarized inputs,
regeneration and legacy contexts; no live model or extraction gold run is involved.
This is a compact operational summary of the current Meeting Lab state.
## Objective
@@ -25,7 +38,28 @@ Implemented:
- Prompt loading from `src/meeting_lab/llm/prompts.py`.
- Meeting Context V1 loading, validation and optional extraction prompt
injection with minimal extraction JSON provenance.
- FFmpeg-backed WAV, FLAC and M4A preparation into a per-run canonical mono
16 kHz signed PCM16 WAV artifact before transcription or diarization. Audio
preparation always runs. Optional loudness normalization defaults to on and
currently uses the isolated FFmpeg filter
`loudnorm=I=-16:LRA=11:TP=-1.5`. This is a conservative speech-recording
default and may be revisited after empirical comparison without changing the
orchestration API.
- Interim Markdown protocol generation in `src/meeting_lab/protocol/`.
- Direct protocol prompt input protection: diarized transcripts are rendered as
compact adjacent-speaker blocks without per-segment timestamps. Every source
segment remains represented in order. A deterministic heuristic enforces a
configurable safe input budget, falls back to complete plain transcript text
when necessary, and fails before any Ollama request if even that input is too
large. Silent head/tail truncation is prohibited.
- The `qwen3.8:27b` direct-protocol stage explicitly requests `num_ctx=32768`
and `think=false`; the practical prompt target is approximately 29,000 tokens.
A 31,038-token synthetic prompt passed, but larger prompts are not assumed safe
from the model's advertised 262,144-token native context alone.
- `regenerate_mvp_protocol` updates the run's validated Meeting Context and
regenerates protocol artifacts from the existing diarized transcript when
available. It never reruns audio preparation, Whisper or Pyannote, and it
preserves anonymous speaker labels in the source transcript.
- Non-LLM unit tests for chunking, extraction helpers, protocol rendering and
gold-test runner validation.
- Meeting Context V1 scaffold and documentation for manually maintained
@@ -139,6 +173,9 @@ departments only when they are explicitly supplied as metadata. It must not be
used to infer responsibilities. In the current implementation this context can
be injected into chunk extraction prompts as authoritative metadata, and only
minimal provenance is written to extraction JSON.
The implemented MVP statuses are exactly `present` and `mentioned_only`.
Legacy entries without a status receive collection-appropriate defaults. Only
present participants may be targets of explicit `SPEAKER_XX` mappings.
A `responsible` or future `owner` / `assignee` value may be recorded only when
source evidence explicitly assigns, accepts or confirms responsibility. If the
+55 -1
View File
@@ -1,5 +1,12 @@
# Meeting Lab
The direct protocol uses saved Meeting Context `meeting.language`: `de` requests
German output and `en` requests English output, for plain and diarized transcripts
and protocol-only regeneration. Meeting Assistant supplies the same selection to
Whisper. Missing context/language retains German output for older runs. Effective
output language is recorded as `output_language` in protocol runtime metadata.
Transcript text, authored context, names and speaker mappings are not translated.
Experimentierumgebung zur Entwicklung eines lokalen Diskussionsanalyzers für Meetingtranskripte.
## Ziel
@@ -10,7 +17,54 @@ Im Mittelpunkt steht nicht die Softwarearchitektur, sondern die Frage:
> **Wie lässt sich aus einem realen Meeting möglichst zuverlässig strukturiertes Wissen extrahieren?**
Neue Ideen werden zunächst hier experimentell umgesetzt. Erst wenn sich ein Ansatz bewährt hat, wird er in den eigentlichen *Meeting Assistant* übernommen.
Meeting Lab ist die experimentelle R&D-Umgebung fuer den zukuenftigen
*Meeting Assistant*. Es dient gleichzeitig als Forschungsplattform,
Architektur-Spielwiese, Regressionsframework, Benchmark-Umgebung und
Prototypimplementierung. Neue Ideen werden hier untersucht und gegen reale
Meeting-Beispiele validiert, bevor sie fuer das Produkt in Betracht kommen.
Der beabsichtigte Reifeprozess ist:
```text
Research idea
->
Meeting Lab experiment
->
Regression tests
->
Stable architecture
->
Meeting Assistant implementation
```
Nur ausreichend ausgereifte und verifizierte Komponenten sollen in den
Meeting Assistant uebernommen werden. Meeting Lab darf bewusst experimentelle
Ansaetze und Entwicklungszweige enthalten, die verworfen werden oder nie den
Assistant erreichen.
## Meeting Assistant und langfristige Produktentwicklung
Der Meeting Assistant ist als Produktionsanwendung vorgesehen. Seine erste
oeffentliche Beta soll auf einem stabilen Meeting-Lab-MVP basieren und eine
polierte User Experience, Installer, Konfiguration und eine produktionsreife
Pipeline bieten. Eine GUI ist optional; experimentelle Funktionen sollen
standardmaessig nicht aktiviert sein.
Meeting Lab entwickelt sich unabhaengig weiter und bleibt der langfristige
Innovationszweig. Der erwartete Transferpfad lautet:
```text
Meeting Lab Alpha
->
Meeting Lab Beta
->
Meeting Assistant Beta
->
Meeting Assistant Release
```
Meeting-Assistant-Releases uebernehmen damit gezielt bewaehrte Meeting-Lab-
Komponenten, waehrend der Assistant den stabilen Produktzweig bildet.
---
+206 -2
View File
@@ -2,13 +2,63 @@
## Purpose
The **Meeting Lab** is an experimental environment for developing and evaluating methods to extract structured knowledge from real meeting transcripts.
The **Meeting Lab** is the experimental R&D environment for developing and
evaluating methods to extract structured knowledge from real meeting
transcripts. It is the research platform, architecture playground, regression
framework, benchmark environment and prototype implementation for the future
Meeting Assistant.
Its purpose is not to build a complete meeting assistant, but to answer a single question:
> **How can knowledge be extracted from real discussions as reliably as possible?**
Successful approaches will later be integrated into the Meeting Assistant project.
Its purpose is to validate ideas before they are promoted into the product.
Only sufficiently mature and verified components should migrate into Meeting
Assistant. Meeting Lab may intentionally contain experiments or development
branches that are rejected, remain inconclusive or never reach the Assistant.
## Meeting Lab and Meeting Assistant lifecycle
Meeting Lab and Meeting Assistant have different long-term responsibilities:
- **Meeting Lab** is the long-term innovation branch. It favors learning,
inspectable experiments, regression evidence, benchmarks and architectural
change.
- **Meeting Assistant** is the stable product branch. It favors a polished user
experience, installation, configuration and a production pipeline.
Architectural promotion follows an evidence-based lifecycle:
```text
Research idea
->
Meeting Lab experiment
->
Regression tests
->
Stable architecture
->
Meeting Assistant implementation
```
The expected release progression is:
```text
Meeting Lab Alpha
->
Meeting Lab Beta
->
Meeting Assistant Beta
->
Meeting Assistant Release
```
The first public Meeting Assistant beta should be based on a stable Meeting
Lab MVP. It should provide a polished user experience, an installer,
configuration and a production-quality pipeline. A GUI is optional.
Experimental features should not be enabled by default. Meeting Lab continues
to evolve independently after components have migrated; promotion does not
turn the Lab itself into the product branch.
---
@@ -356,6 +406,144 @@ The architecture document only describes the overall system.
---
# Version 2 Accepted Architectural Direction
The following topics are accepted architectural goals for Version 2. They
record direction reached through the BUG-011 through BUG-015 investigations;
they are not descriptions of implemented behavior or authorization to change
the current pipeline.
## Speaker diarization before semantic analysis
Version 2 should determine **who is speaking before semantic analysis**.
Speaker identity contains evidence that cannot reliably be reconstructed from
text alone. It helps distinguish, for example, who answers a question, accepts
work, agrees with a proposal, or advances the discussion after another
speaker. It also preserves conversational flow that anonymous transcript text
can erase.
Diarization is therefore a semantic prerequisite in the intended Version 2
architecture, not merely a display enhancement. Its output should remain
traceable to transcript segments so later stages can preserve speaker and
source provenance.
## Persistent speaker identification
Version 2 should add a persistent speaker database and an interactive identity
workflow during import:
```text
Unknown speaker detected
->
Representative audio sample (approximately 20 seconds)
->
User selects an existing identity or creates a new identity
->
Known speaker available for future recognition
```
Automatic recognition may suggest an identity, but user confirmation governs
the persistent association. Over the long term, speaker embeddings rather
than raw meeting recordings should be the persistent recognition
representation. Representative raw audio is an import and confirmation aid,
not the intended durable identity store. Privacy, deletion and false-match
handling require separate design before implementation.
## Meeting Context as a probabilistic prior
Known meeting participants should influence semantic interpretation, but
Meeting Context is a **probabilistic prior**, not a deterministic semantic
rule. It can make one interpretation more plausible and help focus review; it
must never manufacture a commitment, decision or responsibility assignment.
For example, if Marleen is confirmed as present, “Marleen müsste sich mal
äußern” is more likely to be conversation management directed at a current
participant than future project work. Presence alone does not prove this
interpretation, and it does not establish an Action Item or responsibility.
Explicit meeting evidence remains authoritative. This extends, rather than
weakens, the responsibility attribution invariant.
The existing Meeting Context V2 entity direction is documented in
[`adr-meeting-context-v2-entity-registry.md`](adr-meeting-context-v2-entity-registry.md).
Speaker identities and meeting-specific participant confirmation should
eventually feed that context without turning registry metadata into semantic
facts.
## Conversation Management versus Meeting Content
BUG-015 reinforced that not every utterance is protocol-worthy content.
Version 2 should conceptually distinguish:
- **Conversation Management**: utterances that coordinate the meeting itself,
such as asking a present participant to speak, moderation, requesting a
slide, or asking someone to repeat something.
- **Meeting Content**: propositions that may contribute to the meeting's
durable knowledge, including facts, technical findings, decisions, action
items and open questions.
This is an architectural concept, not a currently implemented category or
filter. The distinction should prevent conversational coordination from being
promoted into project commitments while retaining sufficient provenance to
understand dialogue. Context and diarization can inform the distinction, but
neither should act as a deterministic keyword or participant rule.
## Evidence and commitment before protocol eligibility
The BUG-015 design study concludes that semantic state should be classified
before policy determines whether an item is eligible for a protocol. A binary
keep/reject verifier conflates evidence recognition with publication policy
and loses valid intermediate states.
The proposed decision progression is:
```text
idea -> option -> proposal -> preferred option -> tentative agreement -> decision
```
The proposed action progression is:
```text
possible next step -> recommendation -> requested action
-> established action -> ongoing work -> completed
```
Questions combine a communicative **kind** with an independent **resolution
state**, rather than treating every uncertainty or interrogative as an Open
Question. Responsibility remains an independent dimension and may be recorded
only when explicitly assigned, accepted or confirmed. Evidence strength is
also independent: it describes support for a semantic label, not semantic
maturity or protocol eligibility.
After semantic state classification, explicit policy should select Decisions,
Action Items and Open Questions for a particular output view. Renderers should
receive policy-selected semantic content and must not promote proposals or
conversation management into commitments. The complete taxonomy, trade-offs,
architecture interactions and migration questions are recorded in
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md).
## Topic-oriented primary protocol
One of the highest-level Version 2 requirements is:
> **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.**
Semantic categories remain metadata and supporting structure within topics;
they must not dictate the main document structure. The conceptual target flow
is `Transcript -> Evidence Extraction -> Topic Reconstruction -> Semantic
Synthesis -> Protocol Rendering`. The primary renderer should eventually
receive topic-oriented semantic knowledge. Category-oriented Action Item,
Decision, Open Question and management views remain useful derived outputs.
The detailed rationale, example, architectural implications and explicit
deferral of a final Topic Reconstruction schema are documented in
[`design/evidence_commitment_model.md`](design/evidence_commitment_model.md#thematic-protocol-as-the-primary-structure).
Implementation is intentionally postponed until this architecture has been
reviewed. No Version 2 goal in this section changes current extraction,
canonicalization, consolidation or rendering behavior.
---
# Current State
Implemented:
@@ -389,6 +577,22 @@ Action Item may have no known owner, while a named owner requires explicit
assignment, volunteering or acceptance. These are semantic LLM classifications;
deterministic validation must not guess intent from keywords.
### Classification Verifier
Before extraction output is normalized for Canonicalizer input, Decision,
Action Item and Open Question candidates pass through a semantic precision
gate. Facts and technical details pass through unchanged. Each candidate is
reviewed independently with its evidence and bounded local chunk context; the
verifier may only keep or reject the existing candidate. It cannot add or
rewrite semantic items.
Verifier output contains the stable candidate ID, category, `keep|reject`
verdict, evidence-based reason and `responsibility_supported`. A kept Action
Item with an unsupported named owner is retained with its responsibility
cleared. Malformed output fails the extraction verification substage closed,
after preserving candidate input and raw response. Per-candidate results and an
aggregate audit remain traceable before Canonicalizer input is written.
## Semantic Consolidator failure handling
Semantic Consolidator V0 preserves every raw model response before parsing.
+369
View File
@@ -0,0 +1,369 @@
# Evidence / Commitment Model
Status: architecture design study for BUG-015; no implementation recommendation is made here.
## Motivation
Meeting Lab currently asks extraction and the BUG-015 verifier to cross a category boundary in one step: a candidate is either a protocol-worthy Decision, Action Item or Open Question, or it is rejected. That combines two different questions:
1. What does the transcript provide evidence for?
2. Which evidenced states should a particular protocol view publish?
Meetings do not move directly from absence to commitment. An option can become a proposal, then a preferred option, then a tentative agreement and finally a decision. Work can be suggested, requested, assigned, accepted, already underway or completed. A question can be asked and answered, or can remain explicitly unresolved. Rejecting everything below the final protocol threshold discards useful evidence; promoting it creates false commitments.
The proposed model therefore records the evidenced semantic state first. A later policy selects protocol-worthy states. It preserves the existing separation between extraction, deterministic canonicalization, semantic consolidation and rendering, and it does not make rendered output the semantic source of truth.
## Observed BUG-015 failures
The BUG-015 Gold cases expose both sides of the binary-classification problem.
| Progeo-derived evidence | Correct evidence state | Binary failure to avoid |
| --- | --- | --- |
| “Die Option steht im Raum, das Material chemisch recyceln zu lassen.” | option | False Decision |
| “Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden.” | personal preference with a conditional option | False Decision |
| “Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht. … das ist entschieden.” | explicit decision: rejection | False rejection |
| “Die Geometrie kann man vielleicht noch optimieren.” | possible next step | False Action Item |
| “Marleen müsste vielleicht mal äußern …” | suggested/requested action; no accepted or confirmed responsibility | False Action Item and false owner |
| “Wir könnten Dirk Textor vielleicht noch einmal kontaktieren.” | suggestion | False Action Item |
| “Nina, übernimmst du …? – Ja, ich übernehme …” | accepted action with explicit responsibility and deadline | False rejection |
| “Den CET-Artikel erstellen wir bereits …” | ongoing established work; owner not evidenced | Consistent false rejection by the binary verifier |
| “Ob sich das Waschen lohnt, weiß ich nicht.” | uncertainty | False Open Question |
| “Gibt es schon ein Programm? – Ja …” | explicit, resolved question | False Open Question if local resolution is ignored |
| “Welche Daten … dürfen wir veröffentlichen? … weiterhin ungeklärt.” | explicit unresolved question | Consistent false rejection by the binary verifier |
The verifier was reliable on several negative cases and on explicit commitment language, but not on valid states whose evidence did not resemble a fresh agreement: already-established work and an explicitly unresolved information need. This suggests a representation problem, not merely an insufficient keep/reject prompt.
## Model shape
The model is deliberately small and compositional. Each evidence item has:
- a **domain**: decision, action or information need;
- a **semantic state** within that domain;
- an **evidence support level** describing how directly the transcript supports that label;
- preserved transcript evidence and source references;
- domain-specific independent attributes, such as responsibility or resolution;
- no automatic claim that the item belongs in a final protocol.
Semantic maturity and evidence support are independent. An explicitly worded proposal remains a proposal; it is not a weak decision. Conversely, ongoing work can be strongly evidenced without a recorded moment of assignment.
## Semantic state diagrams
The arrows show common progressions, not mandatory workflows. Meetings may enter at any state, skip states, regress, or end without commitment.
### Decision evolution
```text
idea -> option -> proposal -> preferred_option -> tentative_agreement -> decision
| | | | |
+----------+--------------+--------------------+----------> withdrawn
superseded
reopened
```
An **idea** is an undeveloped possibility. An **option** is a candidate alternative. A **proposal** asks the group to adopt an outcome. A **preferred option** expresses comparative preference without settlement. A **tentative agreement** records provisional convergence that is explicitly conditional or awaiting confirmation. A **decision** records an outcome the meeting settled, selected, approved, rejected or committed to. The intermediate preferred and tentative states matter because they prevent likely direction from being confused with commitment.
### Action evolution
```text
possible_next_step -> recommendation -> requested_action -> established_action -> ongoing_work -> completed
| | | | |
+------------------+-----------------+--------------------+----------> cancelled
Responsibility (independent):
unset -> proposed_responsible -> assigned -> accepted/confirmed
```
An action can enter directly as **established_action** through explicit assignment, acceptance or commitment. It can enter directly as **ongoing_work** when the transcript clearly says that the work is already being performed, as in the CET example. `proposed_responsible` is evidence about a suggestion, not permission to populate the protocol's responsible field. Only `assigned`, `accepted` or `confirmed` supports recorded responsibility, and only where the meeting evidence explicitly assigns, accepts or confirms it.
### Information-need evolution
```text
uncertainty ---------> information_need(unresolved) ---------> resolved
curiosity -----------> explicit_question(unresolved) --------> resolved
request_for_clarification(unresolved) -----------------------> resolved
missing_information(unresolved) -----------------------------> resolved
rhetorical_question -> discourse only
```
Uncertainty and curiosity do not automatically create an information need. An explicit question may be resolved immediately. “Open Question” is therefore a policy result derived primarily from a qualifying kind plus `unresolved`, rather than a primitive utterance category.
## Proposed taxonomy
### Common fields
| Field | Proposed values | Purpose |
| --- | --- | --- |
| `domain` | `decision`, `action`, `information_need` | Selects the domain taxonomy. |
| `state` | Domain-specific values below | Records what the evidence says now. |
| `support` | `indirect`, `direct`, `explicit` | Rates support for the chosen state, not protocol importance. |
| `polarity` | `positive`, `negative` where applicable | Preserves approval/rejection and adoption/refusal without rewriting meaning. |
| `evidence` | transcript excerpt(s) | Proves the state and attributes. |
| `source_refs` | stable source references | Retains provenance across stages. |
`support` is intentionally not `weak / medium / strong`. Those terms mix confidence, semantic maturity and number of witnesses. The proposed meanings are narrower:
- `indirect`: the state is supported by context but not stated in a self-contained utterance; downstream commitment policy should normally be conservative.
- `direct`: an utterance directly expresses the state, such as a proposal, preference, ongoing-work statement or question.
- `explicit`: the utterance also names the decisive status, for example “das ist entschieden”, “ich übernehme”, “wir erstellen bereits” or “weiterhin ungeklärt”.
This common axis simplifies evidence auditing and threshold policy, but cannot replace domain state. An explicit suggestion is still not an Action Item, and an explicit uncertainty is still not an Open Question. Model confidence, if retained at all, must be a separate operational field and must not be presented as evidence strength.
### Decision states
| State | Meaning | Default protocol treatment |
| --- | --- | --- |
| `idea` | Undeveloped possibility or brainstorming contribution | Exclude |
| `option` | Candidate alternative under consideration | Exclude |
| `proposal` | Outcome offered for adoption | Exclude |
| `preferred_option` | Expressed preference among alternatives | Exclude |
| `tentative_agreement` | Provisional convergence with an expressed condition or pending confirmation | Exclude from Decisions; potentially expose to an editorial view |
| `decision` | Settled, selected, approved, rejected or committed outcome | Include as Decision when evidence is sufficient |
| `reopened` | Earlier decision explicitly returned to unresolved consideration | Do not render the earlier decision as currently settled without qualification |
| `superseded` | Earlier decision replaced by a later one | Retain provenance; normally render only the current decision |
| `withdrawn` | Candidate state explicitly withdrawn | Exclude as current commitment |
`idea`, `option`, `proposal`, `preferred_option`, `tentative_agreement` and `decision` are necessary distinctions for BUG-015. `reopened`, `superseded`, `withdrawn` are lifecycle states needed to avoid treating historical evidence as current policy; they need not be first-iteration extraction targets.
### Action states and attributes
| State | Meaning | Default protocol treatment |
| --- | --- | --- |
| `possible_next_step` | Hypothetical or exploratory action | Exclude |
| `recommendation` | Action advocated but not established as work | Exclude |
| `requested_action` | Someone asks that work be done, without enough evidence that it is established | Exclude by default |
| `established_action` | Concrete future work established by assignment, acceptance, commitment or confirmation | Include as Action Item |
| `ongoing_work` | Concrete work explicitly already underway | Include when still relevant; do not require a newly witnessed assignment |
| `completed` | Work explicitly reported complete | Exclude from open Action Items; retain as status/history |
| `cancelled` | Work explicitly cancelled or declined | Exclude from open Action Items; retain provenance |
`suggestion` is represented as `possible_next_step` or `recommendation`, depending on whether the speaker advocates it. `accepted_action` and `assigned_work` should not be competing lifecycle states: both establish `established_action`, while the independent commitment basis records `accepted`, `assigned`, `self_committed` or `confirmed_existing`. This avoids an artificial choice when, as with Nina, an assignment and acceptance occur together.
Responsibility is independent:
| Responsibility status | Meaning | May populate `responsible`? |
| --- | --- | --- |
| `unset` | No person/team explicitly tied to ownership | No |
| `proposed` | A person is suggested, asked speculatively or mentioned near the work | No |
| `assigned` | The meeting explicitly assigns the work | Yes |
| `accepted` | The party explicitly accepts or volunteers | Yes |
| `confirmed` | Existing ownership is explicitly confirmed | Yes |
This preserves the project invariant: discussion, expertise, adjacency, organizational role and likely ownership never establish responsibility. An action may be valid with `responsibility_status: unset`, as with established CET work.
### Information-need kinds and resolution
Question form and resolution are separate attributes.
| Kind | Meaning | Can become an Open Question? |
| --- | --- | --- |
| `uncertainty` | Speaker expresses doubt or lack of certainty without establishing a concrete need | No, by itself |
| `curiosity` | Interest without a concrete need requiring follow-up | No, by itself |
| `explicit_question` | Direct interrogative seeking an answer | Yes, if unresolved |
| `request_for_clarification` | Explicit request to clarify a concrete matter | Yes, if unresolved |
| `missing_information` | Concrete required information is stated as absent | Yes, if unresolved |
| `rhetorical_question` | Interrogative used for emphasis rather than an answer | No |
Resolution is one of `unresolved`, `resolved`, or `resolution_unclear`. `resolved_question` is therefore not a separate kind: it is, for example, `explicit_question + resolved`. The publication example is `explicit_question + unresolved`; the event-program example is `explicit_question + resolved`; the washing example is `uncertainty` unless later evidence establishes a concrete unresolved need. `resolution_unclear` preserves evidence but should not silently pass a precision-first Open Question policy.
An “unresolved question” is a derived, protocol-relevant combination rather than a fourth kind. This makes the classification testable: kind answers what communicative act occurred, and resolution answers what remained at the end of the available context.
## Progeo examples under the model
| Evidence | Proposed representation | Working Protocol policy result |
| --- | --- | --- |
| Chemical recycling “Option steht im Raum” | `decision / option / explicit` | Not a Decision |
| Technikum preference | `decision / preferred_option / direct`; personal scope preserved | Not a Decision |
| Schlummer collaboration rejected and “entschieden” | `decision / decision / explicit / negative` | Decision |
| Geometry “kann man vielleicht … optimieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item |
| Marleen “müsste vielleicht mal” | `action / requested_action / direct`, responsibility proposed only | Not an Action Item; no responsible party |
| Textor “könnten … kontaktieren” | `action / possible_next_step / direct`, responsibility unset | Not an Action Item |
| Nina asks and accepts the review by Friday | `action / established_action / explicit`, basis assigned + accepted, responsible Nina, deadline Friday | Action Item |
| CET article “erstellen wir bereits” and later circulation | `action / ongoing_work / explicit`, responsibility unset | Action Item without invented owner |
| Washing “ob sich das lohnt, weiß ich nicht” | `information_need / uncertainty / direct`, resolution unclear | Not an Open Question |
| Event program asked and answered | `information_need / explicit_question / direct`, resolved | Not an Open Question |
| Publishable energy-audit data “weiterhin ungeklärt” | `information_need / explicit_question / explicit`, unresolved | Open Question |
These labels preserve every BUG-015 distinction without treating rejected protocol candidates as meaningless.
## Thematic protocol as the primary structure
The Evidence / Commitment Model supplies semantic distinctions inside a larger
meeting reconstruction. It does not imply that the primary protocol should be
organized as one section per semantic category.
The central Version 2 requirement is:
> **The protocol is primarily a topic-oriented reconstruction of the meeting, not a category-oriented listing of extracted information.**
A primary protocol should reconstruct which topics were discussed, what
relevant information emerged within each topic, how the discussion developed,
which alternatives, ideas, objections or proposals mattered, what outcome or
current state was reached, and which actions or unresolved questions resulted
from that topic.
Semantic categories remain essential, but as metadata and supporting structure
attached to topics. Facts, technical findings, ideas, alternatives, proposals,
objections, decisions, action items and open questions must not dictate the
main document structure.
For example, a human-style topic section may read:
```text
## Trial setup
Several variants for the next trial were discussed.
A thinner carrier material was proposed as one possible alternative.
Concerns were raised regarding its mechanical suitability.
The group therefore decided to continue with the existing setup for the next trial.
Nina will obtain the remaining samples before the next production run.
The publication question remains unresolved.
```
This preserves the relationship between the proposal, objection, decision,
resulting work and unresolved question. A primary document split into separate
Ideas, Objections, Decisions and Action Items sections would lose that thematic
and conversational relationship.
Category-oriented outputs remain valuable as derived secondary views. An
Action Item table, Decision register, Open Questions list or Management summary
can be generated from the same underlying meeting knowledge after thematic
reconstruction. These indexes and summaries do not replace the primary
topic-oriented protocol.
The conceptual Version 2 flow is:
```text
Transcript
->
Evidence Extraction
->
Topic Reconstruction
->
Semantic Synthesis
->
Protocol Rendering
```
Evidence items and their semantic metadata remain attached to topics and
support Semantic Synthesis. The Version 2 renderer should receive
topic-oriented semantic knowledge rather than a flat collection grouped by
category. This direction does not define the final Topic Reconstruction schema,
redesign the current pipeline or remove the existing Working Protocol V2
architecture. It is an accepted target architecture whose implementation is
postponed.
## Architectural impact
### Extraction
Extraction could emit evidence records with richer labels instead of immediately claiming `decision`, `action_item` or `open_question`. This is a bounded increase in extraction vocabulary, not a request for larger context windows, a multi-chunk strategy or a monolithic synthesis prompt. Existing Facts, Positions and Technical Details remain distinct categories; the new domains refine only commitment-sensitive content.
The immediate output would describe the observed state. Protocol category selection would happen later through an explicit policy. Some policy rules can be deterministic once semantic labels are trustworthy, for example:
```text
Decision := domain=decision AND state=decision
Action Item := domain=action AND state IN {established_action, ongoing_work}
Open Question := domain=information_need
AND kind IN {explicit_question, request_for_clarification, missing_information}
AND resolution=unresolved
Responsible := responsibility_status IN {assigned, accepted, confirmed}
```
This keeps downstream selection simple, but does not make semantic extraction deterministic. The LLM still has to distinguish proposals from decisions and uncertainty from unresolved needs.
### LLM stability
The architecture should reduce one source of instability: the model no longer has to erase a supported proposition merely because it falls below a protocol threshold, nor relocate it into another category to retain it. Labels correspond more closely to observable speech acts and lifecycle statements, and downstream thresholds become explicit.
It will not eliminate instability. More labels create adjacent-class boundaries, local context may not reveal resolution, and indirect language remains difficult. Stability depends on small schemas, evidence spans, one best supported state, conservative handling of `resolution_unclear`, and later evaluation against the Gold corpus. A generic evidence-strength score alone would likely worsen instability because it invites subjective grading.
### Canonicalizer
The Canonicalizer would benefit from carrying normalized state names, independent responsibility fields, resolution, polarity and stable source references. It could deterministically validate allowed combinations, normalize aliases, preserve evidence, group exact duplicates and reject structurally impossible combinations. It must not promote a proposal to a decision, infer resolution, infer responsibility or perform uncertain semantic merging.
Richer labels make safe comparisons easier: two `option` records can be recognized as candidates about the same subject without being merged into a `decision`; an `explicit_question + resolved` item is not confused with an unresolved instance. Provenance improves because a later decision can link back to earlier option/proposal evidence rather than overwriting it.
Semantic merging becomes better informed, but not automatically easy. Equivalence and lifecycle transitions remain semantic work. A future consolidator should preserve all source evidence, distinguish duplicate evidence from state evolution, and mark contradictions or uncertainty rather than collapsing them.
### Renderer and output policy
The renderer should not make commitment classification decisions. A
policy/view-selection stage should determine protocol-worthy semantic states
while preserving their topic associations. In the Version 2 target
architecture, the primary protocol renderer should receive topic-oriented
semantic knowledge containing eligible evidence and state transitions, not a
flat category-oriented collection. It must not silently promote raw proposals
to Decisions or present conversational coordination as Meeting Content.
Proposals may still be visible to a renderer for a view explicitly designed to show them, such as a discussion appendix or editorial drafting view. That access must be typed and intentional; a renderer must never silently format a proposal under “Decisions.” Different output views may select different states while sharing the same Canonical Meeting Knowledge.
Category-oriented renderers remain valid for secondary views such as Action
Item tables, Decision registers and Open Questions lists. This accepted
direction supplements rather than removes the current Working Protocol V2
architecture; it does not change current rendering behavior.
### Future BPD-style protocol generation
Richer labels would give editorial generation better material without granting it authority to invent commitment. A BPD-style view could explain the path from options through tentative agreement to decision, separate current obligations from completed work, and describe why an information need remains open. Proposals and preferred options could be useful appendix material when the requested view values discussion history.
The editorial layer may condense or order these records, but it must preserve state, polarity, responsibility status and provenance. Appendix inclusion is a view policy, not a semantic promotion.
## Trade-offs
Benefits:
- preserves valid evidence below the final-protocol threshold;
- separates meeting semantics from publication policy;
- represents established work without inventing an assignment event or owner;
- treats question kind and resolution independently;
- makes responsibility independently auditable;
- improves provenance across lifecycle changes;
- prevents renderers from silently converting proposals into commitments;
- gives future editorial views useful, explicitly non-final material.
Costs and risks:
- expands the extraction schema and Gold specification;
- introduces adjacent semantic labels that require precise definitions;
- requires consolidation to distinguish duplicates from state transitions;
- may retain more evidence records, increasing storage and review volume;
- depends on adequate local context to determine resolution and lifecycle state;
- requires explicit view policies so consumers do not treat every evidence record as protocol-worthy;
- cannot guarantee LLM consistency merely by replacing a binary verdict with a taxonomy.
The taxonomy should remain closed and small. In particular, modality, confidence, support, commitment state and responsibility must not be collapsed into a single scalar.
## Migration strategy
This is a staged design path, not a recommendation to implement it now.
1. Specify the taxonomy and invariants against the existing BUG-015 Gold candidates and the cited Progeo excerpts, without changing current expected outputs.
2. Annotate a design-only mapping from each current Decision, Action Item and Open Question example to evidence state, support and independent attributes. Confirm objectively unique ground truth, especially for `requested_action` versus `possible_next_step` and `tentative_agreement` versus `preferred_option`.
3. Define a versioned evidence-record schema and deterministic protocol-selection policy on paper. Keep responsibility and question resolution independent.
4. Design evaluation measures separately for state accuracy, resolution accuracy, responsibility accuracy and protocol-policy output. A correct lower state must not count as a false protocol commitment.
5. Only after the design is validated, plan an experiment that compares the evidence-first approach with current extraction and the binary verifier. Any prompt or Python changes would be separate, explicitly scoped tasks following the Gold methodology and LLM safety rules.
6. If later adopted, preserve compatibility through an adapter that maps protocol-worthy evidence states to the current canonical categories while Canonical Meeting Knowledge evolves. Do not ask the renderer to interpret raw states.
No database, prompt, code, test or current output-schema migration is proposed by this document.
## Open questions
- Is `requested_action` useful as a distinct state when a direct assignment already creates `established_action`, or should it be limited to unaccepted requests?
- Does `tentative_agreement` have sufficiently unique evidence in the current corpus, or should it remain an annotation until more Gold examples exist?
- Should an ongoing-work item whose relevance to the meeting is unclear pass the Working Protocol policy, or require an explicit continuation/follow-up signal?
- How much local context is required to label a question `resolved` safely when the answer occurs across a technical chunk boundary?
- Should `resolution_unclear` be retained only in Canonical Meeting Knowledge, or also exposed to a review view?
- Should support be limited to `direct / explicit`, treating `indirect` as review-only, to reduce subjective classification?
- How should consolidation represent one subject moving from proposal to decision: linked immutable evidence records, or a current-state object with a preserved event history?
- Which output views, if any, should include withdrawn proposals, cancelled work and resolved questions?
- Can polarity and lifecycle links be normalized deterministically without introducing semantic inference?
## Recommendation
The evidence-first architecture should replace the current binary keep/reject classification approach as the target architecture. The replacement should be conceptual and staged, not implemented yet: extract a small, domain-specific semantic state plus independent support, responsibility and resolution attributes; preserve evidence and provenance; then apply explicit policy to produce protocol categories.
A single common Evidence Strength axis should complement this taxonomy but must not replace it. The decisive improvement is separation of semantic state from protocol eligibility. That separation directly explains the BUG-015 false positives and false negatives, supports the Progeo cases without invented ownership, simplifies view selection, and provides a sounder basis for Canonical Meeting Knowledge and future BPD-style rendering.
+62
View File
@@ -0,0 +1,62 @@
# Optional Speaker Diarization
The direct-protocol MVP keeps speaker diarization disabled by default. Enable
anonymous Community-1 speaker labels with `--diarization auto`, `gpu`, or
`cpu`:
```bash
python3 scripts/run_mvp_meeting.py meeting.wav \
--whisper-model /path/to/ggml-model.bin \
--diarization auto
```
Native mode (the default runtime) requires a compatible local PyTorch and
`pyannote.audio==4.0.7`. For isolated ROCm/CUDA environments, select the
container runtime and provide its image and hardware arguments explicitly:
```bash
python3 scripts/run_mvp_meeting.py meeting.wav \
--whisper-model /path/to/ggml-model.bin \
--diarization gpu \
--diarization-runtime container \
--diarization-container-image IMAGE \
--diarization-container-arg=--device=/dev/kfd \
--diarization-container-arg=--device=/dev/dri \
--diarization-container-arg=--group-add \
--diarization-container-arg=video
```
The container receives `HF_TOKEN` by environment-variable name only. It mounts
the source audio and repository read-only and writes diarization artifacts into
the current run directory. Meeting Lab loads mono 16 kHz PCM16 WAV with
Python's `wave` module and sends an in-memory tensor to pyannote, avoiding its
torchcodec file decoder.
Anonymous `SPEAKER_XX` labels are aligned to Whisper segments by maximum
temporal overlap with Community-1 exclusive diarization. The original Whisper
transcript is preserved; the derived transcript under `diarization/` is used as
the direct-protocol generator's source input.
The full diarized JSON and timestamped text remain immutable audit artifacts,
but their per-segment formatting is too verbose for a full-meeting LLM prompt:
timestamps and repeated speaker labels can more than double input size. For
protocol generation, Meeting Lab deterministically groups only adjacent
segments assigned to the same anonymous speaker and omits timestamps. A later
return by the same speaker starts a new block, and unassigned segments remain
under `SPEAKER_UNASSIGNED`. `protocol/transcript_input.txt` preserves the exact
derived representation sent to prompt construction.
Before contacting Ollama, Meeting Lab conservatively estimates prompt tokens
from UTF-8 byte count without adding a model tokenizer dependency. The default safe
budget is 29,000 estimated tokens within the explicitly configured 32,768-token
Ollama context. The estimate is calibrated against the currently validated
German BPD input and is configurable through
`MvpMeetingConfig.protocol_safe_input_token_budget` or
`--protocol-safe-input-token-budget`.
If compact diarized input exceeds the budget, the generator deterministically
uses the complete plain segment transcript and records the fallback. If that
also exceeds the budget, generation fails before model lookup or generation;
it never truncates, chunks, summarizes, retries, or makes multiple protocol
calls implicitly. Full diarization artifacts are never overwritten by this
selection.
+938
View File
@@ -1350,3 +1350,941 @@ accordance with the Gold Standard methodology.
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
remains Open; no phrase-specific deterministic filter was introduced.
## EXP-0027 — Evidence-near observation extraction
Date: 2026-08-18
Hypothesis: `qwen3.5:9B` can more reliably extract evidence-near linguistic and
semantic properties than directly synthesize protocol-level events, outcomes,
actions and unresolved issues. This isolated experiment stops before semantic
interpretation and does not connect to the production pipeline.
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
source evidence, explicit targets, a five-value relation vocabulary, modality,
temporality, evaluation, agreement, responsibility/person, uncertainty,
clarification need and free-text scope. It contains no protocol-level category
field. The validator requires sequential observation IDs, known evidence IDs,
backward-only valid observation targets, closed categorical vocabularies,
consistent responsibility/person pairs, and one canonical absence form: JSON
null for person and `absent` for scope.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries. All nine cases ran exactly once, for nine LLM
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
responses, parsed observations, validation results, Ollama metadata and
comparisons are preserved under
`/tmp/meeting-lab-evidence-observations-v1-20260818/`.
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
Five cases failed structure because the model represented a single target as a
one-element list, usually `["discussion_subject"]`; the accepted schema permits
a list only for two or more jointly referenced observations. Several responses
also copied the relation label `limits_scope` into the free-text scope field.
These were systematic model-output errors, not transport or parser failures.
The prompt and run were not retried or tuned.
Human semantic review of the preserved raw responses:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override
strict structural failures; it separates useful semantic signal from schema
compliance.
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving
closer to evidence reduced some direct promotion behavior: the geometry mention
did not become work, the washing disadvantage did not become an unresolved
issue, both alternatives in B remained uncommitted, and the publication query
did not become an assignment. However, the important promotion errors did not
disappear. C acquired unsupported acceptance, and G still promoted a personal
preference into rejection. Positive cases were only partly preserved: explicit
rejection was recognized but incorrectly linked; trial-only language was kept
but responsibility was invented; Nina's request and commitment were recognized
but responsibility states were wrong; and the publication issue was recognized
but its uncertainty was contradicted by rejection.
Result: **B — evidence-near extraction is promising, but specific observation
dimensions remain unreliable.** Target/relation selection, scope attachment,
responsibility state/person attribution, and agreement versus uncertainty are
not reliable enough to justify designing the later interpretation stage yet.
No production integration or later interpretation stage was implemented.
## EXP-0028 — Evidence-Near Observation Extraction V2
Date: 2026-08-19
V2 tested whether `qwen3.5:9B` preserves the evidence needed by a later
controlled interpretation stage when direct responsibility, agreement and
semantic graph relations are removed. Responsibility was replaced by explicit
participant/discourse facts (`speaker`, `named_person`, `addressee`, singular
self-reference, collective `we`, and impersonal person reference). Agreement
was replaced by explicit affirmation, explicit negation and determination
statement signals. Graph relations were reduced to nullable scalar
`refers_to`; scope became free-text `qualifier` plus nullable scalar
`limits_target`. No later derivation stage was implemented.
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
human interpretations. It contains no responsibility, agreement, action,
decision, open-question, accepted-trial, rejected-alternative or protocol
eligibility fields. Validation enforces known evidence IDs, sequential unique
observation IDs, backward-only scalar references, closed vocabularies, boolean
participant flags, JSON-nullable participant/qualifier/reference fields and no
string `"null"`.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries or voting. A launch-path defect was corrected
before the live run; the failed launch made zero model calls. A sandbox-blocked
localhost attempt also made zero model calls. The completed run called the
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
Persistent prompts, Gold input and expectations, raw and parsed model output,
validation, automatic comparison, Ollama metadata and human evaluation are in
`artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/`.
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
were schema-invalid. The dominant serialization pattern was use of `present`
instead of the specified `explicit` for affirmation/negation; E additionally
used `none` instead of `absent` for a determination signal, while D emitted the
separate uncertainty concept as an invalid modality. These errors are
contract violations, although most `present`/`explicit` differences are
deterministically normalizable without changing meaning. A and C were valid
JSON/schema outputs but had critical semantic mismatches.
Human semantic review:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; `present` caused schema failure. |
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced
V1 promotion errors: collective speech and speaker identity no longer became
individual responsibility; personal preference no longer became a group-level
rejection field; an information question did not become work; and explicit
negation/affirmation survived as separate evidence. Useful participant evidence
also survived strongly in H and collective-speech evidence in F.
Simplification did not make all evidence-near dimensions reliable. Scalar
references and `limits_target` were almost entirely omitted, qualifiers were
usually omitted, committed modality was missed in E, F and H, and C/G repeated
important modality, uncertainty and participant-feature errors. Some positive
semantic information therefore survived only in free-text `content`, not in
the structural signals a controlled derivation stage would need.
Result: **B — V2 is materially better, but specific evidence-near dimensions
still require refinement.** Direct responsibility, agreement and graph-relation
classification should remain excluded. Before designing the derivation stage,
the next work should examine the minimal reliable representation of explicit
reference/scope limitation, commitment modality and participant deixis. No
production integration, Progeo run or derivation implementation was performed.
## EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
Date: 2026-08-19
Hypothesis: `qwen3.5:9B` is substantially more reliable when the first semantic
stage preserves meeting meaning as atomic natural-language observations with
provenance and only simple participant information, without classifying or
deriving higher-level meeting semantics.
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
observation contains exactly `observation_id`, `evidence_id`, `content`,
`speaker`, nullable `named_person`, and nullable `addressee`. It contains no
modality, temporality, evaluation, affirmation, negation, determination,
uncertainty, clarification, responsibility, agreement, relation, reference,
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
Instead, the prompt asks for conservative atomic content that retains hedges,
conditions, personal/collective/impersonal language, requests, acceptances,
rejections, quantities, deadlines and boundaries in natural language.
Structural validation is intentionally small: exact schema keys, non-empty
observations/content, unique `obs_N` identifiers, known evidence IDs, speaker
matching its evidence, explicit named people/addressees, and no string
`"null"`. Human semantic preservation against per-case requirements is the
primary evaluation; wording differences do not fail a case.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries, voting or per-case tuning. One sandbox-blocked
localhost launch made zero model calls. The completed run made exactly nine
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
structural validation. Persistent source evidence, semantic requirements,
exact prompts, raw and parsed responses, validation, Ollama metadata and human
evaluation are stored under
`artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/`.
Human semantic preservation results:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PASS | Preserved `kann`, `vielleicht`, tentative follow-up, and explicit `Dann` sequence without commitment. |
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, `wenn überhaupt`, impersonal checking need and no decision, but dropped collective `wir` from who would receive contaminated material. |
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
“that we receive contaminated material back” into an impersonal passive phrase;
the risk itself remained hypothetical. H translated `Freitag` to `Friday`, a
harmless wording difference. I retained two compound observations rather than
splitting every proposition, but all required semantic boundaries and
dependencies remained explicit.
Compared with V2, categorical-field removal improved content preservation in
A, G and I: A retained `Dann`; G retained `wenn überhaupt`, personal `Ich` and
impersonal `Man`; I retained publication purpose and all boundaries. It also
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
strengthening into responsibility, group rejection, established work or
assigned clarification. F and H remain sufficiently complete in natural
language for a later interpretation experiment. No useful meaning was shown to
depend on the removed fields; the V3 content retained the useful signals that
V2's fields had attempted to encode.
Result: **A — MINIMAL FIRST STAGE ACCEPTED.** On A-I, minimal atomic content
with evidence provenance and simple participants is sufficiently reliable to
be the candidate first semantic stage. A later bounded experiment may examine
controlled semantic interpretation, but no derivation stage, production
integration or Progeo run was implemented here.
## EXP-0030 — V3 model comparison: qwen3.5:9B vs qwen3.6:35B-A3B
Date: 2026-08-19
This controlled comparison reran the accepted V3 minimal semantic-preservation
experiment unchanged with the locally installed `qwen3.6:35B-A3B`. It used the
same implementation, A-I fixture, evidence, prompt, minimal schema, temperature
0, `think=false`, `num_ctx=16384`, `num_predict=4096`, no retries, no voting and
one call per case. The completed run made exactly nine calls in 96.392 seconds.
Artifacts are preserved under
`artifacts/experiments/evidence_observations_v3/20260819_v3_qwen36_35b_a3b_single_run/`.
Structural validation passed 8/9 cases. D was semantically faithful but invalid
because the model copied transcript speakers Antonius and Martin into
`named_person`, although those names were not explicitly named within their
utterances. Human semantic preservation produced 7 PASS, 1 PARTIAL and 1 FAIL:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PASS | Preserved `can`, `perhaps`, tentative `would`, explicit `then` and no commitment, but changed the content language to English. |
| B | PARTIAL | Preserved both approaches overall, but removed `Oder` from the 20-20 observation and locally strengthened it into collective planned conduct. |
| C | PASS | Preserved Tim's tentative personal Textor contact and possible follow-up without established work. |
| D | PASS | Preserved possibility, uncertainty, process and energy consequence; structural failure was confined to invalid speaker-as-named-person values. |
| E | PASS | Preserved cost, collective rejection/non-pursuit and later determination, though the confirmation dropped explicit `Ja`. |
| F | FAIL | Preserved quantity, timing and trial/series boundaries, but changed collective `wir` into “Martin suggests” and “Tim agrees,” inventing individual proposal/agreement meaning. |
| G | PASS | Preserved hypothetical risk, collective recipient `wir`, Martin's personal stance, `wenn überhaupt`, conditional Technikum, impersonal `man müsste` and no decision/owner. |
| H | PASS | Preserved request, addressee, Friday, explicit acceptance and future personal commitment without a responsibility field; content was English. |
| I | PASS | Preserved local/pure-production and comparison boundaries, five-degree difference, upstream effort, publication purpose, unresolved permission and clarification without assignment. |
Direct comparison:
| Measure | `qwen3.5:9B` | `qwen3.6:35B-A3B` |
| --- | ---: | ---: |
| Structurally valid | 9/9 | 8/9 |
| Human PASS | 8 | 7 |
| Human PARTIAL | 1 | 1 |
| Human FAIL | 0 | 1 |
| Observations | 28 | 30 |
| Runtime | 40.074 s | 96.392 s |
| LLM calls | 9 | 9 |
| Prompt-evaluation tokens | 6,268 | 6,268 |
| Evaluation tokens | 2,590 | 2,718 |
The larger model fixed the 9B weakness in G by preserving collective `wir`, and
it split I's compound production/effort observations more cleanly. Those gains
did not offset regressions: B was locally strengthened, F materially converted
collective conduct into individual agreement, D violated the participant
schema, observation count increased, and runtime was 2.4 times higher. Both
models preserved German consistently in six of nine cases, but in different
cases; the 35B-A3B model changed A, F and H to English, while 9B changed C, F
and H wholly or partly to English.
Result: **D — REGRESSION.** `qwen3.6:35B-A3B` does not materially improve the
accepted minimal V3 first-stage preservation over `qwen3.5:9B`; it is worse on
the A-I comparison because of the F ownership-adjacent strengthening and lower
structural validity. This conclusion applies only to the minimal V3 first
stage and does not determine model choice for any later semantic derivation.
No production integration, derivation implementation or Progeo run occurred.
## EXP-0031 — Controlled Semantic Derivation H V0
Date: 2026-08-19
This isolated experiment tested the first controlled second-stage derivation
using only the accepted `qwen3.5:9B` V3 observations for case H. The derivation
LLM received the two V3 observations, not the transcript or Gold expectation.
Its deliberately narrow task was limited to recognizing whether `obs_1` is a
concrete request and whether `obs_2` explicitly commits its speaker to
substantially the same work. Its strict output schema forbids responsibility,
requested actor, establishment/status, Action Item, protocol, confidence and
generic relation/graph fields.
Deterministic code validates observation/evidence provenance, obtains the
requested actor only from the request observation's addressee, requires the
acceptance to follow the request, requires the accepting speaker to equal that
addressee, and establishes responsibility only after all semantic and
structural gates pass. A bounded weekday normalizer reconciles `Friday` and
`Freitag`, rejects conflicting weekdays, and separates the supported due date
from the normalized action text. No general temporal or action ontology was
introduced.
Twenty focused deterministic tests cover the positive H path and the required
negative invariants: request alone, acknowledgement/non-commitment, tentative
acceptance, different response speaker, different work, reversed order,
speaker/name/addressee alone, conflicting deadlines, unknown observation IDs,
inconsistent evidence provenance, forbidden semantic fields, malformed JSON
and persistent artifacts. The complete non-LLM suite passed 192/192.
Configuration: one `qwen3.5:9B` call, temperature 0, `think=false`,
`num_ctx=16384`, `num_predict=1024`, no retries or voting. The call took 11.765
seconds, with 469 prompt-evaluation and 124 evaluation tokens. The model
returned a valid recognition object: `obs_1` is a concrete request, `obs_2` is
an explicit commitment, and both concern substantially the same work. It
returned no responsibility or establishment judgment.
All deterministic gates passed. The final derived result is an established
action `Prüfung der Messdaten`, requested from and assigned to Nina, due
`Freitag`, supported by request `obs_1/e1` and acceptance `obs_2/e2`. The model
included `bis Friday` in its normalized request text; after the single call, a
deterministic-only bounded correction separated that already-recognized due
phrase from action content without changing the prompt, recognition schema,
semantic result or call count. Focused and complete non-LLM suites still
passed after this correction.
Artifacts are preserved under
`artifacts/experiments/controlled_semantic_derivation_h/20260819_h_qwen35_9b_single_run/`.
Result: the H mechanism succeeded. This establishes only that the narrow
request-plus-explicit-acceptance pattern can be recognized and gated for H; it
does not generalize the derivation architecture to other cases or semantic
categories. No production integration, other case run, semantic graph,
protocol derivation or Progeo run occurred.
## EXP-0032 — Request / Acceptance Gold V0
Status: Experimental; promising with semantic precision gaps
Date: 2026-08-20
This isolated regression experiment tested whether the EXP-0031 mechanism
generalizes beyond H. It used ten short synthetic cases containing only
V3-style observations. Evidence Observation V3 was neither called nor changed,
and the model received no raw transcript or expected result. The fixed
recognition schema permits only a nullable concrete request and nullable later
explicit personal commitment, plus the same-requested-work judgment and
normalized action text. Responsibility, requested actor, established status,
Action Item, protocol, confidence and generic graph fields remain forbidden.
Cases:
- RA-01 explicit positive acceptance: PASS.
- RA-02 paraphrased positive acceptance: PASS.
- RA-03 acknowledgement only: PASS.
- RA-04 tentative response: PASS.
- RA-05 different responder without personal acceptance: PASS.
- RA-06 explicit commitment to different work: PARTIAL. The model returned no
acceptance instead of recognizing a commitment with `same_requested_work`
false. The requested action correctly remained unestablished.
- RA-07 request without response: PASS.
- RA-08 collective commitment: PARTIAL. The model over-recognized the
collective `wir` statement as an explicit commitment, but no request existed
and deterministic gates prevented individual responsibility.
- RA-09 impersonal necessity: PARTIAL. The model over-recognized the impersonal
necessity as a concrete request, but the observation had no addressee and
deterministic gates prevented establishment.
- RA-10 tentative personal suggestion: PASS.
Configuration: exactly ten sequential `qwen3.5:9B` calls, one per case,
temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no retries,
no voting and no prompt change between cases. Summed call time was 23.754
seconds, with 4,826 prompt-evaluation tokens and 793 evaluation tokens. The
strict schema validated every response and no responsibility or establishment
field leaked into model output.
Both positive cases recognized the request, explicit commitment and same-work
relationship, including the paraphrased acceptance, and deterministically
established Clara as responsible with due date `Dienstag`. The model rendered
the normalized action in semantically equivalent English; evaluation therefore
checks the structural deterministic result exactly while treating normalized
action wording as evidence-near semantic text rather than requiring lexical
identity. Acknowledgement and tentative response were not promoted. Every
negative case remained unestablished, and no individual responsibility was
invented.
Recognition-level errors were two false positives (RA-08 commitment and RA-09
request) and one false negative (RA-06 different-work commitment). Final
established-action false positives and false negatives were both zero. The
overall result was seven PASS, three PARTIAL and zero FAIL.
Conclusion: the narrow request-plus-acceptance architecture remains promising
for established individual actions because deterministic addressee, ordering,
speaker, same-work, provenance and deadline gates contained all recognition
errors. The recognition layer is not yet precise enough to generalize: its
handling of collective commitment, impersonal necessity and commitments to
different work needs further isolated study. No production integration or
additional semantic category is justified by this result.
Artifacts are preserved under
`artifacts/experiments/request_acceptance_gold_v0/20260820_qwen35_9b_single_run/`.
## EXP-0033 — Collective Commitment Gold V0
Status: Experimental; architecturally successful with one contained
recognition false positive
Date: 2026-08-20
This isolated second-stage experiment tested whether an explicit collective
first-person commitment can establish an action without inventing an individual
owner. It used ten synthetic cases containing one minimal V3-style observation
each. Evidence Observation V3 was neither called nor changed, and the accepted
Request/Acceptance mechanism remained unchanged and independent.
The strict semantic schema contains exactly `observation_id`,
`commitment_form` and `normalized_action_text`. `commitment_form` is closed to
`individual_first_person`, `collective_first_person` and `none`. The model
cannot output responsibility, ownership, requested actor, establishment,
Action Item, protocol, confidence, relations, graphs, decisions or unresolved
issues. Deterministic code validates schema and provenance, requires collective
commitment plus non-empty action text, applies bounded deadline consistency and
explicit-negation gates, and only then sets `status: established`,
`commitment_scope: collective` and `responsible_person: null`.
Gold results:
- CC-01 explicit collective commitment: PASS; established, due `nächste
Woche`, no person.
- CC-02 individual commitment: PASS; correctly routed out of the collective
path.
- CC-03 tentative collective possibility: PASS; unestablished.
- CC-04 collective suggestion: PASS; unestablished.
- CC-05 impersonal necessity: PASS; unestablished.
- CC-06 passive future statement: PASS; unestablished.
- CC-07 collective rejection: PARTIAL. The model incorrectly returned
`collective_first_person`, but the deterministic negation gate detected
`nicht` and prevented establishment.
- CC-08 qualified collective commitment: PASS; established with `nur im
Technikum` preserved, null due and no person.
- CC-09 collective commitment without deadline: PASS; established with null
due and no person.
- CC-10 speaker ownership trap: PASS; established collectively while Martin
remained only the speaker and was not assigned ownership.
Configuration: exactly ten successful sequential `qwen3.5:9B` calls, one per
case, temperature 0, `think=false`, `num_ctx=16384`, `num_predict=1024`, no
retries, no voting and no prompt change. There were zero technical failed
calls. Aggregate runner time was 10.504 seconds; summed per-call time was 10.500
seconds, with 4,267 prompt-evaluation tokens and 415 evaluation tokens.
The outcome was nine PASS, one PARTIAL and zero FAIL. There was one recognition
false positive and no recognition false negatives. No qualifier was lost, no
individual owner was invented, and no responsibility or status field leaked
into recognition. Bounded due handling preserved `nächste Woche` verbatim and
returned null when no deadline was present.
Conclusion: the collective-commitment path is architecturally successful for
this narrow Gold set. The deterministic negation gate contained the only model
error, and every successful collective result necessarily retained
`responsible_person: null`. This does not justify a generic commitment system,
production integration, group identity inference or another semantic category.
Artifacts are preserved under
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
## EXP-0034 — Explicit Rejection Gold V0
Status: Failed architecturally
Date: 2026-08-20
This isolated Stage-2 experiment tested the narrow evidence fact that a
concrete action, option, proposal or future course was explicitly rejected,
abandoned, discontinued or ruled out. It used twelve synthetic cases containing
one self-contained observation or one local target/rejection pair. Evidence
Observation V3 was not called or changed. The accepted Request/Acceptance and
Collective Commitment paths remained unchanged and were not invoked.
The strict semantic schema contains exactly `rejection_observation_id`,
`target_observation_id`, `rejection_form` and
`normalized_rejected_action_text`. `rejection_form` is closed to
`explicit_action_rejection` and `none`. A positive recognition requires a
known local target and non-empty normalized target; `none` requires both target
and normalized text to be null. Decision, outcome, topic-closure,
responsibility, ownership, protocol, confidence and graph fields are forbidden.
Target resolution is limited to the same observation or one earlier supplied
observation. Deterministic code validates schema, IDs, ordering and complete
provenance before emitting the narrow status `explicitly_rejected`.
`explicitly_rejected` means rejected by the cited evidence only. It is not yet
a final meeting decision or final topic outcome, does not close a topic, and
does not supersede an earlier commitment.
Gold results:
- RJ-01 explicit collective rejection with local target: PASS.
- RJ-02 explicit non-pursuit with paired target: PASS.
- RJ-03 self-contained collaboration rejection: FAIL. The model returned
`none`, producing one recognition false negative.
- RJ-04 personal preference: FAIL. The model promoted the preference to an
explicit rejection and derived an unsupported rejection.
- RJ-05 concern: PASS; remained a non-rejection.
- RJ-06 uncertainty: PASS; remained a non-rejection.
- RJ-07 negative recommendation: FAIL. The model promoted advice to an
explicit rejection and derived an unsupported rejection.
- RJ-08 deferral: PASS; remained a non-rejection.
- RJ-09 factual negation: PASS; remained a non-rejection.
- RJ-10 temporary non-action: FAIL. The model treated `erstmal noch nicht` as
abandonment and derived an unsupported rejection.
- RJ-11 explicit rejection with material scope: PASS. Real-plant and
Druckversuch scope were preserved.
- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option
was rejected; the Technikum alternative was not absorbed.
Configuration: exactly twelve successful sequential `qwen3.5:9B` calls, one
per case, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
technical failed calls. Aggregate runner time was 15.518 seconds; summed
per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681
evaluation tokens.
The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced
three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03).
There were four strict target-field expectation mismatches: three were
consequences of false-positive rejection objects populating otherwise locally
correct antecedents, and one was the missing self-contained RJ-03 target. No
derived positive selected the wrong concrete antecedent. Qualifier-loss count
was zero, positive-alternative absorption count was zero, and no responsibility,
decision, outcome or topic-closure field leaked into model output.
Conclusion: the experiment is not architecturally successful. Deterministic
structural gates cannot contain a semantically well-formed false-positive
rejection with valid local target and provenance. The model did distinguish
concern, uncertainty, deferral and factual negation, and it handled scoped and
alternative-bearing positives correctly, but it did not reliably separate
explicit rejection from personal preference, advice or temporary non-action.
The current binary recognition `explicit_action_rejection | none` is
insufficient for reliable generalization.
No production integration, generic rejection system, prompt tuning or
cross-pattern reconciliation is justified.
Artifacts are preserved under
`artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/`.
## EXP-0035 — Negative Act Form V0
Status: Experimental; successful for form classification with normalization
limitations
Date: 2026-08-20
EXP-0034 failed because the binary `explicit_action_rejection | none` question
collapsed materially different negative acts. It missed self-contained
non-pursuit and promoted personal preference, recommendation and temporary
non-action to rejection. This isolated follow-up tested only whether those
evidence-near forms can be distinguished before any normative derivation. It
does not derive rejection, decision, outcome, topic closure, responsibility or
protocol status, and EXP-0034 remained unchanged.
The strict output schema contains exactly `observation_id`,
`negative_act_form` and `normalized_action_text`. The closed form vocabulary is
`explicit_non_pursuit`, `personal_preference`, `recommendation`,
`temporary_non_action` and `none`. Non-`none` forms require non-empty normalized
action text; `none` requires null. Rejection, status, decision, outcome,
responsibility and other normative fields are forbidden recursively. Local
context may resolve a candidate observation's pronoun, but the schema contains
no target relation and the experiment exposes no derivation function.
Gold results:
- NA-01 explicit non-pursuit: PARTIAL. The form was correct; `working with Dr.
Schlummer` omitted the continuation aspect from normalization.
- NA-02 paraphrased explicit non-pursuit: PASS.
- NA-03 personal preference: PARTIAL. The form was correct, but normalization
repeated `Ich würde das nicht machen` instead of resolving the real-plant
trial target.
- NA-04 negative recommendation: PARTIAL. The form was correct; the normalized
English action used the loose rendering `real asset` for `reale Anlage`.
- NA-05 temporary non-action: PASS.
- NA-06 concern only: PASS with `none` and null action text.
- NA-07 uncertainty: PASS with `none` and null action text.
- NA-08 factual negation: PASS with `none` and null action text.
Expected-versus-actual form confusion was entirely diagonal:
| Expected form | Actual form | Count |
| --- | --- | ---: |
| `explicit_non_pursuit` | `explicit_non_pursuit` | 2 |
| `personal_preference` | `personal_preference` | 1 |
| `recommendation` | `recommendation` | 1 |
| `temporary_non_action` | `temporary_non_action` | 1 |
| `none` | `none` | 3 |
Configuration: exactly eight successful sequential `qwen3.5:9B` calls, one
per case, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
technical failures. Aggregate runner time was 8.688 seconds; summed per-call
time was 8.686 seconds, with 4,183 prompt-evaluation tokens and 310 evaluation
tokens.
The result was five PASS, three PARTIAL and zero FAIL. All eight
`negative_act_form` classifications matched Gold. There was no unsupported
semantic strengthening and no rejection, status, decision, outcome,
responsibility or topic-closure leakage. Normalized action meaning was fully
acceptable in five cases and imperfect in three.
Conclusion: the finer evidence-near form vocabulary successfully distinguished
the four semantic boundaries that defeated the binary rejection experiment in
this small Gold set. The result supports separating negative-act-form
recognition from later normative derivation, but local target normalization is
not yet uniformly reliable. It does not justify modifying EXP-0034, deriving
rejection, production integration or beginning cross-pattern reconciliation.
Artifacts are preserved under
`artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run/`.
## EXP-0036 — Controlled Rejection Derivation V1
Status: Experimental; architecturally unsuccessful
Date: 2026-08-20
This isolated experiment followed the failed binary rejection baseline
(EXP-0034) and successful Negative Act Form classification (EXP-0035). Its V1
hypothesis was to classify the negative act first, resolve its local target in
a separate semantic call, and only then derive `explicitly_rejected`
deterministically. It did not modify either predecessor or any accepted Stage-2
pattern, and it has no production integration.
The target recognizer emitted exactly `candidate_observation_id`,
`target_observation_id`, and `normalized_target_text`. Only
`explicit_non_pursuit` was deterministically eligible. Personal preference,
recommendation, temporary non-action, and `none` could never derive rejection,
even with a valid target. Provenance, local membership, ordering, non-empty
target text, and strict non-normative output were additional gates.
`explicitly_rejected` means rejected by the cited evidence only, not a final
decision, topic outcome, permanent state, or closure.
The run reused five exact accepted Negative Act Form outputs and made three new
Negative Act calls plus eight target-resolution calls. All calls used
`qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=1024`, no retries, voting, or prompt changes.
| Case | Negative Act expected / actual | Target result | Verdict |
| --- | --- | --- | --- |
| CR-01 | `explicit_non_pursuit` / same | Model returned the string `"null"` as an unknown ID; self-contained target was not linked | FAIL |
| CR-02 | `explicit_non_pursuit` / same | `obs_1`, external solution and continuation preserved | PASS |
| CR-03 | `personal_preference` / same | `obs_1`; eligibility gate prevented rejection | PASS |
| CR-04 | `recommendation` / same | `obs_1`; eligibility gate prevented rejection | PASS |
| CR-05 | `temporary_non_action` / same | `obs_1`; eligibility gate prevented rejection | PASS |
| CR-06 | `none` / same | Null target; final non-rejection was correct, but expected local target was unresolved | FAIL |
| CR-07 | `explicit_non_pursuit` / same | `obs_1`; real-plant and pressure-test scope survived, but normalization remained proposition-like | PARTIAL |
| CR-08 | `explicit_non_pursuit` / same | `obs_1`; real-plant scope preserved and Technikum alternative excluded | PASS |
Result: five PASS, one PARTIAL, two FAIL. All eight Negative Act forms were
correct. There were no false-positive rejections: the valid targets in CR-03,
CR-04, and CR-05 could not override their ineligible forms. There was one
false-negative rejection, CR-01, caused by invalid target output. Target
resolution missed two expected links (invalid CR-01 and null CR-06), so the
wrong/unresolved-target count was two. CR-06 exposed a strategy flaw: a
non-eligible Negative Act form should not be required to pass target resolution
when it cannot derive rejection. Qualifier-loss count was zero. CR-08 isolated
the positive alternative successfully. No individual owner, responsibility,
decision, outcome, topic-closure, or LLM-emitted rejection status appeared.
There were 3 new Negative Act calls, 8 target calls, 5 accepted classification
reuses, zero technical call failures, and one structural target-validation
failure. Aggregate runner time was 11.951 seconds.
Conclusion: negative-act-form gating is promising and successfully contains
the semantic false positives that defeated EXP-0034, but the experiment is not
architecturally successful. The current target-resolution strategy failed the
required self-contained positive CR-01 and unnecessarily evaluated the
ineligible CR-06 path; it is not reliable enough for rejection derivation.
Artifacts are preserved under
`artifacts/experiments/controlled_rejection_v1/20260820_qwen35_9b_single_run/`.
## EXP-0037 — Target Resolution V0
Status: Experimental; FAILED for target resolution
Date: 2026-08-20
Controlled Rejection V1 showed that fine-grained Negative Act Form eligibility
contained false-positive rejection, but its target strategy failed a
self-contained positive and unnecessarily resolved a target for an ineligible
`none` form. This isolated experiment tested target resolution only. It
contains no rejection derivation, status, decision, outcome, responsibility,
topic closure, or production integration.
Eligibility was deterministic: only `explicit_non_pursuit` could reach the
resolver. TR-05 personal preference, TR-06 recommendation, TR-07 temporary
non-action, and TR-08 `none` stopped before prompt construction and recorded an
explicit skipped-call artifact. This hard gate worked in all four cases.
The target schema contained exactly `candidate_observation_id`,
`target_observation_id`, and `normalized_target_text`, with local IDs,
same-or-earlier ordering, unique evidence provenance, null consistency, and
recursive normative-field exclusion. TR-01 used the self-contained strategy:
the prompt stated that linkage was deterministically fixed to the candidate and
requested semantic normalization only. TR-02 through TR-04 used paired local
resolution. No original transcript or new Negative Act classification call was
used.
| Case | Form / eligible | Call | Target result | Verdict |
| --- | --- | --- | --- | --- |
| TR-01 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; required same-observation target unresolved | FAIL |
| TR-02 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; `obs_1` unresolved | FAIL |
| TR-03 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; scoped `obs_1` unresolved | FAIL |
| TR-04 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; real-plant target unresolved | FAIL |
| TR-05 | `personal_preference` / no | no | Deterministically skipped | PASS |
| TR-06 | `recommendation` / no | no | Deterministically skipped | PASS |
| TR-07 | `temporary_non_action` / no | no | Deterministically skipped | PASS |
| TR-08 | `none` / no | no | Deterministically skipped | PASS |
Result: four PASS, zero PARTIAL, four FAIL. Exactly four successful Ollama
calls were made, all for eligible cases; there were zero technical call
failures and four structural validation failures. All four raw responses used
the JSON string `"null"` as target ID rather than a supplied observation ID or
JSON null. Wrong-target count and unresolved-target count were therefore four.
No qualifier-preservation claim can be made because no eligible positive target
passed validation. TR-04 alternative isolation likewise could not be
established. No rejection, status, decision, outcome, responsibility, or other
normative leakage occurred, and no rejection derivation was performed.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`,
`num_ctx=16384`, `num_predict=1024`, no retries, voting, or prompt changes.
Aggregate runner time was 4.034 seconds.
Conclusion: eligibility gating is successful and should be retained; it fully
prevents unnecessary target calls for ineligible Negative Act forms. Target
Resolution V0 itself failed structurally across all eligible cases. Neither the
self-contained nor paired strategy produced a valid target, and merely
instructing deterministic self-linkage in the semantic prompt did not make the
linkage structurally deterministic. The repeated `"null"` string pattern
requires diagnosis before changing the architecture or prompt. No rejection
derivation is justified by this result.
Artifacts are preserved under
`artifacts/experiments/target_resolution_v0/20260820_qwen35_9b_single_run/`.
## EXP-0038 — Target Resolution V1 Diagnostic
Status: Experimental; linkage boundary successful, normalization incomplete
Date: 2026-08-20
Forensics on failed Target Resolution V0 found a definite prompt defect: its
illustrative value `"observation ID or null"` placed both alternatives inside
a JSON string. V0 also sent only `format: "json"`, which enforced JSON syntax
but not field types. This isolated diagnostic changed only the linkage/output
boundary. It contains no rejection derivation or normative semantics.
Ollama 0.32.6 accepted a true JSON Schema object in `format`. TR1-V1 removed
target selection from the model output entirely and deterministically linked
the self-contained candidate to itself. TR2-V1 through TR4-V1 used a closed
allowed-ID list, an enum of those IDs plus JSON null, typed positive and null
examples, recursive strict validation, and one fixed paired prompt. Linkage and
normalization were persisted separately.
| Case | Strategy / ID source | Target | Normalized target | Verdict |
| --- | --- | --- | --- | --- |
| TR1-V1 | self-contained / deterministic | `obs_1` | `Mit Dr. Schlummer arbeiten wir nicht weiter.` retained negation instead of a positive action meaning | FAIL |
| TR2-V1 | paired / LLM | `obs_1` | `externe Lösung weiterverfolgen` | PASS |
| TR3-V1 | paired / LLM | `obs_1` | `reale Anlage zur Diskussion` lost `Druckversuch` purpose and the `nutzen` action | FAIL |
| TR4-V1 | paired / LLM | `obs_1` | `Versuch in der realen Anlage durchführen`; Technikum excluded | PASS |
Result: two PASS, zero PARTIAL, two FAIL. All four responses passed their true
JSON Schemas. Every resulting target was `obs_1`; wrong-target and
unresolved-target counts were zero. The string `"null"` recurrence count was
zero, and there were zero structural validation failures. TR1 preserved the
collaboration, person, and continuation wording but failed positive-action
normalization by retaining negation. TR3 had one material scope loss. TR4
preserved real-plant scope and isolated the Technikum alternative. No
normative leakage occurred.
Configuration: exactly four `qwen3.5:9B` calls, temperature 0,
`think=false`, `num_ctx=16384`, `num_predict=1024`, no retries, voting, or
prompt tuning. Aggregate runner time was 4.557 seconds.
Conclusion: the V0 string-null failure was primarily a linkage/output-boundary
failure rather than evidence that observation-ID linkage is semantically
impossible. True typed schemas, closed ID lists, and deterministic self-linkage
eliminated every structural and target-ID failure. The experiment still fails
its complete acceptance criterion because target normalization is not reliably
positive or scope-preserving. These results justify separating linkage from
normalization, but not deriving rejection or integrating a new pipeline.
Artifacts are preserved under
`artifacts/experiments/target_resolution_v1_diagnostic/20260820_qwen35_9b_single_run/`.
## EXP-0039 — Target Normalization V0
Status: Experimental; normalization improved but incomplete
Date: 2026-08-20
Target Resolution V1 established correct linkage for all four narrow cases and
eliminated structural ID failures with deterministic self-linkage, closed ID
lists, and true JSON Schemas. Its remaining failures were normalization-only.
This isolated follow-up therefore accepted candidate and target IDs as fixed
input and tested only reconstruction of the positive German action meaning. It
contains no target selection, Negative Act classification, eligibility logic,
rejection derivation, or production integration.
The strict output schema contained exactly `candidate_observation_id`,
`target_observation_id`, and `normalized_target_text`. Both IDs were constrained
to their supplied values with JSON Schema `const`; normalized text was a
non-empty string and null was disallowed. The one fixed prompt required removal
of negative polarity, preservation of action, continuation, material scope and
source language, and exclusion of separate alternatives.
| Case | Actual normalized target | Verdict |
| --- | --- | --- |
| TN-01 | `Mit Dr. Schlummer zusammenarbeiten` | FAIL: positive polarity and collaboration survived, but continuation was lost |
| TN-02 | `externe Lösung weiterverfolgen` | PASS |
| TN-03 | `reale Anlage für den Druckversuch nutzen` | PASS |
| TN-04 | `Versuch in der realen Anlage durchführen` | PASS; Technikum alternative excluded |
Result: three PASS, zero PARTIAL, one FAIL. All four outputs passed strict
schema validation and copied both fixed IDs exactly, so changed-ID count was
zero. Polarity-error count was zero: even TN-01 removed rejection and negation.
Action/continuation-loss count was one (TN-01); material purpose/location
scope-loss count was zero; alternative-absorption count was zero. There was no
unsupported strengthening or normative leakage.
Configuration: exactly four `qwen3.5:9B` calls, temperature 0,
`think=false`, true JSON Schema, `num_ctx=16384`, `num_predict=1024`, no
retries, voting, or prompt tuning. Aggregate runner time was 5.066 seconds.
Conclusion: isolating normalization solved the polarity and scoped-action
failures seen in Target Resolution V1 for three of four cases, including exact
pressure-test scope and alternative isolation. Continuation semantics remain
unreliable in the self-contained collaboration case, so Target Normalization
V0 does not meet its full acceptance criterion. The result does not justify
rejection derivation or production integration.
Artifacts are preserved under
`artifacts/experiments/target_normalization_v0/20260820_qwen35_9b_single_run/`.
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
Date: 2026-08-11
Hypothesis: the primary protocol should be a topic-oriented reconstruction of
the meeting rather than a category-oriented list of extracted information.
This first isolated prototype does not replace or connect to the production
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
transcript excerpts to `qwen3.5:9B` and requests Discussion Subjects. Each
subject may contain supported discourse events, an outcome with mandatory
scope, resulting actions and unresolved issues. Optional structures must be
omitted when absent. Every semantic object must reference known evidence IDs.
The strict experimental schema validates:
- non-empty subjects and globally unique semantic identifiers;
- a closed discourse-event vocabulary;
- non-empty, known and non-duplicated evidence references;
- outcome text, scope, certainty and evidence;
- action text, JSON-nullable responsibility/deadline and evidence;
- unresolved-issue text and evidence;
- omission rather than null or empty optional structures.
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only,
multiple options, unaccepted proposal, proposal with objection, rejected
alternative, trial-scoped acceptance, no-decision discussion, resulting Action
Item, and outcome plus unresolved issue. Evaluation targets semantic identity,
development, outcome scope, actions, unresolved issues, traceability and
absence of invented commitments rather than exact wording.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`. Each case received exactly one model call; there were no
model retries or prompt iterations. The nine completed calls took 59.251
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
failure artifacts were preserved under
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/` and
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/`. Two earlier launch
attempts made zero LLM calls: one failed on the script import path and one was
blocked by sandbox networking.
Human-reviewed results after correcting two objectively wrong Gold assumptions
without another model call:
| Case | Verdict | Reason |
| --- | --- | --- |
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled `considered_option` rather than `introduced_idea`. |
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case
F: the model distinguished acceptance for a bounded trial from acceptance as a
final solution. It also handled the explicit action in case H well. However,
the experiment failed systematically on sparse structured output, subject
grouping and restraint around absent outcomes/actions/unresolved issues. The
model frequently mirrored optional schema fields as empty/null values, treated
alternatives as outcomes, split one discussion into multiple subjects, or
invented open issues from mere non-selection.
The focused experiment is not promising enough to justify a real Progeo chunk
sanity check. No such run was performed, and no architecture is accepted on
the basis of this prototype. Further work should first analyze whether the
failure comes from the schema/prompt representation, the model's sparse-output
reliability, or the boundary between subject grouping and semantic synthesis.
It should not proceed through repeated prompt tuning against these nine cases.
## EXP-0027 — GTM-Hub real-world protocol regression case
Status: Accepted as a qualitative regression case
Date: 2026-09-09
The private `samples/real_live/gtm_hub_2026-09-07/` case preserves an existing
78-minute German GTM-Hub discussion, its exact protocol input transcript, a
human colleague's concise reference, and two existing Meeting Assistant
protocols (with anonymous speaker labels and with five confirmed mappings).
It is an evaluation dataset, not a prompt/model experiment; no inference was
run to create it.
The evaluation establishes product and architecture learning without changing
the current architecture:
- the detailed/working protocol remains the primary product artifact; a short
distribution protocol is a later separate renderer, not a replacement;
- diarization has particular value for long, multi-speaker,
disagreement-heavy meetings because it improves discourse attribution;
- diarization does not solve semantic commitment extraction: speaker
attribution and normative interpretation are separate problems;
- responsibility and action extraction require stricter evidence than ordinary
discussion summarization, especially when a named speaker is involved.
Evidence and explicit qualitative regression expectations are recorded in
`samples/real_live/gtm_hub_2026-09-07/evaluation.md`. The case does not define
a synthetic gold protocol or a benchmark metric.
+7 -3
View File
@@ -107,8 +107,10 @@ or aliases are corrected.
`department`: Organizational unit. Optional and nullable.
`attendance_status`: `present` for participants. This distinguishes attendees
from mentioned people.
`attendance_status`: exactly `present` for participants or `mentioned_only` for
people who are relevant but did not attend. For backward compatibility, a
missing status defaults to `present` in `participants` and `mentioned_only` in
`mentioned_people`.
`mentioned_people`: People discussed or referenced but not present. They are
not participants and must not be treated as speakers.
@@ -153,7 +155,7 @@ mentioned_people:
aliases: []
role: null
department: null
attendance_status: "not_present"
attendance_status: "mentioned_only"
notes: "Wurde erwaehnt, war aber nicht anwesend."
organization:
@@ -282,6 +284,8 @@ The current validator checks that:
- participant ids and mentioned-person ids do not collide
- referenced departments exist in `organization.departments`
- `attendance_status` values are valid
- speaker mappings reference present participants only; mentioned-only people
cannot be diarized speakers
- participants are marked `present`
- mentioned people are not marked `present`
+229
View File
@@ -0,0 +1,229 @@
# Protocol Generation Decision
## Executive Summary
Meeting Lab tested ten protocol-generation and runtime variants against the
same 93.5-minute reference meeting, `project_process_meeting`. The current
evidence does not support a fully automatic protocol. The best practical local
baseline remains one direct call to `qwen3.6:35B-A3B`, followed by informed
human review. It is fast and produces readable, broadly useful Markdown, but it
still overstates consensus, compresses unresolved process boundaries, misses
some challenge and follow-up paths, and can infer unsafe ownership. Its current
quality verdict is **C — promising but insufficient**.
Additional prompting, review, Meeting Map, hierarchical, diarized, dense-model
and 70B-scale variants did not produce a reliable step change. Some improved
individual dimensions, but none reached a stable B result or removed the need
for substantive review. This is a current evidence-based product choice, not a
permanent architecture decision.
## Experiments Compared
All protocol rows used the complete cleaned transcript and meeting context
unless stated otherwise. A dash means that the artifact did not record the
number; it is not an estimate. Verdicts marked “assessment” are comparative
assessments of saved outputs because those older artifact directories contain
no formal `quality_review.json`.
| Experiment | Model | Architecture | LLM calls | Prompt tokens | Runtime | Human editing | Verdict | Main strength | Main failure | MVP | Research |
| --- | --- | --- | ---: | ---: | ---: | --- | --- | --- | --- | --- | --- |
| Direct one-shot baseline | Qwen3.6 35B-A3B | Direct full-context protocol | 1 | 18,385 | 72.5 s cold; about 38 s inference | Not recorded | C | Fast, readable, broad topic outline | Consensus and ownership overpromotion; missing boundaries and follow-up | **Yes, with review** | Baseline |
| Conservative one-shot | Qwen3.6 35B-A3B | Direct with stronger safety instructions | 1 | 19,120 | 39.5 s warm | Not recorded | C (assessment) | Better uncertainty and pending-feedback language | Still invents or upgrades named follow-up actions | No | Limited |
| Draft → review | Qwen3.6 35B-A3B | Conservative draft plus review call | 2 | 39,071 total | About 110.4 s summed | Not recorded | C (assessment) | Removes some unsafe named attribution | Does not reliably restore omitted content; empty/weak action sections remain | No | Limited |
| Meeting Map → protocol | Qwen3.6 35B-A3B | Semantic map followed by rendering | 2 | 39,129 total | 134.8 s | Not recorded | C (assessment) | Explicit intermediate structure | Map errors propagate: false consensus and named ownership remain | No | Yes |
| Hierarchical notes → protocol | Qwen3.6 35B-A3B | Five chunk-note calls plus synthesis | 6 | 32,145 total | 251.1 s | Not recorded | C (assessment) | Highest recall in several detailed/open topics | Amplifies unsupported speaker/name interpretations and confirmed actions | No | Yes |
| Segment-level anonymous diarization | Qwen3.6 35B-A3B | One-shot over 1,154 labeled segments | 1 | 39,308 | 125.4 s | 25–35 min | C | Preserves some filtered-idea challenge structure | Token count more than doubled; actions and deadlines became less safe | No | No further protocol tests |
| Turn-merged anonymous diarization | Qwen3.6 35B-A3B | One-shot over 435 merged turns | 1 | 22,490 | 101.2 s | 25–35 min | C | Corrected token inflation; recovered some topic and feedback detail | Still did not beat raw input; unsafe actions/deadlines persisted | No | UI/search only |
| Dense one-shot | Qwen3.5 27B | Direct full-context protocol | 1 | 18,385 | 181.2 s | Not recorded | C (assessment) | Somewhat better recall of process details | Much slower; no material overall quality gain | No | No |
| 70B scale one-shot | Llama 3.3 70B Q3_K_S | Dense, 64% CPU / 36% GPU | 1 | 21,156 | 470.1 s warm | 45–60 min | D | Technically proved a 70B hybrid load can run | Severe coverage loss, invented governance, internal contradiction | No | Negative scale result |
| Ollama vs native llama.cpp | Qwen3.6 35B-A3B | Same Q4_K_M GGUF; ROCm/Vulkan servers | 1 per backend | 18,385 | ROCm 38.4 s; Vulkan 42.3 s | N/A | Runtime only | Native ROCm reached 51.38 generated tok/s | No meaningful end-to-end advantage; more operational complexity | Ollama | Runtime reference |
The draft-review total combines the saved conservative draft call and the
saved review call. Its review metadata itself reports only the one new review
call (19,951 prompt tokens and 70.8 seconds). The direct baseline's 72.5-second
wall time includes a 34.5-second cold load; its measured prompt evaluation plus
generation was 37.8 seconds. These distinctions explain apparent runtime
differences between otherwise similar Qwen3.6 calls.
### Recurring quality patterns
- **Topic coverage and factual accuracy:** Direct Qwen3.6 captures the main
process but misses the second review after enrichment, project reporting and
parts of the filtered-idea challenge path. Hierarchical processing recalls
more detail but introduces too many unsupported interpretations. Llama 3.3
loses most of the meeting and invents a governance role for the
Geschäftsführung.
- **Consensus and unresolved boundaries:** Every broad one-shot family remains
vulnerable to turning discussion or a working direction into agreement. The
unresolved boundary between central coordination and autonomous department
work, and the uncertainty around universal filter criteria, are especially
fragile.
- **Visibility, veto and reconsideration:** No approach consistently preserves
initial filtering, later cross-functional challenge, reconsideration after
enrichment and the return through the project cycle together.
- **Stakeholder feedback:** Pending Jovana and Björn feedback is an important
quality probe. Some variants preserve both; segment-level diarization drops
Björn, while Llama 3.3 drops both.
- **Actions and attribution:** Added structure does not guarantee safety.
Conservative, reviewed, Meeting Map, hierarchical and diarized outputs still
promote proposals or expected work into confirmed actions, infer owners from
roles or conversational context, or invent deadlines. Human review remains
mandatory.
## Model Findings
### Qwen3.6:35B-A3B
Qwen3.6 is the best overall local practical baseline. Its Q4_K_M model is
operationally fast on the RX 9070/CPU hybrid setup, follows the requested
Markdown form and usually provides a useful first draft. It remains verdict C:
larger context and fluent synthesis do not reliably protect evidence strength,
responsibility attribution or unresolved process boundaries.
### Qwen3.5:27b dense
The dense 27B run recalled some process details better than the MoE baseline,
but took 181.2 seconds and generated at 8.45 tokens/s. The gains did not amount
to a material overall quality improvement. This result does not prove that
dense models are generally inferior; it shows that this dense model is not a
better product choice on this hardware and meeting.
### Llama 3.3 70B Q3_K_S
Llama 3.3 70B was technically runnable at 32k context with a 42 GB loaded
footprint and a 64% CPU / 36% GPU split. Its 7m50s warm meeting run produced a
very short, materially worse protocol: one critical invented governance claim,
five new major errors and an estimated 45–60 minutes of editing. Raw parameter
count alone is therefore insufficient. The older model generation and
aggressive Q3 quantization are plausible contributors, but this experiment
does not isolate or prove either cause.
## Runtime Findings
The native comparison reused the exact 23,938,321,664-byte Qwen3.6 Q4_K_M GGUF
that Ollama uses. The tested `llama-server` binary was the llama.cpp runtime
shipped with the installed Ollama distribution, not an independent source
build.
Native ROCm processed the reference request in 38.4 seconds and generated at
51.38 tokens/s. Vulkan took 42.3 seconds and generated at 46.73 tokens/s. The
comparable Ollama baseline generated at 44.68 tokens/s, with about 37.8 seconds
of prompt evaluation plus generation when load time is excluded. Output token
counts differed, so generation throughput alone is not an end-to-end quality or
latency comparison.
Native ROCm gained some generation throughput, but did not provide a meaningful
end-to-end advantage for this workload. Vulkan required more host spill and was
not preferable. Ollama already provides the relevant llama.cpp runtime
components, model lifecycle and API integration; it remains the preferred
routine Meeting Lab runtime.
## Diarization Findings
### Technical feasibility
Pyannote `speaker-diarization-community-1` successfully processed the
93.5-minute meeting on CPU in 1,647 seconds (about 27m27s), an RTF of 0.293. It
detected four anonymous clusters and assigned 1,145 of 1,154 Whisper segments
(99.2%). Peak RSS was about 3.3 GiB. This establishes technical feasibility; it
does not establish speaker identity or diarization accuracy against labeled
ground truth.
### Protocol-quality impact
Annotating every Whisper segment increased the Qwen prompt from 18,385 to
39,308 tokens, confounding speaker structure with fragmentation and token
inflation. Deterministic turn merging reduced 1,154 segments to 435 turns and
the complete prompt to 22,490 tokens. That controlled the main representation
confound, but the resulting protocol still did not materially outperform the
raw transcript and remained verdict C.
Anonymous diarization is therefore not justified as a mandatory MVP
protocol-quality feature. This does **not** mean diarization is generally
useless. It may remain valuable for speaker-aware UI, navigation and search,
participation statistics, traceability, or later carefully validated real-name
mapping.
## Semantic Research Findings
The semantic experiments provide architectural evidence, but should not
dominate the product decision:
- **Evidence Observation V3** is a strong evidence-near candidate stage. With
Qwen3.5 9B it achieved 8 PASS, 1 PARTIAL and 0 FAIL while preserving hedges,
alternatives, requests, commitments and boundaries in natural language.
- **Request/Acceptance** and **Collective Commitment** show that narrow semantic
recognition followed by deterministic provenance, ordering, addressee,
negation and deadline gates can safely derive limited consequences. Model
recognition errors were contained without inventing individual ownership.
- **Explicit Rejection** failed when reduced to a coarse binary recognition
problem: semantically valid false positives passed structural gates.
- **Negative Act Form** worked better by distinguishing non-pursuit, personal
preference, recommendation and temporary non-action before any normative
derivation. All eight form classifications matched Gold, although normalized
action text was imperfect in three cases.
- **Target Resolution V0** failed because prompt examples and a weak JSON
boundary encouraged the string `"null"` instead of typed linkage.
**Target Resolution V1** fixed all linkage/ID failures with deterministic
self-linkage, closed ID lists and true JSON Schema, but normalization remained
incomplete.
- **Target Normalization V0** improved polarity and scope preservation to 3/4
PASS, but still lost continuation meaning in the collaboration case.
These findings support Meeting Lab as a research and validation track. They do
not yet justify placing a multi-stage semantic pipeline on the MVP critical
path.
## Current MVP Decision
The current product path is:
```text
Audio
-> transcription
-> direct qwen3.6:35B-A3B protocol generation through Ollama
-> informed human review
-> final protocol
```
The first MVP should treat the generated protocol as an editable draft, not an
authoritative semantic record. Human review must specifically check consensus,
unresolved boundaries, competing positions, action status, owners, deadlines
and pending stakeholder feedback.
Diarization is optional and deferred. The semantic research pipeline remains
in Meeting Lab, outside the MVP critical path. Ollama remains the default local
runtime.
## Rejected / Deferred Directions
- Do not continue Qwen3.6 prompt variants as the main quality strategy.
- Do not add draft-review, Meeting Map or hierarchical generation to the MVP;
their added calls and complexity did not deliver reliable quality gains.
- Do not continue anonymous-diarization protocol experiments. Revisit
diarization for UI, search, statistics or traceability instead.
- Do not use Qwen3.5 27B or Llama 3.3 70B Q3_K_S as the routine protocol model.
- Do not replace Ollama with a manually managed native llama.cpp service for
this workload.
- Retain semantic experiments, but defer production integration and broad
semantic consolidation.
## Open Questions
1. How does a genuinely newer, materially stronger model perform when a useful
quantization fits the available RAM/VRAM without severe swap?
2. If project policy permits, what quality ceiling does the unchanged reference
prompt achieve with a commercial frontier model?
3. What is the measured reviewer time and correction distribution once the
direct Qwen3.6 draft path is exercised in an end-to-end MVP workflow?
4. Which non-protocol product benefits justify revisiting diarization later?
No further Qwen3.6 prompt variants, anonymous-diarization protocol runs or old
70B Q3 scale tests are recommended.
## Recommended Next Product Step
Build the practical end-to-end MVP around direct Qwen3.6 generation and an
explicit human review handoff. Measure reviewer time and correction categories
in real use. Keep the experiment artifacts and semantic Gold work as validation
evidence, but do not block the first product loop on broader research stages.
@@ -0,0 +1,69 @@
# GTM-Hub meeting — 2026-09-07
Durable real-world regression case for a German GTM-Hub meeting. This is a
private real-meeting sample; do not publish or redistribute it outside the
intended development context.
## Case metadata
- **Case ID:** `gtm_hub_2026-09-07`
- **Meeting:** GTM-Hub Meeting, 2026-09-07
- **Duration:** about 78 min 36 s (4,715.52 s from existing diarization metadata)
- **Context:** a biweekly cross-functional board meeting about coordinating
projects for new markets, products, and applications; this discussion focused
initially on administrative questions and process definition.
- **Participants:** five confirmed participants: Malte Schnau, Henning
Ehrenberg, Björn-Erik Falkenau, Martin Tazl, and Jovana Husemann.
## Contents and provenance
`transcript/transcript.txt` is the exact compact diarized transcript
representation supplied to both comparable Meeting Assistant protocol calls;
it is the evidentiary source for evaluation. It is byte-identical to both
source runs' `protocol/transcript_input.txt` files (SHA-256
`42ac52223cac36eeeb7c01e6d6673b2fb50973c4365b5ddaafdd12927b2f12eb`).
The source plain Whisper transcripts were also identical between the two runs,
but are not copied because they were not the exact representation supplied to
protocol generation.
| Case file | Meeting Assistant source | Generation | Speaker status |
| --- | --- | --- | --- |
| `references/meeting_assistant_anonymous_speakers.md` | `data/meetings/gtm-hub-meeting-2026-09-07/runs/b895cda2_2026-09-07_11-18-54_GTM-HUB_20260909_025239/protocol.md` | 2026-09-09 01:00:46 UTC; `qwen3.8:27b` | Diarized with anonymous `SPEAKER_XX` labels; zero confirmed speaker-to-person mappings. |
| `references/meeting_assistant_speaker.md` | `data/meetings/gtm-hub-meeting-2026-09-07/runs/284a47be_2026-09-07_11-18-54_GTM-HUB_20260909_031106/protocol.md` | 2026-09-09 01:30:44 UTC; `qwen3.8:27b` | Diarized with the same anonymous labels plus five confirmed speaker-to-person mappings in the prompt/context. |
| `references/human_protocol_malte.txt` | copied human colleague reference supplied for this case | not generated | author identified as Malte by filename; preserve unchanged. |
Both runs have meeting ID `gtm-hub-meeting-2026-09-07`, title/date
`GTM-Hub Meeting 2026-09-07`, the same source recording name and size
(154,620,897-byte FLAC), the same 4,715.52-second prepared-audio duration, the
same Whisper model (`ggml-large-v3-turbo.bin`), and the same protocol model,
temperature (0), context window (32,768), and compact diarized transcript
selection. Their Meeting Context V1 files identify the same meeting and four
shared confirmed participants; the speaker-aware context additionally includes
Jovana Husemann and the five confirmed mappings. Exact run IDs, hashes, flags,
and paths are in `manifest.json`.
The source runs were separate full-pipeline executions: their audio,
transcription, diarization, and exact protocol-input artifacts are
byte-identical, but the mapped-speaker run is not operational evidence that it
reused the anonymous-speaker run's Whisper or diarization artifacts. This case
therefore compares equivalent diarized input conditions with different speaker
identity mappings only.
## Evaluation role
The transcript is the evidentiary source. The human colleague protocol is
**not** a gold protocol: it is a human reference showing what an experienced
participant considered worth preserving in a concise working protocol. The two
Meeting Assistant files are comparison variants. This case must not force a
system to imitate the human protocol verbatim.
This is a valuable regression case because a long, five-speaker discussion
contains competing mental models, repeated disagreement and clarification, and
only partial convergence. It distinguishes strategic collection/evaluation and
coordination from QMS/process documentation and operational project work. It
also discusses responsibility repeatedly without always creating a concrete
personal commitment.
See `evaluation.md` for the qualitative comparison and explicit regression
expectations. No audio, raw model response, container log, or other unrelated
run artifacts are intentionally included.
@@ -0,0 +1,113 @@
# Qualitative evaluation — GTM-Hub 2026-09-07
## Method and status
This is a qualitative regression case, not a measured accuracy benchmark.
Assess future outputs against `transcript/transcript.txt`, the evidentiary
source. The human reference is a concise colleague-authored view, not gold
labels. The two Meeting Assistant protocols are comparison variants; their
wording is not itself evidence.
The transcript shows a long attempt to align several interpretations of the
GTM-Hub: a cross-functional place to collect, evaluate, coordinate, and then
assign ideas; separate QMS/process-description work; and later operational
project execution. It records repeated requests for clarification and limited
common understanding, not immediate detailed closure.
## Comparison
| Dimension | Human reference | Assistant with anonymous speaker labels | Assistant with confirmed speaker mappings |
| --- | --- | --- | --- |
| Core meeting intent | Captures the organizational outcome and the three aims: responsibility clarity, information flow, described project flow. | Strong GTM-Hub coverage as strategic collection, evaluation, coordination. | Equally strong, with competing framings retained. |
| Topic coverage | Selective and highly compressed. | Strong: purpose, process/QMS, criteria, roles. | Strong; additionally separates threads and proponents. |
| Discussion reconstruction | Loses much of why discussion evolved. | Reconstructs broad process logic but smooths argumentative turns. | Best reconstruction of structure and repeated clarification. |
| Differing viewpoints | Notes different perspectives, not their content or ownership. | Weak at retaining who held which position. | Best preservation: differing positions remain tied to confirmed speakers. |
| Speaker/person attribution | Initials only; not a complete attribution record. | No confirmed person mapping in input context. | Valuable where mappings are confirmed; names make unsupported interpretation more consequential. |
| Position / proposal / agreement / consensus / decision | Compresses them into a basic shared understanding. | Sometimes promotes discussion to “es wurde vereinbart” or “die Teilnehmer waren sich einig.” | Preserves positions better, but still overstates some proposals/open details as agreement, decision, or obligation. |
| Action-item precision | One concise genuine next step: comment on/give feedback on the BD presentation. | Creates broadly attributed implied actions despite absent explicit individual assignment. | Turns suggestions into named commitments, including delegation by Martin Tazl. |
| Responsibility attribution | Focuses on later project-responsibility clarity rather than naming it prematurely. | Describes future assignment reasonably, but implies responsibility in action wording. | Better attribution does not establish responsibility; the named false action is high risk. |
| Open-question preservation | Retains only the core unresolved process issue. | Lists open points, some synthesized rather than clearly preserved as open. | Keeps several open themes but turns others into settled next steps. |
| Useful detail | Good concise working/distribution reference; insufficient alone for later argument reconstruction. | Detailed enough for overall logic, though attribution and restraint are weak. | Most useful detailed reference for complex multi-speaker discussion, provided commitments are evidence-checked. |
| Risk of overinterpretation | Low through editorial compression, at cost of omitted context. | Moderate: consensus and implied collective actions strengthened. | Highest impact when a semantic overinterpretation is attached to a named person. |
| Product suitability | Strong concise working/distribution protocol; not detailed evidentiary reconstruction. | Better detailed reference than short distribution protocol, but needs restraint. | Best detailed reference of the comparison outputs; a separate later renderer should create the concise distribution protocol. |
### Overall conclusion
The human reference is very strongly editorially compressed. It captures the
main organizational outcome well and works as a concise working/distribution
protocol, but loses much argumentative context and can be insufficient weeks
later when reconstructing why the discussion evolved.
The anonymous-speaker Assistant output has strong topic coverage and
reconstructs overall process logic well. It tends to smooth disagreement into
stronger consensus than the transcript supports, is weaker at retaining who
held which position, and can overstate phrases such as “es wurde vereinbart”
or “die Teilnehmer waren sich einig.”
The mapped-speaker Assistant output best preserves differing viewpoints and
discussion structure. It is substantially better for reconstructing who argued
what in this complex multi-speaker discussion; speaker information increases
the value of a detailed protocol. It does not solve semantic interpretation of
commitments/responsibility. A false action is more dangerous when confidently
attached to a named person.
## Explicit regression expectations
These are qualitative expectations, not a synthetic gold protocol.
### R1 — Core purpose
The detailed protocol should preserve that the GTM-Hub is a cross-functional
collection, evaluation, and coordination point for ideas, markets, products,
or initiatives.
### R2 — Different perspectives
Preserve that participants approached the subject from different
perspectives/mental models; do not claim immediate detailed agreement.
### R3 — Strategic vs. operational distinction
Preserve the distinction among strategic evaluation/coordination, formal
process description/QMS, and operational project execution.
### R4 — Responsibility timing
Do not imply concrete project responsibility before group evaluation and
assignment. The transcript explicitly says responsibility does not yet exist at
that stage and is assigned later.
### R5 — Consensus strength
It is valid to report a basic/common understanding around collection,
discussion, and assignment of ideas. Do not turn unresolved process details
into final consensus.
### R6 — Speaker-aware value
For diarized processing, differing positions should remain attributable to
separate speakers/persons where mappings are confirmed.
### R7 — No responsibility invention
Speaker-aware processing must not turn an organizational suggestion into a
personal commitment. In particular, the statement that formal completeness
checking could be done by an administrative office/secretariat must **not**
become “Martin Tazl will delegate the formal review to an administrative
office” without explicit commitment evidence.
### R8 — Proposal vs. action
A suggestion to structure discussion into two parts must not automatically
become a personal action item for the person proposing it.
### R9 — Human-reference role
Detailed-protocol mode must not be forced to become as short as the human
reference.
### R10 — Detailed-protocol product direction
The main protocol should retain enough detail for a participant to reconstruct
the substance several weeks later. A future short/distribution protocol may
deliberately compress much more strongly.
@@ -0,0 +1,42 @@
{
"case_id": "gtm_hub_2026-09-07",
"meeting_date": "2026-09-07",
"meeting_id": "gtm-hub-meeting-2026-09-07",
"duration_seconds": 4715.52,
"transcript": {
"case_path": "transcript/transcript.txt",
"source_paths": [
"/opt/git-projekts/meeting-assistant/data/meetings/gtm-hub-meeting-2026-09-07/runs/b895cda2_2026-09-07_11-18-54_GTM-HUB_20260909_025239/protocol/transcript_input.txt",
"/opt/git-projekts/meeting-assistant/data/meetings/gtm-hub-meeting-2026-09-07/runs/284a47be_2026-09-07_11-18-54_GTM-HUB_20260909_031106/protocol/transcript_input.txt"
],
"sha256": "42ac52223cac36eeeb7c01e6d6673b2fb50973c4365b5ddaafdd12927b2f12eb",
"representation": "exact compact diarized protocol input"
},
"human_reference": { "path": "references/human_protocol_malte.txt" },
"comparison_provenance": {
"source_runs_are_separate_full_pipeline_executions": true,
"relevant_audio_transcription_diarization_and_protocol_input_artifacts_are_byte_identical": true,
"comparison": "equivalent diarized input conditions with anonymous versus confirmed speaker identity mappings",
"does_not_establish_operational_reuse_of_whisper_or_diarization_between_generations": true
},
"anonymous_speaker_reference": {
"case_path": "references/meeting_assistant_anonymous_speakers.md",
"source_path": "/opt/git-projekts/meeting-assistant/data/meetings/gtm-hub-meeting-2026-09-07/runs/b895cda2_2026-09-07_11-18-54_GTM-HUB_20260909_025239/protocol.md",
"source_run_id": "b895cda2_2026-09-07_11-18-54_GTM-HUB_20260909_025239",
"generation_timestamp": "2026-09-09T01:00:46+00:00",
"model": "qwen3.8:27b",
"diarization_enabled": true,
"speaker_mapping_count": 0,
"sha256": "134fe2caa0c798f46248f039a4e179fc5a1024f0469ca09f7f9ca176e5e30264"
},
"speaker_aware_reference": {
"case_path": "references/meeting_assistant_speaker.md",
"source_path": "/opt/git-projekts/meeting-assistant/data/meetings/gtm-hub-meeting-2026-09-07/runs/284a47be_2026-09-07_11-18-54_GTM-HUB_20260909_031106/protocol.md",
"source_run_id": "284a47be_2026-09-07_11-18-54_GTM-HUB_20260909_031106",
"generation_timestamp": "2026-09-09T01:30:44+00:00",
"model": "qwen3.8:27b",
"diarization_enabled": true,
"speaker_mapping_count": 5,
"sha256": "18c77411a7ecf869940e56308db258f177c1d813dbf7495f9d865a72a1fadf3d"
}
}
@@ -0,0 +1,21 @@
Teilnehmer:
EH, MT, BEF, JH, JMS
Thema 1: Prozessaufbau nach Mission 2030
Der Aufbau der Prozesse nach Mission 2030 erfolgt unabhängig von dem im GtM-Hub besprochenen Prozess.
Mission 2030 wird dafür in mehrere Abschnitte unterteilt und neu beschrieben. Die Abbildung dient der Zertifizierung nach ISO 9001.
EH hat den ersten Entwurf vorgestellt und den Unterschied zum GtM-Hub-Prozess erläutert.
Ziel ist es einen Großteil so aufzubauen, wie wir möchten, dass es bearbeitet wird. Es wird natürlich weiterhin Ausnahmen geben.
Thema 2: Verantwortlichkeiten und Projektablauf
Wir haben lange diskutiert und festgestellt, dass wir das Thema aus unterschiedlichen Perspektiven und mit unterschiedlichen Zielsetzungen angegangen sind.
Zielsetzung ist:
Verantwortung: Projektverantwortliche klar benennen.
Informationsfluss: Sicherstellen, dass Themen nicht in einzelnen Abteilungen bearbeitet werden, ohne andere betroffene Bereiche zu informieren.
Projektablauf: Den Projektablauf genauer beschreiben. Systematische Bewertung etc..
Dabei wurde angemerkt, dass es Kriterien braucht, wann dieser Projektablauf greift – nicht jede sinnvolle Idee lässt sich in ein einheitliches Raster pressen.
Grundsätzlich besteht Einigkeit darüber, wie Ideen ins Gremium getragen und dort weiterverarbeitet werden sollen.
Als nächster Schritt wurde festgehalten, dass wir die Präsentation von BD kommentieren und Feedback dazu geben.
@@ -0,0 +1,53 @@
# Meeting Protocol
## Einordnung und Zielsetzung des Gremiums
Die Teilnehmer diskutierten die Rolle und den Zweck des „GTM-Hub“ (Go-To-Market Hub). Es wurde klargestellt, dass das Gremium nicht dazu dient, einzelne Projekte im Detail zu steuern oder als Projektmanagement-Instanz zu fungieren, sondern als strategischer Filter und Koordinationspunkt für neue Ideen, Märkte und Initiativen im Rahmen der „Mission 2030“.
Ein zentraler Punkt der Diskussion war die Abgrenzung zwischen dem strategischen Prozess (Sammeln, Bewerten und Zuordnen von Ideen) und der operativen Projektumsetzung. Es wurde betont, dass die Geschäftsführung die Diskussion über neue Märkte und strategische Innovationen in diesem Gremium wünscht, um abteilungsübergreifende Transparenz zu schaffen und Doppelarbeit zu vermeiden. Die Teilnehmer waren sich einig, dass der Fokus zunächst auf der Definition des gemeinsamen Verständnisses und der groben Prozessschritte liegen muss, bevor in Details der einzelnen Projekttypen (z. B. Zulassung, Produktentwicklung) eingetaucht wird.
## Prozessbeschreibung und Integration in das Prozessmanagement
Es wurde vereinbart, den GTM-Hub-Prozess im bestehenden Prozessmanagement-System (im Kontext der ISO 9000-Anforderungen) abzubilden. Der Prozess soll als Unterprozess oder eigenständiger Prozess beschrieben werden, der die Schritte von der Idee bis zur Zuordnung an einen Verantwortlichen umfasst.
Die Diskussion umfasste folgende Aspekte:
* **Eingangsstufe:** Ideen können aus verschiedenen Quellen (Vertrieb, F&E, Marketing, externe Quellen) kommen. Es wurde vorgeschlagen, eine formale Eingangsstufe (z. B. eine E-Mail-Adresse oder ein Formular) zu definieren, die von einer neutralen Stelle (z. B. Sekretariat oder F&E als Sammelstelle) auf formale Vollständigkeit geprüft wird, um sicherzustellen, dass ausreichend Informationen vorliegen, bevor das Gremium tagt.
* **Entscheidungsstufe:** In der GTM-Hub-Runde wird diskutiert, ob die Idee strategisch passt, welche Abteilung federführend ist und ob weitere Informationen benötigt werden.
* **Zuordnung:** Nach der Entscheidung wird die Idee an den zuständigen Bereich (z. B. Business Development, Produktmanagement, F&E) zur weiteren Bearbeitung übergeben.
Es wurde darauf hingewiesen, dass die Beschreibung im Prozessmanagement so gestaltet sein muss, dass sie für alle Beteiligten nachvollziehbar ist und die Mindestanforderungen der ISO 9000 erfüllt, ohne unnötige bürokratische Hürden zu schaffen. Die Teilnehmer waren sich einig, dass die Beschreibung lebendig bleiben muss und bei Bedarf angepasst werden kann.
## Strategische Ausrichtung und Kriterien
Ein wiederkehrender Punkt war die Notwendigkeit eines strategischen Checks, bevor in die operative Bearbeitung eingestiegen wird. Es wurde kritisiert, dass in früheren Diskussionen zu schnell in Details gegangen wurde, ohne die strategische Passung (z. B. Marktfit, Deckungsbeitrag, strategische Priorität) ausreichend zu prüfen.
Die Teilnehmer diskutierten, welche Kriterien für die Bewertung von Ideen relevant sind. Dabei wurden unterschiedliche Ansätze genannt:
* **Strategische Kriterien:** Passung zur Mission 2030, Marktchancen, Wettbewerbslage.
* **Operative Kriterien:** Technische Machbarkeit, Ressourcenbedarf, Zeitplan.
Es wurde vereinbart, dass die Kriterien nicht für alle Projekttypen identisch sein müssen. Für neue Märkte und strategische Innovationen sind andere Bewertungskriterien relevant als für Produktverbesserungen oder Zulassungen. Die Teilnehmer waren sich einig, dass die Definition dieser Kriterien Teil der Prozessbeschreibung sein muss, aber nicht im Detail im GTM-Hub-Prozess selbst festgelegt werden sollte, sondern in den jeweiligen Fachprozessen.
## Rollen und Verantwortlichkeiten
Die Diskussion um die Rollen und Verantwortlichkeiten ergab folgende Punkte:
* **Sammelstelle:** Es wurde vorgeschlagen, eine neutrale Stelle (z. B. F&E oder Sekretariat) zu benennen, die für die formale Prüfung und das Einsammeln der Ideen verantwortlich ist. Diese Stelle trifft keine inhaltlichen Entscheidungen, sondern stellt sicher, dass die Ideen ausreichend beschrieben sind.
* **GTM-Hub-Runde:** Die Runde besteht aus Vertretern der relevanten Abteilungen (F&E, Marketing, Business Development, Qualitätssicherung, Produktmanagement). Sie trifft die Entscheidung über die Weiterverfolgung und die Zuordnung.
* **Projektverantwortliche:** Nach der Zuordnung liegt die Verantwortung für die operative Umsetzung bei der jeweiligen Fachabteilung. Der GTM-Hub behält die Übersicht und kann bei Bedarf Feedback einholen, ist aber nicht für die Projektsteuerung verantwortlich.
Es wurde betont, dass die Zuordnung an einen Projektverantwortlichen (oder eine Abteilung) ein zentraler Schritt ist, um Transparenz zu schaffen und Doppelarbeit zu vermeiden. Die Teilnehmer waren sich einig, dass die Definition der Verantwortlichkeiten im Prozess klar sein muss, aber die konkrete Person erst im Einzelfall bestimmt wird.
## Offene Punkte und nächste Schritte
* **Präsentation und Feedback:** Es wurde vereinbart, dass die aktuelle Fassung des GTM-Hub-Prozesses (bzw. die Präsentation mit den vorgeschlagenen Schritten) an alle Teilnehmer versendet wird. Die Teilnehmer sollen bis zu einem bestimmten Termin (nicht explizit genannt, aber implizit kurzfristig) Feedback geben.
* **Strategische Kriterien:** Die Definition der strategischen Kriterien für die Bewertung von Ideen ist noch offen. Es wurde vorgeschlagen, dies in einem separaten Schritt oder in Zusammenarbeit mit der Geschäftsführung zu klären.
* **Prozessbeschreibung:** Die formale Beschreibung des GTM-Hub-Prozesses im Prozessmanagement-System ist noch nicht abgeschlossen. Die Teilnehmer werden ihre jeweiligen Inputs einbringen, um die Beschreibung zu finalisieren.
* **Abgrenzung zu anderen Prozessen:** Die Abgrenzung zwischen dem GTM-Hub-Prozess und den bestehenden Prozessen (z. B. Vertriebsprozess, Entwicklungsprozess) muss noch klarer definiert werden. Es wurde vereinbart, dass der GTM-Hub-Prozess als Eingangspunkt für neue Ideen dient und die Zuordnung an die bestehenden Prozesse erfolgt.
## Maßnahmenliste
* **Verantwortlich:** [Nicht explizit zugewiesen, implizit die Teilnehmer des GTM-Hubs]
* **Maßnahme:** Versendung der aktuellen Fassung des GTM-Hub-Prozesses an alle Teilnehmer.
* **Frist:** [Nicht explizit genannt]
* **Verantwortlich:** [Nicht explizit zugewiesen, implizit die Teilnehmer des GTM-Hubs]
* **Maßnahme:** Feedback zur aktuellen Fassung des GTM-Hub-Prozesses geben.
* **Frist:** [Nicht explizit genannt]
* **Verantwortlich:** [Nicht explizit zugewiesen, implizit die Teilnehmer des GTM-Hubs]
* **Maßnahme:** Definition der strategischen Kriterien für die Bewertung von Ideen.
* **Frist:** [Nicht explizit genannt]
* **Verantwortlich:** [Nicht explizit zugewiesen, implizit die Teilnehmer des GTM-Hubs]
* **Maßnahme:** Finalisierung der formalen Beschreibung des GTM-Hub-Prozesses im Prozessmanagement-System.
* **Frist:** [Nicht explizit genannt]
@@ -0,0 +1,45 @@
# Meeting Protocol
## Einordnung und Zielsetzung des GTM-Hubs
Die Teilnehmer haben die ursprüngliche Intention und den aktuellen Stand des „GTM-Hub“-Prozesses (Go-To-Market Hub) diskutiert. Der Prozess wurde im Kontext der Strategie „Mission 2030“ eingeführt, um abteilungsübergreifende Ideen, neue Märkte und Produktinnovationen zu bündeln, zu bewerten und zuzuordnen.
**Diskussionsverlauf und Positionen:**
* **Jovana Husemann** betonte, dass die Runde primär strategischen Charakter hat. Sie verwies darauf, dass die Geschäftsführung die Diskussion über neue Märkte und strategische Innovationen in diesem Gremium gewünscht habe. Sie kritisierte, dass in den letzten Sitzungen zu viel Zeit für prozessuale Details aufgewendet wurde, während der strategische Teil (die Bewertung, ob eine Idee zur Strategie passt) noch nicht abschließend geklärt ist. Sie forderte, dass vor der detaillierten Prozessbeschreibung die strategischen Kriterien und die Zustimmung der Abteilungen zu den definierten Stufen (Stages) erfolgen müssen.
* **Malte Schnau** stellte fest, dass der Fokus bisher zu stark auf der reinen Prozessbeschreibung lag, während die Kernanforderung – die transparente Sammlung von Ideen, die Zuordnung an einen Projektverantwortlichen und die Vermeidung von Doppelarbeit – im Vordergrund stehen sollte. Er argumentierte, dass der Prozess lebendig bleiben muss und nicht im Detail für jede Projektart (z. B. Zulassung vs. Produktentwicklung) vordefiniert werden kann, da die Anforderungen stark variieren.
* **Henning Ehrenberg** ordnete die Diskussion in den Kontext des Qualitätsmanagementsystems (ISO 9000) ein. Er betonte, dass die Beschreibung der Prozesse nicht nur der internen Abstimmung dient, sondern auch der Nachweisführung gegenüber Auditoren (Rezertifizierungsaudit) dienen muss. Er plädierte dafür, die Prozesse so zu beschreiben, dass sie zu 90 % der Realität entsprechen, um Abweichungen im Audit zu minimieren.
* **Björn-Erik Falkenau** hob hervor, dass die Runde dazu dienen soll, die „Geschichte“ und den strategischen Check vor der operativen Umsetzung sicherzustellen. Er kritisierte, dass oft zu schnell in operative Details eingetaucht wird, ohne vorher zu prüfen, ob die Idee zur Unternehmensstrategie passt oder ob bereits andere Abteilungen an ähnlichen Themen arbeiten. Er forderte, den Prozess global in wenigen Hauptschritten zu definieren, bevor Details geklärt werden.
* **Martin Tazl** stellte klar, dass die Forschung & Entwicklung (F&E) nicht die operative Verantwortung für die strategische Bewertung neuer Märkte oder Business-Entwicklung trägt. Die Rolle der F&E beschränkt sich auf die formale Prüfung der Eingangsideen (Vollständigkeit) und die technische Machbarkeit, falls die Idee in den Bereich der Produktentwicklung fällt. Er betonte, dass die Zuordnung der Verantwortung (wer „den Hut aufhat“) erst nach der strategischen Bewertung in der Runde erfolgt.
**Entscheidungsgrundlagen und Konsens:**
* Es besteht Einigkeit, dass der GTM-Hub als zentraler Sammel- und Bewertungspunkt für neue Ideen, Märkte und Produkte dient.
* Die strategische Bewertung (Passt die Idee zur Mission 2030?) muss vor der operativen Zuordnung an eine Fachabteilung erfolgen.
* Die Prozessbeschreibung muss die Anforderungen des Qualitätsmanagements (ISO 9000) erfüllen, darf aber keine unnötigen Fesseln für die operative Arbeit anlegen.
* Die Verantwortung für die inhaltliche Bewertung neuer Märkte liegt nicht primär bei der F&E, sondern erfordert die Einbindung von Business Development, Marketing und Vertrieb.
## Prozessbeschreibung und Integration in das Qualitätsmanagement
Die Teilnehmer diskutierten die technische Umsetzung der Prozessbeschreibung im Rahmen des Qualitätsmanagementsystems.
**Diskussionsverlauf und Positionen:**
* **Henning Ehrenberg** erläuterte, dass die Beschreibung der Prozesse im Qualitätsmanagementsystem (QMS) bereits seit über einem Jahr läuft und unabhängig von der GTM-Hub-Diskussion fortgeführt wird. Er zeigte auf, wie der GTM-Hub als Unterprozess in die bestehende Prozesslandschaft (z. B. Vertriebsprozess, F&E-Prozess) integriert werden kann. Er betonte, dass die Beschreibung so erfolgen muss, dass sie für alle Beteiligten nachvollziehbar ist und die Mindestanforderungen der ISO 9000 erfüllt.
* **Jovana Husemann** äußerte Bedenken, dass die reine Prozessbeschreibung im QMS von den Mitarbeitern nicht gelesen oder beachtet wird. Sie argumentierte, dass die strategische Ausrichtung und die Zustimmung der Abteilungen zur Strategie wichtiger sind als die formale Prozessbeschreibung. Sie forderte, dass die Präsentation der strategischen Stufen (Stages) von allen Abteilungen kommentiert und bestätigt werden muss, bevor die Prozessbeschreibung finalisiert wird.
* **Malte Schnau** schlug vor, die Diskussion in zwei Teile zu gliedern: 1) Die grobe Definition des Ablaufs (Sammlung, Bewertung, Zuordnung) und 2) die detaillierte Definition der Kriterien für verschiedene Projektarten. Er betonte, dass der erste Teil (der Ablauf) schnell abgeschlossen werden sollte, um Fortschritt zu erzielen.
* **Martin Tazl** wies darauf hin, dass die formale Prüfung der Eingangsideen (z. B. Vollständigkeit der Informationen) eine einfache Aufgabe ist, die nicht zwingend von der F&E, sondern von einer administrativen Stelle (z. B. Sekretariat) übernommen werden kann. Die F&E sollte sich nur auf die technische Machbarkeit konzentrieren, wenn die Idee in den Bereich der Produktentwicklung fällt.
**Entscheidungsgrundlagen und Konsens:**
* Die Prozessbeschreibung im QMS muss die strategischen Ziele des GTM-Hubs abbilden und die Anforderungen der ISO 9000 erfüllen.
* Die formale Prüfung der Eingangsideen kann von einer administrativen Stelle übernommen werden, um die F&E zu entlasten.
* Die strategische Bewertung und die Zuordnung der Verantwortung erfolgen in der GTM-Hub-Runde.
## Nächste Schritte und offene Punkte
Die Teilnehmer haben die nächsten Schritte zur Klärung der offenen Punkte besprochen.
**Maßnahmen und Verantwortlichkeiten:**
* **Jovana Husemann** wird die Präsentation der strategischen Stufen (Stages) an die Teilnehmer senden und um Kommentare bitten. Sie erwartet, dass die Abteilungen die strategische Ausrichtung bestätigen oder kommentieren.
* **Malte Schnau** wird die Diskussion in zwei Teile gliedern: 1) Die grobe Definition des Ablaufs und 2) die detaillierte Definition der Kriterien. Er wird vorschlagen, den ersten Teil (den Ablauf) schnell abzuschließen.
* **Henning Ehrenberg** wird die Prozessbeschreibung im QMS so anpassen, dass sie die strategischen Ziele des GTM-Hubs abbildet und die Anforderungen der ISO 9000 erfüllt.
* **Martin Tazl** wird die formale Prüfung der Eingangsideen an eine administrative Stelle delegieren, um die F&E zu entlasten.
**Offene Punkte:**
* Die strategische Bewertung der Ideen muss vor der operativen Zuordnung an eine Fachabteilung erfolgen.
* Die Prozessbeschreibung im QMS muss die Anforderungen der ISO 9000 erfüllen.
* Die formale Prüfung der Eingangsideen muss von einer administrativen Stelle übernommen werden.
@@ -0,0 +1,193 @@
SPEAKER_03: wie das mit Off-Handing-Warten lohnt. Okay. Vielleicht könnt ihr mich noch mal auf Stand holen, Giovanna. Wir hatten das letzte Mal vor drei Wochen zusammengesessen bezüglich der... Also beziehungsweise da war ich nicht mit da, das war im Urlaub. Lars hat da einmal ein Protokoll geschrieben mit der Namensänderung und wie der Ablauf der Projekte... Das war mal ein bisschen... Ah, hallo. Hi. Dass da nochmal über den Ablauf gesprochen wurde. Ich habe ja mit Vorplag gemacht. Darüber wurde dann ja noch mal mit ausgiebig drüber diskutiert, über die möglichen Abläufe. Und mein letzter Stand war jetzt, dass ihr am 28. noch mal sitzen wolltet, August, bezüglich eines Vorschlags. Da stand ein Protokoll von Lars.
SPEAKER_02: Ja, wir haben am 28. zusammengesessen, aber das hatte ein anderes Thema. Da ging es darum, dass wir generell das ganze Thema Mission 2030 im Prozessmanagement abbilden müssen. Und wir sind angefangen mit dem Vertriebsprozess oder haben an dem Vertriebsprozess gearbeitet, wie wir den abbilden wollen. Wir haben das auch ganz bewusst erstmal nicht in einer Riesenrunde gemacht. Es geht auch wirklich nur um den Vertriebsprozess. Also es kommt irgendwas bei Commercial Sales an. Ne, stimmt gar nicht. Pre-Sales. Und wie geht es dann weiter?
SPEAKER_04: Ja, und grundsätzlich zum Termin, als du nicht da warst, also wir haben tatsächlich zwei Themen, meiner Meinung nach. Einmal das sehr strategische Thema, wie wollen wir, dass die Geschäftsführung geklärt hat, die Idee kommt rein, wir wollen die Themen einmal zu klären. Und das haben wir als Development erstmal eine Präsentation gemacht, als du das mitbekommst.
SPEAKER_03: Ja, ich habe es gesehen.
SPEAKER_04: Genau, und das andere Thema ist an diesem Prozessmanagement, das ist eigentlich, also ich kann das wenig zu sagen, da gibt es schon einen Weg, wie du das machst, Henning, oder also, ich glaube, ihr braucht mir da gar nicht so ein Business Development dafür.
SPEAKER_02: Also wir werden die mit Sicherheit brauchen, aber nicht zum jetzigen Zeitpunkt, weil, was wir vorhaben, ist ja, dass wir, also wir haben dieses Mission 2030, was irgendwann mit Business Development losgeht, was über mehrere Phasen dann irgendwann hoffentlich darin endet, dass wir Ware ausgeliefert haben. Und das ist ja ein mehrstufiger Prozess und wir sind jetzt so weit, dass wir gesagt haben, okay, wir werden das zerlegen in Business Development, in Pre-Sales, bis zu dem Zeitpunkt, wo eine Ausschreibung rauskommt, dann ist es im Wesentlichen der Commercial Sales, wenn die Ausschreibung raus ist. Und dann haben wir den Lieferprozess, den Lieferprozess, den haben wir schon, den brauchen wir nicht ändern. Und die anderen drei Prozesse müssen wir beschreiben. Also wir müssen schon beschreiben, wir müssen beschreiben, was dort passiert. Und wir müssen auch gewisse Regeln einhalten. Also das ist jetzt der ISO 9000-Gedanke, da wird es sicherlich die ein oder andere Anforderung geben, warum man sagt, das muss man jetzt so und so machen und man muss es auch so und so, man müsste es in irgendeiner Form auch dokumentieren können, zeigen können, also nicht nur verbal erzählen, wir arbeiten so, sondern es auch zeigen können. Aber das können wir alles relativ light halten und wir haben jetzt beim letzten Mal wirklich angefangen und haben gesagt, okay, wir nehmen uns mal den zweiten Block, wie arbeitet eigentlich Technik mit Precels, wo geht das ganze Spielchen los, bis hin, da kommt jetzt irgendwann die Ausschreibung. Und da sind wir auch beim letzten Mal, wir sind nicht fertig geworden, wir haben ja einfach angefangen. Und genauso werden wir das für die anderen Teilprozesse auch beschreiben müssen, das wird nicht anders gehen.
SPEAKER_04: Ja, das habe ich jetzt gesagt, weil ich sehe, das sind zwei Projekte meiner Meinung nach, weil eine Sache ist diese technikischen Teil, wir wollen die Projekte einmal, also hier in diese, ich weiß nicht, dieses Akademium erstmal, also dokumentieren, auswerten und dann umsetzen und die andere ist dann, wie begleiten wir das für den Prozessmanagement. So sehe ich das.
SPEAKER_02: Ja, also das, also ich habe beim letzten Mal beschrieben, dass wir an dem Prozessmanagement schon seit über einem Jahr arbeiten. Das hat Lars im Protokoll etwas anders dargestellt. Das machen wir sowieso. Das ist auch überhaupt keine Frage, das hat auch nichts damit zu tun, ob wir jetzt diese Runde hier haben oder nicht. Wir müssen beschreiben, wie unsere Prozesse sind.
SPEAKER_04: Aber das läuft schon.
SPEAKER_02: Genau, das läuft schon und das läuft schon seit einem Jahr und das werden wir sukzessive abarbeiten und wir werden sicherlich auch noch, ich war ja letzte Woche bei einem Seminar, was durch die neue ISO 9000 auf uns zukommt, da wird es auch noch ein paar neue Forderungen geben, die müssen wir auch noch berücksichtigen und da werden wir auch alleine durch unsere organisatorischen Änderungen ein paar Sachen etwas detaillierter beschreiben müssen. Eine BBG gab es früher nicht. Das, was früher die alte BBG war, müssen wir in irgendeiner Form auch mit beschreiben. Also das hat aber alles hiermit nichts zu tun.
SPEAKER_04: Ja, und das hat Sie, wenn wir schon da sind, das ist schon vorbereitet für das BBG. Ich kann Sie das zukommen lassen.
SPEAKER_02: Ne, ich meine nicht die neue BBG, ich meine die alte BBG. Die neue BBG hat für uns, für mich jetzt hier mit dem Thema ISO 9000 und Prozessmanagement nichts zu tun, weil das ist ja eine separate Firma. Das, was früher BBG war, was jetzt EDD ist, das haben wir bisher in unseren Prozessen wenig beschrieben. Das werden wir in Zukunft etwas detaillierter beschreiben müssen.
SPEAKER_03: Okay. Aber dann, wie du sagtest, das heißt, das ist einmal das Thema Mission 2030 allgemein nach ISO 9000, dann muss ich gestehen, habe ich jetzt für mich noch nicht verstanden, wo und bei wem gerade der Ball liegt und wie der Stand ist bei dem Part, bei dem anderen Part, den du gerade befinden hast. Also bei dem, wie, dem Ablauf, wie kommen die Ideen hier rein? Ich meine, das fing ja an, das war jetzt nur einmal gerade aus meiner Perspektive. Wir haben das hier besprochen vor, ich glaube, mittlerweile sechs Wochen ungefähr, haben uns doch geeinigt, dass ich einen Vorschlag mache. Da hatten wir dann ja auch noch zusammengesessen. Den Vorschlag, der hat nochmal zur Diskussion geführt, zwei, drei Wochen später und da weiß ich gerade nicht, genau, dann habt ihr die Mail fertig gemacht mit eurem Vorschlag und dann war ich im Urlaub und ab jetzt ist das da für mich halt einmal der Punkt, wo ich gerade nicht weiß, wo das liegt, wie das ist, was da besprochen wurde.
SPEAKER_02: Wir haben hier, da weißt du glaube ich, den Urlaub, da haben wir hier drüber diskutiert und das Ganze ist ja losgegangen an der Diskussion. Wir haben ja im F&E haben wir einen sogenannten Ding-Pool und ich habe dann gesagt, naja, aus meiner Sicht ist das jetzt relativ einfach. Man könnte das, was wir hier vorhaben, einfach dort mit beschreiben als Prozess, als Unterprozess und eigentlich ist es in meinem Augen auch gar nicht viel. Ich habe angefangen zu versuchen, das zu beschreiben. Das können wir uns jetzt gerne angucken. wo wir da jetzt gerade stehen. auf einen kleinen Augenblick.
SPEAKER_03: Ja, weil in dem Protokoll ließ es sich wirklich ein bisschen vermischt von den zwei Punkten, muss ich gestehen.
SPEAKER_04: also deswegen habe ich so angefangen zu erklären, also wir haben tatsächlich, meiner Meinung nach, das sind zwei getrennte Projekte. Und wir haben am Anfang gesagt, dieses Gremium trifft sich einfach um das Projekt auszuwerten und das hat mit Prozessmanagement, meiner Meinung nach, nichts zu tun.
SPEAKER_03: Klar ist das verbunden.
SPEAKER_04: Ja, genau. Also aber die letzten zwei Termine, es ging mehr um Prozessmanagement statt diese...
SPEAKER_02: Also es ging nicht um Prozessmanagement, es ging darum, wo und wie beschreiben wir, was wir tun wollen. Und das gehört ins Prozessmanagement, aber ich kann im Prozessmanagement oder sollte im Prozessmanagement halt das beschreiben, wie wir arbeiten. Punkt. Und das, wie wir arbeiten, sollte gewisse Mindeststandards erfüllen und dann ist alles gut. Vergiss Prozessmanagement, das ist einfach nur das Cool, womit ich es beschreibe. Wir müssen uns darüber einig sein und im Klaren sein, wie wir es eigentlich abbilden, wie wir es leben wollen. Und dann geht es nur noch darum, das, was wir leben wollen, vernünftig abzubilden.
SPEAKER_04: Gut, und das, wie wir das leben wollen, das hat Business Development Vorbereitung, das habe ich ja mal präsentiert
SPEAKER_UNASSIGNED: und das habe ich ja mal präsentiert
SPEAKER_04: und das habe ich mitbekommen.
SPEAKER_03: Okay. Aber ich wusste nicht, ob das jetzt nochmal besprochen wurde oder ob das jetzt für alle so in Ordnung war.
SPEAKER_04: Das haben wir ja nicht besprochen, aber es war ja bei allen anderen hier. Und also das ist jetzt unser Input. Also ich habe auch keinen Antwort bekommen. Ich weiß es auch nicht. Das ist wohl mein letzter Stand. Also ich kann nichts mehr sagen. Weil innerhalb dieser Präsentation habe ich auch klar nochmal deutlich gemacht, dass wir ja tatsächlich jemand brauchen, der dann diese Umsetzung übernimmt. Wer kümmert sich von Projektmanagementseite für die Umsetzung? strategisch, wie wir das machen wollen. Das ist einfach nur, was die Geschäftsführung gesagt hat, einfach nur mal dokumentiert und mit Ideen, welches geht und so weiter. Aber einer muss das umfassen, oder einer muss sich kümmern, dass das wirklich stattfindet, zu koordinieren, zu managen. Wir brauchen einen Projektmanager dafür.
SPEAKER_02: Für was jetzt? Bin ich nicht ganz sicher, was du meinst.
SPEAKER_04: Also das ist jetzt ein bisschen das Problem, weil du warst in die ersten zwei Runden nicht dabei. Also diese Gremium müssen jetzt...
SPEAKER_02: Das sollte kein Problem sein, weil im Idealfall haben wir eine Beschreibung, die so ist, dass jeder sich sofort damit zurechtfinden kann. Wenn das nicht so ist, dann ist die Beschreibung nicht gut. Ja, sorry. Also ich denke jetzt mal im Sinne vom Prozessmanagement.
SPEAKER_04: Ja gut, aber ich habe das nicht beschrieben, weil also das ist... Also wir haben die Protokolle. Hast du die Protokolle gelesen?
SPEAKER_02: Ich habe die Protokolle gelesen, ja.
SPEAKER_04: Also diese Gremium müssen...
SPEAKER_02: Ich habe die nicht mehr alle vor Augen, also müsste ich dann nochmal wieder neu lesen.
SPEAKER_04: Also diese Gremium ist hier, der Stand von Geschäftsführung hat gesagt, wir...
SPEAKER_02: Du brauchst nicht immer betonen, dass das Stand von der Geschäftsführung ist. Wir arbeiten hier zusammen.
SPEAKER_04: Also wenn du willst, ich kann dir die Geschichte von Anfang an erzählen. Wenn du das nicht möchtest, ich das erzähle, dann vielleicht kannst du selber die Protokolle lesen. Das muss ich gar nicht sagen. Ich kann dir nur sagen, warum ich hier sitze.
SPEAKER_02: Das ist der Satz, deswegen der Geschäftsführung. Ich dachte, wir arbeiten hier zusammen. Habe ich da irgendwas falsch verstanden?
SPEAKER_04: Also ich habe es nicht falsch verstanden, aber du hast... Nun, wir würden folgst und was ich sagen. Ich kann das sagen, wie ich das möchte. Ich habe recht dafür, dass du sagen, wie ich das möchte. Ja. Aber ich habe nicht einmal gesagt, ich akzeptiere, was du sagst.
SPEAKER_02: Also, mein Kenntnisstand ist, dass wir beschreiben wollten, wie wir die Projekte hier reinfliegen und wie wir sie dann zuordnen. Nicht, wie wir einzelne Projekte abwickeln. Habe ich dann irgendwas falsch verstanden? Das war mein Kenntnisstand. Nochmal vom Wording. Also, es ging doch darum, dass von irgendwo aus verschiedenen Stellen im Unternehmen oder aus den Prozessbeteiligten Ideen hier eingebracht werden, wo dann diskutiert wird, macht es Sinn, die weiter zu verfolgen oder nicht?
SPEAKER_00: Ja.
SPEAKER_02: Und wenn wir der Meinung sind, dass es Sinn macht, die weiter zu verfolgen, auch zu definieren, wo sie weiter verfolgt werden. Genau.
SPEAKER_03: Eine Zuordnung an einem Projektleiter oder wer noch immer. Das war mein Verstand.
SPEAKER_02: Das war das, was ich verstanden habe, warum es hier jetzt eigentlich gehen soll, dass wir das beschreiben beziehungsweise definieren. Genau.
SPEAKER_03: Und ich hatte mich, weil es einen alten Prozess gab, über den wir gesprochen haben, hatte ich mich im ersten Schritt an diesem alten Prozess, nicht um jetzt einen Prozess abzubilden, aber einfach, da war ja schon mal ganz schön beschrieben, wie man Ideen so durch ein Unternehmen eben durchschleusen kann. Habe den als Aufschlag genommen und was ergänzt, was wir in der Diskussion besprochen hatten. Also das halt einfach, BD-Projekte haben wir fehlt, die mussten dann noch kurz beschrieben werden. Aber ich habe es nicht in dem tiefsten Detailgrad gemacht, wohlwissend, dass wir auch einfach ein Gefühl dafür kriegen müssen, je nachdem was für Projekte da auf den Tisch kommen. Das konnte ja wirklich von wir brauchen einen neuen Aufsatz für den Stapler sein, bis hin zu wir wollen mit Produkt XY auf den neuen Markt oder mit System XY oder mit einem ganz neuen Produktidee in der Region, wo wir noch nicht waren, wie auch immer, das ist jetzt die Funktionssache. Das muss ja alles irgendwo abbildbar sein und das kann man natürlich nicht im Detail beschreiben. Das heißt, ich habe versucht, die Filterkategorien anzupassen an die Sachen, die sich in den letzten zwei Jahren verändert haben und das so zu beschreiben. Genau. Und dann kam halt die Diskussion und das, genau, die Präsentation von dir, was von dem, was ich das gelesen habe, eigentlich in eine ziemlich ähnliche Richtung ging, wie was das andere auch beschrieben hat, mit ein, zwei Punkten, die vom Detailgrad tiefer waren. Aber im Großen und Ganzen ist die Frage, die ich mir jetzt stelle, wir sitzen hier jetzt sechs Wochen ungefähr zusammen seit dem ersten Aufschlag. Ich denke, dass mit dem Abbilden müssen wir meiner Meinung nach einfach zum Ende kommen, uns irgendwie einigen und dann eine Lösung zu finden, dass wir das einfach anfangen können. Das ist zumindest meine Meinung. Ich meine, unabhängig davon, ob man es jetzt als Präsentation oder als Prozess abbildet, das ist ja lebend. Also wenn wir in zwei Monaten merken, dass etwas überhaupt nicht passt, dann müssen wir es ja wieder anpassen, grundsätzlich. Zumindest meine Meinung.
SPEAKER_02: Also ich habe versucht, warte mal, wir machen das jetzt am besten, sollte ich mich über Teams einreden. Ja, ich denke, das ist ein Kastellon. Ja.
SPEAKER_03: Ich habe mich hier das Mikro ein, und das ist ein Kastellon.
SPEAKER_UNASSIGNED: Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_03: Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Und das ist ein Kastellon.
SPEAKER_04: Und das ist ein Kastellon. Und das ist ein Kastellon.
SPEAKER_03: Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_UNASSIGNED: Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_02: Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_01: Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_UNASSIGNED: Das ist ein Kastellon.
SPEAKER_02: Das ist ein Kastellon. Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_UNASSIGNED: Das ist ein Kastellon.
SPEAKER_02: Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_UNASSIGNED: Das ist ein Kastellon. Das ist ein Kastellon.
SPEAKER_02: Ich habe versucht, in dem F&G-Prozess, das ist der Prozess, so wie wir ihn heute fast haben. Ein paar kleine Änderungen schon. Also wir hatten bisher hier vorne einfach das Thema "Ideenpool".
SPEAKER_03: Ja. Willst du mit rüber kommentieren mit rauf gucken, das kann man deutlich besser sehen, Johanna. Willst du mit rüber kommentieren mit rauf gucken, das kann man deutlich besser sehen.
SPEAKER_02: Also statt dass wir hier den Ideenpool haben, habe ich hier jetzt einfach einen Unterprozess eingeführt. Den habe ich genannt "Go-To-Market Hub", so wie wir das ja auch als Namen besprochen haben. Und hier drin in der Beschreibung habe ich im Wesentlichen das reingenommen, was Lars auch schon als Texte verteilt hatte. Sorry, das ist ein bisschen klein. Alle zwei bis vier Wochen findet eine Abstimmung, einen Austausch über neue Themen, neue Initiativen, neue Produktideen, neue Marktideen. Hieran nehmen mindestens teil, F&E, WM, Business Development, Marketing, Qualitätssicherung. In dieser Runde werden Themen besprochen, die perspektivisch marktrelevant sind. Wir haben in der Verbesserung der Kommunikation der vertriebsunterstützenden Abteilungen ein Meeting eingeführt. Koordinierung, vertriebsvorbereiten der vertriebsunterstützenden Prozesse, Koordinierung, Priorisierung. Im Prinzip ist es das, ich sage mal, ich habe das einfach rauskopiert aus dem, was Lars beim letzten Mal verteilt hat. Und daraus waren letztendlich diese vier Themen, wo ich, ich habe es jetzt einfach mal rüberkopiert. Ich bin mir inhaltlich momentan nicht sicher, ob das so sinnvoll ist und so passiv ist. Ich sage mal, als Beispiel, neue Märkte, weiß ich nicht, ist neue Märkte im Entwicklungsprozess, den wir als Entwicklungsprozess unter F&E sehen. Da war ich mir jetzt, nachdem ich das gelesen habe oder neu gelesen habe, nicht mehr sicher, ob das so an der Stelle richtig ist. Also ja, klar, neue Produkte ist mal auf jeden Fall F&E-Prozess, neue Märkte, weiß ich nicht.
SPEAKER_04: Deswegen habe ich das letztes Mal auch gesprochen, deswegen habe ich das zweimal auch gesagt. Das ist kein F&E-Prozess mehr, sondern wir haben gesagt, das bleibt nur in der F&E, wo das war. Weil ich muss nochmal sagen, pünktmal, die gewünscht von Geschäftsführung war, dass die neuen Märkte auch hier diskutiert werden in der Runde. Und du warst auf den ersten Termin nicht dabei und die Geschäftsführung hat gesagt, es gibt die bestehenden Märkte und die neuen Märkte. Bestehende Märkte betreibt weiter Betrieb und für die neuen Märkte, alles, was wir über die neuen Märkte diskutieren, da können wir hier auf den Tisch legen und zusammen diskutieren. Und deswegen ist das jetzt nicht ein F&E-Prozess, sondern ist ein Prozess für die gesamten Unternehmen. Aber dann hat Martin gesagt, und ich habe damit leben, das war für mich in Ordnung, weil das schon war in F&E, sollte in F&E bleiben. Aber grundsätzlich macht das nicht in F&E, das ist was anderes, weil es geht um die Geschäftsentwicklung und Betrieb, dass es hier anfängt, so mit dieser Gremie. Es liegt nicht um die Erweiterung und bestehenden Prozess, das schon existiert und das funktioniert, meiner Meinung nach, sondern es geht um die Umsetzung von Strategie oder von Mission 2030. Und deswegen war der strategische Teil von Geschäftsführung und Erweiterung basierend auf dessen, was wir auf der Strategiesitzung gehört haben.
SPEAKER_02: Ja, so jetzt ist die Frage, ich sag mal, Björn, Martin, wie seht ihr das vom Verständnis her? Also zumindest mal, Björn und ich stellen das zumindest in Frage, ob das hier überhaupt mit Entwicklungsprozess richtig beschrieben ist oder geht es uns? Und das wäre jetzt die Frage, geht es uns nur da, nur in Anführungsstrichen darum, hier in der Runde das sozusagen zuzuordnen, das auf die Liste zu packen von allen möglichen Prozessen, wo dann an der Stelle eben steht, wird nicht in der Entwicklung, sondern wird dann so anders bearbeitet. Also das ist für mich die Grundsache. Also einfach nur, dass wir überhaupt ein gemeinsames Verständnis dafür kriegen.
SPEAKER_01: Genau, so sehe ich das. Also ich will mich überhaupt gar nicht im Entwicklungsprozess, im Entwicklungsprozess im Sinne einer F&E-Entwicklung mit Business Developing beschäftigen. Das ist nicht meine Aufgabe, ist nicht meine Kompetenz, kann ich nicht, brauche ich auch nicht. Nur wir waren uns ja einig, dass der Eingang unter Umständen nicht, beim Eingang von irgendetwas unter Umständen noch nicht klar ist, in welchem Trichter es dann weiterfällt. Und deswegen haben wir gesagt, okay, wir nehmen halt einen bestehenden Prozess, wo bereits Ideen gesammelt werden. Und dann, ob es dann in Swimlanes aufgeteilt wird, wie du das damals genannt hast, glaube ich, oder ob wir dann in Unterprozesse zerfallen, das ist von meiner Seite egal. Oder ob wir den Prozess namentlich auch so verallgemeinern, dass wir eben alles in diesem Prozess machen können, was irgendwie eine Entwicklung im weitesten Sinne ist. Eine Marktentwicklung ist ja auch eine Entwicklung. Es muss ja nicht ein Produkt sein oder ein Design oder so. Also da kann ich mit beiden leben. Aber nein, es ist nicht originär im Entwicklungsprozess, wie er bisher war, richtig angeordnet, aus meiner Sicht.
SPEAKER_03: Ich verstehe die Bezeichnung jetzt auch so. Sorry, nicht, dass ich, das ist immer blöd, wenn man bemerkt, wo man nicht dabei war. Aber ich verstehe das jetzt einfach so, dass es bedeuten soll, wir führen die drei Punkte durch diesen, wie hieß er, auch 0,8. Und alles andere, bestehende Märkte, Anwendungen, Produkte durch diesen Neubau, den wir, wo wir gerade eben drüber gesprochen haben, diesen Punkt 1, durch den Prozess. So habe ich das jetzt verstanden. Das teilt sich halt auch, aber letztendlich, ja, es ist doch einfach nur ein Trichter. Und danach verteilen wir es ja hier gemeinsam an den Projektverantwortlichen. Das, was, so wie ich es verstanden habe, eine der Anforderungen war. Also einmal, dass wir untereinander informiert sind und jeder weiß, welche Projekte unabhängig, wo er im Unternehmen sitzt und wo es bearbeitet wird, weiß, dass was passiert, dass nicht Dinge doppelt bearbeitet werden. Und dass Leuten gut aufgesetzt wird, um dieses Projekt wirklich durchzuführen. Und das war ja jetzt hier sehr zu sagen, das machen wir doch auch.
SPEAKER_04: Ja, aber um das zu kommen, so Prozess zu kommen, müssen wir uns erstmal einigen, sind wir im strategischen Teil einverstanden. Weil Henning hat jetzt gesagt, ist es so, bestimmte Märkte, neue Märkte, neue Anwendungen, strategische Innovation. Also für mich kommt Prozessmanagement am Ende, wenn alles schon hier abgestimmt ist, wie wollen wir weiter zum Unternehmen machen. Dann kann man weiter, weil die kann ich jetzt zum Beispiel Betriebe klären, am Ende mit Prozessmanagement, also was wir hier teilen. Wie kann jetzt Betriebe mitmachen, wenn ich jetzt komme und sage, deswegen war ich tatsächlich nicht einverstanden, warum das Betriebeforschung bleibt. Weil ich finde, es ist irgendwo, weil ich jetzt denke, für mich ist Research und Gewalt, das ist Wissenschaft. Und jetzt reden wir hier, so muss ich die strategische Innovation, das ist Forschung und Entwicklung, aber mit dem Ziel, dass ich Betriebeaktivität optimieren kann. Deswegen würde ich sagen, wir sollten den Prozess kopieren und unter anderem Namen nennen und dann nur strategische Innovation unter Forschung und Entwicklung probieren. Aber bevor wir zum Prozess kommen, muss ich sagen, wir müssen erstmal einig, von strategischer Zeit, der einverstanden, dass der Weg nicht so weit ausgetürt wird.
SPEAKER_02: Aber dadurch, da steht ja nur F&E, weil die Verantwortlichkeit für den Ablauf in der, also ich habe die Lösung, das ist relativ einfach. Also, Johanna sagt, dass im Prinzip, du störst dich daran, dass das jetzt im F&E-Prozess ist. Das ist aber inhaltlich überhaupt kein Problem, weil das, was ich jetzt hier eingefügt habe, ich mache es mal hier auf, dieses Kästchen, was auch immer wieder reinsteigt, Das kann ich zum Beispiel an zwölf verschiedenen Stellen in den Vision 2030 Prozess einbinden, das kann ich an drei Stellen im PM-Prozess einbinden, an fünf Stellen, ich kann das überall, wo ich sage, da können Ideen kommen, bla bla bla und so weiter und so fort. Und nach meinem Verständnis geht es hier jetzt ausschließlich darum, und wir bilden das jetzt gerade hier ab, das können wir hinterher auch genauso woanders machen, geht es darum, dass wir sagen, wie treffen wir die Entscheidung. Hier vorne, hier vorne kommt irgendwas rein, ich wäre jetzt mittlerweile so weit, dass ich sagen würde, wir schmeißen dieses Kästchen hier komplett raus, perspektivisch, aber ich muss ja immer leider ein bisschen rum springen, also dieses Kästchen schmeißen wir hier komplett raus, das ist nur irreführend und ist eigentlich auch in meinen Augen gar nicht hilfreich. So, wir haben dann hier das Kästchen Ideenpool, da kommt irgendwas rein, wir diskutieren das hier in der Runde, wir fragen eventuell, ergänzen wir Informationen an, weil wir sagen, sorry, das ist nicht ausreichend, da können wir noch gar keine Entscheidung treffen. Und dann wird in dieser Runde eine Entscheidung getroffen, wo federführend, in welchem Bereich dieses Thema weiterverfolgt wird.
SPEAKER_03: Genau, das heißt die Verantwortung zu dem Zeitpunkt, die gibt es noch gar nicht. Ja, die wird leider definiert, genau, für das Projekt, genau. Die Verantwortung, die im R&D liegt, ist nur das Einsammeln der Ideen, aber wir können ja genauso in den, also ich als Person, als PM-Person oder du als PD-Person, du hast ja deine eigenen Ideen, aus welchen Inputs auch immer. Es geht ja nur darum, die hier zusammenzutragen und dafür ist das R&D verantwortlich und danach geht es weiter.
SPEAKER_04: Aber ich bin immer noch zur Meinung, dass das nicht der richtige Weg ist, weil ich kann überhaupt nicht mehr vorstellen, dass zum Beispiel jemand von zum Beispiel Vertrieb hat Toll-Idee für die neuen Anwendung. Genau. Oder, ja, Entschuldigung, man macht.
SPEAKER_03: Ja.
SPEAKER_04: Er geht nicht zur Forschung und Entwicklung.
SPEAKER_03: Nein, nein, nein, nein, genau. Er kann das, er kann das, er kann das, grundsätzlich kann er das entweder über diese E-Mail-Adresse melden, die in dem Prozess stand, die DNF, kann jeder dann melden, wo er möchte, oder er spricht mit irgendwen. Und es geht ja nur darum, dass wir das wieder hier sammeln, damit es nicht mehr, damit Raeck als Beispiel, das ist einfach ein gutes Beispiel dafür, nichts wie Raeck, ich mag, dass er so ruhig ist, aber dass er nicht mehr zu Martin gehen kann. Wenn Martin nein sagt, fragt er zwei Tage später mich und ich sage dann vielleicht ja, sondern wenn Martin das gehört hat, ist es spätestens alle zwei Wochen halt in diesem Themenpool mit einer Bewertung.
SPEAKER_02: Was man hier jetzt, was man hier jetzt geschickterweise machen könnte, wäre, man fügt noch eine Schwimmbahn ein, die heißt GPM-Hack.
SPEAKER_04: Genau.
SPEAKER_02: Und dann sagt man, da in der Runde läuft das Thema auf und es kann auch, also das Beispiel, was du gerade gebracht hast, Martin könnte dann sagen, ja, super, Idee-Reit, nämlich wird in die nächste Runde. Genau. So, und dann wird es in der Runde diskutiert und dann wird es entschieden, ob es weiter verfolgt wird oder nicht oder ob wir noch weitere Informationen brauchen und dann geht es eben aus der Runde zurück in irgendeine von diesen anderen Schritten. Stimmenwahn, wer auch immer den Hut aufbricht und der, der den Hut aufbricht, heißt ja auch nur, dass er es verantwortlich macht. Der macht es ja auch nicht alleine, sondern bindet die anderen Abteilungen so, wie es nötig ist, entsprechend mit ein. Und wenn ich es richtig verstanden habe, aus dem heraus, was Lars mit von der ersten Runde erzählt hat, war dann die Idee, dass dann diese Liste der Themen sozusagen ist. Die sollen in der F&E gepflegt werden, weil da am ehesten die Kapazität da ist. Es geht nur um die Liste. Sag mal, was für ein Thema haben wir? Wie ordne ich das möglicherweise zu? Also, was ist das für ein Themenkomplex? Und wer ist verantwortlich? Punkt. Genau, aber es hat keinen Einfluss.
SPEAKER_03: Also, es gibt halt unterschiedliche Wege, wie du das Thema in der NGTN-Kampf reinbringen kannst. Also es geht nur um die Verantwortlichkeit des Sammelns und was halt die F&E auch machen soll. Bis vor einmal schon mal gucken, wie viele Informationen liegen denn vorab. Schon mal einmal diese Korrekturschleifnisse machen mit, komm, gib mir doch mal ein bisschen mehr Input, ein bisschen mehr Futter zu der Idee. Da hatten wir ein, zwei Sachen, wie viel, dass wir dann, wenn wir hier schon zusammengesetzen mit anderen, wenigstens genug Informationen zu der Idee auch vorliegen haben. Wenn die Person das dann in der Lage ist, zusammenzufragen, kommt ja auch noch ein bisschen drauf an.
SPEAKER_01: Ja, aber es geht da ja auch nicht um Inhaltliches, sondern das kann ich ja, da muss ich überhaupt gar nicht wissen, ob das vernünftig ist oder nicht. Aber wenn einer nur mit einem Einsatzidee kommt, dass das nicht ausreicht, das kann einfach jeder entscheiden.
SPEAKER_02: Da würde ich jetzt auch sagen, das müsste jetzt nicht unbedingt irgendjemand aus der F&E machen, der möglicherweise jeder gar nicht dabei ist, sondern im Prinzip könnte das auch eine Standard-Checkliste sein, wo da einfach einen Haken dran machen, was wir wollen. Habe ich die wissen, die Informationen, oder habe ich sie nicht in der Testen, nochmal nachfragen, genau. Bevor ich das hier in der Rumi-Walks-Thema mache, nochmal, nochmal, die Sache, dass F&E da stand, war einfach nur,
SPEAKER_01: dieser jemand ist in unserem Unternehmen recht schwer zu finden. Also im Prinzip habe ich gesagt, okay, zu dem Zeitpunkt hatte ich ja noch zwei Mitarbeiter, die das machen konnten. Die machen das einfach. Die haben eine ganz kurze Checkliste und fragen noch zwei, drei Sachen ab. Und die treffen auch keine Entscheide, sondern die fragen einfach nur, kannst du mir kurz dazu noch sagen, in welchem Markt du dir das vorstellst? Das ist jetzt auch nur ins Blaue gesprochen, in welchem Markt du dir das vorstellst, was auch immer. Und sammeln das zusammen, schreiben das auf einen virtuellen Zettel und dann haben wir ein bisschen mehr als einfach nur eine Einsatzidee. Mehr nicht. Das kann im Prinzip jedes Sekretariat machen. Nur die reißen sich nicht da drum.
SPEAKER_03: Und genau bei dieser Arbeit hatte ich halt gedanklich, wir hatten ja auch ganz am Anfang daran gesprochen, dass der einzige Grund, warum es am Anfang in der Art und die Swimland stand. Und danach geht es ja in diese Runde mit allen zusammen und dann die Zuordnung, je nachdem, wo das Projekt halt hingehört. Das war der ganze Gedanke dahinter. Ich hoffe, dass wir das halt recht einfach weiterhalten können, weil ich meine, da war es noch ein bisschen anfangen.
SPEAKER_02: Ich bin mir nicht ganz sicher, ob ich dich von richtig verstanden habe. Du hast von der Präsentation gesprochen, die du verteilt hast. Die Präsentation beschreibt, meine ich, wenn ich es richtig in Erinnerung habe, im Wesentlichen, wie ihr Projekte in Business Development bearbeitet. Nein. Okay, dann habe ich falsch in Erinnerung.
SPEAKER_04: Die Präsentation beschreibt, wie wir das von Geschäftsverlust bestanden haben, wie die Investition 2030 umsetzen sollen und Vorschlag, wie wir das umsetzen sollen im Unternehmen. Deswegen gibt es diese unterschiedlichen Kategorien, also bestehende Märkte, wo Betriebsverlust ist, neue Märkte, diese Developing, neue Anwendungen, Produktmanagement und dann strategische Innovation ist dann konstrund entwickelt.
SPEAKER_03: Das ist bei uns dann sozusagen der Punkt, wo wir hier in der Runde entscheiden, okay, ob der Plätze sind wir auch, was ist ein Produkt oder ein Projekt, die Idee, die in einer der Runde geht, dann kriegt die Abteilung und das sind sozusagen wieder Kriterienkatalogen und danach dann entscheidest du und wer kriegt die gut auf und gut.
SPEAKER_01: Okay. So von meinem Verständnis. Das Ganze ist übrigens auch Schattenkämpfen, weil die Anzahl an Ideen, die irgendwie aus Out of the Blue kommen, die wird sehr, sehr gering sein. Die allermeisten Projekte werden aus unserer Runde mehr oder minder kommen und dann ist dieser erste Teil sowieso eher irrelevant.
SPEAKER_00: Also ich habe mich jetzt ganz mal ein bisschen zurückgelegt. Ich habe mich jetzt ganz mal ein bisschen auch angeregt und versucht auch so ein bisschen für mich das zu interpretieren, warum gibt es diese Runde überhaupt und was war die ursprüngliche Intention, dass wir hier zusammensitzen und so weiter. Das, was mir komplett in dieser Runde auf sich fehlt, ist im Grunde genommen, was wir immer vergessen, ist die Geschichte, also warum wir überhaupt hier zusammengekommen sind. Sondern der eine Intention war ja, es entstehen hier irgendwo Innovationen, Ideen oder Sonstiges, die dann halt, ich sage mal so, in den einzelnen Abteilungen oder bei den einzelnen Personen weiterverfolgt werden, weiterbearbeitet werden und irgendwann stehe ich das mal bei irgendeinem auf dem Tisch und derjenige muss das weiter bearbeiten. Und das war ja unsere Intention, das so ein bisschen mal zu kanalisieren, mal zu bündeln, wo wir gesagt haben, okay, in dieser Runde treffen wir uns und besprechen dann halt irgendwelche Sachen, die wir links und rechts irgendwie aufgefasst haben oder die wir halt mitbekommen haben oder wo wir involviert sind, schmeißen wir alles mal in den Topf rein, um so ein bisschen Transparenz zu sehen. Das war die erste Sache. Die zweite Sache ist, das was viel läuft in dieser ganzen Diskussion und wenn ich auch nur da drauf gewesen bin, das kennt kein Pro, ich bin ja auch mit dabei, vermisse ist, wir vermissen die Geschäftswertung und vor allen Dingen bei diesen ganzen Sachen, die wir dort überdenken, wir reden hier ja nicht nur über Produktideen oder Innovationen, sondern wir reden hier auch über neue Märkte und so weiter und da fehlt von meiner Seite aus am Anfang überhaupt der gesamte Strategie-Check davor. Weil wir sind hier schon sehr, sehr tief mit drin von wegen, ja, wie kriegen wir die Produktinnovation hier durch, aber wir wissen ja noch gar nicht, ob das in unsere Strategie hineinpasst und wir wissen ja auch alle und das war ja auch mit einer Intention, weswegen wir hier zusammensitzen, dass zum Teil dann, ich sag mal, da oben, von mir aus gesehen, da oben Entscheidungen getroffen werden, die wir dann irgendwann im ersten Halbesjahr später erfahren, weil dann irgendwelche neue Mitarbeiter da stehen oder neue Ideen da sind oder neue Märkte da sind und so weiter. Und dieser Strategie-Check, der muss vorher auch aberfolgen und der muss auch mit der Geschäftswertung erfolgen, weil es macht doch keinen Sinn, dass wir hier tief einsteigen mit Forschung und Entwicklung oder sonst wo, wenn die Geschäftsleitung eine ganz, ganz andere Idee hat oder vielleicht schon irgendwo schon aktiv ist oder halt auch andere Menschen oder Abteilungen davon aktiv sind. Das heißt also, das dürfen wir nicht außen vorlassen. Und ansonsten müssten wir an und für sich diesen Prozess als solches gar nicht so aufsetzen mit, ich verstehe auch nicht, Martin, Entschuldigung, wenn ich etwas falsch sage, aber ich zum Beispiel, die Forschung und Entwicklung ist da irgendwie am Anfang mit drin. Nein, Forschung und Entwicklung ist nur ein Detailbereich. Irgendwo mal da hinten irgendwie sich im Prozess daran teilnimmt. Da sind erstmal ganz, ganz andere Abteilungen oder Wissen halt auch gefragt und das muss in diesen Prozess einfließen. Und ich glaube, das dürfen wir nicht außer Acht lassen. Ich glaube, wir müssen dann vielleicht auch mal in diesem Falle auch mal vielleicht auch unsere Brille absetzen, um halt auch besser zu werden, um halt unsere internen Prozesse, die wir vielleicht über Jahre sich aufgebaut haben, mal abzulösen, um da halt auch mal einen neuen Weg zu gehen, einen professionelleren, strategischen Weg, als das, was momentan halt stattfindet. Und deswegen finde ich so ein bisschen so die Entwicklung, in den letzten Meetings war ich nicht dabei, da war ich krankheitsbedingt raus, beziehungsweise ohneausbedingt raus, aber ich vermisse ich da halt mal ein bisschen dran. Vielleicht sehe ich das auch falsch.
SPEAKER_01: Ja, das siehst du falsch. Zumindest ein Detail davon. F&E hat in diesem Prozess, genau wie du beschreibst, einen Teil, nämlich alles das, was dann wirklich Entwicklung ist. Wir brauchen nur ein Sekretariat, was am Anfang einmal sozusagen die von außen einkommenden Sachen ganz kurz auf formale Vollständigkeit oder zumindest auf minimale Anforderungen prüft und die einfach eingleist. Und das ist nur rein zufällig F&E. Das kann von mir aus auch Martina Nagel machen. Schlechtes Beispiel, weil die geht in Rente. Aber das kann jeder tun. Nur du wirst im Unternehmen keinen finden, der da Bock drauf hat. Deswegen habe ich gesagt, okay, das machen wir vorne weg. Streich da einfach mal das F&E weg. Ist ein Sekretariat. Irgendjemand muss diese fucking E-Mail-Adresse lesen und gucken, was da drin steht.
SPEAKER_00: Mehr ist es nicht. Nee, ich meine das Gleiche. Wir meinen doch das Gleiche. Das ist nur ein Beispiel. Wie gesagt, wir müssen jemanden haben, der sich dort darum kümmert, der das, ich sage mal so, abteilungsübergreifend halt organisiert, ob es ein Projektmanager später ist oder wie halt auch immer. Aber letztendlich, was ich damit halt nur meine, letztendlich, wir sind immer sehr, sehr ganz schnell tief in irgendwelchen Sachen unterwegs. Und ich glaube, wir müssen erstmal, ich sage mal, global den Prozess, und wenn das erstmal nur vier oder fünf Schritte sind, mit anfangen, wie sieht so ein Prozess bei Now aus, oder wie sollte er aussehen? Und da müssen wir halt die Geschäftsleitung auch die Strategie noch mit reinbringen, weil wir sind ganz, ganz schnell sofort in irgendwelchen Details drinnen. Und das ist ja gerade das, was wir an und für sich auch aufheben wollten mit dieser Runde, dass da nicht schon was losgetreten wird und da hinten ist schon jemand unterwegs, das ist schon 20 Kilometer entfernt, sondern wir wollen ja am Anfang dabei sein und darüber entscheiden und das müssen wir ja halt auch, weil halt auch dann später auch alle Abteilungen ja auch dann involviert sind. Später werden es um neue Märkte geht, ist Marketing involviert, ist dann irgendwie auch dann wieder auf FUE dran involviert, ist dann Produktmanagement mit involviert, ist Vertrieb mit involviert, wir müssen erstmal Sachen überhaupt erst mal am ersten Schritt abklären, bevor wir so tief mal einsteigen. Und das war mein Grundgedanke, wo ich einfach sage, jetzt irgendwie, dass es doch den Prozess mal erstmal so gut bleibt, erstmal viele Hauptschritte mal definieren, wie gesagt, dieser strategische Kurzcheck, der sollte meisterachtens am Vorderstelle irgendwo stehen, weil da vielleicht Sachen dann schon irgendwie abgelöst werden, weil es da vielleicht schon ganz, ganz andere Ideen gibt und wie gesagt, das war ja auch so ein bisschen so die Intention halt dieser Runde halt, um halt alles mal zusammenzuschmeißen und dadurch auch mal die Informationen, die wir alle haben oder wo wir was von erfahren, mal so ein bisschen transparent werden. Das müssen wir natürlich halt auch im Prozess dann halt auch abbilden. Ich hoffe,
SPEAKER_02: ihr könnt mich verstehen. Also gehört haben wir dich auf jeden Fall. Verstehen?
SPEAKER_04: Wenn ich jetzt darf, ich sage es mal ganz klar. Ich versuche schon das viertes Mal was zu sagen. Darf ich jetzt bitte?
SPEAKER_00: Ja, ich lasse mir nur vielleicht, Frau Rosen, einen kurzen Satz nur, nur von der Erklärung her. Worum es mir grundsätzlich geht, ist an und für sich, wir sollten uns überlegen, warum wir jetzt zusammengekommen sind, warum es diese Runde überhaupt gibt. Und also, wenn wir ganz ehrlich sind, haben wir selber gemerkt, dass es einmal irgendwie ständig irgendwelche Bestrebungen gab, irgendwelche Aktionen gab, die von hinten aufgeholt werden, wo wir dann diese Runde zusammenwitzen, ich habe gehört, das und das ist jetzt irgendwie unser neuer Plan. Und das war die Intention dieser Runde. Und so sind wir dann später dann halt über F und E und es gibt ja schon einen Prozess und hin und her irgendwie dahin wie hingestolpert. Nur dann sollten wir doch erstmal alles wieder zurückholen und mal wirklich sagen, okay, jetzt endlich, was wollen wir eigentlich erreichen und dann von dort aus an wieder runterbrechen und nicht jetzt schon wieder in Details reingehen, wo wir sagen, ja, okay, wo haben wir hier einen Ideenpool, wer zieht das damit an, welche Abteilung müssen wir damit reinbrechen, weil es bringt uns nicht weiter. Meine Meinung. So, jetzt bin ich fertig. Oh, Mann. Ja, fertig.
SPEAKER_03: So, wir haben keinen Sprechfall.
SPEAKER_04: Alles ganz gut. Also, eine Frage für Sie. Also, wenn Sie reden, wir müssen das tägliche Teil erstmal festlegen. Wir haben Strategiesitzung im Juni gehabt und wir waren alle dabei. Also, haben Sie als Marketingabteilung vielleicht etwas erarbeitet für den täglichen Teil, dass Sie vorschlagen können, weil Geschäftsentwicklung hat das jetzt vor drei, nee, vier Wochen einmal ein erfülltes Vorschlag gemacht und gezeigt, wie wir denken, dass die Strategie, was Geschäftsführung geklärt hat und gesetzt sein sollte. Also, wir haben das gemacht. Keiner hat uns gefragt. Wir haben das, wir haben uns dessen und einfach nochmal definiert, basierend auf dessen, was wir gehört haben, auf Sprechsitzung. Also, ein Dokument existiert. Dann sind die anderen Abteilungen gefragt, bitte, bringt die Kommentaren. Wenn keiner was sagt, heißt es für mich, okay, ich bin einverstanden. Also, das ist Frage Nummer eins. Die andere Kommentare, was ich jetzt habe, ist das jetzt vielleicht, lasse ich nicht dabei. Aber ich muss das so sagen, er hat Hennig mit Prozessmanagement eingebunden, weil ich glaube, er wollte zwei Sachen gleichzeitig regeln. Eine Sache wäre Strategie und andere, wie können wir das im Prozessmanagement erbinden. Wir haben uns tatsächlich jetzt mehr mit Prozessmanagement beschäftigt, obwohl ich denke, wir müssen erstmal strategische Teil festlegen. Das haben wir noch nicht gemacht. Und klar, wir können jetzt alle kritisieren, was alles fehlt, aber die Frage ist, welche Abteilung hast du am Tisch gebracht?
SPEAKER_03: Ich würde gerne noch einmal dazu was ergänzen. So wie ich das verstanden habe und ich lasse jetzt mal bewusst die Abteilung weg, ich gehe jetzt nochmal, auch wenn ich mich jetzt gedanklich an den Prozess dran langhangeln, bitte nagel ich mich jetzt nicht drauf. Aber die Idee und auch mit dem Kernproblem, das Björn jetzt nochmal angesprochen hat, alle sind informiert, wir wollen jemanden den Hut aufsetzen und so weiter. Das ist für mich der Ausgangspunkt, weswegen wir zusammenkommen. Wir sammeln von überall Ideen. Punkt. An ein Wort. Und diese Ideen werden einmal von einer Person überprüft, ob das mehr als einfach nur ein Satz ist, sondern ob noch ein bisschen mehr Information das angemessert werden kann. Damit, wenn sich eine Runde mit hochbezahlten Leuten damit auseinandersetzt, wenigstens eine gewisse Basis an Informationen vorliegt. Dann, und das wird auch so aufgeführt, und ist auch so aufgeführt in den Testen. Gibt es Kriterien? Welche Art von Projekt ist es? Ist es die Management-Projekt? Oder neue Märkte? Wie auch immer, wie rum man die Bezeichnungen jetzt machen wollte. Und da war auch Dinge drin, wie beispielsweise, jetzt nagelt es nicht genau fest, 5.120.000 Jahre oder was hatte Herr Peter gesagt? So ein Pima-Norm oder sowas. Also auch diese Strategien, strategischen Kriterien, die wir mal für Bank haben. Oder die mal da einen anderen Moment wurden. Diese Kriterien, und das war auch so mit abgebildet, oder es ist auch immer noch, stehen da mit drinne. Die überprüfen wir. Und wenn wir dann hier sagen, okay, es geht so und so weiter, dann geht es an der Person weiter. Ich habe nicht das Gefühl, also, oder ich habe hier das Gefühl, besser so gesagt, als sind wir eigentlich, kommt nicht einer Meinung, schafft uns aber gerade nicht, über das
SPEAKER_02: Gleiche wirklich zu reden.
SPEAKER_03: vielleicht ist es einfach nur ein Gefühl von mir. Natürlich kann man manche Sachen detaillierter aufschreiben, aber ich glaube, grundsätzlich wollen wir doch das Gleiche.
SPEAKER_04: Dass wir alle gleich wollen, das ist keine Frage. Klar wollen wir alles gleich. Wir wollen alle ein gutes Unternehmen. Das geht nicht darum, sondern es geht einfach nur, warte bitte, es ist einfach so, du siehst das von Produktmanagementseite, dieser Prozent, das geschrieben ist, ist von Produktmanagementseite und wir haben jetzt gesagt, vielleicht sollten wir einfach nur die Blickweise ändern, dass wir schauen, von drauf, was Markt will, was im Stakeholder war, wo ist jetzt, wo bewegt sich jetzt alle in der ganzen Welt und dann kommen wir langsam, okay, zu Produkt. Früher war das Unternehmen immer so orientiert, Produkt und dann gucke ich mal, jetzt kann ich hier anwenden, hier gehe ich auf den Markt, wo sind die Zwiebeln auch so organisiert und du warst dann mit mir nicht dabei, wenn wir gesagt haben, Prozessmanagement ist an der einen Seite, wir müssen erstmal die ganzen Unternehmen überzeugen, diese Strategie zu befolgen, weil sonst, dass du jetzt erwartest, dass einer sich meldet auf diese Imeradresse und sagt, die Idee, vielleicht werden die Paidee kommen, aber die beste Idee, die werden normalerweise kommen, einer, der mit mir klarkommt, wird mich anrufen und erzählen oder anderer, ruft dem jetzt Cashman an, macht den Tatzel, der ruft dich an, ich weiß nicht, es wird super in die Idee kommen, nicht wirklich über diese Imeradresse, wie wir das organisieren, ich finde, das ist wirklich es mir eigentlich egal, wenn das in Vorstellungen geht. Mir geht es nur darum, dass die Leute erstmal strategische Zeit verstehen und dass wir dann das Prozessmanagement am Ende brauchen und wir sind noch nicht hier bei strategischer Zeit noch nicht fertig, weil, du sagst, du hast das Prozessmanagement erklärt, aber Prozessmanagement liest keiner. Du kannst jetzt einfach komplizieren, nicht sagen, bitte liest mal Prozessmanagement. ich erinnere mich,
SPEAKER_UNASSIGNED: ich erinnere mich,
SPEAKER_04: dass der Prozessmanagement von Marketing damals geschrieben hat, hat das jemand gelesen? Nein, man meldet sich bei Marketing, was muss ich jetzt machen, wer muss ich jetzt mitmelden? Weil keiner hat Prozessmanagement für die Verifizierung und Umsetzung für die Unternehmen. Aber die Strategie liest, und das mussten wir immer klar definieren. Und deswegen ist diese Präsentation, wo das, was Geschäftsführung, weil die Strategiesitzung mitgeteilt hat, wo wir mitdiskutiert haben, jetzt geht es darum, wir mussten das im Unternehmen erst mal präsentieren, aber wir mussten erst mal einschauen, ist das jetzt richtig? Sind wir einmal, seit wir alle einverstanden, ist das jetzt richtig? Oder wollen wir das anpassen?
SPEAKER_03: Das verstehe ich.
SPEAKER_04: Deswegen haben diese Präsentation in die Gruppe gecheckt, gegeben, aber also an dem Tag, glaube ich, du warst nicht dabei, du warst auch nicht dabei, vielleicht nehmt ihr bitte Zeit und gebt die Kommentare. Oder wenn ihr keine Kommentare habt, dann sind wir alle einverstanden.
SPEAKER_03: Ja, wir müssen das dann. Das ist eigentlich, was ich ja schon sagen soll. Nein, nein, nein, nein, nein.
SPEAKER_04: Das ist traditionell bei diesem Vernehmen in die Vergangenheit geführt von Produktmanagementseite. Und das ist prinzipiell nicht falsch. Aber wofür habe ich dann diese Gewährung als Abteilung, wenn das nicht notwendig ist? Wir schauen anders, in Höhe zu, auf den Markt, das gefragt wird, und dann gucke ich mal, kann ich mit meinen Produkten und um dies zu leisten, was das Unternehmen anbieten kann, ein Business-Paket stocken, das ich anbieten kann.
SPEAKER_03: Wahrscheinlich, der Ansatz von der Ideentrennung und der Entwicklung ist zwar auf einer anderen Ebene aber der gleiche Gedankengang. Also sollte ja auch nutzenorientiert und nicht einfach, wir machen jetzt ein Produkt nix absoluter und dann gucken wir mal, wie wir uns verkauft kriegen. Also grundsätzlich sollte vom PM der Ansatz ja auch anders muss sein. Nutzenorientiert, kundenorientiert, etc. Aber ich hatte nicht das Gefühl, dass das jetzt die Aufgabenstellung ist. Ich hatte das Gefühl, die Aufgabenstellung ist, das Wissen zu zentrieren und auch zusammen an verschiedenen Punkten an den gleichen Dingen zu arbeiten. Das einmal als Aufgabenstellung und Projekten einen klaren Projektbeantwortlichen zuzuordnen. Das war vom Schnittstellen-Meeting, so wie es am Anfang hieß, wo Herr Peter hier mit uns saß, die Anforderungen, die an mich oder an uns, ich sage jetzt mich, weil ich es dann auf der Basis ausgearbeitet habe, das war der Gedankengang dahinter. Deswegen haben die Punkte, wie du sagst, hundertprozentig die Ansatzberechtigung und müssen auf jeden Fall auch aufschreiben, wie wir Projekte abarbeiten und so weiter und so fort. Aber das war überhaupt nicht der Gedankengang beim Ausführen in den Prozessen, sondern das war einfach nur, wie und wo sammeln wir die ganze Information, welche Kriterien gehören dahinter, ist es überhaupt Geotechnik und so weiter und so fort und wie gehen wir dann weiter. Das war alles, was da drinnen steckt und das nach diesen Kriterien zu bewerten, dem ist natürlich nicht gerecht. Aber das ist für mich halt auch einfach an einer anderen Stelle. Das ist für mich an der Projekt, also der, der dann den Hut dafür aufhat, wie er das abarbeitet, dass nichts von diesen Aussagen bildet das ab. Aber trotzdem hast du natürlich zu hundertprozent recht mit, dass auch daran gearbeitet werden muss. Aber das habe ich überhaupt nicht umsichtig.
SPEAKER_04: Ich habe das jetzt ein bisschen anders verstanden. Ich habe das richtig verstanden, dass diese Cremie das verdienen sollte, denn wenn du zum Beispiel Geschäftsentwicklung kommt und sagt, wir wollen auf den Markt, so keine Ahnung,
SPEAKER_03: ich weiß nicht, Max,
SPEAKER_04: achten kann.
SPEAKER_03: und dann kommt das jetzt erst mal hier auf den Tisch in irgendeiner Form. so, ob du das jetzt hier reinbringst, ob das per Mail kommt, das ist ja wurscht. Das kommt hier auf den Tisch mit einer gewissen Anzahl an Informationen und dann sehen wir, das ist ein neuer Markt, das ist ein Business Development Projekt. Wer hat ein Business Development den Hut auf? So, dann, was weiß ich, machst du das, der ist eine Person, dann wissen wir hier, dieses Projekt wird gerade bearbeitet, steht in der Liste, du hast den Hut auf und irgendwann wollen wir wieder davon und wenn wir zufällig im Produktmanagement auch das Thema Afrika aus irgendeinem Grund auf den Tisch kommen und jemand zu mir kommt, weiß ich, wir waren aber sogar schon, ich spreche bitte mit dem anderen. Das war für mich die Intention und das soll es abbilden. Aber du hast ja die volle Projektverantwortung dann in dem Sinne, das ist bei euch, das gehört zu euren Kriterien und so weiter, ihr sollt nur laut dem Prozess und das sind die Sachen, die wir diskutieren wollten, irgendwann mal eine Schleife drehen und das hier vielleicht mal wieder vorstellen, was der aktuelle Stand ist und so weiter und so fort. Aber das ist für mich alles, was da abgebildet werden sollte und das ist das, weswegen ich mich auch frage, warum brauchen wir, Entschuldigung, aber sechs, sieben, acht Wochen, das fühlt sich für mich gerade langsam an, einfach.
SPEAKER_04: Ja.
SPEAKER_03: Ja. Ja, aber... Na, da...
SPEAKER_04: Ja, aber die Frage ist jetzt, auch wenn das entschieden wird, okay, wer ist jetzt der Fraordinator? Eigentlich geht es mehr darum, dass ihr Resorten festgelegt werden, weil, wenn ich zum Beispiel sage, ich arbeite jetzt in Afrika, keine Ahnung, in Ghana, alte Komis-Stiftung, dann muss ich, versetze ich es an, auch vom Marketing, also Unterstützung bekommen, ich muss die Unterstützung von WM bekommen, keine Ahnung.
SPEAKER_03: und dann, genau, genau, deswegen sitzen wir hier dann zusammen, kann darüber reden, aber, ja, so hatte ich, das ist einfach mein Verständnis gewesen, letztendlich. So, und ich habe das ja auch extra so gemacht, wir hatten das so ausdiskutiert, so hatte ich es verstanden, deswegen hatte ich es so vorbereitet, in diesem, ich habe einfach den Spendenprozess genutzt, weil es ja eh aufgeschrieben werden muss, und ich habe es ja auch rumgeschickt, bitte in zwei Wochen, bitte Kommentare zurück, und so weiter und so fort, habe keine Kommentare gekriegt und dementsprechend hatte ich angefangen mit der Umsetzung, weil, wie du gerade eben sagtest, kein Kommentar ist Zustimmung, deswegen hatte ich weitergemacht. Und, ja, also ich sehe gerade nicht, dass wir wirklich weiterkommen, das sind die gleichen Punkte, über die wir gerade diskutiert haben, wie vor drei Wochen.
SPEAKER_02: Ich sehe, ehrlich gesagt, das ganze Problem nicht, außer dass man sich jetzt daran aufhängen kann, dass wir hier vermeintlich nur über Prozessmanagement reden, nee, tun wir gar nicht, wir reden einfach nur darüber, wie der Prozess aussehen soll und den bilden wir dann ab, Punkt. So, und wenn das so ist und wenn wir uns da vom Grunde her erstmal einig sind, dann heißt das eigentlich nichts anderes als, da bin ich dann bei Björn, da bin ich auch bei dir, wir haben am Ende des Tages irgendwo, ich sage mal, vier, fünf Schritte zu beschreiben, das können wir in der Präsentation machen, das können wir in der Word-Datei machen, sinnvollerweise machen wir das im Prozessmanagement, okay, bei dir habe ich jetzt rausgehört, Prozessmanagement interessiert sowieso keinen, guckt ja eh keiner rein, sehe ich ein bisschen anders und selbst wenn es so wäre, wäre es trotzdem falsch, das dann eben auch so zu gutieren, sondern da müsste man im zweiten Fall dagegen arbeiten, aber wir sollten es einfach vernünftig beschreiben, Punkt. und ob ich das dann hinterher präsentiere oder ob ich den Trichter präsentiere oder sonstigen, das ist auch noch wieder eine andere Baustelle. Auf jeden Fall habe ich dann, wenn ich das im Prozess einmal beschrieben habe, immer die Möglichkeit, die Kollegen da auch hinzuweisen. Es wäre auch sowieso geschickt, ich sage mal, die Dinge aus dem Prozessmanagement dann für Präsentationen zu nutzen, das sorgt dann vielleicht dafür, dass die Leute eher da reinschauen, als dass sie sich irgendwelche Präsentationen raussuchen, die sie irgendwann mal gekriegt haben. Also es würde Sinn machen, spielt aber alles keine Rolle. Der Punkt ist, ich habe es so verstanden und so habe ich es von Lars verstanden, als wir über diese Runde gesprochen haben, der will mich nicht dabei haben, weil er es im Prozessmanagement geschrieben haben will, sondern der will mich dabei haben, weil ich muss mich um Prüfungen kümmern, ich weiß, welche Regelwerke in den einzelnen Ländern zu tun sind, ich weiß, was ich bei Zulassungen machen muss und all das sind Themen, die wir zwangsweise bei neuen Produkten, neuen Märkten immer mit auf den Tisch stellen. Deswegen wollte er mich dabei haben. Das Prozessmanagement habe ich einfach nur gesagt, es muss vernünftig geschrieben werden. Punkt. Für mich ist es relativ einfach. Es gibt ein Eingangskästchen, das heißt, GTM Hub oder wie auch immer, können wir auch einen anderen Namen nehmen, ist mir völlig egal. Da steht drin, was ist das hier für eine Runde? Was tun wir hier? Was ist unsere Ausgabe? Und dann gibt es ein Kästchen, möglicherweise ein zweites, da ist dann Ideenpool, das können wir irgendwo anhängen, scheißegal, wir können das in der Runde anhängen, wir können das als nächstes Kästchen machen, wir fragen irgendwelche Dinge nach, dann sind wir bei einer Entscheidungsfindung, das ist das, was Björn gesagt hat, was gehört zur Entscheidungsfindung dazu, wir fragen uns strategisch das, wir fragen das, das, das, das sind die Fragen und auf Grundlage der Antworten, die es zu den Fragen gibt, gibt es eine Entscheidung, wem geben wir die Aufgabe.
SPEAKER_03: Genau.
SPEAKER_02: Ich glaube nicht, dass wir an der Stelle schon sagen können, jawohl, fünf Jahre, fünf Millionen, 30 Prozent Lettungsbeitrag, das kommt möglicherweise ein halbes Jahr später, nachdem irgendjemand angefangen hat, eine Basisarbeit zu machen. Unabhängig davon, verteilen wir aber trotzdem eine Aufgabe und sagen, ey, irgendeiner muss sich ja erstmal damit beschäftigen, weil sonst machen wir jede Aufgabe im Ansatz tot, weil das kann realistisch keiner. Wenn ich irgendwo eine Idee habe, ich sage mal ein Beispiel, was Giovanna gerade sagte, keine Ahnung, ob es realistisch ist, in fünf Jahren mit gezielten Aktionen in Ghana fünf Millionen Umsatz mit 30 Prozent Deckungsbeitrag zu machen. Dazu muss ich erstmal ein bisschen Basisarbeit machen. Aber das ist ja das Beispiel. Danach geht es weiter. Und wer die Fachabteilung ist, die sich dann danach kümmert oder den Hut auf hat, das ergibt sich aus den Themen. Klar, wenn es neue Märkte sind, ist es Business Development. Wenn es irgendwelche ganz anderen Themen sind, die bisher mit klassischen Geocoolstocken gar nichts zu tun haben. Ist es möglicherweise auch Business Development oder ist es der PM oder, oder, oder. Das wird man themenabhängig definieren müssen.
SPEAKER_04: Okay, aber wenn du dich erinnerst, was ich hier präsentiert habe, dann haben wir damals sieben Stufen einmal Stages vorgestellt. Dann haben wir gesagt, ich habe das jetzt noch jetzt vom Telefon, kann ich noch mal lesen. Erstmal war diese Opportunity Identification, das wäre zum Beispiel
SPEAKER_01: die Liste.
SPEAKER_04: gelistet, dann kommt die Nischung Screening und die Nischung Screening heißt, ich muss tatsächlich die ersten Daten zusammenfassen, um einfach zu entscheiden. Wir machen weiter, wir pausieren oder wir sagen, nee, geht nicht weiter.
SPEAKER_02: Aber wer ist ich an der Stelle?
SPEAKER_04: Ja, genau. Das ist ja die Frage. Weil in dem Moment muss ich dann entscheiden, in welcher Abteilung geht die Idee.
SPEAKER_02: Und das ist das, was Björn vorhin sagte. Passt das in unsere Strategie? Möglicherweise kann ich das an dem Zeitpunkt noch gar nicht entscheiden, wenn einer die Idee abgegeben hat, sondern erst nach dem zweiten, dritten Step, weil ich erst mal ein bisschen Basic-Arbeit machen muss. Aber all das können wir relativ einfach in Stichpunkten darin beschreiben. Dann weiß jeder, was ist die Aufgabe. Und nochmal, wir können das in FOD reinhängen, wir können das in Produktmanagement reinhängen, wir können das in den Business Development Part reinhängen, wir können das an vielen Stellen können wir das reinhängen. Dann hat das jeder immer vor Augen, kann mit einem Klick da drauf das aufmachen und sieht sofort, ah, was passiert da und wer macht es.
SPEAKER_03: Und das war's. Und wir werden den Ablauf des Projektes auch nicht im Prozessmanagement definieren können. Dafür ist das viel zu unterschiedlich, hat viel zu unterschiedliche Anforderungen, je nachdem, gar nicht.
SPEAKER_02: Nein, aber wir können definieren, wir können definieren bis zu dem Zeitpunkt, wo wir sagen, wir verteilen die Aufgabe und möglicherweise können wir auch definieren, wir haben Prüfpunkte, wo wir ein Feedback in dieser Runde möglicherweise erwarten, das muss man definieren und sagen, jetzt kriegen wir mal was zurückgespielt und jetzt entscheiden wir, geht das Projekt weiter oder geht es nicht weiter oder wir sagen auch, wir sind uns unsicher, müssen wir vielleicht in Geschäftsführung anfangen, wie sie es sieht. Und auch das ist da drinnen schon zu tun. Aber auch um das mal ganz klar zu sagen, ich sitze hier nicht als einer, der sagt, wir müssen den Prozess so oder so oder so leben, beschreiben oder so, beschreiben schon, leben. Das ist nicht meine Aufgabe als Prozessverantwortlicher oder als der, der für das Prozessmanagement verantwortlich ist. Die Aufgabe als Prozessmanagement verantwortlicher ist, zu beschreiben, wie wir arbeiten wollen. Punkt.
SPEAKER_04: Aber wir haben uns noch nicht abgestimmt,
SPEAKER_02: wie wir arbeiten wollen.
SPEAKER_04: Deswegen komme ich einen Schritt vor.
SPEAKER_02: Ja, aber möglicherweise ist auch das einfacher, wenn man mal ein Gruppgerüst stehen hat und sagt, okay, das ist mal der Große und jetzt machen wir hier noch was.
SPEAKER_04: Aber das habe ich präsentiert. Ich habe das präsentiert bei der letzten Termin, wo wir alle zusammen gesammelt haben. Normaler war nicht dabei und der falscher war natürlich dabei.
SPEAKER_02: Das habe ich einmal präsentiert. Habe ich jetzt ehrlich nicht mehr vor Augen, wenn du diesen Funnel präsentiert hast, dann ist das sozusagen wie das Projekt läuft?
SPEAKER_04: Nein, nein, nicht das wie das Projekt. Von der Idee bis die Umsetzung. Welche Schritte müssen wir machen? Weil auch wenn die Idee gut ist, einer muss sich kümmern, wie komme ich mit Ideen auf den Markt? Das habe ich überhaupt nicht definiert. Wie komme ich von der tolle Idee, bis der Betrieb etwas verkaufen kann? Wir haben diesen strategischen Teil und der Prozess definiert noch nicht.
SPEAKER_02: Mag sein.
SPEAKER_04: Nein, haben wir nicht.
SPEAKER_02: Ja. Ich kann es momentan nicht sagen, weil ich habe es jetzt leider nicht vor Augen. Aber die Kunde ist wiederholen.
SPEAKER_04: Ich kann doch mal...
SPEAKER_02: Andersrum, andersrum. Also, vielleicht vom Grunde her sind wir uns einig, was wir wollen, glaube ich. Also, vom Grunde her sind wir uns einig, dass wir eine Runde haben wollen, wo wir sagen, da kommen die Ideen an. Wir reden über die Ideen, sagen, jawohl, an welcher Stufe sind wir womöglich, müssen wir noch Aufgaben verteilen, brauchen wir noch viel Deck oder können wir es komplett verlagern, was auch immer. Und wir entscheiden, jawohl, da und dann geht es weiter. Das ist ja erst mal unsere Kernaufgabe. Da habe ich es zumindest verstanden. Ja, unsere Kernaufgabe ist es möglicherweise, das müssten wir dann definieren, zu sagen, jawohl, diese einzelnen Stufen, die wollen wir sehen oder wir sagen, nee, das liegt in der Verantwortung des Projektverantwortlichen, das können wir definieren. Also ich sage mal, ich habe es bisher nicht so verstanden, dass wir nicht eine Runde sind, wo jetzt einzelne Projektverantwortlichen quasi regelmäßig reporten sollen, wo sie stehen. Nein, das gibt es nicht. Also die Frage ist, wer kümmert sich darum, dass das Projekt dann diese einzelnen Schritte auch beinern wird. Auch das kann man definieren. Wir können im Prozessmanagement, könnten wir so sagen, wenn wir das so wollen, jawohl, das ist ein Projekt, das ist YZ, der kümmert sich darum, hier hast du einen Fahrplan, das sind die einzelnen Steps, die musst du abarbeiten. Irgendwann wollen wir das mal sehen. Vielleicht wollen wir das erst sehen ganz am Ende, aber wir wollen sehen, dass diese einzelnen Schritte eingehalten wurden.
SPEAKER_03: Genau, wir haben halt einfach nur einen einzelnen Loop eingebaut zu den Zeiten, dass einfach, falls es Punkte gibt, die erst, wie du gesagt hast, detaillierter ausgearbeitet werden müssen, beispielsweise, dass es dann halt wieder zurückgeht. Da hat man einen Loop drin, beispielsweise. Aber das ist halt für mich jetzt auch schon Detail, zumindest bei dem Part, weil das ist für mich einfach zwei unterschiedliche Punkte sind. Das eine ist der Ablauf, wie führen wir das jetzt zusammen und durch und das andere ist für mich dann nochmal der detaillierte Projektfahrplan. Ja.
SPEAKER_04: Also ich würde euch alle bitten, also nochmal die Präsentation zu lesen. Ich kann die abgelehnt und die Präsentation schicken und dann, dass wir uns nochmal treffen, das entwürdig zu entscheiden, wer das genau ist.
SPEAKER_02: Die Frage ist, wollen wir in Zukunft jedes Thema, ich sag mal so, hoch hängen? Fragezeichen? Spaß einfach mal. Weil das ist natürlich ein extrem hoher Anspruch oder ist das dann etwa, also ich denke jetzt auch mal an das Thema fünf Millionen Umsatz, wann kann ich das haben, wie viel bringt, ne? Also ich,
SPEAKER_03: wenn man sich überlegen, es wird ja auch Ideen geben,
SPEAKER_UNASSIGNED: wie,
SPEAKER_03: und da gehe ich jetzt mal gerade gedanklich auf das Produkt, drei Prozent Herstellkosten bei Venture Fix durch XY und das hat ja diesen Anspruch dann gar nicht an Detail gemacht und deswegen habe ich es nicht in diesem Detail gerade, in diesen Projekten bei dem Projekt mit aufgeführt, weil man es einfach nicht verallgemeinern kann. Ja, das ist ein gutes Beispiel. Genau. Natürlich bin ich zu 2000 Prozent auf deiner Seite, dass wir anfangen müssen, diesen Detailgrad zu überführen in der Art Projekt mit Dünn kommt fast. Gerade Business Development und auch vermehrt in PR-Anwendungen. Da muss der Gegenstände auch noch...
SPEAKER_04: Und die Gegend sind. Weil es kann sein, dass der Gewerleiter sagt, okay, es ist ein toller Markt, tolle Anwendung und dann, eigentlich die nächste Gegend zum Produktmanager, dann ist das technisch unnötig möglich dann.
SPEAKER_03: Abbruch. Genau, klare Abbruch. Dann geht es nicht da.
SPEAKER_04: Und jetzt ist es so, jeder macht, was er will, oder sie will. Und es ist nicht koordiniert. Und es kann nur koordiniert, wenn alles 100 Prozent funktioniert. Deswegen, ich bitte euch, also nochmal, Präsentation zu lesen. Und ich habe auch verstanden, was jetzt Henning gesagt hat und das wird am Ende von uns dann entscheiden, wie kommen wir in der Freizeit. Weil jetzt diskutieren wir, finde ich, wie jeder das hat verstanden hat. Jeder hat es ein bisschen anders verstanden, das ist auch normal. Aber, letztendlich, wie gesagt, für mich ist klar, wir haben uns vor die Schöpfige-Sitzung getroffen, da war die Schöpfige-Sitzung, da war klar, was zu tun ist. Und jetzt sind wir zwei Momente danach und wir haben uns gesagt, dass wir haben uns nicht bewegt.
SPEAKER_03: Können wir uns, das ist jetzt so ein Vorschlag von meiner Seite, vielleicht darauf einigen, dass wir das gedanklich trennen, dass wir diesen Ablauf, wie wir uns hier zusammensetzen mit den Prozessen und den ersten Kriterien und davon den Gedanken trennen, wie oder welche Kriterien sind für welche Art von Projekten wichtig. Können wir es so nennen, dass wir das von der anderen trennen, dass wir das eine vielleicht endlich zum Abschluss bringen können und dann aber sagen,
SPEAKER_04: in der Präsentation gibt es überhaupt nicht Kriterien. Das sind keine Kriterien genannt.
SPEAKER_03: Wie die Gates, die du gerade angesprochen hast, das ist doch eine Art Kriterien.
SPEAKER_04: Nein, es geht um die Schritte. Vom Anfang, vom Ideen, bis zum Ende, bis zur Umsetzung. Welche Schritte muss ich durchführen?
SPEAKER_03: Ja, das habe ich mir angeguckt, das reicht nicht großartig, also das hat an manchen Punkten einen höheren Detailgrad, aber der Ablauf ist ja nicht großartig anders als das, was wir jetzt auch schon abgebildet haben. Das ist vom Grundgedanken, sehr ähnlich meiner Meinung nach.
SPEAKER_04: Ich glaube nicht, weil ich habe mich mit der Abwägung beschäftigt.
SPEAKER_03: Okay. Ich habe das in Details
SPEAKER_04: viermal gelesen.
SPEAKER_03: Ich habe mir das auch angeguckt, ich habe es so wahrgenommen, ja, dass es vom Grunde her bringt gleich Gedanken.
SPEAKER_02: Also eigentlich ist das relativ einfach und ich sehe auch nicht, also ich bin nicht der Meinung, dass wir uns nicht bewegt haben, im Gegenteil, weil so doof diese Diskussionen manchmal sind, sie sorgen ja dafür, dass man im Gegensatz ein besseres Verständnis hat und das ist ja genau der Fakt, noch vor zwei Monaten haben wir anscheinend alle einander vorbeigeredet und das kommt jetzt mehr und mehr auf den Tisch und jetzt sehen wir langsam, wo die unterschiedlichen Interpretationen dessen sind, was auf den Tisch lag und jetzt müssen wir versuchen, das in eine Richtung zu bringen, damit wir dann wirklich am Ende des Tages auch, ich sage mal, in dieser Runde, aber noch viel näher da draußen, das gleiche kommunizieren denn es bringt ja gar nichts, wenn Giovanna Christberg irgendwas erzählt und zwei Tage später redet er mit dir und du erzählst das Gegenteil. Also das wäre einmal das Allerschlechtste. Dann müssen wir nur die Anforderungen und wenn das dann mal drei Monate dauert, dann dauert es drei Monate. Aber wenn wir dann anschließend mit einer Meinung rausgehen können und die auch gemeinsam mit dem gleichen Wording vertreten nach draußen, dann haben wir richtig was erreicht.
SPEAKER_03: Ja, dann sollten wir meiner Meinung nach wenigstens die Anforderungen nochmal um ein oder zwei Sätze ergänzen, weil für mich kommt das aus dieser ersten Anforderung mit Liste für alle und Projektverantwortlichen ist der Part für mich gedacht, ich habe einfach nie Teil der Institution gewesen. Das ist halt einfach der Punkt. Deswegen, wenn ich ja trotzdem zu 100% bei Giovanna auf der Seite, dass das gemacht werden muss, wirklich wichtig ist, dass wir da auch noch
SPEAKER_02: Nachholbedarf haben. Also für mich wäre es gefühlt gerade relativ einfach, sowohl das, was Giovanna will, abzubilden, als auch das, was möglicherweise an der Stelle too much ist, weil ich glaube, es gibt eine ganze Menge Projekte, wo dieser Anspruchsgedanke möglicherweise zu viel ist oder ich würde damit jedes Thema gleich im Ansatz killen. Auch das kann man glaube ich abbilden, indem wir einfach sagen, in der Runde wird entschieden, wie das A, also A, wer verfolgt wird weiter und B, auf welchem Niveau. Auch das kann man definieren. Und ich sage mal, gerade wenn es um neue Merkel geht, ist es vielleicht relativ gut zu definieren, wenn es um neue Produkte geht auch, wenn es um andere Themen geht, ist es vielleicht nicht mehr so fast zu definieren. Nehmen wir das Beispiel Zulassung, wir diskutieren gerade auch eine Zulassung in einem Land machen, da kann Stand heute keiner realistisch sagen, dass wir damit in 5 Jahren 5 Millionen Umsatz bei 30% Deckungsbeitrag machen. Vielleicht machen wir 5 Millionen Umsatz, aber sicherlich nicht bei 30% Deckungsbeitrag. Lassen wir dann die Zulassung gleich sein? Also ich schicke euch
SPEAKER_04: jetzt die Präsentation weiter und ja, dann
SPEAKER_03: bist du denn aber mit dem restlichen Ablauf einverstanden, wenn wir nochmal darüber gesprochen haben.
SPEAKER_04: auch mit Prozessmanagement, ich habe ja nichts zu sagen, du hast halt so, das soll ich das sagen. Also ich bin, wenn, ist es okay für mich. Auch wenn die Projekte bleiben mit Forschung und Entwicklung, also das ist auch okay. Ich habe nichts dagegen. Für mich ist es wichtig, dass die Wissen-Gewerberkunft integriert ist, weil Wissen-Gewerberkunft irgendwie wo ist und ich habe es okay. Alexandras, was soll ich sagen? Ich stecke mir gar nicht im Detail, dass wir das abgebildet sein muss.
SPEAKER_02: Und es gibt auch kein Muss, es gibt, ja, es gibt schon ein paar Grundanforderungen, aber ich sage mal, wie wir Prozessmanagement abbilden, wie wir Mission 2030 abbilden, wir haben da letztes Jahr mehrfach darüber diskutiert, bis wir zu dem Gedanken gekommen sind, wie wir es jetzt tendenziell machen wollen. Ich kann es auch gerne mal zeigen, wenn ihr es sehen wollt, aber wir sind da wirklich noch am Anfang. Das ist der Gedanke, das ist der Gedanke.
SPEAKER_00: Das ist der Gedanke.
SPEAKER_03: Das ist der Gedanke. Das ist der Gedanke. Das ist der Gedanke. Das ist der Gedanke.
SPEAKER_UNASSIGNED: Das ist der Gedanke. Das ist der Gedanke. Das ist der Gedanke. Das ist der Gedanke.
SPEAKER_02: Danke.
SPEAKER_UNASSIGNED: Danke.
SPEAKER_02: Das ist der Gedanke.
SPEAKER_UNASSIGNED: Das ist der Gedanke. das ist der Gedanke.
SPEAKER_02: Das ist der Gedanke. Das untere ist der Gedanke.
SPEAKER_UNASSIGNED: Das ist der Gedanke. Das ist der Gedanke.
SPEAKER_02: Das untere ist das, wie wir es heute in der Präsentationen im Wesentlichen immer haben, wo man sieht, wer hat mehr oder weniger den Lied und wer sind die beteiligten Abteilungen. Und wir haben das jetzt gedankelich mal so geteilt, dass wir gesagt haben, es gibt eine Phase, bevor der Tender raus ist und es gibt eine Phase, nachdem der Tender raus ist. Und wir sind jetzt angefangen zu gucken, wie wir das im Detail beschreiben können, indem wir gesagt haben, okay, da ist jetzt ein Tender raus. Also das ist hinter den anderen steckt noch nichts dahinter. Das ist ein bestehender Prozess, so wie wir ihn heute haben, nämlich der Prozess, der Auftrag ist da, der Vertrieb hat die Auftrags-Order-Checkliste ausgestellt und jetzt geht es weiter mit Versand und so weiter. An dem Prozess ändert sich erst mal nichts. Insofern haben wir den hier einfach stumpf reinkopiert. So, und dann haben wir angefangen und haben uns den Tender-Prozess angeguckt und das ist jetzt nur ein Entwurf. Den haben wir noch nicht diskutiert. Wir haben das vom Grunde her diskutiert, aber noch nicht vollständig ausdiskutiert und da wird es auch sicherlich noch Änderungen geben, wo wir einfach erst mal gesagt haben, wer ist denn überhaupt in dieser Phase jetzt noch in irgendeiner Form beteiligt. Hauptsachen laufen in Commercial Sales, dann gibt es Technical Pre-Sales, dann gibt es EDD, den PM, den Einkauf, Supply Chain Management, Logistik, Qualitätssicherung und FA war nicht Financial. Ich muss überlegen, was ist der Begriff FA noch mal. Dahinten müsste es dann irgendwann stehen. Finanzbuchhaltung. Also wenn es darum geht, ist jemand kreditwürdig und so weiter. Das heißt, hier detaillieren wir jetzt, beschreiben wir, wie läuft es eigentlich, hier für Kunde schickt ein Tender, was passiert jetzt alles bis zu dem Zeitpunkt, wo wir den Aufrag haben. Und wie gesagt, der basiert jetzt im Wesentlichen an dem alten Prozess, aber eben nicht komplett. Und in einer ähnlichen Form müssen wir das für andere Prozesse auch noch beschreiben. Und hier ist dann zum Beispiel auch ein Unterprozess drin, Kundenzufriedenheit, hier ist ein Unterprozess von Rücknahme von Waren, den wir woanders beschrieben haben, den wir hier jetzt auch noch mal beschreiben. Und genauso können wir es mit dem GTM Hub machen. Wir können hier zum Beispiel einfach reinkopieren, den GTM hat. Und wenn irgendeiner in Commercial Sales eine Idee hat, dann weiß er, ach, so läuft es. An wen wende ich mich jetzt? Ich habe gestern noch mit Malte gesprochen, ich schicke es mal an Malte. Oder ich schicke es an Giovanna oder ich schicke es an Björn oder wen auch immer. Völlig egal, weil von da weiß ich, alles klar, so läuft es weiter. Das ist das, so wie es im Prozessmanagement beschrieben ist. Wir haben hier zum Beispiel Antrag für Sonderartikel, den Prozess gibt es schon. Den habe ich hier einfach reinkopiert, da kannst du hinterher im Prozessmanagement, machst du einen Doppelklick drauf, macht er den auf, Sonderartikel. Und genau so werden wir versuchen, perspektivisch die anderen Prozesse zu beschreiben. Gehe ich jetzt hier nochmal wieder zurück. Business Development, keine Ahnung, da werden wir uns zusammensetzen müssen mit den Leuten, die es betrifft und sagen müssen, wie läuft es. Und zwar müssen wir es so beschreiben, bis wir auf der einen Seite was nachvollziehbar haben und auf der anderen Seite uns aber keine unnötigen Fesseln anlegen. weil es läuft natürlich nicht jedes Mal genau so. Und genauso ist es bei Design Phase until Tender Preparation. Das ist ein neuer Prozess, den haben wir so in der Form. Den werden wir beschreiben, da werden wir mit dem Vertrieb zusammensetzen müssen und sagen müssen, wie soll es denn jetzt eigentlich laufen.
SPEAKER_04: Du hast ja eine Frage, also manchmal ist es tatsächlich so, du hast keinen Geotechnical Engineer, sondern hast nur Specification Engineer. Das ist, was ich meine, du hast da auch keine, das etwas mit Geotechnik zu tun hat. Und der Specification Engineer ist dafür für alles verantwortlich. Achso, ich rede jetzt von diesem Design.
SPEAKER_02: Du meinst das hier.
SPEAKER_04: Ja, genau.
SPEAKER_02: Ja, das ist, also... Das ist, was ich meine,
SPEAKER_04: manchmal fehlt ein Schritt. Oder entfällt ein Schritt.
SPEAKER_02: Ja, aber das ist, das ist an der... Also, ich sag mal, das müssen wir einfach versuchen zu beschreiben. Wir müssen uns, wenn wir sagen, wir haben keinen Geotechnical Engineer, dann müssen wir es halt so beschreiben. Also wir können ja auch hier nicht, wir können ja auch hier schreiben, das ist ein Ingenieurbüro. Neutral. Das hier ist ja, das kommt aus den bisherigen Präsentationen. Das müssen wir jetzt nicht alles eins zu eins übernehmen. Aber darum geht es. Warum das? Weil wir haben nun mal einfach, leider die Situation, dass wir nächstes Jahr wieder ein Rezentifizierungsaudit haben. Und die wissen, die wissen, wir wollen SAP einführen. Die gehen davon, also Stand heute gehen die davon aus, SAP läuft dann schon. Okay, werden sie überrascht sein. Vielleicht auch nicht, dass es dann doch noch nicht läuft. Die haben schon mal was von Salesforce gehört. Das kennen die. Die wissen auch von Alex aus dem letzten Audit, dass wir Salesforce eingeführt haben. Das heißt, die werden solche Sachen hinterfragen. Die haben auch schon bis zum 2030. Also wir müssen es einfach beschreiben. Und hier geht es mir, am Ende geht es mir darum, ich möchte, dass wir die Prozesse nach Möglichkeit so beschreiben, dass wir zu 90% sagen können, jawohl, so ist es, wie wir es haben wollen. Dass es immer mal Ausnahmefälle gibt, wo es anders läuft, aus guten Gründen. Das ist völlig legitim. Nur sollte es nicht genau andersrum sein, dass 10% so laufen, wie es beschrieben ist und 90% laufen anders. Weil dann ist unsere Beschreibung falsch. Wir wollen das beschreiben, wie wir wirklich arbeiten. Und ich muss den Quercheck machen, ob das die Mindestanforderung der ISO 9000 erfüllt. Aber wenn das der Fall ist, haben wir alle Freiheit. Und wenn du sagst, ich, ich, Giovanna, will, dass der Business-Ziel Welt auf meinem Prozess so läuft und genau so und nicht anders, dann beschreiben wir das so und dann musst du halt deinen Leuten das einfach nur entsprechend vermitteln, dass du es genau so haben willst. Sofern es denn die Basics passiert. So läuft es bei allen anderen Prozessen auch. Das ist an der Stelle jetzt auch wirklich Standard.
SPEAKER_04: Deswegen habe ich gesagt, ich kann nichts dafür sagen, weil ich war hier nicht involviert.
SPEAKER_02: Nein, nicht.
SPEAKER_04: Das ist für mich somit okay.
SPEAKER_02: Nein, du bist hier, momentan bist du hier ganz bewusst noch nicht involviert.
SPEAKER_04: Nein.
SPEAKER_02: Weil wir das Thema Business Development ja noch gar nicht bearbeiten. Also es macht auch wenig Sinn, ich sage mal, als Beispiel, wenn wir jetzt Tender to Order besprechen, dass wir da fünf Leute von Pre-Sales dabei haben und so weiter und so viel.
SPEAKER_04: Das ist jetzt auch die Frage, weil das ist die Construction Tender. Aber du hast vor Construction Tender, du hast auch Tender für die, also Final Design, Tender für die Pre-Feasibility Studio, Tender für die Feasibility Studio.
SPEAKER_02: Ja, das ist aber dann mehr Business Development.
SPEAKER_04: Ja, ich wollte nur sagen, so viele Details geben wir hier nicht.
SPEAKER_03: Wissen wir vielleicht auch nicht. Macht man ja auch bewusst ganz häufig nicht, man versucht ja wirklich nur die, den Haupt, oder das häufigste.
SPEAKER_04: Ja, aber deswegen, ja, aber deswegen ist hier, würde ich vereinfachen, und nicht für Technologien Spezification Engineer, sondern Engineer schreiben. Aber manchmal hast du kein Technologien.
SPEAKER_02: Johanna, überhaupt kein Problem. Der untere Part, der ist für mich entscheidend. Den habe ich ganz bewusst dort reinkopiert, der ist auch nicht verlinkt, gar nichts, weil das ist das, was unsere Kollegen aus den verschiedenen Präsentationen kennen. Und ich will einfach nur sicherstellen, dass die sich wiederfinden.
SPEAKER_04: Ja.
SPEAKER_02: dass die wissen, ah, alles klar, ich bin da oben, weil das ist das, was ich kenne, ah, ich muss da reingucken, ich muss da reingucken, darum geht es. Irgendwann werden wir das untere Netz schmeißen, weil das hat kein, außer dem Wiedererkennungswert, hat es keinen Mehrwert. Da drin, ob wir einen Prozess mit drei Schwimmbahnen haben oder mit 20 Schwimmbahnen, ob der gigantisch ist oder klein ist, da haben wir keine Vorgaben. Wir müssen es so beschreiben, wie wir es für richtig halten. Und dabei sollte man immer vor Augen haben, im Idealfall ist das ja auch etwas, was den Kollegen durchaus mal hier und da eine Hilfestemmung sagen kann. Also schon einen vernünftigen Detaillierungsrat, ohne mir auch unnötige Fester anzunehmen.
SPEAKER_03: Ich habe jetzt einmal aufgeschrieben. Wir gucken nochmal über den Business Development Vorschlag. Ich würde auch nochmal mit reiten schreiben, dass das auch nochmal eine andere Zielsetzung ist. Also dass wir diese drei Punkte jetzt, also der Projektablauf, Ideen zusammenführen und das gut aufsetzen, dass das drei Anforderungen sind, damit wir uns da alle vom Grunde her einig sind, was wir damit erreichen wollen und dass wir den Part von Vibana, wie die Projekte abläufen, halt nochmal jetzt neu bewerten, sozusagen, und dass wir den ersten Teil erst mal, wie ist grob der Ablauf, dass wir den so lassen, dass wir uns da eben einig sind, dass wir aber nochmal das beim Projekt selbst, wie ist der Projektablauf, welche Kriterien möglich sind, dass wir da jetzt nochmal detailliert da reinschauen. Seid ihr damit einversprochen?
SPEAKER_04: Ja gut, aber Projekt, je nachdem, wie du das Projekt definierst, weil ein Projekt könnte, also, ja, das Projekt kann alle sein,
SPEAKER_03: und das ist ja die Problematik beim Dachstellen tatsächlich. Das ist ja auch das,
SPEAKER_02: was Henning meinte mit Wir werden nicht... Vielleicht können wir, vielleicht können wir dem, was du ausgearbeitet hast, einen prägnanten Namen geben und einfach sagen, es gibt bestimmte Themen, die arbeiten wir immer so ab und dann ist das genau das und es gibt andere Themen, die könnten, da mag es Sinn machen, die so abzuarbeiten oder es mag auch Gründe geben, warum es keinen Sinn macht, die so abzuarbeiten, dann werden sie eben anders... Also sprich, dass wir immer sagen, wenn wir das so abgearbeitet haben wollen, wie dann der auch immer der Name ist, dann müssen wir das eben auch klar definieren. Aber, dass wir im Umkehrschluss, weil, das Spielchen Ken Martin ja, von F&E, wenn dann der Auditor da sitzt und wir sagen, jedes Projekt wird so abgearbeitet, dann ist das das, was wir selber sagen, dann geht der Auditor da hin und sagt, können Sie mir mal drei Beispiele zeigen? Und wenn du dann zwei Beispiele hast, wie super sind und ein Beispiel, was scheiße ist, dann haben wir eine Abweichung.
SPEAKER_04: Wo wir sehen, heißt es jetzt nicht Projekt, sondern heißt Opportunity und Opportunity muss qualifiziert sein, Projekt zu werden. Aber vielleicht, wenn ich die Präsentation gelesen habe, dann, also einmal vielleicht lesen Okay, danke.
SPEAKER_01: Ganz ehrlich, ich habe das gleiche eben auch
@@ -75,7 +75,7 @@ mentioned_people:
aliases: []
role: null
department: null
attendance_status: "not_present"
attendance_status: "mentioned_only"
notes: null
organization:
@@ -127,4 +127,4 @@ context_rules:
do_not_infer_departments: true
do_not_infer_responsibilities: true
do_not_infer_attendance: true
mentioned_people_are_not_participants: true
mentioned_people_are_not_participants: true
@@ -64,7 +64,7 @@ mentioned_people:
- "Giovana"
role: "Leiterin Business Development"
department_id: "bd"
attendance_status: "not_present"
attendance_status: "mentioned_only"
notes: null
organization:
@@ -0,0 +1,41 @@
# Vanheede–Naue recycled-HDPE meeting — 2026-09-14
Durable private real-world regression and reference case for an English meeting between Naue and Vanheede. The source recording contains non-native English speech; the persisted source metadata does not record speakers' native languages or accents, so those descriptions are retained only as the operator context that motivated preservation.
## Case metadata
- **Case ID:** `vanheede_naue_recycled_hdpe_2026-09-14`
- **Meeting/run:** `Vanheede and Naue, Recycling and Circular Economy` / `d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952`
- **Source meeting date:** 2026-09-14
- **Run completed:** 2026-09-15T14:59:52+02:00
- **Language:** English
- **Duration:** 3,433.92 s (57 m 13.92 s)
- **Participants:** Christian Psiorz, Martin Tazl, Dries Nollet, and Jorik Muylle
## Purpose and scope
This case is an evidentiary reference for English protocol generation from diarized, non-native speech. It preserves two immutable source generations:
| Variant | Source generation | Confirmed mappings in prompt | Role in this case |
| --- | --- | ---: | --- |
| `001` | `protocol/generations/001` | 0 | Tests the limits of name use from participant-list and discourse evidence. |
| `002` | `protocol/generations/002` | 4 | Primary final reference: all speaker identities are explicitly mapped. |
The compact diarized protocol input is byte-identical between the two generations. Their prompts differ because only generation 002 includes the four authoritative mappings. `speaker_identity_analysis.md` deliberately separates direct transcript support from assignments that could result merely from the participant list.
## Contents
- `transcript/source_transcript.txt`: source compact plain transcript.
- `transcript/diarized_transcript.txt`: timestamped diarized transcript.
- `transcript/protocol_input.txt`: exact compact diarized input supplied to both protocol generations.
- `context/meeting_context.yaml`: authoritative participant/context record, including the final explicit mappings.
- `protocol_generation/<NNN>/`: exact prompt, generated protocol, and runtime metadata copied verbatim from immutable source generation `<NNN>`.
- `speaker_identity_analysis.md`: conservative identity-evidence assessment.
- `evaluation.md`: regression expectations and boundaries for evaluation.
- `manifest.json`: source paths and SHA-256 provenance.
No audio, raw model response, container logs, package-install logs, or unrelated run data is copied. The audio reference and its hash are in the manifest.
## Evaluation role
The exact compact diarized protocol input is the evidence source. The mapped generation is a valuable reference protocol, but it is not a gold standard for every wording choice. Operator feedback says later human editing was primarily editorial selection rather than factual correction; that feedback is not a persisted source artifact and must not be converted into unverified factual claims or expected wording.
@@ -0,0 +1,79 @@
schema_version: '1'
meeting:
meeting_id: vanheede-and-naue-recycling-and-circular-economy
title: Vanheede and Naue, Recycling and Circular Economy
language: en
date: '2026-09-14'
objective: ''
notes: Use of recycled HDPE from Post Consumer Sources in Gemomembrane Production,
especially for Naue Carbofol.
participants:
- participant_id: participant-0c5377fb
display_name: Christian Psiorz
aliases: []
role: Sales Direcctor V1
department: naue
attendance_status: present
notes: null
- participant_id: participant-8c7d5a9d
display_name: Martin Tazl
aliases: []
role: Director R&D / Technic
department: naue
attendance_status: present
notes: null
- participant_id: participant-dddd3ddf
display_name: Dries Nollet
aliases: []
role: Sales Person
department: vanheede
attendance_status: present
notes: null
- participant_id: participant-c9280bb9
display_name: Jorik Muylle
aliases: []
role: Technical
department: vanheede
attendance_status: present
notes: null
speaker_mappings:
SPEAKER_00: participant-c9280bb9
SPEAKER_01: participant-8c7d5a9d
SPEAKER_02: participant-0c5377fb
SPEAKER_03: participant-dddd3ddf
mentioned_people: []
organization:
name: null
departments:
- id: naue
name: Naue
aliases: []
- id: vanheede
name: Vanheede
aliases: []
abbreviations: {}
known_entities:
Authoritative terminology:
- 'Bentofix (aliases: Bento Fix)'
- 'Carbofol (aliases: Carbofool, Carbovol, Karbofol)'
- 'Diekmeier (aliases: Diekmeyer, Dikmeier, Dikmeyer)'
- 'GM13 (aliases: GM 13, GM-13)'
- 'GM17 (aliases: GM 17, GM-17)'
- 'GM42 (aliases: GM 42, GM-42)'
- 'Kockmann (aliases: Cockmann, Kogmann)'
- 'Luminy (aliases: Lumini)'
- PBAT
- 'Secugrid (aliases: Sikirgut)'
- 'Secugrid HS (aliases: Sikirgut Heistlöse, Sikirgut-Heistlöse, Sikirgut-Heißlöse)'
- 'Secumat (aliases: Sekumat)'
- Secutex
context_rules:
participant_list_is_authoritative: true
do_not_infer_roles: true
do_not_infer_departments: true
do_not_infer_responsibilities: true
mentioned_people_are_not_participants: true
glossary_canonical_spelling: Use canonical glossary spellings only when the meeting
clearly refers to those terms; do not invent matches or replace unrelated words.
glossary_core_terms: Preserve surrounding context and use canonical core terms inside
compounds where appropriate.
@@ -0,0 +1,27 @@
# Regression expectations
The exact prompts and compact diarized protocol input are the evaluation evidence. Test semantic fidelity and grounded attribution; do not require a byte-identical protocol or treat editorial concision as an extraction error.
| ID | Expectation |
| --- | --- |
| R1 | The resulting protocol is clearly English. |
| R2 | The substantive discussion remains usable despite non-native English speech and the transcription imperfections visible in the evidence. |
| R3 | The four diarized clusters are not casually collapsed or conflated. |
| R4 | With the generation-002 mappings, Christian Psiorz, Martin Tazl, Dries Nollet, and Jorik Muylle remain attached to their explicit clusters. |
| R5 | An unmapped cluster may be named only where the transcript gives positive, unique conversational evidence. Generation 001 supplies strong evidence for Martin, Christian, and Dries. |
| R6 | The participant list alone must not justify a name assignment. In generation 001, the `SPEAKER_00` → Jorik Muylle assignment is not sufficiently explained by persisted evidence. |
| R7 | Where the evidence is not unique, retain anonymity or state uncertainty; do not manufacture a person identity. |
| R8 | Identity resolution must not turn an idea, question, offer, or reference into a personal responsibility. In particular, conditions around lab testing and discussion with polymer partners must retain their stated scope. |
| R9 | Preserve the material technical, regulatory, commercial, and tender context: regulatory limits on external recycled content; durability/testing constraints; technical purity/process discussion; potential certified hybrid materials; and the stated tender follow-up. |
| R10 | A human’s later editorial choice to omit information is distinct from a factual extraction error. No persisted human-edited protocol exists in this case, so this is a review rule rather than a textual golden answer. |
## Attribution-sensitive checks
- `SPEAKER_03` says that the tender will be sent; the generation-002 mapping supports Dries Nollet as the speaker. Verify the wording does not expand this to unrelated ownership.
- Martin offers lab-scale tests **if** Vanheede provides material; preserve the conditional offer rather than a committed unconditional task.
- The discussion says Jorik/Vanheede will discuss requirements with polymer partners in context, but a responsible person must be recorded only where the source explicitly assigns or accepts it.
- The protocol must distinguish current constraints from forecasts about EU standards and market changes.
## Reference outputs
`protocol_generation/002/protocol.md` is the primary mapped-speaker reference. `protocol_generation/001/protocol.md` is intentionally retained as a cautionary unmapped comparison: it is useful for content review, but its name attribution cannot be used as a target wherever the transcript does not support it.
@@ -0,0 +1,150 @@
{
"case_id": "vanheede_naue_recycled_hdpe_2026-09-14",
"privacy": "private real-world meeting sample; do not publish or redistribute outside intended development context",
"source_meeting_assistant_run": {
"path": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952",
"run_id": "d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952",
"timestamp": "2026-09-15T14:59:52+02:00",
"meeting_id": "vanheede-and-naue-recycling-and-circular-economy",
"title": "Vanheede and Naue, Recycling and Circular Economy",
"language": "en",
"full_pipeline_run": true
},
"audio_reference": {
"source_path": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/uploads/d7189066_2026-09-14 16-00-18 Recycled HDPE.flac",
"filename": "d7189066_2026-09-14 16-00-18 Recycled HDPE.flac",
"sha256": "750ef63cb807e47c2f0a4a9bdbff5334e46fc7e31978058daf47bd5c6153206d",
"size_bytes": 114357855,
"copied": false
},
"pipeline": {
"whisper": {
"backend": "whisper.cpp",
"model": "ggml-large-v3-turbo.bin",
"language": "en",
"threads": 6,
"flash_attention": true,
"runtime_seconds": 45.82035218599958
},
"diarization": {
"backend": "pyannote.audio",
"model": "pyannote/speaker-diarization-community-1",
"runtime": "container",
"device": "AMD Radeon RX 9070 via cuda",
"speaker_count": 4,
"speaker_labels": [
"SPEAKER_00",
"SPEAKER_01",
"SPEAKER_02",
"SPEAKER_03"
],
"turn_count": 501,
"exclusive_turn_count": 471,
"runtime_seconds": 66.73990752200007
}
},
"copied_artifacts": [
{
"case_path": "transcript/source_transcript.txt",
"source_path": "transcript/transcript.txt",
"sha256": "c3167524e96e5561fe5d91e6495454d216de8c7df780101a9bb70d2bc52ea8f2"
},
{
"case_path": "transcript/diarized_transcript.txt",
"source_path": "diarization/transcript_diarized.txt",
"sha256": "dfce25ebadb3ca3d247527d95260426953fa14acc8c60319253721ec175fbae9"
},
{
"case_path": "transcript/protocol_input.txt",
"source_path": "protocol/generations/001/transcript_input.txt and protocol/generations/002/transcript_input.txt",
"sha256": "0acab653aec6c50e35ca998ed38c3c39ac78e32aec99548c1832b15a11c8391b",
"representation": "exact compact diarized protocol input; byte-identical across generations"
},
{
"case_path": "context/meeting_context.yaml",
"source_path": "context/meeting_context.yaml",
"sha256": "0fe03f6ec3786e45094fa1e325f3fdec7a1b47c2a425d44e3f4e270074467076"
},
{
"case_path": "source_run_metadata.json",
"source_path": "run_metadata.json",
"sha256": "f5ef3f98bea8871dd7c5b7558faf2d62785bc76974cde41ca9924bc677a82382"
},
{
"case_path": "protocol_generation/001/exact_prompt.txt",
"source_path": "protocol/generations/001/exact_prompt.txt",
"sha256": "c08b561f9b1b524784eda96ada1422d1774be27da4a144bb99b6f757b42522a8"
},
{
"case_path": "protocol_generation/001/protocol.md",
"source_path": "protocol/generations/001/protocol.md",
"sha256": "3f4ba405813c6496aca1349f1337fecaa92f655e67999911bd5d50c2656a1d21"
},
{
"case_path": "protocol_generation/001/runtime_metadata.json",
"source_path": "protocol/generations/001/runtime_metadata.json",
"sha256": "b1904471bcd96662671783de16ee9dc22118c32b049806211b60536bab8a22e4"
},
{
"case_path": "protocol_generation/002/exact_prompt.txt",
"source_path": "protocol/generations/002/exact_prompt.txt",
"sha256": "49e028aeeffd44cba562160d3a120240a127f81bd74de75eddf9e482ec6cbae6"
},
{
"case_path": "protocol_generation/002/protocol.md",
"source_path": "protocol/generations/002/protocol.md",
"sha256": "8551ecd1e41aa06fdd138bdc5aeeed4ad55a77f3ec4ae5e6410dc063d9ea90fc"
},
{
"case_path": "protocol_generation/002/runtime_metadata.json",
"source_path": "protocol/generations/002/runtime_metadata.json",
"sha256": "d626828bc59f2bf291a76a21d4283b39233da37d584238e5bdda85dc91ff715a"
}
],
"protocol_generations": {
"001": {
"model": "qwen3.8:27b",
"temperature": 0,
"think": false,
"num_ctx": 32768,
"num_predict": 8192,
"num_thread": null,
"output_language": "en",
"completion_reason": "stop",
"output_token_count": 1407,
"prompt_token_count": 11665,
"speaker_mapping_count": 0
},
"002": {
"model": "qwen3.8:27b",
"temperature": 0,
"think": false,
"num_ctx": 32768,
"num_predict": 8192,
"num_thread": null,
"output_language": "en",
"completion_reason": "stop",
"output_token_count": 2304,
"prompt_token_count": 11933,
"speaker_mapping_count": 4,
"explicit_speaker_mappings": {
"SPEAKER_00": "Jorik Muylle",
"SPEAKER_01": "Martin Tazl",
"SPEAKER_02": "Christian Psiorz",
"SPEAKER_03": "Dries Nollet"
}
}
},
"exclusions": [
"audio",
"raw model responses",
"container logs",
"package-install logs",
"caches",
"unrelated runtime artifacts"
],
"operator_context": {
"claimed_subsequent_human_editing": "primarily editorial selection rather than factual correction",
"persisted_source_artifact": false
}
}
File diff suppressed because one or more lines are too long
@@ -0,0 +1,37 @@
# Meeting Protocol
## Meeting Overview
The meeting focused on exploring the feasibility of integrating recycled materials from Vanheede into Naue’s production processes, specifically for geosynthetics such as geomembranes. The discussion clarified that while technical processing is possible, regulatory constraints and market requirements currently prohibit the use of external post-consumer or post-industrial recycled materials for high-durability applications. The parties agreed that the primary obstacle is not technical capability but rather compliance with standards requiring extensive lot-by-lot testing for recycled inputs, which is economically unfeasible at scale.
## Regulatory Constraints and Market Limitations
Martin Tazl (Naue) explained that the core issue preventing the use of recycled material is regulatory, not technical. Current standards for geosynthetics (geomembranes, geogrids, nonwovens) prohibit the use of external recycled material for products claiming durability beyond five years. To obtain CE marking or claim long-term durability (e.g., 25–100 years), manufacturers must prove aging behavior and other properties for every single lot of recycled material used.
* **Testing Burden:** These tests take approximately four months and are expensive. This requires storing material for months before processing, creating logistical and financial barriers.
* **Market Reality:** Customers typically require long-term safety (100+ years) for waste containment. Recycled material, even if technically viable, is often perceived as lower quality or carries higher risk. Furthermore, recycled material is frequently more expensive than virgin material, offering no cost advantage.
* **Current Policy:** Naue currently operates under a "zero RC (recycled content) policy" for external materials. They are only permitted to use a limited percentage (10–15%) of internal reuse materials (e.g., side trims).
* **Future Outlook:** Martin Tazl noted that the EU is working on new standards to normalize recycled material usage. The expected outcome is that recycled material will face the same performance requirements as virgin material, but with a shift from lot-by-lot testing to manufacturer-level quality control (QC) and audits. This change is anticipated in the coming years but is not yet in effect.
## Technical Specifications and Processing Capabilities
The discussion addressed the technical requirements for processing recycled HDPE/PE materials in Naue’s extrusion lines.
* **Particle Size:** Jorik Muylle (Vanheede) inquired about the 38-micron specification. Martin Tazl clarified this refers to impurities (remaining dirt), not pellet size. Naue’s extruder filters are set at 38 microns; particles larger than this clog filters and cause wear, while smaller particles are acceptable.
* **Polymer Type:** Naue uses LLDPE or MDPE (not true HDPE polymer) mixed with carbon black to achieve HDPE-like density. This composition ensures high elongation at break (700–800%) and weldability, which are critical for geomembranes. True HDPE polymer is too stiff and prone to stress cracking.
* **Stress Crack Resistance:** Martin Tazl highlighted that adding even small amounts (1–2%) of high-density material to their standard resin can drastically reduce stress crack resistance (from >3000 hours to <50 hours). This non-linear behavior makes mixing different polymer grades risky.
* **Extrusion Technology:** Naue uses flat die cast extrusion. They have lines capable of ABA structures, but for geosynthetics, the material must be homogeneous from top to bottom; ABA structures do not mitigate regulatory or performance issues.
* **PP Content:** The presence of polypropylene (PP) is problematic because landfill applications typically require 100% polyethylene for weldability and chemical resistance.
## Feasibility of Specific Material Streams
* **Post-Consumer Household Packaging:** Jorik Muylle stated that Vanheede’s recycling plant is designed for post-consumer household packaging (bottles, films). He argued that using temporary landfill coverings (HDPE foil) as feedstock is not feasible at scale due to contamination, small lot sizes, and the need for extensive pre-processing (washing, sorting) that Vanheede’s current lines are not optimized for.
* **Geosynthetic Recycling:** Martin Tazl mentioned a research program ("Pro Geo UP") focused on recycling geosynthetics into new geosynthetics. This is technically feasible because the material is consistent, but the volume of available waste geosynthetics is too low to support industrial scale production.
* **Hybrid Materials:** Jorik Muylle suggested that the most viable path for recycled content in the near future is through hybrid materials produced by major polymer producers (e.g., TotalEnergies, Borealis, Ravago). These producers integrate recycled content into virgin polymer, providing certified, ready-to-use material with consistent quality, thereby bypassing the lot-by-lot testing burden for converters like Naue.
## Strategic Outlook and Next Steps
* **Short-Term:** No immediate commercial opportunity exists for using Vanheede’s recycled HDPE in Naue’s geomembrane production due to regulatory and market constraints.
* **Medium-Term (2–4 Years):** Both parties anticipate regulatory changes that may allow for the use of certified recycled content. Jorik Muylle plans to consult with his polymer partners (TotalEnergies, Borealis, Ravago) to understand the technical demands and norms for hybrid materials.
* **Testing Offer:** Martin Tazl offered to conduct lab-scale tests (approx. 50 kg) using a mixture of 20% Vanheede recycled material and 80% Naue virgin material to assess technical feasibility, though he emphasized this would not solve the regulatory issues.
* **Competitive Landscape:** Martin Tazl noted that competitors are also exploring recycling of geosynthetics but are not currently using post-consumer packaging streams. He predicted that recycled polyester (from bottles) may be the first recycled material to gain traction in geosynthetics due to its cleaner, more consistent stream.
## Action Items
* **Jorik Muylle (Vanheede):** To discuss technical demands and regulatory norms with polymer partners (TotalEnergies, Borealis, Ravago) regarding hybrid recycled/virgin materials.
* **Martin Tazl (Naue):** To share information on the "Pro Geo UP" research program and potentially conduct lab-scale tests with Vanheede material if requested.
* **Dries Nollet (Vanheede):** To send out the upcoming tender to Naue, keeping them in copy.
@@ -0,0 +1,63 @@
{
"output_language": "en",
"model": "qwen3.8:27b",
"prompt_token_count": 11665,
"output_token_count": 1407,
"prompt_evaluation_duration_ns": 19633919000,
"generation_duration_ns": 123254843000,
"total_ollama_duration_ns": 147435392290,
"client_wall_time_seconds": 147.45673817200077,
"completion_reason": "stop",
"done": true,
"request_count": 1,
"temperature": 0.0,
"think": false,
"num_ctx": 32768,
"num_predict": 8192,
"num_thread": null,
"selected_transcript_representation": "diarized_compact",
"estimated_input_tokens": 10734,
"safe_input_token_budget": 29000,
"input_token_estimation_method": "utf8_bytes_divided_by_4.4",
"fallback_used": false,
"diarization_enabled": true,
"speaker_attribution_available": true,
"speaker_attribution_loss_reason": null,
"speaker_mapping_count": 0,
"glossary_aliases_configured": {
"Bento Fix": "Bentofix",
"Carbofool": "Carbofol",
"Carbovol": "Carbofol",
"Cockmann": "Kockmann",
"Diekmeyer": "Diekmeier",
"Dikmeier": "Diekmeier",
"Dikmeyer": "Diekmeier",
"GM 13": "GM13",
"GM 17": "GM17",
"GM 42": "GM42",
"GM-13": "GM13",
"GM-17": "GM17",
"GM-42": "GM42",
"Karbofol": "Carbofol",
"Kogmann": "Kockmann",
"Lumini": "Luminy",
"Sekumat": "Secumat",
"Sikirgut": "Secugrid",
"Sikirgut Heistlöse": "Secugrid HS",
"Sikirgut-Heistlöse": "Secugrid HS",
"Sikirgut-Heißlöse": "Secugrid HS"
},
"glossary_replacements": [],
"generation_index": 1,
"generation_timestamp": "2026-09-15T13:05:17+00:00",
"source_run_id": "d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952",
"speaker_mapping": {},
"speaker_mapping_names": {},
"git": {
"meeting_lab": {
"repository": "/opt/git-projekts/meeting-lab",
"sha": "21082e66b3b603049898fb19fab3558e484a2478",
"dirty": true
}
}
}
File diff suppressed because one or more lines are too long
@@ -0,0 +1,82 @@
# Meeting Protocol
## Meeting Overview
**Title:** Vanheede and Naue, Recycling and Circular Economy
**Participants:**
* **Naue:** Christian Psiorz (Sales Director), Martin Tazl (Director R&D / Technic)
* **Vanheede:** Dries Nollet (Sales Person), Jorik Muylle (Technical)
**Context:**
The meeting aimed to explore the feasibility of using recycled plastic materials from Vanheede in Naue’s production processes, specifically for geosynthetics. The discussion clarified that while technical processing is possible, regulatory constraints and market requirements currently prohibit the use of external post-consumer recycled material for high-durability geosynthetics. The parties agreed that future opportunities depend on regulatory changes and the development of certified hybrid materials by major polymer producers.
## Clarification of Scope and Intent
**Discussion:**
Martin Tazl (Naue) clarified a potential misunderstanding regarding the scope of the collaboration. He noted that initial communications suggested a limited reuse of specific waste streams (e.g., HDPE foil) within Naue’s own operations. However, the questions forwarded by Vanheede implied a broader interest in using general recycled material for Naue’s commercial production.
Jorik Muylle (Vanheede) confirmed that the intent was broader: to explore if Vanheede’s general plastics recycling output (not just HDPE foil) could be utilized in Naue’s production. He explained that the initial focus on HDPE foil was deemed difficult due to contamination issues, leading to a broader discussion about using Vanheede’s general recycling capabilities.
**Outcome:**
The discussion shifted from a specific waste reuse case to a general inquiry about the feasibility of using Vanheede’s recycled output for Naue’s geosynthetic products.
## Regulatory Barriers and Market Constraints
**Discussion:**
Martin Tazl (Naue) explained that the primary obstacle to using external recycled material is not technical processability but regulatory compliance.
* **Durability Claims:** To claim a durability of more than five years (standard for landfill applications), manufacturers must prove aging behavior for every lot of recycled material used. Virgin material is exempt from this per-lot testing due to established QC audits.
* **Testing Burden:** Aging tests for recycled material take approximately four months and are expensive. This requires storing material for four months before processing, which is not feasible for large-scale production.
* **Market Limitation:** Without per-lot testing, products can only claim a five-year durability. This limits the market to temporary applications (e.g., temporary capping) where customers are willing to accept this limitation. However, customers typically prefer virgin material because it is cheaper, has no time-limit risk, and avoids the complexity of liability for long-term performance.
* **Current Policy:** Naue currently has a "zero RC (recycled content) policy" for all geotextiles (geomembranes, geogrids, wovens, nonwovens) due to these regulatory hurdles. Only a small percentage (10-15%) of internal side-trim reuse is permitted.
**Position:**
Christian Psiorz (Naue) confirmed that while a circular economy approach is desirable, current regulations prevent large-scale implementation. He noted that temporary capping applications (5-year lifespan) are technically possible but represent a narrow market.
## Technical Specifications and Material Properties
**Discussion:**
Jorik Muylle (Vanheede) and Martin Tazl (Naue) discussed specific technical parameters:
1. **Particle Size:**
* Jorik Muylle asked for clarification on the 38-micron limit.
* Martin Tazl explained that this refers to impurities/dirt, not pellet size. Naue’s extruder filters are set to 38 microns. Particles larger than this cause filter clogging and wear on the extruder, while particles smaller than this cannot be filtered out and may cause failure points in the membrane.
2. **Extrusion Technology:**
* Jorik Muylle asked if Naue uses flat die cast extrusion and if ABA (virgin-recycled-virgin) structures are possible.
* Martin Tazl confirmed they use flat die cast extrusion and have lines capable of ABA structures. However, for geosynthetics, the material must be homogeneous from top to bottom to ensure consistent performance; ABA structures do not offer a regulatory or technical advantage in this context.
3. **Polymer Type and Purity:**
* Jorik Muylle noted that Vanheede’s recycling inputs are primarily post-consumer household packaging (bottles, films), resulting in HDPE/LLDPE/MDPE mixes.
* Martin Tazl clarified that Naue’s "HDPE" geomembranes are actually made from a mix of LLDPE and MDPE (approx. 60% MDPE, 40% LLDPE) with added carbon black to achieve the required density. This mix provides better weldability and elongation at break (target ~800%) compared to true HDPE, which is too stiff and prone to stress cracking.
* **PP Content:** Jorik Muylle highlighted that PP content is critical for weldability. Martin Tazl confirmed that for landfill applications, 100% polyethylene is required; any significant PP content disqualifies the material for these high-end uses.
4. **Stress Crack Resistance:**
* Martin Tazl emphasized that geomembranes require high stress crack resistance (500–800+ hours). Virgin HDPE often fails this test (reaching only ~10 hours). Mixing different polymer grades (e.g., adding high-density material to low-density) can unpredictably degrade stress crack resistance, as demonstrated in Naue’s internal tests where adding 2% high-density material dropped resistance from 3000+ hours to under 50 hours.
## Feasibility of Vanheede’s Recycled Material
**Discussion:**
* **Vanheede’s Input:** Jorik Muylle stated that Vanheede’s plant is designed for post-consumer household waste (bottles, films). Their output is a mix of HDPE, LLDPE, and MDPE. They do not currently produce a pure, high-grade HDPE suitable for Naue’s specifications without further compounding.
* **Naue’s Assessment:** Martin Tazl stated that 100% recycled material from Vanheede’s current streams would likely not meet the specifications for high-end geomembranes due to variability and impurities. However, he acknowledged that if a major polymer producer (e.g., TotalEnergies, Borealis, Ravago) provides a certified hybrid material (e.g., 20% recycled + 80% virgin) with guaranteed quality, Naue could process it.
* **Scale and Availability:** Jorik Muylle noted that for scale, post-consumer streams are necessary. Martin Tazl agreed that availability is a major constraint for recycling geosynthetics themselves (e.g., old geomembranes), as they are designed to last 25–100 years and are not available in large quantities.
**Conclusion:**
Direct use of Vanheede’s current recycled output for Naue’s standard geomembrane production is not feasible due to regulatory and technical consistency issues. The viable path is through certified hybrid materials from major polymer producers.
## Future Outlook and Regulatory Changes
**Discussion:**
* **EU Regulations:** Martin Tazl mentioned that the EU is working on new standards to normalize the use of recycled materials in geosynthetics. He expects that future regulations will require recycled materials to meet the same performance standards as virgin materials, rather than allowing lower standards. This will likely shift the burden of quality assurance to the recycled material producer.
* **Research Programs:** Martin Tazl referenced the "ProGeoUP" research program, which focuses on recycling geosynthetics into new geosynthetics. He noted that while technically possible, the lack of available feedstock (old geosynthetics) limits its commercial viability.
* **Market Drivers:** Christian Psiorz (Naue) noted that financial incentives or regulatory pressure (e.g., mandatory recycled content in packaging by 2030) are needed to drive change. He expressed skepticism about the current market willingness to pay a premium for recycled content in geosynthetics.
* **Competitor Activity:** Martin Tazl stated he is not aware of competitors offering geomembranes made from post-consumer recycled material. He noted that bio-based polymers (e.g., bio-based polyester) are a more likely near-term development for recycled content in geosynthetics.
## Action Items and Next Steps
**Discussion:**
* **Information Exchange:** Jorik Muylle (Vanheede) requested that Naue share the current regulatory norms and technical specifications (MFR, elongation at break, stress crack resistance) for their geomembranes. This will allow Vanheede to discuss with their polymer partners (TotalEnergies, Borealis, Ravago) about developing suitable hybrid materials.
* **Testing:** Martin Tazl (Naue) offered to conduct lab-scale tests (30 cm wide, ~50 kg) with a mixture of 20% Vanheede recycled material and 80% virgin material to assess technical feasibility, if Vanheede provides the material. However, he emphasized that this would not solve the regulatory issues.
* **Tender:** Dries Nollet (Vanheede) mentioned that a tender is expected to be sent out by the end of the current week or the beginning of the next week. Christian Psiorz (Naue) confirmed they will stay in contact regarding this tender.
**Agreed Actions:**
1. **Naue (Martin Tazl):** To provide current regulatory norms and technical specifications for geomembranes to Vanheede.
2. **Vanheede (Jorik Muylle):** To discuss technical demands and potential hybrid material solutions with polymer partners (TotalEnergies, Borealis, Ravago).
3. **Both Parties:** To maintain contact regarding the upcoming tender and future regulatory developments.
## Unresolved Issues
* **Regulatory Compliance:** No current regulatory pathway exists for using external post-consumer recycled material in high-durability geosynthetics without per-lot testing, which is not commercially viable.
* **Material Consistency:** Vanheede’s current recycled output does not meet the purity and consistency requirements for Naue’s standard geomembrane production.
* **Market Demand:** There is currently no significant market demand for geomembranes made from recycled material due to cost and liability concerns.
@@ -0,0 +1,73 @@
{
"output_language": "en",
"model": "qwen3.8:27b",
"prompt_token_count": 11933,
"output_token_count": 2304,
"prompt_evaluation_duration_ns": 20091506000,
"generation_duration_ns": 190776825000,
"total_ollama_duration_ns": 214434255275,
"client_wall_time_seconds": 214.45167524499993,
"completion_reason": "stop",
"done": true,
"request_count": 1,
"temperature": 0.0,
"think": false,
"num_ctx": 32768,
"num_predict": 8192,
"num_thread": null,
"selected_transcript_representation": "diarized_compact",
"estimated_input_tokens": 10973,
"safe_input_token_budget": 29000,
"input_token_estimation_method": "utf8_bytes_divided_by_4.4",
"fallback_used": false,
"diarization_enabled": true,
"speaker_attribution_available": true,
"speaker_attribution_loss_reason": null,
"speaker_mapping_count": 4,
"glossary_aliases_configured": {
"Bento Fix": "Bentofix",
"Carbofool": "Carbofol",
"Carbovol": "Carbofol",
"Cockmann": "Kockmann",
"Diekmeyer": "Diekmeier",
"Dikmeier": "Diekmeier",
"Dikmeyer": "Diekmeier",
"GM 13": "GM13",
"GM 17": "GM17",
"GM 42": "GM42",
"GM-13": "GM13",
"GM-17": "GM17",
"GM-42": "GM42",
"Karbofol": "Carbofol",
"Kogmann": "Kockmann",
"Lumini": "Luminy",
"Sekumat": "Secumat",
"Sikirgut": "Secugrid",
"Sikirgut Heistlöse": "Secugrid HS",
"Sikirgut-Heistlöse": "Secugrid HS",
"Sikirgut-Heißlöse": "Secugrid HS"
},
"glossary_replacements": [],
"generation_index": 2,
"generation_timestamp": "2026-09-15T13:20:10+00:00",
"source_run_id": "d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952",
"speaker_mapping": {
"SPEAKER_00": "participant-c9280bb9",
"SPEAKER_01": "participant-8c7d5a9d",
"SPEAKER_02": "participant-0c5377fb",
"SPEAKER_03": "participant-dddd3ddf"
},
"speaker_mapping_names": {
"SPEAKER_00": "Jorik Muylle",
"SPEAKER_01": "Martin Tazl",
"SPEAKER_02": "Christian Psiorz",
"SPEAKER_03": "Dries Nollet"
},
"git": {
"meeting_lab": {
"repository": "/opt/git-projekts/meeting-lab",
"sha": "21082e66b3b603049898fb19fab3558e484a2478",
"dirty": true
}
}
}
@@ -0,0 +1,64 @@
{
"audio_preparation": {
"canonical_output": {
"bits_per_sample": 16,
"channels": 1,
"codec": "pcm_s16le",
"container": "wav",
"sample_rate_hz": 16000
},
"ffmpeg_executable": "/usr/bin/ffmpeg",
"normalization_enabled": true,
"normalization_filter": "loudnorm=I=-16:LRA=11:TP=-1.5",
"normalization_method": "ffmpeg_loudnorm",
"original_format": "flac",
"original_source_name": "d7189066_2026-09-14 16-00-18 Recycled HDPE.flac",
"original_source_path": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/uploads/d7189066_2026-09-14 16-00-18 Recycled HDPE.flac",
"preparation_method": "ffmpeg",
"prepared_audio_path": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952/audio/prepared.wav"
},
"diarization": {
"actual_device": "cuda",
"backend": "pyannote.audio",
"device_name": "AMD Radeon RX 9070",
"enabled": true,
"metadata_path": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952/diarization/metadata.json",
"model": "pyannote/speaker-diarization-community-1",
"requested_device_mode": "auto",
"runtime": "container",
"runtime_seconds": 66.73990752200007,
"speaker_count": 4,
"transcript_diarized": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952/diarization/transcript_diarized.json",
"transcript_diarized_text": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952/diarization/transcript_diarized.txt"
},
"failure": null,
"input_audio": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/uploads/d7189066_2026-09-14 16-00-18 Recycled HDPE.flac",
"model": "qwen3.8:27b",
"ollama_endpoint": "http://127.0.0.1:11434",
"protocol_output": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952/protocol.md",
"run_id": "d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952",
"stage_runtimes_seconds": {
"audio_preparation": 44.826,
"diarization": 86.676,
"diarization_alignment": 0.02,
"protocol": 147.471,
"setup": 44.827,
"transcript_validation": 0.0,
"validation": 0.0,
"whisper": 45.823
},
"status": "completed",
"timestamp": "2026-09-15T14:59:52+02:00",
"timing": {
"stages_seconds": {
"diarization": 86.69619259699903,
"preparing": 44.82779592799852,
"protocol_generation": 147.4888497869997,
"transcription": 45.823291290998895
},
"total_seconds": 324.8364486570008
},
"total_runtime_seconds": 324.836,
"transcript_output": "/opt/git-projekts/meeting-assistant/data/meetings/vanheede-and-naue-recycling-and-circular-economy/runs/d7189066_2026-09-14_16-00-18_Recycled_HDPE_20260915_145952/transcript/transcript.json",
"whisper_model": "/home/martin/whisper.cpp/whisper.cpp/models/ggml-large-v3-turbo.bin"
}
@@ -0,0 +1,30 @@
# Speaker identity analysis
## Evidence boundary
Generation 001 received the participant list but **no confirmed speaker mappings**. Its prompt also says “Erfinde keine Fakten oder Identitäten” (do not invent facts or identities). It nevertheless uses participant names. The table below assesses whether that use can be justified from the persisted compact transcript. Generation 002 is different: its exact prompt explicitly maps every cluster and instructs the model to use only those mappings.
| Speaker cluster | Identity in generation 002 | Explicit mapping? | Relevant persisted discourse evidence | Assessment |
| --- | --- | --- | --- | --- |
| `SPEAKER_00` | Jorik Muylle | Yes, in 002 only | The speaker says “Well, Martin” and thanks “Christian”; a phonetic-looking `Jurek`/`Jürich` is mentioned by other speakers, but no turn uniquely identifies this speaker as Jorik Muylle. | **NOT EXPLAINABLE FROM PERSISTED EVIDENCE** for 001; **EXPLICITLY MAPPED** for 002. |
| `SPEAKER_01` | Martin Tazl | Yes, in 002 only | `SPEAKER_00` says “Well, Martin, thank you for the explanation”; `SPEAKER_03` refers to “your Martin’s requirements”; the following turn is the technical explanation. | **STRONGLY DISCOURSE-SUPPORTED** for 001; **EXPLICITLY MAPPED** for 002. |
| `SPEAKER_02` | Christian Psiorz | Yes, in 002 only | `SPEAKER_00` thanks “Christian” for what he gave before; `SPEAKER_01` speaks of “Christian and me”; `SPEAKER_02` is the other Naue-side speaker. | **STRONGLY DISCOURSE-SUPPORTED** for 001; **EXPLICITLY MAPPED** for 002. |
| `SPEAKER_03` | Dries Nollet | Yes, in 002 only | Near the close, `SPEAKER_02` addresses “Dries” about the upcoming tender; `SPEAKER_03` then says they will send the tender. | **STRONGLY DISCOURSE-SUPPORTED** for 001; **EXPLICITLY MAPPED** for 002. |
Short excerpts are intentionally kept only to what establishes the link:
- `SPEAKER_00`: “Well, Martin, thank you for the explanation … and Christian …”
- `SPEAKER_03`: “based upon … Martin’s requirements”
- `SPEAKER_02`: “Dries, I think we will stay in contact … regarding the upcoming tender.”
- `SPEAKER_03`: “end of this week … we will send out the tender.”
## Counterexamples and implications
The participant list has exactly four names and diarization has exactly four clusters. That cardinality makes an arbitrary remaining-name assignment easy; it is not positive evidence. In particular, the transcript mentions `Jurek`, `Jürich`, and `Jeroen`, which are not canonical participant names in the context. The available record does not establish whether either phonetic form denotes Jorik Muylle or a different person. Do not use the apparent correct remaining-name match as evidence that generation 001 performed reliable discourse identity resolution.
The case therefore supports two distinct regression checks:
1. With the four explicit mappings in generation 002, preserve every mapped identity correctly.
2. Without an explicit mapping, allow a name only when positive discourse evidence identifies a unique cluster. Participant-list membership and one-to-one count matching alone are insufficient. For this transcript, `SPEAKER_00` must remain anonymous unless new evidence is supplied.
This evidence assessment concerns identity attribution only. It does not make speaker identity a basis for assigning an action or responsibility.
@@ -0,0 +1,575 @@
[00:00:00.000 - 00:00:07.880] SPEAKER_03: Well it's a Monday. Weekend was too short as always.
[00:00:07.880 - 00:00:22.380] SPEAKER_02: Everything's good. How many of your colleagues are joining? Jurek and Jeroen both coming?
[00:00:22.380 - 00:00:36.180] SPEAKER_03: Normally I invite Jeroen. There he comes. I think Jurek will not join.
[00:00:36.180 - 00:00:51.160] SPEAKER_00: Welcome. Hello to everybody. Jurek will not join, that was your question. I see he's in
[00:00:51.160 - 00:00:56.320] SPEAKER_00: a meeting so what he said as it will probably be a technical meeting or a more
[00:00:56.320 - 00:01:02.020] SPEAKER_00: technical meeting I will correct get the info from my colleagues. Yes I think the
[00:01:02.020 - 00:01:07.660] SPEAKER_03: aim of today's meeting is to to get to know more to get to know more about let's
[00:01:07.660 - 00:01:14.800] SPEAKER_03: say the specifications in terms of using the recycled material from Van Heden and
[00:01:14.800 - 00:01:25.660] SPEAKER_03: and if we can let's say find a solution or if we can work together on this front as well since it was currently still at an early stage
[00:01:25.660 - 00:01:36.520] SPEAKER_03: where we in the previous meeting discussed the potential opportunity and then Christian forwarded based upon I think it's your Martin's requirements.
[00:01:36.520 - 00:01:47.380] SPEAKER_03: And then I forwarded it to Jeroen our expert at Van Heden and then he I think forwarded some additional questions.
[00:01:47.380 - 00:01:58.240] SPEAKER_01: Yeah so before we start we should get something straight when Christian and me were starting
[00:01:58.240 - 00:02:05.380] SPEAKER_01: talking about or when Christian told me about your ideas what to do there was I assumed that it was that it would be about a limited a limited amount of material which is on your side which you do recycle and you thought about okay we can do it.
[00:02:05.380 - 00:02:18.380] SPEAKER_01: We can reuse this particular material for our purpose on our own landfills so that's what I understood the questions you sent they are all valid and this is all most of that is well known.
[00:02:18.380 - 00:02:34.380] SPEAKER_01: But to me that sounded like a general offer to use recycled material to use recycled material in our production and so first of all was that was that your intention?
[00:02:34.380 - 00:02:55.380] SPEAKER_01: So I think first we discussed with Christian in the last meeting if it was possible to re-use HDPE foil which we use for a temporary basis and then get back to let's say the recycling of HDPE foil.
[00:02:55.380 - 00:03:18.380] SPEAKER_03: But I think this was deemed to be very difficult since it was also of course polluted and also difficult to let's say yeah you put it on the dirt you have to pull it back it's contaminated.
[00:03:18.380 - 00:03:33.380] SPEAKER_03: And then Jürich I think switched his view to okay at Van Heide we also have plastics recycling which is not only linked to HDPE foil.
[00:03:33.380 - 00:03:45.380] SPEAKER_03: And that was let's say the broader discussion could we be or could the products that we supply or foresee be also utilized in your production process.
[00:03:45.380 - 00:03:52.380] SPEAKER_01: So this is okay if this is a general question then unfortunately no and yes.
[00:03:52.380 - 00:03:59.380] SPEAKER_01: So in general I have to I have to do a little bigger circle.
[00:03:59.380 - 00:04:07.380] SPEAKER_01: It is it is less on the side if we can technically process these kind of materials.
[00:04:07.380 - 00:04:16.380] SPEAKER_01: It's more like we can't make any product of of that for any market we are working on.
[00:04:16.380 - 00:04:20.380] SPEAKER_01: The problem is not the technical process ability.
[00:04:20.380 - 00:04:34.380] SPEAKER_01: So if we follow some standards or and some let's say some requirements like MFR has to be in a certain range and so on and there has to be a minimum of stabilization and whatsoever.
[00:04:34.380 - 00:04:39.380] SPEAKER_01: But that's nothing you can fix which you cannot fix with a master batch for example.
[00:04:39.380 - 00:04:52.380] SPEAKER_01: But all this or but a lot of regulations standards and so on just prohibit using of as by now.
[00:04:52.380 - 00:04:57.380] SPEAKER_01: Yeah, just prohibit the usage of external recycling material.
[00:04:57.380 - 00:05:09.380] SPEAKER_01: So we're just not allowed to use post consumer post industrial recycling for making geo membranes as geo membranes.
[00:05:09.380 - 00:05:21.380] SPEAKER_01: And the main problem about that is whenever you want to claim a durability of more than five years then you have to prove certain things.
[00:05:21.380 - 00:05:22.380] SPEAKER_01: So that's where it begins.
[00:05:22.380 - 00:05:38.380] SPEAKER_01: So to get a CE marking, for example, you have to prove basically every lot you use, you have to prove that the aging behavior and so on.
[00:05:38.380 - 00:05:56.380] SPEAKER_01: So of course, we don't do that for virgin material because for virgin material it is assumed as long as the QC of the manufacturer of the virtual material is checked every it's audited in regular times and so on.
[00:05:56.380 - 00:06:14.380] SPEAKER_01: It can be assumed that the material stays always the same and so on and therefore they elongated this timeframe where you have to recheck this to about five years if I'm not wrong for RC material where you don't actually know what material went into that lot.
[00:06:14.380 - 00:06:22.380] SPEAKER_01: You probably can't tell if this was yogurt bins or whatsoever.
[00:06:22.380 - 00:06:25.380] SPEAKER_01: You technically have to do this for every lot.
[00:06:25.380 - 00:06:29.380] SPEAKER_01: And these tests are not only expensive.
[00:06:29.380 - 00:06:32.380] SPEAKER_01: They are as well taking four months.
[00:06:32.380 - 00:06:41.380] SPEAKER_01: So for us, if you would use for our products recycled material now.
[00:06:41.380 - 00:06:53.380] SPEAKER_01: So it's the regulation nowadays, we would have to store the material for about four months until we get the result of the aging before we can further process.
[00:06:53.380 - 00:06:59.380] SPEAKER_01: That has nothing to do with the quality of the material or the possibility process ability of the material.
[00:06:59.380 - 00:07:10.380] SPEAKER_01: It's just because of the regulations that we have to check every lot for a certain for certain things, especially aging and so on.
[00:07:10.380 - 00:07:19.380] SPEAKER_01: That said, this limits the market to so if you don't do that, you can only claim five years of durability.
[00:07:19.380 - 00:07:23.380] SPEAKER_01: And this limits your market very much.
[00:07:23.380 - 00:07:29.380] SPEAKER_01: You can claim that five years is pretty much enough for temporary capping or whatsoever.
[00:07:29.380 - 00:07:41.380] SPEAKER_01: The problem in the discussion is the customer usually doesn't know by but for sure that the application will be done within five years.
[00:07:41.380 - 00:07:48.380] SPEAKER_01: So what happens if for any reason you have to use it longer longer than the than the five years?
[00:07:48.380 - 00:07:52.380] SPEAKER_01: Who is taking the responsibility for the time after?
[00:07:52.380 - 00:07:59.380] SPEAKER_01: And what happens if you buy the material and the installation is delayed by one year for any reason?
[00:07:59.380 - 00:08:02.380] SPEAKER_01: You know, so and the price difference.
[00:08:02.380 - 00:08:11.380] SPEAKER_01: If there is any to our to our knowledge, usually the recycling material is even more expensive than version material in most of the cases.
[00:08:11.380 - 00:08:16.380] SPEAKER_01: The customers usually just go with the virgin material is cheaper.
[00:08:16.380 - 00:08:20.380] SPEAKER_01: They don't have to worry about these five years timeframe.
[00:08:20.380 - 00:08:23.380] SPEAKER_01: So much more convenient.
[00:08:23.380 - 00:08:24.380] SPEAKER_01: That's the problem.
[00:08:24.380 - 00:08:26.380] SPEAKER_01: One hand side.
[00:08:26.380 - 00:08:32.380] SPEAKER_01: Second hand side is the technical requirements to geomembranes in particular.
[00:08:32.380 - 00:08:37.380] SPEAKER_01: For geomembranes, it's not only to be aging proof and so on.
[00:08:37.380 - 00:08:44.380] SPEAKER_01: It's as well, there are some requirements which are related to the chemicals in the landfill.
[00:08:44.380 - 00:08:47.380] SPEAKER_01: So it's it's called stress crack resistance.
[00:08:47.380 - 00:08:58.380] SPEAKER_01: It's basically you always have some force on a geomembrane in due to the the own weight and due to some capping on top or whatsoever.
[00:08:58.380 - 00:09:06.380] SPEAKER_01: So like when it forms like a mattress or whatsoever, and especially HDPEs.
[00:09:06.380 - 00:09:17.380] SPEAKER_01: So really HDPEs where the density of the polymer exceeds 900, 45, 40, 45, something like that.
[00:09:17.380 - 00:09:21.380] SPEAKER_01: They are very prone to stress crack behavior.
[00:09:23.380 - 00:09:30.380] SPEAKER_01: We have to proof our products usually something like 500 to 800 hours of stress crack resistance.
[00:09:30.380 - 00:09:33.380] SPEAKER_01: You don't have to understand exactly the test.
[00:09:33.380 - 00:09:42.380] SPEAKER_01: But what we know from HDPEs, even from virgin HDPEs, which are by themselves already HDPE.
[00:09:42.380 - 00:09:47.380] SPEAKER_01: We reach something like 10 hours under the same conditions.
[00:09:47.380 - 00:09:54.380] SPEAKER_01: And this really is not suitable for landfill, typical landfill applications.
[00:09:54.380 - 00:10:13.380] SPEAKER_01: So the possibility of using recycled material would need a customer who is willing to pay more, get less and has some restrictions on how to use it.
[00:10:13.380 - 00:10:24.380] SPEAKER_01: So we can discuss about, yeah, if on a capping, which is only temporary and usually just necessary to keep away some water.
[00:10:24.380 - 00:10:31.380] SPEAKER_01: And even if it has a minor leakage, it is still much better than without the capping and so on.
[00:10:31.380 - 00:10:38.380] SPEAKER_01: But that is a very narrow, very narrow market.
[00:10:38.380 - 00:11:00.380] SPEAKER_01: So our concerns are 80, I would say 90% about regulations and market restrictions and only 10% about the, how to say, process stability and the possibility to run the material actually on our lines.
[00:11:00.380 - 00:11:18.380] SPEAKER_01: So for example, the requirement of the ash remains is just a requirement that there must not be any grains in the membrane because these are failure spots.
[00:11:18.380 - 00:11:33.380] SPEAKER_01: So that's when you apply some, when you apply some load to the membrane, then we know that at that places, the membrane will start failing much earlier as within a, if there is a very homogeneous metric.
[00:11:33.380 - 00:11:36.380] SPEAKER_01: So that's the reasons why we are.
[00:11:36.380 - 00:11:37.380] SPEAKER_01: So that's the reasons why we asked for that.
[00:11:37.380 - 00:11:46.380] SPEAKER_01: And if you exceed a certain, let's say a certain grain size and a certain amount, you get some wear and tear in the extruder.
[00:11:46.380 - 00:11:48.380] SPEAKER_01: You probably know that very well.
[00:11:48.380 - 00:11:57.380] SPEAKER_01: So that's the reason why we request these boundary conditions like this 38 micrometers.
[00:11:57.380 - 00:12:04.380] SPEAKER_01: That's just because bigger grain sizes would cause failures in the membrane later on.
[00:12:04.380 - 00:12:09.380] SPEAKER_01: And it usually will create wear and tear on our extruders.
[00:12:09.380 - 00:12:10.380] SPEAKER_01: That's, that's the two things.
[00:12:10.380 - 00:12:12.380] SPEAKER_01: It's not just like, I didn't bring that up.
[00:12:12.380 - 00:12:13.380] SPEAKER_01: Just yeah.
[00:12:13.380 - 00:12:16.380] SPEAKER_01: That's the specification of the virtual material.
[00:12:16.380 - 00:12:25.380] SPEAKER_01: We have some real, we have some real reasons, real reasons for that, which are related to our product.
[00:12:25.380 - 00:12:34.380] SPEAKER_01: So, so to say a general usage of recycled material, and that it's not particular for your product.
[00:12:34.380 - 00:12:50.380] SPEAKER_01: It's in general is mostly the biggest, the biggest obstacle is the, is the, the, the, the standards and so on.
[00:12:50.380 - 00:12:59.380] SPEAKER_01: Second is it, it has to have a certain cleanliness to be processed, but the main thing is the regulations.
[00:12:59.380 - 00:13:13.380] SPEAKER_01: Um, um, as well, you need to, for, for landfill, um, applications or for all ceiling applications, which are where, where people are willing to pay more money than for a PVC film.
[00:13:13.380 - 00:13:21.380] SPEAKER_01: Um, they usually look exactly for the PE, uh, behavior and the resistance of PE.
[00:13:21.380 - 00:13:40.380] SPEAKER_01: And when you, when you have like 5% or 2% or whatever for, from PP or polyester or whatsoever in your mixture, you just out by definition because they usually request 100% polyethylene.
[00:13:40.380 - 00:13:50.380] SPEAKER_01: Um, so that's just, yeah, it's, it's, it's a geomembrane market on that, on that point.
[00:13:50.380 - 00:14:06.380] SPEAKER_01: On the other hand, if there is a, if there is a demand, or if you want to do it really in a closed, in a closed, closed environment or a closed project where everyone is aware of the costs and the limitations.
[00:14:06.380 - 00:14:12.380] SPEAKER_01: Um, it shouldn't be a problem for us to make these products.
[00:14:12.380 - 00:14:18.380] SPEAKER_01: We have some constraints because we have to separate it from our other materials.
[00:14:18.380 - 00:14:22.380] SPEAKER_01: So silo handling and so on will be an issue.
[00:14:22.380 - 00:14:35.380] SPEAKER_01: But from the technical point of view, if the MFR at, um, I think you gave 200, uh, 192.16, right?
[00:14:35.380 - 00:14:39.380] SPEAKER_00: We gave below one on a, on a 2.16 kilogram measurement.
[00:14:39.380 - 00:14:40.380] SPEAKER_01: Yeah.
[00:14:40.380 - 00:14:49.380] SPEAKER_01: So, um, we are usually that that's what we, so one is already fluid for us, but we are in the range.
[00:14:49.380 - 00:14:50.380] SPEAKER_01: I would say we are in the same range.
[00:14:50.380 - 00:14:54.380] SPEAKER_01: So the minimum we use actually is 0.3, if I'm not wrong.
[00:14:54.380 - 00:14:56.380] SPEAKER_01: And yeah, maybe up to one.
[00:14:56.380 - 00:14:59.380] SPEAKER_01: So we are more familiar with five kilogram values.
[00:14:59.380 - 00:15:14.380] SPEAKER_01: So the five kilogram values would be between one and 0.6 to, I would say three grams per 10 minutes or something like that.
[00:15:14.380 - 00:15:24.380] SPEAKER_01: So generally there's a good chance to process that, especially when we mix it with virgin material.
[00:15:24.380 - 00:15:37.380] SPEAKER_01: So, and that's the, but anyways, this will give us much better tech or much more fitting technical parameters, but it will not solve the problem on the regulation side.
[00:15:37.380 - 00:15:43.380] SPEAKER_01: On a regulation side at the moment, there will be, there will be some changes in the future.
[00:15:43.380 - 00:15:50.380] SPEAKER_01: The European union is working on standards, um, to normalize the usage of, um, recycled materials.
[00:15:50.380 - 00:15:55.380] SPEAKER_01: But as for now, and I would say the upcoming five years doesn't matter.
[00:15:55.380 - 00:15:59.380] SPEAKER_01: There is a zero RC policy in the products we make.
[00:15:59.380 - 00:16:03.380] SPEAKER_01: And it's not only geomembranes, it's geo grids.
[00:16:03.380 - 00:16:13.380] SPEAKER_01: It's nonwovens, it's wovens, everything, which is considered, uh, geo textile in its main definition.
[00:16:13.380 - 00:16:15.380] SPEAKER_01: There is this limitations on that.
[00:16:15.380 - 00:16:23.380] SPEAKER_01: So we are only allowed to use a certain percentage of internal reuse materials of side trims and so on.
[00:16:23.380 - 00:16:24.380] SPEAKER_01: That's the only excuse.
[00:16:24.380 - 00:16:26.380] SPEAKER_01: And even that is heavily limited.
[00:16:27.380 - 00:16:30.380] SPEAKER_01: So we can only use, I think 10 or 15% of that.
[00:16:32.380 - 00:16:37.380] SPEAKER_02: So generally speaking, the, there's a certain activity.
[00:16:37.380 - 00:16:38.380] SPEAKER_02: Yep.
[00:16:38.380 - 00:16:45.380] SPEAKER_02: That this, uh, circular economy approach, uh, that we discussed last week is, is, uh, really possible.
[00:16:45.380 - 00:16:56.380] SPEAKER_02: Um, on the other hand, what Martin is just describing, um, it limits our action on the market with regards to all these regulations.
[00:16:56.380 - 00:17:03.380] SPEAKER_02: So we, from, from, from now aside, we cannot think this in a, in a large scale model.
[00:17:03.380 - 00:17:04.380] SPEAKER_02: Yeah.
[00:17:04.380 - 00:17:15.380] SPEAKER_02: Um, but if, uh, wherever there are demands that the material just needs to, uh, survive for five years.
[00:17:15.380 - 00:17:18.380] SPEAKER_02: So like the temporary cappings that you described trees.
[00:17:18.380 - 00:17:19.380] SPEAKER_02: Yeah.
[00:17:19.380 - 00:17:20.380] SPEAKER_02: Then of course, this is possible.
[00:17:20.380 - 00:17:34.380] SPEAKER_00: Well, Martin, thank you for the explanation and, and Christian that you, that you gave before I jump into the regulatory, um, aspects.
[00:17:34.380 - 00:17:39.380] SPEAKER_00: Uh, it's just for me to understand also with the technical and some of the, of the topics you already touched.
[00:17:39.380 - 00:17:54.380] SPEAKER_00: But to understand well on the particle size, do you mean that at the 38 micron, that, that should be the, uh, uh, uh, if we go put it into the melting phase, that no particles smaller than 38 micron can be found.
[00:17:54.380 - 00:18:00.380] SPEAKER_00: So nothing to do with the pellet size or, this is about possible impurities.
[00:18:00.380 - 00:18:02.380] SPEAKER_01: So remaining dirt, I would say.
[00:18:02.380 - 00:18:03.380] SPEAKER_01: Yeah.
[00:18:03.380 - 00:18:10.380] SPEAKER_01: So, um, um, our C set in the extruder is actually a 38 micron.
[00:18:10.380 - 00:18:15.380] SPEAKER_01: So everything which is below that is accepted because we can't even filter it out on our own.
[00:18:15.380 - 00:18:18.380] SPEAKER_01: Everything above that will be filtered out.
[00:18:18.380 - 00:18:22.380] SPEAKER_01: If it is too much, we get in trouble because our filters clock very soon.
[00:18:22.380 - 00:18:23.380] SPEAKER_01: Yeah.
[00:18:23.380 - 00:18:24.380] SPEAKER_01: Okay.
[00:18:24.380 - 00:18:28.380] SPEAKER_00: Um, the technology that you are using, is that cast extrusion or.
[00:18:28.380 - 00:18:30.380] SPEAKER_00: Flat die cast extrusion.
[00:18:30.380 - 00:18:31.380] SPEAKER_01: Flat die cast extrusion.
[00:18:31.380 - 00:18:32.380] SPEAKER_00: Okay.
[00:18:32.380 - 00:18:33.380] SPEAKER_00: From.
[00:18:33.380 - 00:18:42.380] SPEAKER_00: And so that's one single layer, or do you have the option to, to, uh, let's say if it's in the future, it is possible to put recycled contents.
[00:18:42.380 - 00:18:49.380] SPEAKER_00: Are you, uh, in the option of doing ABA, uh, structures and put the recycled contents in the middle?
[00:18:49.380 - 00:18:51.380] SPEAKER_01: We have one line, which does ABA.
[00:18:51.380 - 00:18:56.380] SPEAKER_01: And we have one line, which actually has two extruders, which are extruding next to each other.
[00:18:56.380 - 00:19:05.380] SPEAKER_01: But for the more modern one, we have ABA or, uh, you can switch between A, B, A, B, A, as you want.
[00:19:05.380 - 00:19:06.380] SPEAKER_01: Okay.
[00:19:06.380 - 00:19:14.380] SPEAKER_01: But for our regulations or for our purpose, um, it doesn't make a difference.
[00:19:14.380 - 00:19:15.380] SPEAKER_01: So we can't differentiate.
[00:19:15.380 - 00:19:27.380] SPEAKER_01: It's not like, um, for example, window frames or so where you do the virtual material to the outside to get the pretty look and to get, uh, the colors and whatsoever.
[00:19:27.380 - 00:19:37.380] SPEAKER_01: And for the filling, you use the recycled materials, our material, um, how to say, has to last from the top to the bottom the same.
[00:19:37.380 - 00:19:38.380] SPEAKER_01: Yeah.
[00:19:38.380 - 00:19:39.380] SPEAKER_00: Okay.
[00:19:39.380 - 00:19:40.380] SPEAKER_00: Okay.
[00:19:40.380 - 00:19:41.380] SPEAKER_00: Um, okay.
[00:19:41.380 - 00:19:52.380] SPEAKER_00: And then based on the polymer purity, uh, as, as you're only using virgin material, of course, it's only PE.
[00:19:52.380 - 00:20:02.380] SPEAKER_00: Um, or products are also, um, I mean, the PP content is, uh, an important parameter because of the, um, uh, you, you should be able to still seal the material.
[00:20:02.380 - 00:20:11.380] SPEAKER_00: And especially also on our landfill, uh, if you cannot longer make a good seal between the different, um, products, then we have a problem, a problem.
[00:20:11.380 - 00:20:14.380] SPEAKER_00: That's the same thing, which goes to pipes, for instance.
[00:20:14.380 - 00:20:21.380] SPEAKER_00: So also pipes needs to be well, weldable, uh, still, and PP has a higher melting point compared to, to PE.
[00:20:21.380 - 00:20:31.380] SPEAKER_00: Um, and therefore I know the spec is there around 4% of, of, uh, of PP, but that's maybe things you're not very familiar with as you, as at this point,
[00:20:31.380 - 00:20:35.380] SPEAKER_00: you're only using, um, HD, right?
[00:20:35.380 - 00:20:44.380] SPEAKER_01: Again, we, we, technically we use, um, this is a little bit weird, but technically we use maximum MDPE.
[00:20:44.380 - 00:20:53.380] SPEAKER_01: So, uh, usually we use, uh, linear low, so LLDPE or MDPE from the work, from the polymer side.
[00:20:53.380 - 00:21:00.380] SPEAKER_01: And our material is just called HDPE because we add about two and a half percent carbon black,
[00:21:00.380 - 00:21:04.380] SPEAKER_01: which lifts the density in the range of HDPE.
[00:21:04.380 - 00:21:09.380] SPEAKER_01: So from the polymer side, we are using maximum MDPE.
[00:21:09.380 - 00:21:12.380] SPEAKER_01: And that's the reason why we have the pretty good weldability.
[00:21:12.380 - 00:21:17.380] SPEAKER_01: Um, we have this stress, stress, stress, resistance.
[00:21:17.380 - 00:21:27.380] SPEAKER_01: And some other values like, um, yield elongation in a certain range, uh, maximum elongation.
[00:21:27.380 - 00:21:31.380] SPEAKER_01: You need a break along maximum elongation at break of about 800%.
[00:21:31.380 - 00:21:36.380] SPEAKER_01: I think some, uh, some regulations ask for 700, some for 800.
[00:21:36.380 - 00:21:41.380] SPEAKER_01: And we know that if you use HDPE polymer.
[00:21:41.380 - 00:21:48.380] SPEAKER_01: So where the polymer already is an HDPE, um, you usually can't reach that because it's too stiff.
[00:21:48.380 - 00:21:49.380] SPEAKER_01: Yeah.
[00:21:49.380 - 00:21:50.380] SPEAKER_00: Yeah.
[00:21:50.380 - 00:21:51.380] SPEAKER_01: Yeah.
[00:21:51.380 - 00:21:52.380] SPEAKER_01: So that's the reason.
[00:21:52.380 - 00:21:53.380] SPEAKER_01: Yeah.
[00:21:53.380 - 00:21:54.380] SPEAKER_01: It's a little bit weird.
[00:21:54.380 - 00:22:01.380] SPEAKER_01: So we call it HDPE, but actually it is more like, I would say 60%, um, MDPE and 40% linear low.
[00:22:01.380 - 00:22:02.380] SPEAKER_01: Yeah.
[00:22:02.380 - 00:22:03.380] SPEAKER_00: Okay.
[00:22:03.380 - 00:22:11.380] SPEAKER_00: Well, our recycling plant, the only inputs we are using are, um, uh, post-consumer household packaging.
[00:22:11.380 - 00:22:15.380] SPEAKER_00: And so we are, we're integrated with our group and we do collection of household packaging.
[00:22:15.380 - 00:22:17.380] SPEAKER_00: We have sorting lines for household packaging.
[00:22:17.380 - 00:22:25.380] SPEAKER_00: And then the sorted, uh, bales of, uh, HDPE and PP are then further treated at our, uh, polymer plants.
[00:22:25.380 - 00:22:37.380] SPEAKER_00: And so on, on HDPE, this means, uh, yeah, mainly bottles, uh, HD, but you could also have some LDPE from creams, uh, cream and packaging, things like that.
[00:22:37.380 - 00:22:55.380] SPEAKER_00: Um, and so on HD, we don't only have chemical bottles because the chemical bottles have a high crack, um, stress resistance because you don't want, um, chemical to, to go out, but not in the, in the values that you are mentioning to be clear, uh, because we're still speaking about bottles and canisters.
[00:22:55.380 - 00:22:59.380] SPEAKER_00: Um, so that's the grade we are making.
[00:22:59.380 - 00:23:12.380] SPEAKER_00: So if you're referring to the, uh, to the high allegation that breaks, and then I understand your argument that it's rather an LDPE or MD and not an HD that you're using in the, in the, in the film.
[00:23:12.380 - 00:23:16.380] SPEAKER_00: So, um, so yeah, good to know.
[00:23:16.380 - 00:23:21.380] SPEAKER_00: Um, I think at this point, uh, mainly to regulatory affairs, it's a bit difficult.
[00:23:21.380 - 00:23:24.380] SPEAKER_00: Nevertheless, it's already good to understand the process.
[00:23:24.380 - 00:23:27.380] SPEAKER_00: Also good to understand that there is an ABA structure possible.
[00:23:27.380 - 00:23:40.380] SPEAKER_00: If regulatory will, would change that maybe there, there's possibilities, but then we have to look maybe more at the, at film, film grates rather than, than, than bottle, uh, bottle grates.
[00:23:40.380 - 00:23:41.380] SPEAKER_00: For sure.
[00:23:41.380 - 00:23:45.380] SPEAKER_00: Um, speaking about regulation, we had a bit of similar case.
[00:23:45.380 - 00:24:02.380] SPEAKER_00: So when we are looking, um, at the usage of all consumables within our group, which are, uh, used for, uh, with, with, with virgin material, we had a big, uh, consumption of, uh, medical bins that we are using in the hospitals.
[00:24:02.380 - 00:24:03.380] SPEAKER_00: Hmm.
[00:24:03.380 - 00:24:04.380] SPEAKER_02: Okay.
[00:24:04.380 - 00:24:06.380] SPEAKER_00: So these are PPE bins.
[00:24:06.380 - 00:24:13.380] SPEAKER_00: This was also not possible to use from, um, with, with, with recycled material because of regulatory affairs.
[00:24:13.380 - 00:24:22.380] SPEAKER_00: But we worked for two years on that, uh, project and we, um, we managed to get an exception, uh, on that.
[00:24:22.380 - 00:24:29.380] SPEAKER_00: Uh, why it was difficult because we are in ADR, uh, ADR transport, that's transport of dangerous goods.
[00:24:29.380 - 00:24:38.380] SPEAKER_00: And so in that regulatory, it was mentioned that we could only use virgin material or, uh, material coming from the same application.
[00:24:38.380 - 00:24:41.380] SPEAKER_00: A bit like the case you are, uh, mentioning.
[00:24:41.380 - 00:24:46.380] SPEAKER_00: Um, which state to censor doesn't really make sense as long.
[00:24:46.380 - 00:24:56.380] SPEAKER_00: If you could prove that you can fulfill the application with other origins, it's, uh, in our opinion, it could be used.
[00:24:56.380 - 00:25:08.380] SPEAKER_00: And so in this way, you went into discussion with, uh, European regulatory affairs together with, and the local, uh, legislation here in Belgium before, for transport of dangerous goods.
[00:25:08.380 - 00:25:16.380] SPEAKER_00: Um, and we managed to make, uh, to make boxes, which also had the same technical properties like were mentioned.
[00:25:16.380 - 00:25:26.380] SPEAKER_00: So also a high impact that the boxes don't break when they could fall off a truck and dangerous material from hospitals could be, could be, could be spilled.
[00:25:26.380 - 00:25:33.380] SPEAKER_00: So, um, I don't see a solution within the upcoming short periods to be clear.
[00:25:33.380 - 00:25:42.380] SPEAKER_00: Um, but it would be good for us to, if you could send us over the current regulation, which is in place, uh, which that's a bit the, the norm.
[00:25:42.380 - 00:25:47.380] SPEAKER_00: And next to that, maybe, uh, have a bit, an idea.
[00:25:47.380 - 00:25:54.380] SPEAKER_00: You mentioned some values about MFR, allegation net break, um, the, uh, stress crack resistance.
[00:25:54.380 - 00:25:57.380] SPEAKER_00: If we could have the values values, we can save them.
[00:25:57.380 - 00:26:03.380] SPEAKER_00: And then who knows within two or three years, if regulatory changes that, that we are well prepared,
[00:26:03.380 - 00:26:11.380] SPEAKER_00: that maybe also that part of our consumables within our group, we could replace a part with virgin material, uh, with recycled material, um,
[00:26:11.380 - 00:26:20.380] SPEAKER_00: which, which, which could lower significantly or, uh, or, um, yeah, a carbon footprint from our, from our group.
[00:26:20.380 - 00:26:28.380] SPEAKER_01: I can, uh, I can say that we are part of a research program, which is coming to its end, which has exactly that goal.
[00:26:28.380 - 00:26:38.380] SPEAKER_01: So this, it is the, the goal of this project, this pro geo UP, whenever you Google it, you will probably find some information about that.
[00:26:38.380 - 00:26:44.380] SPEAKER_01: But even though I can share the link, there's a, there's a page where it's roughly describing what we are doing.
[00:26:44.380 - 00:26:58.380] SPEAKER_01: Um, this is, this is especially the idea, especially behind that part we are doing is recycling of geo texts or geo composites, geo textiles, um, grids on whatsoever.
[00:26:58.380 - 00:27:09.380] SPEAKER_01: Um, the, yeah, the goal or the, the attempt there is more in the way, like you just described for these medical, medical bins.
[00:27:09.380 - 00:27:17.380] SPEAKER_01: So taking old geo textiles or geo synthetics and reconvert them to geo geo synthetics.
[00:27:17.380 - 00:27:24.380] SPEAKER_01: This has some advantages because you're using material, which is proper for the, um, for the task.
[00:27:25.380 - 00:27:37.380] SPEAKER_01: The, the problem with the post consumer material is that it in most of the cases is not proper for the geo synthetics case.
[00:27:37.380 - 00:27:46.380] SPEAKER_01: So that, that's a problem with, when you do the recycling within the same, uh, kind of material, let's say we already proved that this is technically possible.
[00:27:46.380 - 00:27:54.380] SPEAKER_01: Uh, we just proved it for a geo grid and we can for sure show it for a geo membrane on a landfill as well.
[00:27:54.380 - 00:28:11.380] SPEAKER_01: So if you, if you, if you would take a geo membrane from the landfill and give it to you and you do the processing, like a rough cleaning and, um, melting and regranulate, we can probably do a geo membrane, which fulfills the task.
[00:28:11.380 - 00:28:17.380] SPEAKER_01: And which technically probably meets the requirements because it's the same material.
[00:28:17.380 - 00:28:21.380] SPEAKER_01: The aging is usually negligible for that time.
[00:28:21.380 - 00:28:31.380] SPEAKER_01: Um, the problem already, uh, always occurs when you're mixing in material from other sources meant for other things.
[00:28:31.380 - 00:28:39.380] SPEAKER_01: So, um, as well, the, the PP content, as you mentioned, the polypropylene content is probably technically not a big issue.
[00:28:40.380 - 00:28:47.380] SPEAKER_01: Um, it's just because for the landfill application, that's most landfill, uh, and a hazardous control.
[00:28:47.380 - 00:28:55.380] SPEAKER_01: This are the 99% of the cases where you use a geo membrane, they require 100% polyethylene.
[00:28:55.380 - 00:28:56.380] SPEAKER_UNASSIGNED: Um, so that's it.
[00:28:56.380 - 00:29:08.380] SPEAKER_01: So even if we can, for the moment, even if we say, okay, yeah, um, we know that this is not an issue and so on, it's just not there.
[00:29:08.380 - 00:29:16.380] SPEAKER_01: Um, but my colleague Henning Ehrenberg is part of the, um, regulation committee at the European Union.
[00:29:16.380 - 00:29:27.380] SPEAKER_01: And from that I know for sure that they are working on changing the regulations for geosynthetics to make it possible to use recycled materials.
[00:29:27.380 - 00:29:39.380] SPEAKER_01: And I can already say that the outcome will be something like, okay, for the recycled materials, the same things apply then for virgin materials.
[00:29:39.380 - 00:29:46.380] SPEAKER_01: So there will be no difference between virgin and recycled as long as you fulfill the given requirements.
[00:29:46.380 - 00:29:49.380] SPEAKER_01: It doesn't matter if it is from recycled or from virgin.
[00:29:49.380 - 00:30:02.380] SPEAKER_01: Um, there will probably not be the, it will probably not be the case that there will be lower requirements for recycled sources for the same purpose.
[00:30:02.380 - 00:30:14.380] SPEAKER_01: So there is always the opportunity to say, okay, we define a special purpose where we, where everyone is aware of using recycled materials with lower, um, properties.
[00:30:14.380 - 00:30:16.380] SPEAKER_01: And that's, that's always possible.
[00:30:16.380 - 00:30:23.380] SPEAKER_01: But if you want to join the, the open market, you have to prove that you can deliver the same quality.
[00:30:23.380 - 00:30:38.380] SPEAKER_01: And there is already the discussion how we can, or we can come to a situation, which is like the virgin materials where you don't have to test each particular lot,
[00:30:38.380 - 00:30:48.380] SPEAKER_01: but as well put some, um, um, QC requirements on you as the, uh, uh, manufacturer or the producer of the RC material.
[00:30:48.380 - 00:31:04.380] SPEAKER_01: So that you with given, with given tests or given audits or whatsoever that everyone can rely on that, that the raw material or the recycled material has a constant quality, whatever this quality is.
[00:31:04.380 - 00:31:14.380] SPEAKER_01: But then it can be assured that the quality is constant and it's not necessary anymore to test every delivery literally.
[00:31:14.380 - 00:31:26.380] SPEAKER_01: So even, even today, so maybe that wasn't clear today, I can do a proper geosynthetics out of recycled material.
[00:31:26.380 - 00:31:27.380] SPEAKER_01: It's just not feasible.
[00:31:27.380 - 00:31:37.380] SPEAKER_01: I, I'm allowed to do that, but the amount of work and, and checking I have to put in that is just not worth doing it.
[00:31:37.380 - 00:31:47.380] SPEAKER_01: I can, if, if I get one delivery from your side and we do all the testing, we wait for four months to get the results for the aging and so on.
[00:31:47.380 - 00:31:55.380] SPEAKER_01: Then with this lot alone, I can do a geosynthetic production, whatever it is, geomembrane or whatsoever.
[00:31:55.380 - 00:31:59.380] SPEAKER_01: And I can put the, the stamp on it and say, okay, yeah, that's, that's okay.
[00:31:59.380 - 00:32:06.380] SPEAKER_01: That's fulfilling all the requirements, but that's not feasible because I can't do that in scale.
[00:32:06.380 - 00:32:07.380] SPEAKER_01: Yeah.
[00:32:07.380 - 00:32:08.380] SPEAKER_01: This is for one truck.
[00:32:08.380 - 00:32:09.380] SPEAKER_01: That's okay.
[00:32:09.380 - 00:32:14.380] SPEAKER_01: But not for, not for a hundred tons or, or 500 tons or so.
[00:32:14.380 - 00:32:18.380] SPEAKER_01: So it's not like it is, it's not like it's forbidden.
[00:32:18.380 - 00:32:28.380] SPEAKER_01: It's just not, it's, it's not possible by the amount of effort you have to put in it to make it work in scale.
[00:32:28.380 - 00:32:29.380] SPEAKER_01: Yeah.
[00:32:29.380 - 00:32:37.380] SPEAKER_01: And that's, that's even worse actually, because the, the answer of everyone is yes, you can do it.
[00:32:37.380 - 00:32:39.380] SPEAKER_01: You just have to do this, this and that.
[00:32:39.380 - 00:32:48.380] SPEAKER_01: Um, unfortunately doing that is making you unflexible because it's taking time.
[00:32:48.380 - 00:32:57.380] SPEAKER_01: We, we would need you or we, some one of us would have to have big silos to store the material,
[00:32:57.380 - 00:33:01.380] SPEAKER_01: which is under testing for the time being.
[00:33:01.380 - 00:33:03.380] SPEAKER_01: And what if it fails?
[00:33:03.380 - 00:33:04.380] SPEAKER_01: Huh?
[00:33:04.380 - 00:33:05.380] SPEAKER_01: No one knows.
[00:33:05.380 - 00:33:16.380] SPEAKER_01: And it is, um, relatively high costs for each of these tests, which even increases the price of the raw material.
[00:33:16.380 - 00:33:22.380] SPEAKER_01: So it's actually not like I can give you a paragraph where it says it's forbidden.
[00:33:22.380 - 00:33:29.380] SPEAKER_01: It, the requirements are just that stupid on RC materials or the, the requirements are easy on version materials.
[00:33:29.380 - 00:33:33.380] SPEAKER_01: And they are applied on RC materials in the same way.
[00:33:33.380 - 00:33:36.380] SPEAKER_01: And this makes it practically impossible.
[00:33:36.380 - 00:33:37.380] SPEAKER_01: Yeah.
[00:33:37.380 - 00:33:38.380] SPEAKER_UNASSIGNED: Yeah.
[00:33:38.380 - 00:33:40.380] SPEAKER_00: I understand, uh, very well.
[00:33:40.380 - 00:33:45.380] SPEAKER_00: And unfortunately, um, this is still the world of today.
[00:33:45.380 - 00:33:54.380] SPEAKER_00: And, um, luckily for instance, in packaging, we have some new regulations, uh, which, which obligate, obligates in 2030, a certain percentage of, of usage.
[00:33:54.380 - 00:34:00.380] SPEAKER_00: And this helps of course, to create also the market because it's not a matter of willing.
[00:34:00.380 - 00:34:01.380] SPEAKER_00: It's, it's a fact.
[00:34:01.380 - 00:34:07.380] SPEAKER_00: And as long as there, it's not there, I understand in other markets, it's quite, uh, quite difficult.
[00:34:07.380 - 00:34:13.380] SPEAKER_00: Nevertheless, um, with our plant, we have some collaborations running with some, um, more and more.
[00:34:13.380 - 00:34:24.380] SPEAKER_00: We have virgin polymer producers who are integrating recycled contents to give a ready to use 30 percentage, uh, range to, to converters.
[00:34:24.380 - 00:34:32.380] SPEAKER_00: So they, they, so that they, so that they don't need to worry about all those, uh, things that you are, uh, mentioning.
[00:34:32.380 - 00:34:40.380] SPEAKER_00: So this goes from, uh, Total Energies, uh, Borealis, uh, also Ravago with the manufacturing that we're having.
[00:34:40.380 - 00:34:47.380] SPEAKER_01: Um, I'm pretty sure it will not take too long to, that this is established.
[00:34:47.380 - 00:34:48.380] SPEAKER_01: Yeah.
[00:34:48.380 - 00:34:58.380] SPEAKER_01: But you need, as you say, you need one of the big players, um, taking the responsibility and doing the paperwork, let's say.
[00:34:58.380 - 00:35:07.380] SPEAKER_01: So yeah, as soon as they can proclaim, okay, we take care that what's leaving our plant has a constant quality.
[00:35:07.380 - 00:35:13.380] SPEAKER_01: So whatever this is, constant quality, and we can ensure that this is always the case.
[00:35:13.380 - 00:35:23.380] SPEAKER_01: Then we, as the mid user, um, don't have any problem as long as the result of the product is feasible.
[00:35:23.380 - 00:35:24.380] SPEAKER_01: Huh?
[00:35:24.380 - 00:35:29.380] SPEAKER_01: So, but nowadays, um, we did some test runs, not on geomembranes, but on other products.
[00:35:29.380 - 00:35:33.380] SPEAKER_01: We did some test runs with chemical recycled material.
[00:35:33.380 - 00:35:36.380] SPEAKER_01: So there is a, I think it was a polyester in that case.
[00:35:36.380 - 00:35:47.380] SPEAKER_01: They added during the, they added deep polymerized, um, monomers of recycled materials to the polymerization process.
[00:35:47.380 - 00:35:50.380] SPEAKER_01: So at the end, it is like virgin material.
[00:35:50.380 - 00:35:55.380] SPEAKER_01: Just, just the, the, the atoms and the molecules have been reused.
[00:35:55.380 - 00:35:55.380] SPEAKER_UNASSIGNED: Yeah.
[00:35:55.380 - 00:35:56.380] SPEAKER_01: It is.
[00:35:56.380 - 00:36:02.380] SPEAKER_01: And even that gives us trouble in the explanation or in, yeah.
[00:36:02.380 - 00:36:07.380] SPEAKER_01: It's technically, you can't divide it from virgin material.
[00:36:07.380 - 00:36:08.380] SPEAKER_00: Yeah.
[00:36:08.380 - 00:36:09.380] SPEAKER_00: I see.
[00:36:09.380 - 00:36:15.380] SPEAKER_00: But that's, that's, that shows the difficulty here because with recycled material, this is just virgin material with the certificate.
[00:36:15.380 - 00:36:16.380] SPEAKER_00: Yeah.
[00:36:16.380 - 00:36:19.380] SPEAKER_00: Because at the end, it's just the same virgin material.
[00:36:19.380 - 00:36:20.380] SPEAKER_00: So, uh, I understand.
[00:36:20.380 - 00:36:28.380] SPEAKER_00: Um, just for, for us, it could, it would be good, uh, if we can get some kind of idea with the, for the geo membranes.
[00:36:28.380 - 00:36:30.380] SPEAKER_00: Uh, uh, uh, okay.
[00:36:30.380 - 00:36:38.380] SPEAKER_00: The virgin material you are using not to make a copy paste, but to have a bit an idea about those allegation and breaks and, uh, and other things.
[00:36:38.380 - 00:36:48.380] SPEAKER_00: And then from our sides, I will just, with our, um, partners that we are making hybrids, virgin and recycle, we'll just, uh, exchange some thoughts.
[00:36:48.380 - 00:36:57.380] SPEAKER_00: Uh, if there would be no regulatory boundary, uh, what is feasible and what is not feasible because for the virgin by parts, they're using some kind of, yeah, we call it boosters.
[00:36:57.380 - 00:37:07.380] SPEAKER_00: It's, it's overqualified virgin material to compensate for, for, uh, yeah, the lack of parameters and, and, and some, some, some RC material.
[00:37:07.380 - 00:37:09.380] SPEAKER_00: This has to be research.
[00:37:09.380 - 00:37:23.380] SPEAKER_01: We have done some, unfortunately mixing of materials, especially of high density and low density material, give some results, which are, which we didn't expect and which were unpredictable for us.
[00:37:23.380 - 00:37:33.380] SPEAKER_01: So we, in our testing, we added very high density material in very low percentage to our typical resin.
[00:37:33.380 - 00:37:43.380] SPEAKER_01: And for example, the stress crack resistance, even with 2% high density material dropped from usually 3000 hours plus.
[00:37:43.380 - 00:37:55.380] SPEAKER_01: So our longest test we stopped at 3000 hours for that test dropped to less than 50 by just adding one or 2% of HD, HD material.
[00:37:55.380 - 00:38:08.380] SPEAKER_01: So it's not like mixing milk and cacao, for example, it is, there is something happening, which is not like they don't behave, um, in relation to the mixture ratio.
[00:38:08.380 - 00:38:11.380] SPEAKER_01: So this is a very interesting field.
[00:38:11.380 - 00:38:14.380] SPEAKER_01: So, um, but, um, what I can mention is yes.
[00:38:14.380 - 00:38:22.380] SPEAKER_01: Having a ready-made material, which is already, let's say 20% recycled together with virgin, with a certificate is always welcome.
[00:38:22.380 - 00:38:27.380] SPEAKER_01: From the technical point of view, we can also mix in our extruders.
[00:38:27.380 - 00:38:31.380] SPEAKER_01: We don't have this double screw mixing extruders.
[00:38:31.380 - 00:38:35.380] SPEAKER_01: We have a single screw extruders, but with, um, barriers and so on.
[00:38:35.380 - 00:38:40.380] SPEAKER_01: So we do regular mixing of same kind of materials.
[00:38:40.380 - 00:38:45.380] SPEAKER_01: So different raw materials in the same area of MFR and so on.
[00:38:45.380 - 00:38:48.380] SPEAKER_01: So we do regular mixing on that.
[00:38:48.380 - 00:38:52.380] SPEAKER_01: Technically, as I said, technically is not the biggest difficulty.
[00:38:52.380 - 00:38:55.380] SPEAKER_01: It is like how it behaves afterwards.
[00:38:55.380 - 00:38:57.380] SPEAKER_01: So is it fit for the purpose?
[00:38:57.380 - 00:39:01.380] SPEAKER_01: And the big thing is regulation things.
[00:39:01.380 - 00:39:02.380] SPEAKER_01: Regulation.
[00:39:02.380 - 00:39:03.380] SPEAKER_03: Yeah.
[00:39:03.380 - 00:39:04.380] SPEAKER_01: Yeah.
[00:39:04.380 - 00:39:13.380] SPEAKER_01: And, um, if you're interested, we can do some quick, quick tests, quick research eventually.
[00:39:13.380 - 00:39:18.380] SPEAKER_01: So every now and then we do some tests on a lab extruder scale.
[00:39:18.380 - 00:39:28.380] SPEAKER_01: So just 30 centimeters wide, a couple of kilograms, um, to get an idea what we can expect from this material.
[00:39:28.380 - 00:39:44.380] SPEAKER_01: Uh, we can incorporate in our next test run, maybe a mixture of 20% of your RC material together with 80% of our, uh, virgin material or even 100% just to get an idea what we would end up with.
[00:39:44.380 - 00:39:46.380] SPEAKER_01: Yeah.
[00:39:46.380 - 00:39:50.380] SPEAKER_01: So then we were talking about 20, let's say 50 kilograms.
[00:39:50.380 - 00:39:51.380] SPEAKER_00: Yeah.
[00:39:51.380 - 00:40:05.380] SPEAKER_00: Well, honestly, I think the only thing what would make sense is, um, is a hybrid product where already 20% is integrated and made by, uh, uh, Total or Borealis or Avago.
[00:40:05.380 - 00:40:09.380] SPEAKER_00: That's, that's three main or Bazel, Bazel could do it as well.
[00:40:09.380 - 00:40:11.380] SPEAKER_00: It's the main players we're working with together.
[00:40:11.380 - 00:40:21.380] SPEAKER_00: Um, uh, 100% pure RC from our sites will never make the specification because it's HD based.
[00:40:21.380 - 00:40:30.380] SPEAKER_00: I'm now more thinking about, uh, uh, postcon, because if you want to do it at scale, you need post-consumer material, which is available everywhere in Europe.
[00:40:30.380 - 00:40:33.380] SPEAKER_00: Otherwise you will never manage to make, to do something on that scale.
[00:40:33.380 - 00:40:34.380] SPEAKER_00: That's the second point.
[00:40:34.380 - 00:40:35.380] SPEAKER_00: Yeah.
[00:40:35.380 - 00:40:43.380] SPEAKER_00: So this means if, if, if, if from our bottles are everywhere and packaging and, but I think it's will be very hard or impossible.
[00:40:43.380 - 00:40:46.380] SPEAKER_00: The films are also everywhere in, um, in Europe.
[00:40:46.380 - 00:40:58.380] SPEAKER_00: So I think we mainly have to look at an, an LD, LLD mixtures from, um, from films, um, to manage the elongation and breaks and this kinds of parameters.
[00:40:58.380 - 00:41:01.380] SPEAKER_00: But, uh, that's, that's something I want to discuss with them first.
[00:41:01.380 - 00:41:02.380] SPEAKER_00: Okay.
[00:41:02.380 - 00:41:12.380] SPEAKER_00: What type of origin could be usable combined with your boosters, uh, from, from Virgin to end up with something which, which comes in the range with the current spec you are using.
[00:41:12.380 - 00:41:19.380] SPEAKER_00: And, uh, and of course, then I want to challenge them as well with, uh, with the regulations and what kind of, uh, things are asked.
[00:41:19.380 - 00:41:32.380] SPEAKER_00: And, uh, they have also people, people working monthly on that because they also want to open markets for high quality recycled material where currently, um, no recycled is being used.
[00:41:32.380 - 00:41:33.380] SPEAKER_01: Yeah.
[00:41:33.380 - 00:41:35.380] SPEAKER_01: That's a, that's the availability is a big point.
[00:41:35.380 - 00:41:44.380] SPEAKER_01: So for our tests of recycling geosynthetics, um, so make out of a, use geosynthetics, make new ones.
[00:41:44.380 - 00:41:49.380] SPEAKER_01: One of the biggest struggles besides all the techniques is the availability.
[00:41:49.380 - 00:41:57.380] SPEAKER_01: So geosynthetics, they are made to last at least 25 years in constructions, some 100 years.
[00:41:57.380 - 00:42:00.380] SPEAKER_01: So they are just not available in numbers on the market.
[00:42:00.380 - 00:42:04.380] SPEAKER_01: There are three tons there, five tons there.
[00:42:06.380 - 00:42:09.380] SPEAKER_00: Man, you can, you will never manage to do something at scale.
[00:42:09.380 - 00:42:10.380] SPEAKER_01: No, no.
[00:42:10.380 - 00:42:11.380] SPEAKER_01: Yeah.
[00:42:11.380 - 00:42:15.380] SPEAKER_00: Of course, consumer is household packaging material, uh, commercial films.
[00:42:15.380 - 00:42:24.380] SPEAKER_00: Um, that's the main markets you will find, which you will found everywhere in every country in Europe, where at least there's some collection.
[00:42:24.380 - 00:42:31.380] SPEAKER_00: So, uh, so if you want to do something at scale, you have to, this, this has to be the basis from, from the recycled content.
[00:42:31.380 - 00:42:32.380] SPEAKER_02: Yeah.
[00:42:32.380 - 00:42:35.380] SPEAKER_02: So scaling is probably the right word.
[00:42:35.380 - 00:42:48.380] SPEAKER_02: Um, we have now learned there's one big column, uh, this part of discussion, which would be a win-win situation in the best case, maybe in a few, in a couple of years.
[00:42:48.380 - 00:42:49.380] SPEAKER_02: Yeah.
[00:42:49.380 - 00:43:00.380] SPEAKER_02: Um, the other part that is, uh, where I actually don't know, is this topic still, uh, of interest would be the column with circular economy.
[00:43:00.380 - 00:43:17.380] SPEAKER_02: Um, please, would this be an option for you to make a reuse of your HDPE for circular economy reasons to build a story out of that?
[00:43:17.380 - 00:43:21.380] SPEAKER_02: Or is this not of interest?
[00:43:21.380 - 00:43:28.380] SPEAKER_03: The one of which we recycle, uh, uh, recycler are temporary covering, you mean?
[00:43:28.380 - 00:43:29.380] SPEAKER_03: Yeah.
[00:43:29.380 - 00:43:34.380] SPEAKER_03: I think at this, at this stage would be very difficult, uh, to do so.
[00:43:34.380 - 00:43:36.380] SPEAKER_03: But I, I think Jeroen, you have a better view on it.
[00:43:36.380 - 00:43:37.380] SPEAKER_03: Yeah.
[00:43:37.380 - 00:43:40.380] SPEAKER_03: But I think in terms of, yeah, feasibility.
[00:43:40.380 - 00:43:43.380] SPEAKER_00: On, on our plants, it's, it's not possible.
[00:43:43.380 - 00:43:47.380] SPEAKER_00: Our plant is built for post-consumer, post-consumer households, uh, waste.
[00:43:47.380 - 00:43:50.380] SPEAKER_00: And so our inputs are bells of sorted bottles.
[00:43:50.380 - 00:43:55.380] SPEAKER_00: So, um, it could be very, the same thing you are mentioning, eh?
[00:43:55.380 - 00:44:01.380] SPEAKER_00: We have to, we would have to do very small lots of, of those temporary film material.
[00:44:01.380 - 00:44:08.380] SPEAKER_00: We should, we would have to do external size reduction or baling or whatever to precondition it so we can feed it into our plant.
[00:44:08.380 - 00:44:17.380] SPEAKER_00: So, I think it's very difficult to, to get that to, um, to, uh, at scale level, uh, at, at our recycling plants.
[00:44:17.380 - 00:44:24.380] SPEAKER_00: Um, we have such types of films sometimes that we collect, that we deliver to, uh, to a local recycle,
[00:44:24.380 - 00:44:30.380] SPEAKER_00: which is using it in very low-end application because getting it to a recycling line is one.
[00:44:30.380 - 00:44:35.380] SPEAKER_00: But getting it to a level where it can be reused for high-end applications is another story.
[00:44:35.380 - 00:44:40.380] SPEAKER_00: I mean, our washing line, we, uh, we built that for a preparation for a PPWR.
[00:44:40.380 - 00:44:43.380] SPEAKER_00: So that's for a new packaging regulation.
[00:44:43.380 - 00:44:45.380] SPEAKER_00: So we have four washing steps.
[00:44:45.380 - 00:44:50.380] SPEAKER_00: We can, uh, we can, um, increase the temperature of, uh, of one of the washing steps.
[00:44:50.380 - 00:44:57.380] SPEAKER_00: We asked caustic soda, we, to prepare as much as possible to get it to higher-end applications
[00:44:57.380 - 00:45:00.380] SPEAKER_00: and packaging and color sorting, all kinds of, of those things.
[00:45:00.380 - 00:45:09.380] SPEAKER_00: And so, um, yeah, at this point, if it goes to lower-end applications, uh, uh, then it's,
[00:45:09.380 - 00:45:12.380] SPEAKER_00: or a plant is not the right plant to, to, to put it through.
[00:45:12.380 - 00:45:18.380] SPEAKER_00: So the circular economy ID, uh, also, I think it's about lower quantities.
[00:45:18.380 - 00:45:21.380] SPEAKER_00: And I don't know the quantities, uh, which, which are coming out of that.
[00:45:21.380 - 00:45:26.380] SPEAKER_00: But, um, my opinion, that's very difficult to do that at scale.
[00:45:26.380 - 00:45:28.380] SPEAKER_00: Getting it recycled is one thing.
[00:45:28.380 - 00:45:30.380] SPEAKER_00: And then to lower-end application.
[00:45:30.380 - 00:45:32.380] SPEAKER_00: And then we should aim as much as possible for.
[00:45:32.380 - 00:45:34.380] SPEAKER_00: So that's not a discussion.
[00:45:34.380 - 00:45:37.380] SPEAKER_00: But to set something up, to reuse it back into your applications,
[00:45:37.380 - 00:45:43.380] SPEAKER_00: it has to come from post-consumer, uh, higher availability streams, um,
[00:45:43.380 - 00:45:48.380] SPEAKER_00: um, and then do the right compounding steps to get the ready-to-use, uh, polymer
[00:45:48.380 - 00:45:53.380] SPEAKER_00: without any stress at converting site with the right, uh, certificates and ever.
[00:45:53.380 - 00:46:01.380] SPEAKER_00: And, and then that's the only way I see to, to get recycled content into your products, to be honest.
[00:46:01.380 - 00:46:02.380] SPEAKER_UNASSIGNED: Yeah.
[00:46:02.380 - 00:46:05.380] SPEAKER_03: Is it a question, Martin and Christian?
[00:46:05.380 - 00:46:12.380] SPEAKER_03: Because we're now at Van Heede, we're questioning if a recycled substitute is available.
[00:46:12.380 - 00:46:15.380] SPEAKER_03: Is it also, and you say, yeah, the market is small.
[00:46:15.380 - 00:46:20.380] SPEAKER_03: Have you ever, let's say, gotten the interest or the question from other parties?
[00:46:20.380 - 00:46:24.380] SPEAKER_03: Because we're looking at it for, let's say, a very narrow scope.
[00:46:24.380 - 00:46:29.380] SPEAKER_03: Let's say temporary cover, smaller than, or used less than five years.
[00:46:29.380 - 00:46:36.380] SPEAKER_03: But I don't know if in your market you've had the question before on, hey, do you also offer a version,
[00:46:36.380 - 00:46:40.380] SPEAKER_03: a version with recycled material?
[00:46:40.380 - 00:46:49.380] SPEAKER_02: Well, in this, in this application, uh, that we focus on here, uh, what you just heard from Martin's explanation,
[00:46:49.380 - 00:46:56.380] SPEAKER_02: we are, we are always forced, uh, to generate solutions that are very long lasting, minimum 100 years.
[00:46:56.380 - 00:47:04.380] SPEAKER_02: Um, this, this is the highest argument usually for, for our customers to be, have the safety that the constructions
[00:47:04.380 - 00:47:06.380] SPEAKER_02: they are building last for 100 years.
[00:47:06.380 - 00:47:16.380] SPEAKER_02: And if you can give, uh, as well, some, uh, some, yeah, research results that can have an outlook on the time after,
[00:47:16.380 - 00:47:22.380] SPEAKER_02: yeah, where we are actually thinking, okay, our solutions will also survive 400 years.
[00:47:22.380 - 00:47:26.380] SPEAKER_02: That's the biggest argument for the moment.
[00:47:26.380 - 00:47:27.380] SPEAKER_02: Yeah.
[00:47:27.380 - 00:47:31.380] SPEAKER_02: Everything what we are discussing here now is, is of course, I would say part of the future.
[00:47:31.380 - 00:47:31.380] SPEAKER_UNASSIGNED: Yeah.
[00:47:31.380 - 00:47:49.380] SPEAKER_02: That you also think about how can we bring in recycled material for those solutions, but, uh, in the, in the same moment, uh, staying on the same, on the safe side and, uh, making something out of it, uh, that is really protecting also the third and the fourth generation after us.
[00:47:49.380 - 00:47:50.380] SPEAKER_02: Yeah.
[00:47:50.380 - 00:48:16.380] SPEAKER_01: That's one of the, what Christian describes is the, this geomembranes are one of the most sensitive, um, applications of geosynthetics because you're containing biohazards and really no one is willing to take any risk on that.
[00:48:16.380 - 00:48:25.380] SPEAKER_01: Uh, if you're talking about non-wovens, for example, which act as a filter layer or something like that, the risk of failure is much less.
[00:48:25.380 - 00:48:31.380] SPEAKER_01: So people, uh, so there, it would be much easier to convince people to take some risk.
[00:48:31.380 - 00:48:38.380] SPEAKER_01: Even if we are convinced there is none, but there is still these, uh, it is lower quality.
[00:48:38.380 - 00:48:48.380] SPEAKER_01: Um, then the outcome of, of failure would be probably not even noticeable in, in most of the cases, but for geomembranes is really, it is waste containment.
[00:48:48.380 - 00:48:52.380] SPEAKER_01: No one wants to even think about that.
[00:48:52.380 - 00:48:57.380] SPEAKER_01: There is a hole or a crack or, um, uh, a bad welding.
[00:48:57.380 - 00:49:06.380] SPEAKER_01: Um, if I can have the better, uh, material in, in air quotes for the same price.
[00:49:06.380 - 00:49:22.380] SPEAKER_01: And that's actually, if, if you would be half the price with the recycling, then probably some would think for some applications about it, but just think it from, from the, from your own perspective.
[00:49:22.380 - 00:49:36.380] SPEAKER_01: Um, you're buying a material with lower quality or lower, um, promise of quality for a temporary job, which is also more expensive.
[00:49:37.380 - 00:49:38.380] SPEAKER_01: Yeah.
[00:49:38.380 - 00:49:48.380] SPEAKER_01: No one there's, there's just a very low, uh, amount of people who is willing to pay extra for these carbon footprint.
[00:49:48.380 - 00:49:53.380] SPEAKER_01: Most will go with the lowest price, unfortunately.
[00:49:53.380 - 00:50:00.380] SPEAKER_01: So there are, there are parts in geosynthetics, which are probably a little bit easier from a technical point of view and where the risk is much lower.
[00:50:00.380 - 00:50:06.380] SPEAKER_01: And even those are not for, for the moment, they're not accepting the extra costs.
[00:50:08.380 - 00:50:15.380] SPEAKER_01: Because in the, in the, in the mind of the people, usually recycling is combined with lower quality and lower costs.
[00:50:18.380 - 00:50:22.380] SPEAKER_01: Um, in the Netherlands, you're, you're, you're Belgium, right?
[00:50:23.380 - 00:50:23.380] SPEAKER_UNASSIGNED: Belgium.
[00:50:23.380 - 00:50:23.380] SPEAKER_UNASSIGNED: Yeah.
[00:50:23.380 - 00:50:24.380] SPEAKER_01: Yeah.
[00:50:24.380 - 00:50:39.380] SPEAKER_01: In the Netherlands, I hear every now and then about RC material to be used, but even still with the same requirements, which is for the moment, just impossible.
[00:50:41.380 - 00:50:42.380] SPEAKER_01: Okay.
[00:50:42.380 - 00:50:45.380] SPEAKER_01: But they're more on the, so recycled.
[00:50:45.380 - 00:50:46.380] SPEAKER_01: Yes.
[00:50:46.380 - 00:50:51.380] SPEAKER_01: But they're more on interestingly bio based standard polymers.
[00:50:51.380 - 00:51:01.380] SPEAKER_01: So a bio based polyester or bio based polyethylene, um, which at the end, which at the end is the same as virgin material.
[00:51:01.380 - 00:51:02.380] SPEAKER_01: Yeah.
[00:51:02.380 - 00:51:03.380] SPEAKER_01: Yeah.
[00:51:03.380 - 00:51:04.380] SPEAKER_01: Yeah.
[00:51:04.380 - 00:51:05.380] SPEAKER_01: Yeah.
[00:51:05.380 - 00:51:06.380] SPEAKER_01: It's the product itself.
[00:51:06.380 - 00:51:09.380] SPEAKER_01: You can't tell if it's from, from corn or from, from crude oil.
[00:51:09.380 - 00:51:10.380] SPEAKER_01: So.
[00:51:10.380 - 00:51:11.380] SPEAKER_UNASSIGNED: Yeah.
[00:51:11.380 - 00:51:12.380] SPEAKER_UNASSIGNED: Yeah.
[00:51:12.380 - 00:51:13.380] SPEAKER_UNASSIGNED: Okay.
[00:51:13.380 - 00:51:19.380] SPEAKER_00: But for me, it was interesting because geo membranes is for me a new application of recycled content.
[00:51:19.380 - 00:51:21.380] SPEAKER_00: So I learned already a lot.
[00:51:21.380 - 00:51:42.380] SPEAKER_00: At this point I want to dig a bit with my polymer producers into the technical demands and then the norms of itself.
[00:51:42.380 - 00:51:54.380] SPEAKER_00: And then who knows, within two, three, four years, because the world is spinning very, very fast around and at least we are prepared.
[00:51:54.380 - 00:52:02.380] SPEAKER_00: And yeah, for me that was a good first intro session and who knows where we will be in a couple of years.
[00:52:02.380 - 00:52:06.380] SPEAKER_01: As you said, just think about the bottles, the polyester bottles.
[00:52:06.380 - 00:52:14.380] SPEAKER_01: Maybe there will be some rule like you have to use 30% of recycled materials, then the whole game changes.
[00:52:14.380 - 00:52:16.380] SPEAKER_03: Yeah, yeah.
[00:52:18.380 - 00:52:21.380] SPEAKER_03: Are there also other competitors already working on this?
[00:52:21.380 - 00:52:29.380] SPEAKER_03: And let's say as a vision we want to be the first one offering a recycled solution or are you not aware of?
[00:52:29.380 - 00:52:43.380] SPEAKER_01: Not that I'm aware. So there are some, to be frank, there are some competitors who are showing example cases where they did what we as well did.
[00:52:43.380 - 00:52:56.380] SPEAKER_01: So using a geotextile or geosynthetic and recycled this particular geosynthetic in a same kind of geosynthetic.
[00:52:56.380 - 00:52:58.380] SPEAKER_01: So that's what I'm aware of.
[00:52:58.380 - 00:53:08.380] SPEAKER_01: So buying, especially consumer grade recycling materials, I'm not aware of, let's say it like that.
[00:53:08.380 - 00:53:09.380] SPEAKER_03: Okay.
[00:53:09.380 - 00:53:10.380] SPEAKER_03: We have.
[00:53:10.380 - 00:53:11.380] SPEAKER_UNASSIGNED: All right.
[00:53:11.380 - 00:53:25.380] SPEAKER_01: Typically the assumption, and maybe that's the last word, the assumption of using recycled polyester is that this might be easier.
[00:53:25.380 - 00:53:45.380] SPEAKER_01: Because as far as we have been described is that the recycling materials stream is much cleaner because it's the, if you use the bottles, the bottles are very well sorted by the consumer himself.
[00:53:45.380 - 00:54:02.380] SPEAKER_01: And the material is very consistent and so on, which is just not the case for what you just described, PP and PE, which is more like packaging yogurt bins and so on.
[00:54:02.380 - 00:54:03.380] SPEAKER_01: Yeah.
[00:54:03.380 - 00:54:12.380] SPEAKER_01: So you will probably, the first, the first appearing recycled material product will probably be polyester.
[00:54:12.380 - 00:54:13.380] SPEAKER_03: Yeah.
[00:54:13.380 - 00:54:17.380] SPEAKER_01: If I would bet some money, that would be my bet.
[00:54:17.380 - 00:54:18.380] SPEAKER_00: Yeah.
[00:54:18.380 - 00:54:26.380] SPEAKER_00: But I can align as well that the post-consumer PT bottles are, yeah, it's only from bottles.
[00:54:26.380 - 00:54:31.380] SPEAKER_00: So the, the variability of parameters are very, it's very narrow.
[00:54:31.380 - 00:54:36.380] SPEAKER_00: Um, so I can understand what you're saying, Mark.
[00:54:36.380 - 00:54:37.380] SPEAKER_00: Yeah.
[00:54:37.380 - 00:54:38.380] SPEAKER_00: Yeah.
[00:54:38.380 - 00:54:39.380] SPEAKER_01: Yeah.
[00:54:39.380 - 00:54:40.380] SPEAKER_01: It's plus and minus, huh?
[00:54:40.380 - 00:54:44.380] SPEAKER_01: So it's very narrow, but you know what you get.
[00:54:44.380 - 00:54:45.380] SPEAKER_01: Yeah.
[00:54:45.380 - 00:54:46.380] SPEAKER_00: Yeah.
[00:54:46.380 - 00:54:47.380] SPEAKER_UNASSIGNED: Yeah.
[00:54:47.380 - 00:54:48.380] SPEAKER_UNASSIGNED: Yeah.
[00:54:48.380 - 00:54:49.380] SPEAKER_UNASSIGNED: Yeah.
[00:54:49.380 - 00:54:50.380] SPEAKER_01: Yeah.
[00:54:50.380 - 00:54:51.380] SPEAKER_01: Very interesting for me as well.
[00:54:51.380 - 00:54:52.380] SPEAKER_01: Thank you very much.
[00:54:52.380 - 00:54:58.380] SPEAKER_03: No, thanks to, uh, the, the four of you, uh, for the interesting brainstorm.
[00:54:58.380 - 00:55:01.380] SPEAKER_03: Although I have to admit, I was very astonished about the knowledge.
[00:55:01.380 - 00:55:03.380] SPEAKER_03: For me, it was a bit of Chinese.
[00:55:03.380 - 00:55:08.380] SPEAKER_03: So, um, uh, but, uh, you, you explained it very well.
[00:55:08.380 - 00:55:10.380] SPEAKER_03: So, um, thanks for that.
[00:55:10.380 - 00:55:12.380] SPEAKER_03: That's when you bring in two experts.
[00:55:12.380 - 00:55:13.380] SPEAKER_03: Yeah.
[00:55:13.380 - 00:55:17.380] SPEAKER_03: And we were like, no, but, uh, it's good.
[00:55:17.380 - 00:55:23.380] SPEAKER_03: Uh, it's good to have at least a view on the market and where it would be going, uh,
[00:55:23.380 - 00:55:24.380] SPEAKER_03: within the next couple of years.
[00:55:24.380 - 00:55:29.380] SPEAKER_03: And indeed, as you said, if we're ready from our point of view or our stand, I think it's
[00:55:29.380 - 00:55:30.380] SPEAKER_03: good.
[00:55:30.380 - 00:55:34.380] SPEAKER_01: I'm pretty sure that there will be changes in the next upcoming years.
[00:55:34.380 - 00:55:39.380] SPEAKER_01: Um, so on the regulation side, on the technical side, everything is developing fast.
[00:55:41.380 - 00:55:47.380] SPEAKER_01: Um, with the situation in Middle East, maybe there will be a little more consequence in doing
[00:55:47.380 - 00:55:50.380] SPEAKER_01: whatever net is necessary for recycling.
[00:55:50.380 - 00:55:59.380] SPEAKER_01: If the oil is not that available anymore, that we are used to, or, um, in the worst, or even worse.
[00:55:59.380 - 00:56:05.380] SPEAKER_01: The, the polymers made in Middle East can't leave Middle East.
[00:56:05.380 - 00:56:12.380] SPEAKER_01: So yes, I expect that there will be some acceleration on that in the upcoming years.
[00:56:12.380 - 00:56:18.380] SPEAKER_02: It always needs financial advantages or pressure to change things.
[00:56:19.380 - 00:56:21.380] SPEAKER_02: Unfortunately, that's our world.
[00:56:21.380 - 00:56:22.380] SPEAKER_03: Yep.
[00:56:24.380 - 00:56:25.380] SPEAKER_UNASSIGNED: All right.
[00:56:25.380 - 00:56:26.380] SPEAKER_03: Okay.
[00:56:26.380 - 00:56:27.380] SPEAKER_02: We won't.
[00:56:27.380 - 00:56:28.380] SPEAKER_03: Thank you very much, everybody.
[00:56:28.380 - 00:56:29.380] SPEAKER_02: Yes.
[00:56:29.380 - 00:56:30.380] SPEAKER_02: Indeed.
[00:56:30.380 - 00:56:31.380] SPEAKER_02: Yeah.
[00:56:31.380 - 00:56:32.380] SPEAKER_02: Thanks for this interesting exchange.
[00:56:32.380 - 00:56:36.380] SPEAKER_02: Let's, uh, stay connected with each other.
[00:56:36.380 - 00:56:37.380] SPEAKER_02: Yes.
[00:56:37.380 - 00:56:42.380] SPEAKER_02: And if there's some interest, uh, to, to prove some things, yeah, please do not hesitate to
[00:56:42.380 - 00:56:43.380] SPEAKER_02: come back to us.
[00:56:44.380 - 00:56:51.380] SPEAKER_02: And, uh, yeah, Dries, I think we will stay in contact, uh, now, anyhow, regarding the, uh,
[00:56:51.380 - 00:56:52.380] SPEAKER_02: The upcoming tender.
[00:56:52.380 - 00:56:53.380] SPEAKER_02: The upcoming tenders.
[00:56:53.380 - 00:56:54.380] SPEAKER_02: Yeah.
[00:56:54.380 - 00:56:57.380] SPEAKER_02: And to see what, where, where it brings us together then.
[00:56:57.380 - 00:56:58.380] SPEAKER_02: Yes.
[00:56:58.380 - 00:56:59.380] SPEAKER_03: Yes.
[00:56:59.380 - 00:56:59.380] SPEAKER_UNASSIGNED: Indeed.
[00:56:59.380 - 00:57:06.380] SPEAKER_03: Normally end of this week, latest, beginning of next week, we will send out, uh, the tender.
[00:57:06.380 - 00:57:09.380] SPEAKER_03: So we will keep you in, uh, in copy.
[00:57:09.380 - 00:57:10.380] SPEAKER_03: All right.
[00:57:10.380 - 00:57:11.380] SPEAKER_02: Good.
[00:57:11.380 - 00:57:12.380] SPEAKER_02: All right.
[00:57:12.380 - 00:57:13.380] SPEAKER_03: Good.
[00:57:13.380 - 00:57:13.920] SPEAKER_UNASSIGNED: All right.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,575 @@
[00:00:00.000 - 00:00:07.880] Well it's a Monday. Weekend was too short as always.
[00:00:07.880 - 00:00:22.380] Everything's good. How many of your colleagues are joining? Jurek and Jeroen both coming?
[00:00:22.380 - 00:00:36.180] Normally I invite Jeroen. There he comes. I think Jurek will not join.
[00:00:36.180 - 00:00:51.160] Welcome. Hello to everybody. Jurek will not join, that was your question. I see he's in
[00:00:51.160 - 00:00:56.320] a meeting so what he said as it will probably be a technical meeting or a more
[00:00:56.320 - 00:01:02.020] technical meeting I will correct get the info from my colleagues. Yes I think the
[00:01:02.020 - 00:01:07.660] aim of today's meeting is to to get to know more to get to know more about let's
[00:01:07.660 - 00:01:14.800] say the specifications in terms of using the recycled material from Van Heden and
[00:01:14.800 - 00:01:25.660] and if we can let's say find a solution or if we can work together on this front as well since it was currently still at an early stage
[00:01:25.660 - 00:01:36.520] where we in the previous meeting discussed the potential opportunity and then Christian forwarded based upon I think it's your Martin's requirements.
[00:01:36.520 - 00:01:47.380] And then I forwarded it to Jeroen our expert at Van Heden and then he I think forwarded some additional questions.
[00:01:47.380 - 00:01:58.240] Yeah so before we start we should get something straight when Christian and me were starting
[00:01:58.240 - 00:02:05.380] talking about or when Christian told me about your ideas what to do there was I assumed that it was that it would be about a limited a limited amount of material which is on your side which you do recycle and you thought about okay we can do it.
[00:02:05.380 - 00:02:18.380] We can reuse this particular material for our purpose on our own landfills so that's what I understood the questions you sent they are all valid and this is all most of that is well known.
[00:02:18.380 - 00:02:34.380] But to me that sounded like a general offer to use recycled material to use recycled material in our production and so first of all was that was that your intention?
[00:02:34.380 - 00:02:55.380] So I think first we discussed with Christian in the last meeting if it was possible to re-use HDPE foil which we use for a temporary basis and then get back to let's say the recycling of HDPE foil.
[00:02:55.380 - 00:03:18.380] But I think this was deemed to be very difficult since it was also of course polluted and also difficult to let's say yeah you put it on the dirt you have to pull it back it's contaminated.
[00:03:18.380 - 00:03:33.380] And then Jürich I think switched his view to okay at Van Heide we also have plastics recycling which is not only linked to HDPE foil.
[00:03:33.380 - 00:03:45.380] And that was let's say the broader discussion could we be or could the products that we supply or foresee be also utilized in your production process.
[00:03:45.380 - 00:03:52.380] So this is okay if this is a general question then unfortunately no and yes.
[00:03:52.380 - 00:03:59.380] So in general I have to I have to do a little bigger circle.
[00:03:59.380 - 00:04:07.380] It is it is less on the side if we can technically process these kind of materials.
[00:04:07.380 - 00:04:16.380] It's more like we can't make any product of of that for any market we are working on.
[00:04:16.380 - 00:04:20.380] The problem is not the technical process ability.
[00:04:20.380 - 00:04:34.380] So if we follow some standards or and some let's say some requirements like MFR has to be in a certain range and so on and there has to be a minimum of stabilization and whatsoever.
[00:04:34.380 - 00:04:39.380] But that's nothing you can fix which you cannot fix with a master batch for example.
[00:04:39.380 - 00:04:52.380] But all this or but a lot of regulations standards and so on just prohibit using of as by now.
[00:04:52.380 - 00:04:57.380] Yeah, just prohibit the usage of external recycling material.
[00:04:57.380 - 00:05:09.380] So we're just not allowed to use post consumer post industrial recycling for making geo membranes as geo membranes.
[00:05:09.380 - 00:05:21.380] And the main problem about that is whenever you want to claim a durability of more than five years then you have to prove certain things.
[00:05:21.380 - 00:05:22.380] So that's where it begins.
[00:05:22.380 - 00:05:38.380] So to get a CE marking, for example, you have to prove basically every lot you use, you have to prove that the aging behavior and so on.
[00:05:38.380 - 00:05:56.380] So of course, we don't do that for virgin material because for virgin material it is assumed as long as the QC of the manufacturer of the virtual material is checked every it's audited in regular times and so on.
[00:05:56.380 - 00:06:14.380] It can be assumed that the material stays always the same and so on and therefore they elongated this timeframe where you have to recheck this to about five years if I'm not wrong for RC material where you don't actually know what material went into that lot.
[00:06:14.380 - 00:06:22.380] You probably can't tell if this was yogurt bins or whatsoever.
[00:06:22.380 - 00:06:25.380] You technically have to do this for every lot.
[00:06:25.380 - 00:06:29.380] And these tests are not only expensive.
[00:06:29.380 - 00:06:32.380] They are as well taking four months.
[00:06:32.380 - 00:06:41.380] So for us, if you would use for our products recycled material now.
[00:06:41.380 - 00:06:53.380] So it's the regulation nowadays, we would have to store the material for about four months until we get the result of the aging before we can further process.
[00:06:53.380 - 00:06:59.380] That has nothing to do with the quality of the material or the possibility process ability of the material.
[00:06:59.380 - 00:07:10.380] It's just because of the regulations that we have to check every lot for a certain for certain things, especially aging and so on.
[00:07:10.380 - 00:07:19.380] That said, this limits the market to so if you don't do that, you can only claim five years of durability.
[00:07:19.380 - 00:07:23.380] And this limits your market very much.
[00:07:23.380 - 00:07:29.380] You can claim that five years is pretty much enough for temporary capping or whatsoever.
[00:07:29.380 - 00:07:41.380] The problem in the discussion is the customer usually doesn't know by but for sure that the application will be done within five years.
[00:07:41.380 - 00:07:48.380] So what happens if for any reason you have to use it longer longer than the than the five years?
[00:07:48.380 - 00:07:52.380] Who is taking the responsibility for the time after?
[00:07:52.380 - 00:07:59.380] And what happens if you buy the material and the installation is delayed by one year for any reason?
[00:07:59.380 - 00:08:02.380] You know, so and the price difference.
[00:08:02.380 - 00:08:11.380] If there is any to our to our knowledge, usually the recycling material is even more expensive than version material in most of the cases.
[00:08:11.380 - 00:08:16.380] The customers usually just go with the virgin material is cheaper.
[00:08:16.380 - 00:08:20.380] They don't have to worry about these five years timeframe.
[00:08:20.380 - 00:08:23.380] So much more convenient.
[00:08:23.380 - 00:08:24.380] That's the problem.
[00:08:24.380 - 00:08:26.380] One hand side.
[00:08:26.380 - 00:08:32.380] Second hand side is the technical requirements to geomembranes in particular.
[00:08:32.380 - 00:08:37.380] For geomembranes, it's not only to be aging proof and so on.
[00:08:37.380 - 00:08:44.380] It's as well, there are some requirements which are related to the chemicals in the landfill.
[00:08:44.380 - 00:08:47.380] So it's it's called stress crack resistance.
[00:08:47.380 - 00:08:58.380] It's basically you always have some force on a geomembrane in due to the the own weight and due to some capping on top or whatsoever.
[00:08:58.380 - 00:09:06.380] So like when it forms like a mattress or whatsoever, and especially HDPEs.
[00:09:06.380 - 00:09:17.380] So really HDPEs where the density of the polymer exceeds 900, 45, 40, 45, something like that.
[00:09:17.380 - 00:09:21.380] They are very prone to stress crack behavior.
[00:09:23.380 - 00:09:30.380] We have to proof our products usually something like 500 to 800 hours of stress crack resistance.
[00:09:30.380 - 00:09:33.380] You don't have to understand exactly the test.
[00:09:33.380 - 00:09:42.380] But what we know from HDPEs, even from virgin HDPEs, which are by themselves already HDPE.
[00:09:42.380 - 00:09:47.380] We reach something like 10 hours under the same conditions.
[00:09:47.380 - 00:09:54.380] And this really is not suitable for landfill, typical landfill applications.
[00:09:54.380 - 00:10:13.380] So the possibility of using recycled material would need a customer who is willing to pay more, get less and has some restrictions on how to use it.
[00:10:13.380 - 00:10:24.380] So we can discuss about, yeah, if on a capping, which is only temporary and usually just necessary to keep away some water.
[00:10:24.380 - 00:10:31.380] And even if it has a minor leakage, it is still much better than without the capping and so on.
[00:10:31.380 - 00:10:38.380] But that is a very narrow, very narrow market.
[00:10:38.380 - 00:11:00.380] So our concerns are 80, I would say 90% about regulations and market restrictions and only 10% about the, how to say, process stability and the possibility to run the material actually on our lines.
[00:11:00.380 - 00:11:18.380] So for example, the requirement of the ash remains is just a requirement that there must not be any grains in the membrane because these are failure spots.
[00:11:18.380 - 00:11:33.380] So that's when you apply some, when you apply some load to the membrane, then we know that at that places, the membrane will start failing much earlier as within a, if there is a very homogeneous metric.
[00:11:33.380 - 00:11:36.380] So that's the reasons why we are.
[00:11:36.380 - 00:11:37.380] So that's the reasons why we asked for that.
[00:11:37.380 - 00:11:46.380] And if you exceed a certain, let's say a certain grain size and a certain amount, you get some wear and tear in the extruder.
[00:11:46.380 - 00:11:48.380] You probably know that very well.
[00:11:48.380 - 00:11:57.380] So that's the reason why we request these boundary conditions like this 38 micrometers.
[00:11:57.380 - 00:12:04.380] That's just because bigger grain sizes would cause failures in the membrane later on.
[00:12:04.380 - 00:12:09.380] And it usually will create wear and tear on our extruders.
[00:12:09.380 - 00:12:10.380] That's, that's the two things.
[00:12:10.380 - 00:12:12.380] It's not just like, I didn't bring that up.
[00:12:12.380 - 00:12:13.380] Just yeah.
[00:12:13.380 - 00:12:16.380] That's the specification of the virtual material.
[00:12:16.380 - 00:12:25.380] We have some real, we have some real reasons, real reasons for that, which are related to our product.
[00:12:25.380 - 00:12:34.380] So, so to say a general usage of recycled material, and that it's not particular for your product.
[00:12:34.380 - 00:12:50.380] It's in general is mostly the biggest, the biggest obstacle is the, is the, the, the, the standards and so on.
[00:12:50.380 - 00:12:59.380] Second is it, it has to have a certain cleanliness to be processed, but the main thing is the regulations.
[00:12:59.380 - 00:13:13.380] Um, um, as well, you need to, for, for landfill, um, applications or for all ceiling applications, which are where, where people are willing to pay more money than for a PVC film.
[00:13:13.380 - 00:13:21.380] Um, they usually look exactly for the PE, uh, behavior and the resistance of PE.
[00:13:21.380 - 00:13:40.380] And when you, when you have like 5% or 2% or whatever for, from PP or polyester or whatsoever in your mixture, you just out by definition because they usually request 100% polyethylene.
[00:13:40.380 - 00:13:50.380] Um, so that's just, yeah, it's, it's, it's a geomembrane market on that, on that point.
[00:13:50.380 - 00:14:06.380] On the other hand, if there is a, if there is a demand, or if you want to do it really in a closed, in a closed, closed environment or a closed project where everyone is aware of the costs and the limitations.
[00:14:06.380 - 00:14:12.380] Um, it shouldn't be a problem for us to make these products.
[00:14:12.380 - 00:14:18.380] We have some constraints because we have to separate it from our other materials.
[00:14:18.380 - 00:14:22.380] So silo handling and so on will be an issue.
[00:14:22.380 - 00:14:35.380] But from the technical point of view, if the MFR at, um, I think you gave 200, uh, 192.16, right?
[00:14:35.380 - 00:14:39.380] We gave below one on a, on a 2.16 kilogram measurement.
[00:14:39.380 - 00:14:40.380] Yeah.
[00:14:40.380 - 00:14:49.380] So, um, we are usually that that's what we, so one is already fluid for us, but we are in the range.
[00:14:49.380 - 00:14:50.380] I would say we are in the same range.
[00:14:50.380 - 00:14:54.380] So the minimum we use actually is 0.3, if I'm not wrong.
[00:14:54.380 - 00:14:56.380] And yeah, maybe up to one.
[00:14:56.380 - 00:14:59.380] So we are more familiar with five kilogram values.
[00:14:59.380 - 00:15:14.380] So the five kilogram values would be between one and 0.6 to, I would say three grams per 10 minutes or something like that.
[00:15:14.380 - 00:15:24.380] So generally there's a good chance to process that, especially when we mix it with virgin material.
[00:15:24.380 - 00:15:37.380] So, and that's the, but anyways, this will give us much better tech or much more fitting technical parameters, but it will not solve the problem on the regulation side.
[00:15:37.380 - 00:15:43.380] On a regulation side at the moment, there will be, there will be some changes in the future.
[00:15:43.380 - 00:15:50.380] The European union is working on standards, um, to normalize the usage of, um, recycled materials.
[00:15:50.380 - 00:15:55.380] But as for now, and I would say the upcoming five years doesn't matter.
[00:15:55.380 - 00:15:59.380] There is a zero RC policy in the products we make.
[00:15:59.380 - 00:16:03.380] And it's not only geomembranes, it's geo grids.
[00:16:03.380 - 00:16:13.380] It's nonwovens, it's wovens, everything, which is considered, uh, geo textile in its main definition.
[00:16:13.380 - 00:16:15.380] There is this limitations on that.
[00:16:15.380 - 00:16:23.380] So we are only allowed to use a certain percentage of internal reuse materials of side trims and so on.
[00:16:23.380 - 00:16:24.380] That's the only excuse.
[00:16:24.380 - 00:16:26.380] And even that is heavily limited.
[00:16:27.380 - 00:16:30.380] So we can only use, I think 10 or 15% of that.
[00:16:32.380 - 00:16:37.380] So generally speaking, the, there's a certain activity.
[00:16:37.380 - 00:16:38.380] Yep.
[00:16:38.380 - 00:16:45.380] That this, uh, circular economy approach, uh, that we discussed last week is, is, uh, really possible.
[00:16:45.380 - 00:16:56.380] Um, on the other hand, what Martin is just describing, um, it limits our action on the market with regards to all these regulations.
[00:16:56.380 - 00:17:03.380] So we, from, from, from now aside, we cannot think this in a, in a large scale model.
[00:17:03.380 - 00:17:04.380] Yeah.
[00:17:04.380 - 00:17:15.380] Um, but if, uh, wherever there are demands that the material just needs to, uh, survive for five years.
[00:17:15.380 - 00:17:18.380] So like the temporary cappings that you described trees.
[00:17:18.380 - 00:17:19.380] Yeah.
[00:17:19.380 - 00:17:20.380] Then of course, this is possible.
[00:17:20.380 - 00:17:34.380] Well, Martin, thank you for the explanation and, and Christian that you, that you gave before I jump into the regulatory, um, aspects.
[00:17:34.380 - 00:17:39.380] Uh, it's just for me to understand also with the technical and some of the, of the topics you already touched.
[00:17:39.380 - 00:17:54.380] But to understand well on the particle size, do you mean that at the 38 micron, that, that should be the, uh, uh, uh, if we go put it into the melting phase, that no particles smaller than 38 micron can be found.
[00:17:54.380 - 00:18:00.380] So nothing to do with the pellet size or, this is about possible impurities.
[00:18:00.380 - 00:18:02.380] So remaining dirt, I would say.
[00:18:02.380 - 00:18:03.380] Yeah.
[00:18:03.380 - 00:18:10.380] So, um, um, our C set in the extruder is actually a 38 micron.
[00:18:10.380 - 00:18:15.380] So everything which is below that is accepted because we can't even filter it out on our own.
[00:18:15.380 - 00:18:18.380] Everything above that will be filtered out.
[00:18:18.380 - 00:18:22.380] If it is too much, we get in trouble because our filters clock very soon.
[00:18:22.380 - 00:18:23.380] Yeah.
[00:18:23.380 - 00:18:24.380] Okay.
[00:18:24.380 - 00:18:28.380] Um, the technology that you are using, is that cast extrusion or.
[00:18:28.380 - 00:18:30.380] Flat die cast extrusion.
[00:18:30.380 - 00:18:31.380] Flat die cast extrusion.
[00:18:31.380 - 00:18:32.380] Okay.
[00:18:32.380 - 00:18:33.380] From.
[00:18:33.380 - 00:18:42.380] And so that's one single layer, or do you have the option to, to, uh, let's say if it's in the future, it is possible to put recycled contents.
[00:18:42.380 - 00:18:49.380] Are you, uh, in the option of doing ABA, uh, structures and put the recycled contents in the middle?
[00:18:49.380 - 00:18:51.380] We have one line, which does ABA.
[00:18:51.380 - 00:18:56.380] And we have one line, which actually has two extruders, which are extruding next to each other.
[00:18:56.380 - 00:19:05.380] But for the more modern one, we have ABA or, uh, you can switch between A, B, A, B, A, as you want.
[00:19:05.380 - 00:19:06.380] Okay.
[00:19:06.380 - 00:19:14.380] But for our regulations or for our purpose, um, it doesn't make a difference.
[00:19:14.380 - 00:19:15.380] So we can't differentiate.
[00:19:15.380 - 00:19:27.380] It's not like, um, for example, window frames or so where you do the virtual material to the outside to get the pretty look and to get, uh, the colors and whatsoever.
[00:19:27.380 - 00:19:37.380] And for the filling, you use the recycled materials, our material, um, how to say, has to last from the top to the bottom the same.
[00:19:37.380 - 00:19:38.380] Yeah.
[00:19:38.380 - 00:19:39.380] Okay.
[00:19:39.380 - 00:19:40.380] Okay.
[00:19:40.380 - 00:19:41.380] Um, okay.
[00:19:41.380 - 00:19:52.380] And then based on the polymer purity, uh, as, as you're only using virgin material, of course, it's only PE.
[00:19:52.380 - 00:20:02.380] Um, or products are also, um, I mean, the PP content is, uh, an important parameter because of the, um, uh, you, you should be able to still seal the material.
[00:20:02.380 - 00:20:11.380] And especially also on our landfill, uh, if you cannot longer make a good seal between the different, um, products, then we have a problem, a problem.
[00:20:11.380 - 00:20:14.380] That's the same thing, which goes to pipes, for instance.
[00:20:14.380 - 00:20:21.380] So also pipes needs to be well, weldable, uh, still, and PP has a higher melting point compared to, to PE.
[00:20:21.380 - 00:20:31.380] Um, and therefore I know the spec is there around 4% of, of, uh, of PP, but that's maybe things you're not very familiar with as you, as at this point,
[00:20:31.380 - 00:20:35.380] you're only using, um, HD, right?
[00:20:35.380 - 00:20:44.380] Again, we, we, technically we use, um, this is a little bit weird, but technically we use maximum MDPE.
[00:20:44.380 - 00:20:53.380] So, uh, usually we use, uh, linear low, so LLDPE or MDPE from the work, from the polymer side.
[00:20:53.380 - 00:21:00.380] And our material is just called HDPE because we add about two and a half percent carbon black,
[00:21:00.380 - 00:21:04.380] which lifts the density in the range of HDPE.
[00:21:04.380 - 00:21:09.380] So from the polymer side, we are using maximum MDPE.
[00:21:09.380 - 00:21:12.380] And that's the reason why we have the pretty good weldability.
[00:21:12.380 - 00:21:17.380] Um, we have this stress, stress, stress, resistance.
[00:21:17.380 - 00:21:27.380] And some other values like, um, yield elongation in a certain range, uh, maximum elongation.
[00:21:27.380 - 00:21:31.380] You need a break along maximum elongation at break of about 800%.
[00:21:31.380 - 00:21:36.380] I think some, uh, some regulations ask for 700, some for 800.
[00:21:36.380 - 00:21:41.380] And we know that if you use HDPE polymer.
[00:21:41.380 - 00:21:48.380] So where the polymer already is an HDPE, um, you usually can't reach that because it's too stiff.
[00:21:48.380 - 00:21:49.380] Yeah.
[00:21:49.380 - 00:21:50.380] Yeah.
[00:21:50.380 - 00:21:51.380] Yeah.
[00:21:51.380 - 00:21:52.380] So that's the reason.
[00:21:52.380 - 00:21:53.380] Yeah.
[00:21:53.380 - 00:21:54.380] It's a little bit weird.
[00:21:54.380 - 00:22:01.380] So we call it HDPE, but actually it is more like, I would say 60%, um, MDPE and 40% linear low.
[00:22:01.380 - 00:22:02.380] Yeah.
[00:22:02.380 - 00:22:03.380] Okay.
[00:22:03.380 - 00:22:11.380] Well, our recycling plant, the only inputs we are using are, um, uh, post-consumer household packaging.
[00:22:11.380 - 00:22:15.380] And so we are, we're integrated with our group and we do collection of household packaging.
[00:22:15.380 - 00:22:17.380] We have sorting lines for household packaging.
[00:22:17.380 - 00:22:25.380] And then the sorted, uh, bales of, uh, HDPE and PP are then further treated at our, uh, polymer plants.
[00:22:25.380 - 00:22:37.380] And so on, on HDPE, this means, uh, yeah, mainly bottles, uh, HD, but you could also have some LDPE from creams, uh, cream and packaging, things like that.
[00:22:37.380 - 00:22:55.380] Um, and so on HD, we don't only have chemical bottles because the chemical bottles have a high crack, um, stress resistance because you don't want, um, chemical to, to go out, but not in the, in the values that you are mentioning to be clear, uh, because we're still speaking about bottles and canisters.
[00:22:55.380 - 00:22:59.380] Um, so that's the grade we are making.
[00:22:59.380 - 00:23:12.380] So if you're referring to the, uh, to the high allegation that breaks, and then I understand your argument that it's rather an LDPE or MD and not an HD that you're using in the, in the, in the film.
[00:23:12.380 - 00:23:16.380] So, um, so yeah, good to know.
[00:23:16.380 - 00:23:21.380] Um, I think at this point, uh, mainly to regulatory affairs, it's a bit difficult.
[00:23:21.380 - 00:23:24.380] Nevertheless, it's already good to understand the process.
[00:23:24.380 - 00:23:27.380] Also good to understand that there is an ABA structure possible.
[00:23:27.380 - 00:23:40.380] If regulatory will, would change that maybe there, there's possibilities, but then we have to look maybe more at the, at film, film grates rather than, than, than bottle, uh, bottle grates.
[00:23:40.380 - 00:23:41.380] For sure.
[00:23:41.380 - 00:23:45.380] Um, speaking about regulation, we had a bit of similar case.
[00:23:45.380 - 00:24:02.380] So when we are looking, um, at the usage of all consumables within our group, which are, uh, used for, uh, with, with, with virgin material, we had a big, uh, consumption of, uh, medical bins that we are using in the hospitals.
[00:24:02.380 - 00:24:03.380] Hmm.
[00:24:03.380 - 00:24:04.380] Okay.
[00:24:04.380 - 00:24:06.380] So these are PPE bins.
[00:24:06.380 - 00:24:13.380] This was also not possible to use from, um, with, with, with recycled material because of regulatory affairs.
[00:24:13.380 - 00:24:22.380] But we worked for two years on that, uh, project and we, um, we managed to get an exception, uh, on that.
[00:24:22.380 - 00:24:29.380] Uh, why it was difficult because we are in ADR, uh, ADR transport, that's transport of dangerous goods.
[00:24:29.380 - 00:24:38.380] And so in that regulatory, it was mentioned that we could only use virgin material or, uh, material coming from the same application.
[00:24:38.380 - 00:24:41.380] A bit like the case you are, uh, mentioning.
[00:24:41.380 - 00:24:46.380] Um, which state to censor doesn't really make sense as long.
[00:24:46.380 - 00:24:56.380] If you could prove that you can fulfill the application with other origins, it's, uh, in our opinion, it could be used.
[00:24:56.380 - 00:25:08.380] And so in this way, you went into discussion with, uh, European regulatory affairs together with, and the local, uh, legislation here in Belgium before, for transport of dangerous goods.
[00:25:08.380 - 00:25:16.380] Um, and we managed to make, uh, to make boxes, which also had the same technical properties like were mentioned.
[00:25:16.380 - 00:25:26.380] So also a high impact that the boxes don't break when they could fall off a truck and dangerous material from hospitals could be, could be, could be spilled.
[00:25:26.380 - 00:25:33.380] So, um, I don't see a solution within the upcoming short periods to be clear.
[00:25:33.380 - 00:25:42.380] Um, but it would be good for us to, if you could send us over the current regulation, which is in place, uh, which that's a bit the, the norm.
[00:25:42.380 - 00:25:47.380] And next to that, maybe, uh, have a bit, an idea.
[00:25:47.380 - 00:25:54.380] You mentioned some values about MFR, allegation net break, um, the, uh, stress crack resistance.
[00:25:54.380 - 00:25:57.380] If we could have the values values, we can save them.
[00:25:57.380 - 00:26:03.380] And then who knows within two or three years, if regulatory changes that, that we are well prepared,
[00:26:03.380 - 00:26:11.380] that maybe also that part of our consumables within our group, we could replace a part with virgin material, uh, with recycled material, um,
[00:26:11.380 - 00:26:20.380] which, which, which could lower significantly or, uh, or, um, yeah, a carbon footprint from our, from our group.
[00:26:20.380 - 00:26:28.380] I can, uh, I can say that we are part of a research program, which is coming to its end, which has exactly that goal.
[00:26:28.380 - 00:26:38.380] So this, it is the, the goal of this project, this pro geo UP, whenever you Google it, you will probably find some information about that.
[00:26:38.380 - 00:26:44.380] But even though I can share the link, there's a, there's a page where it's roughly describing what we are doing.
[00:26:44.380 - 00:26:58.380] Um, this is, this is especially the idea, especially behind that part we are doing is recycling of geo texts or geo composites, geo textiles, um, grids on whatsoever.
[00:26:58.380 - 00:27:09.380] Um, the, yeah, the goal or the, the attempt there is more in the way, like you just described for these medical, medical bins.
[00:27:09.380 - 00:27:17.380] So taking old geo textiles or geo synthetics and reconvert them to geo geo synthetics.
[00:27:17.380 - 00:27:24.380] This has some advantages because you're using material, which is proper for the, um, for the task.
[00:27:25.380 - 00:27:37.380] The, the problem with the post consumer material is that it in most of the cases is not proper for the geo synthetics case.
[00:27:37.380 - 00:27:46.380] So that, that's a problem with, when you do the recycling within the same, uh, kind of material, let's say we already proved that this is technically possible.
[00:27:46.380 - 00:27:54.380] Uh, we just proved it for a geo grid and we can for sure show it for a geo membrane on a landfill as well.
[00:27:54.380 - 00:28:11.380] So if you, if you, if you would take a geo membrane from the landfill and give it to you and you do the processing, like a rough cleaning and, um, melting and regranulate, we can probably do a geo membrane, which fulfills the task.
[00:28:11.380 - 00:28:17.380] And which technically probably meets the requirements because it's the same material.
[00:28:17.380 - 00:28:21.380] The aging is usually negligible for that time.
[00:28:21.380 - 00:28:31.380] Um, the problem already, uh, always occurs when you're mixing in material from other sources meant for other things.
[00:28:31.380 - 00:28:39.380] So, um, as well, the, the PP content, as you mentioned, the polypropylene content is probably technically not a big issue.
[00:28:40.380 - 00:28:47.380] Um, it's just because for the landfill application, that's most landfill, uh, and a hazardous control.
[00:28:47.380 - 00:28:55.380] This are the 99% of the cases where you use a geo membrane, they require 100% polyethylene.
[00:28:55.380 - 00:28:56.380] Um, so that's it.
[00:28:56.380 - 00:29:08.380] So even if we can, for the moment, even if we say, okay, yeah, um, we know that this is not an issue and so on, it's just not there.
[00:29:08.380 - 00:29:16.380] Um, but my colleague Henning Ehrenberg is part of the, um, regulation committee at the European Union.
[00:29:16.380 - 00:29:27.380] And from that I know for sure that they are working on changing the regulations for geosynthetics to make it possible to use recycled materials.
[00:29:27.380 - 00:29:39.380] And I can already say that the outcome will be something like, okay, for the recycled materials, the same things apply then for virgin materials.
[00:29:39.380 - 00:29:46.380] So there will be no difference between virgin and recycled as long as you fulfill the given requirements.
[00:29:46.380 - 00:29:49.380] It doesn't matter if it is from recycled or from virgin.
[00:29:49.380 - 00:30:02.380] Um, there will probably not be the, it will probably not be the case that there will be lower requirements for recycled sources for the same purpose.
[00:30:02.380 - 00:30:14.380] So there is always the opportunity to say, okay, we define a special purpose where we, where everyone is aware of using recycled materials with lower, um, properties.
[00:30:14.380 - 00:30:16.380] And that's, that's always possible.
[00:30:16.380 - 00:30:23.380] But if you want to join the, the open market, you have to prove that you can deliver the same quality.
[00:30:23.380 - 00:30:38.380] And there is already the discussion how we can, or we can come to a situation, which is like the virgin materials where you don't have to test each particular lot,
[00:30:38.380 - 00:30:48.380] but as well put some, um, um, QC requirements on you as the, uh, uh, manufacturer or the producer of the RC material.
[00:30:48.380 - 00:31:04.380] So that you with given, with given tests or given audits or whatsoever that everyone can rely on that, that the raw material or the recycled material has a constant quality, whatever this quality is.
[00:31:04.380 - 00:31:14.380] But then it can be assured that the quality is constant and it's not necessary anymore to test every delivery literally.
[00:31:14.380 - 00:31:26.380] So even, even today, so maybe that wasn't clear today, I can do a proper geosynthetics out of recycled material.
[00:31:26.380 - 00:31:27.380] It's just not feasible.
[00:31:27.380 - 00:31:37.380] I, I'm allowed to do that, but the amount of work and, and checking I have to put in that is just not worth doing it.
[00:31:37.380 - 00:31:47.380] I can, if, if I get one delivery from your side and we do all the testing, we wait for four months to get the results for the aging and so on.
[00:31:47.380 - 00:31:55.380] Then with this lot alone, I can do a geosynthetic production, whatever it is, geomembrane or whatsoever.
[00:31:55.380 - 00:31:59.380] And I can put the, the stamp on it and say, okay, yeah, that's, that's okay.
[00:31:59.380 - 00:32:06.380] That's fulfilling all the requirements, but that's not feasible because I can't do that in scale.
[00:32:06.380 - 00:32:07.380] Yeah.
[00:32:07.380 - 00:32:08.380] This is for one truck.
[00:32:08.380 - 00:32:09.380] That's okay.
[00:32:09.380 - 00:32:14.380] But not for, not for a hundred tons or, or 500 tons or so.
[00:32:14.380 - 00:32:18.380] So it's not like it is, it's not like it's forbidden.
[00:32:18.380 - 00:32:28.380] It's just not, it's, it's not possible by the amount of effort you have to put in it to make it work in scale.
[00:32:28.380 - 00:32:29.380] Yeah.
[00:32:29.380 - 00:32:37.380] And that's, that's even worse actually, because the, the answer of everyone is yes, you can do it.
[00:32:37.380 - 00:32:39.380] You just have to do this, this and that.
[00:32:39.380 - 00:32:48.380] Um, unfortunately doing that is making you unflexible because it's taking time.
[00:32:48.380 - 00:32:57.380] We, we would need you or we, some one of us would have to have big silos to store the material,
[00:32:57.380 - 00:33:01.380] which is under testing for the time being.
[00:33:01.380 - 00:33:03.380] And what if it fails?
[00:33:03.380 - 00:33:04.380] Huh?
[00:33:04.380 - 00:33:05.380] No one knows.
[00:33:05.380 - 00:33:16.380] And it is, um, relatively high costs for each of these tests, which even increases the price of the raw material.
[00:33:16.380 - 00:33:22.380] So it's actually not like I can give you a paragraph where it says it's forbidden.
[00:33:22.380 - 00:33:29.380] It, the requirements are just that stupid on RC materials or the, the requirements are easy on version materials.
[00:33:29.380 - 00:33:33.380] And they are applied on RC materials in the same way.
[00:33:33.380 - 00:33:36.380] And this makes it practically impossible.
[00:33:36.380 - 00:33:37.380] Yeah.
[00:33:37.380 - 00:33:38.380] Yeah.
[00:33:38.380 - 00:33:40.380] I understand, uh, very well.
[00:33:40.380 - 00:33:45.380] And unfortunately, um, this is still the world of today.
[00:33:45.380 - 00:33:54.380] And, um, luckily for instance, in packaging, we have some new regulations, uh, which, which obligate, obligates in 2030, a certain percentage of, of usage.
[00:33:54.380 - 00:34:00.380] And this helps of course, to create also the market because it's not a matter of willing.
[00:34:00.380 - 00:34:01.380] It's, it's a fact.
[00:34:01.380 - 00:34:07.380] And as long as there, it's not there, I understand in other markets, it's quite, uh, quite difficult.
[00:34:07.380 - 00:34:13.380] Nevertheless, um, with our plant, we have some collaborations running with some, um, more and more.
[00:34:13.380 - 00:34:24.380] We have virgin polymer producers who are integrating recycled contents to give a ready to use 30 percentage, uh, range to, to converters.
[00:34:24.380 - 00:34:32.380] So they, they, so that they, so that they don't need to worry about all those, uh, things that you are, uh, mentioning.
[00:34:32.380 - 00:34:40.380] So this goes from, uh, Total Energies, uh, Borealis, uh, also Ravago with the manufacturing that we're having.
[00:34:40.380 - 00:34:47.380] Um, I'm pretty sure it will not take too long to, that this is established.
[00:34:47.380 - 00:34:48.380] Yeah.
[00:34:48.380 - 00:34:58.380] But you need, as you say, you need one of the big players, um, taking the responsibility and doing the paperwork, let's say.
[00:34:58.380 - 00:35:07.380] So yeah, as soon as they can proclaim, okay, we take care that what's leaving our plant has a constant quality.
[00:35:07.380 - 00:35:13.380] So whatever this is, constant quality, and we can ensure that this is always the case.
[00:35:13.380 - 00:35:23.380] Then we, as the mid user, um, don't have any problem as long as the result of the product is feasible.
[00:35:23.380 - 00:35:24.380] Huh?
[00:35:24.380 - 00:35:29.380] So, but nowadays, um, we did some test runs, not on geomembranes, but on other products.
[00:35:29.380 - 00:35:33.380] We did some test runs with chemical recycled material.
[00:35:33.380 - 00:35:36.380] So there is a, I think it was a polyester in that case.
[00:35:36.380 - 00:35:47.380] They added during the, they added deep polymerized, um, monomers of recycled materials to the polymerization process.
[00:35:47.380 - 00:35:50.380] So at the end, it is like virgin material.
[00:35:50.380 - 00:35:55.380] Just, just the, the, the atoms and the molecules have been reused.
[00:35:55.380 - 00:35:55.380] Yeah.
[00:35:55.380 - 00:35:56.380] It is.
[00:35:56.380 - 00:36:02.380] And even that gives us trouble in the explanation or in, yeah.
[00:36:02.380 - 00:36:07.380] It's technically, you can't divide it from virgin material.
[00:36:07.380 - 00:36:08.380] Yeah.
[00:36:08.380 - 00:36:09.380] I see.
[00:36:09.380 - 00:36:15.380] But that's, that's, that shows the difficulty here because with recycled material, this is just virgin material with the certificate.
[00:36:15.380 - 00:36:16.380] Yeah.
[00:36:16.380 - 00:36:19.380] Because at the end, it's just the same virgin material.
[00:36:19.380 - 00:36:20.380] So, uh, I understand.
[00:36:20.380 - 00:36:28.380] Um, just for, for us, it could, it would be good, uh, if we can get some kind of idea with the, for the geo membranes.
[00:36:28.380 - 00:36:30.380] Uh, uh, uh, okay.
[00:36:30.380 - 00:36:38.380] The virgin material you are using not to make a copy paste, but to have a bit an idea about those allegation and breaks and, uh, and other things.
[00:36:38.380 - 00:36:48.380] And then from our sides, I will just, with our, um, partners that we are making hybrids, virgin and recycle, we'll just, uh, exchange some thoughts.
[00:36:48.380 - 00:36:57.380] Uh, if there would be no regulatory boundary, uh, what is feasible and what is not feasible because for the virgin by parts, they're using some kind of, yeah, we call it boosters.
[00:36:57.380 - 00:37:07.380] It's, it's overqualified virgin material to compensate for, for, uh, yeah, the lack of parameters and, and, and some, some, some RC material.
[00:37:07.380 - 00:37:09.380] This has to be research.
[00:37:09.380 - 00:37:23.380] We have done some, unfortunately mixing of materials, especially of high density and low density material, give some results, which are, which we didn't expect and which were unpredictable for us.
[00:37:23.380 - 00:37:33.380] So we, in our testing, we added very high density material in very low percentage to our typical resin.
[00:37:33.380 - 00:37:43.380] And for example, the stress crack resistance, even with 2% high density material dropped from usually 3000 hours plus.
[00:37:43.380 - 00:37:55.380] So our longest test we stopped at 3000 hours for that test dropped to less than 50 by just adding one or 2% of HD, HD material.
[00:37:55.380 - 00:38:08.380] So it's not like mixing milk and cacao, for example, it is, there is something happening, which is not like they don't behave, um, in relation to the mixture ratio.
[00:38:08.380 - 00:38:11.380] So this is a very interesting field.
[00:38:11.380 - 00:38:14.380] So, um, but, um, what I can mention is yes.
[00:38:14.380 - 00:38:22.380] Having a ready-made material, which is already, let's say 20% recycled together with virgin, with a certificate is always welcome.
[00:38:22.380 - 00:38:27.380] From the technical point of view, we can also mix in our extruders.
[00:38:27.380 - 00:38:31.380] We don't have this double screw mixing extruders.
[00:38:31.380 - 00:38:35.380] We have a single screw extruders, but with, um, barriers and so on.
[00:38:35.380 - 00:38:40.380] So we do regular mixing of same kind of materials.
[00:38:40.380 - 00:38:45.380] So different raw materials in the same area of MFR and so on.
[00:38:45.380 - 00:38:48.380] So we do regular mixing on that.
[00:38:48.380 - 00:38:52.380] Technically, as I said, technically is not the biggest difficulty.
[00:38:52.380 - 00:38:55.380] It is like how it behaves afterwards.
[00:38:55.380 - 00:38:57.380] So is it fit for the purpose?
[00:38:57.380 - 00:39:01.380] And the big thing is regulation things.
[00:39:01.380 - 00:39:02.380] Regulation.
[00:39:02.380 - 00:39:03.380] Yeah.
[00:39:03.380 - 00:39:04.380] Yeah.
[00:39:04.380 - 00:39:13.380] And, um, if you're interested, we can do some quick, quick tests, quick research eventually.
[00:39:13.380 - 00:39:18.380] So every now and then we do some tests on a lab extruder scale.
[00:39:18.380 - 00:39:28.380] So just 30 centimeters wide, a couple of kilograms, um, to get an idea what we can expect from this material.
[00:39:28.380 - 00:39:44.380] Uh, we can incorporate in our next test run, maybe a mixture of 20% of your RC material together with 80% of our, uh, virgin material or even 100% just to get an idea what we would end up with.
[00:39:44.380 - 00:39:46.380] Yeah.
[00:39:46.380 - 00:39:50.380] So then we were talking about 20, let's say 50 kilograms.
[00:39:50.380 - 00:39:51.380] Yeah.
[00:39:51.380 - 00:40:05.380] Well, honestly, I think the only thing what would make sense is, um, is a hybrid product where already 20% is integrated and made by, uh, uh, Total or Borealis or Avago.
[00:40:05.380 - 00:40:09.380] That's, that's three main or Bazel, Bazel could do it as well.
[00:40:09.380 - 00:40:11.380] It's the main players we're working with together.
[00:40:11.380 - 00:40:21.380] Um, uh, 100% pure RC from our sites will never make the specification because it's HD based.
[00:40:21.380 - 00:40:30.380] I'm now more thinking about, uh, uh, postcon, because if you want to do it at scale, you need post-consumer material, which is available everywhere in Europe.
[00:40:30.380 - 00:40:33.380] Otherwise you will never manage to make, to do something on that scale.
[00:40:33.380 - 00:40:34.380] That's the second point.
[00:40:34.380 - 00:40:35.380] Yeah.
[00:40:35.380 - 00:40:43.380] So this means if, if, if, if from our bottles are everywhere and packaging and, but I think it's will be very hard or impossible.
[00:40:43.380 - 00:40:46.380] The films are also everywhere in, um, in Europe.
[00:40:46.380 - 00:40:58.380] So I think we mainly have to look at an, an LD, LLD mixtures from, um, from films, um, to manage the elongation and breaks and this kinds of parameters.
[00:40:58.380 - 00:41:01.380] But, uh, that's, that's something I want to discuss with them first.
[00:41:01.380 - 00:41:02.380] Okay.
[00:41:02.380 - 00:41:12.380] What type of origin could be usable combined with your boosters, uh, from, from Virgin to end up with something which, which comes in the range with the current spec you are using.
[00:41:12.380 - 00:41:19.380] And, uh, and of course, then I want to challenge them as well with, uh, with the regulations and what kind of, uh, things are asked.
[00:41:19.380 - 00:41:32.380] And, uh, they have also people, people working monthly on that because they also want to open markets for high quality recycled material where currently, um, no recycled is being used.
[00:41:32.380 - 00:41:33.380] Yeah.
[00:41:33.380 - 00:41:35.380] That's a, that's the availability is a big point.
[00:41:35.380 - 00:41:44.380] So for our tests of recycling geosynthetics, um, so make out of a, use geosynthetics, make new ones.
[00:41:44.380 - 00:41:49.380] One of the biggest struggles besides all the techniques is the availability.
[00:41:49.380 - 00:41:57.380] So geosynthetics, they are made to last at least 25 years in constructions, some 100 years.
[00:41:57.380 - 00:42:00.380] So they are just not available in numbers on the market.
[00:42:00.380 - 00:42:04.380] There are three tons there, five tons there.
[00:42:06.380 - 00:42:09.380] Man, you can, you will never manage to do something at scale.
[00:42:09.380 - 00:42:10.380] No, no.
[00:42:10.380 - 00:42:11.380] Yeah.
[00:42:11.380 - 00:42:15.380] Of course, consumer is household packaging material, uh, commercial films.
[00:42:15.380 - 00:42:24.380] Um, that's the main markets you will find, which you will found everywhere in every country in Europe, where at least there's some collection.
[00:42:24.380 - 00:42:31.380] So, uh, so if you want to do something at scale, you have to, this, this has to be the basis from, from the recycled content.
[00:42:31.380 - 00:42:32.380] Yeah.
[00:42:32.380 - 00:42:35.380] So scaling is probably the right word.
[00:42:35.380 - 00:42:48.380] Um, we have now learned there's one big column, uh, this part of discussion, which would be a win-win situation in the best case, maybe in a few, in a couple of years.
[00:42:48.380 - 00:42:49.380] Yeah.
[00:42:49.380 - 00:43:00.380] Um, the other part that is, uh, where I actually don't know, is this topic still, uh, of interest would be the column with circular economy.
[00:43:00.380 - 00:43:17.380] Um, please, would this be an option for you to make a reuse of your HDPE for circular economy reasons to build a story out of that?
[00:43:17.380 - 00:43:21.380] Or is this not of interest?
[00:43:21.380 - 00:43:28.380] The one of which we recycle, uh, uh, recycler are temporary covering, you mean?
[00:43:28.380 - 00:43:29.380] Yeah.
[00:43:29.380 - 00:43:34.380] I think at this, at this stage would be very difficult, uh, to do so.
[00:43:34.380 - 00:43:36.380] But I, I think Jeroen, you have a better view on it.
[00:43:36.380 - 00:43:37.380] Yeah.
[00:43:37.380 - 00:43:40.380] But I think in terms of, yeah, feasibility.
[00:43:40.380 - 00:43:43.380] On, on our plants, it's, it's not possible.
[00:43:43.380 - 00:43:47.380] Our plant is built for post-consumer, post-consumer households, uh, waste.
[00:43:47.380 - 00:43:50.380] And so our inputs are bells of sorted bottles.
[00:43:50.380 - 00:43:55.380] So, um, it could be very, the same thing you are mentioning, eh?
[00:43:55.380 - 00:44:01.380] We have to, we would have to do very small lots of, of those temporary film material.
[00:44:01.380 - 00:44:08.380] We should, we would have to do external size reduction or baling or whatever to precondition it so we can feed it into our plant.
[00:44:08.380 - 00:44:17.380] So, I think it's very difficult to, to get that to, um, to, uh, at scale level, uh, at, at our recycling plants.
[00:44:17.380 - 00:44:24.380] Um, we have such types of films sometimes that we collect, that we deliver to, uh, to a local recycle,
[00:44:24.380 - 00:44:30.380] which is using it in very low-end application because getting it to a recycling line is one.
[00:44:30.380 - 00:44:35.380] But getting it to a level where it can be reused for high-end applications is another story.
[00:44:35.380 - 00:44:40.380] I mean, our washing line, we, uh, we built that for a preparation for a PPWR.
[00:44:40.380 - 00:44:43.380] So that's for a new packaging regulation.
[00:44:43.380 - 00:44:45.380] So we have four washing steps.
[00:44:45.380 - 00:44:50.380] We can, uh, we can, um, increase the temperature of, uh, of one of the washing steps.
[00:44:50.380 - 00:44:57.380] We asked caustic soda, we, to prepare as much as possible to get it to higher-end applications
[00:44:57.380 - 00:45:00.380] and packaging and color sorting, all kinds of, of those things.
[00:45:00.380 - 00:45:09.380] And so, um, yeah, at this point, if it goes to lower-end applications, uh, uh, then it's,
[00:45:09.380 - 00:45:12.380] or a plant is not the right plant to, to, to put it through.
[00:45:12.380 - 00:45:18.380] So the circular economy ID, uh, also, I think it's about lower quantities.
[00:45:18.380 - 00:45:21.380] And I don't know the quantities, uh, which, which are coming out of that.
[00:45:21.380 - 00:45:26.380] But, um, my opinion, that's very difficult to do that at scale.
[00:45:26.380 - 00:45:28.380] Getting it recycled is one thing.
[00:45:28.380 - 00:45:30.380] And then to lower-end application.
[00:45:30.380 - 00:45:32.380] And then we should aim as much as possible for.
[00:45:32.380 - 00:45:34.380] So that's not a discussion.
[00:45:34.380 - 00:45:37.380] But to set something up, to reuse it back into your applications,
[00:45:37.380 - 00:45:43.380] it has to come from post-consumer, uh, higher availability streams, um,
[00:45:43.380 - 00:45:48.380] um, and then do the right compounding steps to get the ready-to-use, uh, polymer
[00:45:48.380 - 00:45:53.380] without any stress at converting site with the right, uh, certificates and ever.
[00:45:53.380 - 00:46:01.380] And, and then that's the only way I see to, to get recycled content into your products, to be honest.
[00:46:01.380 - 00:46:02.380] Yeah.
[00:46:02.380 - 00:46:05.380] Is it a question, Martin and Christian?
[00:46:05.380 - 00:46:12.380] Because we're now at Van Heede, we're questioning if a recycled substitute is available.
[00:46:12.380 - 00:46:15.380] Is it also, and you say, yeah, the market is small.
[00:46:15.380 - 00:46:20.380] Have you ever, let's say, gotten the interest or the question from other parties?
[00:46:20.380 - 00:46:24.380] Because we're looking at it for, let's say, a very narrow scope.
[00:46:24.380 - 00:46:29.380] Let's say temporary cover, smaller than, or used less than five years.
[00:46:29.380 - 00:46:36.380] But I don't know if in your market you've had the question before on, hey, do you also offer a version,
[00:46:36.380 - 00:46:40.380] a version with recycled material?
[00:46:40.380 - 00:46:49.380] Well, in this, in this application, uh, that we focus on here, uh, what you just heard from Martin's explanation,
[00:46:49.380 - 00:46:56.380] we are, we are always forced, uh, to generate solutions that are very long lasting, minimum 100 years.
[00:46:56.380 - 00:47:04.380] Um, this, this is the highest argument usually for, for our customers to be, have the safety that the constructions
[00:47:04.380 - 00:47:06.380] they are building last for 100 years.
[00:47:06.380 - 00:47:16.380] And if you can give, uh, as well, some, uh, some, yeah, research results that can have an outlook on the time after,
[00:47:16.380 - 00:47:22.380] yeah, where we are actually thinking, okay, our solutions will also survive 400 years.
[00:47:22.380 - 00:47:26.380] That's the biggest argument for the moment.
[00:47:26.380 - 00:47:27.380] Yeah.
[00:47:27.380 - 00:47:31.380] Everything what we are discussing here now is, is of course, I would say part of the future.
[00:47:31.380 - 00:47:31.380] Yeah.
[00:47:31.380 - 00:47:49.380] That you also think about how can we bring in recycled material for those solutions, but, uh, in the, in the same moment, uh, staying on the same, on the safe side and, uh, making something out of it, uh, that is really protecting also the third and the fourth generation after us.
[00:47:49.380 - 00:47:50.380] Yeah.
[00:47:50.380 - 00:48:16.380] That's one of the, what Christian describes is the, this geomembranes are one of the most sensitive, um, applications of geosynthetics because you're containing biohazards and really no one is willing to take any risk on that.
[00:48:16.380 - 00:48:25.380] Uh, if you're talking about non-wovens, for example, which act as a filter layer or something like that, the risk of failure is much less.
[00:48:25.380 - 00:48:31.380] So people, uh, so there, it would be much easier to convince people to take some risk.
[00:48:31.380 - 00:48:38.380] Even if we are convinced there is none, but there is still these, uh, it is lower quality.
[00:48:38.380 - 00:48:48.380] Um, then the outcome of, of failure would be probably not even noticeable in, in most of the cases, but for geomembranes is really, it is waste containment.
[00:48:48.380 - 00:48:52.380] No one wants to even think about that.
[00:48:52.380 - 00:48:57.380] There is a hole or a crack or, um, uh, a bad welding.
[00:48:57.380 - 00:49:06.380] Um, if I can have the better, uh, material in, in air quotes for the same price.
[00:49:06.380 - 00:49:22.380] And that's actually, if, if you would be half the price with the recycling, then probably some would think for some applications about it, but just think it from, from the, from your own perspective.
[00:49:22.380 - 00:49:36.380] Um, you're buying a material with lower quality or lower, um, promise of quality for a temporary job, which is also more expensive.
[00:49:37.380 - 00:49:38.380] Yeah.
[00:49:38.380 - 00:49:48.380] No one there's, there's just a very low, uh, amount of people who is willing to pay extra for these carbon footprint.
[00:49:48.380 - 00:49:53.380] Most will go with the lowest price, unfortunately.
[00:49:53.380 - 00:50:00.380] So there are, there are parts in geosynthetics, which are probably a little bit easier from a technical point of view and where the risk is much lower.
[00:50:00.380 - 00:50:06.380] And even those are not for, for the moment, they're not accepting the extra costs.
[00:50:08.380 - 00:50:15.380] Because in the, in the, in the mind of the people, usually recycling is combined with lower quality and lower costs.
[00:50:18.380 - 00:50:22.380] Um, in the Netherlands, you're, you're, you're Belgium, right?
[00:50:23.380 - 00:50:23.380] Belgium.
[00:50:23.380 - 00:50:23.380] Yeah.
[00:50:23.380 - 00:50:24.380] Yeah.
[00:50:24.380 - 00:50:39.380] In the Netherlands, I hear every now and then about RC material to be used, but even still with the same requirements, which is for the moment, just impossible.
[00:50:41.380 - 00:50:42.380] Okay.
[00:50:42.380 - 00:50:45.380] But they're more on the, so recycled.
[00:50:45.380 - 00:50:46.380] Yes.
[00:50:46.380 - 00:50:51.380] But they're more on interestingly bio based standard polymers.
[00:50:51.380 - 00:51:01.380] So a bio based polyester or bio based polyethylene, um, which at the end, which at the end is the same as virgin material.
[00:51:01.380 - 00:51:02.380] Yeah.
[00:51:02.380 - 00:51:03.380] Yeah.
[00:51:03.380 - 00:51:04.380] Yeah.
[00:51:04.380 - 00:51:05.380] Yeah.
[00:51:05.380 - 00:51:06.380] It's the product itself.
[00:51:06.380 - 00:51:09.380] You can't tell if it's from, from corn or from, from crude oil.
[00:51:09.380 - 00:51:10.380] So.
[00:51:10.380 - 00:51:11.380] Yeah.
[00:51:11.380 - 00:51:12.380] Yeah.
[00:51:12.380 - 00:51:13.380] Okay.
[00:51:13.380 - 00:51:19.380] But for me, it was interesting because geo membranes is for me a new application of recycled content.
[00:51:19.380 - 00:51:21.380] So I learned already a lot.
[00:51:21.380 - 00:51:42.380] At this point I want to dig a bit with my polymer producers into the technical demands and then the norms of itself.
[00:51:42.380 - 00:51:54.380] And then who knows, within two, three, four years, because the world is spinning very, very fast around and at least we are prepared.
[00:51:54.380 - 00:52:02.380] And yeah, for me that was a good first intro session and who knows where we will be in a couple of years.
[00:52:02.380 - 00:52:06.380] As you said, just think about the bottles, the polyester bottles.
[00:52:06.380 - 00:52:14.380] Maybe there will be some rule like you have to use 30% of recycled materials, then the whole game changes.
[00:52:14.380 - 00:52:16.380] Yeah, yeah.
[00:52:18.380 - 00:52:21.380] Are there also other competitors already working on this?
[00:52:21.380 - 00:52:29.380] And let's say as a vision we want to be the first one offering a recycled solution or are you not aware of?
[00:52:29.380 - 00:52:43.380] Not that I'm aware. So there are some, to be frank, there are some competitors who are showing example cases where they did what we as well did.
[00:52:43.380 - 00:52:56.380] So using a geotextile or geosynthetic and recycled this particular geosynthetic in a same kind of geosynthetic.
[00:52:56.380 - 00:52:58.380] So that's what I'm aware of.
[00:52:58.380 - 00:53:08.380] So buying, especially consumer grade recycling materials, I'm not aware of, let's say it like that.
[00:53:08.380 - 00:53:09.380] Okay.
[00:53:09.380 - 00:53:10.380] We have.
[00:53:10.380 - 00:53:11.380] All right.
[00:53:11.380 - 00:53:25.380] Typically the assumption, and maybe that's the last word, the assumption of using recycled polyester is that this might be easier.
[00:53:25.380 - 00:53:45.380] Because as far as we have been described is that the recycling materials stream is much cleaner because it's the, if you use the bottles, the bottles are very well sorted by the consumer himself.
[00:53:45.380 - 00:54:02.380] And the material is very consistent and so on, which is just not the case for what you just described, PP and PE, which is more like packaging yogurt bins and so on.
[00:54:02.380 - 00:54:03.380] Yeah.
[00:54:03.380 - 00:54:12.380] So you will probably, the first, the first appearing recycled material product will probably be polyester.
[00:54:12.380 - 00:54:13.380] Yeah.
[00:54:13.380 - 00:54:17.380] If I would bet some money, that would be my bet.
[00:54:17.380 - 00:54:18.380] Yeah.
[00:54:18.380 - 00:54:26.380] But I can align as well that the post-consumer PT bottles are, yeah, it's only from bottles.
[00:54:26.380 - 00:54:31.380] So the, the variability of parameters are very, it's very narrow.
[00:54:31.380 - 00:54:36.380] Um, so I can understand what you're saying, Mark.
[00:54:36.380 - 00:54:37.380] Yeah.
[00:54:37.380 - 00:54:38.380] Yeah.
[00:54:38.380 - 00:54:39.380] Yeah.
[00:54:39.380 - 00:54:40.380] It's plus and minus, huh?
[00:54:40.380 - 00:54:44.380] So it's very narrow, but you know what you get.
[00:54:44.380 - 00:54:45.380] Yeah.
[00:54:45.380 - 00:54:46.380] Yeah.
[00:54:46.380 - 00:54:47.380] Yeah.
[00:54:47.380 - 00:54:48.380] Yeah.
[00:54:48.380 - 00:54:49.380] Yeah.
[00:54:49.380 - 00:54:50.380] Yeah.
[00:54:50.380 - 00:54:51.380] Very interesting for me as well.
[00:54:51.380 - 00:54:52.380] Thank you very much.
[00:54:52.380 - 00:54:58.380] No, thanks to, uh, the, the four of you, uh, for the interesting brainstorm.
[00:54:58.380 - 00:55:01.380] Although I have to admit, I was very astonished about the knowledge.
[00:55:01.380 - 00:55:03.380] For me, it was a bit of Chinese.
[00:55:03.380 - 00:55:08.380] So, um, uh, but, uh, you, you explained it very well.
[00:55:08.380 - 00:55:10.380] So, um, thanks for that.
[00:55:10.380 - 00:55:12.380] That's when you bring in two experts.
[00:55:12.380 - 00:55:13.380] Yeah.
[00:55:13.380 - 00:55:17.380] And we were like, no, but, uh, it's good.
[00:55:17.380 - 00:55:23.380] Uh, it's good to have at least a view on the market and where it would be going, uh,
[00:55:23.380 - 00:55:24.380] within the next couple of years.
[00:55:24.380 - 00:55:29.380] And indeed, as you said, if we're ready from our point of view or our stand, I think it's
[00:55:29.380 - 00:55:30.380] good.
[00:55:30.380 - 00:55:34.380] I'm pretty sure that there will be changes in the next upcoming years.
[00:55:34.380 - 00:55:39.380] Um, so on the regulation side, on the technical side, everything is developing fast.
[00:55:41.380 - 00:55:47.380] Um, with the situation in Middle East, maybe there will be a little more consequence in doing
[00:55:47.380 - 00:55:50.380] whatever net is necessary for recycling.
[00:55:50.380 - 00:55:59.380] If the oil is not that available anymore, that we are used to, or, um, in the worst, or even worse.
[00:55:59.380 - 00:56:05.380] The, the polymers made in Middle East can't leave Middle East.
[00:56:05.380 - 00:56:12.380] So yes, I expect that there will be some acceleration on that in the upcoming years.
[00:56:12.380 - 00:56:18.380] It always needs financial advantages or pressure to change things.
[00:56:19.380 - 00:56:21.380] Unfortunately, that's our world.
[00:56:21.380 - 00:56:22.380] Yep.
[00:56:24.380 - 00:56:25.380] All right.
[00:56:25.380 - 00:56:26.380] Okay.
[00:56:26.380 - 00:56:27.380] We won't.
[00:56:27.380 - 00:56:28.380] Thank you very much, everybody.
[00:56:28.380 - 00:56:29.380] Yes.
[00:56:29.380 - 00:56:30.380] Indeed.
[00:56:30.380 - 00:56:31.380] Yeah.
[00:56:31.380 - 00:56:32.380] Thanks for this interesting exchange.
[00:56:32.380 - 00:56:36.380] Let's, uh, stay connected with each other.
[00:56:36.380 - 00:56:37.380] Yes.
[00:56:37.380 - 00:56:42.380] And if there's some interest, uh, to, to prove some things, yeah, please do not hesitate to
[00:56:42.380 - 00:56:43.380] come back to us.
[00:56:44.380 - 00:56:51.380] And, uh, yeah, Dries, I think we will stay in contact, uh, now, anyhow, regarding the, uh,
[00:56:51.380 - 00:56:52.380] The upcoming tender.
[00:56:52.380 - 00:56:53.380] The upcoming tenders.
[00:56:53.380 - 00:56:54.380] Yeah.
[00:56:54.380 - 00:56:57.380] And to see what, where, where it brings us together then.
[00:56:57.380 - 00:56:58.380] Yes.
[00:56:58.380 - 00:56:59.380] Yes.
[00:56:59.380 - 00:56:59.380] Indeed.
[00:56:59.380 - 00:57:06.380] Normally end of this week, latest, beginning of next week, we will send out, uh, the tender.
[00:57:06.380 - 00:57:09.380] So we will keep you in, uh, in copy.
[00:57:09.380 - 00:57:10.380] All right.
[00:57:10.380 - 00:57:11.380] Good.
[00:57:11.380 - 00:57:12.380] All right.
[00:57:12.380 - 00:57:13.380] Good.
[00:57:13.380 - 00:57:13.920] All right.
@@ -4,6 +4,8 @@
schema_version: "1"
meeting:
# Stable identifier used for provenance across corrections and later runs.
meeting_id: ""
# Human-readable title for the meeting.
title: ""
# Dominant meeting language, for example "de" or "en".
@@ -28,6 +30,12 @@ participants:
attendance_status: "present"
notes: null
# Optional authoritative mapping from diarization labels to actual participants.
# Add entries only after a human or trusted external process confirms identity.
# Never infer mappings from conversational context. Unmapped labels stay anonymous.
speaker_mappings: {}
# SPEAKER_00: "participant-id"
mentioned_people:
# People discussed or referenced but not present in the meeting.
# Mentioned people are not speakers and must not become responsible persons
@@ -37,7 +45,7 @@ mentioned_people:
aliases: []
role: null
department: null
attendance_status: "not_present"
attendance_status: "mentioned_only"
notes: null
organization:
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the collective-commitment Gold experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_collective import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,7 @@
#!/usr/bin/env python3
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
if str(ROOT) not in sys.path: sys.path.insert(0, str(ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection_v1 import main
if __name__ == "__main__": raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the H-only controlled derivation experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_h import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
+153
View File
@@ -0,0 +1,153 @@
#!/usr/bin/env python3
"""Run the one-call direct protocol MVP from compact Whisper JSON."""
from __future__ import annotations
import argparse
import json
import re
import shutil
import sys
import time
from datetime import datetime
from pathlib import Path
from typing import Any, Callable
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT # noqa: E402
from src.meeting_lab.protocol.generate_direct_protocol import ( # noqa: E402
DEFAULT_MODEL,
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
DirectProtocolResult,
generate_direct_protocol,
)
DEFAULT_OUTPUT_ROOT = Path("meeting_data/runs")
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Generate one direct protocol from compact Whisper JSON.")
parser.add_argument("transcript", type=Path)
parser.add_argument("--context", type=Path)
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument(
"--safe-input-token-budget",
type=int,
default=DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
)
return parser.parse_args(argv)
def create_unique_run_dir(
output_root: Path,
transcript_stem: str,
now: Callable[[], datetime] = datetime.now,
) -> Path:
safe_stem = re.sub(r"[^A-Za-z0-9_.-]+", "_", transcript_stem).strip("._-") or "meeting"
base = output_root / f"{safe_stem}_{now().strftime('%Y%m%d_%H%M%S')}"
candidate = base
suffix = 1
while candidate.exists():
candidate = output_root / f"{base.name}_{suffix:02d}"
suffix += 1
candidate.mkdir(parents=True)
return candidate
def write_json(path: Path, data: Any) -> None:
path.write_text(json.dumps(data, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def persist_result(run_dir: Path, result: DirectProtocolResult) -> Path:
protocol_dir = run_dir / "protocol"
protocol_dir.mkdir()
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
write_json(protocol_dir / "raw_response.json", result.raw_response)
write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
transcript_input = getattr(result, "transcript_input", None)
if transcript_input is not None:
(protocol_dir / "transcript_input.txt").write_text(
transcript_input, encoding="utf-8"
)
protocol_path = run_dir / "protocol.md"
protocol_path.write_text(result.protocol_text, encoding="utf-8")
return protocol_path
def run(args: argparse.Namespace) -> tuple[int, Path, Path | None]:
run_dir = create_unique_run_dir(args.output_root, args.transcript.stem)
timestamp = datetime.now().astimezone().isoformat(timespec="seconds")
started = time.perf_counter()
protocol_path: Path | None = None
metadata: dict[str, Any] = {
"run_id": run_dir.name,
"timestamp": timestamp,
"transcript_path": str(args.transcript.resolve()),
"context_path": str(args.context.resolve()) if args.context else None,
"model": args.model,
"ollama_endpoint": args.ollama_endpoint,
"status": "running",
"total_runtime_seconds": None,
"final_protocol_path": None,
}
try:
transcript_dir = run_dir / "transcript"
transcript_dir.mkdir()
if not args.transcript.is_file():
raise FileNotFoundError(f"Transcript file does not exist: {args.transcript}")
preserved_transcript = transcript_dir / "transcript.json"
shutil.copy2(args.transcript, preserved_transcript)
preserved_context: Path | None = None
if args.context is not None:
if not args.context.is_file():
raise FileNotFoundError(f"Meeting Context file does not exist: {args.context}")
context_dir = run_dir / "context"
context_dir.mkdir()
preserved_context = context_dir / "meeting_context.yaml"
shutil.copy2(args.context, preserved_context)
write_json(
run_dir / "input_manifest.json",
{
"transcript_source": str(args.transcript.resolve()),
"transcript_copy": str(preserved_transcript.resolve()),
"context_source": str(args.context.resolve()) if args.context else None,
"context_copy": str(preserved_context.resolve()) if preserved_context else None,
},
)
result = generate_direct_protocol(
preserved_transcript,
preserved_context,
model=args.model,
endpoint=args.ollama_endpoint,
safe_input_token_budget=args.safe_input_token_budget,
)
protocol_path = persist_result(run_dir, result)
metadata["status"] = "completed"
metadata["final_protocol_path"] = str(protocol_path.resolve())
except Exception as exc:
metadata["status"] = "failed"
metadata["failure"] = f"{type(exc).__name__}: {exc}"
print(f"Error: {metadata['failure']}", file=sys.stderr)
finally:
metadata["total_runtime_seconds"] = round(time.perf_counter() - started, 3)
write_json(run_dir / "run_metadata.json", metadata)
return (0 if metadata["status"] == "completed" else 2), run_dir, protocol_path
def main(argv: list[str] | None = None) -> int:
code, _run_dir, protocol_path = run(parse_args(argv))
if protocol_path is not None:
print(protocol_path)
return code
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the evidence-near observation experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.evidence_observations.experiment import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for evidence-near observation experiment V2."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.evidence_observations_v2.experiment import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for evidence-near observation experiment V3."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.evidence_observations_v3.experiment import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the explicit-rejection Gold experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
+151
View File
@@ -0,0 +1,151 @@
#!/usr/bin/env python3
"""CLI adapter for the reusable Meeting Lab MVP orchestration API."""
from __future__ import annotations
import argparse
import json
import sys
from pathlib import Path
from typing import Any
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT # noqa: E402
from src.meeting_lab.models.meeting_context import MeetingContext # noqa: E402
from src.meeting_lab.orchestration.mvp import ( # noqa: E402
DEFAULT_DIARIZATION_MODEL,
DEFAULT_MODEL,
DEFAULT_OUTPUT_ROOT,
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
MvpMeetingConfig,
create_unique_run_dir,
run_mvp_meeting,
)
from src.meeting_lab.progress import ProgressSink # noqa: E402
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Transcribe one meeting and generate one direct protocol."
)
parser.add_argument("audio_file", type=Path)
parser.add_argument("--whisper-model", type=Path, required=True)
parser.add_argument("--whisper-executable", default="whisper-cli")
parser.add_argument("--ffmpeg-executable", default="ffmpeg")
parser.add_argument(
"--audio-normalization",
action=argparse.BooleanOptionalAction,
default=True,
help=(
"Enable FFmpeg loudness normalization during canonical audio preparation "
"(default: enabled)."
),
)
parser.add_argument("--context", type=Path)
parser.add_argument("--output-root", type=Path, default=DEFAULT_OUTPUT_ROOT)
parser.add_argument("--language", default="de")
parser.add_argument(
"--threads",
default="auto",
help="Thread count or 'auto' for physical CPU cores (default: auto).",
)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--ollama-endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument(
"--protocol-safe-input-token-budget",
type=int,
default=DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
help="Conservative estimated prompt-token limit before any Ollama request.",
)
parser.add_argument(
"--diarization",
choices=("auto", "gpu", "cpu", "off"),
default="off",
help="Optional Community-1 diarization device mode (default: off).",
)
parser.add_argument(
"--diarization-runtime",
choices=("native", "container"),
default="native",
help="Run pyannote in this Python environment or an explicit container.",
)
parser.add_argument(
"--diarization-container-image",
help="Container image required with --diarization-runtime container.",
)
parser.add_argument(
"--diarization-container-arg",
action="append",
default=[],
help="Additional docker argument; repeat and use = for values beginning with --.",
)
return parser.parse_args(argv)
def config_from_args(args: argparse.Namespace) -> MvpMeetingConfig:
return MvpMeetingConfig(
audio_file=args.audio_file,
whisper_model=args.whisper_model,
whisper_executable=args.whisper_executable,
ffmpeg_executable=args.ffmpeg_executable,
audio_normalization=args.audio_normalization,
context_file=args.context,
output_root=args.output_root,
language=args.language,
threads=args.threads,
model=args.model,
ollama_endpoint=args.ollama_endpoint,
protocol_safe_input_token_budget=args.protocol_safe_input_token_budget,
diarization=args.diarization,
diarization_runtime=args.diarization_runtime,
diarization_container_image=args.diarization_container_image,
diarization_container_args=tuple(args.diarization_container_arg),
)
def run(
args: argparse.Namespace,
*,
context_override: MeetingContext | dict[str, Any] | None = None,
progress_sink: ProgressSink | None = None,
) -> tuple[int, Path | None, Path | None]:
"""Compatibility wrapper for existing Python callers of the CLI module."""
result = run_mvp_meeting(
config_from_args(args),
meeting_context=context_override,
progress_sink=progress_sink,
)
return result.exit_code, result.run_dir, result.protocol_path
def main(argv: list[str] | None = None) -> int:
args = parse_args(argv)
if args.diarization == "off":
print("Diarization: disabled")
else:
print(
f"Diarization: enabled; backend=pyannote.audio; "
f"model={DEFAULT_DIARIZATION_MODEL}; requested_device={args.diarization}; "
f"runtime={args.diarization_runtime}"
)
code, run_dir, protocol_path = run(args)
if run_dir is not None and args.diarization != "off":
metadata_path = run_dir / "diarization" / "metadata.json"
if metadata_path.is_file():
details = json.loads(metadata_path.read_text(encoding="utf-8"))
print(
f"Diarization result: device={details.get('actual_device')}; "
f"device_name={details.get('device_name') or 'n/a'}; "
f"runtime={details.get('runtime_seconds'):.3f}s; "
f"speakers={details.get('speaker_count')}; artifacts={metadata_path.parent}"
)
if protocol_path is not None:
print(protocol_path)
return code
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the Negative Act Form experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_negative_act import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the request/acceptance Gold experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_gold import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the isolated semantic synthesis experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.semantic_synthesis.experiment import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,7 @@
#!/usr/bin/env python3
import sys
from pathlib import Path
ROOT=Path(__file__).resolve().parents[1]
if str(ROOT) not in sys.path: sys.path.insert(0,str(ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_target_normalization import main
if __name__=="__main__": raise SystemExit(main())
@@ -0,0 +1,7 @@
#!/usr/bin/env python3
import sys
from pathlib import Path
ROOT = Path(__file__).resolve().parents[1]
if str(ROOT) not in sys.path: sys.path.insert(0, str(ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution import main
if __name__ == "__main__": raise SystemExit(main())
@@ -0,0 +1,7 @@
#!/usr/bin/env python3
import sys
from pathlib import Path
ROOT=Path(__file__).resolve().parents[1]
if str(ROOT) not in sys.path: sys.path.insert(0,str(ROOT))
from src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution_v1 import main
if __name__=="__main__": raise SystemExit(main())
@@ -0,0 +1,16 @@
#!/usr/bin/env python3
"""Repository entry point for the isolated topic reconstruction experiment."""
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.topic_reconstruction.experiment import main # noqa: E402
if __name__ == "__main__":
raise SystemExit(main())
+51
View File
@@ -0,0 +1,51 @@
#!/usr/bin/env python3
"""Transcribe one audio file with whisper.cpp; do not generate a protocol."""
from __future__ import annotations
import argparse
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
if str(REPO_ROOT) not in sys.path:
sys.path.insert(0, str(REPO_ROOT))
from src.meeting_lab.transcription.whisper import ( # noqa: E402
TranscriptionError,
transcribe_audio,
)
def parse_args(argv: list[str] | None = None) -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Create a compact Meeting Lab transcript with whisper.cpp.")
parser.add_argument("audio_file", type=Path)
parser.add_argument("--model", type=Path, required=True, help="Path to a whisper.cpp GGML model.")
parser.add_argument("--output-dir", type=Path, required=True)
parser.add_argument("--language", default="auto", help="Language code or 'auto' (default: auto).")
parser.add_argument("--threads", default="auto", help="Thread count or 'auto' for physical CPU cores (default: auto).")
parser.add_argument("--whisper-executable", default="whisper-cli", help="whisper.cpp CLI executable (default: whisper-cli).")
return parser.parse_args(argv)
def main(argv: list[str] | None = None) -> int:
args = parse_args(argv)
try:
result = transcribe_audio(
args.audio_file,
args.model,
args.output_dir,
args.language,
executable=args.whisper_executable,
threads=args.threads,
)
except TranscriptionError as exc:
print(f"Error: {exc}")
return 1
print(f"Transcript: {result.transcript_json}")
print(f"Runtime: {result.runtime_seconds:.3f} seconds")
return 0
if __name__ == "__main__":
raise SystemExit(main())
+5
View File
@@ -0,0 +1,5 @@
"""Canonical audio preparation boundary."""
from .preparation import AudioPreparationError, PreparedAudio, prepare_audio
__all__ = ["AudioPreparationError", "PreparedAudio", "prepare_audio"]
+190
View File
@@ -0,0 +1,190 @@
"""Prepare supported recordings for deterministic downstream processing."""
from __future__ import annotations
import os
import shutil
import subprocess
import wave
from collections.abc import Callable, Sequence
from dataclasses import dataclass
from pathlib import Path
SUPPORTED_EXTENSIONS = {".wav", ".flac", ".m4a"}
CANONICAL_SAMPLE_RATE = 16_000
CANONICAL_CHANNELS = 1
CANONICAL_SAMPLE_WIDTH_BYTES = 2
CANONICAL_CODEC = "pcm_s16le"
DEFAULT_NORMALIZATION_FILTER = "loudnorm=I=-16:LRA=11:TP=-1.5"
DEFAULT_NORMALIZATION_METHOD = "ffmpeg_loudnorm"
class AudioPreparationError(RuntimeError):
"""Raised when source audio cannot be prepared as canonical WAV."""
@dataclass(frozen=True)
class PreparedAudio:
source_path: Path
source_format: str
prepared_path: Path
method: str
ffmpeg_executable: str
normalization_enabled: bool = True
normalization_method: str | None = DEFAULT_NORMALIZATION_METHOD
normalization_filter: str | None = DEFAULT_NORMALIZATION_FILTER
def metadata(self) -> dict[str, object]:
return {
"original_source_path": str(self.source_path.resolve()),
"original_source_name": self.source_path.name,
"original_format": self.source_format,
"prepared_audio_path": str(self.prepared_path.resolve()),
"preparation_method": self.method,
"ffmpeg_executable": self.ffmpeg_executable,
"normalization_enabled": self.normalization_enabled,
"normalization_method": self.normalization_method,
"normalization_filter": self.normalization_filter,
"canonical_output": {
"container": "wav",
"codec": CANONICAL_CODEC,
"channels": CANONICAL_CHANNELS,
"sample_rate_hz": CANONICAL_SAMPLE_RATE,
"bits_per_sample": CANONICAL_SAMPLE_WIDTH_BYTES * 8,
},
}
Runner = Callable[..., subprocess.CompletedProcess[str]]
def prepare_audio(
source_path: Path,
prepared_path: Path,
*,
ffmpeg_executable: str = "ffmpeg",
normalization_enabled: bool = True,
runner: Runner = subprocess.run,
) -> PreparedAudio:
"""Create and validate a canonical mono 16 kHz signed PCM16 WAV artifact."""
source_path = Path(source_path)
prepared_path = Path(prepared_path)
source_format = source_path.suffix.lower()
if not source_path.is_file():
raise AudioPreparationError(f"Source audio does not exist: {source_path}")
if source_format not in SUPPORTED_EXTENSIONS:
supported = ", ".join(sorted(SUPPORTED_EXTENSIONS))
raise AudioPreparationError(
f"Unsupported audio format {source_format or '<none>'!r}; supported: {supported}."
)
resolved_executable = _resolve_executable(ffmpeg_executable)
prepared_path.parent.mkdir(parents=True, exist_ok=True)
temporary_path = prepared_path.with_name(f".{prepared_path.name}.tmp.wav")
command_parts = [
resolved_executable,
"-nostdin",
"-hide_banner",
"-loglevel",
"error",
"-y",
"-i",
str(source_path),
"-map_metadata",
"-1",
"-vn",
]
if normalization_enabled:
command_parts.extend(("-af", DEFAULT_NORMALIZATION_FILTER))
command_parts.extend(
(
"-ac",
str(CANONICAL_CHANNELS),
"-ar",
str(CANONICAL_SAMPLE_RATE),
"-c:a",
CANONICAL_CODEC,
"-fflags",
"+bitexact",
str(temporary_path),
)
)
command: Sequence[str] = tuple(command_parts)
try:
completed = runner(command, capture_output=True, text=True, check=False)
except OSError as exc:
raise AudioPreparationError(f"Could not run FFmpeg: {exc}") from exc
if completed.returncode != 0:
detail = (
completed.stderr or completed.stdout or "no diagnostic output"
).strip()
raise AudioPreparationError(
f"FFmpeg failed to prepare {source_path.name} (exit {completed.returncode}): "
f"{detail}"
)
try:
_validate_canonical_wav(temporary_path)
os.replace(temporary_path, prepared_path)
except Exception:
temporary_path.unlink(missing_ok=True)
raise
return PreparedAudio(
source_path=source_path,
source_format=source_format.removeprefix("."),
prepared_path=prepared_path,
method="ffmpeg",
ffmpeg_executable=resolved_executable,
normalization_enabled=normalization_enabled,
normalization_method=(
DEFAULT_NORMALIZATION_METHOD if normalization_enabled else None
),
normalization_filter=(
DEFAULT_NORMALIZATION_FILTER if normalization_enabled else None
),
)
def _resolve_executable(executable: str) -> str:
value = executable.strip()
if not value:
raise AudioPreparationError("FFmpeg executable must not be empty.")
if Path(value).parent != Path("."):
path = Path(value)
if path.is_file() and os.access(path, os.X_OK):
return str(path)
raise AudioPreparationError(f"FFmpeg executable is not available: {value}")
resolved = shutil.which(value)
if resolved is None:
raise AudioPreparationError(
f"FFmpeg executable {value!r} was not found on PATH. Install FFmpeg or "
"configure its executable path."
)
return resolved
def _validate_canonical_wav(path: Path) -> None:
try:
with wave.open(str(path), "rb") as recording:
properties = (
recording.getnchannels(),
recording.getframerate(),
recording.getsampwidth(),
recording.getcomptype(),
)
except (OSError, EOFError, wave.Error) as exc:
raise AudioPreparationError(
f"FFmpeg did not produce a readable WAV file: {path}: {exc}"
) from exc
expected = (
CANONICAL_CHANNELS,
CANONICAL_SAMPLE_RATE,
CANONICAL_SAMPLE_WIDTH_BYTES,
"NONE",
)
if properties != expected:
raise AudioPreparationError(
"Prepared audio is not canonical mono 16 kHz PCM16 WAV: "
f"channels={properties[0]}, sample_rate={properties[1]}, "
f"sample_width={properties[2]}, compression={properties[3]}."
)
@@ -0,0 +1 @@
"""Isolated controlled semantic derivation experiments."""
@@ -0,0 +1,349 @@
#!/usr/bin/env python3
"""Isolated collective-commitment Gold reliability experiment."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
from .experiment_h import (
DEFAULT_ENDPOINT,
DEFAULT_MODEL,
DerivationValidationError,
OBSERVATION_KEYS,
call_ollama,
)
GOLD_SCHEMA_VERSION = "experimental-collective-commitment-gold-v0"
RECOGNITION_KEYS = {"observation_id", "commitment_form", "normalized_action_text"}
COMMITMENT_FORMS = {"individual_first_person", "collective_first_person", "none"}
FORBIDDEN_LLM_KEYS = {
"responsible_person", "responsibility", "responsibility_scope",
"requested_actor", "owner", "ownership", "assignee", "status",
"established", "action_item", "protocol", "protocol_category", "decision",
"unresolved_issue", "confidence", "relation", "relations", "graph",
}
WEEKDAYS = {
"monday": "Montag", "montag": "Montag", "tuesday": "Dienstag",
"dienstag": "Dienstag", "wednesday": "Mittwoch", "mittwoch": "Mittwoch",
"thursday": "Donnerstag", "donnerstag": "Donnerstag", "friday": "Freitag",
"freitag": "Freitag", "saturday": "Samstag", "samstag": "Samstag",
"sunday": "Sonntag", "sonntag": "Sonntag",
}
PROMPT_TEMPLATE = """Recognize only the explicit first-person commitment form and concise action meaning in the supplied single V3-style observation.
Answer only:
1. What explicit first-person commitment form is present?
- individual_first_person: the speaker explicitly commits themself personally.
- collective_first_person: the speaker explicitly commits a "we" group.
- none: there is no explicit first-person commitment.
2. What is the concise normalized action meaning?
Tentative possibility is not commitment. Suggestion or recommendation is not commitment. Impersonal necessity is not commitment. Passive future wording is not commitment. Rejection or negation is not positive commitment. Speaker identity does not convert collective "we" into individual commitment.
Preserve material limitations such as "nur im Technikum", "nur als Versuch", or "nur 20 Meter" in normalized_action_text. Keep normalized action text in the observation language. When commitment_form is not "none", normalized_action_text must be a non-empty string. When commitment_form is "none", normalized_action_text may be a non-empty action meaning or null.
Do not infer who is responsible. Do not decide whether an action is established. Do not output responsibility, responsibility scope, requested actor, owner, assignee, status, established, Action Item, protocol, confidence, semantic relations, graphs, decisions, or unresolved issues.
Return exactly this JSON shape and no additional fields:
{{
"observation_id": "observation ID",
"commitment_form": "individual_first_person | collective_first_person | none",
"normalized_action_text": "concise action meaning" | null
}}
V3-style observation:
{observation_json}
"""
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing = required - value.keys()
unknown = value.keys() - required
if missing:
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _nonempty_text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise DerivationValidationError(f"{location} must be a non-empty string")
return value.strip()
def _validate_observations(observations: Any) -> None:
if not isinstance(observations, list) or not observations:
raise DerivationValidationError("observations must be a non-empty list")
seen_observations: set[str] = set()
seen_evidence: set[str] = set()
for index, observation in enumerate(observations):
location = f"observations[{index}]"
if not isinstance(observation, dict):
raise DerivationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
if observation_id in seen_observations or evidence_id in seen_evidence:
raise DerivationValidationError("observation and evidence provenance must be unique")
seen_observations.add(observation_id)
seen_evidence.add(evidence_id)
_nonempty_text(observation["content"], f"{location}.content")
_nonempty_text(observation["speaker"], f"{location}.speaker")
for field in ("named_person", "addressee"):
if observation[field] is not None:
_nonempty_text(observation[field], f"{location}.{field}")
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict):
raise DerivationValidationError("Gold fixture must be an object")
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
if data["schema_version"] != GOLD_SCHEMA_VERSION:
raise DerivationValidationError("unexpected Gold fixture schema_version")
cases = data["cases"]
if not isinstance(cases, list) or not cases:
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
seen: set[str] = set()
for case in cases:
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
if case_id in seen:
raise DerivationValidationError(f"duplicate case ID: {case_id}")
seen.add(case_id)
_validate_observations(case["observations"])
if len(case["observations"]) != 1:
raise DerivationValidationError("collective Gold cases require exactly one observation")
return cases
def build_prompt(observations: list[dict[str, Any]]) -> str:
_validate_observations(observations)
if len(observations) != 1:
raise DerivationValidationError("collective recognition requires exactly one observation")
return PROMPT_TEMPLATE.format(
observation_json=json.dumps(observations[0], ensure_ascii=False, indent=2)
)
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
return data
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
if isinstance(value, dict):
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
if forbidden:
raise DerivationValidationError(
f"{location} contains forbidden semantic keys: {sorted(forbidden)}"
)
for key, item in value.items():
_reject_forbidden_keys(item, f"{location}.{key}")
elif isinstance(value, list):
for index, item in enumerate(value):
_reject_forbidden_keys(item, f"{location}[{index}]")
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
_validate_observations(observations)
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
_reject_forbidden_keys(data)
_exact_keys(data, RECOGNITION_KEYS, "output")
observation_id = _nonempty_text(data["observation_id"], "output.observation_id")
if observation_id not in {item["observation_id"] for item in observations}:
raise DerivationValidationError("recognition references unknown observation")
if data["commitment_form"] not in COMMITMENT_FORMS:
raise DerivationValidationError("commitment_form has an unsupported value")
action_text = data["normalized_action_text"]
if action_text is not None:
_nonempty_text(action_text, "output.normalized_action_text")
if data["commitment_form"] != "none" and action_text is None:
raise DerivationValidationError("non-none commitment requires normalized_action_text")
return data
def _bounded_due(observations: list[dict[str, Any]]) -> tuple[str | None, bool]:
due_forms: set[str] = set()
for observation in observations:
content = observation["content"].casefold()
for token in re.findall(r"\b[A-Za-zÄÖÜäöü]+\b", content):
if token in WEEKDAYS:
due_forms.add(WEEKDAYS[token])
if re.search(r"\bnächste\s+woche\b", content):
due_forms.add("nächste Woche")
return (next(iter(due_forms)) if len(due_forms) == 1 else None, len(due_forms) <= 1)
def _has_explicit_negation(observations: list[dict[str, Any]]) -> bool:
return any(
re.search(r"\b(?:nicht|kein(?:e|en|er|es)?|nein|not|no)\b", item["content"], re.IGNORECASE)
for item in observations
)
def derive_collective_action(
observations: list[dict[str, Any]], recognition: dict[str, Any]
) -> tuple[dict[str, bool], dict[str, Any] | None]:
validate_recognition(recognition, observations)
by_id = {item["observation_id"]: item for item in observations}
observation = by_id.get(recognition["observation_id"])
due, deadline_consistent = _bounded_due(observations)
gates = {
"recognition_schema_valid": True,
"observation_exists": observation is not None,
"provenance_valid_and_unique": observation is not None and len({item["evidence_id"] for item in observations}) == len(observations),
"collective_commitment_form": recognition["commitment_form"] == "collective_first_person",
"normalized_action_present": isinstance(recognition["normalized_action_text"], str) and bool(recognition["normalized_action_text"].strip()),
"deadline_supported_and_consistent": deadline_consistent,
"no_explicit_negation": not _has_explicit_negation(observations),
}
if not all(gates.values()):
return gates, None
return gates, {
"action_id": "action_1",
"content": recognition["normalized_action_text"].strip(),
"status": "established",
"commitment_scope": "collective",
"responsible_person": None,
"due": due,
"support": {
"commitment": {
"observation_id": observation["observation_id"],
"evidence_id": observation["evidence_id"],
}
},
}
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
if not concepts:
return True
if not isinstance(text, str):
return False
folded = text.casefold()
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
validate_recognition(recognition, case["observations"])
gates, result = derive_collective_action(case["observations"], recognition)
expected_recognition = case["expected_recognition"]
expected_result = case["expected_result"]
form_correct = recognition["commitment_form"] == expected_recognition["commitment_form"]
action_correct = _concepts_present(recognition["normalized_action_text"], expected_recognition["action_concepts"])
qualifier_preserved = _concepts_present(recognition["normalized_action_text"], expected_recognition["qualifier_concepts"])
established = result is not None
final_correct = established == expected_result["established"]
if result is not None:
final_correct = final_correct and result["due"] == expected_result["due"] and result["responsible_person"] is None and result["commitment_scope"] == "collective"
owner_correct = result is None or result["responsible_person"] is None
automatic_failure = (established and not expected_result["established"]) or not owner_correct or (established and not qualifier_preserved)
semantic_correct = form_correct and action_correct and qualifier_preserved
classification = "FAIL" if automatic_failure or not final_correct else ("PASS" if semantic_correct else "PARTIAL")
return {
"case_id": case["case_id"], "classification": classification,
"commitment_form_correct": form_correct,
"normalized_action_meaning_correct": action_correct,
"material_qualifier_preserved": qualifier_preserved,
"deterministic_gates_correct": final_correct,
"final_result_correct": final_correct,
"responsible_person_correctly_null": owner_correct,
"due_correct": result is None or result["due"] == expected_result["due"],
"unsupported_semantic_strengthening": recognition["commitment_form"] == "collective_first_person" and expected_recognition["commitment_form"] != "collective_first_person",
"responsibility_status_leakage": False,
"gates": gates, "result": result,
}
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
cases = load_gold_cases(args.cases)
args.output.mkdir(parents=True, exist_ok=False)
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
evaluations: list[dict[str, Any]] = []
successful_calls = 0
technical_failures = 0
started = time.perf_counter()
for case in cases:
case_dir = args.output / case["case_id"].lower()
case_dir.mkdir()
observations = case["observations"]
_write_json(case_dir / "v3_style_input_observations.json", observations)
prompt = build_prompt(observations)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
try:
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
successful_calls += 1
except Exception as exc: # one recorded attempt; never retry
technical_failures += 1
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
_write_json(case_dir / "deterministic_gate_results.json", {})
_write_json(case_dir / "final_derived_result.json", None)
_write_json(case_dir / "evaluation.json", failure)
evaluations.append(failure)
continue
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
try:
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
evaluation = evaluate_case(case, parsed)
validation = {"valid": True, "error": None}
gates, result = derive_collective_action(observations, parsed)
except (DerivationValidationError, json.JSONDecodeError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "responsibility_status_leakage": "forbidden" in str(exc)}
gates, result = {}, None
_write_json(case_dir / "structural_validation.json", validation)
_write_json(case_dir / "deterministic_gate_results.json", gates)
_write_json(case_dir / "final_derived_result.json", result)
_write_json(case_dir / "evaluation.json", evaluation)
evaluations.append(evaluation)
summary = {
"experiment": "collective_commitment_gold_v0", "model": args.model,
"successful_llm_call_count": successful_calls,
"technical_failed_call_count": technical_failures,
"runtime_seconds": round(time.perf_counter() - started, 3),
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
"evaluations": evaluations,
}
_write_json(args.output / "summary.json", summary)
return summary
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run isolated collective-commitment Gold experiment")
parser.add_argument("cases", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=1024)
return parser.parse_args()
def main() -> int:
summary = run_gold(parse_args())
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,367 @@
#!/usr/bin/env python3
"""Isolated request/acceptance Gold reliability experiment."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
from .experiment_h import (
DEFAULT_ENDPOINT,
DEFAULT_MODEL,
DerivationValidationError,
FORBIDDEN_LLM_KEYS,
OBSERVATION_KEYS,
call_ollama,
)
GOLD_SCHEMA_VERSION = "experimental-request-acceptance-gold-v0"
RECOGNITION_SCHEMA_VERSION = "experimental-request-acceptance-recognition-v0"
REQUEST_KEYS = {"observation_id", "is_concrete_request", "normalized_action_text"}
ACCEPTANCE_KEYS = {
"observation_id", "is_explicit_commitment", "same_requested_work",
"normalized_action_text",
}
WEEKDAYS = {
"monday": "Montag", "montag": "Montag", "tuesday": "Dienstag",
"dienstag": "Dienstag", "wednesday": "Mittwoch", "mittwoch": "Mittwoch",
"thursday": "Donnerstag", "donnerstag": "Donnerstag", "friday": "Freitag",
"freitag": "Freitag", "saturday": "Samstag", "samstag": "Samstag",
"sunday": "Sonntag", "sonntag": "Sonntag",
}
PROMPT_TEMPLATE = """Recognize only a concrete directed request and a later explicit personal commitment in the supplied V3-style observations.
The input is observations only, not a transcript. Identify:
1. A concrete request directed to the observation's explicit addressee, if one exists.
2. A later response that explicitly commits its speaker to work, if one exists.
3. Whether that explicit commitment concerns substantially the same requested work.
Lexical identity is not required: a contextual paraphrase may denote the same work. Mere acknowledgement, tentative or conditional language, collective "we" statements, impersonal necessity, suggestions, and statements that work should be done are not explicit personal commitments. A commitment to different work is an explicit commitment but not the same requested work.
Do not decide or output responsibility, requested actor, established status, Action Item status, protocol eligibility, confidence, semantic relations, or graphs. Do not answer who is responsible. Deterministic code applies those gates later.
Return exactly this JSON shape and no other fields. Use null for request or acceptance when no qualifying observation exists:
{{
"schema_version": "experimental-request-acceptance-recognition-v0",
"request": null | {{
"observation_id": "observation ID",
"is_concrete_request": true,
"normalized_action_text": "concise requested work"
}},
"acceptance": null | {{
"observation_id": "observation ID",
"is_explicit_commitment": true,
"same_requested_work": true,
"normalized_action_text": "concise committed work"
}}
}}
V3-style observations:
{observations_json}
"""
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing = required - value.keys()
unknown = value.keys() - required
if missing:
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _nonempty_text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise DerivationValidationError(f"{location} must be a non-empty string")
return value.strip()
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict):
raise DerivationValidationError("Gold fixture must be an object")
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
if data["schema_version"] != GOLD_SCHEMA_VERSION:
raise DerivationValidationError("unexpected Gold fixture schema_version")
cases = data["cases"]
if not isinstance(cases, list) or not cases:
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
seen_cases: set[str] = set()
for case in cases:
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
case_id = _nonempty_text(case["case_id"], "case_id")
if case_id in seen_cases:
raise DerivationValidationError(f"duplicate case ID: {case_id}")
seen_cases.add(case_id)
_validate_observations(case["observations"])
return cases
def _validate_observations(observations: Any) -> None:
if not isinstance(observations, list) or not observations:
raise DerivationValidationError("observations must be a non-empty list")
seen_ids: set[str] = set()
seen_evidence: set[str] = set()
for index, observation in enumerate(observations):
location = f"observations[{index}]"
if not isinstance(observation, dict):
raise DerivationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
if observation_id in seen_ids or evidence_id in seen_evidence:
raise DerivationValidationError("observation and evidence IDs must be unique")
seen_ids.add(observation_id)
seen_evidence.add(evidence_id)
_nonempty_text(observation["content"], f"{location}.content")
_nonempty_text(observation["speaker"], f"{location}.speaker")
for field in ("named_person", "addressee"):
if observation[field] is not None:
_nonempty_text(observation[field], f"{location}.{field}")
def build_prompt(observations: list[dict[str, Any]]) -> str:
_validate_observations(observations)
return PROMPT_TEMPLATE.format(
observations_json=json.dumps(observations, ensure_ascii=False, indent=2)
)
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
return data
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
if isinstance(value, dict):
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
if forbidden:
raise DerivationValidationError(
f"{location} contains forbidden semantic keys: {sorted(forbidden)}"
)
for key, item in value.items():
_reject_forbidden_keys(item, f"{location}.{key}")
elif isinstance(value, list):
for index, item in enumerate(value):
_reject_forbidden_keys(item, f"{location}[{index}]")
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
_reject_forbidden_keys(data)
_exact_keys(data, {"schema_version", "request", "acceptance"}, "output")
if data["schema_version"] != RECOGNITION_SCHEMA_VERSION:
raise DerivationValidationError("unexpected recognition schema_version")
known_ids = {item["observation_id"] for item in observations}
request = data["request"]
acceptance = data["acceptance"]
if request is not None:
if not isinstance(request, dict):
raise DerivationValidationError("output.request must be an object or null")
_exact_keys(request, REQUEST_KEYS, "output.request")
if request["observation_id"] not in known_ids:
raise DerivationValidationError("request references unknown observation")
if not isinstance(request["is_concrete_request"], bool):
raise DerivationValidationError("is_concrete_request must be boolean")
_nonempty_text(request["normalized_action_text"], "request.normalized_action_text")
if acceptance is not None:
if not isinstance(acceptance, dict):
raise DerivationValidationError("output.acceptance must be an object or null")
_exact_keys(acceptance, ACCEPTANCE_KEYS, "output.acceptance")
if acceptance["observation_id"] not in known_ids:
raise DerivationValidationError("acceptance references unknown observation")
for field in ("is_explicit_commitment", "same_requested_work"):
if not isinstance(acceptance[field], bool):
raise DerivationValidationError(f"{field} must be boolean")
_nonempty_text(acceptance["normalized_action_text"], "acceptance.normalized_action_text")
if request is not None and acceptance is not None and request["observation_id"] == acceptance["observation_id"]:
raise DerivationValidationError("request and acceptance must reference different observations")
return data
def _bounded_due(observations: list[dict[str, Any]]) -> tuple[str | None, bool]:
forms: set[str] = set()
for observation in observations:
for token in re.findall(r"\b[A-Za-zÄÖÜäöü]+\b", observation["content"].casefold()):
if token in WEEKDAYS:
forms.add(WEEKDAYS[token])
return (next(iter(forms)) if len(forms) == 1 else None, len(forms) <= 1)
def _strip_due(action_text: str) -> str:
weekday = "|".join(re.escape(value) for value in WEEKDAYS)
result = re.sub(rf"\s+(?:bis|by)\s+(?:{weekday})\b", "", action_text, flags=re.IGNORECASE)
return result.strip(" .,:;-") or action_text.strip()
def derive_action(
observations: list[dict[str, Any]], recognition: dict[str, Any]
) -> tuple[dict[str, bool], dict[str, Any] | None]:
_validate_observations(observations)
validate_recognition(recognition, observations)
by_id = {item["observation_id"]: item for item in observations}
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
request_semantic = recognition["request"]
acceptance_semantic = recognition["acceptance"]
request = by_id.get(request_semantic["observation_id"]) if request_semantic else None
acceptance = by_id.get(acceptance_semantic["observation_id"]) if acceptance_semantic else None
due, deadline_consistent = _bounded_due(observations)
gates = {
"request_semantic_positive": request_semantic is not None and request_semantic["is_concrete_request"] is True,
"request_observation_exists": request is not None,
"request_has_addressee": request is not None and isinstance(request["addressee"], str) and bool(request["addressee"].strip()),
"acceptance_semantic_positive": acceptance_semantic is not None and acceptance_semantic["is_explicit_commitment"] is True,
"same_requested_work": acceptance_semantic is not None and acceptance_semantic["same_requested_work"] is True,
"acceptance_observation_exists": acceptance is not None,
"acceptance_after_request": request is not None and acceptance is not None and positions[acceptance["observation_id"]] > positions[request["observation_id"]],
"acceptance_speaker_matches_addressee": request is not None and acceptance is not None and acceptance["speaker"] == request["addressee"],
"provenance_valid_and_consistent": request is not None and acceptance is not None and request["evidence_id"] in {item["evidence_id"] for item in observations} and acceptance["evidence_id"] in {item["evidence_id"] for item in observations},
"deadline_consistent": deadline_consistent,
}
if not all(gates.values()):
return gates, None
return gates, {
"action_id": "action_1",
"content": _strip_due(request_semantic["normalized_action_text"]),
"status": "established",
"requested_actor": request["addressee"],
"responsible_person": acceptance["speaker"],
"due": due,
"support": {
"request": {"observation_id": request["observation_id"], "evidence_id": request["evidence_id"]},
"acceptance": {"observation_id": acceptance["observation_id"], "evidence_id": acceptance["evidence_id"]},
},
}
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
observations = case["observations"]
validate_recognition(recognition, observations)
gates, result = derive_action(observations, recognition)
expected_recognition = case["expected_recognition"]
request_correct = (recognition["request"] is not None and recognition["request"]["is_concrete_request"]) == expected_recognition["request"]
commitment_correct = (recognition["acceptance"] is not None and recognition["acceptance"]["is_explicit_commitment"]) == expected_recognition["commitment"]
same_work_correct = (recognition["acceptance"] is not None and recognition["acceptance"]["same_requested_work"]) == expected_recognition["same_work"]
expected = case["expected_result"]
actual_established = result is not None
final_correct = actual_established == expected["established"]
if result is not None:
final_correct = final_correct and all(
result[key] == expected[key]
for key in ("requested_actor", "responsible_person", "due")
)
else:
final_correct = final_correct and expected["responsible_person"] is None
semantic_correct = request_correct and commitment_correct and same_work_correct
classification = "PASS" if final_correct and semantic_correct else ("PARTIAL" if final_correct else "FAIL")
return {
"case_id": case["case_id"], "classification": classification,
"request_correct": request_correct, "commitment_correct": commitment_correct,
"same_work_correct": same_work_correct, "deterministic_gates_correct": final_correct,
"final_result_correct": final_correct, "result": result, "gates": gates,
"unsupported_semantic_strengthening": not semantic_correct and actual_established,
"responsibility_status_leakage": False,
}
def reevaluate_existing(args: argparse.Namespace) -> dict[str, Any]:
cases = load_gold_cases(args.cases)
evaluations = []
for case in cases:
case_dir = args.output / case["case_id"].lower()
parsed = json.loads(
(case_dir / "parsed_semantic_recognition.json").read_text(encoding="utf-8")
)
evaluation = evaluate_case(case, parsed)
_write_json(case_dir / "evaluation.json", evaluation)
evaluations.append(evaluation)
metadata = [
json.loads(
(args.output / case["case_id"].lower() / "ollama_metadata.json").read_text(
encoding="utf-8"
)
)
for case in cases
]
summary = {
"experiment": "request_acceptance_gold_v0", "model": args.model,
"llm_call_count": len(cases),
"runtime_seconds": round(sum(item["elapsed_seconds"] for item in metadata), 3),
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
"evaluations": evaluations,
}
_write_json(args.output / "summary.json", summary)
return summary
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
cases = load_gold_cases(args.cases)
args.output.mkdir(parents=True, exist_ok=False)
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
evaluations = []
total_started = time.perf_counter()
for case in cases:
case_dir = args.output / case["case_id"].lower()
case_dir.mkdir()
_write_json(case_dir / "v3_style_input_observations.json", case["observations"])
prompt = build_prompt(case["observations"])
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
try:
evaluation = evaluate_case(case, parsed)
validation = {"valid": True, "error": None}
except (DerivationValidationError, json.JSONDecodeError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "responsibility_status_leakage": "forbidden" in str(exc)}
_write_json(case_dir / "structural_validation.json", validation)
_write_json(case_dir / "evaluation.json", evaluation)
evaluations.append(evaluation)
summary = {
"experiment": "request_acceptance_gold_v0", "model": args.model,
"llm_call_count": len(cases), "runtime_seconds": round(time.perf_counter() - total_started, 3),
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
"evaluations": evaluations,
}
_write_json(args.output / "summary.json", summary)
return summary
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run isolated request/acceptance Gold experiment")
parser.add_argument("cases", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=1024)
parser.add_argument("--reevaluate-existing", action="store_true")
return parser.parse_args()
def main() -> int:
args = parse_args()
summary = reevaluate_existing(args) if args.reevaluate_existing else run_gold(args)
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["counts"]["FAIL"] == 0 else 1
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,320 @@
#!/usr/bin/env python3
"""H-only request/acceptance recognition and deterministic action derivation."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
import requests
SCHEMA_VERSION = "experimental-controlled-semantic-recognition-h-v0"
DEFAULT_MODEL = "qwen3.5:9B"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
EXPECTED_PROVENANCE = {"obs_1": "e1", "obs_2": "e2"}
OBSERVATION_KEYS = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
SEMANTIC_KEYS = {"schema_version", "request", "acceptance"}
REQUEST_KEYS = {"observation_id", "is_concrete_request", "normalized_action_text"}
ACCEPTANCE_KEYS = {"observation_id", "is_explicit_commitment", "same_requested_work", "normalized_action_text"}
FORBIDDEN_LLM_KEYS = {
"responsible_person", "responsibility", "requested_actor", "status", "established",
"action_item", "protocol_section", "protocol_category", "confidence", "relation",
"relations", "graph", "decision", "open_question", "unresolved_issue",
}
class DerivationValidationError(ValueError):
"""Raised when experiment input or LLM recognition violates the contract."""
PROMPT_TEMPLATE = """Recognize only two narrow semantic facts in the supplied V3 observations.
The input contains V3 observations only, not a transcript. Answer only:
1. Is obs_1 a concrete request directed to its recorded addressee?
2. Does obs_2 explicitly commit its speaker to substantially the same requested work?
Lexical identity is not required. Conversational paraphrases such as "Prüfung der
Messdaten" and "die Prüfung" may denote the same work when the supplied observation
sequence clearly supports that reading.
Do not decide or output responsibility, requested actor, established status, Action
Item status, protocol eligibility, confidence, semantic relations, or graphs. Do not
answer who is responsible. Deterministic code will apply those gates later.
Return exactly this JSON shape and no other fields:
{{
"schema_version": "experimental-controlled-semantic-recognition-h-v0",
"request": {{
"observation_id": "obs_1",
"is_concrete_request": true,
"normalized_action_text": "concise requested work in the observation language"
}},
"acceptance": {{
"observation_id": "obs_2",
"is_explicit_commitment": true,
"same_requested_work": true,
"normalized_action_text": "concise accepted work in the observation language"
}}
}}
V3 observations:
{observations_json}
"""
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run the H-only controlled semantic derivation experiment.")
parser.add_argument("observations", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=1024)
return parser.parse_args()
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing, unknown = required - value.keys(), value.keys() - required
if missing:
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise DerivationValidationError(f"{location} must be a non-empty string")
result = value.strip()
if result.casefold() == "null":
raise DerivationValidationError(f"{location} must not be the string 'null'")
return result
def load_v3_observations(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict):
raise DerivationValidationError("V3 input must be an object")
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "V3 input")
if data["schema_version"] != "experimental-evidence-observations-v3":
raise DerivationValidationError("V3 input has an unexpected schema_version")
observations = data["observations"]
if not isinstance(observations, list) or not observations:
raise DerivationValidationError("V3 observations must be a non-empty list")
seen: set[str] = set()
for index, observation in enumerate(observations):
location = f"V3 observations[{index}]"
if not isinstance(observation, dict):
raise DerivationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
if observation_id in seen:
raise DerivationValidationError(f"duplicate observation ID: {observation_id}")
seen.add(observation_id)
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
if observation_id not in EXPECTED_PROVENANCE:
raise DerivationValidationError(f"unknown H observation ID: {observation_id}")
if EXPECTED_PROVENANCE[observation_id] != evidence_id:
raise DerivationValidationError(f"inconsistent evidence provenance for {observation_id}")
_text(observation["content"], f"{location}.content")
_text(observation["speaker"], f"{location}.speaker")
for field in ("named_person", "addressee"):
if observation[field] is not None:
_text(observation[field], f"{location}.{field}")
if seen != set(EXPECTED_PROVENANCE):
raise DerivationValidationError("H input must contain exactly obs_1/e1 and obs_2/e2")
return observations
def build_prompt(observations: list[dict[str, Any]]) -> str:
validate_observation_sequence(observations)
return PROMPT_TEMPLATE.format(observations_json=json.dumps(observations, ensure_ascii=False, indent=2))
def validate_observation_sequence(observations: list[dict[str, Any]]) -> None:
if [item.get("observation_id") for item in observations] != ["obs_1", "obs_2"]:
raise DerivationValidationError("H observations must be ordered obs_1, obs_2")
for observation in observations:
if EXPECTED_PROVENANCE.get(observation.get("observation_id")) != observation.get("evidence_id"):
raise DerivationValidationError("H observation provenance is inconsistent")
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
return data
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
if isinstance(value, dict):
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
if forbidden:
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
for key, item in value.items():
_reject_forbidden_keys(item, f"{location}.{key}")
elif isinstance(value, list):
for index, item in enumerate(value):
_reject_forbidden_keys(item, f"{location}[{index}]")
def validate_semantic_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
_reject_forbidden_keys(data)
_exact_keys(data, SEMANTIC_KEYS, "output")
if data["schema_version"] != SCHEMA_VERSION:
raise DerivationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
request, acceptance = data["request"], data["acceptance"]
if not isinstance(request, dict) or not isinstance(acceptance, dict):
raise DerivationValidationError("request and acceptance must be objects")
_exact_keys(request, REQUEST_KEYS, "output.request")
_exact_keys(acceptance, ACCEPTANCE_KEYS, "output.acceptance")
known_ids = {item["observation_id"] for item in observations}
for location, item in (("output.request", request), ("output.acceptance", acceptance)):
observation_id = _text(item["observation_id"], f"{location}.observation_id")
if observation_id not in known_ids:
raise DerivationValidationError(f"{location} references unknown observation: {observation_id}")
_text(item["normalized_action_text"], f"{location}.normalized_action_text")
for field, value in (
("output.request.is_concrete_request", request["is_concrete_request"]),
("output.acceptance.is_explicit_commitment", acceptance["is_explicit_commitment"]),
("output.acceptance.same_requested_work", acceptance["same_requested_work"]),
):
if not isinstance(value, bool):
raise DerivationValidationError(f"{field} must be boolean")
if request["observation_id"] == acceptance["observation_id"]:
raise DerivationValidationError("request and acceptance must reference different observations")
return data
def _bounded_due(observations: list[dict[str, Any]]) -> tuple[str | None, bool]:
weekday_forms = {
"monday": "Montag", "montag": "Montag",
"tuesday": "Dienstag", "dienstag": "Dienstag",
"wednesday": "Mittwoch", "mittwoch": "Mittwoch",
"thursday": "Donnerstag", "donnerstag": "Donnerstag",
"friday": "Freitag", "freitag": "Freitag",
"saturday": "Samstag", "samstag": "Samstag",
"sunday": "Sonntag", "sonntag": "Sonntag",
}
forms: set[str] = set()
for observation in observations:
for token in re.findall(r"\b[A-Za-zÄÖÜäöü]+\b", observation["content"].casefold()):
if token in weekday_forms:
forms.add(weekday_forms[token])
return (next(iter(forms)) if len(forms) == 1 else None, len(forms) <= 1)
def _remove_bounded_due_from_action(action_text: str) -> str:
result = re.sub(
r"\s+(?:bis|by)\s+(?:Friday|Freitag)\b", "", action_text,
flags=re.IGNORECASE,
).strip(" .,:;-")
return result or action_text.strip()
def derive_action(
observations: list[dict[str, Any]], recognition: dict[str, Any]
) -> tuple[dict[str, bool], dict[str, Any] | None]:
by_id = {item["observation_id"]: item for item in observations}
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
request_semantic = recognition["request"]
acceptance_semantic = recognition["acceptance"]
request = by_id.get(request_semantic["observation_id"])
acceptance = by_id.get(acceptance_semantic["observation_id"])
due, deadline_consistent = _bounded_due(observations)
gates = {
"request_semantic_positive": request_semantic["is_concrete_request"] is True,
"request_observation_exists": request is not None,
"request_has_addressee": request is not None and isinstance(request.get("addressee"), str) and bool(request["addressee"].strip()),
"acceptance_semantic_positive": acceptance_semantic["is_explicit_commitment"] is True,
"same_requested_work": acceptance_semantic["same_requested_work"] is True,
"acceptance_observation_exists": acceptance is not None,
"acceptance_after_request": request is not None and acceptance is not None and positions[acceptance["observation_id"]] > positions[request["observation_id"]],
"acceptance_speaker_matches_addressee": request is not None and acceptance is not None and acceptance["speaker"] == request["addressee"],
"provenance_valid_and_consistent": request is not None and acceptance is not None and EXPECTED_PROVENANCE.get(request["observation_id"]) == request["evidence_id"] and EXPECTED_PROVENANCE.get(acceptance["observation_id"]) == acceptance["evidence_id"],
"deadline_consistent": deadline_consistent,
}
if not all(gates.values()):
return gates, None
action_text = _remove_bounded_due_from_action(
request_semantic["normalized_action_text"]
)
result = {
"action_id": "action_1",
"content": action_text,
"status": "established",
"requested_actor": request["addressee"],
"responsible_person": acceptance["speaker"],
"due": due,
"support": {
"request": {"observation_id": request["observation_id"], "evidence_id": request["evidence_id"]},
"acceptance": {"observation_id": acceptance["observation_id"], "evidence_id": acceptance["evidence_id"]},
},
}
return gates, result
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
started = time.perf_counter()
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
elapsed = time.perf_counter() - started
response.raise_for_status()
body = response.json()
raw = body.get("response") if isinstance(body, dict) else None
if not isinstance(raw, str) or not raw.strip():
raise ValueError("Ollama returned no usable response text")
metadata = {"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3), "total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"), "prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"), "eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"), "configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0}}
return raw.strip(), metadata
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
observations = load_v3_observations(args.observations)
args.output.mkdir(parents=True, exist_ok=False)
_write_json(args.output / "v3_input_observations.json", observations)
prompt = build_prompt(observations)
(args.output / "prompt.txt").write_text(prompt, encoding="utf-8")
started = time.perf_counter()
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
(args.output / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(args.output / "ollama_metadata.json", metadata)
parsed = parse_model_json(raw)
_write_json(args.output / "parsed_semantic_recognition.json", parsed)
try:
validate_semantic_recognition(parsed, observations)
validation = {"valid": True, "error": None}
gates, result = derive_action(observations, parsed)
except DerivationValidationError as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
gates, result = {}, None
_write_json(args.output / "structural_validation.json", validation)
_write_json(args.output / "deterministic_gate_results.json", gates)
_write_json(args.output / "final_derived_result.json", result)
summary = {"experiment": "controlled_semantic_derivation_h_v0", "model": args.model, "llm_call_count": 1, "runtime_seconds": round(time.perf_counter() - started, 3), "semantic_recognition_valid": validation["valid"], "all_gates_passed": bool(gates) and all(gates.values()), "action_established": result is not None}
_write_json(args.output / "summary.json", summary)
return summary
def main() -> int:
args = parse_args()
summary = run_experiment(args)
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["action_established"] else 1
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,277 @@
#!/usr/bin/env python3
"""Isolated evidence-near Negative Act Form classification experiment."""
from __future__ import annotations
import argparse
import json
import time
from pathlib import Path
from typing import Any
from .experiment_h import (
DEFAULT_ENDPOINT,
DEFAULT_MODEL,
DerivationValidationError,
OBSERVATION_KEYS,
build_ollama_payload,
call_ollama,
)
GOLD_SCHEMA_VERSION = "experimental-negative-act-form-gold-v0"
RECOGNITION_KEYS = {"observation_id", "negative_act_form", "normalized_action_text"}
NEGATIVE_ACT_FORMS = {
"explicit_non_pursuit", "personal_preference", "recommendation",
"temporary_non_action", "none",
}
FORBIDDEN_LLM_KEYS = {
"rejection_form", "explicitly_rejected", "status", "decision", "outcome",
"topic_status", "responsible_person", "responsibility", "owner",
"requested_actor", "action_item", "protocol", "protocol_category",
"confidence", "relation", "relations", "graph", "unresolved_issue",
}
PROMPT_TEMPLATE = """Classify only the negative semantic form expressed by the candidate observation, using earlier supplied V3-style observations only as local context for pronouns or shortened references.
The candidate observation is {candidate_observation_id}.
Choose exactly one negative_act_form:
- explicit_non_pursuit: explicitly states that an action, option, collaboration, or course will not be continued or pursued. This is stronger than preference, advice, or temporary delay.
- personal_preference: the speaker states what they personally would or would not do, without establishing collective non-pursuit.
- recommendation: the speaker advises for or against an action without establishing abandonment.
- temporary_non_action: the action is postponed, deferred, or explicitly not done for now without abandonment.
- none: none of those four forms is present, including mere concern, uncertainty, negative sentiment, or factual negation.
Do not collapse non-pursuit into temporary non-action. Do not convert a personal conditional preference into collective non-pursuit. Do not convert advice into non-pursuit. Speaker identity does not change personal preference into collective non-pursuit.
When the form is not none, return concise normalized action meaning. Resolve a pronoun only from the supplied local context. If its target is genuinely ambiguous, return none rather than guessing. When the form is none, normalized_action_text must be null. Keep normalized action text in the observation language.
Do not derive or output rejection, status, decision, outcome, topic closure, responsibility, ownership, Action Item, protocol category, confidence, relations, graphs, or unresolved issues.
Return exactly this JSON shape and no additional fields:
{{
"observation_id": "{candidate_observation_id}",
"negative_act_form": "explicit_non_pursuit | personal_preference | recommendation | temporary_non_action | none",
"normalized_action_text": "concise action meaning" | null
}}
V3-style observations:
{observations_json}
"""
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing = required - value.keys()
unknown = value.keys() - required
if missing:
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _nonempty_text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise DerivationValidationError(f"{location} must be a non-empty string")
return value.strip()
def _validate_observations(observations: Any) -> None:
if not isinstance(observations, list) or not observations:
raise DerivationValidationError("observations must be a non-empty list")
seen_observations: set[str] = set()
seen_evidence: set[str] = set()
for index, observation in enumerate(observations):
location = f"observations[{index}]"
if not isinstance(observation, dict):
raise DerivationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
if observation_id in seen_observations or evidence_id in seen_evidence:
raise DerivationValidationError("observation and evidence provenance must be unique")
seen_observations.add(observation_id)
seen_evidence.add(evidence_id)
_nonempty_text(observation["content"], f"{location}.content")
_nonempty_text(observation["speaker"], f"{location}.speaker")
for field in ("named_person", "addressee"):
if observation[field] is not None:
_nonempty_text(observation[field], f"{location}.{field}")
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict):
raise DerivationValidationError("Gold fixture must be an object")
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
if data["schema_version"] != GOLD_SCHEMA_VERSION:
raise DerivationValidationError("unexpected Gold fixture schema_version")
cases = data["cases"]
if not isinstance(cases, list) or not cases:
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
seen: set[str] = set()
for case in cases:
_exact_keys(case, {"case_id", "description", "observations", "expected"}, "Gold case")
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
if case_id in seen:
raise DerivationValidationError(f"duplicate case ID: {case_id}")
seen.add(case_id)
_validate_observations(case["observations"])
if len(case["observations"]) not in (1, 2):
raise DerivationValidationError("Negative Act cases require one or two observations")
return cases
def build_prompt(case: dict[str, Any]) -> str:
observations = case["observations"]
_validate_observations(observations)
candidate_id = observations[-1]["observation_id"]
return PROMPT_TEMPLATE.format(
candidate_observation_id=candidate_id,
observations_json=json.dumps(observations, ensure_ascii=False, indent=2),
)
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise DerivationValidationError("semantic classification must be an object")
return data
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
if isinstance(value, dict):
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
if forbidden:
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
for key, item in value.items():
_reject_forbidden_keys(item, f"{location}.{key}")
elif isinstance(value, list):
for index, item in enumerate(value):
_reject_forbidden_keys(item, f"{location}[{index}]")
def validate_classification(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
_validate_observations(observations)
if not isinstance(data, dict):
raise DerivationValidationError("semantic classification must be an object")
_reject_forbidden_keys(data)
_exact_keys(data, RECOGNITION_KEYS, "output")
observation_id = _nonempty_text(data["observation_id"], "output.observation_id")
if observation_id not in {item["observation_id"] for item in observations}:
raise DerivationValidationError("classification references unknown observation")
form = data["negative_act_form"]
if form not in NEGATIVE_ACT_FORMS:
raise DerivationValidationError("negative_act_form has an unsupported value")
action_text = data["normalized_action_text"]
if form == "none":
if action_text is not None:
raise DerivationValidationError("none form requires null normalized_action_text")
else:
_nonempty_text(action_text, "output.normalized_action_text")
return data
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
if not concepts:
return text is None
if not isinstance(text, str):
return False
folded = text.casefold()
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
def evaluate_case(case: dict[str, Any], classification: dict[str, Any]) -> dict[str, Any]:
validate_classification(classification, case["observations"])
expected = case["expected"]
observation_correct = classification["observation_id"] == expected["observation_id"]
form_correct = classification["negative_act_form"] == expected["negative_act_form"]
action_correct = _concepts_present(classification["normalized_action_text"], expected["action_concepts"])
unsupported_strengthening = expected["negative_act_form"] == "none" and classification["negative_act_form"] != "none"
classification_label = "PASS" if observation_correct and form_correct and action_correct else ("PARTIAL" if observation_correct and form_correct else "FAIL")
return {
"case_id": case["case_id"], "classification": classification_label,
"expected_negative_act_form": expected["negative_act_form"],
"actual_negative_act_form": classification["negative_act_form"],
"observation_id_correct": observation_correct,
"normalized_action_meaning_correct": action_correct,
"unsupported_semantic_strengthening": unsupported_strengthening,
"normative_leakage": False,
}
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
cases = load_gold_cases(args.cases)
args.output.mkdir(parents=True, exist_ok=False)
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
evaluations: list[dict[str, Any]] = []
successful_calls = 0
technical_failures = 0
started = time.perf_counter()
for case in cases:
case_dir = args.output / case["case_id"].lower()
case_dir.mkdir()
observations = case["observations"]
_write_json(case_dir / "v3_style_input_observations.json", observations)
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
try:
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
successful_calls += 1
except Exception as exc: # one recorded attempt; never retry
technical_failures += 1
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
_write_json(case_dir / "evaluation.json", failure)
evaluations.append(failure)
continue
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
try:
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_semantic_classification.json", parsed)
evaluation = evaluate_case(case, parsed)
validation = {"valid": True, "error": None}
except (DerivationValidationError, json.JSONDecodeError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)}
_write_json(case_dir / "structural_validation.json", validation)
_write_json(case_dir / "evaluation.json", evaluation)
evaluations.append(evaluation)
summary = {
"experiment": "negative_act_form_v0", "model": args.model,
"successful_llm_call_count": successful_calls,
"technical_failed_call_count": technical_failures,
"runtime_seconds": round(time.perf_counter() - started, 3),
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
"evaluations": evaluations,
}
_write_json(args.output / "summary.json", summary)
return summary
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run isolated Negative Act Form experiment")
parser.add_argument("cases", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=1024)
return parser.parse_args()
def main() -> int:
summary = run_experiment(parse_args())
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,347 @@
#!/usr/bin/env python3
"""Isolated explicit-action-rejection Gold reliability experiment."""
from __future__ import annotations
import argparse
import json
import time
from pathlib import Path
from typing import Any
from .experiment_h import (
DEFAULT_ENDPOINT,
DEFAULT_MODEL,
DerivationValidationError,
OBSERVATION_KEYS,
call_ollama,
)
GOLD_SCHEMA_VERSION = "experimental-explicit-rejection-gold-v0"
RECOGNITION_KEYS = {
"rejection_observation_id", "target_observation_id", "rejection_form",
"normalized_rejected_action_text",
}
REJECTION_FORMS = {"explicit_action_rejection", "none"}
FORBIDDEN_LLM_KEYS = {
"decision", "decision_status", "outcome", "topic_status", "closed",
"agreement", "responsible_person", "responsibility", "responsibility_scope",
"owner", "ownership", "assignee", "requested_actor", "status",
"explicitly_rejected", "action_item", "protocol", "protocol_category",
"confidence", "relation", "relations", "graph", "unresolved_issue",
}
PROMPT_TEMPLATE = """Recognize only whether the candidate rejection observation explicitly rejects a concrete action, option, proposal, or future course of action in this small local set of V3-style observations.
Answer only:
1. Does the candidate rejection observation explicitly reject, abandon, discontinue, or rule out a concrete action, option, proposal, or future course of action?
2. If yes, which supplied observation identifies the rejected target?
3. What is the concise normalized meaning of the rejected action or option?
The candidate rejection observation is {rejection_observation_id}.
Use explicit_action_rejection only for an asserted rejection, abandonment, discontinuation, or non-pursuit with a concrete locally resolvable target. Personal preference is not meeting-level explicit rejection. Concern or objection without refusal is not rejection. Uncertainty is not rejection. Negative recommendation or advice is not established rejection. Deferral is not rejection. "Not yet" or temporary non-action is not abandonment. Factual negation is not action rejection. Lack of commitment is not rejection.
The rejected target may be self-contained in the candidate observation or introduced by one earlier supplied observation. Choose only among supplied observation IDs. If the target is ambiguous or unresolved, return rejection_form none. Preserve material scope limitations in normalized_rejected_action_text. Ignore a separate positive alternative when describing the rejected target. Keep normalized text in the observation language.
Do not infer responsibility, ownership, decision status, final outcome, topic closure, protocol status, confidence, relations, graphs, or unresolved issues. Do not answer whether this was finally decided, what the meeting outcome was, who is responsible, or whether the topic is closed.
Return exactly this JSON shape and no additional fields:
{{
"rejection_observation_id": "{rejection_observation_id}",
"target_observation_id": "supplied observation ID" | null,
"rejection_form": "explicit_action_rejection | none",
"normalized_rejected_action_text": "concise rejected target" | null
}}
For rejection_form none, target_observation_id and normalized_rejected_action_text must both be null.
V3-style observations:
{observations_json}
"""
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing = required - value.keys()
unknown = value.keys() - required
if missing:
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _nonempty_text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise DerivationValidationError(f"{location} must be a non-empty string")
return value.strip()
def _validate_observations(observations: Any) -> None:
if not isinstance(observations, list) or not observations:
raise DerivationValidationError("observations must be a non-empty list")
seen_observations: set[str] = set()
seen_evidence: set[str] = set()
for index, observation in enumerate(observations):
location = f"observations[{index}]"
if not isinstance(observation, dict):
raise DerivationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
if observation_id in seen_observations:
raise DerivationValidationError("observation IDs must be unique")
if evidence_id in seen_evidence:
raise DerivationValidationError("evidence provenance must be unique and consistent")
seen_observations.add(observation_id)
seen_evidence.add(evidence_id)
_nonempty_text(observation["content"], f"{location}.content")
_nonempty_text(observation["speaker"], f"{location}.speaker")
for field in ("named_person", "addressee"):
if observation[field] is not None:
_nonempty_text(observation[field], f"{location}.{field}")
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict):
raise DerivationValidationError("Gold fixture must be an object")
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
if data["schema_version"] != GOLD_SCHEMA_VERSION:
raise DerivationValidationError("unexpected Gold fixture schema_version")
cases = data["cases"]
if not isinstance(cases, list) or not cases:
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
seen: set[str] = set()
for case in cases:
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
if case_id in seen:
raise DerivationValidationError(f"duplicate case ID: {case_id}")
seen.add(case_id)
_validate_observations(case["observations"])
if len(case["observations"]) not in (1, 2):
raise DerivationValidationError("rejection Gold cases require one or two observations")
return cases
def build_prompt(case: dict[str, Any]) -> str:
observations = case["observations"]
_validate_observations(observations)
rejection_observation_id = observations[-1]["observation_id"]
return PROMPT_TEMPLATE.format(
rejection_observation_id=rejection_observation_id,
observations_json=json.dumps(observations, ensure_ascii=False, indent=2),
)
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
return data
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
if isinstance(value, dict):
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
if forbidden:
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
for key, item in value.items():
_reject_forbidden_keys(item, f"{location}.{key}")
elif isinstance(value, list):
for index, item in enumerate(value):
_reject_forbidden_keys(item, f"{location}[{index}]")
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
_validate_observations(observations)
if not isinstance(data, dict):
raise DerivationValidationError("semantic recognition must be an object")
_reject_forbidden_keys(data)
_exact_keys(data, RECOGNITION_KEYS, "output")
rejection_id = _nonempty_text(data["rejection_observation_id"], "output.rejection_observation_id")
known_ids = {item["observation_id"] for item in observations}
if rejection_id not in known_ids:
raise DerivationValidationError("unknown rejection observation ID")
form = data["rejection_form"]
if form not in REJECTION_FORMS:
raise DerivationValidationError("rejection_form has an unsupported value")
target_id = data["target_observation_id"]
action_text = data["normalized_rejected_action_text"]
if form == "none":
if target_id is not None:
raise DerivationValidationError("none rejection must have null target_observation_id")
if action_text is not None:
raise DerivationValidationError("none rejection must have null normalized_rejected_action_text")
else:
target_id = _nonempty_text(target_id, "output.target_observation_id")
if target_id not in known_ids:
raise DerivationValidationError("unknown target observation ID")
_nonempty_text(action_text, "output.normalized_rejected_action_text")
return data
def derive_rejection(
observations: list[dict[str, Any]], recognition: dict[str, Any]
) -> tuple[dict[str, bool], dict[str, Any] | None]:
validate_recognition(recognition, observations)
by_id = {item["observation_id"]: item for item in observations}
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
rejection = by_id.get(recognition["rejection_observation_id"])
target_id = recognition["target_observation_id"]
target = by_id.get(target_id) if target_id is not None else None
gates = {
"recognition_schema_valid": True,
"explicit_action_rejection": recognition["rejection_form"] == "explicit_action_rejection",
"rejection_observation_exists": rejection is not None,
"target_observation_exists": target is not None,
"observation_ids_valid_and_unique": len(by_id) == len(observations),
"evidence_provenance_valid_unique_consistent": len({item["evidence_id"] for item in observations}) == len(observations),
"target_same_or_before_rejection": target is not None and rejection is not None and positions[target["observation_id"]] <= positions[rejection["observation_id"]],
"normalized_rejected_action_present": isinstance(recognition["normalized_rejected_action_text"], str) and bool(recognition["normalized_rejected_action_text"].strip()),
"target_local_to_case": target_id in by_id if target_id is not None else False,
"schema_state_consistent": recognition["rejection_form"] == "explicit_action_rejection" and target_id is not None,
"referenced_provenance_available": target is not None and rejection is not None and bool(target["evidence_id"]) and bool(rejection["evidence_id"]),
}
if not all(gates.values()):
return gates, None
return gates, {
"rejection_id": "rejection_1",
"content": recognition["normalized_rejected_action_text"].strip(),
"status": "explicitly_rejected",
"support": {
"target": {"observation_id": target["observation_id"], "evidence_id": target["evidence_id"]},
"rejection": {"observation_id": rejection["observation_id"], "evidence_id": rejection["evidence_id"]},
},
}
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
if not concepts:
return True
if not isinstance(text, str):
return False
folded = text.casefold()
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
def _contains_forbidden_concept(text: str | None, concepts: list[str]) -> bool:
return isinstance(text, str) and any(concept.casefold() in text.casefold() for concept in concepts)
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
validate_recognition(recognition, case["observations"])
gates, result = derive_rejection(case["observations"], recognition)
expected = case["expected_recognition"]
expected_result = case["expected_result"]
form_correct = recognition["rejection_form"] == expected["rejection_form"]
rejection_observation_correct = recognition["rejection_observation_id"] == expected["rejection_observation_id"]
target_correct = recognition["target_observation_id"] == expected["target_observation_id"]
action_correct = _concepts_present(recognition["normalized_rejected_action_text"], expected["action_concepts"])
qualifier_preserved = _concepts_present(recognition["normalized_rejected_action_text"], expected["qualifier_concepts"])
alternative_absorbed = _contains_forbidden_concept(recognition["normalized_rejected_action_text"], expected["forbidden_action_concepts"])
derived = result is not None
final_correct = derived == expected_result["explicitly_rejected"]
if result is not None:
final_correct = final_correct and result["status"] == "explicitly_rejected"
semantic_correct = form_correct and rejection_observation_correct and target_correct and action_correct and qualifier_preserved and not alternative_absorbed
automatic_failure = (derived and not expected_result["explicitly_rejected"]) or (derived and not target_correct) or (derived and not qualifier_preserved) or alternative_absorbed
classification = "FAIL" if automatic_failure or not final_correct else ("PASS" if semantic_correct else "PARTIAL")
return {
"case_id": case["case_id"], "classification": classification,
"rejection_form_correct": form_correct,
"rejection_observation_correct": rejection_observation_correct,
"target_observation_correct": target_correct,
"normalized_rejected_action_correct": action_correct,
"material_qualifiers_preserved": qualifier_preserved,
"positive_alternative_absorbed": alternative_absorbed,
"deterministic_gates_correct": final_correct,
"final_result_correct": final_correct,
"unsupported_semantic_strengthening": recognition["rejection_form"] == "explicit_action_rejection" and expected["rejection_form"] == "none",
"normative_leakage": False,
"gates": gates, "result": result,
}
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
cases = load_gold_cases(args.cases)
args.output.mkdir(parents=True, exist_ok=False)
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
evaluations: list[dict[str, Any]] = []
successful_calls = 0
technical_failures = 0
started = time.perf_counter()
for case in cases:
case_dir = args.output / case["case_id"].lower()
case_dir.mkdir()
observations = case["observations"]
_write_json(case_dir / "v3_style_input_observations.json", observations)
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
try:
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
successful_calls += 1
except Exception as exc: # one recorded attempt; never retry
technical_failures += 1
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
_write_json(case_dir / "deterministic_gate_results.json", {})
_write_json(case_dir / "final_derived_result.json", None)
_write_json(case_dir / "evaluation.json", failure)
evaluations.append(failure)
continue
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
try:
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
evaluation = evaluate_case(case, parsed)
validation = {"valid": True, "error": None}
gates, result = derive_rejection(observations, parsed)
except (DerivationValidationError, json.JSONDecodeError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)}
gates, result = {}, None
_write_json(case_dir / "structural_validation.json", validation)
_write_json(case_dir / "deterministic_gate_results.json", gates)
_write_json(case_dir / "final_derived_result.json", result)
_write_json(case_dir / "evaluation.json", evaluation)
evaluations.append(evaluation)
summary = {
"experiment": "explicit_rejection_gold_v0", "model": args.model,
"successful_llm_call_count": successful_calls,
"technical_failed_call_count": technical_failures,
"runtime_seconds": round(time.perf_counter() - started, 3),
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
"evaluations": evaluations,
}
_write_json(args.output / "summary.json", summary)
return summary
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run isolated explicit-rejection Gold experiment")
parser.add_argument("cases", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=1024)
return parser.parse_args()
def main() -> int:
summary = run_gold(parse_args())
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,100 @@
#!/usr/bin/env python3
"""Isolated controlled rejection derivation V1 experiment."""
from __future__ import annotations
import argparse, json, time
from pathlib import Path
from typing import Any
from .experiment_h import DEFAULT_ENDPOINT, DEFAULT_MODEL, DerivationValidationError, OBSERVATION_KEYS, call_ollama
from .experiment_negative_act import build_prompt as build_negative_prompt, parse_model_json, validate_classification
SCHEMA_VERSION="experimental-controlled-rejection-v1"
TARGET_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
FORBIDDEN={"rejection_form","negative_act_form","explicitly_rejected","status","decision","outcome","topic_status","closed","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
PROMPT="""Resolve only the concrete local action or option referred to by the candidate negative act. The candidate is {candidate}. Choose only a supplied observation ID. Use the same observation for a self-contained target. If no unique local target exists, return null for both target fields. Preserve source-language meaning and material scope such as purpose and location. Preserve continuation when non-pursuit concerns continuing something. Ignore any separate positive alternative. Do not classify the negative act and do not output rejection, status, decision, outcome, responsibility, protocol concepts, confidence, relations, or graphs. Return exactly JSON with candidate_observation_id, target_observation_id, normalized_target_text and no other fields.\nObservations:\n{observations}"""
def _keys(v,r,loc):
if not isinstance(v,dict): raise DerivationValidationError(f"{loc} must be an object")
if set(v)!=r: raise DerivationValidationError(f"{loc} keys invalid: missing={sorted(r-set(v))}, unknown={sorted(set(v)-r)}")
def _text(v,loc):
if not isinstance(v,str) or not v.strip(): raise DerivationValidationError(f"{loc} must be non-empty")
return v.strip()
def _forbidden(v,loc="output"):
if isinstance(v,dict):
bad=FORBIDDEN & set(v)
if bad: raise DerivationValidationError(f"{loc} contains forbidden fields: {sorted(bad)}")
for k,x in v.items(): _forbidden(x,f"{loc}.{k}")
elif isinstance(v,list):
for i,x in enumerate(v): _forbidden(x,f"{loc}[{i}]")
def validate_observations(obs):
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
ids=set(); evid=set()
for i,o in enumerate(obs):
_keys(o,OBSERVATION_KEYS,f"observations[{i}]"); oid=_text(o["observation_id"],"observation_id"); eid=_text(o["evidence_id"],"evidence_id")
if oid in ids or eid in evid: raise DerivationValidationError("provenance must be unique")
ids.add(oid); evid.add(eid); _text(o["content"],"content"); _text(o["speaker"],"speaker")
return ids
def validate_target(data,obs):
ids=validate_observations(obs); _forbidden(data); _keys(data,TARGET_KEYS,"target output")
candidate=_text(data["candidate_observation_id"],"candidate_observation_id")
if candidate not in ids: raise DerivationValidationError("unknown candidate observation")
target=data["target_observation_id"]; normalized=data["normalized_target_text"]
if target is None:
if normalized is not None: raise DerivationValidationError("null target requires null text")
else:
target=_text(target,"target_observation_id")
if target not in ids: raise DerivationValidationError("unknown target observation")
_text(normalized,"normalized_target_text")
return data
def build_target_prompt(case):
validate_observations(case["observations"])
return PROMPT.format(candidate=case["expected"]["candidate_observation_id"],observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
def derive(obs,negative,target):
ids=validate_observations(obs); validate_classification(negative,obs); validate_target(target,obs)
candidate=negative["observation_id"]
if target["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate outputs disagree")
positions={o["observation_id"]:i for i,o in enumerate(obs)}; tid=target["target_observation_id"]
gates={"negative_act_valid":True,"eligible_explicit_non_pursuit":negative["negative_act_form"]=="explicit_non_pursuit","candidate_exists":candidate in ids,"target_valid":True,"target_present":tid is not None,"target_exists":tid in ids if tid else False,"target_not_after_candidate":positions[tid]<=positions[candidate] if tid else False,"provenance_valid_unique":True,"normalized_target_nonempty":bool(target["normalized_target_text"] and target["normalized_target_text"].strip()),"same_isolated_case":tid in ids if tid else False,"no_forbidden_fields":True}
established=all(gates.values())
result=None
if established:
byid={o["observation_id"]:o for o in obs}
result={"rejection_id":"rejection_1","content":target["normalized_target_text"].strip(),"status":"explicitly_rejected","support":{"target":{"observation_id":tid,"evidence_id":byid[tid]["evidence_id"]},"negative_act":{"observation_id":candidate,"evidence_id":byid[candidate]["evidence_id"]}}}
return {"gates":gates,"derived_result":result}
def load_cases(path):
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
return data["cases"]
def _concepts(text,groups):
folded=(text or "").casefold(); return all(any(x.casefold() in folded for x in g) for g in groups)
def evaluate(case,negative,target,derivation):
e=case["expected"]; text=target["normalized_target_text"]
form=negative["negative_act_form"]==e["negative_act_form"]; target_ok=target["target_observation_id"]==e["target_observation_id"]
action=_concepts(text,e["action_concepts"]); material=_concepts(text,e["material_concepts"]); forbidden=any(x.casefold() in (text or "").casefold() for x in e["forbidden_concepts"])
final=(derivation["derived_result"] is not None)==e["explicitly_rejected"]
label="PASS" if form and target_ok and action and material and not forbidden and final else ("PARTIAL" if form and target_ok and material and not forbidden and final else "FAIL")
return {"case_id":case["case_id"],"classification":label,"expected_negative_act_form":e["negative_act_form"],"actual_negative_act_form":negative["negative_act_form"],"expected_target_observation_id":e["target_observation_id"],"actual_target_observation_id":target["target_observation_id"],"normalized_target_text":text,"normalized_action_correct":action,"material_scope_preserved":material,"alternative_absorbed":forbidden,"final_rejection_correct":final}
def _write(p,v): p.write_text(json.dumps(v,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
def run(args):
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases})
evals=[]; naf_calls=target_calls=technical_failures=structural_failures=0; start=time.perf_counter()
for case in cases:
d=args.output/case["case_id"].lower(); d.mkdir(); obs=case["observations"]; _write(d/"v3_style_input_observations.json",obs)
try:
source=case["negative_act_source"]
if source=="live":
np=build_negative_prompt(case); (d/"negative_act_prompt.txt").write_text(np,encoding="utf-8"); raw,nmeta=call_ollama(args.endpoint,args.model,np,args.timeout,args.num_ctx,args.num_predict); naf_calls+=1; (d/"negative_act_raw_response.txt").write_text(raw+"\n",encoding="utf-8"); negative=parse_model_json(raw)
_write(d/"negative_act_ollama_metadata.json",nmeta); _write(d/"negative_act_source.json",{"kind":"live_call"})
else:
sd=args.negative_act_artifacts/source.lower(); accepted=json.loads((sd/"v3_style_input_observations.json").read_text());
if accepted!=obs: raise DerivationValidationError(f"{source} observations do not exactly match")
negative=json.loads((sd/"parsed_semantic_classification.json").read_text()); _write(d/"negative_act_source.json",{"kind":"accepted_artifact_reuse","case_id":source,"path":str(sd)})
validate_classification(negative,obs); _write(d/"negative_act_classification.json",negative)
tp=build_target_prompt(case); (d/"target_prompt.txt").write_text(tp,encoding="utf-8"); traw,tmeta=call_ollama(args.endpoint,args.model,tp,args.timeout,args.num_ctx,args.num_predict); target_calls+=1; (d/"target_raw_response.txt").write_text(traw+"\n",encoding="utf-8"); _write(d/"target_ollama_metadata.json",tmeta); target=parse_model_json(traw); _write(d/"target_recognition.json",target); validate_target(target,obs)
derivation=derive(obs,negative,target); _write(d/"deterministic_gate_results.json",derivation["gates"]); _write(d/"final_derived_result.json",derivation["derived_result"]); ev=evaluate(case,negative,target,derivation)
_write(d/"structural_validation.json",{"valid":True})
except Exception as exc:
structural_failures+=1; ev={"case_id":case["case_id"],"classification":"FAIL","technical_or_validation_failure":str(exc)}; _write(d/"structural_validation.json",{"valid":False,"error":str(exc)})
_write(d/"evaluation.json",ev); evals.append(ev)
summary={"experiment":"controlled_rejection_v1","model":args.model,"negative_act_llm_call_count":naf_calls,"reused_negative_act_count":len(cases)-naf_calls,"target_resolution_llm_call_count":target_calls,"technical_failed_call_count":technical_failures,"structural_validation_failure_count":structural_failures,"runtime_seconds":round(time.perf_counter()-start,3),"counts":{x:sum(e["classification"]==x for e in evals) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evals}; _write(args.output/"summary.json",summary); return summary
def main():
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--negative-act-artifacts",type=Path,default=Path("artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run")); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); args=p.parse_args(); print(json.dumps(run(args),ensure_ascii=False,indent=2)); return 0
@@ -0,0 +1,81 @@
#!/usr/bin/env python3
"""Target Normalization V0: reconstruct target text with fixed linkage."""
from __future__ import annotations
import argparse,json,time
from pathlib import Path
from typing import Any,Callable
import requests
from .experiment_h import DEFAULT_ENDPOINT,DEFAULT_MODEL,DerivationValidationError,OBSERVATION_KEYS
SCHEMA_VERSION="experimental-target-normalization-v0"
OUTPUT_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
LINK_KEYS={"candidate_observation_id","target_observation_id"}
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","topic_status","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
PROMPT="""The candidate and target observation IDs below are already resolved. Copy both IDs exactly; do not perform target selection. Reconstruct only the concrete POSITIVE action or option meaning targeted by the negative act. Remove rejection and negation polarity while preserving the underlying positive action. Preserve German source language, material qualifiers, purpose, location, named people, and continuation. Exclude separate positive alternatives. Do not summarize the discussion or infer rejection, decision, outcome, status, responsibility, ownership, protocol relevance, confidence, relations, graphs, or topic state. Return only the JSON-Schema-conforming object; null is not permitted.\n\nExample A observations: [{{"observation_id":"obs_a","content":"Mit Frau Beispiel arbeiten wir nicht weiter."}}]\nFixed IDs: candidate=obs_a, target=obs_a\nOutput: {{"candidate_observation_id":"obs_a","target_observation_id":"obs_a","normalized_target_text":"Zusammenarbeit mit Frau Beispiel fortsetzen"}}\n\nExample B observations: [{{"observation_id":"obs_a","content":"Für den Druckversuch steht die reale Anlage zur Diskussion."}},{{"observation_id":"obs_b","content":"Die reale Anlage nutzen wir dafür nicht."}}]\nFixed IDs: candidate=obs_b, target=obs_a\nOutput: {{"candidate_observation_id":"obs_b","target_observation_id":"obs_a","normalized_target_text":"reale Anlage für den Druckversuch nutzen"}}\n\nFixed candidate_observation_id: {candidate}\nFixed target_observation_id: {target}\nV3-style observations:\n{observations}"""
def _keys(value,required,where):
if not isinstance(value,dict): raise DerivationValidationError(f"{where} must be an object")
if set(value)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(value))}, unknown={sorted(set(value)-required)}")
def _text(value,where):
if not isinstance(value,str) or not value.strip(): raise DerivationValidationError(f"{where} must be non-empty")
return value.strip()
def _reject_forbidden(value,where="output"):
if isinstance(value,dict):
bad=FORBIDDEN & set(value)
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
for key,item in value.items(): _reject_forbidden(item,f"{where}.{key}")
elif isinstance(value,list):
for index,item in enumerate(value): _reject_forbidden(item,f"{where}[{index}]")
def validate_observations(observations):
if not isinstance(observations,list) or not observations: raise DerivationValidationError("observations must be non-empty")
ids=set(); evidence=set()
for index,item in enumerate(observations):
_keys(item,OBSERVATION_KEYS,f"observations[{index}]"); oid=_text(item["observation_id"],"observation_id"); eid=_text(item["evidence_id"],"evidence_id")
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
ids.add(oid); evidence.add(eid); _text(item["content"],"content"); _text(item["speaker"],"speaker")
return ids
def validate_linkage(case):
ids=validate_observations(case["observations"]); linkage=case["fixed_linkage"]; _keys(linkage,LINK_KEYS,"fixed_linkage")
for field in LINK_KEYS:
if _text(linkage[field],field) not in ids: raise DerivationValidationError(f"{field} is unknown")
return linkage
def output_schema(case):
link=validate_linkage(case)
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","target_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":link["candidate_observation_id"]},"target_observation_id":{"const":link["target_observation_id"]},"normalized_target_text":{"type":"string","minLength":1}}}
def build_prompt(case):
link=validate_linkage(case)
return PROMPT.format(candidate=link["candidate_observation_id"],target=link["target_observation_id"],observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
def validate_output(data,case):
_reject_forbidden(data); _keys(data,OUTPUT_KEYS,"output"); link=validate_linkage(case)
if data["candidate_observation_id"]!=link["candidate_observation_id"]: raise DerivationValidationError("candidate ID changed")
if data["target_observation_id"]!=link["target_observation_id"]: raise DerivationValidationError("target ID changed")
_text(data["normalized_target_text"],"normalized_target_text"); return data
def build_payload(model,prompt,schema,num_ctx,num_predict): return {"model":model,"prompt":prompt,"think":False,"stream":False,"format":schema,"options":{"temperature":0,"num_ctx":num_ctx,"num_predict":num_predict}}
def call_schema(endpoint,model,prompt,schema,timeout,num_ctx,num_predict):
started=time.perf_counter(); response=requests.post(endpoint,json=build_payload(model,prompt,schema,num_ctx,num_predict),timeout=timeout); elapsed=time.perf_counter()-started; response.raise_for_status(); body=response.json(); raw=body.get("response")
if not isinstance(raw,str) or not raw.strip(): raise ValueError("Ollama returned no usable response")
return raw.strip(),{"model":body.get("model",model),"elapsed_seconds":round(elapsed,3),"total_duration_ns":body.get("total_duration"),"prompt_eval_count":body.get("prompt_eval_count"),"eval_count":body.get("eval_count"),"configuration":{"temperature":0,"think":False,"format":"json_schema_object","num_ctx":num_ctx,"num_predict":num_predict,"retries":0}}
def _concepts(text,groups):
folded=text.casefold(); return all(any(alias.casefold() in folded for alias in group) for group in groups)
def evaluate(case,output):
validate_output(output,case); expected=case["expected"]; text=output["normalized_target_text"]; folded=text.casefold(); action=_concepts(text,expected["action_concepts"]); scope=_concepts(text,expected["material_concepts"]); forbidden=[x for x in expected["forbidden_concepts"] if x.casefold() in folded]; positive=not any(x in forbidden for x in ("nicht","beenden")); german=any(x.casefold() in folded for x in expected["german_markers"]); alternative=not any(x.casefold() in folded for x in ("technikum","stattdessen")); strengthening=False
label="PASS" if positive and action and scope and german and alternative and not forbidden and not strengthening else "FAIL"
return {"case_id":case["case_id"],"classification":label,"expected_normalized_target_text":expected["normalized_target_text"],"actual_normalized_target_text":text,"positive_polarity_correct":positive,"action_semantics_preserved":action,"material_scope_preserved":scope,"source_language_preserved":german,"separate_alternative_excluded":alternative,"forbidden_semantics_present":forbidden,"unsupported_strengthening":strengthening,"normative_leakage":False,"candidate_id_unchanged":True,"target_id_unchanged":True}
def load_cases(path):
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
for case in data["cases"]: validate_linkage(case)
return data["cases"]
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
def run(args,caller:Callable=call_schema):
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases}); evaluations=[]; calls=failures=0; started=time.perf_counter()
for case in cases:
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"fixed_linkage.json",case["fixed_linkage"]); schema=output_schema(case); _write(folder/"ollama_json_schema.json",schema); prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
try:
raw,metadata=caller(args.endpoint,args.model,prompt,schema,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",metadata); parsed=json.loads(raw); _write(folder/"parsed_response.json",parsed); validate_output(parsed,case); _write(folder/"structural_validation.json",{"valid":True}); _write(folder/"normalized_target_result.json",{"normalized_target_text":parsed["normalized_target_text"]}); evaluation=evaluate(case,parsed)
except Exception as exc:
failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); evaluation={"case_id":case["case_id"],"classification":"FAIL","error":str(exc),"normative_leakage":False}
_write(folder/"evaluation.json",evaluation); evaluations.append(evaluation)
summary={"experiment":"target_normalization_v0","model":args.model,"llm_call_count":calls,"structural_validation_failure_count":failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
def main():
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
@@ -0,0 +1,92 @@
#!/usr/bin/env python3
"""Isolated local target-resolution experiment; performs no rejection derivation."""
from __future__ import annotations
import argparse, json, time
from pathlib import Path
from typing import Any, Callable
from .experiment_h import DEFAULT_ENDPOINT, DEFAULT_MODEL, DerivationValidationError, OBSERVATION_KEYS, call_ollama
from .experiment_negative_act import validate_classification
SCHEMA_VERSION="experimental-target-resolution-v0"
TARGET_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","topic_status","closed","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
PROMPT="""Resolve and normalize only the concrete action or option referred to by the candidate negative act. Candidate: {candidate}. Strategy: {instruction} Choose only a supplied observation ID. If no unique local target exists, use JSON null for both target fields. Preserve German source meaning, continuation, purpose, location, and other material scope. Do not translate Anlage as asset. Ignore separate positive alternatives. Do not output negative-act form, rejection, status, decision, outcome, responsibility, topic closure, protocol concepts, confidence, relations, or graphs. Return exactly this JSON object with no additional fields: {{"candidate_observation_id":"{candidate}","target_observation_id":"observation ID or null","normalized_target_text":"concise positive action in German or null"}}\nObservations:\n{observations}"""
def _keys(v:Any, required:set[str], where:str):
if not isinstance(v,dict): raise DerivationValidationError(f"{where} must be an object")
if set(v)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(v))}, unknown={sorted(set(v)-required)}")
def _text(v:Any, where:str):
if not isinstance(v,str) or not v.strip(): raise DerivationValidationError(f"{where} must be non-empty")
return v.strip()
def _reject_forbidden(v:Any, where="output"):
if isinstance(v,dict):
bad=FORBIDDEN & set(v)
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
for k,x in v.items(): _reject_forbidden(x,f"{where}.{k}")
elif isinstance(v,list):
for i,x in enumerate(v): _reject_forbidden(x,f"{where}[{i}]")
def validate_observations(obs):
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
ids=set(); evidence=set()
for i,o in enumerate(obs):
_keys(o,OBSERVATION_KEYS,f"observations[{i}]"); oid=_text(o["observation_id"],"observation_id"); eid=_text(o["evidence_id"],"evidence_id")
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
ids.add(oid); evidence.add(eid); _text(o["content"],"content"); _text(o["speaker"],"speaker")
return ids
def eligibility(negative,obs):
validate_classification(negative,obs)
eligible=negative["negative_act_form"]=="explicit_non_pursuit"
return {"eligible_for_target_resolution":eligible,"reason":None if eligible else "negative_act_form_not_explicit_non_pursuit"}
def validate_target(data,obs,candidate):
ids=validate_observations(obs); _reject_forbidden(data); _keys(data,TARGET_KEYS,"target output")
if _text(data["candidate_observation_id"],"candidate_observation_id")!=candidate: raise DerivationValidationError("candidate observation mismatch")
if candidate not in ids: raise DerivationValidationError("unknown candidate observation")
target=data["target_observation_id"]; normalized=data["normalized_target_text"]
if target is None:
if normalized is not None: raise DerivationValidationError("null target requires null text")
else:
target=_text(target,"target_observation_id")
if target not in ids: raise DerivationValidationError("unknown target observation")
if [o["observation_id"] for o in obs].index(target)>[o["observation_id"] for o in obs].index(candidate): raise DerivationValidationError("target must not occur after candidate")
_text(normalized,"normalized_target_text")
return data
def build_prompt(case):
gate=eligibility(case["negative_act"],case["observations"])
if not gate["eligible_for_target_resolution"]: raise DerivationValidationError("ineligible case must not build a target prompt")
candidate=case["negative_act"]["observation_id"]
instruction=(f"The target linkage is deterministically fixed to {candidate}; output that exact target ID and only normalize its positive action meaning." if case["strategy"]=="self_contained" else "Resolve the unique preceding local observation that supplies the referenced action.")
return PROMPT.format(candidate=candidate,instruction=instruction,observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
def _concepts(text,groups):
folded=(text or "").casefold(); return all(any(alias.casefold() in folded for alias in group) for group in groups)
def evaluate(case,gate,called,target):
e=case["expected"]; eligible=gate["eligible_for_target_resolution"]==e["eligible"]; call_ok=called==e["eligible"]
if not e["eligible"]:
label="PASS" if eligible and call_ok and target is None else "FAIL"
return {"case_id":case["case_id"],"classification":label,"negative_act_form":case["negative_act"]["negative_act_form"],"eligible":gate["eligible_for_target_resolution"],"target_resolution_call_made":called,"expected_target_observation_id":None,"actual_target_observation_id":None,"normalized_target_text":None,"material_scope_preserved":True,"alternative_isolation":True,"normative_leakage":False}
text=target["normalized_target_text"]; target_ok=target["target_observation_id"]==e["target_observation_id"]; concepts=_concepts(text,e["concepts"]); material=_concepts(text,e["material_concepts"]); isolated=not any(x.casefold() in (text or "").casefold() for x in e["forbidden_concepts"])
label="PASS" if eligible and call_ok and target_ok and concepts and material and isolated else ("PARTIAL" if eligible and call_ok and target_ok and material and isolated else "FAIL")
return {"case_id":case["case_id"],"classification":label,"negative_act_form":case["negative_act"]["negative_act_form"],"eligible":gate["eligible_for_target_resolution"],"target_resolution_call_made":called,"expected_target_observation_id":e["target_observation_id"],"actual_target_observation_id":target["target_observation_id"],"normalized_target_text":text,"normalized_action_correct":concepts,"material_scope_preserved":material,"alternative_isolation":isolated,"normative_leakage":False}
def load_cases(path):
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("unexpected schema version")
for case in data["cases"]: validate_observations(case["observations"]); validate_classification(case["negative_act"],case["observations"])
return data["cases"]
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
def run(args, resolver:Callable=call_ollama):
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases})
evaluations=[]; calls=technical_failures=structural_failures=0; started=time.perf_counter()
for case in cases:
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"negative_act_form.json",case["negative_act"])
gate=eligibility(case["negative_act"],case["observations"]); _write(folder/"eligibility.json",gate)
if not gate["eligible_for_target_resolution"]:
skipped={"call_made":False,"reason":gate["reason"]}; _write(folder/"target_resolution_skipped.json",skipped); ev=evaluate(case,gate,False,None)
else:
prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
try:
raw,meta=resolver(args.endpoint,args.model,prompt,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",meta); parsed=json.loads(raw); _write(folder/"parsed_target_resolution.json",parsed); validate_target(parsed,case["observations"],case["negative_act"]["observation_id"]); _write(folder/"structural_validation.json",{"valid":True}); ev=evaluate(case,gate,True,parsed)
except Exception as exc:
structural_failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); ev={"case_id":case["case_id"],"classification":"FAIL","negative_act_form":case["negative_act"]["negative_act_form"],"eligible":True,"target_resolution_call_made":True,"error":str(exc)}
_write(folder/"evaluation.json",ev); evaluations.append(ev)
summary={"experiment":"target_resolution_v0","model":args.model,"target_resolution_llm_call_count":calls,"technical_failed_call_count":technical_failures,"structural_validation_failure_count":structural_failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
def main():
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
@@ -0,0 +1,107 @@
#!/usr/bin/env python3
"""Target Resolution V1 diagnostic: linkage and normalization only."""
from __future__ import annotations
import argparse,json,time
from pathlib import Path
from typing import Any,Callable
import requests
from .experiment_h import DEFAULT_ENDPOINT,DEFAULT_MODEL,DerivationValidationError,OBSERVATION_KEYS
from .experiment_negative_act import validate_classification
SCHEMA_VERSION="experimental-target-resolution-v1-diagnostic"
SELF_KEYS={"candidate_observation_id","normalized_target_text"}; PAIRED_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","topic_status","closed","unresolved_issue"}
SELF_PROMPT="""Normalize only the concrete positive action meaning in the self-contained candidate observation. The target linkage is already deterministic and is not your task. Preserve German, collaboration, named people, and continuation meaning. Do not output a target ID, rejection, status, decision, outcome, responsibility, ownership, protocol concepts, confidence, relations, graphs, or topic closure. Return only the schema-conforming object.\nCandidate observation ID: {candidate}\nObservation:\n{observations}"""
PAIRED_PROMPT="""Resolve and normalize only the concrete local action or option referred to by the candidate negative act. Choose exactly one listed allowed target observation ID, or use JSON null only when no unique local target exists. Never return the string \"null\". Preserve German source language and all material purpose/location scope. Do not absorb a separate positive alternative. Do not output negative-act form, rejection, status, decision, outcome, responsibility, ownership, protocol concepts, confidence, relations, graphs, or topic closure.\nAllowed target observation IDs:\n{allowed}\nConcrete positive typed example:\n{{"candidate_observation_id":"obs_2","target_observation_id":"obs_1","normalized_target_text":"externe Lösung weiterverfolgen"}}\nActual JSON-null example:\n{{"candidate_observation_id":"obs_2","target_observation_id":null,"normalized_target_text":null}}\nReturn only the schema-conforming object.\nCandidate observation ID: {candidate}\nObservations:\n{observations}"""
def _keys(value,required,where):
if not isinstance(value,dict): raise DerivationValidationError(f"{where} must be an object")
if set(value)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(value))}, unknown={sorted(set(value)-required)}")
def _text(value,where):
if not isinstance(value,str) or not value.strip(): raise DerivationValidationError(f"{where} must be non-empty")
return value.strip()
def _forbidden(value,where="output"):
if isinstance(value,dict):
bad=FORBIDDEN & set(value)
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
for key,item in value.items(): _forbidden(item,f"{where}.{key}")
elif isinstance(value,list):
for index,item in enumerate(value): _forbidden(item,f"{where}[{index}]")
def validate_observations(obs):
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
ids=[]; evidence=set()
for index,item in enumerate(obs):
_keys(item,OBSERVATION_KEYS,f"observations[{index}]"); oid=_text(item["observation_id"],"observation_id"); eid=_text(item["evidence_id"],"evidence_id")
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
ids.append(oid); evidence.add(eid); _text(item["content"],"content"); _text(item["speaker"],"speaker")
return ids
def allowed_ids(case): return validate_observations(case["observations"])
def deterministic_self_link(case):
if case["strategy"]!="self_contained": raise DerivationValidationError("self-linkage requires self-contained strategy")
ids=validate_observations(case["observations"]); candidate=case["negative_act"]["observation_id"]
if candidate not in ids: raise DerivationValidationError("unknown candidate")
return {"linkage_source":"deterministic","candidate_observation_id":candidate,"target_observation_id":candidate}
def output_schema(case):
candidate=case["negative_act"]["observation_id"]
if case["strategy"]=="self_contained":
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":candidate},"normalized_target_text":{"type":"string","minLength":1}}}
ids=allowed_ids(case)
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","target_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":candidate},"target_observation_id":{"enum":ids+[None]},"normalized_target_text":{"type":["string","null"]}},"allOf":[{"if":{"properties":{"target_observation_id":{"type":"null"}}},"then":{"properties":{"normalized_target_text":{"type":"null"}}},"else":{"properties":{"normalized_target_text":{"type":"string","minLength":1}}}}]}
def build_prompt(case):
validate_classification(case["negative_act"],case["observations"]); candidate=case["negative_act"]["observation_id"]
if case["strategy"]=="self_contained": return SELF_PROMPT.format(candidate=candidate,observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
return PAIRED_PROMPT.format(candidate=candidate,allowed=json.dumps(allowed_ids(case),ensure_ascii=False),observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
def validate_semantic_output(data,case):
_forbidden(data); candidate=case["negative_act"]["observation_id"]
if case["strategy"]=="self_contained":
_keys(data,SELF_KEYS,"self output")
if data["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate mismatch")
_text(data["normalized_target_text"],"normalized_target_text")
else:
_keys(data,PAIRED_KEYS,"paired output")
if data["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate mismatch")
target=data["target_observation_id"]
if target=="null": raise DerivationValidationError('string "null" is forbidden')
if target is None:
if data["normalized_target_text"] is not None: raise DerivationValidationError("null target requires null text")
else:
if target not in allowed_ids(case): raise DerivationValidationError("target is not an allowed ID")
ids=allowed_ids(case)
if ids.index(target)>ids.index(candidate): raise DerivationValidationError("target must not occur after candidate")
_text(data["normalized_target_text"],"normalized_target_text")
return data
def combine(case,semantic):
validate_semantic_output(semantic,case)
if case["strategy"]=="self_contained":
link=deterministic_self_link(case); return {**link,"normalized_target_text":semantic["normalized_target_text"]}
return {"linkage_source":"llm","candidate_observation_id":semantic["candidate_observation_id"],"target_observation_id":semantic["target_observation_id"],"normalized_target_text":semantic["normalized_target_text"]}
def build_payload(model,prompt,schema,num_ctx,num_predict):
return {"model":model,"prompt":prompt,"think":False,"stream":False,"format":schema,"options":{"temperature":0,"num_ctx":num_ctx,"num_predict":num_predict}}
def call_schema(endpoint,model,prompt,schema,timeout,num_ctx,num_predict):
started=time.perf_counter(); response=requests.post(endpoint,json=build_payload(model,prompt,schema,num_ctx,num_predict),timeout=timeout); elapsed=time.perf_counter()-started; response.raise_for_status(); body=response.json(); raw=body.get("response")
if not isinstance(raw,str) or not raw.strip(): raise ValueError("Ollama returned no usable response")
meta={"model":body.get("model",model),"elapsed_seconds":round(elapsed,3),"total_duration_ns":body.get("total_duration"),"prompt_eval_count":body.get("prompt_eval_count"),"eval_count":body.get("eval_count"),"configuration":{"temperature":0,"think":False,"format":"json_schema_object","num_ctx":num_ctx,"num_predict":num_predict,"retries":0}}
return raw.strip(),meta
def _concepts(text,groups):
folded=(text or "").casefold(); return all(any(x.casefold() in folded for x in group) for group in groups)
def evaluate(case,semantic,combined):
expected=case["expected"]; text=combined["normalized_target_text"]; target=combined["target_observation_id"]; concepts=_concepts(text,expected["concepts"]); material=_concepts(text,expected["material_concepts"]); isolated=not any(x.casefold() in (text or "").casefold() for x in expected["forbidden_concepts"]); recurrence=semantic.get("target_observation_id")=="null"; correct=target==expected["target_observation_id"]
label="PASS" if correct and concepts and material and isolated and not recurrence else ("PARTIAL" if correct and material and isolated and not recurrence else "FAIL")
return {"case_id":case["case_id"],"classification":label,"strategy":case["strategy"],"target_id_decision_source":combined["linkage_source"],"expected_target_observation_id":expected["target_observation_id"],"actual_target_observation_id":target,"normalized_target_text":text,"continuation_or_action_preserved":concepts,"material_scope_preserved":material,"alternative_isolated":isolated,"schema_valid":True,"string_null_recurrence":recurrence,"normative_leakage":False}
def load_cases(path):
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
return data["cases"]
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
def run(args,caller:Callable=call_schema):
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases}); evaluations=[]; calls=failures=0; started=time.perf_counter()
for case in cases:
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"negative_act_form.json",case["negative_act"]); _write(folder/"eligibility.json",{"eligible_for_target_resolution":True,"reason":None}); _write(folder/"deterministic_strategy.json",{"strategy":case["strategy"],"target_id_decision_source":"deterministic" if case["strategy"]=="self_contained" else "llm"}); _write(folder/"allowed_target_ids.json",allowed_ids(case)); schema=output_schema(case); _write(folder/"ollama_json_schema.json",schema); prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
try:
raw,meta=caller(args.endpoint,args.model,prompt,schema,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",meta); semantic=json.loads(raw); _write(folder/"parsed_semantic_output.json",semantic); validate_semantic_output(semantic,case); combined=combine(case,semantic); _write(folder/"structural_validation.json",{"valid":True}); _write(folder/"deterministic_linkage_result.json",{k:combined[k] for k in ("linkage_source","candidate_observation_id","target_observation_id")}); _write(folder/"normalized_target_result.json",{"normalized_target_text":combined["normalized_target_text"]}); evaluation=evaluate(case,semantic,combined)
except Exception as exc:
failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); evaluation={"case_id":case["case_id"],"classification":"FAIL","strategy":case["strategy"],"schema_valid":False,"error":str(exc)}
_write(folder/"evaluation.json",evaluation); evaluations.append(evaluation)
summary={"experiment":"target_resolution_v1_diagnostic","model":args.model,"llm_call_count":calls,"structural_validation_failure_count":failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
def main():
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
+20
View File
@@ -0,0 +1,20 @@
"""Optional speaker diarization and transcript alignment."""
from src.meeting_lab.diarization.alignment import align_transcript, write_diarized_transcript
from src.meeting_lab.diarization.backend import (
DEFAULT_MODEL,
DiarizationError,
DiarizationResult,
diarize_audio,
select_device,
)
__all__ = [
"DEFAULT_MODEL",
"DiarizationError",
"DiarizationResult",
"align_transcript",
"diarize_audio",
"select_device",
"write_diarized_transcript",
]
+129
View File
@@ -0,0 +1,129 @@
"""Deterministic Whisper-segment alignment to anonymous diarization turns."""
from __future__ import annotations
import json
from pathlib import Path
from typing import Any
class AlignmentError(ValueError):
"""Raised when transcript or diarization inputs are malformed."""
def _number(value: Any, description: str) -> float:
if not isinstance(value, (int, float)) or isinstance(value, bool):
raise AlignmentError(f"{description} must be a number.")
return float(value)
def _validated_turns(turns: list[dict[str, Any]]) -> list[dict[str, Any]]:
validated = []
for index, turn in enumerate(turns):
if not isinstance(turn, dict):
raise AlignmentError(f"Diarization turn {index} must be an object.")
start = _number(turn.get("start"), f"Diarization turn {index} start")
end = _number(turn.get("end"), f"Diarization turn {index} end")
speaker = turn.get("speaker_id", turn.get("speaker"))
if end < start:
raise AlignmentError(f"Diarization turn {index} ends before it starts.")
if not isinstance(speaker, str) or not speaker.startswith("SPEAKER_"):
raise AlignmentError(
f"Diarization turn {index} must have an anonymous SPEAKER_ label."
)
validated.append({"start": start, "end": end, "speaker_id": speaker})
return validated
def align_transcript(
transcript: dict[str, Any], exclusive_turns: list[dict[str, Any]]
) -> dict[str, Any]:
"""Return a derived transcript using maximum exclusive-turn overlap per segment."""
if not isinstance(transcript, dict) or not isinstance(transcript.get("segments"), list):
raise AlignmentError("Whisper transcript must contain a 'segments' list.")
turns = _validated_turns(exclusive_turns)
aligned_segments: list[dict[str, Any]] = []
for index, source in enumerate(transcript["segments"]):
if not isinstance(source, dict):
raise AlignmentError(f"Transcript segment {index} must be an object.")
start = _number(source.get("start"), f"Transcript segment {index} start")
end = _number(source.get("end"), f"Transcript segment {index} end")
if end < start:
raise AlignmentError(f"Transcript segment {index} ends before it starts.")
overlap_by_speaker: dict[str, float] = {}
for turn in turns:
overlap = max(0.0, min(end, turn["end"]) - max(start, turn["start"]))
if overlap:
speaker = turn["speaker_id"]
overlap_by_speaker[speaker] = overlap_by_speaker.get(speaker, 0.0) + overlap
speaker_id = None
overlap_seconds = 0.0
if overlap_by_speaker:
speaker_id, overlap_seconds = min(
overlap_by_speaker.items(), key=lambda item: (-item[1], item[0])
)
duration = end - start
aligned = dict(source)
aligned.update(
{
"speaker_id": speaker_id,
"speaker_overlap_seconds": round(overlap_seconds, 6),
"speaker_overlap_ratio": round(
overlap_seconds / duration if duration > 0 else 0.0, 6
),
}
)
aligned_segments.append(aligned)
return {
"text": diarized_transcript_text(aligned_segments, include_end=True),
"segments": aligned_segments,
"speaker_labels_anonymous": True,
"alignment_source": "exclusive_diarization",
}
def _timestamp(seconds: float) -> str:
milliseconds = int(round(seconds * 1000))
hours, remainder = divmod(milliseconds, 3_600_000)
minutes, remainder = divmod(remainder, 60_000)
secs, millis = divmod(remainder, 1000)
return f"{hours:02d}:{minutes:02d}:{secs:02d}.{millis:03d}"
def diarized_transcript_text(
segments: list[dict[str, Any]], *, include_end: bool = True
) -> str:
lines = []
for segment in segments:
start = _timestamp(float(segment["start"]))
end = _timestamp(float(segment["end"]))
speaker = segment.get("speaker_id") or "SPEAKER_UNASSIGNED"
timestamp = f"[{start} - {end}]" if include_end else f"[{start}]"
lines.append(f"{timestamp} {speaker}: {str(segment.get('text', '')).strip()}")
return "\n".join(lines) + ("\n" if lines else "")
def write_diarized_transcript(
transcript_path: Path,
exclusive_turns_path: Path,
output_dir: Path,
) -> tuple[Path, Path]:
"""Read source artifacts and write a separate speaker-aware transcript pair."""
transcript = json.loads(Path(transcript_path).read_text(encoding="utf-8-sig"))
turns = json.loads(Path(exclusive_turns_path).read_text(encoding="utf-8"))
if not isinstance(turns, list):
raise AlignmentError("Exclusive diarization turns must contain a JSON list.")
derived = align_transcript(transcript, turns)
output_dir = Path(output_dir)
output_dir.mkdir(parents=True, exist_ok=True)
json_path = output_dir / "transcript_diarized.json"
text_path = output_dir / "transcript_diarized.txt"
json_path.write_text(
json.dumps(derived, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
text_path.write_text(
diarized_transcript_text(derived["segments"], include_end=True), encoding="utf-8"
)
return json_path, text_path
+314
View File
@@ -0,0 +1,314 @@
"""pyannote Community-1 backend with native and isolated-container runtimes."""
from __future__ import annotations
import importlib.metadata
import json
import os
import subprocess
import time
import wave
from dataclasses import dataclass
from pathlib import Path
from typing import Any, Callable, Literal, Sequence
DEFAULT_MODEL = "pyannote/speaker-diarization-community-1"
PYANNOTE_VERSION = "4.0.7"
DeviceMode = Literal["auto", "gpu", "cpu"]
RuntimeMode = Literal["native", "container"]
class DiarizationError(RuntimeError):
"""Raised when diarization configuration or execution fails."""
@dataclass(frozen=True)
class DiarizationResult:
output_dir: Path
metadata_path: Path
ordinary_rttm: Path
exclusive_rttm: Path
turns_json: Path
exclusive_turns_json: Path
metadata: dict[str, Any]
def _host_uid() -> int:
getter = getattr(os, "getuid", None)
if getter is None:
raise DiarizationError("Container diarization requires host UID discovery.")
return int(getter())
def _host_gid() -> int:
getter = getattr(os, "getgid", None)
if getter is None:
raise DiarizationError("Container diarization requires host GID discovery.")
return int(getter())
def select_device(mode: DeviceMode, torch_module: Any) -> tuple[Any, str | None]:
"""Resolve CPU/GPU without depending on the GPU vendor."""
if mode == "cpu":
return torch_module.device("cpu"), None
if mode not in ("auto", "gpu"):
raise DiarizationError(f"Unsupported diarization device mode: {mode}")
try:
available = bool(torch_module.cuda.is_available())
if available:
name = str(torch_module.cuda.get_device_name(0))
probe = torch_module.zeros(1, device="cuda")
del probe
return torch_module.device("cuda"), name
except Exception as exc:
if mode == "gpu":
raise DiarizationError(f"Requested PyTorch GPU is not usable: {exc}") from exc
if mode == "gpu":
raise DiarizationError("Requested PyTorch GPU is unavailable.")
return torch_module.device("cpu"), None
def _load_pcm_wave(audio_path: Path, torch_module: Any) -> tuple[Any, int, float, dict[str, Any]]:
try:
with wave.open(str(audio_path), "rb") as source:
channels = source.getnchannels()
sample_rate = source.getframerate()
sample_width = source.getsampwidth()
frame_count = source.getnframes()
pcm = bytearray(source.readframes(frame_count))
except (OSError, wave.Error) as exc:
raise DiarizationError(f"Cannot read PCM WAV input {audio_path}: {exc}") from exc
if channels != 1 or sample_rate != 16000 or sample_width != 2:
raise DiarizationError(
"Diarization currently requires mono 16 kHz signed 16-bit PCM WAV; "
f"got channels={channels}, sample_rate={sample_rate}, sample_width={sample_width}."
)
waveform = torch_module.frombuffer(pcm, dtype=torch_module.int16).to(
torch_module.float32
)
waveform = (waveform / 32768.0).reshape(channels, frame_count)
duration = frame_count / sample_rate
validation = {
"waveform_dtype": str(waveform.dtype),
"waveform_shape": list(waveform.shape),
"sample_rate": sample_rate,
"sample_count": frame_count,
"duration_seconds": duration,
"min_sample_value": waveform.min().item(),
"max_sample_value": waveform.max().item(),
"audio_loading": "python_wave_pcm16",
}
return waveform, sample_rate, duration, validation
def _turns(annotation: Any) -> list[dict[str, Any]]:
return [
{
"start": segment.start,
"end": segment.end,
"speaker_id": speaker,
}
for segment, _track, speaker in annotation.itertracks(yield_label=True)
]
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def _result_from_output(output_dir: Path) -> DiarizationResult:
metadata_path = output_dir / "metadata.json"
try:
metadata = json.loads(metadata_path.read_text(encoding="utf-8"))
except (OSError, json.JSONDecodeError) as exc:
raise DiarizationError(f"Cannot read diarization metadata: {exc}") from exc
return DiarizationResult(
output_dir=output_dir,
metadata_path=metadata_path,
ordinary_rttm=output_dir / "diarization.rttm",
exclusive_rttm=output_dir / "exclusive_diarization.rttm",
turns_json=output_dir / "turns.json",
exclusive_turns_json=output_dir / "exclusive_turns.json",
metadata=metadata,
)
def _require_writable_output(output_dir: Path) -> None:
unwritable = [
path
for path in (output_dir, *output_dir.rglob("*"))
if not os.access(path, os.W_OK)
]
if unwritable:
rendered = ", ".join(str(path) for path in unwritable[:3])
if len(unwritable) > 3:
rendered += f", and {len(unwritable) - 3} more"
raise DiarizationError(
f"Container diarization artifacts are not writable by the host user: {rendered}"
)
def run_native_pyannote(
audio_path: Path,
output_dir: Path,
device_mode: DeviceMode,
*,
model: str = DEFAULT_MODEL,
) -> DiarizationResult:
"""Run one local pyannote inference using an in-memory waveform mapping."""
try:
import torch
from pyannote.audio import Pipeline
except ImportError as exc:
raise DiarizationError(
f"Native diarization requires pyannote.audio=={PYANNOTE_VERSION} and PyTorch."
) from exc
token = os.environ.get("HF_TOKEN")
if not token:
raise DiarizationError("HF_TOKEN is required for the pyannote model.")
output_dir.mkdir(parents=True, exist_ok=True)
waveform, sample_rate, duration, audio_metadata = _load_pcm_wave(audio_path, torch)
device, device_name = select_device(device_mode, torch)
try:
pipeline = Pipeline.from_pretrained(model, token=token)
pipeline.to(device)
started = time.perf_counter()
output = pipeline(
{
"waveform": waveform,
"sample_rate": sample_rate,
"uri": audio_path.stem,
}
)
runtime = time.perf_counter() - started
except Exception as exc:
raise DiarizationError(f"pyannote diarization failed: {type(exc).__name__}: {exc}") from exc
ordinary = getattr(output, "speaker_diarization", output)
exclusive = getattr(output, "exclusive_speaker_diarization", None)
if exclusive is None:
raise DiarizationError("Community-1 did not return exclusive diarization.")
ordinary_turns = _turns(ordinary)
exclusive_turns = _turns(exclusive)
with (output_dir / "diarization.rttm").open("w", encoding="utf-8") as handle:
ordinary.write_rttm(handle)
with (output_dir / "exclusive_diarization.rttm").open(
"w", encoding="utf-8"
) as handle:
exclusive.write_rttm(handle)
_write_json(output_dir / "turns.json", ordinary_turns)
_write_json(output_dir / "exclusive_turns.json", exclusive_turns)
speakers = sorted({turn["speaker_id"] for turn in ordinary_turns})
actual_device = str(device)
metadata = {
"backend": "pyannote.audio",
"model": model,
"pyannote_version": importlib.metadata.version("pyannote.audio"),
"torch_version": torch.__version__,
"hip_version": getattr(torch.version, "hip", None),
"cuda_version": getattr(torch.version, "cuda", None),
"runtime_adapter": "native",
"requested_device_mode": device_mode,
"actual_device": actual_device,
"device_name": device_name if actual_device == "cuda" else None,
"audio_duration_seconds": duration,
"runtime_seconds": runtime,
"rtf": runtime / duration,
"speaker_count": len(speakers),
"speaker_labels": speakers,
"turn_count": len(ordinary_turns),
"exclusive_turn_count": len(exclusive_turns),
"audio": audio_metadata,
"credentials_persisted": False,
"output_files": {
"ordinary_rttm": "diarization.rttm",
"exclusive_rttm": "exclusive_diarization.rttm",
"turns": "turns.json",
"exclusive_turns": "exclusive_turns.json",
},
}
_write_json(output_dir / "metadata.json", metadata)
return _result_from_output(output_dir)
def run_container_pyannote(
audio_path: Path,
output_dir: Path,
device_mode: DeviceMode,
*,
image: str,
container_args: Sequence[str] = (),
runner: Callable[..., subprocess.CompletedProcess[str]] = subprocess.run,
uid_getter: Callable[[], int] = _host_uid,
gid_getter: Callable[[], int] = _host_gid,
) -> DiarizationResult:
"""Run the same backend in an explicitly configured disposable container."""
if not image.strip():
raise DiarizationError("A diarization container image is required.")
output_dir.mkdir(parents=True, exist_ok=True)
host_uid = uid_getter()
host_gid = gid_getter()
if host_uid < 0 or host_gid < 0:
raise DiarizationError("Host UID and GID must be non-negative integers.")
command = [
"docker", "run", "--rm", "--ipc=host", "--shm-size=8g", "-e", "HF_TOKEN",
*container_args,
"-v", f"{Path(audio_path).resolve()}:/input/audio.wav:ro",
"-v", f"{output_dir.resolve()}:/output:rw",
"-v", f"{Path(__file__).resolve().parents[3]}:/work/meeting-lab:ro",
"-w", "/work/meeting-lab",
image,
"/bin/bash", "-lc",
(
"inference_status=0; "
f"python -m pip install --disable-pip-version-check pyannote.audio=={PYANNOTE_VERSION} "
"> /output/pip-install.log 2>&1 && "
"python -m src.meeting_lab.diarization.container_entry "
f"/input/audio.wav /output --device {device_mode} || inference_status=$?; "
f"chown -R {host_uid}:{host_gid} /output || exit $?; "
"chmod -R u+rwX /output || exit $?; "
'exit "$inference_status"'
),
]
try:
completed = runner(command, check=False, capture_output=True, text=True)
except OSError as exc:
raise DiarizationError(f"Could not start diarization container: {exc}") from exc
(output_dir / "container_stdout.log").write_text(completed.stdout, encoding="utf-8")
(output_dir / "container_stderr.log").write_text(completed.stderr, encoding="utf-8")
if completed.returncode != 0:
detail = completed.stderr.strip() or completed.stdout.strip() or "no diagnostic output"
raise DiarizationError(
f"Diarization container failed with exit code {completed.returncode}: {detail}"
)
_require_writable_output(output_dir)
result = _result_from_output(output_dir)
metadata = dict(result.metadata)
metadata["runtime_adapter"] = "container"
_write_json(result.metadata_path, metadata)
return _result_from_output(output_dir)
def diarize_audio(
audio_path: Path,
output_dir: Path,
device_mode: DeviceMode,
*,
runtime: RuntimeMode = "native",
container_image: str | None = None,
container_args: Sequence[str] = (),
) -> DiarizationResult:
if runtime == "native":
return run_native_pyannote(audio_path, output_dir, device_mode)
if runtime == "container":
return run_container_pyannote(
audio_path,
output_dir,
device_mode,
image=container_image or "",
container_args=container_args,
)
raise DiarizationError(f"Unsupported diarization runtime: {runtime}")
@@ -0,0 +1,22 @@
"""Internal entry point for the isolated pyannote container adapter."""
from __future__ import annotations
import argparse
from pathlib import Path
from src.meeting_lab.diarization.backend import run_native_pyannote
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("audio", type=Path)
parser.add_argument("output", type=Path)
parser.add_argument("--device", choices=("auto", "gpu", "cpu"), required=True)
args = parser.parse_args()
run_native_pyannote(args.audio, args.output, args.device)
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1 @@
"""Isolated evidence-near observation experiment."""
@@ -0,0 +1,395 @@
#!/usr/bin/env python3
"""Extract evidence-near observations for a fixed Discussion Subject."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
import requests
SCHEMA_VERSION = "experimental-evidence-observations-v1"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
DEFAULT_MODEL = "qwen3.5:9B"
DEFAULT_TIMEOUT = 300
DEFAULT_NUM_CTX = 16384
DEFAULT_NUM_PREDICT = 4096
RELATIONS = {"none", "supports", "opposes", "qualifies", "limits_scope"}
MODALITIES = {
"factual",
"possible",
"suggested",
"interpersonal_request",
"impersonal_necessity",
"information_question",
"committed",
}
TEMPORALITIES = {"existing", "future", "completed", "unspecified"}
EVALUATIONS = {"positive", "negative", "none"}
AGREEMENTS = {"accepted", "rejected", "unclear", "none"}
RESPONSIBILITIES = {"none", "named", "accepted"}
UNCERTAINTIES = {"present", "absent"}
CLARIFICATION_NEEDS = {"explicit", "implicit", "none"}
OBSERVATION_ID_RE = re.compile(r"^obs_[1-9][0-9]*$")
class ObservationValidationError(ValueError):
"""Raised when an experimental fixture or model output is invalid."""
PROMPT_TEMPLATE = """You extract atomic, evidence-near observations for one fixed Discussion Subject.
Stop before protocol interpretation. Never classify anything as an idea, proposal,
objection, decision, action item, or open question. Do not determine protocol
eligibility, reconstruct topics, generate a protocol, or invent missing stages.
Split an evidence unit into multiple observations when it directly contains multiple
propositions. Preserve every observation's source evidence ID. Use concise content in
the evidence language.
Return exactly one JSON object with this shape:
{{
"schema_version": "experimental-evidence-observations-v1",
"subject_id": "copy exactly",
"subject": "copy exactly",
"observations": [
{{
"observation_id": "obs_1",
"evidence_id": "e1",
"content": "directly supported atomic observation",
"target": "discussion_subject",
"relation": "none",
"modality": "factual",
"temporality": "existing",
"evaluation": "none",
"agreement": "none",
"responsibility": "none",
"person": null,
"uncertainty": "absent",
"clarification_need": "none",
"scope": "absent"
}}
]
}}
Rules:
- Number observation_id sequentially as obs_1, obs_2, ... in evidence order.
- target is "discussion_subject", one earlier observation_id, or a non-empty list of
earlier observation_ids only when the evidence jointly refers to them.
- relation is only none, supports, opposes, qualifies, or limits_scope.
- modality is only factual, possible, suggested, interpersonal_request,
impersonal_necessity, information_question, or committed.
- interpersonal_request is a direct request to another person.
- impersonal_necessity says something needs to happen without assigning it.
- information_question expresses missing information without assigning work.
- temporality is only existing, future, completed, or unspecified.
- evaluation is positive, negative, or none. Do not infer evaluation from world
knowledge. A bare cost or technical fact normally has evaluation none.
- agreement is only accepted, rejected, unclear, or none and applies to target.
- responsibility is none, named, or accepted. Use named only for an explicitly
addressed candidate and accepted only for explicit acceptance/commitment.
- person is the explicit person's name for named/accepted responsibility; otherwise
use JSON null. Mentioning or speaking in first person does not establish ownership.
- uncertainty is present or absent.
- clarification_need is explicit, implicit, or none.
- scope is an evidence-grounded qualifier, or exactly "absent". Never use null or the
string "null" anywhere.
- Confirmation of a rejection targets the rejection observation, not the option.
- A trial-only qualification targets and limits the accepted trial.
- A negative consequence can oppose another observation without requiring
clarification.
- Personal preference is not group rejection.
- Collective "we" does not name an individual owner.
Fixed Gold input:
{input_json}
"""
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description="Run the evidence-observation Gold experiment.")
parser.add_argument("fixture", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT)
parser.add_argument("--num-ctx", type=int, default=DEFAULT_NUM_CTX)
parser.add_argument("--num-predict", type=int, default=DEFAULT_NUM_PREDICT)
return parser.parse_args()
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing = required - value.keys()
unknown = value.keys() - required
if missing:
raise ObservationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise ObservationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise ObservationValidationError(f"{location} must be a non-empty string")
result = value.strip()
if result.casefold() == "null":
raise ObservationValidationError(f"{location} must not be the string 'null'")
return result
OBSERVATION_KEYS = {
"observation_id", "evidence_id", "content", "target", "relation", "modality",
"temporality", "evaluation", "agreement", "responsibility", "person",
"uncertainty", "clarification_need", "scope",
}
def _validate_target(value: Any, location: str, earlier: set[str]) -> None:
if isinstance(value, str):
target = _text(value, location)
if target != "discussion_subject" and target not in earlier:
raise ObservationValidationError(f"{location} references unknown or later observation: {target}")
return
if not isinstance(value, list) or not value:
raise ObservationValidationError(f"{location} must be discussion_subject, an earlier observation ID, or a non-empty list")
if len(value) < 2:
raise ObservationValidationError(f"{location} list must contain at least two jointly referenced observations")
seen: set[str] = set()
for index, item in enumerate(value):
target = _text(item, f"{location}[{index}]")
if target not in earlier:
raise ObservationValidationError(f"{location}[{index}] references unknown or later observation: {target}")
if target in seen:
raise ObservationValidationError(f"{location} contains duplicate target: {target}")
seen.add(target)
def validate_observations(data: Any, case: dict[str, Any]) -> dict[str, Any]:
validate_case(case)
if not isinstance(data, dict):
raise ObservationValidationError("output must be an object")
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "output")
if data["schema_version"] != SCHEMA_VERSION:
raise ObservationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
if data["subject_id"] != case["subject_id"] or data["subject"] != case["subject"]:
raise ObservationValidationError("model changed the fixed Discussion Subject")
observations = data["observations"]
if not isinstance(observations, list) or not observations:
raise ObservationValidationError("output.observations must be a non-empty array")
known_evidence = {item["evidence_id"] for item in case["evidence"]}
earlier: set[str] = set()
for index, observation in enumerate(observations, start=1):
location = f"output.observations[{index - 1}]"
if not isinstance(observation, dict):
raise ObservationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
if not OBSERVATION_ID_RE.fullmatch(observation_id) or observation_id != f"obs_{index}":
raise ObservationValidationError(f"{location}.observation_id must be obs_{index}")
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
if evidence_id not in known_evidence:
raise ObservationValidationError(f"{location}.evidence_id references unknown evidence: {evidence_id}")
_text(observation["content"], f"{location}.content")
_validate_target(observation["target"], f"{location}.target", earlier)
for field, values in (
("relation", RELATIONS), ("modality", MODALITIES),
("temporality", TEMPORALITIES), ("evaluation", EVALUATIONS),
("agreement", AGREEMENTS), ("responsibility", RESPONSIBILITIES),
("uncertainty", UNCERTAINTIES), ("clarification_need", CLARIFICATION_NEEDS),
):
if observation[field] not in values:
raise ObservationValidationError(f"{location}.{field} is invalid: {observation[field]!r}")
person = observation["person"]
if observation["responsibility"] == "none":
if person is not None:
raise ObservationValidationError(f"{location}.person must be JSON null when responsibility is none")
else:
_text(person, f"{location}.person")
scope = _text(observation["scope"], f"{location}.scope")
if scope.casefold() == "null":
raise ObservationValidationError(f"{location}.scope must use 'absent', not 'null'")
earlier.add(observation_id)
return data
def validate_case(case: Any) -> dict[str, Any]:
if not isinstance(case, dict):
raise ObservationValidationError("case must be an object")
_exact_keys(case, {"case_id", "description", "subject_id", "subject", "evidence", "expected_observations"}, "case")
for field in ("case_id", "description", "subject_id", "subject"):
_text(case[field], f"case.{field}")
evidence = case["evidence"]
if not isinstance(evidence, list) or not evidence:
raise ObservationValidationError("case.evidence must be a non-empty array")
seen: set[str] = set()
for index, unit in enumerate(evidence):
location = f"case.evidence[{index}]"
if not isinstance(unit, dict):
raise ObservationValidationError(f"{location} must be an object")
_exact_keys(unit, {"evidence_id", "text"}, location)
evidence_id = _text(unit["evidence_id"], f"{location}.evidence_id")
if evidence_id in seen:
raise ObservationValidationError(f"duplicate evidence ID: {evidence_id}")
seen.add(evidence_id)
_text(unit["text"], f"{location}.text")
expected = case["expected_observations"]
if not isinstance(expected, list) or not expected:
raise ObservationValidationError("case.expected_observations must be a non-empty array")
return case
def validate_fixture_case(case: dict[str, Any]) -> dict[str, Any]:
validate_case(case)
data = {"schema_version": SCHEMA_VERSION, "subject_id": case["subject_id"], "subject": case["subject"], "observations": case["expected_observations"]}
validate_observations(data, case)
return case
def build_prompt(case: dict[str, Any]) -> str:
validate_fixture_case(case)
model_input = {"subject_id": case["subject_id"], "subject": case["subject"], "evidence": case["evidence"]}
return PROMPT_TEMPLATE.format(input_json=json.dumps(model_input, ensure_ascii=False, indent=2))
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise ObservationValidationError("model response JSON must be an object")
return data
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
started = time.perf_counter()
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
elapsed = time.perf_counter() - started
response.raise_for_status()
body = response.json()
raw_text = body.get("response") if isinstance(body, dict) else None
if not isinstance(raw_text, str) or not raw_text.strip():
raise ValueError("Ollama returned no usable response text")
metadata = {
"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3),
"total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"),
"prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"),
"eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"),
"configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0},
}
return raw_text.strip(), metadata
COMPARE_FIELDS = ("evidence_id", "target", "relation", "modality", "temporality", "evaluation", "agreement", "responsibility", "person", "uncertainty", "clarification_need")
def _scope_matches(actual: str, expected: str) -> bool:
if expected == "absent":
return actual == "absent"
expected_terms = [term.strip().casefold() for term in expected.split("|")]
folded = actual.casefold()
return any(term in folded for term in expected_terms)
def evaluate_observations(data: dict[str, Any], expected: list[dict[str, Any]]) -> dict[str, Any]:
actual = data["observations"]
checks: list[dict[str, Any]] = []
pair_count = min(len(actual), len(expected))
checks.append({"name": "observation_count", "passed": len(actual) == len(expected), "critical": False})
categories = {"missing_observations": max(0, len(expected) - len(actual)), "invented_observations": max(0, len(actual) - len(expected)), "stronger_commitment": 0, "weaker_commitment": 0, "incorrect_targets_relations": 0, "incorrect_responsibility": 0, "incorrect_uncertainty_clarification": 0}
commitment_rank = {"factual": 0, "possible": 1, "suggested": 1, "information_question": 1, "impersonal_necessity": 2, "interpersonal_request": 2, "committed": 3}
for index in range(pair_count):
got, want = actual[index], expected[index]
for field in COMPARE_FIELDS:
passed = got[field] == want[field]
checks.append({"name": f"obs_{index + 1}:{field}", "passed": passed, "critical": field in {"evidence_id", "target", "relation", "modality", "agreement", "responsibility", "person"}})
if not passed:
if field in {"target", "relation"}: categories["incorrect_targets_relations"] += 1
if field in {"responsibility", "person"}: categories["incorrect_responsibility"] += 1
if field in {"uncertainty", "clarification_need"}: categories["incorrect_uncertainty_clarification"] += 1
scope_ok = _scope_matches(got["scope"], want["scope"])
checks.append({"name": f"obs_{index + 1}:scope", "passed": scope_ok, "critical": False})
got_rank, want_rank = commitment_rank[got["modality"]], commitment_rank[want["modality"]]
if got_rank > want_rank or (want["agreement"] == "none" and got["agreement"] in {"accepted", "rejected"}): categories["stronger_commitment"] += 1
if got_rank < want_rank or (want["agreement"] in {"accepted", "rejected"} and got["agreement"] == "none"): categories["weaker_commitment"] += 1
passed_count = sum(check["passed"] for check in checks)
critical_failures = [check["name"] for check in checks if check["critical"] and not check["passed"]]
ratio = passed_count / len(checks)
if ratio == 1:
verdict = "PASS"
elif ratio >= 0.7 and categories["stronger_commitment"] == 0 and categories["incorrect_responsibility"] == 0:
verdict = "PARTIAL"
else:
verdict = "FAIL"
return {"verdict": verdict, "matched_checks": passed_count, "check_count": len(checks), "match_ratio": round(ratio, 3), "critical_failures": critical_failures, "error_categories": categories, "checks": checks}
def load_fixture(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict) or set(data) != {"cases"} or not isinstance(data["cases"], list) or not data["cases"]:
raise ObservationValidationError("fixture must contain exactly one non-empty cases list")
seen: set[str] = set()
for case in data["cases"]:
validate_fixture_case(case)
if case["case_id"] in seen:
raise ObservationValidationError(f"duplicate case ID: {case['case_id']}")
seen.add(case["case_id"])
return data["cases"]
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_case(case: dict[str, Any], output_root: Path, endpoint: str, model: str, timeout: int, num_ctx: int, num_predict: int) -> dict[str, Any]:
case_dir = output_root / case["case_id"]
case_dir.mkdir(parents=True, exist_ok=False)
_write_json(case_dir / "gold_input.json", {key: case[key] for key in ("case_id", "description", "subject_id", "subject", "evidence")})
_write_json(case_dir / "gold_expected_observations.json", case["expected_observations"])
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
started = time.perf_counter()
try:
raw, metadata = call_ollama(endpoint, model, prompt, timeout, num_ctx, num_predict)
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_observations.json", parsed)
validate_observations(parsed, case)
validation = {"valid": True, "error": None}
evaluation = evaluate_observations(parsed, case["expected_observations"])
except requests.RequestException:
raise
except (json.JSONDecodeError, ObservationValidationError, ValueError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
evaluation = {"verdict": "FAIL", "matched_checks": 0, "check_count": 0, "match_ratio": 0, "critical_failures": ["schema_validation"], "error_categories": {}, "checks": []}
_write_json(case_dir / "validation_result.json", validation)
result = {"case_id": case["case_id"], **evaluation, "elapsed_seconds": round(time.perf_counter() - started, 3)}
_write_json(case_dir / "evaluation_result.json", result)
return result
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
cases = load_fixture(args.fixture)
args.output.mkdir(parents=True, exist_ok=False)
started = time.perf_counter()
results = []
for index, case in enumerate(cases, start=1):
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
results.append(run_case(case, args.output, args.endpoint, args.model, args.timeout, args.num_ctx, args.num_predict))
summary = {"experiment": "evidence_near_observation_extraction", "schema_version": SCHEMA_VERSION, "model": args.model, "temperature": 0, "think": False, "retries": 0, "case_count": len(cases), "llm_call_count": len(results), "runtime_seconds": round(time.perf_counter() - started, 3), "verdict_counts": {v: sum(r["verdict"] == v for r in results) for v in ("PASS", "PARTIAL", "FAIL")}, "results": results}
_write_json(args.output / "summary.json", summary)
return summary
def main() -> int:
args = parse_args()
summary = run_experiment(args)
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["verdict_counts"]["FAIL"] == 0 else 1
@@ -0,0 +1 @@
"""Reduced-semantic-load evidence observation experiment."""
@@ -0,0 +1,344 @@
#!/usr/bin/env python3
"""Extract reduced-semantic-load evidence-near observations."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
import requests
SCHEMA_VERSION = "experimental-evidence-observations-v2"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
DEFAULT_MODEL = "qwen3.5:9B"
MODALITIES = {"factual", "possible", "suggested", "interpersonal_request", "impersonal_necessity", "information_question", "committed"}
TEMPORALITIES = {"existing", "future", "completed", "unspecified"}
EVALUATIONS = {"positive", "negative", "none"}
BINARY_SIGNALS = {"explicit", "absent"}
PRESENCE_SIGNALS = {"present", "absent"}
CLARIFICATION_NEEDS = {"explicit", "implicit", "none"}
OBSERVATION_ID_RE = re.compile(r"^obs_[1-9][0-9]*$")
class ObservationValidationError(ValueError):
"""Raised for invalid fixtures or model output."""
PROMPT_TEMPLATE = """You extract atomic linguistic and discourse observations for one fixed Discussion Subject.
Preserve only facts directly expressed by the evidence. Do not derive responsibility,
agreement, decisions, action items, open questions, accepted trials, rejected
alternatives, established actions, or protocol eligibility. Speaker identity, a name,
an addressee, first-person language, collective "we", and impersonal "man" never by
themselves establish responsibility.
Return exactly one JSON object with this shape:
{{
"schema_version": "experimental-evidence-observations-v2",
"subject_id": "copy exactly",
"subject": "copy exactly",
"observations": [
{{
"observation_id": "obs_1",
"evidence_id": "e1",
"content": "directly supported atomic observation",
"refers_to": null,
"speaker": "name copied from evidence or null",
"named_person": null,
"addressee": null,
"self_reference": false,
"collective_we": false,
"impersonal_person_reference": false,
"modality": "factual",
"temporality": "existing",
"evaluation": "none",
"affirmation": "absent",
"negation": "absent",
"determination_statement": "absent",
"uncertainty": "absent",
"clarification_need": "none",
"qualifier": null,
"limits_target": null
}}
]
}}
Rules:
- Produce multiple observations for distinct propositions in one evidence unit, but do
not fragment a single proposition unnecessarily.
- observation_id is sequential in evidence order. evidence_id must be copied exactly.
- refers_to is null or one earlier observation_id when the utterance explicitly refers
to it. Never use arrays. Preserve joint-reference utterances without inventing a
multi-target graph.
- speaker is the explicit transcript speaker. named_person is a person explicitly
named in the proposition. addressee is a person explicitly addressed.
- self_reference marks singular first-person self-reference. collective_we marks
collective first-person language. impersonal_person_reference marks impersonal
person expressions such as German "man".
- modality is factual, possible, suggested, interpersonal_request,
impersonal_necessity, information_question, or committed.
- temporality is existing, future, completed, or unspecified.
- evaluation is positive, negative, or none, only when linguistically supported.
- affirmation is explicit only for an explicit affirmative discourse signal such as
"ja". negation is explicit only for directly expressed negation/rejection.
- determination_statement is present only when the utterance explicitly says a
determination has been made.
- uncertainty is present or absent. clarification_need is explicit, implicit, or none.
- qualifier is null or concise evidence-grounded qualifying text.
- limits_target is null or one earlier observation explicitly limited in validity or
scope by this observation.
- Use JSON null, never the string "null". Output no fields beyond the schema.
Fixed Gold input:
{input_json}
"""
OBSERVATION_KEYS = {
"observation_id", "evidence_id", "content", "refers_to", "speaker",
"named_person", "addressee", "self_reference", "collective_we",
"impersonal_person_reference", "modality", "temporality", "evaluation",
"affirmation", "negation", "determination_statement", "uncertainty",
"clarification_need", "qualifier", "limits_target",
}
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("fixture", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=4096)
return parser.parse_args()
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing, unknown = required - value.keys(), value.keys() - required
if missing:
raise ObservationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise ObservationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise ObservationValidationError(f"{location} must be a non-empty string")
result = value.strip()
if result.casefold() == "null":
raise ObservationValidationError(f"{location} must not be the string 'null'")
return result
def _nullable_text(value: Any, location: str) -> None:
if value is not None:
_text(value, location)
def _prior_reference(value: Any, location: str, earlier: set[str]) -> None:
if value is None:
return
reference = _text(value, location)
if reference not in earlier:
raise ObservationValidationError(f"{location} references unknown or later observation: {reference}")
def validate_observations(data: Any, case: dict[str, Any]) -> dict[str, Any]:
validate_case(case)
if not isinstance(data, dict):
raise ObservationValidationError("output must be an object")
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "output")
if data["schema_version"] != SCHEMA_VERSION:
raise ObservationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
if data["subject_id"] != case["subject_id"] or data["subject"] != case["subject"]:
raise ObservationValidationError("model changed the fixed Discussion Subject")
observations = data["observations"]
if not isinstance(observations, list) or not observations:
raise ObservationValidationError("output.observations must be a non-empty array")
known_evidence = {item["evidence_id"] for item in case["evidence"]}
earlier: set[str] = set()
for index, observation in enumerate(observations, 1):
location = f"output.observations[{index - 1}]"
if not isinstance(observation, dict):
raise ObservationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
if not OBSERVATION_ID_RE.fullmatch(observation_id) or observation_id != f"obs_{index}":
raise ObservationValidationError(f"{location}.observation_id must be obs_{index}")
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
if evidence_id not in known_evidence:
raise ObservationValidationError(f"{location}.evidence_id references unknown evidence: {evidence_id}")
_text(observation["content"], f"{location}.content")
_prior_reference(observation["refers_to"], f"{location}.refers_to", earlier)
_prior_reference(observation["limits_target"], f"{location}.limits_target", earlier)
for field in ("speaker", "named_person", "addressee", "qualifier"):
_nullable_text(observation[field], f"{location}.{field}")
for field in ("self_reference", "collective_we", "impersonal_person_reference"):
if not isinstance(observation[field], bool):
raise ObservationValidationError(f"{location}.{field} must be boolean")
for field, values in (
("modality", MODALITIES), ("temporality", TEMPORALITIES),
("evaluation", EVALUATIONS), ("affirmation", BINARY_SIGNALS),
("negation", BINARY_SIGNALS), ("determination_statement", PRESENCE_SIGNALS),
("uncertainty", PRESENCE_SIGNALS), ("clarification_need", CLARIFICATION_NEEDS),
):
if observation[field] not in values:
raise ObservationValidationError(f"{location}.{field} is invalid: {observation[field]!r}")
earlier.add(observation_id)
return data
def validate_case(case: Any) -> dict[str, Any]:
if not isinstance(case, dict):
raise ObservationValidationError("case must be an object")
_exact_keys(case, {"case_id", "description", "subject_id", "subject", "evidence", "expected_observations"}, "case")
for field in ("case_id", "description", "subject_id", "subject"):
_text(case[field], f"case.{field}")
if not isinstance(case["evidence"], list) or not case["evidence"]:
raise ObservationValidationError("case.evidence must be a non-empty array")
seen: set[str] = set()
for index, unit in enumerate(case["evidence"]):
_exact_keys(unit, {"evidence_id", "text"}, f"case.evidence[{index}]")
evidence_id = _text(unit["evidence_id"], f"case.evidence[{index}].evidence_id")
if evidence_id in seen:
raise ObservationValidationError(f"duplicate evidence ID: {evidence_id}")
seen.add(evidence_id)
_text(unit["text"], f"case.evidence[{index}].text")
if not isinstance(case["expected_observations"], list) or not case["expected_observations"]:
raise ObservationValidationError("case.expected_observations must be a non-empty array")
return case
def validate_fixture_case(case: dict[str, Any]) -> dict[str, Any]:
validate_case(case)
validate_observations({"schema_version": SCHEMA_VERSION, "subject_id": case["subject_id"], "subject": case["subject"], "observations": case["expected_observations"]}, case)
return case
def build_prompt(case: dict[str, Any]) -> str:
validate_fixture_case(case)
model_input = {key: case[key] for key in ("subject_id", "subject", "evidence")}
return PROMPT_TEMPLATE.format(input_json=json.dumps(model_input, ensure_ascii=False, indent=2))
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise ObservationValidationError("model response JSON must be an object")
return data
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
started = time.perf_counter()
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
elapsed = time.perf_counter() - started
response.raise_for_status()
body = response.json()
raw = body.get("response") if isinstance(body, dict) else None
if not isinstance(raw, str) or not raw.strip():
raise ValueError("Ollama returned no usable response text")
metadata = {"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3), "total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"), "prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"), "eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"), "configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0}}
return raw.strip(), metadata
COMPARE_FIELDS = tuple(sorted(OBSERVATION_KEYS - {"observation_id", "content", "qualifier"}))
def _qualifier_matches(actual: str | None, expected: str | None) -> bool:
if expected is None:
return actual is None
if actual is None:
return False
return any(term.strip().casefold() in actual.casefold() for term in expected.split("|"))
def evaluate_observations(data: dict[str, Any], expected: list[dict[str, Any]]) -> dict[str, Any]:
actual = data["observations"]
checks = [{"name": "observation_count", "passed": len(actual) == len(expected), "critical": False}]
for index, (got, want) in enumerate(zip(actual, expected), 1):
for field in COMPARE_FIELDS:
checks.append({"name": f"obs_{index}:{field}", "passed": got[field] == want[field], "critical": field in {"evidence_id", "refers_to", "limits_target", "modality", "affirmation", "negation", "determination_statement"}})
checks.append({"name": f"obs_{index}:qualifier", "passed": _qualifier_matches(got["qualifier"], want["qualifier"]), "critical": False})
passed = sum(check["passed"] for check in checks)
ratio = passed / len(checks)
critical = [check["name"] for check in checks if check["critical"] and not check["passed"]]
verdict = "PASS" if ratio == 1 else "PARTIAL" if ratio >= 0.75 and not critical else "FAIL"
return {"verdict": verdict, "matched_checks": passed, "check_count": len(checks), "match_ratio": round(ratio, 3), "critical_failures": critical, "checks": checks}
def load_fixture(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict) or set(data) != {"cases"} or not isinstance(data["cases"], list) or not data["cases"]:
raise ObservationValidationError("fixture must contain exactly one non-empty cases list")
seen: set[str] = set()
for case in data["cases"]:
validate_fixture_case(case)
if case["case_id"] in seen:
raise ObservationValidationError(f"duplicate case ID: {case['case_id']}")
seen.add(case["case_id"])
return data["cases"]
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_case(case: dict[str, Any], output_root: Path, endpoint: str, model: str, timeout: int, num_ctx: int, num_predict: int) -> dict[str, Any]:
case_dir = output_root / case["case_id"]
case_dir.mkdir(parents=True, exist_ok=False)
_write_json(case_dir / "gold_input.json", {key: case[key] for key in ("case_id", "description", "subject_id", "subject", "evidence")})
_write_json(case_dir / "gold_expected_observations.json", case["expected_observations"])
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
started = time.perf_counter()
raw, metadata = call_ollama(endpoint, model, prompt, timeout, num_ctx, num_predict)
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
try:
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_observations.json", parsed)
validate_observations(parsed, case)
validation = {"valid": True, "error": None}
evaluation = evaluate_observations(parsed, case["expected_observations"])
except (json.JSONDecodeError, ObservationValidationError, ValueError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
evaluation = {"verdict": "FAIL", "matched_checks": 0, "check_count": 0, "match_ratio": 0, "critical_failures": ["schema_validation"], "checks": []}
_write_json(case_dir / "validation_result.json", validation)
result = {"case_id": case["case_id"], **evaluation, "elapsed_seconds": round(time.perf_counter() - started, 3)}
_write_json(case_dir / "evaluation_result.json", result)
return result
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
cases = load_fixture(args.fixture)
args.output.mkdir(parents=True, exist_ok=False)
started = time.perf_counter()
results = []
for index, case in enumerate(cases, 1):
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
results.append(run_case(case, args.output, args.endpoint, args.model, args.timeout, args.num_ctx, args.num_predict))
summary = {"experiment": "evidence_near_observation_extraction_v2", "schema_version": SCHEMA_VERSION, "model": args.model, "temperature": 0, "think": False, "retries": 0, "case_count": len(cases), "llm_call_count": len(results), "runtime_seconds": round(time.perf_counter() - started, 3), "verdict_counts": {verdict: sum(result["verdict"] == verdict for result in results) for verdict in ("PASS", "PARTIAL", "FAIL")}, "results": results}
_write_json(args.output / "summary.json", summary)
return summary
def main() -> int:
args = parse_args()
summary = run_experiment(args)
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["verdict_counts"]["FAIL"] == 0 else 1
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1 @@
"""Minimal semantic-preservation observation experiment."""
@@ -0,0 +1,271 @@
#!/usr/bin/env python3
"""Preserve meeting meaning as minimal atomic natural-language observations."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
import requests
SCHEMA_VERSION = "experimental-evidence-observations-v3"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
DEFAULT_MODEL = "qwen3.5:9B"
OBSERVATION_ID_RE = re.compile(r"^obs_[1-9][0-9]*$")
OBSERVATION_KEYS = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
class ObservationValidationError(ValueError):
"""Raised for invalid fixtures or model output."""
PROMPT_TEMPLATE = """Preserve the meeting meaning in atomic natural-language observations.
This is semantic preservation, not classification or summarization. Return only facts
faithfully contributed by the evidence. Conservative wording is more important than
elegant prose. When in doubt, preserve the source wording closely.
Return exactly one JSON object:
{{
"schema_version": "experimental-evidence-observations-v3",
"subject_id": "copy exactly",
"subject": "copy exactly",
"observations": [
{{
"observation_id": "obs_1",
"evidence_id": "e1",
"content": "atomic, semantically faithful observation",
"speaker": "speaker copied from evidence",
"named_person": null,
"addressee": null
}}
]
}}
Rules:
- Use only the six observation fields shown. Do not output classifications, labels,
relations, scope fields, responsibility, agreement, decisions, actions, questions,
eligibility, or any other field.
- observation_id is sequential in evidence order. Copy evidence_id and speaker.
- named_person is null or a person explicitly named in that observation's evidence.
- addressee is null or a person explicitly addressed in that observation's evidence.
- A name, speaker, or addressee never implies responsibility, acceptance, ownership,
or assignment.
- content is not a summary. Preserve distinctions needed for later interpretation:
maybe/perhaps; can/could; should/must; personal, collective, or impersonal wording;
explicit requests, acceptances, and rejections; uncertainty and unresolved status;
conditions such as "if at all"; quantities; deadlines; trial/process/comparison
boundaries; "not yet"; and sequence such as "then".
- Never strengthen modality, weaken uncertainty, turn possibility into fact, turn a
preference into group rejection, turn a request into established work, turn "we"
into individual ownership, remove conditions/limits, generalize, or invent relations.
- Split one evidence unit only when it contributes propositions that may later require
different interpretations. Do not split merely because it has several clauses.
- Do not emit observation-ID relations. When evidence clearly makes an observation
depend on the immediately preceding proposition, state that dependency naturally in
content, without inventing an antecedent.
- Preserve content in the evidence language. Use JSON null, never the string "null".
Fixed input:
{input_json}
"""
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser()
parser.add_argument("fixture", type=Path)
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=300)
parser.add_argument("--num-ctx", type=int, default=16384)
parser.add_argument("--num-predict", type=int, default=4096)
return parser.parse_args()
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
missing, unknown = required - value.keys(), value.keys() - required
if missing:
raise ObservationValidationError(f"{location} missing required keys: {sorted(missing)}")
if unknown:
raise ObservationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
def _text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise ObservationValidationError(f"{location} must be a non-empty string")
result = value.strip()
if result.casefold() == "null":
raise ObservationValidationError(f"{location} must not be the string 'null'")
return result
def _explicit_people(text: str) -> set[str]:
prefix = text.split(":", 1)[0].strip() if ":" in text else ""
candidates = set(re.findall(r"\b(?:Dr\.\s+)?[A-ZÄÖÜ][A-Za-zÄÖÜäöüß-]+(?:\s+[A-ZÄÖÜ][A-Za-zÄÖÜäöüß-]+)*", text))
candidates.discard(prefix)
return candidates
def validate_observations(data: Any, case: dict[str, Any]) -> dict[str, Any]:
validate_case(case)
if not isinstance(data, dict):
raise ObservationValidationError("output must be an object")
_exact_keys(data, {"schema_version", "subject_id", "subject", "observations"}, "output")
if data["schema_version"] != SCHEMA_VERSION:
raise ObservationValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
if data["subject_id"] != case["subject_id"] or data["subject"] != case["subject"]:
raise ObservationValidationError("model changed the fixed Discussion Subject")
observations = data["observations"]
if not isinstance(observations, list) or not observations:
raise ObservationValidationError("output.observations must be a non-empty list")
evidence = {item["evidence_id"]: item["text"] for item in case["evidence"]}
seen: set[str] = set()
for index, observation in enumerate(observations):
location = f"output.observations[{index}]"
if not isinstance(observation, dict):
raise ObservationValidationError(f"{location} must be an object")
_exact_keys(observation, OBSERVATION_KEYS, location)
observation_id = _text(observation["observation_id"], f"{location}.observation_id")
if not OBSERVATION_ID_RE.fullmatch(observation_id) or observation_id in seen:
raise ObservationValidationError(f"{location}.observation_id must be unique and match obs_N")
seen.add(observation_id)
evidence_id = _text(observation["evidence_id"], f"{location}.evidence_id")
if evidence_id not in evidence:
raise ObservationValidationError(f"{location}.evidence_id references unknown evidence: {evidence_id}")
source = evidence[evidence_id]
source_speaker = source.split(":", 1)[0].strip()
speaker = _text(observation["speaker"], f"{location}.speaker")
if speaker != source_speaker:
raise ObservationValidationError(f"{location}.speaker must match evidence speaker {source_speaker!r}")
_text(observation["content"], f"{location}.content")
explicit_people = _explicit_people(source)
for field in ("named_person", "addressee"):
person = observation[field]
if person is not None:
person = _text(person, f"{location}.{field}")
if person not in explicit_people:
raise ObservationValidationError(f"{location}.{field} is not an explicit person in evidence: {person!r}")
return data
def validate_case(case: Any) -> dict[str, Any]:
required = {"case_id", "description", "subject_id", "subject", "evidence", "semantic_requirements"}
if not isinstance(case, dict):
raise ObservationValidationError("case must be an object")
_exact_keys(case, required, "case")
for field in ("case_id", "description", "subject_id", "subject"):
_text(case[field], f"case.{field}")
if not isinstance(case["evidence"], list) or not case["evidence"]:
raise ObservationValidationError("case.evidence must be a non-empty list")
evidence_ids: set[str] = set()
for index, unit in enumerate(case["evidence"]):
_exact_keys(unit, {"evidence_id", "text"}, f"case.evidence[{index}]")
evidence_id = _text(unit["evidence_id"], f"case.evidence[{index}].evidence_id")
if evidence_id in evidence_ids:
raise ObservationValidationError(f"duplicate evidence ID: {evidence_id}")
evidence_ids.add(evidence_id)
_text(unit["text"], f"case.evidence[{index}].text")
if not isinstance(case["semantic_requirements"], list) or not case["semantic_requirements"]:
raise ObservationValidationError("case.semantic_requirements must be a non-empty list")
for index, requirement in enumerate(case["semantic_requirements"]):
_text(requirement, f"case.semantic_requirements[{index}]")
return case
def build_prompt(case: dict[str, Any]) -> str:
validate_case(case)
model_input = {key: case[key] for key in ("subject_id", "subject", "evidence")}
return PROMPT_TEMPLATE.format(input_json=json.dumps(model_input, ensure_ascii=False, indent=2))
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise ObservationValidationError("model response JSON must be an object")
return data
def build_ollama_payload(model: str, prompt: str, num_ctx: int, num_predict: int) -> dict[str, Any]:
return {"model": model, "prompt": prompt, "think": False, "stream": False, "format": "json", "options": {"temperature": 0, "num_ctx": num_ctx, "num_predict": num_predict}}
def call_ollama(endpoint: str, model: str, prompt: str, timeout: int, num_ctx: int, num_predict: int) -> tuple[str, dict[str, Any]]:
started = time.perf_counter()
response = requests.post(endpoint, json=build_ollama_payload(model, prompt, num_ctx, num_predict), timeout=timeout)
elapsed = time.perf_counter() - started
response.raise_for_status()
body = response.json()
raw = body.get("response") if isinstance(body, dict) else None
if not isinstance(raw, str) or not raw.strip():
raise ValueError("Ollama returned no usable response text")
metadata = {"model": body.get("model", model), "elapsed_seconds": round(elapsed, 3), "total_duration_ns": body.get("total_duration"), "load_duration_ns": body.get("load_duration"), "prompt_eval_count": body.get("prompt_eval_count"), "prompt_eval_duration_ns": body.get("prompt_eval_duration"), "eval_count": body.get("eval_count"), "eval_duration_ns": body.get("eval_duration"), "configuration": {"temperature": 0, "think": False, "num_ctx": num_ctx, "num_predict": num_predict, "retries": 0}}
return raw.strip(), metadata
def load_fixture(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict) or set(data) != {"cases"} or not isinstance(data["cases"], list) or not data["cases"]:
raise ObservationValidationError("fixture must contain exactly one non-empty cases list")
seen: set[str] = set()
for case in data["cases"]:
validate_case(case)
if case["case_id"] in seen:
raise ObservationValidationError(f"duplicate case ID: {case['case_id']}")
seen.add(case["case_id"])
return data["cases"]
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def run_case(case: dict[str, Any], output_root: Path, endpoint: str, model: str, timeout: int, num_ctx: int, num_predict: int) -> dict[str, Any]:
case_dir = output_root / case["case_id"]
case_dir.mkdir(parents=True, exist_ok=False)
_write_json(case_dir / "source_evidence.json", {key: case[key] for key in ("case_id", "description", "subject_id", "subject", "evidence")})
_write_json(case_dir / "gold_semantic_requirements.json", case["semantic_requirements"])
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
started = time.perf_counter()
raw, metadata = call_ollama(endpoint, model, prompt, timeout, num_ctx, num_predict)
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
_write_json(case_dir / "ollama_metadata.json", metadata)
try:
parsed = parse_model_json(raw)
_write_json(case_dir / "parsed_observations.json", parsed)
validate_observations(parsed, case)
validation = {"valid": True, "error": None}
except (json.JSONDecodeError, ObservationValidationError, ValueError) as exc:
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
_write_json(case_dir / "structural_validation.json", validation)
return {"case_id": case["case_id"], "structurally_valid": validation["valid"], "elapsed_seconds": round(time.perf_counter() - started, 3)}
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
cases = load_fixture(args.fixture)
args.output.mkdir(parents=True, exist_ok=False)
started = time.perf_counter()
results = []
for index, case in enumerate(cases, 1):
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
results.append(run_case(case, args.output, args.endpoint, args.model, args.timeout, args.num_ctx, args.num_predict))
summary = {"experiment": "evidence_near_observation_extraction_v3", "schema_version": SCHEMA_VERSION, "model": args.model, "temperature": 0, "think": False, "retries": 0, "case_count": len(cases), "llm_call_count": len(results), "runtime_seconds": round(time.perf_counter() - started, 3), "structurally_valid_count": sum(result["structurally_valid"] for result in results), "results": results}
_write_json(args.output / "summary.json", summary)
return summary
def main() -> int:
args = parse_args()
summary = run_experiment(args)
print(json.dumps(summary, ensure_ascii=False, indent=2))
return 0 if summary["structurally_valid_count"] == summary["case_count"] else 1
if __name__ == "__main__":
raise SystemExit(main())
+95
View File
@@ -0,0 +1,95 @@
"""Minimal Ollama client behavior used by the direct protocol MVP."""
from __future__ import annotations
import time
from dataclasses import dataclass
from typing import Any
import requests
DEFAULT_ENDPOINT = "http://127.0.0.1:11434"
class OllamaError(RuntimeError):
"""Raised when Ollama cannot safely complete the requested operation."""
@dataclass(frozen=True)
class OllamaGeneration:
raw_response: dict[str, Any]
text: str
client_wall_time_seconds: float
def ollama_base_url(endpoint: str) -> str:
endpoint = endpoint.rstrip("/")
return endpoint.rsplit("/api/", 1)[0] if "/api/" in endpoint else endpoint
def generate_url(endpoint: str) -> str:
return f"{ollama_base_url(endpoint)}/api/generate"
def require_model(endpoint: str, model: str, timeout: int = 10) -> dict[str, Any]:
base_url = ollama_base_url(endpoint)
try:
response = requests.get(f"{base_url}/api/tags", timeout=timeout)
response.raise_for_status()
data = response.json()
except (requests.RequestException, ValueError) as exc:
raise OllamaError(f"Ollama endpoint is not reachable at {base_url}: {exc}") from exc
models = data.get("models") if isinstance(data, dict) else None
if not isinstance(models, list):
raise OllamaError("Ollama /api/tags returned a malformed response.")
installed = {
item.get("name")
for item in models
if isinstance(item, dict) and isinstance(item.get("name"), str)
}
if model not in installed:
raise OllamaError(f"Requested model is not installed in Ollama: {model}")
return {"base_url": base_url, "model": model, "installed": True}
def generate_once(
endpoint: str,
model: str,
prompt: str,
*,
timeout: int,
num_ctx: int,
num_predict: int,
) -> OllamaGeneration:
payload = {
"model": model,
"prompt": prompt,
"think": False,
"stream": False,
"options": {
"temperature": 0.0,
"num_ctx": num_ctx,
"num_predict": num_predict,
},
}
started = time.perf_counter()
try:
response = requests.post(generate_url(endpoint), json=payload, timeout=timeout)
response.raise_for_status()
data = response.json()
except requests.RequestException as exc:
raise OllamaError(f"Ollama generation request failed: {exc}") from exc
except ValueError as exc:
raise OllamaError("Ollama generation response is not valid JSON.") from exc
wall_time = time.perf_counter() - started
if not isinstance(data, dict):
raise OllamaError("Ollama generation response must be a JSON object.")
text = data.get("response")
if not isinstance(text, str):
raise OllamaError("Ollama generation response has no string 'response' field.")
if not text.strip():
raise OllamaError("Ollama returned an empty protocol.")
return OllamaGeneration(data, text, wall_time)
+119 -11
View File
@@ -3,13 +3,15 @@
from __future__ import annotations
import ast
import copy
import re
from dataclasses import dataclass
from pathlib import Path
from typing import Any
SUPPORTED_SCHEMA_VERSIONS = {"1"}
VALID_ATTENDANCE_STATUSES = {"present", "not_present", "absent"}
VALID_ATTENDANCE_STATUSES = {"present", "mentioned_only"}
class MeetingContextValidationError(ValueError):
@@ -29,6 +31,21 @@ class MeetingContext:
def meeting_id(self) -> str:
return str(self.data["meeting"]["meeting_id"])
@property
def speaker_mappings(self) -> dict[str, str]:
mappings = self.data.get("speaker_mappings")
return dict(mappings) if isinstance(mappings, dict) else {}
def participant_for_speaker(self, speaker_label: str) -> dict[str, Any] | None:
"""Resolve only an explicit authoritative mapping; never infer identity."""
participant_id = self.speaker_mappings.get(speaker_label)
if participant_id is None:
return None
for participant in self.data.get("participants", []):
if participant.get("participant_id") == participant_id:
return participant
return None
def provenance(self) -> dict[str, str]:
return {
"meeting_id": self.meeting_id,
@@ -42,8 +59,43 @@ def load_meeting_context(path: Path) -> MeetingContext:
if not isinstance(loaded, dict):
raise MeetingContextValidationError("Meeting Context must be a YAML object.")
validate_meeting_context(loaded)
return MeetingContext(data=loaded, source_file=path)
normalized = _with_attendance_defaults(loaded)
validate_meeting_context(normalized)
return MeetingContext(data=normalized, source_file=path)
def create_meeting_context(
data: dict[str, Any], *, source_file: Path = Path("<generated>")
) -> MeetingContext:
"""Validate structured data and return an immutable context boundary."""
validated = _with_attendance_defaults(data)
validate_meeting_context(validated)
return MeetingContext(data=validated, source_file=source_file)
def serialize_meeting_context_yaml(context: MeetingContext) -> str:
"""Serialize validated Meeting Context data deterministically as YAML."""
validate_meeting_context(context.data)
try:
import yaml # type: ignore[import-not-found]
except ModuleNotFoundError as exc:
raise MeetingContextValidationError(
"PyYAML is required to write Meeting Context YAML."
) from exc
return yaml.safe_dump(
context.data,
allow_unicode=True,
sort_keys=False,
default_flow_style=False,
)
def write_meeting_context(context: MeetingContext, path: Path) -> Path:
"""Persist a validated context without changing its schema or semantics."""
path = Path(path)
path.parent.mkdir(parents=True, exist_ok=True)
path.write_text(serialize_meeting_context_yaml(context), encoding="utf-8")
return path
def validate_meeting_context(data: dict[str, Any]) -> None:
@@ -56,7 +108,9 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
meeting = _require_mapping(data, "meeting")
_require_non_empty_string(meeting, "meeting.meeting_id")
_require_non_empty_string(meeting, "meeting.title")
_require_non_empty_string(meeting, "meeting.language")
# Legacy contexts may omit language; protocol generation defaults to German.
if "language" in meeting:
_require_non_empty_string(meeting, "meeting.language")
organization = _optional_mapping(data.get("organization"), "organization")
departments = _optional_list(organization.get("departments"), "organization.departments")
@@ -74,10 +128,26 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
+ ", ".join(collisions)
)
speaker_mappings = _optional_mapping(
data.get("speaker_mappings"), "speaker_mappings"
)
for speaker_label, participant_id in speaker_mappings.items():
if not isinstance(speaker_label, str) or re.fullmatch(
r"SPEAKER_\d+", speaker_label
) is None:
raise MeetingContextValidationError(
f"Invalid diarization speaker label: {speaker_label!r}."
)
if not isinstance(participant_id, str) or participant_id not in participant_ids:
raise MeetingContextValidationError(
f"speaker_mappings.{speaker_label} references unknown participant: "
f"{participant_id!r}."
)
for index, participant in enumerate(participants):
item_path = f"participants[{index}]"
_validate_attendance(participant, item_path)
if participant.get("attendance_status") != "present":
status = _validate_attendance(participant, item_path, default="present")
if status != "present":
raise MeetingContextValidationError(
f"{item_path}.attendance_status must be 'present'."
)
@@ -85,10 +155,10 @@ def validate_meeting_context(data: dict[str, Any]) -> None:
for index, person in enumerate(mentioned_people):
item_path = f"mentioned_people[{index}]"
_validate_attendance(person, item_path)
if person.get("attendance_status") == "present":
status = _validate_attendance(person, item_path, default="mentioned_only")
if status != "mentioned_only":
raise MeetingContextValidationError(
f"{item_path}.attendance_status must not be 'present'."
f"{item_path}.attendance_status must be 'mentioned_only'."
)
_validate_department_reference(person, item_path, department_ids)
@@ -131,6 +201,28 @@ def render_meeting_context_for_prompt(context: MeetingContext) -> str:
for participant in participants:
lines.append(_render_person_line(participant, "participant_id", departments_by_id))
speaker_mappings = context.speaker_mappings
if speaker_mappings:
participants_by_id = {
participant["participant_id"]: participant
for participant in participants
if isinstance(participant, dict) and participant.get("participant_id")
}
lines.extend(
[
"",
"Confirmed diarization speaker mappings (authoritative):",
"- Use only these explicit mappings. Never infer identities for other speaker labels.",
"- Unmapped SPEAKER_XX labels must remain anonymous.",
]
)
for speaker_label, participant_id in sorted(speaker_mappings.items()):
participant = participants_by_id[participant_id]
lines.append(
f"- {speaker_label}: {_text(participant.get('display_name'))} "
f"(participant_id: {participant_id})"
)
mentioned_people = _optional_list(data.get("mentioned_people"), "mentioned_people")
if mentioned_people:
lines.extend(["", "Mentioned but absent people:"])
@@ -245,12 +337,28 @@ def _collect_unique_ids(items: list[Any], key: str, path: str) -> set[str]:
return ids
def _validate_attendance(item: dict[str, Any], path: str) -> None:
status = item.get("attendance_status")
def _validate_attendance(item: dict[str, Any], path: str, *, default: str) -> str:
status = item.get("attendance_status", default)
if status not in VALID_ATTENDANCE_STATUSES:
raise MeetingContextValidationError(
f"{path}.attendance_status has invalid value: {status!r}."
)
return status
def _with_attendance_defaults(data: dict[str, Any]) -> dict[str, Any]:
normalized = copy.deepcopy(data)
participants = normalized.get("participants")
if isinstance(participants, list):
for participant in participants:
if isinstance(participant, dict):
participant.setdefault("attendance_status", "present")
mentioned_people = normalized.get("mentioned_people")
if isinstance(mentioned_people, list):
for person in mentioned_people:
if isinstance(person, dict):
person.setdefault("attendance_status", "mentioned_only")
return normalized
def _validate_department_reference(
+17
View File
@@ -0,0 +1,17 @@
"""Reusable orchestration APIs for Meeting Lab applications and CLIs."""
from src.meeting_lab.orchestration.mvp import (
DEFAULT_OUTPUT_ROOT,
MvpMeetingConfig,
MvpRunResult,
create_unique_run_dir,
run_mvp_meeting,
)
__all__ = [
"DEFAULT_OUTPUT_ROOT",
"MvpMeetingConfig",
"MvpRunResult",
"create_unique_run_dir",
"run_mvp_meeting",
]
+467
View File
@@ -0,0 +1,467 @@
"""Reusable audio-to-direct-protocol MVP orchestration."""
from __future__ import annotations
import json
import re
import shutil
import sys
import time
from collections.abc import Callable, Mapping, Sequence
from dataclasses import dataclass
from datetime import datetime
from pathlib import Path
from typing import Any
from src.meeting_lab.audio import prepare_audio
from src.meeting_lab.diarization import (
DEFAULT_MODEL as DEFAULT_DIARIZATION_MODEL,
diarize_audio,
write_diarized_transcript,
)
from src.meeting_lab.llm.ollama import DEFAULT_ENDPOINT
from src.meeting_lab.models.meeting_context import (
MeetingContext,
create_meeting_context,
load_meeting_context,
validate_meeting_context,
write_meeting_context,
)
from src.meeting_lab.progress import ProgressEvent, ProgressSink, ProgressStatus
from src.meeting_lab.protocol.generate_direct_protocol import (
DEFAULT_MODEL,
DEFAULT_NUM_CTX,
DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
DirectProtocolResult,
generate_direct_protocol,
load_compact_transcript,
)
from src.meeting_lab.transcription.whisper import transcribe_audio
DEFAULT_OUTPUT_ROOT = Path("meeting_data/runs")
ContextInput = MeetingContext | Mapping[str, Any]
@dataclass(frozen=True)
class MvpMeetingConfig:
audio_file: Path
whisper_model: Path
whisper_executable: str = "whisper-cli"
ffmpeg_executable: str = "ffmpeg"
audio_normalization: bool = True
context_file: Path | None = None
output_root: Path = DEFAULT_OUTPUT_ROOT
language: str = "de"
threads: str | int = "auto"
model: str = DEFAULT_MODEL
ollama_endpoint: str = DEFAULT_ENDPOINT
protocol_num_ctx: int = DEFAULT_NUM_CTX
protocol_safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET
diarization: str = "off"
diarization_runtime: str = "native"
diarization_container_image: str | None = None
diarization_container_args: Sequence[str] = ()
@dataclass(frozen=True)
class MvpRunResult:
exit_code: int
run_dir: Path | None
protocol_path: Path | None
def regenerate_mvp_protocol(
run_dir: Path,
*,
meeting_context: ContextInput,
model: str = DEFAULT_MODEL,
ollama_endpoint: str = DEFAULT_ENDPOINT,
protocol_num_ctx: int = DEFAULT_NUM_CTX,
protocol_safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
progress_sink: ProgressSink | None = None,
) -> MvpRunResult:
"""Regenerate only protocol artifacts from an existing completed run."""
started = time.perf_counter()
run_dir = Path(run_dir)
context = _effective_context(meeting_context)
if context is None:
raise ValueError("Meeting Context is required for protocol regeneration.")
if protocol_num_ctx <= 0:
raise ValueError("Protocol Ollama context size must be positive.")
if protocol_safe_input_token_budget <= 0:
raise ValueError("Protocol safe input token budget must be positive.")
diarized_transcript = run_dir / "diarization" / "transcript_diarized.json"
plain_transcript = run_dir / "transcript" / "transcript.json"
transcript_path = (
diarized_transcript if diarized_transcript.is_file() else plain_transcript
)
if not transcript_path.is_file():
raise FileNotFoundError(
f"Existing run has no protocol transcript artifact: {run_dir}"
)
context_path = run_dir / "context" / "meeting_context.yaml"
write_meeting_context(context, context_path)
_emit(progress_sink, "protocol_generation", "started", started)
try:
result = generate_direct_protocol(
transcript_path,
context_path,
model=model,
endpoint=ollama_endpoint,
num_ctx=protocol_num_ctx,
safe_input_token_budget=protocol_safe_input_token_budget,
)
protocol_path = _persist_protocol(run_dir, result)
except Exception as exc:
_emit(
progress_sink,
"failed",
"failed",
started,
message=f"protocol_generation: {type(exc).__name__}: {exc}",
)
raise
_emit(progress_sink, "protocol_generation", "completed", started)
_emit(progress_sink, "completed", "completed", started)
return MvpRunResult(0, run_dir, protocol_path)
def create_unique_run_dir(
output_root: Path,
meeting_name: str,
now: Callable[[], datetime] = datetime.now,
) -> Path:
safe_name = re.sub(r"[^A-Za-z0-9_.-]+", "_", meeting_name).strip("._-") or "meeting"
base = output_root / f"{safe_name}_{now().strftime('%Y%m%d_%H%M%S')}"
candidate = base
suffix = 1
while candidate.exists():
candidate = output_root / f"{base.name}_{suffix:02d}"
suffix += 1
candidate.mkdir(parents=True)
return candidate
def _write_json(path: Path, value: Any) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def _effective_context(value: ContextInput | None) -> MeetingContext | None:
if value is None:
return None
if isinstance(value, MeetingContext):
validate_meeting_context(value.data)
return value
if isinstance(value, Mapping):
return create_meeting_context(dict(value), source_file=Path("<programmatic>"))
raise TypeError("meeting_context must be MeetingContext, mapping, or None.")
def _validate_inputs(
config: MvpMeetingConfig, meeting_context: MeetingContext | None
) -> None:
if not config.audio_file.is_file():
raise FileNotFoundError(f"Audio file does not exist: {config.audio_file}")
if not config.whisper_model.is_file():
raise FileNotFoundError(f"Whisper model does not exist: {config.whisper_model}")
if meeting_context is not None and config.context_file is not None:
raise ValueError("Use either a context file or a programmatic Meeting Context, not both.")
if meeting_context is not None:
validate_meeting_context(meeting_context.data)
elif config.context_file is not None:
if not config.context_file.is_file():
raise FileNotFoundError(
f"Meeting Context file does not exist: {config.context_file}"
)
load_meeting_context(config.context_file)
if config.diarization not in ("off", "auto", "gpu", "cpu"):
raise ValueError(f"Unsupported diarization mode: {config.diarization}")
if config.diarization_runtime not in ("native", "container"):
raise ValueError(
f"Unsupported diarization runtime: {config.diarization_runtime}"
)
if (
config.diarization != "off"
and config.diarization_runtime == "container"
and not config.diarization_container_image
):
raise ValueError("A diarization container image is required.")
if config.protocol_safe_input_token_budget <= 0:
raise ValueError("Protocol safe input token budget must be positive.")
if config.protocol_num_ctx <= 0:
raise ValueError("Protocol Ollama context size must be positive.")
def _emit(
sink: ProgressSink | None,
stage: str,
status: ProgressStatus,
overall_started: float,
*,
message: str | None = None,
) -> None:
if sink is not None:
sink(
ProgressEvent(
stage=stage,
status=status,
elapsed_seconds=time.perf_counter() - overall_started,
message=message,
)
)
def _persist_protocol(run_dir: Path, result: DirectProtocolResult) -> Path:
protocol_dir = run_dir / "protocol"
protocol_dir.mkdir(exist_ok=True)
(protocol_dir / "exact_prompt.txt").write_text(result.exact_prompt, encoding="utf-8")
_write_json(protocol_dir / "raw_response.json", result.raw_response)
_write_json(protocol_dir / "runtime_metadata.json", result.runtime_metadata)
transcript_input = getattr(result, "transcript_input", None)
if transcript_input is not None:
(protocol_dir / "transcript_input.txt").write_text(
transcript_input, encoding="utf-8"
)
protocol_path = run_dir / "protocol.md"
protocol_path.write_text(result.protocol_text, encoding="utf-8")
return protocol_path
def run_mvp_meeting(
config: MvpMeetingConfig,
*,
meeting_context: ContextInput | None = None,
progress_sink: ProgressSink | None = None,
) -> MvpRunResult:
"""Run the existing MVP directly, without subprocess or GUI dependencies."""
overall_started = time.perf_counter()
validation_started = time.perf_counter()
_emit(progress_sink, "preparing", "started", overall_started)
try:
effective_context = _effective_context(meeting_context)
_validate_inputs(config, effective_context)
except Exception as exc:
_emit(
progress_sink,
"failed",
"failed",
overall_started,
message=f"preparing: {type(exc).__name__}: {exc}",
)
print(f"Error: {type(exc).__name__}: {exc}", file=sys.stderr)
return MvpRunResult(2, None, None)
validation_runtime = time.perf_counter() - validation_started
run_dir = create_unique_run_dir(config.output_root, config.audio_file.stem)
timestamp = datetime.now().astimezone().isoformat(timespec="seconds")
transcript_path = run_dir / "transcript" / "transcript.json"
protocol_path = run_dir / "protocol.md"
stage_runtimes: dict[str, float | None] = {
"validation": round(validation_runtime, 3),
"setup": None,
"audio_preparation": None,
"whisper": None,
"transcript_validation": None,
"protocol": None,
}
if config.diarization != "off":
stage_runtimes["diarization"] = None
stage_runtimes["diarization_alignment"] = None
metadata: dict[str, Any] = {
"run_id": run_dir.name,
"timestamp": timestamp,
"input_audio": str(config.audio_file.resolve()),
"audio_preparation": None,
"transcript_output": str(transcript_path.resolve()),
"protocol_output": str(protocol_path.resolve()),
"whisper_model": str(config.whisper_model.resolve()),
"model": config.model,
"ollama_endpoint": config.ollama_endpoint,
"status": "running",
"stage_runtimes_seconds": stage_runtimes,
"total_runtime_seconds": None,
"failure": None,
"diarization": {
"enabled": config.diarization != "off",
"backend": "pyannote.audio" if config.diarization != "off" else None,
"model": DEFAULT_DIARIZATION_MODEL if config.diarization != "off" else None,
"requested_device_mode": config.diarization,
"runtime": config.diarization_runtime if config.diarization != "off" else None,
"metadata_path": None,
"transcript_diarized": None,
},
}
current_stage = "preparing"
stage_started = time.perf_counter()
try:
audio_dir = run_dir / "audio"
transcript_dir = run_dir / "transcript"
context_dir = run_dir / "context"
protocol_dir = run_dir / "protocol"
audio_dir.mkdir()
transcript_dir.mkdir()
context_dir.mkdir()
protocol_dir.mkdir()
_write_json(
audio_dir / "input_manifest.json",
{
"source_file": str(config.audio_file.resolve()),
"filename": config.audio_file.name,
"size_bytes": config.audio_file.stat().st_size,
},
)
preparation_started = time.perf_counter()
current_stage = "audio_preparation"
stage_started = preparation_started
prepared_audio = prepare_audio(
config.audio_file,
audio_dir / "prepared.wav",
ffmpeg_executable=config.ffmpeg_executable,
normalization_enabled=config.audio_normalization,
)
stage_runtimes["audio_preparation"] = round(
time.perf_counter() - preparation_started, 3
)
current_stage = "preparing"
preparation_metadata = prepared_audio.metadata()
metadata["audio_preparation"] = preparation_metadata
_write_json(audio_dir / "preparation_metadata.json", preparation_metadata)
_write_json(
audio_dir / "input_manifest.json",
{
"source_file": str(config.audio_file.resolve()),
"filename": config.audio_file.name,
"size_bytes": config.audio_file.stat().st_size,
"format": config.audio_file.suffix.lower().removeprefix("."),
"prepared_audio": preparation_metadata,
},
)
preserved_context: Path | None = None
if effective_context is not None:
preserved_context = context_dir / "meeting_context.yaml"
write_meeting_context(effective_context, preserved_context)
elif config.context_file is not None:
preserved_context = context_dir / "meeting_context.yaml"
shutil.copy2(config.context_file, preserved_context)
stage_runtimes["setup"] = round(time.perf_counter() - stage_started, 3)
_emit(progress_sink, "preparing", "completed", overall_started)
current_stage = "transcription"
stage_started = time.perf_counter()
_emit(progress_sink, "transcription", "started", overall_started)
transcription = transcribe_audio(
prepared_audio.prepared_path,
config.whisper_model,
transcript_dir,
config.language,
executable=config.whisper_executable,
threads=config.threads,
)
stage_runtimes["whisper"] = round(time.perf_counter() - stage_started, 3)
_emit(progress_sink, "transcription", "completed", overall_started)
stage_started = time.perf_counter()
load_compact_transcript(transcription.transcript_json)
stage_runtimes["transcript_validation"] = round(
time.perf_counter() - stage_started, 3
)
protocol_transcript = transcription.transcript_json
if config.diarization != "off":
current_stage = "diarization"
stage_started = time.perf_counter()
_emit(progress_sink, "diarization", "started", overall_started)
diarization_dir = run_dir / "diarization"
diarization = diarize_audio(
prepared_audio.prepared_path,
diarization_dir,
config.diarization,
runtime=config.diarization_runtime,
container_image=config.diarization_container_image,
container_args=config.diarization_container_args,
)
stage_runtimes["diarization"] = round(
time.perf_counter() - stage_started, 3
)
metadata["diarization"].update(
{
"actual_device": diarization.metadata.get("actual_device"),
"device_name": diarization.metadata.get("device_name"),
"runtime_seconds": diarization.metadata.get("runtime_seconds"),
"speaker_count": diarization.metadata.get("speaker_count"),
"metadata_path": str(diarization.metadata_path.resolve()),
}
)
stage_started = time.perf_counter()
protocol_transcript, diarized_text = write_diarized_transcript(
transcription.transcript_json,
diarization.exclusive_turns_json,
diarization_dir,
)
load_compact_transcript(protocol_transcript)
stage_runtimes["diarization_alignment"] = round(
time.perf_counter() - stage_started, 3
)
metadata["diarization"].update(
{
"transcript_diarized": str(protocol_transcript.resolve()),
"transcript_diarized_text": str(diarized_text.resolve()),
}
)
_emit(progress_sink, "diarization", "completed", overall_started)
current_stage = "protocol_generation"
stage_started = time.perf_counter()
_emit(progress_sink, "protocol_generation", "started", overall_started)
result = generate_direct_protocol(
protocol_transcript,
preserved_context,
model=config.model,
endpoint=config.ollama_endpoint,
num_ctx=config.protocol_num_ctx,
safe_input_token_budget=config.protocol_safe_input_token_budget,
)
stage_runtimes["protocol"] = round(time.perf_counter() - stage_started, 3)
protocol_path = _persist_protocol(run_dir, result)
_emit(progress_sink, "protocol_generation", "completed", overall_started)
metadata["status"] = "completed"
_emit(progress_sink, "completed", "completed", overall_started)
except Exception as exc:
metadata_stage = {
"preparing": "setup",
"audio_preparation": "audio_preparation",
"transcription": "whisper",
"diarization": "diarization",
"protocol_generation": "protocol",
}.get(current_stage, current_stage)
runtime_key = metadata_stage
if runtime_key in stage_runtimes and stage_runtimes[runtime_key] is None:
stage_runtimes[runtime_key] = round(time.perf_counter() - stage_started, 3)
metadata["status"] = "failed"
metadata["failure"] = {
"stage": metadata_stage,
"type": type(exc).__name__,
"message": str(exc),
}
protocol_path = None
_emit(
progress_sink,
"failed",
"failed",
overall_started,
message=f"{current_stage}: {type(exc).__name__}: {exc}",
)
print(f"Error: {type(exc).__name__}: {exc}", file=sys.stderr)
finally:
metadata["total_runtime_seconds"] = round(time.perf_counter() - overall_started, 3)
_write_json(run_dir / "run_metadata.json", metadata)
exit_code = 0 if metadata["status"] == "completed" else 2
return MvpRunResult(exit_code, run_dir, protocol_path)
+21
View File
@@ -0,0 +1,21 @@
"""Small observer boundary for long-running Meeting Lab operations."""
from __future__ import annotations
from dataclasses import dataclass
from typing import Callable, Literal
ProgressStatus = Literal["started", "completed", "failed"]
@dataclass(frozen=True)
class ProgressEvent:
stage: str
status: ProgressStatus
elapsed_seconds: float
progress: float | None = None
message: str | None = None
ProgressSink = Callable[[ProgressEvent], None]
@@ -0,0 +1,36 @@
"""Prompt construction for the direct transcript-to-protocol MVP."""
from __future__ import annotations
DIRECT_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und dem Meeting-Kontext ein vollständiges, strukturiertes und professionelles internes Besprechungsprotokoll.
Das Protokoll muss themenorientiert sein, nicht chronologisch und nicht nach technischen Kategorien gegliedert. Beginne mit # Meeting Protocol. Verwende für jedes kohärente Thema eine Überschrift ## <Thema> und darunter eine strukturierte Synthese der Diskussion. Bewahre relevante Diskussionsverläufe, unterschiedliche Positionen, offene Punkte und Entscheidungsgrundlagen. Dokumentiere die wesentlichen Inhalte nachvollziehbar und fasse Themenblöcke so zusammen, dass auch Personen, die nicht am Meeting teilgenommen haben, den Kontext und die Entwicklung der Diskussion verstehen können. Nenne Entscheidungen oder abgestimmte Positionen nur, wenn sie tatsächlich belegt sind. Führe Maßnahmen nur auf, wenn eine konkrete zukünftige Handlung gestützt ist; nenne verantwortliche Personen und Fristen ausschließlich bei expliziter Zuweisung, Annahme oder Bestätigung im Transkript. Vorschläge, Einwände, Möglichkeiten und vorläufige Ideen sind keine Entscheidungen oder Verpflichtungen. Bewahre relevante Einschränkungen und ungelöste Meinungsverschiedenheiten. Nenne offene Punkte nur, wenn sie wirklich offen bleiben. Nicht jedes Thema benötigt Entscheidungen, Maßnahmen oder offene Punkte.
Erzeuge keine reine Wiedergabe des Transkripts und verlängere das Protokoll nicht unnötig durch Wiederholungen. Synthetisiere zusammengehörige Aussagen, entferne Füllwörter und Gesprächsrauschen und erfinde keine Fakten, Entscheidungen, Zustimmungen, Verantwortlichen oder Fristen. Gib kein JSON, keine internen Labels und keine Analyse oder Denkprotokolle aus. Das Ergebnis soll als Markdown-Protokoll nach geringfügiger menschlicher Redaktion intern versendbar sein. Eine kompakte themenübergreifende Maßnahmenliste am Ende ist optional, wenn sie nützlich und vollständig belegt ist."""
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION = """Erstelle aus dem vollständigen Transkript und Meeting-Kontext ein vollständiges, professionelles internes Besprechungsprotokoll. Das Transkript ist in aufeinanderfolgende anonyme Sprecherblöcke gegliedert.
Beginne mit # Meeting Protocol. Gliedere themenorientiert mit ## <Thema> und synthetisiere je Thema den relevanten Diskussionsverlauf, Kontext, unterschiedliche Positionen, Entscheidungsgrundlagen, Einschränkungen und ungelöste Meinungsverschiedenheiten so, dass Dritte ihn nachvollziehen können. Nenne Entscheidungen nur bei Beleg. Nenne Maßnahmen, Verantwortliche und Fristen nur bei expliziter Zuweisung, Annahme oder Bestätigung; Vorschläge sind keine Verpflichtungen.
Entferne nur Wiederholungen, Füllwörter und Gesprächsrauschen. Erfinde keine Fakten oder Identitäten. Gib kein JSON, keine Sprecherlabels und kein Denkprotokoll aus. Eine belegte themenübergreifende Maßnahmenliste am Ende ist optional."""
MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION = """Nutze die autoritativen SPEAKER_XX-zu-Teilnehmer-Zuordnungen im Meeting-Kontext, um ausdrücklich belegte Aussagen, Positionen, Entscheidungen, Zuweisungen und angenommene persönliche Verpflichtungen namentlich zuzuordnen. Eine ausdrückliche Ich-Zusage eines zugeordneten Sprechers belegt persönliche Verantwortung. Unterscheide stets den Sprecher einer Aussage von darin nur erwähnten Personen. Leite für nicht zugeordnete Sprecher keine Identität ab und erfinde keine persönliche Verantwortung. Gib die technischen SPEAKER_XX-Bezeichnungen nicht im nutzerseitigen Protokoll aus."""
def build_direct_protocol_prompt(
transcript: str,
meeting_context: str | None = None,
*,
instruction: str = DIRECT_PROTOCOL_INSTRUCTION,
meeting_language: str = "de",
) -> str:
context = meeting_context.strip() if meeting_context else "Kein Meeting-Kontext bereitgestellt."
language = {"de": "German", "en": "English"}.get(meeting_language, meeting_language)
return (
f"{instruction}\n\n"
f"Write the meeting protocol in {language}. "
"Preserve speaker and person names exactly as supplied.\n\n"
f"MEETING-KONTEXT:\n{context}\n\n"
f"VOLLSTAENDIGES TRANSKRIPT:\n{transcript.strip()}\n"
)
@@ -0,0 +1,246 @@
"""One-call direct protocol generation from a compact Whisper transcript."""
from __future__ import annotations
import json
from dataclasses import dataclass
from pathlib import Path
from typing import Any, Callable
from src.meeting_lab.llm.ollama import (
DEFAULT_ENDPOINT,
OllamaGeneration,
generate_once,
require_model,
)
from src.meeting_lab.models.meeting_context import (
MeetingContext,
load_meeting_context,
render_meeting_context_for_prompt,
)
from src.meeting_lab.protocol.direct_protocol_prompt import (
COMPACT_DIARIZED_PROTOCOL_INSTRUCTION,
MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION,
build_direct_protocol_prompt,
)
from src.meeting_lab.protocol.transcript_input import (
TranscriptInputError,
compact_diarized_transcript,
plain_segment_transcript,
)
DEFAULT_MODEL = "qwen3.6:35B-A3B"
DEFAULT_NUM_CTX = 32768
DEFAULT_NUM_PREDICT = 8192
DEFAULT_TIMEOUT = 1800
DEFAULT_SAFE_INPUT_TOKEN_BUDGET = 29_000
ESTIMATED_UTF8_BYTES_PER_TOKEN = 4.4
class DirectProtocolError(ValueError):
"""Raised for invalid direct-protocol inputs or model output."""
@dataclass(frozen=True)
class DirectProtocolResult:
protocol_text: str
exact_prompt: str
model_metadata: dict[str, Any]
runtime_metadata: dict[str, Any]
raw_response: dict[str, Any]
transcript_input: str | None = None
@dataclass(frozen=True)
class SelectedTranscriptInput:
text: str
prompt: str
representation: str
estimated_input_tokens: int
safe_input_token_budget: int
fallback_used: bool
diarization_enabled: bool
def load_compact_transcript(path: Path) -> str:
data = _load_transcript_document(path)
text = data.get("text")
if not isinstance(text, str) or not text.strip():
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
return text
def _load_transcript_document(path: Path) -> dict[str, Any]:
if not path.is_file():
raise DirectProtocolError(f"Transcript file does not exist: {path}")
try:
data = json.loads(path.read_text(encoding="utf-8-sig"))
except json.JSONDecodeError as exc:
raise DirectProtocolError(f"Transcript is not valid JSON: {path}: {exc}") from exc
if not isinstance(data, dict):
raise DirectProtocolError("Transcript JSON must contain a top-level object.")
if "text" not in data:
raise DirectProtocolError("Transcript JSON must contain top-level 'text'.")
return data
def estimate_input_tokens(prompt: str) -> int:
"""Estimate tokens without adding a model-specific tokenizer dependency."""
byte_count = len(prompt.encode("utf-8"))
return max(1, int(byte_count / ESTIMATED_UTF8_BYTES_PER_TOKEN + 0.999999))
def select_transcript_input(
transcript: dict[str, Any],
rendered_context: str | None,
*,
safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
meeting_language: str = "de",
) -> SelectedTranscriptInput:
"""Select complete prompt input without allowing silent tail truncation."""
if safe_input_token_budget <= 0:
raise DirectProtocolError("Safe protocol input token budget must be positive.")
diarization_enabled = transcript.get("speaker_labels_anonymous") is True
if diarization_enabled:
try:
compact = compact_diarized_transcript(transcript.get("segments"))
plain_text = plain_segment_transcript(transcript.get("segments"))
except TranscriptInputError as exc:
raise DirectProtocolError(str(exc)) from exc
instruction = COMPACT_DIARIZED_PROTOCOL_INSTRUCTION
if (
rendered_context
and "Confirmed diarization speaker mappings" in rendered_context
):
instruction = f"{instruction}\n\n{MAPPED_SPEAKER_ATTRIBUTION_INSTRUCTION}"
compact_prompt = build_direct_protocol_prompt(
compact.text,
rendered_context,
instruction=instruction,
meeting_language=meeting_language,
)
compact_estimate = estimate_input_tokens(compact_prompt)
if compact_estimate <= safe_input_token_budget:
return SelectedTranscriptInput(
text=compact.text,
prompt=compact_prompt,
representation="diarized_compact",
estimated_input_tokens=compact_estimate,
safe_input_token_budget=safe_input_token_budget,
fallback_used=False,
diarization_enabled=True,
)
representation = "plain_transcript_fallback"
fallback_used = True
else:
plain_text = transcript.get("text")
if not isinstance(plain_text, str) or not plain_text.strip():
raise DirectProtocolError("Transcript top-level 'text' must be a non-empty string.")
representation = "plain_transcript"
fallback_used = False
plain_prompt = build_direct_protocol_prompt(
plain_text, rendered_context, meeting_language=meeting_language
)
plain_estimate = estimate_input_tokens(plain_prompt)
if plain_estimate > safe_input_token_budget:
raise DirectProtocolError(
"Protocol prompt/input is too large for the configured safe input budget "
f"({plain_estimate} estimated tokens > {safe_input_token_budget}). "
"No LLM request was made; silent truncation is not allowed."
)
return SelectedTranscriptInput(
text=plain_text,
prompt=plain_prompt,
representation=representation,
estimated_input_tokens=plain_estimate,
safe_input_token_budget=safe_input_token_budget,
fallback_used=fallback_used,
diarization_enabled=diarization_enabled,
)
def generate_direct_protocol(
transcript_path: Path,
context_path: Path | None = None,
*,
model: str = DEFAULT_MODEL,
endpoint: str = DEFAULT_ENDPOINT,
timeout: int = DEFAULT_TIMEOUT,
num_ctx: int = DEFAULT_NUM_CTX,
num_predict: int = DEFAULT_NUM_PREDICT,
safe_input_token_budget: int = DEFAULT_SAFE_INPUT_TOKEN_BUDGET,
model_check: Callable[[str, str, int], dict[str, Any]] = require_model,
generation_call: Callable[..., OllamaGeneration] = generate_once,
) -> DirectProtocolResult:
transcript = _load_transcript_document(transcript_path)
context: MeetingContext | None = (
load_meeting_context(context_path) if context_path is not None else None
)
rendered_context = render_meeting_context_for_prompt(context) if context else None
# Older runs without a meeting language retain the historical German output.
meeting_language = (
str(context.data["meeting"].get("language") or "de") if context else "de"
)
selected = select_transcript_input(
transcript,
rendered_context,
safe_input_token_budget=safe_input_token_budget,
meeting_language=meeting_language,
)
model_metadata = model_check(endpoint, model, 10)
generation = generation_call(
endpoint,
model,
selected.prompt,
timeout=timeout,
num_ctx=num_ctx,
num_predict=num_predict,
)
data = generation.raw_response
runtime_metadata = {
"output_language": meeting_language,
"model": model,
"prompt_token_count": data.get("prompt_eval_count"),
"output_token_count": data.get("eval_count"),
"prompt_evaluation_duration_ns": data.get("prompt_eval_duration"),
"generation_duration_ns": data.get("eval_duration"),
"total_ollama_duration_ns": data.get("total_duration"),
"client_wall_time_seconds": generation.client_wall_time_seconds,
"completion_reason": data.get("done_reason"),
"done": data.get("done"),
"request_count": 1,
"temperature": 0.0,
"think": False,
"num_ctx": num_ctx,
"num_predict": num_predict,
"selected_transcript_representation": selected.representation,
"estimated_input_tokens": selected.estimated_input_tokens,
"safe_input_token_budget": selected.safe_input_token_budget,
"input_token_estimation_method": "utf8_bytes_divided_by_4.4",
"fallback_used": selected.fallback_used,
"diarization_enabled": selected.diarization_enabled,
"speaker_attribution_available": (
True
if selected.representation == "diarized_compact"
else False
if selected.representation == "plain_transcript_fallback"
else None
),
"speaker_attribution_loss_reason": (
"plain_transcript_fallback"
if selected.representation == "plain_transcript_fallback"
else None
),
"speaker_mapping_count": len(context.speaker_mappings) if context else 0,
}
return DirectProtocolResult(
protocol_text=generation.text,
exact_prompt=selected.prompt,
model_metadata=model_metadata,
runtime_metadata=runtime_metadata,
raw_response=data,
transcript_input=selected.text,
)
@@ -0,0 +1,97 @@
"""Deterministic transcript representations for one-call protocol prompts."""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any
class TranscriptInputError(ValueError):
"""Raised when a transcript cannot be represented without content loss."""
@dataclass(frozen=True)
class SpeakerBlock:
"""One contiguous run of transcript segments assigned to one speaker."""
speaker_id: str
segment_texts: tuple[str, ...]
@dataclass(frozen=True)
class CompactDiarizedTranscript:
"""Compact prompt text plus structural evidence of segment preservation."""
text: str
blocks: tuple[SpeakerBlock, ...]
source_segment_count: int
@property
def represented_segment_count(self) -> int:
return sum(len(block.segment_texts) for block in self.blocks)
@property
def segment_texts(self) -> tuple[str, ...]:
return tuple(text for block in self.blocks for text in block.segment_texts)
def normalize_segment_text(value: Any, index: int) -> str:
"""Normalize formatting whitespace while retaining all semantic text."""
if not isinstance(value, str):
raise TranscriptInputError(f"Transcript segment {index} text must be a string.")
return " ".join(value.split())
def compact_diarized_transcript(segments: Any) -> CompactDiarizedTranscript:
"""Group only adjacent same-speaker segments and omit repeated timestamps."""
if not isinstance(segments, list) or not segments:
raise TranscriptInputError(
"Diarized transcript must contain a non-empty 'segments' list."
)
mutable_blocks: list[tuple[str, list[str]]] = []
source_texts: list[str] = []
for index, segment in enumerate(segments):
if not isinstance(segment, dict):
raise TranscriptInputError(f"Transcript segment {index} must be an object.")
speaker = segment.get("speaker_id") or "SPEAKER_UNASSIGNED"
if not isinstance(speaker, str) or not speaker.startswith("SPEAKER_"):
raise TranscriptInputError(
f"Transcript segment {index} must use an anonymous SPEAKER_ label."
)
text = normalize_segment_text(segment.get("text"), index)
source_texts.append(text)
if mutable_blocks and mutable_blocks[-1][0] == speaker:
mutable_blocks[-1][1].append(text)
else:
mutable_blocks.append((speaker, [text]))
blocks = tuple(
SpeakerBlock(speaker_id=speaker, segment_texts=tuple(texts))
for speaker, texts in mutable_blocks
)
rendered = "\n".join(
f"{block.speaker_id}: {' '.join(block.segment_texts)}" for block in blocks
)
result = CompactDiarizedTranscript(
text=rendered + "\n",
blocks=blocks,
source_segment_count=len(segments),
)
if result.represented_segment_count != len(segments):
raise TranscriptInputError("Compact diarized transcript lost source segments.")
if result.segment_texts != tuple(source_texts):
raise TranscriptInputError("Compact diarized transcript changed segment order or text.")
return result
def plain_segment_transcript(segments: Any) -> str:
"""Reconstruct plain transcript text from every segment in source order."""
if not isinstance(segments, list) or not segments:
raise TranscriptInputError("Transcript must contain a non-empty 'segments' list.")
texts = []
for index, segment in enumerate(segments):
if not isinstance(segment, dict):
raise TranscriptInputError(f"Transcript segment {index} must be an object.")
texts.append(normalize_segment_text(segment.get("text"), index))
return " ".join(texts)
@@ -0,0 +1 @@
"""Isolated experimental semantic synthesis for known discussion subjects."""
@@ -0,0 +1,640 @@
#!/usr/bin/env python3
"""Run semantic synthesis with subject detection and evidence assignment fixed."""
from __future__ import annotations
import argparse
import json
import time
from pathlib import Path
from typing import Any
import requests
SCHEMA_VERSION = "experimental-semantic-synthesis-v1"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
DEFAULT_MODEL = "qwen3.5:9B"
DEFAULT_TIMEOUT = 300
DEFAULT_NUM_CTX = 8192
DEFAULT_NUM_PREDICT = 2048
EVENT_TYPES = {
"idea",
"option",
"proposal",
"objection",
"supporting_argument",
"clarification",
"rejection",
"scoped_acceptance",
"fact",
"technical_finding",
}
OUTCOME_STATUSES = {"established", "rejected", "scoped_acceptance", "tentative"}
class SynthesisValidationError(ValueError):
"""Raised when isolated semantic synthesis output is structurally invalid."""
PROMPT_TEMPLATE = """You perform semantic synthesis for one already known discussion subject.
The subject boundary and evidence assignment are fixed and complete. Do not discover,
split, merge, rename, or omit the subject. Do not assign evidence to another subject.
Interpret only what the supplied evidence semantically establishes.
Semantic distinctions:
- idea: mentioned possibility without stronger commitment
- option: alternative considered without commitment
- proposal: suggested course of action not yet established as work
- objection: argument or concern against something; not automatically unresolved
- rejection: an alternative is explicitly rejected
- scoped_acceptance: accepted only for the stated test, trial, condition, or scope
- proposal is not an action
- no decision is not a tentative decision
- mention is not an unresolved issue
- an action requires explicit assignment, acceptance, commitment, or established work
- an unresolved issue requires a concrete need explicitly left unresolved
Preserve explicit rejection, explicit accepted work, explicit unresolved questions,
and all limits on an outcome. Never generalize trial acceptance into final acceptance.
Use only supplied evidence IDs. Keep concise semantic text in the evidence language.
Return exactly one JSON object. Always include these fields:
{{
"schema_version": "experimental-semantic-synthesis-v1",
"subject_id": "copy the supplied subject_id exactly",
"subject": "copy the supplied subject exactly",
"events": [
{{
"type": "idea|option|proposal|objection|supporting_argument|clarification|rejection|scoped_acceptance|fact|technical_finding",
"text": "supported semantic event",
"evidence_ids": ["e1"]
}}
],
"actions": [
{{
"text": "established action",
"responsible": null,
"due": null,
"evidence_ids": ["e2"]
}}
],
"unresolved_issues": [
{{
"text": "explicitly unresolved issue",
"evidence_ids": ["e3"]
}}
]
}}
The three arrays are structurally required; use [] when none exist.
Add "outcome" only when an outcome was actually established:
{{
"status": "established|rejected|scoped_acceptance|tentative",
"text": "what was actually established",
"scope": "the exact scope, condition, or limit",
"evidence_ids": ["e2"]
}}
Omit outcome completely when there is none. Never use null for outcome. Never use the
string "null"; use JSON null only for unknown responsible or due values.
Fixed Gold input:
{input_json}
"""
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Run the isolated semantic-synthesis Gold experiment."
)
parser.add_argument("fixture", type=Path, help="Fixed-subject Gold bundle JSON.")
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT)
parser.add_argument("--num-ctx", type=int, default=DEFAULT_NUM_CTX)
parser.add_argument("--num-predict", type=int, default=DEFAULT_NUM_PREDICT)
return parser.parse_args()
def _exact_keys(
value: dict[str, Any], required: set[str], optional: set[str], location: str
) -> None:
missing = required - value.keys()
unknown = value.keys() - required - optional
if missing:
raise SynthesisValidationError(
f"{location} missing required keys: {sorted(missing)}"
)
if unknown:
raise SynthesisValidationError(
f"{location} has unknown keys: {sorted(unknown)}"
)
def _text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise SynthesisValidationError(f"{location} must be a non-empty string")
return value.strip()
def validate_bundle(case: Any) -> dict[str, Any]:
if not isinstance(case, dict):
raise SynthesisValidationError("case must be an object")
_exact_keys(
case,
{
"case_id",
"description",
"subject_id",
"subject",
"evidence",
"allowed_responsible",
"expected",
},
set(),
"case",
)
_text(case["case_id"], "case.case_id")
_text(case["description"], "case.description")
_text(case["subject_id"], "case.subject_id")
_text(case["subject"], "case.subject")
evidence = case["evidence"]
if not isinstance(evidence, list) or not evidence:
raise SynthesisValidationError("case.evidence must be a non-empty list")
seen: set[str] = set()
for index, item in enumerate(evidence):
location = f"case.evidence[{index}]"
if not isinstance(item, dict):
raise SynthesisValidationError(f"{location} must be an object")
_exact_keys(item, {"evidence_id", "text"}, set(), location)
evidence_id = _text(item["evidence_id"], f"{location}.evidence_id")
if evidence_id in seen:
raise SynthesisValidationError(f"duplicate evidence ID: {evidence_id}")
seen.add(evidence_id)
_text(item["text"], f"{location}.text")
allowed = case["allowed_responsible"]
if not isinstance(allowed, list) or any(
not isinstance(value, str) or not value.strip() for value in allowed
):
raise SynthesisValidationError(
"case.allowed_responsible must be a list of non-empty strings"
)
if len(set(allowed)) != len(allowed):
raise SynthesisValidationError("case.allowed_responsible contains duplicates")
if not isinstance(case["expected"], dict):
raise SynthesisValidationError("case.expected must be an object")
return case
def _evidence_ids(value: Any, location: str, known: set[str]) -> list[str]:
if not isinstance(value, list) or not value:
raise SynthesisValidationError(f"{location} must be a non-empty list")
result: list[str] = []
for index, evidence_id in enumerate(value):
evidence_id = _text(evidence_id, f"{location}[{index}]")
if evidence_id not in known:
raise SynthesisValidationError(
f"{location}[{index}] references unknown evidence ID: {evidence_id}"
)
if evidence_id in result:
raise SynthesisValidationError(
f"{location} contains duplicate evidence ID: {evidence_id}"
)
result.append(evidence_id)
return result
def _nullable_text(value: Any, location: str) -> str | None:
if value is None:
return None
result = _text(value, location)
if result.casefold() == "null":
raise SynthesisValidationError(
f"{location} must use JSON null, not the string 'null'"
)
return result
def validate_synthesis(data: Any, case: dict[str, Any]) -> dict[str, Any]:
validate_bundle(case)
if not isinstance(data, dict):
raise SynthesisValidationError("output must be an object")
_exact_keys(
data,
{
"schema_version",
"subject_id",
"subject",
"events",
"actions",
"unresolved_issues",
},
{"outcome"},
"output",
)
if data["schema_version"] != SCHEMA_VERSION:
raise SynthesisValidationError(f"schema_version must be {SCHEMA_VERSION!r}")
if data["subject_id"] != case["subject_id"]:
raise SynthesisValidationError("model changed fixed subject_id")
if data["subject"] != case["subject"]:
raise SynthesisValidationError("model changed fixed subject")
known = {item["evidence_id"] for item in case["evidence"]}
events = data["events"]
if not isinstance(events, list):
raise SynthesisValidationError("output.events must be an array")
for index, event in enumerate(events):
location = f"output.events[{index}]"
if not isinstance(event, dict):
raise SynthesisValidationError(f"{location} must be an object")
_exact_keys(event, {"type", "text", "evidence_ids"}, set(), location)
if event["type"] not in EVENT_TYPES:
raise SynthesisValidationError(f"{location}.type is invalid")
_text(event["text"], f"{location}.text")
_evidence_ids(event["evidence_ids"], f"{location}.evidence_ids", known)
if "outcome" in data:
outcome = data["outcome"]
if not isinstance(outcome, dict):
raise SynthesisValidationError(
"output.outcome must be an object when present; omit it when absent"
)
_exact_keys(
outcome, {"status", "text", "scope", "evidence_ids"}, set(), "output.outcome"
)
if outcome["status"] not in OUTCOME_STATUSES:
raise SynthesisValidationError("output.outcome.status is invalid")
_text(outcome["text"], "output.outcome.text")
_text(outcome["scope"], "output.outcome.scope")
_evidence_ids(outcome["evidence_ids"], "output.outcome.evidence_ids", known)
actions = data["actions"]
if not isinstance(actions, list):
raise SynthesisValidationError("output.actions must be an array")
allowed = set(case["allowed_responsible"])
for index, action in enumerate(actions):
location = f"output.actions[{index}]"
if not isinstance(action, dict):
raise SynthesisValidationError(f"{location} must be an object")
_exact_keys(
action,
{"text", "responsible", "due", "evidence_ids"},
set(),
location,
)
_text(action["text"], f"{location}.text")
responsible = _nullable_text(action["responsible"], f"{location}.responsible")
if responsible is not None and responsible not in allowed:
raise SynthesisValidationError(
f"{location}.responsible is not allowed: {responsible}"
)
_nullable_text(action["due"], f"{location}.due")
_evidence_ids(action["evidence_ids"], f"{location}.evidence_ids", known)
issues = data["unresolved_issues"]
if not isinstance(issues, list):
raise SynthesisValidationError("output.unresolved_issues must be an array")
for index, issue in enumerate(issues):
location = f"output.unresolved_issues[{index}]"
if not isinstance(issue, dict):
raise SynthesisValidationError(f"{location} must be an object")
_exact_keys(issue, {"text", "evidence_ids"}, set(), location)
_text(issue["text"], f"{location}.text")
_evidence_ids(issue["evidence_ids"], f"{location}.evidence_ids", known)
return data
def build_prompt(case: dict[str, Any]) -> str:
validate_bundle(case)
model_input = {
"subject_id": case["subject_id"],
"subject": case["subject"],
"evidence": case["evidence"],
}
return PROMPT_TEMPLATE.format(
input_json=json.dumps(model_input, ensure_ascii=False, indent=2)
)
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise SynthesisValidationError("model response JSON must be an object")
return data
def build_ollama_payload(
model: str, prompt: str, num_ctx: int, num_predict: int
) -> dict[str, Any]:
return {
"model": model,
"prompt": prompt,
"think": False,
"stream": False,
"format": "json",
"options": {
"temperature": 0,
"num_ctx": num_ctx,
"num_predict": num_predict,
},
}
def call_ollama(
endpoint: str,
model: str,
prompt: str,
timeout: int,
num_ctx: int,
num_predict: int,
) -> tuple[str, dict[str, Any]]:
payload = build_ollama_payload(model, prompt, num_ctx, num_predict)
started = time.perf_counter()
response = requests.post(endpoint, json=payload, timeout=timeout)
elapsed = time.perf_counter() - started
response.raise_for_status()
body = response.json()
if not isinstance(body, dict):
raise ValueError("Ollama response must be an object")
raw_text = body.get("response")
if not isinstance(raw_text, str) or not raw_text.strip():
raise ValueError("Ollama returned no usable response text")
metadata = {
"model": body.get("model", model),
"elapsed_seconds": round(elapsed, 3),
"total_duration_ns": body.get("total_duration"),
"load_duration_ns": body.get("load_duration"),
"prompt_eval_count": body.get("prompt_eval_count"),
"prompt_eval_duration_ns": body.get("prompt_eval_duration"),
"eval_count": body.get("eval_count"),
"eval_duration_ns": body.get("eval_duration"),
"configuration": {
"temperature": 0,
"think": False,
"num_ctx": num_ctx,
"num_predict": num_predict,
},
}
return raw_text.strip(), metadata
def _contains(text: str, terms: list[str]) -> bool:
folded = text.casefold()
return any(term.casefold() in folded for term in terms)
def _refs_cover(items: list[dict[str, Any]], expected: list[str]) -> bool:
actual = {
evidence_id
for item in items
for evidence_id in item.get("evidence_ids", [])
}
return set(expected).issubset(actual)
def evaluate_synthesis(data: dict[str, Any], expected: dict[str, Any]) -> dict[str, Any]:
checks: list[dict[str, Any]] = []
def add(name: str, passed: bool, critical: bool = False) -> None:
checks.append({"name": name, "passed": passed, "critical": critical})
events = data["events"]
event_types = [item["type"] for item in events]
for event_type, minimum in expected.get("event_type_minimums", {}).items():
add(f"event:{event_type}", event_types.count(event_type) >= minimum)
allowed_types = set(expected.get("allowed_event_types", EVENT_TYPES))
add("no_unexpected_event_types", set(event_types).issubset(allowed_types))
add(
"event_evidence",
_refs_cover(events, expected.get("event_evidence_ids", [])),
)
outcome_expected = expected["outcome"]
outcome = data.get("outcome")
add(
"outcome_presence",
(outcome is not None) == outcome_expected["required"],
critical=True,
)
if outcome_expected["required"] and outcome is not None:
add("outcome_status", outcome["status"] in outcome_expected["statuses"])
combined = f"{outcome['text']} {outcome['scope']}"
add("outcome_meaning", _contains(combined, outcome_expected["terms"]))
add(
"outcome_scope",
_contains(combined, outcome_expected["scope_terms"]),
critical=True,
)
add(
"outcome_evidence",
set(outcome_expected["evidence_ids"]).issubset(outcome["evidence_ids"]),
critical=True,
)
actions = data["actions"]
expected_actions = expected["actions"]
add(
"action_count",
len(actions) == expected_actions["count"],
critical=True,
)
if expected_actions["count"] and actions:
action_text = " ".join(item["text"] for item in actions)
add("action_meaning", _contains(action_text, expected_actions["terms"]))
if "responsible" in expected_actions:
add(
"action_responsible",
any(item["responsible"] == expected_actions["responsible"] for item in actions),
critical=True,
)
if expected_actions.get("due_terms"):
due_text = " ".join(str(item["due"] or "") for item in actions)
add("action_due", _contains(due_text, expected_actions["due_terms"]))
add(
"action_evidence",
_refs_cover(actions, expected_actions["evidence_ids"]),
critical=True,
)
issues = data["unresolved_issues"]
expected_issues = expected["unresolved_issues"]
add(
"unresolved_count",
len(issues) == expected_issues["count"],
critical=True,
)
if expected_issues["count"] and issues:
issue_text = " ".join(item["text"] for item in issues)
add("unresolved_meaning", _contains(issue_text, expected_issues["terms"]))
add(
"unresolved_evidence",
_refs_cover(issues, expected_issues["evidence_ids"]),
critical=True,
)
passed = sum(item["passed"] for item in checks)
critical_failures = [
item["name"] for item in checks if item["critical"] and not item["passed"]
]
ratio = passed / len(checks)
if ratio == 1:
verdict = "PASS"
elif ratio >= 0.7 and not critical_failures:
verdict = "PARTIAL"
else:
verdict = "FAIL"
failed = [item["name"] for item in checks if not item["passed"]]
return {
"verdict": verdict,
"reason": "All semantic checks passed." if not failed else "Failed: " + ", ".join(failed),
"passed_checks": passed,
"check_count": len(checks),
"critical_failures": critical_failures,
"checks": checks,
}
def load_fixture(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict) or set(data) != {"cases"}:
raise SynthesisValidationError("fixture must contain exactly a cases list")
cases = data["cases"]
if not isinstance(cases, list) or not cases:
raise SynthesisValidationError("fixture cases must be a non-empty list")
seen: set[str] = set()
for case in cases:
validate_bundle(case)
if case["case_id"] in seen:
raise SynthesisValidationError(f"duplicate case ID: {case['case_id']}")
seen.add(case["case_id"])
return cases
def run_case(
case: dict[str, Any],
output_root: Path,
endpoint: str,
model: str,
timeout: int,
num_ctx: int,
num_predict: int,
) -> dict[str, Any]:
case_dir = output_root / case["case_id"]
case_dir.mkdir(parents=True, exist_ok=False)
gold_input = {
"case_id": case["case_id"],
"description": case["description"],
"subject_id": case["subject_id"],
"subject": case["subject"],
"evidence": case["evidence"],
}
(case_dir / "gold_input.json").write_text(
json.dumps(gold_input, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
started = time.perf_counter()
try:
raw_text, metadata = call_ollama(
endpoint, model, prompt, timeout, num_ctx, num_predict
)
(case_dir / "raw_model_response.txt").write_text(raw_text + "\n", encoding="utf-8")
(case_dir / "ollama_metadata.json").write_text(
json.dumps(metadata, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
parsed = parse_model_json(raw_text)
(case_dir / "parsed_response.json").write_text(
json.dumps(parsed, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
validated = validate_synthesis(parsed, case)
validation = {"valid": True, "error": None}
evaluation = evaluate_synthesis(validated, case["expected"])
except requests.RequestException:
raise
except (json.JSONDecodeError, SynthesisValidationError, ValueError) as exc:
validation = {
"valid": False,
"error_type": type(exc).__name__,
"error": str(exc),
}
evaluation = {
"verdict": "FAIL",
"reason": f"Schema validation failed: {exc}",
"passed_checks": 0,
"check_count": 0,
"critical_failures": ["schema_validation"],
"checks": [],
}
(case_dir / "validation_result.json").write_text(
json.dumps(validation, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
result = {
"case_id": case["case_id"],
"description": case["description"],
**evaluation,
"elapsed_seconds": round(time.perf_counter() - started, 3),
}
(case_dir / "evaluation_result.json").write_text(
json.dumps(result, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
return result
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
cases = load_fixture(args.fixture)
args.output.mkdir(parents=True, exist_ok=False)
started = time.perf_counter()
results: list[dict[str, Any]] = []
for index, case in enumerate(cases, start=1):
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
results.append(
run_case(
case,
args.output,
args.endpoint,
args.model,
args.timeout,
args.num_ctx,
args.num_predict,
)
)
summary = {
"experiment": "semantic_synthesis_isolation",
"schema_version": SCHEMA_VERSION,
"model": args.model,
"temperature": 0,
"think": False,
"num_ctx": args.num_ctx,
"num_predict": args.num_predict,
"case_count": len(cases),
"llm_call_count": len(results),
"runtime_seconds": round(time.perf_counter() - started, 3),
"verdict_counts": {
verdict: sum(result["verdict"] == verdict for result in results)
for verdict in ("PASS", "PARTIAL", "FAIL")
},
"results": results,
}
(args.output / "summary.json").write_text(
json.dumps(summary, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
)
return summary
def main() -> int:
args = parse_args()
try:
summary = run_experiment(args)
except (OSError, ValueError, requests.RequestException) as exc:
print(f"Error: {exc}")
return 1
print(json.dumps(summary["verdict_counts"], sort_keys=True))
print(f"Artifacts: {args.output.resolve()}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1 @@
"""Experimental topic-oriented discussion reconstruction."""
@@ -0,0 +1,721 @@
#!/usr/bin/env python3
"""Run an isolated Discussion Subject reconstruction experiment with Ollama."""
from __future__ import annotations
import argparse
import json
import re
import time
from pathlib import Path
from typing import Any
import requests
SCHEMA_VERSION = "experimental-discussion-subjects-v1"
DEFAULT_ENDPOINT = "http://127.0.0.1:11434/api/generate"
DEFAULT_MODEL = "qwen3.5:9B"
DEFAULT_TIMEOUT = 300
DEFAULT_NUM_CTX = 16384
DEFAULT_NUM_PREDICT = 4096
EVENT_TYPES = {
"introduced_idea",
"considered_option",
"proposal",
"supporting_argument",
"objection",
"clarification",
"modification",
"fact",
"technical_finding",
}
OUTCOME_CERTAINTIES = {"established", "tentative", "conditional", "rejected"}
IDENTIFIER_RE = re.compile(r"^[a-z][a-z0-9_]*$")
class ReconstructionValidationError(ValueError):
"""Raised when experimental reconstruction output violates the schema."""
PROMPT_TEMPLATE = """You reconstruct discussion subjects from meeting evidence.
This is semantic reconstruction, not protocol writing and not flat category extraction.
Group evidence by what participants are actually discussing. For each subject, record
only supported discourse events and, when present, the actual outcome, resulting
actions, and genuinely unresolved issues.
Important distinctions:
- discussed is not necessarily proposed
- proposed is not necessarily preferred or accepted
- preferred is not accepted
- accepted for a trial is not accepted as a final solution
- mentioned is not an unresolved question
- an outcome must preserve its scope, conditions, polarity, and uncertainty
- do not infer responsibility from mention, expertise, adjacency, or likely role
- do not invent missing stages or emit empty optional structures
Evidence discipline:
- Use only the supplied evidence IDs in evidence_refs.
- Every subject, event, outcome, action, and unresolved issue needs at least one
evidence reference.
- Keep statements concise; do not copy long evidence passages.
- A subject may consist only of one introduced idea.
Return one JSON object with exactly:
{{
"schema_version": "experimental-discussion-subjects-v1",
"subjects": [
{{
"subject_id": "subject_1",
"title": "concise discussion subject",
"evidence_refs": ["e1"],
"development": [
{{
"event_id": "event_1",
"type": "introduced_idea|considered_option|proposal|supporting_argument|objection|clarification|modification|fact|technical_finding",
"text": "what happened in the discussion",
"evidence_refs": ["e1"]
}}
],
"outcome": {{
"text": "only what was established",
"scope": "explicit limit or full scope of the outcome",
"certainty": "established|tentative|conditional|rejected",
"evidence_refs": ["e2"]
}},
"actions": [
{{
"action_id": "action_1",
"text": "established work only",
"responsible": "explicitly supported name or null",
"deadline": "explicitly supported deadline or null",
"evidence_refs": ["e3"]
}}
],
"unresolved_issues": [
{{
"issue_id": "issue_1",
"text": "concrete unresolved issue",
"evidence_refs": ["e4"]
}}
]
}}
]
}}
Only subject_id, title, evidence_refs are required for each subject. Omit
development, outcome, actions, or unresolved_issues when absent. Never emit null
or an empty optional list/object.
Case ID: {case_id}
Evidence units:
{evidence_json}
"""
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Run the isolated topic-reconstruction Gold experiment."
)
parser.add_argument("fixture", type=Path, help="Focused Gold cases JSON.")
parser.add_argument("-o", "--output", type=Path, required=True)
parser.add_argument("--model", default=DEFAULT_MODEL)
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
parser.add_argument("--timeout", type=int, default=DEFAULT_TIMEOUT)
parser.add_argument("--num-ctx", type=int, default=DEFAULT_NUM_CTX)
parser.add_argument("--num-predict", type=int, default=DEFAULT_NUM_PREDICT)
parser.add_argument(
"--case", action="append", dest="case_ids", help="Run only this case ID."
)
return parser.parse_args()
def _expect_exact_keys(
value: dict[str, Any], required: set[str], optional: set[str], location: str
) -> None:
missing = required - value.keys()
unknown = value.keys() - required - optional
if missing:
raise ReconstructionValidationError(
f"{location} missing required keys: {sorted(missing)}"
)
if unknown:
raise ReconstructionValidationError(
f"{location} has unknown keys: {sorted(unknown)}"
)
def _nonempty_text(value: Any, location: str) -> str:
if not isinstance(value, str) or not value.strip():
raise ReconstructionValidationError(f"{location} must be a non-empty string")
return value.strip()
def _identifier(value: Any, location: str, seen: set[str]) -> str:
text = _nonempty_text(value, location)
if not IDENTIFIER_RE.fullmatch(text):
raise ReconstructionValidationError(f"{location} is not a valid identifier")
if text in seen:
raise ReconstructionValidationError(f"duplicate identifier: {text}")
seen.add(text)
return text
def _nullable_text(value: Any, location: str) -> str | None:
if value is None:
return None
text = _nonempty_text(value, location)
if text.casefold() == "null":
raise ReconstructionValidationError(
f"{location} must use JSON null, not the string 'null'"
)
return text
def _evidence_refs(value: Any, location: str, known: set[str]) -> list[str]:
if not isinstance(value, list) or not value:
raise ReconstructionValidationError(f"{location} must be a non-empty list")
refs: list[str] = []
for index, ref in enumerate(value):
ref = _nonempty_text(ref, f"{location}[{index}]")
if ref not in known:
raise ReconstructionValidationError(
f"{location}[{index}] references unknown evidence ID: {ref}"
)
if ref in refs:
raise ReconstructionValidationError(
f"{location} contains duplicate evidence reference: {ref}"
)
refs.append(ref)
return refs
def validate_evidence_units(evidence_units: Any) -> set[str]:
if not isinstance(evidence_units, list) or not evidence_units:
raise ReconstructionValidationError("evidence_units must be a non-empty list")
known: set[str] = set()
for index, unit in enumerate(evidence_units):
location = f"evidence_units[{index}]"
if not isinstance(unit, dict):
raise ReconstructionValidationError(f"{location} must be an object")
_expect_exact_keys(unit, {"evidence_id", "text"}, set(), location)
evidence_id = _nonempty_text(unit["evidence_id"], f"{location}.evidence_id")
if evidence_id in known:
raise ReconstructionValidationError(
f"duplicate input evidence identifier: {evidence_id}"
)
known.add(evidence_id)
_nonempty_text(unit["text"], f"{location}.text")
return known
def validate_reconstruction(data: Any, evidence_units: Any) -> dict[str, Any]:
known = validate_evidence_units(evidence_units)
if not isinstance(data, dict):
raise ReconstructionValidationError("model output must be an object")
_expect_exact_keys(data, {"schema_version", "subjects"}, set(), "output")
if data["schema_version"] != SCHEMA_VERSION:
raise ReconstructionValidationError(
f"schema_version must be {SCHEMA_VERSION!r}"
)
subjects = data["subjects"]
if not isinstance(subjects, list) or not subjects:
raise ReconstructionValidationError("subjects must be a non-empty list")
seen: set[str] = set()
for subject_index, subject in enumerate(subjects):
location = f"subjects[{subject_index}]"
if not isinstance(subject, dict):
raise ReconstructionValidationError(f"{location} must be an object")
_expect_exact_keys(
subject,
{"subject_id", "title", "evidence_refs"},
{"development", "outcome", "actions", "unresolved_issues"},
location,
)
_identifier(subject["subject_id"], f"{location}.subject_id", seen)
_nonempty_text(subject["title"], f"{location}.title")
_evidence_refs(subject["evidence_refs"], f"{location}.evidence_refs", known)
if "development" in subject:
events = subject["development"]
if not isinstance(events, list) or not events:
raise ReconstructionValidationError(
f"{location}.development must be a non-empty list when present"
)
for event_index, event in enumerate(events):
event_location = f"{location}.development[{event_index}]"
if not isinstance(event, dict):
raise ReconstructionValidationError(
f"{event_location} must be an object"
)
_expect_exact_keys(
event,
{"event_id", "type", "text", "evidence_refs"},
set(),
event_location,
)
_identifier(event["event_id"], f"{event_location}.event_id", seen)
if event["type"] not in EVENT_TYPES:
raise ReconstructionValidationError(
f"{event_location}.type is invalid: {event['type']!r}"
)
_nonempty_text(event["text"], f"{event_location}.text")
_evidence_refs(
event["evidence_refs"], f"{event_location}.evidence_refs", known
)
if "outcome" in subject:
outcome = subject["outcome"]
outcome_location = f"{location}.outcome"
if not isinstance(outcome, dict):
raise ReconstructionValidationError(
f"{outcome_location} must be a non-empty object when present"
)
_expect_exact_keys(
outcome,
{"text", "scope", "certainty", "evidence_refs"},
set(),
outcome_location,
)
_nonempty_text(outcome["text"], f"{outcome_location}.text")
_nonempty_text(outcome["scope"], f"{outcome_location}.scope")
if outcome["certainty"] not in OUTCOME_CERTAINTIES:
raise ReconstructionValidationError(
f"{outcome_location}.certainty is invalid: {outcome['certainty']!r}"
)
_evidence_refs(
outcome["evidence_refs"], f"{outcome_location}.evidence_refs", known
)
if "actions" in subject:
actions = subject["actions"]
if not isinstance(actions, list) or not actions:
raise ReconstructionValidationError(
f"{location}.actions must be a non-empty list when present"
)
for action_index, action in enumerate(actions):
action_location = f"{location}.actions[{action_index}]"
if not isinstance(action, dict):
raise ReconstructionValidationError(
f"{action_location} must be an object"
)
_expect_exact_keys(
action,
{"action_id", "text", "responsible", "deadline", "evidence_refs"},
set(),
action_location,
)
_identifier(action["action_id"], f"{action_location}.action_id", seen)
_nonempty_text(action["text"], f"{action_location}.text")
for field in ("responsible", "deadline"):
_nullable_text(action[field], f"{action_location}.{field}")
_evidence_refs(
action["evidence_refs"], f"{action_location}.evidence_refs", known
)
if "unresolved_issues" in subject:
issues = subject["unresolved_issues"]
if not isinstance(issues, list) or not issues:
raise ReconstructionValidationError(
f"{location}.unresolved_issues must be a non-empty list when present"
)
for issue_index, issue in enumerate(issues):
issue_location = f"{location}.unresolved_issues[{issue_index}]"
if not isinstance(issue, dict):
raise ReconstructionValidationError(
f"{issue_location} must be an object"
)
_expect_exact_keys(
issue,
{"issue_id", "text", "evidence_refs"},
set(),
issue_location,
)
_identifier(issue["issue_id"], f"{issue_location}.issue_id", seen)
_nonempty_text(issue["text"], f"{issue_location}.text")
_evidence_refs(
issue["evidence_refs"], f"{issue_location}.evidence_refs", known
)
return data
def build_prompt(case: dict[str, Any]) -> str:
evidence_units = case["evidence_units"]
validate_evidence_units(evidence_units)
return PROMPT_TEMPLATE.format(
case_id=case["case_id"],
evidence_json=json.dumps(evidence_units, ensure_ascii=False, indent=2),
)
def parse_model_json(raw_text: str) -> dict[str, Any]:
data = json.loads(raw_text)
if not isinstance(data, dict):
raise ReconstructionValidationError("model response JSON must be an object")
return data
def build_ollama_payload(
model: str, prompt: str, num_ctx: int, num_predict: int
) -> dict[str, Any]:
return {
"model": model,
"prompt": prompt,
"think": False,
"stream": False,
"format": "json",
"options": {
"temperature": 0,
"num_ctx": num_ctx,
"num_predict": num_predict,
},
}
def call_ollama(
endpoint: str,
model: str,
prompt: str,
timeout: int,
num_ctx: int,
num_predict: int,
) -> tuple[str, dict[str, Any]]:
payload = build_ollama_payload(model, prompt, num_ctx, num_predict)
started = time.perf_counter()
response = requests.post(endpoint, json=payload, timeout=timeout)
elapsed = time.perf_counter() - started
response.raise_for_status()
data = response.json()
if not isinstance(data, dict):
raise ValueError("Ollama response must be a JSON object")
raw_text = data.get("response")
if not isinstance(raw_text, str) or not raw_text.strip():
raise ValueError("Ollama returned no usable response text")
metadata = {
"model": data.get("model", model),
"elapsed_seconds": round(elapsed, 3),
"total_duration_ns": data.get("total_duration"),
"load_duration_ns": data.get("load_duration"),
"prompt_eval_count": data.get("prompt_eval_count"),
"prompt_eval_duration_ns": data.get("prompt_eval_duration"),
"eval_count": data.get("eval_count"),
"eval_duration_ns": data.get("eval_duration"),
"configuration": {
"temperature": 0,
"think": False,
"num_ctx": num_ctx,
"num_predict": num_predict,
},
}
return raw_text.strip(), metadata
def _all_text(subjects: list[dict[str, Any]]) -> str:
parts: list[str] = []
for subject in subjects:
parts.append(subject["title"])
for event in subject.get("development", []):
parts.append(event["text"])
outcome = subject.get("outcome")
if outcome:
parts.extend((outcome["text"], outcome["scope"]))
for action in subject.get("actions", []):
parts.append(action["text"])
for issue in subject.get("unresolved_issues", []):
parts.append(issue["text"])
return " ".join(parts).casefold()
def _contains_any(text: str, terms: list[str]) -> bool:
return any(term.casefold() in text for term in terms)
def evaluate_reconstruction(
reconstruction: dict[str, Any], expected: dict[str, Any]
) -> dict[str, Any]:
subjects = reconstruction["subjects"]
combined = _all_text(subjects)
events = [event for subject in subjects for event in subject.get("development", [])]
outcomes = [subject["outcome"] for subject in subjects if "outcome" in subject]
actions = [action for subject in subjects for action in subject.get("actions", [])]
issues = [issue for subject in subjects for issue in subject.get("unresolved_issues", [])]
checks: list[dict[str, Any]] = []
def add(name: str, passed: bool, critical: bool = False) -> None:
checks.append({"name": name, "passed": passed, "critical": critical})
add("subject_count", len(subjects) == expected.get("subject_count", 1))
add("subject_identity", _contains_any(combined, expected["subject_terms"]))
event_types = {event["type"] for event in events}
for event_type in expected.get("required_event_types", []):
add(f"event_type:{event_type}", event_type in event_types)
expected_outcome = expected.get("outcome", {})
outcome_required = expected_outcome.get("required", False)
add(
"outcome_presence",
bool(outcomes) is outcome_required,
critical=not outcome_required and bool(outcomes),
)
if outcome_required and outcomes:
outcome_text = " ".join(
f"{item['text']} {item['scope']}" for item in outcomes
).casefold()
add("outcome_meaning", _contains_any(outcome_text, expected_outcome["terms"]))
add(
"outcome_scope",
_contains_any(outcome_text, expected_outcome.get("scope_terms", [])),
critical=True,
)
add(
"outcome_certainty",
any(
item["certainty"] in expected_outcome.get("certainties", [])
for item in outcomes
),
)
expected_actions = expected.get("actions", {})
minimum_actions = expected_actions.get("minimum", 0)
add(
"action_count",
len(actions) >= minimum_actions if minimum_actions else not actions,
critical=minimum_actions == 0 and bool(actions),
)
if minimum_actions and actions:
action_text = " ".join(item["text"] for item in actions).casefold()
add("action_meaning", _contains_any(action_text, expected_actions["terms"]))
if "responsible" in expected_actions:
add(
"action_responsibility",
any(
item["responsible"] == expected_actions["responsible"]
for item in actions
),
critical=True,
)
expected_issues = expected.get("unresolved", {})
minimum_issues = expected_issues.get("minimum", 0)
add(
"unresolved_count",
len(issues) >= minimum_issues if minimum_issues else not issues,
critical=minimum_issues == 0 and bool(issues),
)
if minimum_issues and issues:
issue_text = " ".join(item["text"] for item in issues).casefold()
add("unresolved_meaning", _contains_any(issue_text, expected_issues["terms"]))
passed = sum(check["passed"] for check in checks)
critical_failures = [
check["name"] for check in checks if check["critical"] and not check["passed"]
]
ratio = passed / len(checks)
if ratio == 1:
verdict = "PASS"
elif ratio >= 0.6 and not critical_failures:
verdict = "PARTIAL"
else:
verdict = "FAIL"
failed = [check["name"] for check in checks if not check["passed"]]
reason = "All semantic checks passed." if not failed else "Failed: " + ", ".join(failed)
return {
"verdict": verdict,
"reason": reason,
"passed_checks": passed,
"check_count": len(checks),
"critical_failures": critical_failures,
"checks": checks,
}
def load_fixture(path: Path) -> list[dict[str, Any]]:
data = json.loads(path.read_text(encoding="utf-8-sig"))
if not isinstance(data, dict) or set(data) != {"cases"}:
raise ValueError("fixture must contain exactly one 'cases' list")
cases = data["cases"]
if not isinstance(cases, list) or not cases:
raise ValueError("fixture cases must be a non-empty list")
seen: set[str] = set()
for index, case in enumerate(cases):
if not isinstance(case, dict):
raise ValueError(f"cases[{index}] must be an object")
required = {"case_id", "description", "evidence_units", "expected"}
if set(case) != required:
raise ValueError(f"cases[{index}] must contain exactly {sorted(required)}")
case_id = _nonempty_text(case["case_id"], f"cases[{index}].case_id")
if case_id in seen:
raise ValueError(f"duplicate case_id: {case_id}")
seen.add(case_id)
_nonempty_text(case["description"], f"cases[{index}].description")
validate_evidence_units(case["evidence_units"])
if not isinstance(case["expected"], dict):
raise ValueError(f"cases[{index}].expected must be an object")
return cases
def run_case(
case: dict[str, Any],
output_root: Path,
endpoint: str,
model: str,
timeout: int,
num_ctx: int,
num_predict: int,
) -> dict[str, Any]:
case_dir = output_root / case["case_id"]
case_dir.mkdir(parents=True, exist_ok=False)
input_payload = {
"case_id": case["case_id"],
"description": case["description"],
"evidence_units": case["evidence_units"],
}
(case_dir / "input.json").write_text(
json.dumps(input_payload, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
prompt = build_prompt(case)
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
started = time.perf_counter()
try:
raw_text, metadata = call_ollama(
endpoint, model, prompt, timeout, num_ctx, num_predict
)
(case_dir / "raw_model_response.txt").write_text(
raw_text + "\n", encoding="utf-8"
)
(case_dir / "ollama_metadata.json").write_text(
json.dumps(metadata, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
parsed = parse_model_json(raw_text)
(case_dir / "parsed_output.json").write_text(
json.dumps(parsed, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
validated = validate_reconstruction(parsed, case["evidence_units"])
evaluation = evaluate_reconstruction(validated, case["expected"])
except requests.RequestException as exc:
failure = {
"case_id": case["case_id"],
"error_type": type(exc).__name__,
"error": str(exc),
"elapsed_seconds": round(time.perf_counter() - started, 3),
}
(case_dir / "validation_failure.json").write_text(
json.dumps(failure, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
raise
except (json.JSONDecodeError, ReconstructionValidationError, ValueError) as exc:
elapsed = round(time.perf_counter() - started, 3)
failure = {
"case_id": case["case_id"],
"error_type": type(exc).__name__,
"error": str(exc),
"elapsed_seconds": elapsed,
}
(case_dir / "validation_failure.json").write_text(
json.dumps(failure, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
result = {
"case_id": case["case_id"],
"description": case["description"],
"verdict": "FAIL",
"reason": f"{type(exc).__name__}: {exc}",
"passed_checks": 0,
"check_count": 0,
"critical_failures": ["schema_validation"],
"checks": [],
"elapsed_seconds": elapsed,
"subject_titles": [],
}
(case_dir / "evaluation.json").write_text(
json.dumps(result, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
return result
result = {
"case_id": case["case_id"],
"description": case["description"],
**evaluation,
"elapsed_seconds": metadata["elapsed_seconds"],
"subject_titles": [item["title"] for item in validated["subjects"]],
}
(case_dir / "evaluation.json").write_text(
json.dumps(result, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
return result
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
cases = load_fixture(args.fixture)
selected = set(args.case_ids or [])
if selected:
known = {case["case_id"] for case in cases}
unknown = selected - known
if unknown:
raise ValueError(f"unknown requested case IDs: {sorted(unknown)}")
cases = [case for case in cases if case["case_id"] in selected]
args.output.mkdir(parents=True, exist_ok=False)
results: list[dict[str, Any]] = []
started = time.perf_counter()
for index, case in enumerate(cases, start=1):
print(f"[{index}/{len(cases)}] {case['case_id']}", flush=True)
results.append(
run_case(
case,
args.output,
args.endpoint,
args.model,
args.timeout,
args.num_ctx,
args.num_predict,
)
)
summary = {
"experiment": "topic_reconstruction_v2",
"schema_version": SCHEMA_VERSION,
"model": args.model,
"temperature": 0,
"think": False,
"case_count": len(cases),
"llm_call_count": len(results),
"runtime_seconds": round(time.perf_counter() - started, 3),
"verdict_counts": {
verdict: sum(item["verdict"] == verdict for item in results)
for verdict in ("PASS", "PARTIAL", "FAIL")
},
"results": results,
}
(args.output / "summary.json").write_text(
json.dumps(summary, ensure_ascii=False, indent=2) + "\n",
encoding="utf-8",
)
return summary
def main() -> int:
args = parse_args()
try:
summary = run_experiment(args)
except (OSError, ValueError, requests.RequestException) as exc:
print(f"Error: {exc}")
return 1
print(json.dumps(summary["verdict_counts"], sort_keys=True))
print(f"Artifacts: {args.output.resolve()}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,5 @@
"""Audio transcription support for the direct-protocol MVP."""
from .whisper import TranscriptionError, TranscriptionResult, transcribe_audio
__all__ = ["TranscriptionError", "TranscriptionResult", "transcribe_audio"]
+290
View File
@@ -0,0 +1,290 @@
"""Isolated whisper.cpp wrapper producing Meeting Lab compact transcripts."""
from __future__ import annotations
import json
import os
import platform
import re
import shutil
import subprocess
import tempfile
import time
from dataclasses import dataclass
from datetime import datetime, timezone
from pathlib import Path
from typing import Any, Callable, Sequence
BACKEND = "whisper.cpp"
RAW_FILENAME = "whisper_raw.json"
TRANSCRIPT_FILENAME = "transcript.json"
TEXT_FILENAME = "transcript.txt"
METADATA_FILENAME = "runtime_metadata.json"
DEFAULT_THREADS = "auto"
class TranscriptionError(RuntimeError):
"""Raised when parameters, Whisper execution, or output are invalid."""
@dataclass(frozen=True)
class TranscriptionResult:
output_dir: Path
raw_output: Path
transcript_json: Path
transcript_text: Path
runtime_metadata: Path
runtime_seconds: float
def _logical_cpu_count() -> int:
"""Return the available logical CPU count as a last-resort fallback."""
if hasattr(os, "sched_getaffinity"):
try:
count = len(os.sched_getaffinity(0))
if count > 0:
return count
except OSError:
pass
return os.cpu_count() or 1
def _linux_physical_core_count() -> int | None:
affinity = None
if hasattr(os, "sched_getaffinity"):
try:
affinity = os.sched_getaffinity(0)
except OSError:
pass
cores: set[tuple[str, str]] = set()
for cpu_dir in Path("/sys/devices/system/cpu").glob("cpu[0-9]*"):
try:
cpu_number = int(cpu_dir.name[3:])
if affinity is not None and cpu_number not in affinity:
continue
topology = cpu_dir / "topology"
package = (topology / "physical_package_id").read_text().strip()
core = (topology / "core_id").read_text().strip()
cores.add((package, core))
except (OSError, ValueError):
continue
return len(cores) or None
def _darwin_physical_core_count() -> int | None:
try:
completed = subprocess.run(
("sysctl", "-n", "hw.physicalcpu"),
check=False,
capture_output=True,
text=True,
)
count = int(completed.stdout.strip())
return count if completed.returncode == 0 and count > 0 else None
except (OSError, ValueError):
return None
def physical_core_count() -> int:
"""Detect physical cores where supported, falling back to available threads."""
system = platform.system()
detected = _linux_physical_core_count() if system == "Linux" else None
if system == "Darwin":
detected = _darwin_physical_core_count()
return detected or _logical_cpu_count()
def resolve_threads(
threads: str | int,
detector: Callable[[], int] = physical_core_count,
) -> int:
if isinstance(threads, bool):
raise TranscriptionError("Threads must be 'auto' or a positive integer.")
if threads == "auto":
count = detector()
else:
try:
count = int(threads)
except (TypeError, ValueError) as exc:
raise TranscriptionError("Threads must be 'auto' or a positive integer.") from exc
if count <= 0:
raise TranscriptionError("Threads must be 'auto' or a positive integer.")
return count
def _vulkan_support(raw: dict[str, Any]) -> bool | None:
system_info = raw.get("systeminfo")
if not isinstance(system_info, str) or "VULKAN" not in system_info.upper():
return None
return re.search(r"VULKAN\s*=\s*1", system_info, re.IGNORECASE) is not None
def _json_object(path: Path) -> dict[str, Any]:
try:
data = json.loads(path.read_text(encoding="utf-8"))
except (OSError, UnicodeError, json.JSONDecodeError) as exc:
raise TranscriptionError(f"Cannot read Whisper JSON output {path}: {exc}") from exc
if not isinstance(data, dict):
raise TranscriptionError("Whisper JSON output must contain a top-level object.")
return data
def compact_transcript(raw: dict[str, Any]) -> dict[str, Any]:
"""Convert whisper.cpp JSON without linguistic cleanup or reordering."""
entries = raw.get("transcription")
if not isinstance(entries, list):
raise TranscriptionError("Whisper JSON output must contain a 'transcription' list.")
segments: list[dict[str, Any]] = []
for index, entry in enumerate(entries):
if not isinstance(entry, dict):
raise TranscriptionError(f"transcription[{index}] must be an object.")
offsets = entry.get("offsets")
if not isinstance(offsets, dict):
raise TranscriptionError(f"transcription[{index}].offsets must be an object.")
start_ms = offsets.get("from")
end_ms = offsets.get("to")
if not isinstance(start_ms, (int, float)) or isinstance(start_ms, bool):
raise TranscriptionError(f"transcription[{index}].offsets.from must be a number.")
if not isinstance(end_ms, (int, float)) or isinstance(end_ms, bool):
raise TranscriptionError(f"transcription[{index}].offsets.to must be a number.")
if end_ms < start_ms:
raise TranscriptionError(
f"transcription[{index}].offsets.to must be greater than or equal to offsets.from."
)
text_value = entry.get("text", "")
if not isinstance(text_value, str):
raise TranscriptionError(f"transcription[{index}].text must be a string.")
text = text_value.strip()
if text:
segments.append(
{
"id": len(segments),
"start": float(start_ms) / 1000.0,
"end": float(end_ms) / 1000.0,
"text": text,
}
)
return {"text": " ".join(item["text"] for item in segments), "segments": segments}
def _timestamp(seconds: float) -> str:
milliseconds = int(round(seconds * 1000))
hours, remainder = divmod(milliseconds, 3_600_000)
minutes, remainder = divmod(remainder, 60_000)
secs, millis = divmod(remainder, 1000)
return f"{hours:02d}:{minutes:02d}:{secs:02d}.{millis:03d}"
def transcript_text(transcript: dict[str, Any]) -> str:
lines = [
f"[{_timestamp(item['start'])} - {_timestamp(item['end'])}] {item['text']}"
for item in transcript["segments"]
]
return "\n".join(lines) + ("\n" if lines else "")
def _write_json(path: Path, value: dict[str, Any]) -> None:
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
def transcribe_audio(
audio_path: Path,
model_path: Path,
output_dir: Path,
language: str = "auto",
*,
executable: str = "whisper-cli",
threads: str | int = DEFAULT_THREADS,
thread_detector: Callable[[], int] = physical_core_count,
runner: Callable[..., subprocess.CompletedProcess[str]] = subprocess.run,
monotonic: Callable[[], float] = time.monotonic,
now: Callable[[], datetime] = lambda: datetime.now(timezone.utc),
) -> TranscriptionResult:
"""Run one whisper.cpp call and write raw, compact, text, and metadata outputs."""
audio_path = Path(audio_path)
model_path = Path(model_path)
output_dir = Path(output_dir)
if not audio_path.is_file():
raise TranscriptionError(f"Audio file does not exist: {audio_path}")
if not model_path.is_file():
raise TranscriptionError(f"Whisper model does not exist: {model_path}")
if not isinstance(language, str) or not language.strip():
raise TranscriptionError("Language must be a non-empty string.")
if not executable.strip():
raise TranscriptionError("Whisper executable must be a non-empty string.")
if output_dir.exists() and not output_dir.is_dir():
raise TranscriptionError(f"Output directory path is not a directory: {output_dir}")
thread_count = resolve_threads(threads, thread_detector)
output_dir.mkdir(parents=True, exist_ok=True)
raw_output = output_dir / RAW_FILENAME
started_at = now().astimezone(timezone.utc)
started = monotonic()
with tempfile.TemporaryDirectory(prefix=".whisper-", dir=output_dir) as temp_name:
temporary_prefix = Path(temp_name) / "whisper_raw"
command: Sequence[str] = (
executable,
"-m", str(model_path),
"-f", str(audio_path),
"-l", language.strip(),
"-t", str(thread_count),
"-fa",
"-oj",
"-of", str(temporary_prefix),
)
try:
completed = runner(command, check=False, capture_output=True, text=True)
except OSError as exc:
raise TranscriptionError(f"Could not start {BACKEND}: {exc}") from exc
runtime_seconds = monotonic() - started
temporary_raw = temporary_prefix.with_suffix(".json")
if temporary_raw.is_file():
shutil.copyfile(temporary_raw, raw_output)
if completed.returncode != 0:
detail = completed.stderr.strip() or completed.stdout.strip() or "no diagnostic output"
raise TranscriptionError(
f"{BACKEND} failed with exit code {completed.returncode}: {detail}"
)
if not raw_output.is_file():
raise TranscriptionError(f"{BACKEND} completed without producing JSON output.")
raw_data = _json_object(raw_output)
transcript = compact_transcript(raw_data)
transcript_json_path = output_dir / TRANSCRIPT_FILENAME
transcript_text_path = output_dir / TEXT_FILENAME
metadata_path = output_dir / METADATA_FILENAME
_write_json(transcript_json_path, transcript)
transcript_text_path.write_text(transcript_text(transcript), encoding="utf-8")
duration = max((item["end"] for item in transcript["segments"]), default=None)
metadata = {
"input_file": str(audio_path.resolve()),
"model": str(model_path.resolve()),
"backend": BACKEND,
"whisper_executable": executable,
"language": language.strip(),
"threads": thread_count,
"threads_option": str(threads),
"flash_attention": True,
"vulkan_support_detected": _vulkan_support(raw_data),
"duration_seconds": duration,
"runtime_seconds": runtime_seconds,
"timestamp": started_at.isoformat(),
"output_files": {
"whisper_raw": RAW_FILENAME,
"transcript_json": TRANSCRIPT_FILENAME,
"transcript_text": TEXT_FILENAME,
"runtime_metadata": METADATA_FILENAME,
},
}
_write_json(metadata_path, metadata)
return TranscriptionResult(
output_dir=output_dir,
raw_output=raw_output,
transcript_json=transcript_json_path,
transcript_text=transcript_text_path,
runtime_metadata=metadata_path,
runtime_seconds=runtime_seconds,
)
@@ -0,0 +1,95 @@
{
"schema_version": "experimental-collective-commitment-gold-v0",
"cases": [
{
"case_id": "CC-01",
"description": "Explicit collective commitment",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": true, "due": "nächste Woche"}
},
{
"case_id": "CC-02",
"description": "Individual commitment",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, ich teste nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "individual_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": false, "due": null}
},
{
"case_id": "CC-03",
"description": "Tentative collective possibility",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten nächste Woche 20 Meter testen.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": false, "due": null}
},
{
"case_id": "CC-04",
"description": "Collective suggestion",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Vielleicht sollten wir nächste Woche 20 Meter testen.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": false, "due": null}
},
{
"case_id": "CC-05",
"description": "Impersonal necessity",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Man müsste nächste Woche 20 Meter testen.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": false, "due": null}
},
{
"case_id": "CC-06",
"description": "Passive future statement",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Nächste Woche werden 20 Meter getestet.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "none", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": false, "due": null}
},
{
"case_id": "CC-07",
"description": "Collective rejection",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Nein, das testen wir nächste Woche nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "none", "action_concepts": [], "qualifier_concepts": []},
"expected_result": {"established": false, "due": null}
},
{
"case_id": "CC-08",
"description": "Collective commitment with qualifier",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen 20 Meter, aber nur im Technikum.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": [["nur", "only"], ["technikum", "technical facility", "technical center", "technical centre"]]},
"expected_result": {"established": true, "due": null}
},
{
"case_id": "CC-09",
"description": "Collective commitment without deadline",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": true, "due": null}
},
{
"case_id": "CC-10",
"description": "Speaker ownership trap",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"commitment_form": "collective_first_person", "action_concepts": [["20"], ["meter", "metre"]], "qualifier_concepts": []},
"expected_result": {"established": true, "due": "nächste Woche"}
}
]
}
@@ -0,0 +1,13 @@
{
"schema_version": "experimental-controlled-rejection-v1",
"cases": [
{"case_id":"CR-01","description":"self-contained non-pursuit","negative_act_source":"NA-01","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_1","target_observation_id":"obs_1","action_concepts":[["Schlummer"],["Zusammenarbeit","arbeiten"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":true}},
{"case_id":"CR-02","description":"paired non-pursuit","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":true}},
{"case_id":"CR-03","description":"personal preference","negative_act_source":"NA-03","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage für den Versuch nutzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde das nicht machen.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"personal_preference","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["reale Anlage"],["Versuch"],["nutzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
{"case_id":"CR-04","description":"recommendation","negative_act_source":"NA-04","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage verwenden.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde eher davon abraten.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"recommendation","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["reale Anlage"],["verwenden","nutzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
{"case_id":"CR-05","description":"temporary non-action","negative_act_source":"NA-05","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die Waschstufe einbauen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir erstmal noch nicht.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"temporary_non_action","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Waschstufe"],["einbauen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
{"case_id":"CR-06","description":"concern","negative_act_source":"NA-06","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten das neue Material einsetzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das wäre kritisch.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"none","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["neue Material","neues Material"],["einsetzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
{"case_id":"CR-07","description":"scoped explicit rejection","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[],"explicitly_rejected":true}},
{"case_id":"CR-08","description":"rejection plus alternative","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"],"explicitly_rejected":true}}
]
}
@@ -0,0 +1,147 @@
{
"cases": [
{
"case_id": "a_idea_only",
"description": "Possible geometry optimization without commitment.",
"subject_id": "subject_a",
"subject": "Optimierung der Geometrie",
"evidence": [{"evidence_id": "e1", "text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Geometrie kann vielleicht optimiert werden.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"positive","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"absent"},
{"observation_id":"obs_2","evidence_id":"e1","content":"Danach könnte betrachtet werden, was herauskommt.","target":"obs_1","relation":"qualifies","modality":"suggested","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"implicit","scope":"nach der Optimierung|danach"}
]
},
{
"case_id": "b_multiple_options",
"description": "Two alternatives for insufficient grid strength.",
"subject_id": "subject_b",
"subject": "Umgang mit unzureichender Festigkeit des 40-40-Gitters",
"evidence": [
{"evidence_id":"e1","text":"Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},
{"evidence_id":"e2","text":"Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},
{"evidence_id":"e3","text":"Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Festigkeit reicht noch nicht aus.","target":"discussion_subject","relation":"none","modality":"factual","temporality":"existing","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"40-40-Gitter|40-40"},
{"observation_id":"obs_2","evidence_id":"e2","content":"Mehr Masse könnte für die gleiche Festigkeit eingesetzt werden.","target":"obs_1","relation":"qualifies","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"mehr Masse|gleiche Festigkeit"},
{"observation_id":"obs_3","evidence_id":"e3","content":"Das Produkt könnte als 20-20 statt 40-40 ausgeführt werden.","target":"obs_1","relation":"qualifies","modality":"suggested","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"20-20|statt 40-40"},
{"observation_id":"obs_4","evidence_id":"e3","content":"Die vorherigen Möglichkeiten sind die zwei Ansätze.","target":["obs_2","obs_3"],"relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"absent"}
]
},
{
"case_id": "c_unaccepted_proposal",
"description": "Suggested Textor contact without established work.",
"subject_id": "subject_c",
"subject": "Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
"evidence": [
{"evidence_id":"e1","text":"Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},
{"evidence_id":"e2","text":"Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Tim erwägt, Dirk Textor erneut zu kontaktieren und nach seiner Einschätzung zu fragen.","target":"discussion_subject","relation":"none","modality":"suggested","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"Dirk Textors Einschätzung|erneut kontaktieren"},
{"observation_id":"obs_2","evidence_id":"e2","content":"Eine erneute Rückkopplung mit Dirk Textor ist möglich.","target":"obs_1","relation":"supports","modality":"possible","temporality":"future","evaluation":"positive","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Rückkopplung mit Dirk Textor|mit ihm"}
]
},
{
"case_id": "d_proposal_with_objection",
"description": "Washing possibility and explicit energy disadvantage.",
"subject_id": "subject_d",
"subject": "Waschen des Materials vor der weiteren Verarbeitung",
"evidence": [
{"evidence_id":"e1","text":"Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},
{"evidence_id":"e2","text":"Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Material könnte gewaschen werden.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"vor der weiteren Verarbeitung"},
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin weiß nicht, ob sich das Waschen lohnt.","target":"obs_1","relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"Nutzen des Waschens|ob es sich lohnt"},
{"observation_id":"obs_3","evidence_id":"e2","content":"Waschen umfasst Nassmachen und erneutes Trocknen.","target":"obs_1","relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Waschprozess|Nassmachen und Trocknen"},
{"observation_id":"obs_4","evidence_id":"e2","content":"Waschen und Trocknen verursachen einen sehr hohen Energieaufwand.","target":"obs_1","relation":"opposes","modality":"factual","temporality":"existing","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Energieaufwand des Waschens|Waschen und Trocknen"}
]
},
{
"case_id": "e_rejected_alternative",
"description": "Explicit rejection followed by confirmation of that rejection.",
"subject_id": "subject_e",
"subject": "Zusammenarbeit mit Dr. Schlummer für Versuche",
"evidence": [
{"evidence_id":"e1","text":"Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},
{"evidence_id":"e2","text":"Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},
{"evidence_id":"e3","text":"Antonius: Ja, das ist entschieden."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Angebot kostet 30.000 Euro.","target":"discussion_subject","relation":"none","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Angebot für die Versuche|30.000 Euro"},
{"observation_id":"obs_2","evidence_id":"e2","content":"Die Zusammenarbeit mit Dr. Schlummer wird nicht durchgeführt.","target":"discussion_subject","relation":"none","modality":"committed","temporality":"future","evaluation":"none","agreement":"rejected","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Zusammenarbeit für die Versuche|Dr. Schlummer"},
{"observation_id":"obs_3","evidence_id":"e3","content":"Die vorherige Ablehnung ist entschieden.","target":"obs_2","relation":"supports","modality":"factual","temporality":"completed","evaluation":"none","agreement":"accepted","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"absent"}
]
},
{
"case_id": "f_trial_only_acceptance",
"description": "Acceptance limited to a 20-metre trial.",
"subject_id": "subject_f",
"subject": "20-Prozent-Variante im Versuch am kleinen Extruder",
"evidence": [
{"evidence_id":"e1","text":"Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},
{"evidence_id":"e2","text":"Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},
{"evidence_id":"e3","text":"Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Die 20-Prozent-Variante könnte am kleinen Extruder nachgestellt werden.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"kleiner Extruder"},
{"observation_id":"obs_2","evidence_id":"e2","content":"20 Meter der Variante werden beim nächsten Versuch getestet.","target":"obs_1","relation":"supports","modality":"committed","temporality":"future","evaluation":"none","agreement":"accepted","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"20 Meter beim nächsten Versuch|20 Meter"},
{"observation_id":"obs_3","evidence_id":"e3","content":"Die Zusage gilt nur für einen Versuch.","target":"obs_2","relation":"limits_scope","modality":"factual","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"nur ein Versuch|Versuch"},
{"observation_id":"obs_4","evidence_id":"e3","content":"Die Variante ist noch nicht als Serienlösung festgelegt.","target":"discussion_subject","relation":"qualifies","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"implicit","scope":"Serienlösung|finale Produktion"}
]
},
{
"case_id": "g_no_decision",
"description": "Preference, alternative, and impersonal checking need without decision.",
"subject_id": "subject_g",
"subject": "Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
"evidence": [
{"evidence_id":"e1","text":"Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},
{"evidence_id":"e2","text":"Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},
{"evidence_id":"e3","text":"Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Eine reale Recyclinganlage birgt das Risiko kontaminierten Rückmaterials.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"reale Recyclinganlage|kontaminiertes Material"},
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin würde nicht in eine reale Anlage gehen.","target":"discussion_subject","relation":"opposes","modality":"suggested","temporality":"future","evaluation":"negative","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"Martins persönliche Präferenz|reale Anlage"},
{"observation_id":"obs_3","evidence_id":"e2","content":"Ein Technikum bleibt als bedingte Möglichkeit im Gespräch.","target":"discussion_subject","relation":"none","modality":"possible","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"none","scope":"wenn überhaupt|Technikum"},
{"observation_id":"obs_4","evidence_id":"e3","content":"Zunächst muss geprüft werden, welcher Reinigungsansatz verfügbar ist.","target":"discussion_subject","relation":"qualifies","modality":"impersonal_necessity","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"zunächst|verfügbarer Reinigungsansatz"}
]
},
{
"case_id": "h_resulting_action",
"description": "Interpersonal request followed by accepted responsibility.",
"subject_id": "subject_h",
"subject": "Prüfung der Messdaten bis Freitag",
"evidence": [
{"evidence_id":"e1","text":"Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},
{"evidence_id":"e2","text":"Nina: Ja, ich übernehme die Prüfung bis Freitag."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Antonius bittet Nina um die Prüfung der Messdaten.","target":"discussion_subject","relation":"none","modality":"interpersonal_request","temporality":"future","evaluation":"none","agreement":"none","responsibility":"named","person":"Nina","uncertainty":"absent","clarification_need":"none","scope":"bis Freitag|Freitag"},
{"observation_id":"obs_2","evidence_id":"e2","content":"Nina übernimmt die Prüfung.","target":"obs_1","relation":"supports","modality":"committed","temporality":"future","evaluation":"none","agreement":"accepted","responsibility":"accepted","person":"Nina","uncertainty":"absent","clarification_need":"none","scope":"bis Freitag|Freitag"}
]
},
{
"case_id": "i_outcome_and_unresolved",
"description": "Bounded production finding and unresolved publication information.",
"subject_id": "subject_i",
"subject": "Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
"evidence": [
{"evidence_id":"e1","text":"Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},
{"evidence_id":"e2","text":"Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},
{"evidence_id":"e3","text":"Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},
{"evidence_id":"e4","text":"Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}
],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Bei der reinen Produktion gab es praktisch keine Änderung gegenüber dem Standardprodukt.","target":"discussion_subject","relation":"none","modality":"factual","temporality":"completed","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"eigene Anlage, reine Produktion, Standardprodukt|reine Produktion"},
{"observation_id":"obs_2","evidence_id":"e1","content":"Die Produktion erfolgte fünf Grad kälter.","target":"obs_1","relation":"qualifies","modality":"factual","temporality":"completed","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"fünf Grad kälter|5 Grad"},
{"observation_id":"obs_3","evidence_id":"e2","content":"Gegenüber Virgin Material ist bei reiner Produktion kein zusätzlicher Aufwand notwendig.","target":"obs_1","relation":"supports","modality":"factual","temporality":"existing","evaluation":"none","agreement":"accepted","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"reine Produktion gegenüber Virgin Material|Virgin Material"},
{"observation_id":"obs_4","evidence_id":"e2","content":"Vor der reinen Produktion entsteht Aufwand.","target":"obs_3","relation":"limits_scope","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"absent","clarification_need":"none","scope":"vor der reinen Produktion|davor"},
{"observation_id":"obs_5","evidence_id":"e3","content":"Es wird gefragt, welche Energieaudit-Daten veröffentlicht werden dürfen.","target":"discussion_subject","relation":"none","modality":"information_question","temporality":"unspecified","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"Veröffentlichung von Energieaudit-Daten|Energieaudit"},
{"observation_id":"obs_6","evidence_id":"e4","content":"Die Veröffentlichungserlaubnis ist weiterhin ungeklärt.","target":"obs_5","relation":"supports","modality":"factual","temporality":"existing","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"Veröffentlichungserlaubnis|Freigabe"},
{"observation_id":"obs_7","evidence_id":"e4","content":"Die Freigabe muss noch geklärt werden.","target":"obs_5","relation":"supports","modality":"impersonal_necessity","temporality":"future","evaluation":"none","agreement":"none","responsibility":"none","person":null,"uncertainty":"present","clarification_need":"explicit","scope":"Freigabe zur Veröffentlichung|Freigabe"}
]
}
]
}
@@ -0,0 +1,99 @@
{
"cases": [
{
"case_id": "a_idea_only", "description": "Possible geometry optimization without commitment.",
"subject_id": "subject_a", "subject": "Optimierung der Geometrie",
"evidence": [{"evidence_id": "e1", "text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Geometrie kann vielleicht optimiert werden.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"positive","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":null,"limits_target":null},
{"observation_id":"obs_2","evidence_id":"e1","content":"Danach könnte betrachtet werden, was herauskommt.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"suggested","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"implicit","qualifier":"danach","limits_target":null}
]
},
{
"case_id": "b_multiple_options", "description": "Two alternatives for insufficient grid strength.",
"subject_id": "subject_b", "subject": "Umgang mit unzureichender Festigkeit des 40-40-Gitters",
"evidence": [{"evidence_id":"e1","text":"Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},{"evidence_id":"e2","text":"Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},{"evidence_id":"e3","text":"Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Die Festigkeit des 40-40-Gitters reicht noch nicht aus.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"negative","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"40-40-Gitter","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"Mehr Masse könnte für die gleiche Festigkeit eingesetzt werden.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"mehr Masse für die gleiche Festigkeit","limits_target":null},
{"observation_id":"obs_3","evidence_id":"e3","content":"Das Produkt könnte als 20-20 statt 40-40 ausgeführt werden.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"suggested","temporality":"future","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"20-20 statt 40-40","limits_target":null},
{"observation_id":"obs_4","evidence_id":"e3","content":"Die vorherigen Möglichkeiten sind die zwei Ansätze.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"zwei Ansätze","limits_target":null}
]
},
{
"case_id": "c_unaccepted_proposal", "description": "Suggested Textor contact without established work.",
"subject_id": "subject_c", "subject": "Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
"evidence": [{"evidence_id":"e1","text":"Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},{"evidence_id":"e2","text":"Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Tim erwägt, Dirk Textor erneut zu kontaktieren und nach seiner Einschätzung zu fragen.","refers_to":null,"speaker":"Tim","named_person":"Dirk Textor","addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"suggested","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"erneut; Dirk Textors Einschätzung","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"Eine erneute Rückkopplung mit Dirk Textor ist möglich.","refers_to":"obs_1","speaker":"Tim","named_person":"Dirk Textor","addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"positive","affirmation":"explicit","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"noch einmal mit ihm","limits_target":null}
]
},
{
"case_id": "d_proposal_with_objection", "description": "Washing possibility and explicit energy disadvantage.",
"subject_id": "subject_d", "subject": "Waschen des Materials vor der weiteren Verarbeitung",
"evidence": [{"evidence_id":"e1","text":"Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},{"evidence_id":"e2","text":"Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Material könnte gewaschen werden.","refers_to":null,"speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"vor der weiteren Verarbeitung","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin weiß nicht, ob sich das Waschen lohnt.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"ob es sich lohnt","limits_target":null},
{"observation_id":"obs_3","evidence_id":"e2","content":"Waschen umfasst Nassmachen und erneutes Trocknen.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"nass machen und wieder trocknen","limits_target":null},
{"observation_id":"obs_4","evidence_id":"e2","content":"Waschen und Trocknen verursachen einen sehr hohen Energieaufwand.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"negative","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"Waschen und Trocknen","limits_target":null}
]
},
{
"case_id": "e_rejected_alternative", "description": "Explicit negation followed by confirmation of that determination.",
"subject_id": "subject_e", "subject": "Zusammenarbeit mit Dr. Schlummer für Versuche",
"evidence": [{"evidence_id":"e1","text":"Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},{"evidence_id":"e2","text":"Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},{"evidence_id":"e3","text":"Antonius: Ja, das ist entschieden."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Das Angebot von Dr. Schlummer kostet 30.000 Euro.","refers_to":null,"speaker":"Antonius","named_person":"Dr. Schlummer","addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"für die Versuche; 30.000 Euro","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"Die Zusammenarbeit mit Dr. Schlummer wird nicht durchgeführt.","refers_to":null,"speaker":"Tim","named_person":"Dr. Schlummer","addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"committed","temporality":"future","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"present","uncertainty":"absent","clarification_need":"none","qualifier":"Zusammenarbeit für die Versuche","limits_target":null},
{"observation_id":"obs_3","evidence_id":"e3","content":"Die vorherige Festlegung ist entschieden.","refers_to":"obs_2","speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"completed","evaluation":"none","affirmation":"explicit","negation":"absent","determination_statement":"present","uncertainty":"absent","clarification_need":"none","qualifier":null,"limits_target":null}
]
},
{
"case_id": "f_trial_only_acceptance", "description": "Affirmed commitment limited to a 20-metre trial.",
"subject_id": "subject_f", "subject": "20-Prozent-Variante im Versuch am kleinen Extruder",
"evidence": [{"evidence_id":"e1","text":"Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},{"evidence_id":"e2","text":"Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},{"evidence_id":"e3","text":"Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Die 20-Prozent-Variante könnte am kleinen Extruder nachgestellt werden.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"am kleinen Extruder","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"20 Meter der Variante werden beim nächsten Versuch getestet.","refers_to":"obs_1","speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"committed","temporality":"future","evaluation":"none","affirmation":"explicit","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"20 Meter beim nächsten Versuch","limits_target":null},
{"observation_id":"obs_3","evidence_id":"e3","content":"Die Zusage gilt nur für einen Versuch.","refers_to":"obs_2","speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"nur ein Versuch","limits_target":"obs_2"},
{"observation_id":"obs_4","evidence_id":"e3","content":"Die Variante ist noch nicht als Serienlösung festgelegt.","refers_to":"obs_2","speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"present","uncertainty":"present","clarification_need":"implicit","qualifier":"als Serienlösung","limits_target":null}
]
},
{
"case_id": "g_no_decision", "description": "Preference, alternative, and impersonal checking need without decision.",
"subject_id": "subject_g", "subject": "Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
"evidence": [{"evidence_id":"e1","text":"Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},{"evidence_id":"e2","text":"Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},{"evidence_id":"e3","text":"Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Eine reale Recyclinganlage birgt das Risiko kontaminierten Rückmaterials.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"possible","temporality":"future","evaluation":"negative","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"reale Recyclinganlage; kontaminiertes Material","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"Martin würde persönlich nicht in eine reale Anlage gehen.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"suggested","temporality":"future","evaluation":"negative","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"reale Anlage","limits_target":null},
{"observation_id":"obs_3","evidence_id":"e2","content":"Ein Technikum bleibt als bedingte Möglichkeit im Gespräch.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"possible","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"none","qualifier":"wenn überhaupt; Technikum","limits_target":null},
{"observation_id":"obs_4","evidence_id":"e3","content":"Zunächst muss geprüft werden, welcher Reinigungsansatz verfügbar ist.","refers_to":null,"speaker":"Tim","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":true,"modality":"impersonal_necessity","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"zunächst; verfügbarer Reinigungsansatz","limits_target":null}
]
},
{
"case_id": "h_resulting_action", "description": "Interpersonal request followed by explicit personal acceptance.",
"subject_id": "subject_h", "subject": "Prüfung der Messdaten bis Freitag",
"evidence": [{"evidence_id":"e1","text":"Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},{"evidence_id":"e2","text":"Nina: Ja, ich übernehme die Prüfung bis Freitag."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Antonius richtet an Nina die Bitte, die Messdaten zu prüfen.","refers_to":null,"speaker":"Antonius","named_person":"Nina","addressee":"Nina","self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"interpersonal_request","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"bis Freitag","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e2","content":"Nina sagt zu, die Prüfung zu übernehmen.","refers_to":"obs_1","speaker":"Nina","named_person":null,"addressee":null,"self_reference":true,"collective_we":false,"impersonal_person_reference":false,"modality":"committed","temporality":"future","evaluation":"none","affirmation":"explicit","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"bis Freitag","limits_target":null}
]
},
{
"case_id": "i_outcome_and_unresolved", "description": "Bounded production finding and unresolved publication information.",
"subject_id": "subject_i", "subject": "Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
"evidence": [{"evidence_id":"e1","text":"Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},{"evidence_id":"e2","text":"Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},{"evidence_id":"e3","text":"Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},{"evidence_id":"e4","text":"Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}],
"expected_observations": [
{"observation_id":"obs_1","evidence_id":"e1","content":"Bei der reinen Produktion gab es praktisch keine Änderung gegenüber dem Standardprodukt.","refers_to":null,"speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"completed","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"an unserer Anlage; reine Produktion; gegenüber dem Standardprodukt","limits_target":null},
{"observation_id":"obs_2","evidence_id":"e1","content":"Die Produktion erfolgte fünf Grad kälter.","refers_to":"obs_1","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"factual","temporality":"completed","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"fünf Grad kälter","limits_target":null},
{"observation_id":"obs_3","evidence_id":"e2","content":"Gegenüber Virgin Material ist bei reiner Produktion kein zusätzlicher Aufwand notwendig.","refers_to":"obs_1","speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"explicit","determination_statement":"present","uncertainty":"absent","clarification_need":"none","qualifier":"bei reiner Produktion; gegenüber Virgin Material","limits_target":null},
{"observation_id":"obs_4","evidence_id":"e2","content":"Vor der reinen Produktion entsteht Aufwand.","refers_to":"obs_3","speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"absent","clarification_need":"none","qualifier":"davor","limits_target":"obs_3"},
{"observation_id":"obs_5","evidence_id":"e3","content":"Es wird gefragt, welche Energieaudit-Daten veröffentlicht werden dürfen.","refers_to":null,"speaker":"Antonius","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"information_question","temporality":"unspecified","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"Veröffentlichung von Energieaudit-Daten","limits_target":null},
{"observation_id":"obs_6","evidence_id":"e4","content":"Die Veröffentlichungserlaubnis ist weiterhin ungeklärt.","refers_to":"obs_5","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":false,"impersonal_person_reference":false,"modality":"factual","temporality":"existing","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"weiterhin","limits_target":null},
{"observation_id":"obs_7","evidence_id":"e4","content":"Die Freigabe muss noch geklärt werden.","refers_to":"obs_5","speaker":"Martin","named_person":null,"addressee":null,"self_reference":false,"collective_we":true,"impersonal_person_reference":false,"modality":"impersonal_necessity","temporality":"future","evaluation":"none","affirmation":"absent","negation":"absent","determination_statement":"absent","uncertainty":"present","clarification_need":"explicit","qualifier":"noch; Freigabe zur Veröffentlichung","limits_target":null}
]
}
]
}
@@ -0,0 +1,49 @@
{
"cases": [
{
"case_id":"a_idea_only","description":"Possible geometry optimization without commitment.","subject_id":"subject_a","subject":"Optimierung der Geometrie",
"evidence":[{"evidence_id":"e1","text":"Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}],
"semantic_requirements":["Geometry optimization remains possible and tentative.","Subsequent checking remains conditional and tentative.","The then/sequential dependency survives.","No commitment or owner is introduced."]
},
{
"case_id":"b_multiple_options","description":"Two alternatives for insufficient grid strength.","subject_id":"subject_b","subject":"Umgang mit unzureichender Festigkeit des 40-40-Gitters",
"evidence":[{"evidence_id":"e1","text":"Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},{"evidence_id":"e2","text":"Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},{"evidence_id":"e3","text":"Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}],
"semantic_requirements":["Insufficient 40-40 strength survives.","Additional mass remains one alternative.","20-20 remains another alternative.","Both remain alternatives and neither is selected."]
},
{
"case_id":"c_unaccepted_proposal","description":"Suggested Textor contact without established work.","subject_id":"subject_c","subject":"Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
"evidence":[{"evidence_id":"e1","text":"Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},{"evidence_id":"e2","text":"Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}],
"semantic_requirements":["Contacting Dirk Textor remains Tim's tentative personal suggestion.","The follow-up remains possible and relates to that contact.","No established work or responsibility is introduced."]
},
{
"case_id":"d_proposal_with_objection","description":"Washing possibility and explicit energy disadvantage.","subject_id":"subject_d","subject":"Waschen des Materials vor der weiteren Verarbeitung",
"evidence":[{"evidence_id":"e1","text":"Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},{"evidence_id":"e2","text":"Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}],
"semantic_requirements":["Washing before further processing remains possible.","Martin's uncertainty whether washing is worthwhile survives.","The washing and drying process survives.","The high energy consequence survives.","No unresolved task is invented."]
},
{
"case_id":"e_rejected_alternative","description":"Explicit rejection followed by confirmation of that determination.","subject_id":"subject_e","subject":"Zusammenarbeit mit Dr. Schlummer für Versuche",
"evidence":[{"evidence_id":"e1","text":"Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},{"evidence_id":"e2","text":"Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},{"evidence_id":"e3","text":"Antonius: Ja, das ist entschieden."}],
"semantic_requirements":["The offer cost survives without inferred evaluation.","Collaboration is explicitly not to be pursued.","The explicit rejection survives.","The later statement confirms that the preceding determination has been made."]
},
{
"case_id":"f_trial_only_acceptance","description":"Collective commitment limited to a 20-metre trial.","subject_id":"subject_f","subject":"20-Prozent-Variante im Versuch am kleinen Extruder",
"evidence":[{"evidence_id":"e1","text":"Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},{"evidence_id":"e2","text":"Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},{"evidence_id":"e3","text":"Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}],
"semantic_requirements":["The 20-percent variant at the small extruder remains initially possible.","The later statement collectively commits to a test.","Twenty metres and next-trial timing survive.","The test remains limited to a trial.","Series adoption remains explicitly not yet established.","No individual owner is invented."]
},
{
"case_id":"g_no_decision","description":"Preference, conditional alternative, and impersonal checking need.","subject_id":"subject_g","subject":"Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
"evidence":[{"evidence_id":"e1","text":"Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},{"evidence_id":"e2","text":"Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},{"evidence_id":"e3","text":"Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}],
"semantic_requirements":["Contamination remains a risk rather than a fact.","Martin's negative stance remains personal.","The Technikum remains conditional and if-at-all survives.","Cleaning-method availability still needs to be checked.","The need remains impersonal.","No group decision or owner is invented."]
},
{
"case_id":"h_resulting_action","description":"Interpersonal request followed by explicit personal acceptance.","subject_id":"subject_h","subject":"Prüfung der Messdaten bis Freitag",
"evidence":[{"evidence_id":"e1","text":"Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},{"evidence_id":"e2","text":"Nina: Ja, ich übernehme die Prüfung bis Freitag."}],
"semantic_requirements":["Antonius requests measurement-data review from Nina.","The Friday deadline survives.","Nina explicitly accepts the preceding request.","Nina's response expresses future personal commitment.","No responsibility field or unsupported inference is introduced."]
},
{
"case_id":"i_outcome_and_unresolved","description":"Bounded production finding and unresolved publication information.","subject_id":"subject_i","subject":"Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
"evidence":[{"evidence_id":"e1","text":"Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},{"evidence_id":"e2","text":"Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},{"evidence_id":"e3","text":"Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},{"evidence_id":"e4","text":"Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}],
"semantic_requirements":["The pure-production finding remains bounded to the local plant and standard-product comparison.","The five-degree difference survives.","No-extra-effort remains bounded to pure production compared with Virgin material.","Upstream effort before that production stage survives.","The publication purpose of the energy-audit question remains explicit.","Publication permission remains unresolved and clarification remains necessary.","No assigned work is invented."]
}
]
}
+111
View File
@@ -0,0 +1,111 @@
{
"schema_version": "experimental-explicit-rejection-gold-v0",
"cases": [
{
"case_id": "RJ-01", "description": "Explicit collective rejection with local target",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Nein, das machen wir nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": true}
},
{
"case_id": "RJ-02", "description": "Explicit non-pursuit with paired target",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": true}
},
{
"case_id": "RJ-03", "description": "Self-contained collaboration rejection",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null}
],
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_1", "target_observation_id": "obs_1", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": true}
},
{
"case_id": "RJ-04", "description": "Personal preference",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-05", "description": "Concern",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-06", "description": "Uncertainty",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-07", "description": "Negative recommendation",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-08", "description": "Deferral",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die externe Lösung einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das entscheiden wir nächste Woche.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-09", "description": "Factual negation",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_1", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-10", "description": "Temporary non-action",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": false}
},
{
"case_id": "RJ-11", "description": "Explicit rejection with material scope",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Die reale Anlage nutzen wir dafür nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"]], "qualifier_concepts": [["druckversuch", "dafür", "pressure test"]], "forbidden_action_concepts": []},
"expected_result": {"explicitly_rejected": true}
},
{
"case_id": "RJ-12", "description": "Rejection plus positive alternative",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten den Versuch in der realen Anlage durchführen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir nicht; wir testen stattdessen im Technikum.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["versuch", "trial", "test"], ["real"], ["anlage", "plant"]], "qualifier_concepts": [["real"], ["anlage", "plant"]], "forbidden_action_concepts": ["technikum", "technical facility", "technical center", "technical centre"]},
"expected_result": {"explicitly_rejected": true}
}
]
}
@@ -0,0 +1,66 @@
{
"schema_version": "experimental-negative-act-form-gold-v0",
"cases": [
{
"case_id": "NA-01", "description": "Explicit non-pursuit",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null}
],
"expected": {"observation_id": "obs_1", "negative_act_form": "explicit_non_pursuit", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]]}
},
{
"case_id": "NA-02", "description": "Explicit non-pursuit paraphrase",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Die externe Lösung verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_1", "negative_act_form": "explicit_non_pursuit", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]]}
},
{
"case_id": "NA-03", "description": "Personal preference with local context",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_2", "negative_act_form": "personal_preference", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]]}
},
{
"case_id": "NA-04", "description": "Negative recommendation",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_2", "negative_act_form": "recommendation", "action_concepts": [["real"], ["anlage", "plant"], ["verwend", "use"]]}
},
{
"case_id": "NA-05", "description": "Temporary non-action",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_2", "negative_act_form": "temporary_non_action", "action_concepts": [["waschstufe", "washing stage"], ["einbau", "install"]]}
},
{
"case_id": "NA-06", "description": "Concern only",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_2", "negative_act_form": "none", "action_concepts": []}
},
{
"case_id": "NA-07", "description": "Uncertainty",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_2", "negative_act_form": "none", "action_concepts": []}
},
{
"case_id": "NA-08", "description": "Factual negation",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected": {"observation_id": "obs_1", "negative_act_form": "none", "action_concepts": []}
}
]
}
+101
View File
@@ -0,0 +1,101 @@
{
"schema_version": "experimental-request-acceptance-gold-v0",
"cases": [
{
"case_id": "RA-01",
"description": "Explicit positive acceptance",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja, ich übernehme die Auswertung bis Dienstag.", "speaker": "Clara", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": true, "commitment": true, "same_work": true},
"expected_result": {"established": true, "content": "Auswertung der Messwerte", "requested_actor": "Clara", "responsible_person": "Clara", "due": "Dienstag"}
},
{
"case_id": "RA-02",
"description": "Paraphrased positive acceptance",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja. Ich kümmere mich darum und habe die Auswertung bis Dienstag fertig.", "speaker": "Clara", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": true, "commitment": true, "same_work": true},
"expected_result": {"established": true, "content": "Auswertung der Messwerte", "requested_actor": "Clara", "responsible_person": "Clara", "due": "Dienstag"}
},
{
"case_id": "RA-03",
"description": "Acknowledgement only",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja, ich habe verstanden, worum es geht.", "speaker": "Clara", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-04",
"description": "Tentative response",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ich schaue mal, ob ich das schaffe.", "speaker": "Clara", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-05",
"description": "Different responder without personal acceptance",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ja, das sollte gemacht werden.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-06",
"description": "Explicit commitment to different work",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"},
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Clara: Ja, ich kümmere mich um die Präsentation.", "speaker": "Clara", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": true, "commitment": true, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-07",
"description": "Request without response",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Antonius: Clara, kannst du die Messwerte bis Dienstag auswerten?", "speaker": "Antonius", "named_person": "Clara", "addressee": "Clara"}
],
"expected_recognition": {"request": true, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-08",
"description": "Collective commitment",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ja, wir testen nächste Woche 20 Meter.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": false, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-09",
"description": "Impersonal necessity",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Man müsste zuerst prüfen, welches Reinigungsverfahren verfügbar ist.", "speaker": "Martin", "named_person": null, "addressee": null}
],
"expected_recognition": {"request": false, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
},
{
"case_id": "RA-10",
"description": "Personal suggestion",
"observations": [
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Ich würde vielleicht Dirk Textor kontaktieren.", "speaker": "Martin", "named_person": "Dirk Textor", "addressee": null}
],
"expected_recognition": {"request": false, "commitment": false, "same_work": false},
"expected_result": {"established": false, "content": null, "requested_actor": null, "responsible_person": null, "due": null}
}
]
}
@@ -0,0 +1,11 @@
# semantic_synthesis_isolation
Isolation Gold set derived from the existing Topic Reconstruction V2 A-I
cases. Every case supplies one manually fixed Discussion Subject and the
complete original evidence bundle. The model performs Semantic Synthesis only;
subject detection, subject grouping, and evidence assignment are outside the
experiment.
Expected criteria evaluate semantic event distinctions, outcomes and scope,
actions, unresolved issues, and supporting evidence. They do not evaluate
subject discovery or exact wording.
@@ -0,0 +1,227 @@
{
"cases": [
{
"case_id": "a_idea_only",
"description": "Idea mentioned without stronger commitment.",
"subject_id": "subject_a",
"subject": "Optimierung der Geometrie",
"evidence": [
{"evidence_id": "e1", "text": "Martin: Die Geometrie kann man vielleicht noch optimieren. Dann würde man mal gucken, was herauskommt."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"idea": 1},
"allowed_event_types": ["idea"],
"event_evidence_ids": ["e1"],
"outcome": {"required": false},
"actions": {"count": 0},
"unresolved_issues": {"count": 0}
}
},
{
"case_id": "b_multiple_options",
"description": "Two alternatives for insufficient 40-40 grid strength.",
"subject_id": "subject_b",
"subject": "Umgang mit unzureichender Festigkeit des 40-40-Gitters",
"evidence": [
{"evidence_id": "e1", "text": "Martin: Die Festigkeit reicht für das 40-40-Gitter noch nicht aus."},
{"evidence_id": "e2", "text": "Martin: Man könnte mehr Masse für die gleiche Festigkeit einsetzen."},
{"evidence_id": "e3", "text": "Martin: Oder wir verkaufen es nicht als 40-40-Gitter, sondern machen ein 20-20 daraus. Das wären die zwei Ansätze."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"option": 2},
"allowed_event_types": ["technical_finding", "fact", "option"],
"event_evidence_ids": ["e1", "e2", "e3"],
"outcome": {"required": false},
"actions": {"count": 0},
"unresolved_issues": {"count": 0}
}
},
{
"case_id": "c_unaccepted_proposal",
"description": "Possible Textor contact remains a proposal only.",
"subject_id": "subject_c",
"subject": "Erneute Kontaktaufnahme mit Dirk Textor zur Einschätzung",
"evidence": [
{"evidence_id": "e1", "text": "Tim: Ich würde vielleicht Dirk Textor noch einmal kontaktieren und fragen, wie er das einschätzt."},
{"evidence_id": "e2", "text": "Tim: Das kann man ja mit ihm einfach noch einmal rückkoppeln."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"proposal": 1},
"allowed_event_types": ["proposal"],
"event_evidence_ids": ["e1", "e2"],
"outcome": {"required": false},
"actions": {"count": 0},
"unresolved_issues": {"count": 0}
}
},
{
"case_id": "d_proposal_with_objection",
"description": "Washing proposal with energy objection but no unresolved issue.",
"subject_id": "subject_d",
"subject": "Waschen des Materials vor der weiteren Verarbeitung",
"evidence": [
{"evidence_id": "e1", "text": "Antonius: Man könnte das Material vor der weiteren Verarbeitung waschen."},
{"evidence_id": "e2", "text": "Martin: Ob sich das lohnt, weiß ich nicht. Waschen heißt nass machen und wieder trocknen; das ist ein wahnsinniger Energieaufwand."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"proposal": 1, "objection": 1},
"allowed_event_types": ["proposal", "objection"],
"event_evidence_ids": ["e1", "e2"],
"outcome": {"required": false},
"actions": {"count": 0},
"unresolved_issues": {"count": 0}
}
},
{
"case_id": "e_rejected_alternative",
"description": "Explicit rejection of Schlummer collaboration.",
"subject_id": "subject_e",
"subject": "Zusammenarbeit mit Dr. Schlummer für Versuche",
"evidence": [
{"evidence_id": "e1", "text": "Antonius: Das Angebot von Dr. Schlummer für die Versuche kostet 30.000 Euro."},
{"evidence_id": "e2", "text": "Tim: Dann haben wir gesagt: Nein, die Zusammenarbeit mit Dr. Schlummer machen wir nicht."},
{"evidence_id": "e3", "text": "Antonius: Ja, das ist entschieden."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"fact": 1, "rejection": 1},
"allowed_event_types": ["fact", "rejection", "clarification"],
"event_evidence_ids": ["e1", "e2", "e3"],
"outcome": {
"required": true,
"statuses": ["rejected"],
"terms": ["nicht", "abgelehnt", "keine"],
"scope_terms": ["zusammenarbeit", "versuch", "schlummer"],
"evidence_ids": ["e2", "e3"]
},
"actions": {"count": 0},
"unresolved_issues": {"count": 0}
}
},
{
"case_id": "f_trial_only_acceptance",
"description": "Acceptance limited to a 20-metre trial.",
"subject_id": "subject_f",
"subject": "20-Prozent-Variante im Versuch am kleinen Extruder",
"evidence": [
{"evidence_id": "e1", "text": "Martin: Wir könnten die 20-Prozent-Variante am kleinen Extruder nachstellen."},
{"evidence_id": "e2", "text": "Tim: Ja, wir testen 20 Meter dieser Variante beim nächsten Versuch."},
{"evidence_id": "e3", "text": "Tim: Das ist nur ein Versuch; damit ist die Variante noch nicht als Serienlösung festgelegt."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"proposal": 1, "scoped_acceptance": 1, "clarification": 1},
"allowed_event_types": ["proposal", "scoped_acceptance", "clarification"],
"event_evidence_ids": ["e1", "e2", "e3"],
"outcome": {
"required": true,
"statuses": ["scoped_acceptance"],
"terms": ["test", "versuch"],
"scope_terms": ["20 meter", "20 m", "nur", "begrenzt"],
"evidence_ids": ["e2", "e3"]
},
"actions": {
"count": 1,
"terms": ["test", "versuch", "20 meter"],
"evidence_ids": ["e2"],
"due_terms": ["nächsten versuch", "next trial"]
},
"unresolved_issues": {
"count": 1,
"terms": ["serienlösung", "final", "serie", "festgelegt"],
"evidence_ids": ["e3"]
}
}
},
{
"case_id": "g_no_decision",
"description": "Plant versus Technikum discussion ending without a decision.",
"subject_id": "subject_g",
"subject": "Reale Recyclinganlage oder Technikum und verfügbarer Reinigungsansatz",
"evidence": [
{"evidence_id": "e1", "text": "Martin: Eine reale Recyclinganlage hätte das Risiko, dass wir kontaminiertes Material zurückbekommen."},
{"evidence_id": "e2", "text": "Martin: Ich würde nicht in eine reale Anlage gehen. Wenn überhaupt, können wir über ein Technikum reden."},
{"evidence_id": "e3", "text": "Tim: Man müsste zunächst prüfen, welcher Reinigungsansatz überhaupt verfügbar ist."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"objection": 1, "option": 1},
"allowed_event_types": ["objection", "option", "proposal", "clarification"],
"event_evidence_ids": ["e1", "e2", "e3"],
"outcome": {"required": false},
"actions": {"count": 0},
"unresolved_issues": {
"count": 1,
"terms": ["reinigungsansatz", "verfügbar", "prüfen", "reinigung"],
"evidence_ids": ["e3"]
}
}
},
{
"case_id": "h_resulting_action",
"description": "Explicitly accepted action with owner and deadline.",
"subject_id": "subject_h",
"subject": "Prüfung der Messdaten bis Freitag",
"evidence": [
{"evidence_id": "e1", "text": "Antonius: Nina, übernimmst du die Prüfung der Messdaten bis Freitag?"},
{"evidence_id": "e2", "text": "Nina: Ja, ich übernehme die Prüfung bis Freitag."}
],
"allowed_responsible": ["Nina"],
"expected": {
"event_type_minimums": {},
"allowed_event_types": ["proposal", "clarification", "scoped_acceptance"],
"event_evidence_ids": [],
"outcome": {
"required": true,
"statuses": ["established"],
"terms": ["übernimmt", "prüf", "accepted", "review", "assigned"],
"scope_terms": ["messdaten", "prüfung", "measurement", "review"],
"evidence_ids": ["e2"]
},
"actions": {
"count": 1,
"terms": ["messdaten", "prüf", "measurement", "review"],
"responsible": "Nina",
"due_terms": ["freitag", "friday"],
"evidence_ids": ["e2"]
},
"unresolved_issues": {"count": 0}
}
},
{
"case_id": "i_outcome_and_unresolved",
"description": "Bounded production outcome and unresolved publication question.",
"subject_id": "subject_i",
"subject": "Produktionsaufwand und Veröffentlichung von Energieaudit-Daten",
"evidence": [
{"evidence_id": "e1", "text": "Martin: An unserer Anlage gab es bei der reinen Produktion gegenüber dem Standardprodukt praktisch keine Änderung; wir waren nur fünf Grad kälter."},
{"evidence_id": "e2", "text": "Antonius: Dann können wir mindestens festhalten: Gegenüber Virgin Material ist bei der reinen Produktion kein zusätzlicher Aufwand notwendig. Davor entsteht natürlich Aufwand."},
{"evidence_id": "e3", "text": "Antonius: Welche Daten aus dem Energieaudit dürfen wir veröffentlichen?"},
{"evidence_id": "e4", "text": "Martin: Das ist weiterhin ungeklärt. Wir müssen die Freigabe noch klären."}
],
"allowed_responsible": [],
"expected": {
"event_type_minimums": {"fact": 1},
"allowed_event_types": ["technical_finding", "fact", "clarification"],
"event_evidence_ids": ["e1", "e2"],
"outcome": {
"required": true,
"statuses": ["established"],
"terms": ["kein zusätzlicher", "keine zusätzliche", "unverändert"],
"scope_terms": ["reine produktion", "produktion", "virgin"],
"evidence_ids": ["e1", "e2"]
},
"actions": {"count": 0},
"unresolved_issues": {
"count": 1,
"terms": ["veröffentlich", "freigabe", "energieaudit", "daten"],
"evidence_ids": ["e3", "e4"]
}
}
}
]
}
@@ -0,0 +1,9 @@
{
"schema_version":"experimental-target-normalization-v0",
"cases":[
{"case_id":"TN-01","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_1","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"Zusammenarbeit mit Dr. Schlummer fortsetzen","action_concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":["nicht","beenden"],"german_markers":["Zusammenarbeit","arbeiten","fortsetzen"]}},
{"case_id":"TN-02","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"externe Lösung weiterverfolgen","action_concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":["nicht"],"german_markers":["Lösung","verfolgen"]}},
{"case_id":"TN-03","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"reale Anlage für den Druckversuch nutzen","action_concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":["nicht","zur Diskussion"],"german_markers":["Anlage","Druckversuch","nutzen"]}},
{"case_id":"TN-04","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"Versuch in der realen Anlage durchführen","action_concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["nicht","Technikum","stattdessen"],"german_markers":["Versuch","Anlage","durchführen"]}}
]
}
@@ -0,0 +1,13 @@
{
"schema_version": "experimental-target-resolution-v0",
"cases": [
{"case_id":"TR-01","description":"self-contained continuation target","strategy":"self_contained","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"negative_act":{"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"working with Dr. Schlummer"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":[]}},
{"case_id":"TR-02","description":"paired pronoun target","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"verfolgen wir nicht weiter"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[]}},
{"case_id":"TR-03","description":"scoped location and purpose target","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"reale Anlage dafür nicht nutzen"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[]}},
{"case_id":"TR-04","description":"rejection plus alternative","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"Versuch in der realen Anlage nicht durchführen"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"]}},
{"case_id":"TR-05","description":"personal preference","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage für den Versuch nutzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde das nicht machen.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"personal_preference","normalized_action_text":"Ich würde das nicht machen"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
{"case_id":"TR-06","description":"recommendation","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage verwenden.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde eher davon abraten.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"recommendation","normalized_action_text":"advise against using the real asset"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
{"case_id":"TR-07","description":"temporary non-action","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die Waschstufe einbauen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir erstmal noch nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"temporary_non_action","normalized_action_text":"install the washing stage"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
{"case_id":"TR-08","description":"concern only","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten das neue Material einsetzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das wäre kritisch.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"none","normalized_action_text":null},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}}
]
}
@@ -0,0 +1,9 @@
{
"schema_version":"experimental-target-resolution-v1-diagnostic",
"cases":[
{"case_id":"TR1-V1","strategy":"self_contained","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"negative_act":{"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"working with Dr. Schlummer"},"expected":{"target_observation_id":"obs_1","concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":["nicht"]}},
{"case_id":"TR2-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"verfolgen wir nicht weiter"},"expected":{"target_observation_id":"obs_1","concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[]}},
{"case_id":"TR3-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"reale Anlage dafür nicht nutzen"},"expected":{"target_observation_id":"obs_1","concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[]}},
{"case_id":"TR4-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"Versuch in der realen Anlage nicht durchführen"},"expected":{"target_observation_id":"obs_1","concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"]}}
]
}

Some files were not shown because too many files have changed in this diff Show More