Compare commits
6
Commits
1c36fa76fb
...
3918c0b1c4
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
3918c0b1c4 | ||
|
|
7fa771a7e4 | ||
|
|
3229786b5c | ||
|
|
8ca62fbd92 | ||
|
|
0d4b426021 | ||
|
|
97a22f3ebb |
@@ -1811,6 +1811,382 @@ production integration, group identity inference or another semantic category.
|
|||||||
Artifacts are preserved under
|
Artifacts are preserved under
|
||||||
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
|
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
|
## EXP-0034 — Explicit Rejection Gold V0
|
||||||
|
|
||||||
|
Status: Failed architecturally
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
|
||||||
|
This isolated Stage-2 experiment tested the narrow evidence fact that a
|
||||||
|
concrete action, option, proposal or future course was explicitly rejected,
|
||||||
|
abandoned, discontinued or ruled out. It used twelve synthetic cases containing
|
||||||
|
one self-contained observation or one local target/rejection pair. Evidence
|
||||||
|
Observation V3 was not called or changed. The accepted Request/Acceptance and
|
||||||
|
Collective Commitment paths remained unchanged and were not invoked.
|
||||||
|
|
||||||
|
The strict semantic schema contains exactly `rejection_observation_id`,
|
||||||
|
`target_observation_id`, `rejection_form` and
|
||||||
|
`normalized_rejected_action_text`. `rejection_form` is closed to
|
||||||
|
`explicit_action_rejection` and `none`. A positive recognition requires a
|
||||||
|
known local target and non-empty normalized target; `none` requires both target
|
||||||
|
and normalized text to be null. Decision, outcome, topic-closure,
|
||||||
|
responsibility, ownership, protocol, confidence and graph fields are forbidden.
|
||||||
|
Target resolution is limited to the same observation or one earlier supplied
|
||||||
|
observation. Deterministic code validates schema, IDs, ordering and complete
|
||||||
|
provenance before emitting the narrow status `explicitly_rejected`.
|
||||||
|
|
||||||
|
`explicitly_rejected` means rejected by the cited evidence only. It is not yet
|
||||||
|
a final meeting decision or final topic outcome, does not close a topic, and
|
||||||
|
does not supersede an earlier commitment.
|
||||||
|
|
||||||
|
Gold results:
|
||||||
|
|
||||||
|
- RJ-01 explicit collective rejection with local target: PASS.
|
||||||
|
- RJ-02 explicit non-pursuit with paired target: PASS.
|
||||||
|
- RJ-03 self-contained collaboration rejection: FAIL. The model returned
|
||||||
|
`none`, producing one recognition false negative.
|
||||||
|
- RJ-04 personal preference: FAIL. The model promoted the preference to an
|
||||||
|
explicit rejection and derived an unsupported rejection.
|
||||||
|
- RJ-05 concern: PASS; remained a non-rejection.
|
||||||
|
- RJ-06 uncertainty: PASS; remained a non-rejection.
|
||||||
|
- RJ-07 negative recommendation: FAIL. The model promoted advice to an
|
||||||
|
explicit rejection and derived an unsupported rejection.
|
||||||
|
- RJ-08 deferral: PASS; remained a non-rejection.
|
||||||
|
- RJ-09 factual negation: PASS; remained a non-rejection.
|
||||||
|
- RJ-10 temporary non-action: FAIL. The model treated `erstmal noch nicht` as
|
||||||
|
abandonment and derived an unsupported rejection.
|
||||||
|
- RJ-11 explicit rejection with material scope: PASS. Real-plant and
|
||||||
|
Druckversuch scope were preserved.
|
||||||
|
- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option
|
||||||
|
was rejected; the Technikum alternative was not absorbed.
|
||||||
|
|
||||||
|
Configuration: exactly twelve successful sequential `qwen3.5:9B` calls, one
|
||||||
|
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
||||||
|
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
||||||
|
technical failed calls. Aggregate runner time was 15.518 seconds; summed
|
||||||
|
per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681
|
||||||
|
evaluation tokens.
|
||||||
|
|
||||||
|
The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced
|
||||||
|
three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03).
|
||||||
|
There were four strict target-field expectation mismatches: three were
|
||||||
|
consequences of false-positive rejection objects populating otherwise locally
|
||||||
|
correct antecedents, and one was the missing self-contained RJ-03 target. No
|
||||||
|
derived positive selected the wrong concrete antecedent. Qualifier-loss count
|
||||||
|
was zero, positive-alternative absorption count was zero, and no responsibility,
|
||||||
|
decision, outcome or topic-closure field leaked into model output.
|
||||||
|
|
||||||
|
Conclusion: the experiment is not architecturally successful. Deterministic
|
||||||
|
structural gates cannot contain a semantically well-formed false-positive
|
||||||
|
rejection with valid local target and provenance. The model did distinguish
|
||||||
|
concern, uncertainty, deferral and factual negation, and it handled scoped and
|
||||||
|
alternative-bearing positives correctly, but it did not reliably separate
|
||||||
|
explicit rejection from personal preference, advice or temporary non-action.
|
||||||
|
The current binary recognition `explicit_action_rejection | none` is
|
||||||
|
insufficient for reliable generalization.
|
||||||
|
No production integration, generic rejection system, prompt tuning or
|
||||||
|
cross-pattern reconciliation is justified.
|
||||||
|
|
||||||
|
Artifacts are preserved under
|
||||||
|
`artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
|
## EXP-0035 — Negative Act Form V0
|
||||||
|
|
||||||
|
Status: Experimental; successful for form classification with normalization
|
||||||
|
limitations
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
|
||||||
|
EXP-0034 failed because the binary `explicit_action_rejection | none` question
|
||||||
|
collapsed materially different negative acts. It missed self-contained
|
||||||
|
non-pursuit and promoted personal preference, recommendation and temporary
|
||||||
|
non-action to rejection. This isolated follow-up tested only whether those
|
||||||
|
evidence-near forms can be distinguished before any normative derivation. It
|
||||||
|
does not derive rejection, decision, outcome, topic closure, responsibility or
|
||||||
|
protocol status, and EXP-0034 remained unchanged.
|
||||||
|
|
||||||
|
The strict output schema contains exactly `observation_id`,
|
||||||
|
`negative_act_form` and `normalized_action_text`. The closed form vocabulary is
|
||||||
|
`explicit_non_pursuit`, `personal_preference`, `recommendation`,
|
||||||
|
`temporary_non_action` and `none`. Non-`none` forms require non-empty normalized
|
||||||
|
action text; `none` requires null. Rejection, status, decision, outcome,
|
||||||
|
responsibility and other normative fields are forbidden recursively. Local
|
||||||
|
context may resolve a candidate observation's pronoun, but the schema contains
|
||||||
|
no target relation and the experiment exposes no derivation function.
|
||||||
|
|
||||||
|
Gold results:
|
||||||
|
|
||||||
|
- NA-01 explicit non-pursuit: PARTIAL. The form was correct; `working with Dr.
|
||||||
|
Schlummer` omitted the continuation aspect from normalization.
|
||||||
|
- NA-02 paraphrased explicit non-pursuit: PASS.
|
||||||
|
- NA-03 personal preference: PARTIAL. The form was correct, but normalization
|
||||||
|
repeated `Ich würde das nicht machen` instead of resolving the real-plant
|
||||||
|
trial target.
|
||||||
|
- NA-04 negative recommendation: PARTIAL. The form was correct; the normalized
|
||||||
|
English action used the loose rendering `real asset` for `reale Anlage`.
|
||||||
|
- NA-05 temporary non-action: PASS.
|
||||||
|
- NA-06 concern only: PASS with `none` and null action text.
|
||||||
|
- NA-07 uncertainty: PASS with `none` and null action text.
|
||||||
|
- NA-08 factual negation: PASS with `none` and null action text.
|
||||||
|
|
||||||
|
Expected-versus-actual form confusion was entirely diagonal:
|
||||||
|
|
||||||
|
| Expected form | Actual form | Count |
|
||||||
|
| --- | --- | ---: |
|
||||||
|
| `explicit_non_pursuit` | `explicit_non_pursuit` | 2 |
|
||||||
|
| `personal_preference` | `personal_preference` | 1 |
|
||||||
|
| `recommendation` | `recommendation` | 1 |
|
||||||
|
| `temporary_non_action` | `temporary_non_action` | 1 |
|
||||||
|
| `none` | `none` | 3 |
|
||||||
|
|
||||||
|
Configuration: exactly eight successful sequential `qwen3.5:9B` calls, one
|
||||||
|
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
||||||
|
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
||||||
|
technical failures. Aggregate runner time was 8.688 seconds; summed per-call
|
||||||
|
time was 8.686 seconds, with 4,183 prompt-evaluation tokens and 310 evaluation
|
||||||
|
tokens.
|
||||||
|
|
||||||
|
The result was five PASS, three PARTIAL and zero FAIL. All eight
|
||||||
|
`negative_act_form` classifications matched Gold. There was no unsupported
|
||||||
|
semantic strengthening and no rejection, status, decision, outcome,
|
||||||
|
responsibility or topic-closure leakage. Normalized action meaning was fully
|
||||||
|
acceptable in five cases and imperfect in three.
|
||||||
|
|
||||||
|
Conclusion: the finer evidence-near form vocabulary successfully distinguished
|
||||||
|
the four semantic boundaries that defeated the binary rejection experiment in
|
||||||
|
this small Gold set. The result supports separating negative-act-form
|
||||||
|
recognition from later normative derivation, but local target normalization is
|
||||||
|
not yet uniformly reliable. It does not justify modifying EXP-0034, deriving
|
||||||
|
rejection, production integration or beginning cross-pattern reconciliation.
|
||||||
|
|
||||||
|
Artifacts are preserved under
|
||||||
|
`artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
|
## EXP-0036 — Controlled Rejection Derivation V1
|
||||||
|
|
||||||
|
Status: Experimental; architecturally unsuccessful
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
|
||||||
|
This isolated experiment followed the failed binary rejection baseline
|
||||||
|
(EXP-0034) and successful Negative Act Form classification (EXP-0035). Its V1
|
||||||
|
hypothesis was to classify the negative act first, resolve its local target in
|
||||||
|
a separate semantic call, and only then derive `explicitly_rejected`
|
||||||
|
deterministically. It did not modify either predecessor or any accepted Stage-2
|
||||||
|
pattern, and it has no production integration.
|
||||||
|
|
||||||
|
The target recognizer emitted exactly `candidate_observation_id`,
|
||||||
|
`target_observation_id`, and `normalized_target_text`. Only
|
||||||
|
`explicit_non_pursuit` was deterministically eligible. Personal preference,
|
||||||
|
recommendation, temporary non-action, and `none` could never derive rejection,
|
||||||
|
even with a valid target. Provenance, local membership, ordering, non-empty
|
||||||
|
target text, and strict non-normative output were additional gates.
|
||||||
|
`explicitly_rejected` means rejected by the cited evidence only, not a final
|
||||||
|
decision, topic outcome, permanent state, or closure.
|
||||||
|
|
||||||
|
The run reused five exact accepted Negative Act Form outputs and made three new
|
||||||
|
Negative Act calls plus eight target-resolution calls. All calls used
|
||||||
|
`qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||||
|
`num_predict=1024`, no retries, voting, or prompt changes.
|
||||||
|
|
||||||
|
| Case | Negative Act expected / actual | Target result | Verdict |
|
||||||
|
| --- | --- | --- | --- |
|
||||||
|
| CR-01 | `explicit_non_pursuit` / same | Model returned the string `"null"` as an unknown ID; self-contained target was not linked | FAIL |
|
||||||
|
| CR-02 | `explicit_non_pursuit` / same | `obs_1`, external solution and continuation preserved | PASS |
|
||||||
|
| CR-03 | `personal_preference` / same | `obs_1`; eligibility gate prevented rejection | PASS |
|
||||||
|
| CR-04 | `recommendation` / same | `obs_1`; eligibility gate prevented rejection | PASS |
|
||||||
|
| CR-05 | `temporary_non_action` / same | `obs_1`; eligibility gate prevented rejection | PASS |
|
||||||
|
| CR-06 | `none` / same | Null target; final non-rejection was correct, but expected local target was unresolved | FAIL |
|
||||||
|
| CR-07 | `explicit_non_pursuit` / same | `obs_1`; real-plant and pressure-test scope survived, but normalization remained proposition-like | PARTIAL |
|
||||||
|
| CR-08 | `explicit_non_pursuit` / same | `obs_1`; real-plant scope preserved and Technikum alternative excluded | PASS |
|
||||||
|
|
||||||
|
Result: five PASS, one PARTIAL, two FAIL. All eight Negative Act forms were
|
||||||
|
correct. There were no false-positive rejections: the valid targets in CR-03,
|
||||||
|
CR-04, and CR-05 could not override their ineligible forms. There was one
|
||||||
|
false-negative rejection, CR-01, caused by invalid target output. Target
|
||||||
|
resolution missed two expected links (invalid CR-01 and null CR-06), so the
|
||||||
|
wrong/unresolved-target count was two. CR-06 exposed a strategy flaw: a
|
||||||
|
non-eligible Negative Act form should not be required to pass target resolution
|
||||||
|
when it cannot derive rejection. Qualifier-loss count was zero. CR-08 isolated
|
||||||
|
the positive alternative successfully. No individual owner, responsibility,
|
||||||
|
decision, outcome, topic-closure, or LLM-emitted rejection status appeared.
|
||||||
|
|
||||||
|
There were 3 new Negative Act calls, 8 target calls, 5 accepted classification
|
||||||
|
reuses, zero technical call failures, and one structural target-validation
|
||||||
|
failure. Aggregate runner time was 11.951 seconds.
|
||||||
|
|
||||||
|
Conclusion: negative-act-form gating is promising and successfully contains
|
||||||
|
the semantic false positives that defeated EXP-0034, but the experiment is not
|
||||||
|
architecturally successful. The current target-resolution strategy failed the
|
||||||
|
required self-contained positive CR-01 and unnecessarily evaluated the
|
||||||
|
ineligible CR-06 path; it is not reliable enough for rejection derivation.
|
||||||
|
|
||||||
|
Artifacts are preserved under
|
||||||
|
`artifacts/experiments/controlled_rejection_v1/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
|
## EXP-0037 — Target Resolution V0
|
||||||
|
|
||||||
|
Status: Experimental; FAILED for target resolution
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
|
||||||
|
Controlled Rejection V1 showed that fine-grained Negative Act Form eligibility
|
||||||
|
contained false-positive rejection, but its target strategy failed a
|
||||||
|
self-contained positive and unnecessarily resolved a target for an ineligible
|
||||||
|
`none` form. This isolated experiment tested target resolution only. It
|
||||||
|
contains no rejection derivation, status, decision, outcome, responsibility,
|
||||||
|
topic closure, or production integration.
|
||||||
|
|
||||||
|
Eligibility was deterministic: only `explicit_non_pursuit` could reach the
|
||||||
|
resolver. TR-05 personal preference, TR-06 recommendation, TR-07 temporary
|
||||||
|
non-action, and TR-08 `none` stopped before prompt construction and recorded an
|
||||||
|
explicit skipped-call artifact. This hard gate worked in all four cases.
|
||||||
|
|
||||||
|
The target schema contained exactly `candidate_observation_id`,
|
||||||
|
`target_observation_id`, and `normalized_target_text`, with local IDs,
|
||||||
|
same-or-earlier ordering, unique evidence provenance, null consistency, and
|
||||||
|
recursive normative-field exclusion. TR-01 used the self-contained strategy:
|
||||||
|
the prompt stated that linkage was deterministically fixed to the candidate and
|
||||||
|
requested semantic normalization only. TR-02 through TR-04 used paired local
|
||||||
|
resolution. No original transcript or new Negative Act classification call was
|
||||||
|
used.
|
||||||
|
|
||||||
|
| Case | Form / eligible | Call | Target result | Verdict |
|
||||||
|
| --- | --- | --- | --- | --- |
|
||||||
|
| TR-01 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; required same-observation target unresolved | FAIL |
|
||||||
|
| TR-02 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; `obs_1` unresolved | FAIL |
|
||||||
|
| TR-03 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; scoped `obs_1` unresolved | FAIL |
|
||||||
|
| TR-04 | `explicit_non_pursuit` / yes | yes | Returned string `"null"`; real-plant target unresolved | FAIL |
|
||||||
|
| TR-05 | `personal_preference` / no | no | Deterministically skipped | PASS |
|
||||||
|
| TR-06 | `recommendation` / no | no | Deterministically skipped | PASS |
|
||||||
|
| TR-07 | `temporary_non_action` / no | no | Deterministically skipped | PASS |
|
||||||
|
| TR-08 | `none` / no | no | Deterministically skipped | PASS |
|
||||||
|
|
||||||
|
Result: four PASS, zero PARTIAL, four FAIL. Exactly four successful Ollama
|
||||||
|
calls were made, all for eligible cases; there were zero technical call
|
||||||
|
failures and four structural validation failures. All four raw responses used
|
||||||
|
the JSON string `"null"` as target ID rather than a supplied observation ID or
|
||||||
|
JSON null. Wrong-target count and unresolved-target count were therefore four.
|
||||||
|
No qualifier-preservation claim can be made because no eligible positive target
|
||||||
|
passed validation. TR-04 alternative isolation likewise could not be
|
||||||
|
established. No rejection, status, decision, outcome, responsibility, or other
|
||||||
|
normative leakage occurred, and no rejection derivation was performed.
|
||||||
|
|
||||||
|
Configuration: `qwen3.5:9B`, temperature 0, `think=false`,
|
||||||
|
`num_ctx=16384`, `num_predict=1024`, no retries, voting, or prompt changes.
|
||||||
|
Aggregate runner time was 4.034 seconds.
|
||||||
|
|
||||||
|
Conclusion: eligibility gating is successful and should be retained; it fully
|
||||||
|
prevents unnecessary target calls for ineligible Negative Act forms. Target
|
||||||
|
Resolution V0 itself failed structurally across all eligible cases. Neither the
|
||||||
|
self-contained nor paired strategy produced a valid target, and merely
|
||||||
|
instructing deterministic self-linkage in the semantic prompt did not make the
|
||||||
|
linkage structurally deterministic. The repeated `"null"` string pattern
|
||||||
|
requires diagnosis before changing the architecture or prompt. No rejection
|
||||||
|
derivation is justified by this result.
|
||||||
|
|
||||||
|
Artifacts are preserved under
|
||||||
|
`artifacts/experiments/target_resolution_v0/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
|
## EXP-0038 — Target Resolution V1 Diagnostic
|
||||||
|
|
||||||
|
Status: Experimental; linkage boundary successful, normalization incomplete
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
|
||||||
|
Forensics on failed Target Resolution V0 found a definite prompt defect: its
|
||||||
|
illustrative value `"observation ID or null"` placed both alternatives inside
|
||||||
|
a JSON string. V0 also sent only `format: "json"`, which enforced JSON syntax
|
||||||
|
but not field types. This isolated diagnostic changed only the linkage/output
|
||||||
|
boundary. It contains no rejection derivation or normative semantics.
|
||||||
|
|
||||||
|
Ollama 0.32.6 accepted a true JSON Schema object in `format`. TR1-V1 removed
|
||||||
|
target selection from the model output entirely and deterministically linked
|
||||||
|
the self-contained candidate to itself. TR2-V1 through TR4-V1 used a closed
|
||||||
|
allowed-ID list, an enum of those IDs plus JSON null, typed positive and null
|
||||||
|
examples, recursive strict validation, and one fixed paired prompt. Linkage and
|
||||||
|
normalization were persisted separately.
|
||||||
|
|
||||||
|
| Case | Strategy / ID source | Target | Normalized target | Verdict |
|
||||||
|
| --- | --- | --- | --- | --- |
|
||||||
|
| TR1-V1 | self-contained / deterministic | `obs_1` | `Mit Dr. Schlummer arbeiten wir nicht weiter.` retained negation instead of a positive action meaning | FAIL |
|
||||||
|
| TR2-V1 | paired / LLM | `obs_1` | `externe Lösung weiterverfolgen` | PASS |
|
||||||
|
| TR3-V1 | paired / LLM | `obs_1` | `reale Anlage zur Diskussion` lost `Druckversuch` purpose and the `nutzen` action | FAIL |
|
||||||
|
| TR4-V1 | paired / LLM | `obs_1` | `Versuch in der realen Anlage durchführen`; Technikum excluded | PASS |
|
||||||
|
|
||||||
|
Result: two PASS, zero PARTIAL, two FAIL. All four responses passed their true
|
||||||
|
JSON Schemas. Every resulting target was `obs_1`; wrong-target and
|
||||||
|
unresolved-target counts were zero. The string `"null"` recurrence count was
|
||||||
|
zero, and there were zero structural validation failures. TR1 preserved the
|
||||||
|
collaboration, person, and continuation wording but failed positive-action
|
||||||
|
normalization by retaining negation. TR3 had one material scope loss. TR4
|
||||||
|
preserved real-plant scope and isolated the Technikum alternative. No
|
||||||
|
normative leakage occurred.
|
||||||
|
|
||||||
|
Configuration: exactly four `qwen3.5:9B` calls, temperature 0,
|
||||||
|
`think=false`, `num_ctx=16384`, `num_predict=1024`, no retries, voting, or
|
||||||
|
prompt tuning. Aggregate runner time was 4.557 seconds.
|
||||||
|
|
||||||
|
Conclusion: the V0 string-null failure was primarily a linkage/output-boundary
|
||||||
|
failure rather than evidence that observation-ID linkage is semantically
|
||||||
|
impossible. True typed schemas, closed ID lists, and deterministic self-linkage
|
||||||
|
eliminated every structural and target-ID failure. The experiment still fails
|
||||||
|
its complete acceptance criterion because target normalization is not reliably
|
||||||
|
positive or scope-preserving. These results justify separating linkage from
|
||||||
|
normalization, but not deriving rejection or integrating a new pipeline.
|
||||||
|
|
||||||
|
Artifacts are preserved under
|
||||||
|
`artifacts/experiments/target_resolution_v1_diagnostic/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
|
## EXP-0039 — Target Normalization V0
|
||||||
|
|
||||||
|
Status: Experimental; normalization improved but incomplete
|
||||||
|
|
||||||
|
Date: 2026-08-20
|
||||||
|
|
||||||
|
Target Resolution V1 established correct linkage for all four narrow cases and
|
||||||
|
eliminated structural ID failures with deterministic self-linkage, closed ID
|
||||||
|
lists, and true JSON Schemas. Its remaining failures were normalization-only.
|
||||||
|
This isolated follow-up therefore accepted candidate and target IDs as fixed
|
||||||
|
input and tested only reconstruction of the positive German action meaning. It
|
||||||
|
contains no target selection, Negative Act classification, eligibility logic,
|
||||||
|
rejection derivation, or production integration.
|
||||||
|
|
||||||
|
The strict output schema contained exactly `candidate_observation_id`,
|
||||||
|
`target_observation_id`, and `normalized_target_text`. Both IDs were constrained
|
||||||
|
to their supplied values with JSON Schema `const`; normalized text was a
|
||||||
|
non-empty string and null was disallowed. The one fixed prompt required removal
|
||||||
|
of negative polarity, preservation of action, continuation, material scope and
|
||||||
|
source language, and exclusion of separate alternatives.
|
||||||
|
|
||||||
|
| Case | Actual normalized target | Verdict |
|
||||||
|
| --- | --- | --- |
|
||||||
|
| TN-01 | `Mit Dr. Schlummer zusammenarbeiten` | FAIL: positive polarity and collaboration survived, but continuation was lost |
|
||||||
|
| TN-02 | `externe Lösung weiterverfolgen` | PASS |
|
||||||
|
| TN-03 | `reale Anlage für den Druckversuch nutzen` | PASS |
|
||||||
|
| TN-04 | `Versuch in der realen Anlage durchführen` | PASS; Technikum alternative excluded |
|
||||||
|
|
||||||
|
Result: three PASS, zero PARTIAL, one FAIL. All four outputs passed strict
|
||||||
|
schema validation and copied both fixed IDs exactly, so changed-ID count was
|
||||||
|
zero. Polarity-error count was zero: even TN-01 removed rejection and negation.
|
||||||
|
Action/continuation-loss count was one (TN-01); material purpose/location
|
||||||
|
scope-loss count was zero; alternative-absorption count was zero. There was no
|
||||||
|
unsupported strengthening or normative leakage.
|
||||||
|
|
||||||
|
Configuration: exactly four `qwen3.5:9B` calls, temperature 0,
|
||||||
|
`think=false`, true JSON Schema, `num_ctx=16384`, `num_predict=1024`, no
|
||||||
|
retries, voting, or prompt tuning. Aggregate runner time was 5.066 seconds.
|
||||||
|
|
||||||
|
Conclusion: isolating normalization solved the polarity and scoped-action
|
||||||
|
failures seen in Target Resolution V1 for three of four cases, including exact
|
||||||
|
pressure-test scope and alternative isolation. Continuation semantics remain
|
||||||
|
unreliable in the self-contained collaboration case, so Target Normalization
|
||||||
|
V0 does not meet its full acceptance criterion. The result does not justify
|
||||||
|
rejection derivation or production integration.
|
||||||
|
|
||||||
|
Artifacts are preserved under
|
||||||
|
`artifacts/experiments/target_normalization_v0/20260820_qwen35_9b_single_run/`.
|
||||||
|
|
||||||
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
|
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
|
||||||
|
|
||||||
Date: 2026-08-11
|
Date: 2026-08-11
|
||||||
|
|||||||
@@ -0,0 +1,7 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
if str(ROOT) not in sys.path: sys.path.insert(0, str(ROOT))
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection_v1 import main
|
||||||
|
if __name__ == "__main__": raise SystemExit(main())
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Repository entry point for the explicit-rejection Gold experiment."""
|
||||||
|
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
if str(REPO_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(REPO_ROOT))
|
||||||
|
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import main # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,16 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Repository entry point for the Negative Act Form experiment."""
|
||||||
|
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
REPO_ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
if str(REPO_ROOT) not in sys.path:
|
||||||
|
sys.path.insert(0, str(REPO_ROOT))
|
||||||
|
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_negative_act import main # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
ROOT=Path(__file__).resolve().parents[1]
|
||||||
|
if str(ROOT) not in sys.path: sys.path.insert(0,str(ROOT))
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_target_normalization import main
|
||||||
|
if __name__=="__main__": raise SystemExit(main())
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
ROOT = Path(__file__).resolve().parents[1]
|
||||||
|
if str(ROOT) not in sys.path: sys.path.insert(0, str(ROOT))
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution import main
|
||||||
|
if __name__ == "__main__": raise SystemExit(main())
|
||||||
@@ -0,0 +1,7 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
ROOT=Path(__file__).resolve().parents[1]
|
||||||
|
if str(ROOT) not in sys.path: sys.path.insert(0,str(ROOT))
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution_v1 import main
|
||||||
|
if __name__=="__main__": raise SystemExit(main())
|
||||||
@@ -0,0 +1,277 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Isolated evidence-near Negative Act Form classification experiment."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import time
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from .experiment_h import (
|
||||||
|
DEFAULT_ENDPOINT,
|
||||||
|
DEFAULT_MODEL,
|
||||||
|
DerivationValidationError,
|
||||||
|
OBSERVATION_KEYS,
|
||||||
|
build_ollama_payload,
|
||||||
|
call_ollama,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
GOLD_SCHEMA_VERSION = "experimental-negative-act-form-gold-v0"
|
||||||
|
RECOGNITION_KEYS = {"observation_id", "negative_act_form", "normalized_action_text"}
|
||||||
|
NEGATIVE_ACT_FORMS = {
|
||||||
|
"explicit_non_pursuit", "personal_preference", "recommendation",
|
||||||
|
"temporary_non_action", "none",
|
||||||
|
}
|
||||||
|
FORBIDDEN_LLM_KEYS = {
|
||||||
|
"rejection_form", "explicitly_rejected", "status", "decision", "outcome",
|
||||||
|
"topic_status", "responsible_person", "responsibility", "owner",
|
||||||
|
"requested_actor", "action_item", "protocol", "protocol_category",
|
||||||
|
"confidence", "relation", "relations", "graph", "unresolved_issue",
|
||||||
|
}
|
||||||
|
|
||||||
|
PROMPT_TEMPLATE = """Classify only the negative semantic form expressed by the candidate observation, using earlier supplied V3-style observations only as local context for pronouns or shortened references.
|
||||||
|
|
||||||
|
The candidate observation is {candidate_observation_id}.
|
||||||
|
|
||||||
|
Choose exactly one negative_act_form:
|
||||||
|
- explicit_non_pursuit: explicitly states that an action, option, collaboration, or course will not be continued or pursued. This is stronger than preference, advice, or temporary delay.
|
||||||
|
- personal_preference: the speaker states what they personally would or would not do, without establishing collective non-pursuit.
|
||||||
|
- recommendation: the speaker advises for or against an action without establishing abandonment.
|
||||||
|
- temporary_non_action: the action is postponed, deferred, or explicitly not done for now without abandonment.
|
||||||
|
- none: none of those four forms is present, including mere concern, uncertainty, negative sentiment, or factual negation.
|
||||||
|
|
||||||
|
Do not collapse non-pursuit into temporary non-action. Do not convert a personal conditional preference into collective non-pursuit. Do not convert advice into non-pursuit. Speaker identity does not change personal preference into collective non-pursuit.
|
||||||
|
|
||||||
|
When the form is not none, return concise normalized action meaning. Resolve a pronoun only from the supplied local context. If its target is genuinely ambiguous, return none rather than guessing. When the form is none, normalized_action_text must be null. Keep normalized action text in the observation language.
|
||||||
|
|
||||||
|
Do not derive or output rejection, status, decision, outcome, topic closure, responsibility, ownership, Action Item, protocol category, confidence, relations, graphs, or unresolved issues.
|
||||||
|
|
||||||
|
Return exactly this JSON shape and no additional fields:
|
||||||
|
{{
|
||||||
|
"observation_id": "{candidate_observation_id}",
|
||||||
|
"negative_act_form": "explicit_non_pursuit | personal_preference | recommendation | temporary_non_action | none",
|
||||||
|
"normalized_action_text": "concise action meaning" | null
|
||||||
|
}}
|
||||||
|
|
||||||
|
V3-style observations:
|
||||||
|
{observations_json}
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||||
|
missing = required - value.keys()
|
||||||
|
unknown = value.keys() - required
|
||||||
|
if missing:
|
||||||
|
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||||
|
if unknown:
|
||||||
|
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||||
|
|
||||||
|
|
||||||
|
def _nonempty_text(value: Any, location: str) -> str:
|
||||||
|
if not isinstance(value, str) or not value.strip():
|
||||||
|
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||||
|
return value.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_observations(observations: Any) -> None:
|
||||||
|
if not isinstance(observations, list) or not observations:
|
||||||
|
raise DerivationValidationError("observations must be a non-empty list")
|
||||||
|
seen_observations: set[str] = set()
|
||||||
|
seen_evidence: set[str] = set()
|
||||||
|
for index, observation in enumerate(observations):
|
||||||
|
location = f"observations[{index}]"
|
||||||
|
if not isinstance(observation, dict):
|
||||||
|
raise DerivationValidationError(f"{location} must be an object")
|
||||||
|
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||||
|
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
|
||||||
|
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
|
||||||
|
if observation_id in seen_observations or evidence_id in seen_evidence:
|
||||||
|
raise DerivationValidationError("observation and evidence provenance must be unique")
|
||||||
|
seen_observations.add(observation_id)
|
||||||
|
seen_evidence.add(evidence_id)
|
||||||
|
_nonempty_text(observation["content"], f"{location}.content")
|
||||||
|
_nonempty_text(observation["speaker"], f"{location}.speaker")
|
||||||
|
for field in ("named_person", "addressee"):
|
||||||
|
if observation[field] is not None:
|
||||||
|
_nonempty_text(observation[field], f"{location}.{field}")
|
||||||
|
|
||||||
|
|
||||||
|
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
|
||||||
|
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise DerivationValidationError("Gold fixture must be an object")
|
||||||
|
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
|
||||||
|
if data["schema_version"] != GOLD_SCHEMA_VERSION:
|
||||||
|
raise DerivationValidationError("unexpected Gold fixture schema_version")
|
||||||
|
cases = data["cases"]
|
||||||
|
if not isinstance(cases, list) or not cases:
|
||||||
|
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
|
||||||
|
seen: set[str] = set()
|
||||||
|
for case in cases:
|
||||||
|
_exact_keys(case, {"case_id", "description", "observations", "expected"}, "Gold case")
|
||||||
|
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
|
||||||
|
if case_id in seen:
|
||||||
|
raise DerivationValidationError(f"duplicate case ID: {case_id}")
|
||||||
|
seen.add(case_id)
|
||||||
|
_validate_observations(case["observations"])
|
||||||
|
if len(case["observations"]) not in (1, 2):
|
||||||
|
raise DerivationValidationError("Negative Act cases require one or two observations")
|
||||||
|
return cases
|
||||||
|
|
||||||
|
|
||||||
|
def build_prompt(case: dict[str, Any]) -> str:
|
||||||
|
observations = case["observations"]
|
||||||
|
_validate_observations(observations)
|
||||||
|
candidate_id = observations[-1]["observation_id"]
|
||||||
|
return PROMPT_TEMPLATE.format(
|
||||||
|
candidate_observation_id=candidate_id,
|
||||||
|
observations_json=json.dumps(observations, ensure_ascii=False, indent=2),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||||
|
data = json.loads(raw_text)
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise DerivationValidationError("semantic classification must be an object")
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||||
|
if forbidden:
|
||||||
|
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
|
||||||
|
for key, item in value.items():
|
||||||
|
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||||
|
elif isinstance(value, list):
|
||||||
|
for index, item in enumerate(value):
|
||||||
|
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||||
|
|
||||||
|
|
||||||
|
def validate_classification(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||||
|
_validate_observations(observations)
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise DerivationValidationError("semantic classification must be an object")
|
||||||
|
_reject_forbidden_keys(data)
|
||||||
|
_exact_keys(data, RECOGNITION_KEYS, "output")
|
||||||
|
observation_id = _nonempty_text(data["observation_id"], "output.observation_id")
|
||||||
|
if observation_id not in {item["observation_id"] for item in observations}:
|
||||||
|
raise DerivationValidationError("classification references unknown observation")
|
||||||
|
form = data["negative_act_form"]
|
||||||
|
if form not in NEGATIVE_ACT_FORMS:
|
||||||
|
raise DerivationValidationError("negative_act_form has an unsupported value")
|
||||||
|
action_text = data["normalized_action_text"]
|
||||||
|
if form == "none":
|
||||||
|
if action_text is not None:
|
||||||
|
raise DerivationValidationError("none form requires null normalized_action_text")
|
||||||
|
else:
|
||||||
|
_nonempty_text(action_text, "output.normalized_action_text")
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
|
||||||
|
if not concepts:
|
||||||
|
return text is None
|
||||||
|
if not isinstance(text, str):
|
||||||
|
return False
|
||||||
|
folded = text.casefold()
|
||||||
|
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
|
||||||
|
|
||||||
|
|
||||||
|
def evaluate_case(case: dict[str, Any], classification: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
expected = case["expected"]
|
||||||
|
observation_correct = classification["observation_id"] == expected["observation_id"]
|
||||||
|
form_correct = classification["negative_act_form"] == expected["negative_act_form"]
|
||||||
|
action_correct = _concepts_present(classification["normalized_action_text"], expected["action_concepts"])
|
||||||
|
unsupported_strengthening = expected["negative_act_form"] == "none" and classification["negative_act_form"] != "none"
|
||||||
|
classification_label = "PASS" if observation_correct and form_correct and action_correct else ("PARTIAL" if observation_correct and form_correct else "FAIL")
|
||||||
|
return {
|
||||||
|
"case_id": case["case_id"], "classification": classification_label,
|
||||||
|
"expected_negative_act_form": expected["negative_act_form"],
|
||||||
|
"actual_negative_act_form": classification["negative_act_form"],
|
||||||
|
"observation_id_correct": observation_correct,
|
||||||
|
"normalized_action_meaning_correct": action_correct,
|
||||||
|
"unsupported_semantic_strengthening": unsupported_strengthening,
|
||||||
|
"normative_leakage": False,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _write_json(path: Path, value: Any) -> None:
|
||||||
|
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def run_experiment(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
cases = load_gold_cases(args.cases)
|
||||||
|
args.output.mkdir(parents=True, exist_ok=False)
|
||||||
|
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
|
||||||
|
evaluations: list[dict[str, Any]] = []
|
||||||
|
successful_calls = 0
|
||||||
|
technical_failures = 0
|
||||||
|
started = time.perf_counter()
|
||||||
|
for case in cases:
|
||||||
|
case_dir = args.output / case["case_id"].lower()
|
||||||
|
case_dir.mkdir()
|
||||||
|
observations = case["observations"]
|
||||||
|
_write_json(case_dir / "v3_style_input_observations.json", observations)
|
||||||
|
prompt = build_prompt(case)
|
||||||
|
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||||
|
try:
|
||||||
|
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||||
|
successful_calls += 1
|
||||||
|
except Exception as exc: # one recorded attempt; never retry
|
||||||
|
technical_failures += 1
|
||||||
|
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
|
||||||
|
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
|
||||||
|
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
|
||||||
|
_write_json(case_dir / "evaluation.json", failure)
|
||||||
|
evaluations.append(failure)
|
||||||
|
continue
|
||||||
|
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||||
|
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||||
|
try:
|
||||||
|
parsed = parse_model_json(raw)
|
||||||
|
_write_json(case_dir / "parsed_semantic_classification.json", parsed)
|
||||||
|
evaluation = evaluate_case(case, parsed)
|
||||||
|
validation = {"valid": True, "error": None}
|
||||||
|
except (DerivationValidationError, json.JSONDecodeError) as exc:
|
||||||
|
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||||
|
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)}
|
||||||
|
_write_json(case_dir / "structural_validation.json", validation)
|
||||||
|
_write_json(case_dir / "evaluation.json", evaluation)
|
||||||
|
evaluations.append(evaluation)
|
||||||
|
summary = {
|
||||||
|
"experiment": "negative_act_form_v0", "model": args.model,
|
||||||
|
"successful_llm_call_count": successful_calls,
|
||||||
|
"technical_failed_call_count": technical_failures,
|
||||||
|
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||||
|
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||||
|
"evaluations": evaluations,
|
||||||
|
}
|
||||||
|
_write_json(args.output / "summary.json", summary)
|
||||||
|
return summary
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser(description="Run isolated Negative Act Form experiment")
|
||||||
|
parser.add_argument("cases", type=Path)
|
||||||
|
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||||
|
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||||
|
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||||
|
parser.add_argument("--timeout", type=int, default=300)
|
||||||
|
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||||
|
parser.add_argument("--num-predict", type=int, default=1024)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
summary = run_experiment(parse_args())
|
||||||
|
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||||
|
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,347 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Isolated explicit-action-rejection Gold reliability experiment."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import time
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from .experiment_h import (
|
||||||
|
DEFAULT_ENDPOINT,
|
||||||
|
DEFAULT_MODEL,
|
||||||
|
DerivationValidationError,
|
||||||
|
OBSERVATION_KEYS,
|
||||||
|
call_ollama,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
GOLD_SCHEMA_VERSION = "experimental-explicit-rejection-gold-v0"
|
||||||
|
RECOGNITION_KEYS = {
|
||||||
|
"rejection_observation_id", "target_observation_id", "rejection_form",
|
||||||
|
"normalized_rejected_action_text",
|
||||||
|
}
|
||||||
|
REJECTION_FORMS = {"explicit_action_rejection", "none"}
|
||||||
|
FORBIDDEN_LLM_KEYS = {
|
||||||
|
"decision", "decision_status", "outcome", "topic_status", "closed",
|
||||||
|
"agreement", "responsible_person", "responsibility", "responsibility_scope",
|
||||||
|
"owner", "ownership", "assignee", "requested_actor", "status",
|
||||||
|
"explicitly_rejected", "action_item", "protocol", "protocol_category",
|
||||||
|
"confidence", "relation", "relations", "graph", "unresolved_issue",
|
||||||
|
}
|
||||||
|
|
||||||
|
PROMPT_TEMPLATE = """Recognize only whether the candidate rejection observation explicitly rejects a concrete action, option, proposal, or future course of action in this small local set of V3-style observations.
|
||||||
|
|
||||||
|
Answer only:
|
||||||
|
1. Does the candidate rejection observation explicitly reject, abandon, discontinue, or rule out a concrete action, option, proposal, or future course of action?
|
||||||
|
2. If yes, which supplied observation identifies the rejected target?
|
||||||
|
3. What is the concise normalized meaning of the rejected action or option?
|
||||||
|
|
||||||
|
The candidate rejection observation is {rejection_observation_id}.
|
||||||
|
|
||||||
|
Use explicit_action_rejection only for an asserted rejection, abandonment, discontinuation, or non-pursuit with a concrete locally resolvable target. Personal preference is not meeting-level explicit rejection. Concern or objection without refusal is not rejection. Uncertainty is not rejection. Negative recommendation or advice is not established rejection. Deferral is not rejection. "Not yet" or temporary non-action is not abandonment. Factual negation is not action rejection. Lack of commitment is not rejection.
|
||||||
|
|
||||||
|
The rejected target may be self-contained in the candidate observation or introduced by one earlier supplied observation. Choose only among supplied observation IDs. If the target is ambiguous or unresolved, return rejection_form none. Preserve material scope limitations in normalized_rejected_action_text. Ignore a separate positive alternative when describing the rejected target. Keep normalized text in the observation language.
|
||||||
|
|
||||||
|
Do not infer responsibility, ownership, decision status, final outcome, topic closure, protocol status, confidence, relations, graphs, or unresolved issues. Do not answer whether this was finally decided, what the meeting outcome was, who is responsible, or whether the topic is closed.
|
||||||
|
|
||||||
|
Return exactly this JSON shape and no additional fields:
|
||||||
|
{{
|
||||||
|
"rejection_observation_id": "{rejection_observation_id}",
|
||||||
|
"target_observation_id": "supplied observation ID" | null,
|
||||||
|
"rejection_form": "explicit_action_rejection | none",
|
||||||
|
"normalized_rejected_action_text": "concise rejected target" | null
|
||||||
|
}}
|
||||||
|
|
||||||
|
For rejection_form none, target_observation_id and normalized_rejected_action_text must both be null.
|
||||||
|
|
||||||
|
V3-style observations:
|
||||||
|
{observations_json}
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None:
|
||||||
|
missing = required - value.keys()
|
||||||
|
unknown = value.keys() - required
|
||||||
|
if missing:
|
||||||
|
raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}")
|
||||||
|
if unknown:
|
||||||
|
raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}")
|
||||||
|
|
||||||
|
|
||||||
|
def _nonempty_text(value: Any, location: str) -> str:
|
||||||
|
if not isinstance(value, str) or not value.strip():
|
||||||
|
raise DerivationValidationError(f"{location} must be a non-empty string")
|
||||||
|
return value.strip()
|
||||||
|
|
||||||
|
|
||||||
|
def _validate_observations(observations: Any) -> None:
|
||||||
|
if not isinstance(observations, list) or not observations:
|
||||||
|
raise DerivationValidationError("observations must be a non-empty list")
|
||||||
|
seen_observations: set[str] = set()
|
||||||
|
seen_evidence: set[str] = set()
|
||||||
|
for index, observation in enumerate(observations):
|
||||||
|
location = f"observations[{index}]"
|
||||||
|
if not isinstance(observation, dict):
|
||||||
|
raise DerivationValidationError(f"{location} must be an object")
|
||||||
|
_exact_keys(observation, OBSERVATION_KEYS, location)
|
||||||
|
observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id")
|
||||||
|
evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id")
|
||||||
|
if observation_id in seen_observations:
|
||||||
|
raise DerivationValidationError("observation IDs must be unique")
|
||||||
|
if evidence_id in seen_evidence:
|
||||||
|
raise DerivationValidationError("evidence provenance must be unique and consistent")
|
||||||
|
seen_observations.add(observation_id)
|
||||||
|
seen_evidence.add(evidence_id)
|
||||||
|
_nonempty_text(observation["content"], f"{location}.content")
|
||||||
|
_nonempty_text(observation["speaker"], f"{location}.speaker")
|
||||||
|
for field in ("named_person", "addressee"):
|
||||||
|
if observation[field] is not None:
|
||||||
|
_nonempty_text(observation[field], f"{location}.{field}")
|
||||||
|
|
||||||
|
|
||||||
|
def load_gold_cases(path: Path) -> list[dict[str, Any]]:
|
||||||
|
data = json.loads(path.read_text(encoding="utf-8-sig"))
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise DerivationValidationError("Gold fixture must be an object")
|
||||||
|
_exact_keys(data, {"schema_version", "cases"}, "Gold fixture")
|
||||||
|
if data["schema_version"] != GOLD_SCHEMA_VERSION:
|
||||||
|
raise DerivationValidationError("unexpected Gold fixture schema_version")
|
||||||
|
cases = data["cases"]
|
||||||
|
if not isinstance(cases, list) or not cases:
|
||||||
|
raise DerivationValidationError("Gold fixture cases must be a non-empty list")
|
||||||
|
seen: set[str] = set()
|
||||||
|
for case in cases:
|
||||||
|
_exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case")
|
||||||
|
case_id = _nonempty_text(case["case_id"], "Gold case.case_id")
|
||||||
|
if case_id in seen:
|
||||||
|
raise DerivationValidationError(f"duplicate case ID: {case_id}")
|
||||||
|
seen.add(case_id)
|
||||||
|
_validate_observations(case["observations"])
|
||||||
|
if len(case["observations"]) not in (1, 2):
|
||||||
|
raise DerivationValidationError("rejection Gold cases require one or two observations")
|
||||||
|
return cases
|
||||||
|
|
||||||
|
|
||||||
|
def build_prompt(case: dict[str, Any]) -> str:
|
||||||
|
observations = case["observations"]
|
||||||
|
_validate_observations(observations)
|
||||||
|
rejection_observation_id = observations[-1]["observation_id"]
|
||||||
|
return PROMPT_TEMPLATE.format(
|
||||||
|
rejection_observation_id=rejection_observation_id,
|
||||||
|
observations_json=json.dumps(observations, ensure_ascii=False, indent=2),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_model_json(raw_text: str) -> dict[str, Any]:
|
||||||
|
data = json.loads(raw_text)
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise DerivationValidationError("semantic recognition must be an object")
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def _reject_forbidden_keys(value: Any, location: str = "output") -> None:
|
||||||
|
if isinstance(value, dict):
|
||||||
|
forbidden = FORBIDDEN_LLM_KEYS.intersection(value)
|
||||||
|
if forbidden:
|
||||||
|
raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}")
|
||||||
|
for key, item in value.items():
|
||||||
|
_reject_forbidden_keys(item, f"{location}.{key}")
|
||||||
|
elif isinstance(value, list):
|
||||||
|
for index, item in enumerate(value):
|
||||||
|
_reject_forbidden_keys(item, f"{location}[{index}]")
|
||||||
|
|
||||||
|
|
||||||
|
def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]:
|
||||||
|
_validate_observations(observations)
|
||||||
|
if not isinstance(data, dict):
|
||||||
|
raise DerivationValidationError("semantic recognition must be an object")
|
||||||
|
_reject_forbidden_keys(data)
|
||||||
|
_exact_keys(data, RECOGNITION_KEYS, "output")
|
||||||
|
rejection_id = _nonempty_text(data["rejection_observation_id"], "output.rejection_observation_id")
|
||||||
|
known_ids = {item["observation_id"] for item in observations}
|
||||||
|
if rejection_id not in known_ids:
|
||||||
|
raise DerivationValidationError("unknown rejection observation ID")
|
||||||
|
form = data["rejection_form"]
|
||||||
|
if form not in REJECTION_FORMS:
|
||||||
|
raise DerivationValidationError("rejection_form has an unsupported value")
|
||||||
|
target_id = data["target_observation_id"]
|
||||||
|
action_text = data["normalized_rejected_action_text"]
|
||||||
|
if form == "none":
|
||||||
|
if target_id is not None:
|
||||||
|
raise DerivationValidationError("none rejection must have null target_observation_id")
|
||||||
|
if action_text is not None:
|
||||||
|
raise DerivationValidationError("none rejection must have null normalized_rejected_action_text")
|
||||||
|
else:
|
||||||
|
target_id = _nonempty_text(target_id, "output.target_observation_id")
|
||||||
|
if target_id not in known_ids:
|
||||||
|
raise DerivationValidationError("unknown target observation ID")
|
||||||
|
_nonempty_text(action_text, "output.normalized_rejected_action_text")
|
||||||
|
return data
|
||||||
|
|
||||||
|
|
||||||
|
def derive_rejection(
|
||||||
|
observations: list[dict[str, Any]], recognition: dict[str, Any]
|
||||||
|
) -> tuple[dict[str, bool], dict[str, Any] | None]:
|
||||||
|
validate_recognition(recognition, observations)
|
||||||
|
by_id = {item["observation_id"]: item for item in observations}
|
||||||
|
positions = {item["observation_id"]: index for index, item in enumerate(observations)}
|
||||||
|
rejection = by_id.get(recognition["rejection_observation_id"])
|
||||||
|
target_id = recognition["target_observation_id"]
|
||||||
|
target = by_id.get(target_id) if target_id is not None else None
|
||||||
|
gates = {
|
||||||
|
"recognition_schema_valid": True,
|
||||||
|
"explicit_action_rejection": recognition["rejection_form"] == "explicit_action_rejection",
|
||||||
|
"rejection_observation_exists": rejection is not None,
|
||||||
|
"target_observation_exists": target is not None,
|
||||||
|
"observation_ids_valid_and_unique": len(by_id) == len(observations),
|
||||||
|
"evidence_provenance_valid_unique_consistent": len({item["evidence_id"] for item in observations}) == len(observations),
|
||||||
|
"target_same_or_before_rejection": target is not None and rejection is not None and positions[target["observation_id"]] <= positions[rejection["observation_id"]],
|
||||||
|
"normalized_rejected_action_present": isinstance(recognition["normalized_rejected_action_text"], str) and bool(recognition["normalized_rejected_action_text"].strip()),
|
||||||
|
"target_local_to_case": target_id in by_id if target_id is not None else False,
|
||||||
|
"schema_state_consistent": recognition["rejection_form"] == "explicit_action_rejection" and target_id is not None,
|
||||||
|
"referenced_provenance_available": target is not None and rejection is not None and bool(target["evidence_id"]) and bool(rejection["evidence_id"]),
|
||||||
|
}
|
||||||
|
if not all(gates.values()):
|
||||||
|
return gates, None
|
||||||
|
return gates, {
|
||||||
|
"rejection_id": "rejection_1",
|
||||||
|
"content": recognition["normalized_rejected_action_text"].strip(),
|
||||||
|
"status": "explicitly_rejected",
|
||||||
|
"support": {
|
||||||
|
"target": {"observation_id": target["observation_id"], "evidence_id": target["evidence_id"]},
|
||||||
|
"rejection": {"observation_id": rejection["observation_id"], "evidence_id": rejection["evidence_id"]},
|
||||||
|
},
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool:
|
||||||
|
if not concepts:
|
||||||
|
return True
|
||||||
|
if not isinstance(text, str):
|
||||||
|
return False
|
||||||
|
folded = text.casefold()
|
||||||
|
return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts)
|
||||||
|
|
||||||
|
|
||||||
|
def _contains_forbidden_concept(text: str | None, concepts: list[str]) -> bool:
|
||||||
|
return isinstance(text, str) and any(concept.casefold() in text.casefold() for concept in concepts)
|
||||||
|
|
||||||
|
|
||||||
|
def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]:
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
gates, result = derive_rejection(case["observations"], recognition)
|
||||||
|
expected = case["expected_recognition"]
|
||||||
|
expected_result = case["expected_result"]
|
||||||
|
form_correct = recognition["rejection_form"] == expected["rejection_form"]
|
||||||
|
rejection_observation_correct = recognition["rejection_observation_id"] == expected["rejection_observation_id"]
|
||||||
|
target_correct = recognition["target_observation_id"] == expected["target_observation_id"]
|
||||||
|
action_correct = _concepts_present(recognition["normalized_rejected_action_text"], expected["action_concepts"])
|
||||||
|
qualifier_preserved = _concepts_present(recognition["normalized_rejected_action_text"], expected["qualifier_concepts"])
|
||||||
|
alternative_absorbed = _contains_forbidden_concept(recognition["normalized_rejected_action_text"], expected["forbidden_action_concepts"])
|
||||||
|
derived = result is not None
|
||||||
|
final_correct = derived == expected_result["explicitly_rejected"]
|
||||||
|
if result is not None:
|
||||||
|
final_correct = final_correct and result["status"] == "explicitly_rejected"
|
||||||
|
semantic_correct = form_correct and rejection_observation_correct and target_correct and action_correct and qualifier_preserved and not alternative_absorbed
|
||||||
|
automatic_failure = (derived and not expected_result["explicitly_rejected"]) or (derived and not target_correct) or (derived and not qualifier_preserved) or alternative_absorbed
|
||||||
|
classification = "FAIL" if automatic_failure or not final_correct else ("PASS" if semantic_correct else "PARTIAL")
|
||||||
|
return {
|
||||||
|
"case_id": case["case_id"], "classification": classification,
|
||||||
|
"rejection_form_correct": form_correct,
|
||||||
|
"rejection_observation_correct": rejection_observation_correct,
|
||||||
|
"target_observation_correct": target_correct,
|
||||||
|
"normalized_rejected_action_correct": action_correct,
|
||||||
|
"material_qualifiers_preserved": qualifier_preserved,
|
||||||
|
"positive_alternative_absorbed": alternative_absorbed,
|
||||||
|
"deterministic_gates_correct": final_correct,
|
||||||
|
"final_result_correct": final_correct,
|
||||||
|
"unsupported_semantic_strengthening": recognition["rejection_form"] == "explicit_action_rejection" and expected["rejection_form"] == "none",
|
||||||
|
"normative_leakage": False,
|
||||||
|
"gates": gates, "result": result,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _write_json(path: Path, value: Any) -> None:
|
||||||
|
path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
|
||||||
|
|
||||||
|
|
||||||
|
def run_gold(args: argparse.Namespace) -> dict[str, Any]:
|
||||||
|
cases = load_gold_cases(args.cases)
|
||||||
|
args.output.mkdir(parents=True, exist_ok=False)
|
||||||
|
_write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases})
|
||||||
|
evaluations: list[dict[str, Any]] = []
|
||||||
|
successful_calls = 0
|
||||||
|
technical_failures = 0
|
||||||
|
started = time.perf_counter()
|
||||||
|
for case in cases:
|
||||||
|
case_dir = args.output / case["case_id"].lower()
|
||||||
|
case_dir.mkdir()
|
||||||
|
observations = case["observations"]
|
||||||
|
_write_json(case_dir / "v3_style_input_observations.json", observations)
|
||||||
|
prompt = build_prompt(case)
|
||||||
|
(case_dir / "prompt.txt").write_text(prompt, encoding="utf-8")
|
||||||
|
try:
|
||||||
|
raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict)
|
||||||
|
successful_calls += 1
|
||||||
|
except Exception as exc: # one recorded attempt; never retry
|
||||||
|
technical_failures += 1
|
||||||
|
failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)}
|
||||||
|
_write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure})
|
||||||
|
_write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)})
|
||||||
|
_write_json(case_dir / "deterministic_gate_results.json", {})
|
||||||
|
_write_json(case_dir / "final_derived_result.json", None)
|
||||||
|
_write_json(case_dir / "evaluation.json", failure)
|
||||||
|
evaluations.append(failure)
|
||||||
|
continue
|
||||||
|
(case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8")
|
||||||
|
_write_json(case_dir / "ollama_metadata.json", metadata)
|
||||||
|
try:
|
||||||
|
parsed = parse_model_json(raw)
|
||||||
|
_write_json(case_dir / "parsed_semantic_recognition.json", parsed)
|
||||||
|
evaluation = evaluate_case(case, parsed)
|
||||||
|
validation = {"valid": True, "error": None}
|
||||||
|
gates, result = derive_rejection(observations, parsed)
|
||||||
|
except (DerivationValidationError, json.JSONDecodeError) as exc:
|
||||||
|
validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)}
|
||||||
|
evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)}
|
||||||
|
gates, result = {}, None
|
||||||
|
_write_json(case_dir / "structural_validation.json", validation)
|
||||||
|
_write_json(case_dir / "deterministic_gate_results.json", gates)
|
||||||
|
_write_json(case_dir / "final_derived_result.json", result)
|
||||||
|
_write_json(case_dir / "evaluation.json", evaluation)
|
||||||
|
evaluations.append(evaluation)
|
||||||
|
summary = {
|
||||||
|
"experiment": "explicit_rejection_gold_v0", "model": args.model,
|
||||||
|
"successful_llm_call_count": successful_calls,
|
||||||
|
"technical_failed_call_count": technical_failures,
|
||||||
|
"runtime_seconds": round(time.perf_counter() - started, 3),
|
||||||
|
"counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")},
|
||||||
|
"evaluations": evaluations,
|
||||||
|
}
|
||||||
|
_write_json(args.output / "summary.json", summary)
|
||||||
|
return summary
|
||||||
|
|
||||||
|
|
||||||
|
def parse_args() -> argparse.Namespace:
|
||||||
|
parser = argparse.ArgumentParser(description="Run isolated explicit-rejection Gold experiment")
|
||||||
|
parser.add_argument("cases", type=Path)
|
||||||
|
parser.add_argument("-o", "--output", type=Path, required=True)
|
||||||
|
parser.add_argument("--model", default=DEFAULT_MODEL)
|
||||||
|
parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT)
|
||||||
|
parser.add_argument("--timeout", type=int, default=300)
|
||||||
|
parser.add_argument("--num-ctx", type=int, default=16384)
|
||||||
|
parser.add_argument("--num-predict", type=int, default=1024)
|
||||||
|
return parser.parse_args()
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
summary = run_gold(parse_args())
|
||||||
|
print(json.dumps(summary, ensure_ascii=False, indent=2))
|
||||||
|
return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
@@ -0,0 +1,100 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Isolated controlled rejection derivation V1 experiment."""
|
||||||
|
from __future__ import annotations
|
||||||
|
import argparse, json, time
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
from .experiment_h import DEFAULT_ENDPOINT, DEFAULT_MODEL, DerivationValidationError, OBSERVATION_KEYS, call_ollama
|
||||||
|
from .experiment_negative_act import build_prompt as build_negative_prompt, parse_model_json, validate_classification
|
||||||
|
|
||||||
|
SCHEMA_VERSION="experimental-controlled-rejection-v1"
|
||||||
|
TARGET_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||||
|
FORBIDDEN={"rejection_form","negative_act_form","explicitly_rejected","status","decision","outcome","topic_status","closed","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
|
||||||
|
PROMPT="""Resolve only the concrete local action or option referred to by the candidate negative act. The candidate is {candidate}. Choose only a supplied observation ID. Use the same observation for a self-contained target. If no unique local target exists, return null for both target fields. Preserve source-language meaning and material scope such as purpose and location. Preserve continuation when non-pursuit concerns continuing something. Ignore any separate positive alternative. Do not classify the negative act and do not output rejection, status, decision, outcome, responsibility, protocol concepts, confidence, relations, or graphs. Return exactly JSON with candidate_observation_id, target_observation_id, normalized_target_text and no other fields.\nObservations:\n{observations}"""
|
||||||
|
|
||||||
|
def _keys(v,r,loc):
|
||||||
|
if not isinstance(v,dict): raise DerivationValidationError(f"{loc} must be an object")
|
||||||
|
if set(v)!=r: raise DerivationValidationError(f"{loc} keys invalid: missing={sorted(r-set(v))}, unknown={sorted(set(v)-r)}")
|
||||||
|
def _text(v,loc):
|
||||||
|
if not isinstance(v,str) or not v.strip(): raise DerivationValidationError(f"{loc} must be non-empty")
|
||||||
|
return v.strip()
|
||||||
|
def _forbidden(v,loc="output"):
|
||||||
|
if isinstance(v,dict):
|
||||||
|
bad=FORBIDDEN & set(v)
|
||||||
|
if bad: raise DerivationValidationError(f"{loc} contains forbidden fields: {sorted(bad)}")
|
||||||
|
for k,x in v.items(): _forbidden(x,f"{loc}.{k}")
|
||||||
|
elif isinstance(v,list):
|
||||||
|
for i,x in enumerate(v): _forbidden(x,f"{loc}[{i}]")
|
||||||
|
def validate_observations(obs):
|
||||||
|
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
|
||||||
|
ids=set(); evid=set()
|
||||||
|
for i,o in enumerate(obs):
|
||||||
|
_keys(o,OBSERVATION_KEYS,f"observations[{i}]"); oid=_text(o["observation_id"],"observation_id"); eid=_text(o["evidence_id"],"evidence_id")
|
||||||
|
if oid in ids or eid in evid: raise DerivationValidationError("provenance must be unique")
|
||||||
|
ids.add(oid); evid.add(eid); _text(o["content"],"content"); _text(o["speaker"],"speaker")
|
||||||
|
return ids
|
||||||
|
def validate_target(data,obs):
|
||||||
|
ids=validate_observations(obs); _forbidden(data); _keys(data,TARGET_KEYS,"target output")
|
||||||
|
candidate=_text(data["candidate_observation_id"],"candidate_observation_id")
|
||||||
|
if candidate not in ids: raise DerivationValidationError("unknown candidate observation")
|
||||||
|
target=data["target_observation_id"]; normalized=data["normalized_target_text"]
|
||||||
|
if target is None:
|
||||||
|
if normalized is not None: raise DerivationValidationError("null target requires null text")
|
||||||
|
else:
|
||||||
|
target=_text(target,"target_observation_id")
|
||||||
|
if target not in ids: raise DerivationValidationError("unknown target observation")
|
||||||
|
_text(normalized,"normalized_target_text")
|
||||||
|
return data
|
||||||
|
def build_target_prompt(case):
|
||||||
|
validate_observations(case["observations"])
|
||||||
|
return PROMPT.format(candidate=case["expected"]["candidate_observation_id"],observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||||
|
def derive(obs,negative,target):
|
||||||
|
ids=validate_observations(obs); validate_classification(negative,obs); validate_target(target,obs)
|
||||||
|
candidate=negative["observation_id"]
|
||||||
|
if target["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate outputs disagree")
|
||||||
|
positions={o["observation_id"]:i for i,o in enumerate(obs)}; tid=target["target_observation_id"]
|
||||||
|
gates={"negative_act_valid":True,"eligible_explicit_non_pursuit":negative["negative_act_form"]=="explicit_non_pursuit","candidate_exists":candidate in ids,"target_valid":True,"target_present":tid is not None,"target_exists":tid in ids if tid else False,"target_not_after_candidate":positions[tid]<=positions[candidate] if tid else False,"provenance_valid_unique":True,"normalized_target_nonempty":bool(target["normalized_target_text"] and target["normalized_target_text"].strip()),"same_isolated_case":tid in ids if tid else False,"no_forbidden_fields":True}
|
||||||
|
established=all(gates.values())
|
||||||
|
result=None
|
||||||
|
if established:
|
||||||
|
byid={o["observation_id"]:o for o in obs}
|
||||||
|
result={"rejection_id":"rejection_1","content":target["normalized_target_text"].strip(),"status":"explicitly_rejected","support":{"target":{"observation_id":tid,"evidence_id":byid[tid]["evidence_id"]},"negative_act":{"observation_id":candidate,"evidence_id":byid[candidate]["evidence_id"]}}}
|
||||||
|
return {"gates":gates,"derived_result":result}
|
||||||
|
def load_cases(path):
|
||||||
|
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||||
|
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
|
||||||
|
return data["cases"]
|
||||||
|
def _concepts(text,groups):
|
||||||
|
folded=(text or "").casefold(); return all(any(x.casefold() in folded for x in g) for g in groups)
|
||||||
|
def evaluate(case,negative,target,derivation):
|
||||||
|
e=case["expected"]; text=target["normalized_target_text"]
|
||||||
|
form=negative["negative_act_form"]==e["negative_act_form"]; target_ok=target["target_observation_id"]==e["target_observation_id"]
|
||||||
|
action=_concepts(text,e["action_concepts"]); material=_concepts(text,e["material_concepts"]); forbidden=any(x.casefold() in (text or "").casefold() for x in e["forbidden_concepts"])
|
||||||
|
final=(derivation["derived_result"] is not None)==e["explicitly_rejected"]
|
||||||
|
label="PASS" if form and target_ok and action and material and not forbidden and final else ("PARTIAL" if form and target_ok and material and not forbidden and final else "FAIL")
|
||||||
|
return {"case_id":case["case_id"],"classification":label,"expected_negative_act_form":e["negative_act_form"],"actual_negative_act_form":negative["negative_act_form"],"expected_target_observation_id":e["target_observation_id"],"actual_target_observation_id":target["target_observation_id"],"normalized_target_text":text,"normalized_action_correct":action,"material_scope_preserved":material,"alternative_absorbed":forbidden,"final_rejection_correct":final}
|
||||||
|
def _write(p,v): p.write_text(json.dumps(v,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||||
|
def run(args):
|
||||||
|
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases})
|
||||||
|
evals=[]; naf_calls=target_calls=technical_failures=structural_failures=0; start=time.perf_counter()
|
||||||
|
for case in cases:
|
||||||
|
d=args.output/case["case_id"].lower(); d.mkdir(); obs=case["observations"]; _write(d/"v3_style_input_observations.json",obs)
|
||||||
|
try:
|
||||||
|
source=case["negative_act_source"]
|
||||||
|
if source=="live":
|
||||||
|
np=build_negative_prompt(case); (d/"negative_act_prompt.txt").write_text(np,encoding="utf-8"); raw,nmeta=call_ollama(args.endpoint,args.model,np,args.timeout,args.num_ctx,args.num_predict); naf_calls+=1; (d/"negative_act_raw_response.txt").write_text(raw+"\n",encoding="utf-8"); negative=parse_model_json(raw)
|
||||||
|
_write(d/"negative_act_ollama_metadata.json",nmeta); _write(d/"negative_act_source.json",{"kind":"live_call"})
|
||||||
|
else:
|
||||||
|
sd=args.negative_act_artifacts/source.lower(); accepted=json.loads((sd/"v3_style_input_observations.json").read_text());
|
||||||
|
if accepted!=obs: raise DerivationValidationError(f"{source} observations do not exactly match")
|
||||||
|
negative=json.loads((sd/"parsed_semantic_classification.json").read_text()); _write(d/"negative_act_source.json",{"kind":"accepted_artifact_reuse","case_id":source,"path":str(sd)})
|
||||||
|
validate_classification(negative,obs); _write(d/"negative_act_classification.json",negative)
|
||||||
|
tp=build_target_prompt(case); (d/"target_prompt.txt").write_text(tp,encoding="utf-8"); traw,tmeta=call_ollama(args.endpoint,args.model,tp,args.timeout,args.num_ctx,args.num_predict); target_calls+=1; (d/"target_raw_response.txt").write_text(traw+"\n",encoding="utf-8"); _write(d/"target_ollama_metadata.json",tmeta); target=parse_model_json(traw); _write(d/"target_recognition.json",target); validate_target(target,obs)
|
||||||
|
derivation=derive(obs,negative,target); _write(d/"deterministic_gate_results.json",derivation["gates"]); _write(d/"final_derived_result.json",derivation["derived_result"]); ev=evaluate(case,negative,target,derivation)
|
||||||
|
_write(d/"structural_validation.json",{"valid":True})
|
||||||
|
except Exception as exc:
|
||||||
|
structural_failures+=1; ev={"case_id":case["case_id"],"classification":"FAIL","technical_or_validation_failure":str(exc)}; _write(d/"structural_validation.json",{"valid":False,"error":str(exc)})
|
||||||
|
_write(d/"evaluation.json",ev); evals.append(ev)
|
||||||
|
summary={"experiment":"controlled_rejection_v1","model":args.model,"negative_act_llm_call_count":naf_calls,"reused_negative_act_count":len(cases)-naf_calls,"target_resolution_llm_call_count":target_calls,"technical_failed_call_count":technical_failures,"structural_validation_failure_count":structural_failures,"runtime_seconds":round(time.perf_counter()-start,3),"counts":{x:sum(e["classification"]==x for e in evals) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evals}; _write(args.output/"summary.json",summary); return summary
|
||||||
|
def main():
|
||||||
|
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--negative-act-artifacts",type=Path,default=Path("artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run")); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); args=p.parse_args(); print(json.dumps(run(args),ensure_ascii=False,indent=2)); return 0
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Target Normalization V0: reconstruct target text with fixed linkage."""
|
||||||
|
from __future__ import annotations
|
||||||
|
import argparse,json,time
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any,Callable
|
||||||
|
import requests
|
||||||
|
from .experiment_h import DEFAULT_ENDPOINT,DEFAULT_MODEL,DerivationValidationError,OBSERVATION_KEYS
|
||||||
|
|
||||||
|
SCHEMA_VERSION="experimental-target-normalization-v0"
|
||||||
|
OUTPUT_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||||
|
LINK_KEYS={"candidate_observation_id","target_observation_id"}
|
||||||
|
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","topic_status","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
|
||||||
|
PROMPT="""The candidate and target observation IDs below are already resolved. Copy both IDs exactly; do not perform target selection. Reconstruct only the concrete POSITIVE action or option meaning targeted by the negative act. Remove rejection and negation polarity while preserving the underlying positive action. Preserve German source language, material qualifiers, purpose, location, named people, and continuation. Exclude separate positive alternatives. Do not summarize the discussion or infer rejection, decision, outcome, status, responsibility, ownership, protocol relevance, confidence, relations, graphs, or topic state. Return only the JSON-Schema-conforming object; null is not permitted.\n\nExample A observations: [{{"observation_id":"obs_a","content":"Mit Frau Beispiel arbeiten wir nicht weiter."}}]\nFixed IDs: candidate=obs_a, target=obs_a\nOutput: {{"candidate_observation_id":"obs_a","target_observation_id":"obs_a","normalized_target_text":"Zusammenarbeit mit Frau Beispiel fortsetzen"}}\n\nExample B observations: [{{"observation_id":"obs_a","content":"Für den Druckversuch steht die reale Anlage zur Diskussion."}},{{"observation_id":"obs_b","content":"Die reale Anlage nutzen wir dafür nicht."}}]\nFixed IDs: candidate=obs_b, target=obs_a\nOutput: {{"candidate_observation_id":"obs_b","target_observation_id":"obs_a","normalized_target_text":"reale Anlage für den Druckversuch nutzen"}}\n\nFixed candidate_observation_id: {candidate}\nFixed target_observation_id: {target}\nV3-style observations:\n{observations}"""
|
||||||
|
|
||||||
|
def _keys(value,required,where):
|
||||||
|
if not isinstance(value,dict): raise DerivationValidationError(f"{where} must be an object")
|
||||||
|
if set(value)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(value))}, unknown={sorted(set(value)-required)}")
|
||||||
|
def _text(value,where):
|
||||||
|
if not isinstance(value,str) or not value.strip(): raise DerivationValidationError(f"{where} must be non-empty")
|
||||||
|
return value.strip()
|
||||||
|
def _reject_forbidden(value,where="output"):
|
||||||
|
if isinstance(value,dict):
|
||||||
|
bad=FORBIDDEN & set(value)
|
||||||
|
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
|
||||||
|
for key,item in value.items(): _reject_forbidden(item,f"{where}.{key}")
|
||||||
|
elif isinstance(value,list):
|
||||||
|
for index,item in enumerate(value): _reject_forbidden(item,f"{where}[{index}]")
|
||||||
|
def validate_observations(observations):
|
||||||
|
if not isinstance(observations,list) or not observations: raise DerivationValidationError("observations must be non-empty")
|
||||||
|
ids=set(); evidence=set()
|
||||||
|
for index,item in enumerate(observations):
|
||||||
|
_keys(item,OBSERVATION_KEYS,f"observations[{index}]"); oid=_text(item["observation_id"],"observation_id"); eid=_text(item["evidence_id"],"evidence_id")
|
||||||
|
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
|
||||||
|
ids.add(oid); evidence.add(eid); _text(item["content"],"content"); _text(item["speaker"],"speaker")
|
||||||
|
return ids
|
||||||
|
def validate_linkage(case):
|
||||||
|
ids=validate_observations(case["observations"]); linkage=case["fixed_linkage"]; _keys(linkage,LINK_KEYS,"fixed_linkage")
|
||||||
|
for field in LINK_KEYS:
|
||||||
|
if _text(linkage[field],field) not in ids: raise DerivationValidationError(f"{field} is unknown")
|
||||||
|
return linkage
|
||||||
|
def output_schema(case):
|
||||||
|
link=validate_linkage(case)
|
||||||
|
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","target_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":link["candidate_observation_id"]},"target_observation_id":{"const":link["target_observation_id"]},"normalized_target_text":{"type":"string","minLength":1}}}
|
||||||
|
def build_prompt(case):
|
||||||
|
link=validate_linkage(case)
|
||||||
|
return PROMPT.format(candidate=link["candidate_observation_id"],target=link["target_observation_id"],observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||||
|
def validate_output(data,case):
|
||||||
|
_reject_forbidden(data); _keys(data,OUTPUT_KEYS,"output"); link=validate_linkage(case)
|
||||||
|
if data["candidate_observation_id"]!=link["candidate_observation_id"]: raise DerivationValidationError("candidate ID changed")
|
||||||
|
if data["target_observation_id"]!=link["target_observation_id"]: raise DerivationValidationError("target ID changed")
|
||||||
|
_text(data["normalized_target_text"],"normalized_target_text"); return data
|
||||||
|
def build_payload(model,prompt,schema,num_ctx,num_predict): return {"model":model,"prompt":prompt,"think":False,"stream":False,"format":schema,"options":{"temperature":0,"num_ctx":num_ctx,"num_predict":num_predict}}
|
||||||
|
def call_schema(endpoint,model,prompt,schema,timeout,num_ctx,num_predict):
|
||||||
|
started=time.perf_counter(); response=requests.post(endpoint,json=build_payload(model,prompt,schema,num_ctx,num_predict),timeout=timeout); elapsed=time.perf_counter()-started; response.raise_for_status(); body=response.json(); raw=body.get("response")
|
||||||
|
if not isinstance(raw,str) or not raw.strip(): raise ValueError("Ollama returned no usable response")
|
||||||
|
return raw.strip(),{"model":body.get("model",model),"elapsed_seconds":round(elapsed,3),"total_duration_ns":body.get("total_duration"),"prompt_eval_count":body.get("prompt_eval_count"),"eval_count":body.get("eval_count"),"configuration":{"temperature":0,"think":False,"format":"json_schema_object","num_ctx":num_ctx,"num_predict":num_predict,"retries":0}}
|
||||||
|
def _concepts(text,groups):
|
||||||
|
folded=text.casefold(); return all(any(alias.casefold() in folded for alias in group) for group in groups)
|
||||||
|
def evaluate(case,output):
|
||||||
|
validate_output(output,case); expected=case["expected"]; text=output["normalized_target_text"]; folded=text.casefold(); action=_concepts(text,expected["action_concepts"]); scope=_concepts(text,expected["material_concepts"]); forbidden=[x for x in expected["forbidden_concepts"] if x.casefold() in folded]; positive=not any(x in forbidden for x in ("nicht","beenden")); german=any(x.casefold() in folded for x in expected["german_markers"]); alternative=not any(x.casefold() in folded for x in ("technikum","stattdessen")); strengthening=False
|
||||||
|
label="PASS" if positive and action and scope and german and alternative and not forbidden and not strengthening else "FAIL"
|
||||||
|
return {"case_id":case["case_id"],"classification":label,"expected_normalized_target_text":expected["normalized_target_text"],"actual_normalized_target_text":text,"positive_polarity_correct":positive,"action_semantics_preserved":action,"material_scope_preserved":scope,"source_language_preserved":german,"separate_alternative_excluded":alternative,"forbidden_semantics_present":forbidden,"unsupported_strengthening":strengthening,"normative_leakage":False,"candidate_id_unchanged":True,"target_id_unchanged":True}
|
||||||
|
def load_cases(path):
|
||||||
|
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||||
|
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
|
||||||
|
for case in data["cases"]: validate_linkage(case)
|
||||||
|
return data["cases"]
|
||||||
|
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||||
|
def run(args,caller:Callable=call_schema):
|
||||||
|
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases}); evaluations=[]; calls=failures=0; started=time.perf_counter()
|
||||||
|
for case in cases:
|
||||||
|
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"fixed_linkage.json",case["fixed_linkage"]); schema=output_schema(case); _write(folder/"ollama_json_schema.json",schema); prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
|
||||||
|
try:
|
||||||
|
raw,metadata=caller(args.endpoint,args.model,prompt,schema,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",metadata); parsed=json.loads(raw); _write(folder/"parsed_response.json",parsed); validate_output(parsed,case); _write(folder/"structural_validation.json",{"valid":True}); _write(folder/"normalized_target_result.json",{"normalized_target_text":parsed["normalized_target_text"]}); evaluation=evaluate(case,parsed)
|
||||||
|
except Exception as exc:
|
||||||
|
failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); evaluation={"case_id":case["case_id"],"classification":"FAIL","error":str(exc),"normative_leakage":False}
|
||||||
|
_write(folder/"evaluation.json",evaluation); evaluations.append(evaluation)
|
||||||
|
summary={"experiment":"target_normalization_v0","model":args.model,"llm_call_count":calls,"structural_validation_failure_count":failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
|
||||||
|
def main():
|
||||||
|
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
|
||||||
@@ -0,0 +1,92 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Isolated local target-resolution experiment; performs no rejection derivation."""
|
||||||
|
from __future__ import annotations
|
||||||
|
import argparse, json, time
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any, Callable
|
||||||
|
from .experiment_h import DEFAULT_ENDPOINT, DEFAULT_MODEL, DerivationValidationError, OBSERVATION_KEYS, call_ollama
|
||||||
|
from .experiment_negative_act import validate_classification
|
||||||
|
|
||||||
|
SCHEMA_VERSION="experimental-target-resolution-v0"
|
||||||
|
TARGET_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||||
|
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","topic_status","closed","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","unresolved_issue"}
|
||||||
|
PROMPT="""Resolve and normalize only the concrete action or option referred to by the candidate negative act. Candidate: {candidate}. Strategy: {instruction} Choose only a supplied observation ID. If no unique local target exists, use JSON null for both target fields. Preserve German source meaning, continuation, purpose, location, and other material scope. Do not translate Anlage as asset. Ignore separate positive alternatives. Do not output negative-act form, rejection, status, decision, outcome, responsibility, topic closure, protocol concepts, confidence, relations, or graphs. Return exactly this JSON object with no additional fields: {{"candidate_observation_id":"{candidate}","target_observation_id":"observation ID or null","normalized_target_text":"concise positive action in German or null"}}\nObservations:\n{observations}"""
|
||||||
|
|
||||||
|
def _keys(v:Any, required:set[str], where:str):
|
||||||
|
if not isinstance(v,dict): raise DerivationValidationError(f"{where} must be an object")
|
||||||
|
if set(v)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(v))}, unknown={sorted(set(v)-required)}")
|
||||||
|
def _text(v:Any, where:str):
|
||||||
|
if not isinstance(v,str) or not v.strip(): raise DerivationValidationError(f"{where} must be non-empty")
|
||||||
|
return v.strip()
|
||||||
|
def _reject_forbidden(v:Any, where="output"):
|
||||||
|
if isinstance(v,dict):
|
||||||
|
bad=FORBIDDEN & set(v)
|
||||||
|
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
|
||||||
|
for k,x in v.items(): _reject_forbidden(x,f"{where}.{k}")
|
||||||
|
elif isinstance(v,list):
|
||||||
|
for i,x in enumerate(v): _reject_forbidden(x,f"{where}[{i}]")
|
||||||
|
def validate_observations(obs):
|
||||||
|
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
|
||||||
|
ids=set(); evidence=set()
|
||||||
|
for i,o in enumerate(obs):
|
||||||
|
_keys(o,OBSERVATION_KEYS,f"observations[{i}]"); oid=_text(o["observation_id"],"observation_id"); eid=_text(o["evidence_id"],"evidence_id")
|
||||||
|
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
|
||||||
|
ids.add(oid); evidence.add(eid); _text(o["content"],"content"); _text(o["speaker"],"speaker")
|
||||||
|
return ids
|
||||||
|
def eligibility(negative,obs):
|
||||||
|
validate_classification(negative,obs)
|
||||||
|
eligible=negative["negative_act_form"]=="explicit_non_pursuit"
|
||||||
|
return {"eligible_for_target_resolution":eligible,"reason":None if eligible else "negative_act_form_not_explicit_non_pursuit"}
|
||||||
|
def validate_target(data,obs,candidate):
|
||||||
|
ids=validate_observations(obs); _reject_forbidden(data); _keys(data,TARGET_KEYS,"target output")
|
||||||
|
if _text(data["candidate_observation_id"],"candidate_observation_id")!=candidate: raise DerivationValidationError("candidate observation mismatch")
|
||||||
|
if candidate not in ids: raise DerivationValidationError("unknown candidate observation")
|
||||||
|
target=data["target_observation_id"]; normalized=data["normalized_target_text"]
|
||||||
|
if target is None:
|
||||||
|
if normalized is not None: raise DerivationValidationError("null target requires null text")
|
||||||
|
else:
|
||||||
|
target=_text(target,"target_observation_id")
|
||||||
|
if target not in ids: raise DerivationValidationError("unknown target observation")
|
||||||
|
if [o["observation_id"] for o in obs].index(target)>[o["observation_id"] for o in obs].index(candidate): raise DerivationValidationError("target must not occur after candidate")
|
||||||
|
_text(normalized,"normalized_target_text")
|
||||||
|
return data
|
||||||
|
def build_prompt(case):
|
||||||
|
gate=eligibility(case["negative_act"],case["observations"])
|
||||||
|
if not gate["eligible_for_target_resolution"]: raise DerivationValidationError("ineligible case must not build a target prompt")
|
||||||
|
candidate=case["negative_act"]["observation_id"]
|
||||||
|
instruction=(f"The target linkage is deterministically fixed to {candidate}; output that exact target ID and only normalize its positive action meaning." if case["strategy"]=="self_contained" else "Resolve the unique preceding local observation that supplies the referenced action.")
|
||||||
|
return PROMPT.format(candidate=candidate,instruction=instruction,observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||||
|
def _concepts(text,groups):
|
||||||
|
folded=(text or "").casefold(); return all(any(alias.casefold() in folded for alias in group) for group in groups)
|
||||||
|
def evaluate(case,gate,called,target):
|
||||||
|
e=case["expected"]; eligible=gate["eligible_for_target_resolution"]==e["eligible"]; call_ok=called==e["eligible"]
|
||||||
|
if not e["eligible"]:
|
||||||
|
label="PASS" if eligible and call_ok and target is None else "FAIL"
|
||||||
|
return {"case_id":case["case_id"],"classification":label,"negative_act_form":case["negative_act"]["negative_act_form"],"eligible":gate["eligible_for_target_resolution"],"target_resolution_call_made":called,"expected_target_observation_id":None,"actual_target_observation_id":None,"normalized_target_text":None,"material_scope_preserved":True,"alternative_isolation":True,"normative_leakage":False}
|
||||||
|
text=target["normalized_target_text"]; target_ok=target["target_observation_id"]==e["target_observation_id"]; concepts=_concepts(text,e["concepts"]); material=_concepts(text,e["material_concepts"]); isolated=not any(x.casefold() in (text or "").casefold() for x in e["forbidden_concepts"])
|
||||||
|
label="PASS" if eligible and call_ok and target_ok and concepts and material and isolated else ("PARTIAL" if eligible and call_ok and target_ok and material and isolated else "FAIL")
|
||||||
|
return {"case_id":case["case_id"],"classification":label,"negative_act_form":case["negative_act"]["negative_act_form"],"eligible":gate["eligible_for_target_resolution"],"target_resolution_call_made":called,"expected_target_observation_id":e["target_observation_id"],"actual_target_observation_id":target["target_observation_id"],"normalized_target_text":text,"normalized_action_correct":concepts,"material_scope_preserved":material,"alternative_isolation":isolated,"normative_leakage":False}
|
||||||
|
def load_cases(path):
|
||||||
|
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||||
|
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("unexpected schema version")
|
||||||
|
for case in data["cases"]: validate_observations(case["observations"]); validate_classification(case["negative_act"],case["observations"])
|
||||||
|
return data["cases"]
|
||||||
|
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||||
|
def run(args, resolver:Callable=call_ollama):
|
||||||
|
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases})
|
||||||
|
evaluations=[]; calls=technical_failures=structural_failures=0; started=time.perf_counter()
|
||||||
|
for case in cases:
|
||||||
|
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"negative_act_form.json",case["negative_act"])
|
||||||
|
gate=eligibility(case["negative_act"],case["observations"]); _write(folder/"eligibility.json",gate)
|
||||||
|
if not gate["eligible_for_target_resolution"]:
|
||||||
|
skipped={"call_made":False,"reason":gate["reason"]}; _write(folder/"target_resolution_skipped.json",skipped); ev=evaluate(case,gate,False,None)
|
||||||
|
else:
|
||||||
|
prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
|
||||||
|
try:
|
||||||
|
raw,meta=resolver(args.endpoint,args.model,prompt,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",meta); parsed=json.loads(raw); _write(folder/"parsed_target_resolution.json",parsed); validate_target(parsed,case["observations"],case["negative_act"]["observation_id"]); _write(folder/"structural_validation.json",{"valid":True}); ev=evaluate(case,gate,True,parsed)
|
||||||
|
except Exception as exc:
|
||||||
|
structural_failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); ev={"case_id":case["case_id"],"classification":"FAIL","negative_act_form":case["negative_act"]["negative_act_form"],"eligible":True,"target_resolution_call_made":True,"error":str(exc)}
|
||||||
|
_write(folder/"evaluation.json",ev); evaluations.append(ev)
|
||||||
|
summary={"experiment":"target_resolution_v0","model":args.model,"target_resolution_llm_call_count":calls,"technical_failed_call_count":technical_failures,"structural_validation_failure_count":structural_failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
|
||||||
|
def main():
|
||||||
|
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
|
||||||
@@ -0,0 +1,107 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Target Resolution V1 diagnostic: linkage and normalization only."""
|
||||||
|
from __future__ import annotations
|
||||||
|
import argparse,json,time
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any,Callable
|
||||||
|
import requests
|
||||||
|
from .experiment_h import DEFAULT_ENDPOINT,DEFAULT_MODEL,DerivationValidationError,OBSERVATION_KEYS
|
||||||
|
from .experiment_negative_act import validate_classification
|
||||||
|
|
||||||
|
SCHEMA_VERSION="experimental-target-resolution-v1-diagnostic"
|
||||||
|
SELF_KEYS={"candidate_observation_id","normalized_target_text"}; PAIRED_KEYS={"candidate_observation_id","target_observation_id","normalized_target_text"}
|
||||||
|
FORBIDDEN={"negative_act_form","rejection_form","explicitly_rejected","status","decision","outcome","responsible_person","responsibility","owner","requested_actor","action_item","protocol_category","confidence","relation","relations","graph","topic_status","closed","unresolved_issue"}
|
||||||
|
SELF_PROMPT="""Normalize only the concrete positive action meaning in the self-contained candidate observation. The target linkage is already deterministic and is not your task. Preserve German, collaboration, named people, and continuation meaning. Do not output a target ID, rejection, status, decision, outcome, responsibility, ownership, protocol concepts, confidence, relations, graphs, or topic closure. Return only the schema-conforming object.\nCandidate observation ID: {candidate}\nObservation:\n{observations}"""
|
||||||
|
PAIRED_PROMPT="""Resolve and normalize only the concrete local action or option referred to by the candidate negative act. Choose exactly one listed allowed target observation ID, or use JSON null only when no unique local target exists. Never return the string \"null\". Preserve German source language and all material purpose/location scope. Do not absorb a separate positive alternative. Do not output negative-act form, rejection, status, decision, outcome, responsibility, ownership, protocol concepts, confidence, relations, graphs, or topic closure.\nAllowed target observation IDs:\n{allowed}\nConcrete positive typed example:\n{{"candidate_observation_id":"obs_2","target_observation_id":"obs_1","normalized_target_text":"externe Lösung weiterverfolgen"}}\nActual JSON-null example:\n{{"candidate_observation_id":"obs_2","target_observation_id":null,"normalized_target_text":null}}\nReturn only the schema-conforming object.\nCandidate observation ID: {candidate}\nObservations:\n{observations}"""
|
||||||
|
|
||||||
|
def _keys(value,required,where):
|
||||||
|
if not isinstance(value,dict): raise DerivationValidationError(f"{where} must be an object")
|
||||||
|
if set(value)!=required: raise DerivationValidationError(f"{where} keys invalid: missing={sorted(required-set(value))}, unknown={sorted(set(value)-required)}")
|
||||||
|
def _text(value,where):
|
||||||
|
if not isinstance(value,str) or not value.strip(): raise DerivationValidationError(f"{where} must be non-empty")
|
||||||
|
return value.strip()
|
||||||
|
def _forbidden(value,where="output"):
|
||||||
|
if isinstance(value,dict):
|
||||||
|
bad=FORBIDDEN & set(value)
|
||||||
|
if bad: raise DerivationValidationError(f"{where} contains forbidden fields: {sorted(bad)}")
|
||||||
|
for key,item in value.items(): _forbidden(item,f"{where}.{key}")
|
||||||
|
elif isinstance(value,list):
|
||||||
|
for index,item in enumerate(value): _forbidden(item,f"{where}[{index}]")
|
||||||
|
def validate_observations(obs):
|
||||||
|
if not isinstance(obs,list) or not obs: raise DerivationValidationError("observations must be non-empty")
|
||||||
|
ids=[]; evidence=set()
|
||||||
|
for index,item in enumerate(obs):
|
||||||
|
_keys(item,OBSERVATION_KEYS,f"observations[{index}]"); oid=_text(item["observation_id"],"observation_id"); eid=_text(item["evidence_id"],"evidence_id")
|
||||||
|
if oid in ids or eid in evidence: raise DerivationValidationError("observation/evidence provenance must be unique")
|
||||||
|
ids.append(oid); evidence.add(eid); _text(item["content"],"content"); _text(item["speaker"],"speaker")
|
||||||
|
return ids
|
||||||
|
def allowed_ids(case): return validate_observations(case["observations"])
|
||||||
|
def deterministic_self_link(case):
|
||||||
|
if case["strategy"]!="self_contained": raise DerivationValidationError("self-linkage requires self-contained strategy")
|
||||||
|
ids=validate_observations(case["observations"]); candidate=case["negative_act"]["observation_id"]
|
||||||
|
if candidate not in ids: raise DerivationValidationError("unknown candidate")
|
||||||
|
return {"linkage_source":"deterministic","candidate_observation_id":candidate,"target_observation_id":candidate}
|
||||||
|
def output_schema(case):
|
||||||
|
candidate=case["negative_act"]["observation_id"]
|
||||||
|
if case["strategy"]=="self_contained":
|
||||||
|
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":candidate},"normalized_target_text":{"type":"string","minLength":1}}}
|
||||||
|
ids=allowed_ids(case)
|
||||||
|
return {"type":"object","additionalProperties":False,"required":["candidate_observation_id","target_observation_id","normalized_target_text"],"properties":{"candidate_observation_id":{"const":candidate},"target_observation_id":{"enum":ids+[None]},"normalized_target_text":{"type":["string","null"]}},"allOf":[{"if":{"properties":{"target_observation_id":{"type":"null"}}},"then":{"properties":{"normalized_target_text":{"type":"null"}}},"else":{"properties":{"normalized_target_text":{"type":"string","minLength":1}}}}]}
|
||||||
|
def build_prompt(case):
|
||||||
|
validate_classification(case["negative_act"],case["observations"]); candidate=case["negative_act"]["observation_id"]
|
||||||
|
if case["strategy"]=="self_contained": return SELF_PROMPT.format(candidate=candidate,observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||||
|
return PAIRED_PROMPT.format(candidate=candidate,allowed=json.dumps(allowed_ids(case),ensure_ascii=False),observations=json.dumps(case["observations"],ensure_ascii=False,indent=2))
|
||||||
|
def validate_semantic_output(data,case):
|
||||||
|
_forbidden(data); candidate=case["negative_act"]["observation_id"]
|
||||||
|
if case["strategy"]=="self_contained":
|
||||||
|
_keys(data,SELF_KEYS,"self output")
|
||||||
|
if data["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate mismatch")
|
||||||
|
_text(data["normalized_target_text"],"normalized_target_text")
|
||||||
|
else:
|
||||||
|
_keys(data,PAIRED_KEYS,"paired output")
|
||||||
|
if data["candidate_observation_id"]!=candidate: raise DerivationValidationError("candidate mismatch")
|
||||||
|
target=data["target_observation_id"]
|
||||||
|
if target=="null": raise DerivationValidationError('string "null" is forbidden')
|
||||||
|
if target is None:
|
||||||
|
if data["normalized_target_text"] is not None: raise DerivationValidationError("null target requires null text")
|
||||||
|
else:
|
||||||
|
if target not in allowed_ids(case): raise DerivationValidationError("target is not an allowed ID")
|
||||||
|
ids=allowed_ids(case)
|
||||||
|
if ids.index(target)>ids.index(candidate): raise DerivationValidationError("target must not occur after candidate")
|
||||||
|
_text(data["normalized_target_text"],"normalized_target_text")
|
||||||
|
return data
|
||||||
|
def combine(case,semantic):
|
||||||
|
validate_semantic_output(semantic,case)
|
||||||
|
if case["strategy"]=="self_contained":
|
||||||
|
link=deterministic_self_link(case); return {**link,"normalized_target_text":semantic["normalized_target_text"]}
|
||||||
|
return {"linkage_source":"llm","candidate_observation_id":semantic["candidate_observation_id"],"target_observation_id":semantic["target_observation_id"],"normalized_target_text":semantic["normalized_target_text"]}
|
||||||
|
def build_payload(model,prompt,schema,num_ctx,num_predict):
|
||||||
|
return {"model":model,"prompt":prompt,"think":False,"stream":False,"format":schema,"options":{"temperature":0,"num_ctx":num_ctx,"num_predict":num_predict}}
|
||||||
|
def call_schema(endpoint,model,prompt,schema,timeout,num_ctx,num_predict):
|
||||||
|
started=time.perf_counter(); response=requests.post(endpoint,json=build_payload(model,prompt,schema,num_ctx,num_predict),timeout=timeout); elapsed=time.perf_counter()-started; response.raise_for_status(); body=response.json(); raw=body.get("response")
|
||||||
|
if not isinstance(raw,str) or not raw.strip(): raise ValueError("Ollama returned no usable response")
|
||||||
|
meta={"model":body.get("model",model),"elapsed_seconds":round(elapsed,3),"total_duration_ns":body.get("total_duration"),"prompt_eval_count":body.get("prompt_eval_count"),"eval_count":body.get("eval_count"),"configuration":{"temperature":0,"think":False,"format":"json_schema_object","num_ctx":num_ctx,"num_predict":num_predict,"retries":0}}
|
||||||
|
return raw.strip(),meta
|
||||||
|
def _concepts(text,groups):
|
||||||
|
folded=(text or "").casefold(); return all(any(x.casefold() in folded for x in group) for group in groups)
|
||||||
|
def evaluate(case,semantic,combined):
|
||||||
|
expected=case["expected"]; text=combined["normalized_target_text"]; target=combined["target_observation_id"]; concepts=_concepts(text,expected["concepts"]); material=_concepts(text,expected["material_concepts"]); isolated=not any(x.casefold() in (text or "").casefold() for x in expected["forbidden_concepts"]); recurrence=semantic.get("target_observation_id")=="null"; correct=target==expected["target_observation_id"]
|
||||||
|
label="PASS" if correct and concepts and material and isolated and not recurrence else ("PARTIAL" if correct and material and isolated and not recurrence else "FAIL")
|
||||||
|
return {"case_id":case["case_id"],"classification":label,"strategy":case["strategy"],"target_id_decision_source":combined["linkage_source"],"expected_target_observation_id":expected["target_observation_id"],"actual_target_observation_id":target,"normalized_target_text":text,"continuation_or_action_preserved":concepts,"material_scope_preserved":material,"alternative_isolated":isolated,"schema_valid":True,"string_null_recurrence":recurrence,"normative_leakage":False}
|
||||||
|
def load_cases(path):
|
||||||
|
data=json.loads(path.read_text(encoding="utf-8")); _keys(data,{"schema_version","cases"},"fixture")
|
||||||
|
if data["schema_version"]!=SCHEMA_VERSION: raise DerivationValidationError("wrong schema version")
|
||||||
|
return data["cases"]
|
||||||
|
def _write(path,value): path.write_text(json.dumps(value,ensure_ascii=False,indent=2)+"\n",encoding="utf-8")
|
||||||
|
def run(args,caller:Callable=call_schema):
|
||||||
|
cases=load_cases(args.cases); args.output.mkdir(parents=True,exist_ok=False); _write(args.output/"gold_cases.json",{"schema_version":SCHEMA_VERSION,"cases":cases}); evaluations=[]; calls=failures=0; started=time.perf_counter()
|
||||||
|
for case in cases:
|
||||||
|
folder=args.output/case["case_id"].lower(); folder.mkdir(); _write(folder/"v3_style_input_observations.json",case["observations"]); _write(folder/"negative_act_form.json",case["negative_act"]); _write(folder/"eligibility.json",{"eligible_for_target_resolution":True,"reason":None}); _write(folder/"deterministic_strategy.json",{"strategy":case["strategy"],"target_id_decision_source":"deterministic" if case["strategy"]=="self_contained" else "llm"}); _write(folder/"allowed_target_ids.json",allowed_ids(case)); schema=output_schema(case); _write(folder/"ollama_json_schema.json",schema); prompt=build_prompt(case); (folder/"prompt.txt").write_text(prompt,encoding="utf-8")
|
||||||
|
try:
|
||||||
|
raw,meta=caller(args.endpoint,args.model,prompt,schema,args.timeout,args.num_ctx,args.num_predict); calls+=1; (folder/"raw_model_response.txt").write_text(raw+"\n",encoding="utf-8"); _write(folder/"ollama_metadata.json",meta); semantic=json.loads(raw); _write(folder/"parsed_semantic_output.json",semantic); validate_semantic_output(semantic,case); combined=combine(case,semantic); _write(folder/"structural_validation.json",{"valid":True}); _write(folder/"deterministic_linkage_result.json",{k:combined[k] for k in ("linkage_source","candidate_observation_id","target_observation_id")}); _write(folder/"normalized_target_result.json",{"normalized_target_text":combined["normalized_target_text"]}); evaluation=evaluate(case,semantic,combined)
|
||||||
|
except Exception as exc:
|
||||||
|
failures+=1; _write(folder/"structural_validation.json",{"valid":False,"error":str(exc)}); evaluation={"case_id":case["case_id"],"classification":"FAIL","strategy":case["strategy"],"schema_valid":False,"error":str(exc)}
|
||||||
|
_write(folder/"evaluation.json",evaluation); evaluations.append(evaluation)
|
||||||
|
summary={"experiment":"target_resolution_v1_diagnostic","model":args.model,"llm_call_count":calls,"structural_validation_failure_count":failures,"runtime_seconds":round(time.perf_counter()-started,3),"counts":{x:sum(e["classification"]==x for e in evaluations) for x in ["PASS","PARTIAL","FAIL"]},"evaluations":evaluations}; _write(args.output/"summary.json",summary); return summary
|
||||||
|
def main():
|
||||||
|
p=argparse.ArgumentParser(); p.add_argument("cases",type=Path); p.add_argument("-o","--output",type=Path,required=True); p.add_argument("--model",default=DEFAULT_MODEL); p.add_argument("--endpoint",default=DEFAULT_ENDPOINT); p.add_argument("--timeout",type=int,default=300); p.add_argument("--num-ctx",type=int,default=16384); p.add_argument("--num-predict",type=int,default=1024); print(json.dumps(run(p.parse_args()),ensure_ascii=False,indent=2)); return 0
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "experimental-controlled-rejection-v1",
|
||||||
|
"cases": [
|
||||||
|
{"case_id":"CR-01","description":"self-contained non-pursuit","negative_act_source":"NA-01","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_1","target_observation_id":"obs_1","action_concepts":[["Schlummer"],["Zusammenarbeit","arbeiten"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":true}},
|
||||||
|
{"case_id":"CR-02","description":"paired non-pursuit","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":true}},
|
||||||
|
{"case_id":"CR-03","description":"personal preference","negative_act_source":"NA-03","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage für den Versuch nutzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde das nicht machen.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"personal_preference","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["reale Anlage"],["Versuch"],["nutzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||||
|
{"case_id":"CR-04","description":"recommendation","negative_act_source":"NA-04","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage verwenden.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde eher davon abraten.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"recommendation","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["reale Anlage"],["verwenden","nutzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||||
|
{"case_id":"CR-05","description":"temporary non-action","negative_act_source":"NA-05","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die Waschstufe einbauen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir erstmal noch nicht.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"temporary_non_action","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Waschstufe"],["einbauen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||||
|
{"case_id":"CR-06","description":"concern","negative_act_source":"NA-06","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten das neue Material einsetzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das wäre kritisch.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"none","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["neue Material","neues Material"],["einsetzen"]],"material_concepts":[],"forbidden_concepts":[],"explicitly_rejected":false}},
|
||||||
|
{"case_id":"CR-07","description":"scoped explicit rejection","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[],"explicitly_rejected":true}},
|
||||||
|
{"case_id":"CR-08","description":"rejection plus alternative","negative_act_source":"live","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"expected":{"negative_act_form":"explicit_non_pursuit","candidate_observation_id":"obs_2","target_observation_id":"obs_1","action_concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"],"explicitly_rejected":true}}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,111 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "experimental-explicit-rejection-gold-v0",
|
||||||
|
"cases": [
|
||||||
|
{
|
||||||
|
"case_id": "RJ-01", "description": "Explicit collective rejection with local target",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Nein, das machen wir nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": true}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-02", "description": "Explicit non-pursuit with paired target",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": true}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-03", "description": "Self-contained collaboration rejection",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_1", "target_observation_id": "obs_1", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": true}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-04", "description": "Personal preference",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-05", "description": "Concern",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-06", "description": "Uncertainty",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-07", "description": "Negative recommendation",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-08", "description": "Deferral",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die externe Lösung einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das entscheiden wir nächste Woche.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-09", "description": "Factual negation",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_1", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-10", "description": "Temporary non-action",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": false}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-11", "description": "Explicit rejection with material scope",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Die reale Anlage nutzen wir dafür nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"]], "qualifier_concepts": [["druckversuch", "dafür", "pressure test"]], "forbidden_action_concepts": []},
|
||||||
|
"expected_result": {"explicitly_rejected": true}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "RJ-12", "description": "Rejection plus positive alternative",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten den Versuch in der realen Anlage durchführen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir nicht; wir testen stattdessen im Technikum.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["versuch", "trial", "test"], ["real"], ["anlage", "plant"]], "qualifier_concepts": [["real"], ["anlage", "plant"]], "forbidden_action_concepts": ["technikum", "technical facility", "technical center", "technical centre"]},
|
||||||
|
"expected_result": {"explicitly_rejected": true}
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "experimental-negative-act-form-gold-v0",
|
||||||
|
"cases": [
|
||||||
|
{
|
||||||
|
"case_id": "NA-01", "description": "Explicit non-pursuit",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_1", "negative_act_form": "explicit_non_pursuit", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]]}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-02", "description": "Explicit non-pursuit paraphrase",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Die externe Lösung verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_1", "negative_act_form": "explicit_non_pursuit", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]]}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-03", "description": "Personal preference with local context",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_2", "negative_act_form": "personal_preference", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]]}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-04", "description": "Negative recommendation",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_2", "negative_act_form": "recommendation", "action_concepts": [["real"], ["anlage", "plant"], ["verwend", "use"]]}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-05", "description": "Temporary non-action",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_2", "negative_act_form": "temporary_non_action", "action_concepts": [["waschstufe", "washing stage"], ["einbau", "install"]]}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-06", "description": "Concern only",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_2", "negative_act_form": "none", "action_concepts": []}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-07", "description": "Uncertainty",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null},
|
||||||
|
{"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_2", "negative_act_form": "none", "action_concepts": []}
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"case_id": "NA-08", "description": "Factual negation",
|
||||||
|
"observations": [
|
||||||
|
{"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null}
|
||||||
|
],
|
||||||
|
"expected": {"observation_id": "obs_1", "negative_act_form": "none", "action_concepts": []}
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
{
|
||||||
|
"schema_version":"experimental-target-normalization-v0",
|
||||||
|
"cases":[
|
||||||
|
{"case_id":"TN-01","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_1","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"Zusammenarbeit mit Dr. Schlummer fortsetzen","action_concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":["nicht","beenden"],"german_markers":["Zusammenarbeit","arbeiten","fortsetzen"]}},
|
||||||
|
{"case_id":"TN-02","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"externe Lösung weiterverfolgen","action_concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":["nicht"],"german_markers":["Lösung","verfolgen"]}},
|
||||||
|
{"case_id":"TN-03","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"reale Anlage für den Druckversuch nutzen","action_concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":["nicht","zur Diskussion"],"german_markers":["Anlage","Druckversuch","nutzen"]}},
|
||||||
|
{"case_id":"TN-04","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"fixed_linkage":{"candidate_observation_id":"obs_2","target_observation_id":"obs_1"},"expected":{"normalized_target_text":"Versuch in der realen Anlage durchführen","action_concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["nicht","Technikum","stattdessen"],"german_markers":["Versuch","Anlage","durchführen"]}}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,13 @@
|
|||||||
|
{
|
||||||
|
"schema_version": "experimental-target-resolution-v0",
|
||||||
|
"cases": [
|
||||||
|
{"case_id":"TR-01","description":"self-contained continuation target","strategy":"self_contained","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"negative_act":{"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"working with Dr. Schlummer"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR-02","description":"paired pronoun target","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"verfolgen wir nicht weiter"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR-03","description":"scoped location and purpose target","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"reale Anlage dafür nicht nutzen"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR-04","description":"rejection plus alternative","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"Versuch in der realen Anlage nicht durchführen"},"expected":{"eligible":true,"target_observation_id":"obs_1","concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"]}},
|
||||||
|
{"case_id":"TR-05","description":"personal preference","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage für den Versuch nutzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde das nicht machen.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"personal_preference","normalized_action_text":"Ich würde das nicht machen"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR-06","description":"recommendation","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die reale Anlage verwenden.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Ich würde eher davon abraten.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"recommendation","normalized_action_text":"advise against using the real asset"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR-07","description":"temporary non-action","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten die Waschstufe einbauen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir erstmal noch nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"temporary_non_action","normalized_action_text":"install the washing stage"},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR-08","description":"concern only","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten das neue Material einsetzen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das wäre kritisch.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"none","normalized_action_text":null},"expected":{"eligible":false,"target_observation_id":null,"concepts":[],"material_concepts":[],"forbidden_concepts":[]}}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,9 @@
|
|||||||
|
{
|
||||||
|
"schema_version":"experimental-target-resolution-v1-diagnostic",
|
||||||
|
"cases":[
|
||||||
|
{"case_id":"TR1-V1","strategy":"self_contained","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.","speaker":"Martin","named_person":"Dr. Schlummer","addressee":null}],"negative_act":{"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"working with Dr. Schlummer"},"expected":{"target_observation_id":"obs_1","concepts":[["Zusammenarbeit","arbeiten"],["Schlummer"],["fortsetzen","weiter"]],"material_concepts":[],"forbidden_concepts":["nicht"]}},
|
||||||
|
{"case_id":"TR2-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das verfolgen wir nicht weiter.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"verfolgen wir nicht weiter"},"expected":{"target_observation_id":"obs_1","concepts":[["externe Lösung"],["weiterverfolgen","weiter verfolgen"]],"material_concepts":[],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR3-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Die reale Anlage nutzen wir dafür nicht.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"reale Anlage dafür nicht nutzen"},"expected":{"target_observation_id":"obs_1","concepts":[["Anlage"],["nutzen"]],"material_concepts":[["real"],["Druckversuch"]],"forbidden_concepts":[]}},
|
||||||
|
{"case_id":"TR4-V1","strategy":"paired","observations":[{"observation_id":"obs_1","evidence_id":"e1","content":"Martin: Wir könnten den Versuch in der realen Anlage durchführen.","speaker":"Martin","named_person":null,"addressee":null},{"observation_id":"obs_2","evidence_id":"e2","content":"Martin: Das machen wir nicht; wir testen stattdessen im Technikum.","speaker":"Martin","named_person":null,"addressee":null}],"negative_act":{"observation_id":"obs_2","negative_act_form":"explicit_non_pursuit","normalized_action_text":"Versuch in der realen Anlage nicht durchführen"},"expected":{"target_observation_id":"obs_1","concepts":[["Versuch"],["durchführen"]],"material_concepts":[["real"],["Anlage"]],"forbidden_concepts":["Technikum"]}}
|
||||||
|
]
|
||||||
|
}
|
||||||
@@ -0,0 +1,82 @@
|
|||||||
|
import copy, json, unittest
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection_v1 import (
|
||||||
|
build_target_prompt, derive, load_cases, validate_target,
|
||||||
|
)
|
||||||
|
|
||||||
|
CASES=load_cases(Path("tests/gold/controlled_rejection_v1/cases.json"))
|
||||||
|
BY_ID={c["case_id"]:c for c in CASES}
|
||||||
|
|
||||||
|
def negative(case, form=None):
|
||||||
|
return {"observation_id":case["expected"]["candidate_observation_id"],"negative_act_form":form or case["expected"]["negative_act_form"],"normalized_action_text":None if (form or case["expected"]["negative_act_form"])=="none" else "semantische Aktion"}
|
||||||
|
def target(case, oid=None, text="konkrete Zielhandlung"):
|
||||||
|
return {"candidate_observation_id":case["expected"]["candidate_observation_id"],"target_observation_id":oid or case["expected"]["target_observation_id"],"normalized_target_text":text}
|
||||||
|
|
||||||
|
class ControlledRejectionV1Tests(unittest.TestCase):
|
||||||
|
def test_positive_forms_derive_and_provenance_survives(self):
|
||||||
|
for cid in ("CR-01","CR-02","CR-07","CR-08"):
|
||||||
|
c=BY_ID[cid]; out=derive(c["observations"],negative(c),target(c))
|
||||||
|
self.assertEqual(out["derived_result"]["status"],"explicitly_rejected")
|
||||||
|
self.assertEqual(out["derived_result"]["support"]["target"]["evidence_id"],"e1")
|
||||||
|
self.assertNotIn("responsible_person",json.dumps(out["derived_result"]))
|
||||||
|
|
||||||
|
def test_noneligible_form_never_derives_even_with_target(self):
|
||||||
|
for form in ["personal_preference","recommendation","temporary_non_action","none"]:
|
||||||
|
c=BY_ID["CR-03"]; self.assertIsNone(derive(c["observations"],negative(c,form),target(c))["derived_result"])
|
||||||
|
|
||||||
|
def test_missing_target_prevents_derivation(self):
|
||||||
|
c=BY_ID["CR-02"]; t=target(c); t.update(target_observation_id=None,normalized_target_text=None)
|
||||||
|
self.assertIsNone(derive(c["observations"],negative(c),t)["derived_result"])
|
||||||
|
|
||||||
|
def test_unknown_ids_rejected(self):
|
||||||
|
for field in ["candidate_observation_id","target_observation_id"]:
|
||||||
|
c=BY_ID["CR-02"]; t=target(c); t[field]="obs_unknown"
|
||||||
|
with self.assertRaises(DerivationValidationError): validate_target(t,c["observations"])
|
||||||
|
|
||||||
|
def test_target_after_candidate_cannot_derive(self):
|
||||||
|
c=copy.deepcopy(BY_ID["CR-02"]); n={"observation_id":"obs_1","negative_act_form":"explicit_non_pursuit","normalized_action_text":"x"}; t={"candidate_observation_id":"obs_1","target_observation_id":"obs_2","normalized_target_text":"x"}
|
||||||
|
self.assertIsNone(derive(c["observations"],n,t)["derived_result"])
|
||||||
|
|
||||||
|
def test_same_observation_target_allowed(self):
|
||||||
|
c=BY_ID["CR-01"]; self.assertIsNotNone(derive(c["observations"],negative(c),target(c))["derived_result"])
|
||||||
|
|
||||||
|
def test_duplicate_provenance_rejected(self):
|
||||||
|
for field in ["observation_id","evidence_id"]:
|
||||||
|
c=copy.deepcopy(BY_ID["CR-02"]); c["observations"][1][field]=c["observations"][0][field]
|
||||||
|
with self.assertRaises(DerivationValidationError): derive(c["observations"],negative(BY_ID["CR-02"]),target(BY_ID["CR-02"]))
|
||||||
|
|
||||||
|
def test_target_text_constraints(self):
|
||||||
|
c=BY_ID["CR-02"]
|
||||||
|
with self.assertRaises(DerivationValidationError): validate_target(target(c,text=""),c["observations"])
|
||||||
|
t=target(c); t["target_observation_id"]=None
|
||||||
|
with self.assertRaises(DerivationValidationError): validate_target(t,c["observations"])
|
||||||
|
|
||||||
|
def test_null_target_accepts_only_null_text(self):
|
||||||
|
c=BY_ID["CR-02"]; t=target(c); t.update(target_observation_id=None,normalized_target_text=None)
|
||||||
|
self.assertEqual(validate_target(t,c["observations"]),t)
|
||||||
|
|
||||||
|
def test_forbidden_and_unknown_fields_rejected(self):
|
||||||
|
for extra in [{"status":"rejected"},{"nested":{"decision":True}},{"extra":1}]:
|
||||||
|
c=BY_ID["CR-02"]; t=target(c); t.update(extra)
|
||||||
|
with self.assertRaises(DerivationValidationError): validate_target(t,c["observations"])
|
||||||
|
|
||||||
|
def test_candidate_outputs_must_agree(self):
|
||||||
|
c=BY_ID["CR-02"]; t=target(c); t["candidate_observation_id"]="obs_1"
|
||||||
|
with self.assertRaises(DerivationValidationError): derive(c["observations"],negative(c),t)
|
||||||
|
|
||||||
|
def test_scope_and_alternative_fixture_contract(self):
|
||||||
|
self.assertLessEqual({"real","Druckversuch"},{x for group in BY_ID["CR-07"]["expected"]["material_concepts"] for x in group})
|
||||||
|
self.assertEqual(BY_ID["CR-08"]["expected"]["forbidden_concepts"],["Technikum"])
|
||||||
|
|
||||||
|
def test_target_prompt_is_semantic_only_and_fixed(self):
|
||||||
|
prompt=build_target_prompt(BY_ID["CR-08"])
|
||||||
|
self.assertIn("Ignore any separate positive alternative",prompt)
|
||||||
|
self.assertIn("Do not classify the negative act",prompt)
|
||||||
|
|
||||||
|
def test_accepted_negative_act_inputs_are_exactly_reused(self):
|
||||||
|
root=Path("artifacts/experiments/negative_act_form_v0/20260820_qwen35_9b_single_run")
|
||||||
|
for cid,nid in (("CR-01","NA-01"),("CR-03","NA-03"),("CR-04","NA-04"),("CR-05","NA-05"),("CR-06","NA-06")):
|
||||||
|
accepted=json.loads((root/nid.lower()/"v3_style_input_observations.json").read_text())
|
||||||
|
self.assertEqual(BY_ID[cid]["observations"],accepted)
|
||||||
@@ -0,0 +1,210 @@
|
|||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from copy import deepcopy
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import (
|
||||||
|
DerivationValidationError,
|
||||||
|
build_prompt,
|
||||||
|
derive_rejection,
|
||||||
|
evaluate_case,
|
||||||
|
load_gold_cases,
|
||||||
|
validate_recognition,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
GOLD_PATH = Path("tests/gold/explicit_rejection_v0/cases.json")
|
||||||
|
|
||||||
|
|
||||||
|
POSITIVE_TEXT = {
|
||||||
|
"RJ-01": "reale Anlage für den Versuch nutzen",
|
||||||
|
"RJ-02": "externe Lösung weiterverfolgen",
|
||||||
|
"RJ-03": "Zusammenarbeit mit Dr. Schlummer fortsetzen",
|
||||||
|
"RJ-11": "reale Anlage für den Druckversuch nutzen",
|
||||||
|
"RJ-12": "Versuch in der realen Anlage durchführen",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def recognition_for(case):
|
||||||
|
expected = case["expected_recognition"]
|
||||||
|
positive = expected["rejection_form"] == "explicit_action_rejection"
|
||||||
|
return {
|
||||||
|
"rejection_observation_id": expected["rejection_observation_id"],
|
||||||
|
"target_observation_id": expected["target_observation_id"] if positive else None,
|
||||||
|
"rejection_form": expected["rejection_form"],
|
||||||
|
"normalized_rejected_action_text": POSITIVE_TEXT.get(case["case_id"]) if positive else None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class ExplicitRejectionGoldExperimentTests(unittest.TestCase):
|
||||||
|
@classmethod
|
||||||
|
def setUpClass(cls):
|
||||||
|
cls.cases = load_gold_cases(GOLD_PATH)
|
||||||
|
cls.by_id = {case["case_id"]: case for case in cls.cases}
|
||||||
|
|
||||||
|
def test_fixture_contains_exactly_rj_01_through_rj_12(self):
|
||||||
|
self.assertEqual(list(self.by_id), [f"RJ-{number:02d}" for number in range(1, 13)])
|
||||||
|
|
||||||
|
def test_cases_use_only_minimal_v3_style_observations(self):
|
||||||
|
keys = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"}
|
||||||
|
for case in self.cases:
|
||||||
|
with self.subTest(case=case["case_id"]):
|
||||||
|
self.assertIn(len(case["observations"]), (1, 2))
|
||||||
|
self.assertTrue(all(set(item) == keys for item in case["observations"]))
|
||||||
|
|
||||||
|
def test_rj_01_derives_target_and_both_provenance_paths(self):
|
||||||
|
case = self.by_id["RJ-01"]
|
||||||
|
gates, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
self.assertTrue(all(gates.values()))
|
||||||
|
self.assertEqual(result["status"], "explicitly_rejected")
|
||||||
|
self.assertEqual(result["support"]["target"], {"observation_id": "obs_1", "evidence_id": "e1"})
|
||||||
|
self.assertEqual(result["support"]["rejection"], {"observation_id": "obs_2", "evidence_id": "e2"})
|
||||||
|
|
||||||
|
def test_rj_02_requires_paired_target_and_derives_abandonment(self):
|
||||||
|
case = self.by_id["RJ-02"]
|
||||||
|
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
self.assertIn("externe Lösung", result["content"])
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unknown target"):
|
||||||
|
derive_rejection(case["observations"][1:], recognition_for(case))
|
||||||
|
|
||||||
|
def test_rj_03_supports_same_observation_target_and_rejection(self):
|
||||||
|
case = self.by_id["RJ-03"]
|
||||||
|
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
self.assertEqual(result["support"]["target"], result["support"]["rejection"])
|
||||||
|
self.assertIn("Dr. Schlummer", result["content"])
|
||||||
|
|
||||||
|
def test_all_required_negative_cases_remain_non_rejections(self):
|
||||||
|
for case_id in ("RJ-04", "RJ-05", "RJ-06", "RJ-07", "RJ-08", "RJ-09", "RJ-10"):
|
||||||
|
case = self.by_id[case_id]
|
||||||
|
gates, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
with self.subTest(case=case_id):
|
||||||
|
self.assertFalse(gates["explicit_action_rejection"])
|
||||||
|
self.assertIsNone(result)
|
||||||
|
|
||||||
|
def test_rj_11_derives_and_preserves_location_and_purpose_scope(self):
|
||||||
|
case = self.by_id["RJ-11"]
|
||||||
|
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
self.assertIsNotNone(result)
|
||||||
|
self.assertIn("reale Anlage", result["content"])
|
||||||
|
self.assertIn("Druckversuch", result["content"])
|
||||||
|
|
||||||
|
def test_rj_12_rejects_only_real_plant_action_and_not_alternative(self):
|
||||||
|
case = self.by_id["RJ-12"]
|
||||||
|
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
self.assertIn("realen Anlage", result["content"])
|
||||||
|
self.assertNotIn("Technikum", result["content"])
|
||||||
|
|
||||||
|
def test_separate_target_cannot_follow_rejection(self):
|
||||||
|
case = deepcopy(self.by_id["RJ-01"])
|
||||||
|
case["observations"].reverse()
|
||||||
|
gates, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
self.assertFalse(gates["target_same_or_before_rejection"])
|
||||||
|
self.assertIsNone(result)
|
||||||
|
|
||||||
|
def test_unknown_target_observation_id_is_rejected(self):
|
||||||
|
case = self.by_id["RJ-01"]
|
||||||
|
recognition = recognition_for(case)
|
||||||
|
recognition["target_observation_id"] = "obs_99"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unknown target"):
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
|
||||||
|
def test_unknown_rejection_observation_id_is_rejected(self):
|
||||||
|
case = self.by_id["RJ-01"]
|
||||||
|
recognition = recognition_for(case)
|
||||||
|
recognition["rejection_observation_id"] = "obs_99"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unknown rejection"):
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
|
||||||
|
def test_duplicate_observation_ids_are_rejected(self):
|
||||||
|
fixture = json.loads(GOLD_PATH.read_text())
|
||||||
|
fixture["cases"][0]["observations"][1]["observation_id"] = "obs_1"
|
||||||
|
self._assert_bad_fixture(fixture, "observation IDs must be unique")
|
||||||
|
|
||||||
|
def test_inconsistent_evidence_provenance_is_rejected(self):
|
||||||
|
fixture = json.loads(GOLD_PATH.read_text())
|
||||||
|
fixture["cases"][0]["observations"][1]["evidence_id"] = "e1"
|
||||||
|
self._assert_bad_fixture(fixture, "evidence provenance must be unique")
|
||||||
|
|
||||||
|
def test_none_rejects_populated_target_or_action(self):
|
||||||
|
case = self.by_id["RJ-04"]
|
||||||
|
for field, value, message in (
|
||||||
|
("target_observation_id", "obs_1", "null target"),
|
||||||
|
("normalized_rejected_action_text", "Anlage nutzen", "null normalized"),
|
||||||
|
):
|
||||||
|
recognition = recognition_for(case)
|
||||||
|
recognition[field] = value
|
||||||
|
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, message):
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
|
||||||
|
def test_explicit_rejection_requires_normalized_target_text(self):
|
||||||
|
case = self.by_id["RJ-01"]
|
||||||
|
for value in (None, ""):
|
||||||
|
recognition = recognition_for(case)
|
||||||
|
recognition["normalized_rejected_action_text"] = value
|
||||||
|
with self.subTest(value=value), self.assertRaises(DerivationValidationError):
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
|
||||||
|
def test_unknown_schema_fields_are_rejected(self):
|
||||||
|
case = self.by_id["RJ-01"]
|
||||||
|
recognition = recognition_for(case)
|
||||||
|
recognition["explanation"] = "extra"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unknown keys"):
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
|
||||||
|
def test_forbidden_normative_fields_are_rejected_recursively(self):
|
||||||
|
case = self.by_id["RJ-01"]
|
||||||
|
fields = (
|
||||||
|
"decision", "decision_status", "outcome", "topic_status", "closed",
|
||||||
|
"agreement", "responsible_person", "responsibility", "responsibility_scope",
|
||||||
|
"owner", "ownership", "assignee", "requested_actor", "status",
|
||||||
|
"explicitly_rejected", "action_item", "protocol_category", "confidence",
|
||||||
|
"relation", "relations", "graph", "unresolved_issue",
|
||||||
|
)
|
||||||
|
for field in fields:
|
||||||
|
recognition = recognition_for(case)
|
||||||
|
recognition["wrapper"] = {field: "forbidden"}
|
||||||
|
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||||
|
validate_recognition(recognition, case["observations"])
|
||||||
|
|
||||||
|
def test_speaker_identity_creates_no_ownership_or_responsibility(self):
|
||||||
|
case = deepcopy(self.by_id["RJ-03"])
|
||||||
|
for speaker in ("Martin", "Clara", "Antonius"):
|
||||||
|
case["observations"][0]["speaker"] = speaker
|
||||||
|
_, result = derive_rejection(case["observations"], recognition_for(case))
|
||||||
|
with self.subTest(speaker=speaker):
|
||||||
|
self.assertNotIn("responsible_person", result)
|
||||||
|
self.assertNotIn("owner", result)
|
||||||
|
|
||||||
|
def test_all_expected_recognitions_evaluate_as_pass(self):
|
||||||
|
for case in self.cases:
|
||||||
|
evaluation = evaluate_case(case, recognition_for(case))
|
||||||
|
with self.subTest(case=case["case_id"]):
|
||||||
|
self.assertEqual(evaluation["classification"], "PASS")
|
||||||
|
|
||||||
|
def test_rj_11_qualifier_loss_and_rj_12_alternative_absorption_fail(self):
|
||||||
|
rj11 = self.by_id["RJ-11"]
|
||||||
|
recognition = recognition_for(rj11)
|
||||||
|
recognition["normalized_rejected_action_text"] = "reale Anlage nutzen"
|
||||||
|
self.assertEqual(evaluate_case(rj11, recognition)["classification"], "FAIL")
|
||||||
|
rj12 = self.by_id["RJ-12"]
|
||||||
|
recognition = recognition_for(rj12)
|
||||||
|
recognition["normalized_rejected_action_text"] += "; stattdessen im Technikum testen"
|
||||||
|
self.assertEqual(evaluate_case(rj12, recognition)["classification"], "FAIL")
|
||||||
|
|
||||||
|
def test_prompt_is_fixed_narrow_and_does_not_expose_gold_expectation(self):
|
||||||
|
prompt = build_prompt(self.by_id["RJ-01"])
|
||||||
|
self.assertIn("candidate rejection observation is obs_2", prompt)
|
||||||
|
self.assertNotIn("expected_result", prompt)
|
||||||
|
self.assertNotIn("Who is responsible", prompt)
|
||||||
|
|
||||||
|
def _assert_bad_fixture(self, fixture, message):
|
||||||
|
with tempfile.TemporaryDirectory() as temporary:
|
||||||
|
path = Path(temporary) / "cases.json"
|
||||||
|
path.write_text(json.dumps(fixture), encoding="utf-8")
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, message):
|
||||||
|
load_gold_cases(path)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -0,0 +1,162 @@
|
|||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from copy import deepcopy
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch
|
||||||
|
|
||||||
|
import src.meeting_lab.controlled_semantic_derivation.experiment_negative_act as module
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_negative_act import (
|
||||||
|
DerivationValidationError,
|
||||||
|
build_ollama_payload,
|
||||||
|
build_prompt,
|
||||||
|
evaluate_case,
|
||||||
|
load_gold_cases,
|
||||||
|
parse_model_json,
|
||||||
|
run_experiment,
|
||||||
|
validate_classification,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
GOLD_PATH = Path("tests/gold/negative_act_form_v0/cases.json")
|
||||||
|
|
||||||
|
|
||||||
|
FORM_TEXT = {
|
||||||
|
"NA-01": "Zusammenarbeit mit Dr. Schlummer fortsetzen",
|
||||||
|
"NA-02": "externe Lösung weiterverfolgen",
|
||||||
|
"NA-03": "reale Anlage für den Versuch nutzen",
|
||||||
|
"NA-04": "reale Anlage verwenden",
|
||||||
|
"NA-05": "Waschstufe einbauen",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def classification_for(case):
|
||||||
|
expected = case["expected"]
|
||||||
|
return {
|
||||||
|
"observation_id": expected["observation_id"],
|
||||||
|
"negative_act_form": expected["negative_act_form"],
|
||||||
|
"normalized_action_text": FORM_TEXT.get(case["case_id"]),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class NegativeActFormExperimentTests(unittest.TestCase):
|
||||||
|
@classmethod
|
||||||
|
def setUpClass(cls):
|
||||||
|
cls.cases = load_gold_cases(GOLD_PATH)
|
||||||
|
cls.by_id = {case["case_id"]: case for case in cls.cases}
|
||||||
|
|
||||||
|
def test_fixture_contains_exactly_na_01_through_na_08(self):
|
||||||
|
self.assertEqual(list(self.by_id), [f"NA-{number:02d}" for number in range(1, 9)])
|
||||||
|
|
||||||
|
def test_exact_schema_is_accepted(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
self.assertEqual(validate_classification(classification_for(case), case["observations"]), classification_for(case))
|
||||||
|
|
||||||
|
def test_unknown_field_is_rejected(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
classification = classification_for(case)
|
||||||
|
classification["explanation"] = "extra"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unknown keys"):
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
|
||||||
|
def test_invalid_enum_is_rejected(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
classification = classification_for(case)
|
||||||
|
classification["negative_act_form"] = "rejection"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unsupported value"):
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
|
||||||
|
def test_non_none_requires_normalized_action_text(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
for value in (None, ""):
|
||||||
|
classification = classification_for(case)
|
||||||
|
classification["normalized_action_text"] = value
|
||||||
|
with self.subTest(value=value), self.assertRaises(DerivationValidationError):
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
|
||||||
|
def test_none_requires_null_normalized_action_text(self):
|
||||||
|
case = self.by_id["NA-06"]
|
||||||
|
classification = classification_for(case)
|
||||||
|
self.assertIsNone(classification["normalized_action_text"])
|
||||||
|
classification["normalized_action_text"] = "Material einsetzen"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "requires null"):
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
|
||||||
|
def test_forbidden_normative_fields_are_rejected_recursively(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
fields = (
|
||||||
|
"rejection_form", "explicitly_rejected", "status", "decision", "outcome",
|
||||||
|
"topic_status", "responsible_person", "responsibility", "owner",
|
||||||
|
"requested_actor", "action_item", "protocol_category", "confidence",
|
||||||
|
"relation", "relations", "graph", "unresolved_issue",
|
||||||
|
)
|
||||||
|
for field in fields:
|
||||||
|
classification = classification_for(case)
|
||||||
|
classification["wrapper"] = {field: "forbidden"}
|
||||||
|
with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"):
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
|
||||||
|
def test_unknown_observation_id_is_rejected(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
classification = classification_for(case)
|
||||||
|
classification["observation_id"] = "obs_99"
|
||||||
|
with self.assertRaisesRegex(DerivationValidationError, "unknown observation"):
|
||||||
|
validate_classification(classification, case["observations"])
|
||||||
|
|
||||||
|
def test_malformed_json_is_rejected(self):
|
||||||
|
with self.assertRaises(json.JSONDecodeError):
|
||||||
|
parse_model_json("{bad json")
|
||||||
|
|
||||||
|
def test_all_expected_classifications_evaluate_as_pass(self):
|
||||||
|
for case in self.cases:
|
||||||
|
evaluation = evaluate_case(case, classification_for(case))
|
||||||
|
with self.subTest(case=case["case_id"]):
|
||||||
|
self.assertEqual(evaluation["classification"], "PASS")
|
||||||
|
|
||||||
|
def test_fixed_prompt_contains_candidate_and_no_gold_expectation(self):
|
||||||
|
prompt = build_prompt(self.by_id["NA-03"])
|
||||||
|
self.assertIn("candidate observation is obs_2", prompt)
|
||||||
|
self.assertNotIn("expected", prompt)
|
||||||
|
self.assertNotIn("Who is responsible", prompt)
|
||||||
|
|
||||||
|
def test_fixed_model_configuration(self):
|
||||||
|
payload = build_ollama_payload("qwen3.5:9B", "prompt", 16384, 1024)
|
||||||
|
self.assertFalse(payload["think"])
|
||||||
|
self.assertFalse(payload["stream"])
|
||||||
|
self.assertEqual(payload["options"]["temperature"], 0)
|
||||||
|
|
||||||
|
def test_no_rejection_or_status_derivation_function_exists(self):
|
||||||
|
public_names = {name for name in dir(module) if not name.startswith("_")}
|
||||||
|
self.assertNotIn("derive_rejection", public_names)
|
||||||
|
self.assertFalse(any(name.startswith("derive_") for name in public_names))
|
||||||
|
|
||||||
|
def test_artifacts_preserve_semantic_classification_only(self):
|
||||||
|
case = self.by_id["NA-01"]
|
||||||
|
raw = json.dumps(classification_for(case), ensure_ascii=False)
|
||||||
|
with tempfile.TemporaryDirectory() as temporary:
|
||||||
|
output = Path(temporary) / "run"
|
||||||
|
args = argparse.Namespace(
|
||||||
|
cases=GOLD_PATH, output=output, model="qwen3.5:9B",
|
||||||
|
endpoint="http://unused", timeout=1, num_ctx=16384, num_predict=1024,
|
||||||
|
)
|
||||||
|
with patch.object(module, "load_gold_cases", return_value=[deepcopy(case)]), patch.object(
|
||||||
|
module, "call_ollama", return_value=(raw, {"model": "qwen3.5:9B"})
|
||||||
|
):
|
||||||
|
summary = run_experiment(args)
|
||||||
|
self.assertEqual(summary["successful_llm_call_count"], 1)
|
||||||
|
case_dir = output / "na-01"
|
||||||
|
for filename in (
|
||||||
|
"v3_style_input_observations.json", "prompt.txt", "raw_model_response.txt",
|
||||||
|
"parsed_semantic_classification.json", "structural_validation.json",
|
||||||
|
"evaluation.json", "ollama_metadata.json",
|
||||||
|
):
|
||||||
|
self.assertTrue((case_dir / filename).is_file(), filename)
|
||||||
|
self.assertFalse((case_dir / "final_derived_result.json").exists())
|
||||||
|
self.assertFalse((case_dir / "deterministic_gate_results.json").exists())
|
||||||
|
parsed = json.loads((case_dir / "parsed_semantic_classification.json").read_text())
|
||||||
|
self.assertEqual(set(parsed), {"observation_id", "negative_act_form", "normalized_action_text"})
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -0,0 +1,51 @@
|
|||||||
|
import argparse,copy,json,tempfile,unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import Mock
|
||||||
|
import src.meeting_lab.controlled_semantic_derivation.experiment_target_normalization as module
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||||
|
|
||||||
|
CASES=module.load_cases(Path("tests/gold/target_normalization_v0/cases.json")); BY={c["case_id"]:c for c in CASES}
|
||||||
|
def output(case,text=None):
|
||||||
|
link=case["fixed_linkage"]; return {"candidate_observation_id":link["candidate_observation_id"],"target_observation_id":link["target_observation_id"],"normalized_target_text":text or case["expected"]["normalized_target_text"]}
|
||||||
|
|
||||||
|
class TargetNormalizationTests(unittest.TestCase):
|
||||||
|
def test_exact_id_copying_accepted(self):
|
||||||
|
for case in CASES: self.assertEqual(module.validate_output(output(case),case),output(case))
|
||||||
|
def test_changed_candidate_rejected(self):
|
||||||
|
case=BY["TN-02"]; data=output(case); data["candidate_observation_id"]="obs_1"
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||||
|
def test_changed_target_rejected(self):
|
||||||
|
case=BY["TN-02"]; data=output(case); data["target_observation_id"]="obs_2"
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||||
|
def test_empty_and_null_text_rejected(self):
|
||||||
|
case=BY["TN-01"]
|
||||||
|
for value in ("",None):
|
||||||
|
data=output(case); data["normalized_target_text"]=value
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||||
|
def test_unknown_and_forbidden_fields_rejected(self):
|
||||||
|
case=BY["TN-01"]
|
||||||
|
for extra in ({"extra":1},{"status":"x"},{"nested":{"decision":True}}):
|
||||||
|
data=output(case); data.update(extra)
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_output(data,case)
|
||||||
|
def test_true_schema_fixes_both_ids_and_disallows_null(self):
|
||||||
|
case=BY["TN-02"]; schema=module.output_schema(case); self.assertEqual(schema["properties"]["candidate_observation_id"]["const"],"obs_2"); self.assertEqual(schema["properties"]["target_observation_id"]["const"],"obs_1"); self.assertEqual(schema["properties"]["normalized_target_text"]["type"],"string"); self.assertFalse(schema["additionalProperties"])
|
||||||
|
def test_payload_uses_schema_object(self):
|
||||||
|
schema=module.output_schema(BY["TN-01"]); payload=module.build_payload("qwen3.5:9B","p",schema,16384,1024); self.assertIs(payload["format"],schema); self.assertIsInstance(payload["format"],dict)
|
||||||
|
def test_duplicate_ids_and_evidence_rejected(self):
|
||||||
|
for field in ("observation_id","evidence_id"):
|
||||||
|
case=copy.deepcopy(BY["TN-02"]); case["observations"][1][field]=case["observations"][0][field]
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_linkage(case)
|
||||||
|
def test_prompt_is_fixed_normalization_only(self):
|
||||||
|
first=module.build_prompt(BY["TN-01"]); second=module.build_prompt(BY["TN-02"]); self.assertIn("do not perform target selection",first); self.assertIn("concrete POSITIVE action",first); self.assertEqual(first.split("Fixed candidate_observation_id:")[0],second.split("Fixed candidate_observation_id:")[0])
|
||||||
|
def test_no_target_selection_or_rejection_derivation_exists(self):
|
||||||
|
self.assertFalse(hasattr(module,"select_target")); self.assertFalse(hasattr(module,"derive")); self.assertNotIn("explicitly_rejected",module.OUTPUT_KEYS); self.assertNotIn("status",module.OUTPUT_KEYS)
|
||||||
|
def test_artifacts_preserve_fixed_linkage(self):
|
||||||
|
case=BY["TN-01"]; fixture={"schema_version":module.SCHEMA_VERSION,"cases":[case]}; caller=Mock(return_value=(json.dumps(output(case)),{"model":"qwen3.5:9B"}))
|
||||||
|
with tempfile.TemporaryDirectory() as tmp:
|
||||||
|
root=Path(tmp); path=root/"cases.json"; path.write_text(json.dumps(fixture)); out=root/"out"; args=argparse.Namespace(cases=path,output=out,endpoint="x",model="qwen3.5:9B",timeout=1,num_ctx=16384,num_predict=1024); summary=module.run(args,caller); self.assertEqual(summary["llm_call_count"],1); self.assertEqual(json.loads((out/"tn-01"/"fixed_linkage.json").read_text()),case["fixed_linkage"]); self.assertTrue((out/"tn-01"/"normalized_target_result.json").exists())
|
||||||
|
def test_negative_polarity_fails_evaluation(self):
|
||||||
|
case=BY["TN-01"]; self.assertEqual(module.evaluate(case,output(case,"Mit Dr. Schlummer arbeiten wir nicht weiter"))["classification"],"FAIL")
|
||||||
|
def test_material_scope_and_alternative_contract(self):
|
||||||
|
self.assertEqual(module.evaluate(BY["TN-03"],output(BY["TN-03"]))["classification"],"PASS"); self.assertEqual(module.evaluate(BY["TN-04"],output(BY["TN-04"]))["classification"],"PASS")
|
||||||
|
|
||||||
|
if __name__=="__main__": unittest.main()
|
||||||
@@ -0,0 +1,75 @@
|
|||||||
|
import argparse, copy, json, tempfile, unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import Mock
|
||||||
|
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||||
|
import src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution as module
|
||||||
|
|
||||||
|
GOLD=Path("tests/gold/target_resolution_v0/cases.json")
|
||||||
|
CASES=module.load_cases(GOLD); BY_ID={c["case_id"]:c for c in CASES}
|
||||||
|
|
||||||
|
def target(case, target_id=None, text="konkrete Zielhandlung"):
|
||||||
|
return {"candidate_observation_id":case["negative_act"]["observation_id"],"target_observation_id":target_id if target_id is not None else case["expected"]["target_observation_id"],"normalized_target_text":text}
|
||||||
|
|
||||||
|
class TargetResolutionTests(unittest.TestCase):
|
||||||
|
def test_eligibility_enum_boundary(self):
|
||||||
|
for cid in ("TR-01","TR-02","TR-03","TR-04"):
|
||||||
|
self.assertTrue(module.eligibility(BY_ID[cid]["negative_act"],BY_ID[cid]["observations"])["eligible_for_target_resolution"])
|
||||||
|
for cid in ("TR-05","TR-06","TR-07","TR-08"):
|
||||||
|
gate=module.eligibility(BY_ID[cid]["negative_act"],BY_ID[cid]["observations"])
|
||||||
|
self.assertFalse(gate["eligible_for_target_resolution"]); self.assertEqual(gate["reason"],"negative_act_form_not_explicit_non_pursuit")
|
||||||
|
|
||||||
|
def test_ineligible_cases_never_call_resolver_and_record_skip(self):
|
||||||
|
fixture={"schema_version":module.SCHEMA_VERSION,"cases":[BY_ID[x] for x in ("TR-05","TR-06","TR-07","TR-08")]}; resolver=Mock()
|
||||||
|
with tempfile.TemporaryDirectory() as tmp:
|
||||||
|
root=Path(tmp); cases=root/"cases.json"; cases.write_text(json.dumps(fixture)); out=root/"out"
|
||||||
|
summary=module.run(argparse.Namespace(cases=cases,output=out,endpoint="x",model="qwen3.5:9B",timeout=1,num_ctx=16384,num_predict=1024),resolver)
|
||||||
|
self.assertEqual(summary["target_resolution_llm_call_count"],0); resolver.assert_not_called()
|
||||||
|
for cid in ("tr-05","tr-06","tr-07","tr-08"):
|
||||||
|
skipped=json.loads((out/cid/"target_resolution_skipped.json").read_text()); self.assertFalse(skipped["call_made"])
|
||||||
|
|
||||||
|
def test_self_contained_target_equals_candidate(self):
|
||||||
|
c=BY_ID["TR-01"]; self.assertEqual(module.validate_target(target(c,text="Zusammenarbeit mit Dr. Schlummer fortsetzen"),c["observations"],"obs_1")["target_observation_id"],"obs_1")
|
||||||
|
|
||||||
|
def test_paired_target_precedes_candidate(self):
|
||||||
|
c=BY_ID["TR-02"]; self.assertEqual(module.validate_target(target(c,text="externe Lösung weiterverfolgen"),c["observations"],"obs_2")["target_observation_id"],"obs_1")
|
||||||
|
|
||||||
|
def test_target_after_candidate_rejected(self):
|
||||||
|
c=copy.deepcopy(BY_ID["TR-02"]); data={"candidate_observation_id":"obs_1","target_observation_id":"obs_2","normalized_target_text":"x"}
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_target(data,c["observations"],"obs_1")
|
||||||
|
|
||||||
|
def test_unknown_candidate_and_target_rejected(self):
|
||||||
|
c=BY_ID["TR-02"]
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_target({"candidate_observation_id":"missing","target_observation_id":"obs_1","normalized_target_text":"x"},c["observations"],"missing")
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_target({"candidate_observation_id":"obs_2","target_observation_id":"missing","normalized_target_text":"x"},c["observations"],"obs_2")
|
||||||
|
|
||||||
|
def test_duplicate_observation_and_evidence_ids_rejected(self):
|
||||||
|
for field in ("observation_id","evidence_id"):
|
||||||
|
obs=copy.deepcopy(BY_ID["TR-02"]["observations"]); obs[1][field]=obs[0][field]
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_observations(obs)
|
||||||
|
|
||||||
|
def test_null_and_non_null_text_constraints(self):
|
||||||
|
c=BY_ID["TR-02"]
|
||||||
|
valid={"candidate_observation_id":"obs_2","target_observation_id":None,"normalized_target_text":None}; self.assertEqual(module.validate_target(valid,c["observations"],"obs_2"),valid)
|
||||||
|
for bad in ({"candidate_observation_id":"obs_2","target_observation_id":None,"normalized_target_text":"x"},{"candidate_observation_id":"obs_2","target_observation_id":"obs_1","normalized_target_text":""}):
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_target(bad,c["observations"],"obs_2")
|
||||||
|
|
||||||
|
def test_unknown_and_recursive_forbidden_fields_rejected(self):
|
||||||
|
c=BY_ID["TR-02"]
|
||||||
|
for extra in ({"extra":1},{"nested":{"status":"rejected"}}):
|
||||||
|
data=target(c,text="x"); data.update(extra)
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_target(data,c["observations"],"obs_2")
|
||||||
|
|
||||||
|
def test_self_contained_prompt_fixes_linkage_deterministically(self):
|
||||||
|
prompt=module.build_prompt(BY_ID["TR-01"]); self.assertIn("deterministically fixed to obs_1",prompt); self.assertIn("only normalize",prompt)
|
||||||
|
|
||||||
|
def test_ineligible_prompt_is_impossible(self):
|
||||||
|
with self.assertRaises(DerivationValidationError): module.build_prompt(BY_ID["TR-05"])
|
||||||
|
|
||||||
|
def test_experiment_has_no_rejection_derivation(self):
|
||||||
|
self.assertFalse(hasattr(module,"derive")); self.assertNotIn("explicitly_rejected",module.TARGET_KEYS); self.assertNotIn("status",module.TARGET_KEYS)
|
||||||
|
|
||||||
|
def test_evaluation_preserves_scope_and_alternative_contract(self):
|
||||||
|
c=BY_ID["TR-04"]; gate=module.eligibility(c["negative_act"],c["observations"]); result=module.evaluate(c,gate,True,target(c,text="Versuch in der realen Anlage durchführen")); self.assertEqual(result["classification"],"PASS"); self.assertTrue(result["alternative_isolation"])
|
||||||
|
|
||||||
|
if __name__=="__main__": unittest.main()
|
||||||
@@ -0,0 +1,59 @@
|
|||||||
|
import argparse,copy,json,tempfile,unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import Mock
|
||||||
|
import src.meeting_lab.controlled_semantic_derivation.experiment_target_resolution_v1 as module
|
||||||
|
from src.meeting_lab.controlled_semantic_derivation.experiment_h import DerivationValidationError
|
||||||
|
|
||||||
|
CASES=module.load_cases(Path("tests/gold/target_resolution_v1/cases.json")); BY={c["case_id"]:c for c in CASES}
|
||||||
|
def self_output(text="Zusammenarbeit mit Dr. Schlummer fortsetzen"): return {"candidate_observation_id":"obs_1","normalized_target_text":text}
|
||||||
|
def paired(case,target="obs_1",text="externe Lösung weiterverfolgen"): return {"candidate_observation_id":case["negative_act"]["observation_id"],"target_observation_id":target,"normalized_target_text":text}
|
||||||
|
|
||||||
|
class TargetResolutionV1Tests(unittest.TestCase):
|
||||||
|
def test_self_linkage_is_deterministic_and_equals_candidate(self):
|
||||||
|
result=module.deterministic_self_link(BY["TR1-V1"]); self.assertEqual(result["linkage_source"],"deterministic"); self.assertEqual(result["target_observation_id"],result["candidate_observation_id"])
|
||||||
|
def test_self_llm_output_has_no_target_id(self):
|
||||||
|
self.assertEqual(module.SELF_KEYS,{"candidate_observation_id","normalized_target_text"}); self.assertNotIn("target_observation_id",module.output_schema(BY["TR1-V1"])["properties"])
|
||||||
|
def test_self_combination_records_link_and_normalization_separately(self):
|
||||||
|
combined=module.combine(BY["TR1-V1"],self_output()); self.assertEqual(combined["target_observation_id"],"obs_1"); self.assertIn("fortsetzen",combined["normalized_target_text"])
|
||||||
|
def test_self_link_rejects_paired_strategy(self):
|
||||||
|
with self.assertRaises(DerivationValidationError): module.deterministic_self_link(BY["TR2-V1"])
|
||||||
|
def test_paired_schema_enumerates_allowed_ids_and_null(self):
|
||||||
|
schema=module.output_schema(BY["TR2-V1"]); self.assertEqual(schema["properties"]["target_observation_id"]["enum"],["obs_1","obs_2",None]); self.assertFalse(schema["additionalProperties"])
|
||||||
|
def test_allowed_obs1_and_obs2_are_structurally_accepted(self):
|
||||||
|
case=BY["TR2-V1"]
|
||||||
|
module.validate_semantic_output(paired(case,"obs_1"),case); module.validate_semantic_output(paired(case,"obs_2"),case)
|
||||||
|
def test_unknown_and_string_null_targets_rejected(self):
|
||||||
|
case=BY["TR2-V1"]
|
||||||
|
for value in ("obs_9","null"):
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(paired(case,value),case)
|
||||||
|
def test_json_null_accepted_and_requires_null_text(self):
|
||||||
|
case=BY["TR2-V1"]; valid=paired(case,None,None); self.assertEqual(module.validate_semantic_output(valid,case),valid)
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(paired(case,None,"x"),case)
|
||||||
|
def test_non_null_requires_nonempty_text(self):
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(paired(BY["TR2-V1"],"obs_1",""),BY["TR2-V1"])
|
||||||
|
def test_target_after_candidate_rejected(self):
|
||||||
|
case=copy.deepcopy(BY["TR2-V1"]); case["negative_act"]["observation_id"]="obs_1"
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_semantic_output({"candidate_observation_id":"obs_1","target_observation_id":"obs_2","normalized_target_text":"x"},case)
|
||||||
|
def test_duplicate_ids_and_evidence_rejected(self):
|
||||||
|
for field in ("observation_id","evidence_id"):
|
||||||
|
obs=copy.deepcopy(BY["TR2-V1"]["observations"]); obs[1][field]=obs[0][field]
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_observations(obs)
|
||||||
|
def test_forbidden_and_unknown_fields_rejected(self):
|
||||||
|
case=BY["TR2-V1"]
|
||||||
|
for extra in ({"status":"x"},{"nested":{"decision":True}},{"extra":1}):
|
||||||
|
data=paired(case); data.update(extra)
|
||||||
|
with self.assertRaises(DerivationValidationError): module.validate_semantic_output(data,case)
|
||||||
|
def test_payload_uses_true_schema_object(self):
|
||||||
|
schema=module.output_schema(BY["TR2-V1"]); payload=module.build_payload("qwen3.5:9B","p",schema,16384,1024); self.assertIs(payload["format"],schema); self.assertIsInstance(payload["format"],dict); self.assertEqual(payload["options"]["temperature"],0)
|
||||||
|
def test_prompt_has_typed_examples_and_allowed_ids(self):
|
||||||
|
prompt=module.build_prompt(BY["TR2-V1"]); self.assertIn('["obs_1", "obs_2"]',prompt); self.assertIn('"target_observation_id":null',prompt); self.assertIn('Never return the string "null"',prompt); self.assertNotIn('observation ID or null',prompt)
|
||||||
|
def test_self_normalization_must_be_positive(self):
|
||||||
|
case=BY["TR1-V1"]; semantic=self_output("Mit Dr. Schlummer arbeiten wir nicht weiter."); combined=module.combine(case,semantic); self.assertEqual(module.evaluate(case,semantic,combined)["classification"],"FAIL")
|
||||||
|
def test_no_rejection_derivation_exists(self):
|
||||||
|
self.assertFalse(hasattr(module,"derive")); self.assertNotIn("status",module.PAIRED_KEYS); self.assertNotIn("explicitly_rejected",module.PAIRED_KEYS)
|
||||||
|
def test_runner_artifacts_distinguish_linkage_and_normalization(self):
|
||||||
|
case=BY["TR1-V1"]; fixture={"schema_version":module.SCHEMA_VERSION,"cases":[case]}; caller=Mock(return_value=(json.dumps(self_output()),{"model":"qwen3.5:9B"}))
|
||||||
|
with tempfile.TemporaryDirectory() as tmp:
|
||||||
|
root=Path(tmp); path=root/"cases.json"; path.write_text(json.dumps(fixture)); out=root/"out"; args=argparse.Namespace(cases=path,output=out,endpoint="x",model="qwen3.5:9B",timeout=1,num_ctx=16384,num_predict=1024); summary=module.run(args,caller); self.assertEqual(summary["llm_call_count"],1); self.assertTrue((out/"tr1-v1"/"deterministic_linkage_result.json").exists()); self.assertTrue((out/"tr1-v1"/"normalized_target_result.json").exists())
|
||||||
|
|
||||||
|
if __name__=="__main__": unittest.main()
|
||||||
Reference in New Issue
Block a user