diff --git a/docs/experiments.md b/docs/experiments.md index 0dc029f..d6cf429 100644 --- a/docs/experiments.md +++ b/docs/experiments.md @@ -1811,6 +1811,85 @@ production integration, group identity inference or another semantic category. Artifacts are preserved under `artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`. +## EXP-0034 — Explicit Rejection Gold V0 + +Status: Failed architecturally + +Date: 2026-08-20 + +This isolated Stage-2 experiment tested the narrow evidence fact that a +concrete action, option, proposal or future course was explicitly rejected, +abandoned, discontinued or ruled out. It used twelve synthetic cases containing +one self-contained observation or one local target/rejection pair. Evidence +Observation V3 was not called or changed. The accepted Request/Acceptance and +Collective Commitment paths remained unchanged and were not invoked. + +The strict semantic schema contains exactly `rejection_observation_id`, +`target_observation_id`, `rejection_form` and +`normalized_rejected_action_text`. `rejection_form` is closed to +`explicit_action_rejection` and `none`. A positive recognition requires a +known local target and non-empty normalized target; `none` requires both target +and normalized text to be null. Decision, outcome, topic-closure, +responsibility, ownership, protocol, confidence and graph fields are forbidden. +Target resolution is limited to the same observation or one earlier supplied +observation. Deterministic code validates schema, IDs, ordering and complete +provenance before emitting the narrow status `explicitly_rejected`. + +`explicitly_rejected` means rejected by the cited evidence only. It is not yet +a final meeting decision or final topic outcome, does not close a topic, and +does not supersede an earlier commitment. + +Gold results: + +- RJ-01 explicit collective rejection with local target: PASS. +- RJ-02 explicit non-pursuit with paired target: PASS. +- RJ-03 self-contained collaboration rejection: FAIL. The model returned + `none`, producing one recognition false negative. +- RJ-04 personal preference: FAIL. The model promoted the preference to an + explicit rejection and derived an unsupported rejection. +- RJ-05 concern: PASS; remained a non-rejection. +- RJ-06 uncertainty: PASS; remained a non-rejection. +- RJ-07 negative recommendation: FAIL. The model promoted advice to an + explicit rejection and derived an unsupported rejection. +- RJ-08 deferral: PASS; remained a non-rejection. +- RJ-09 factual negation: PASS; remained a non-rejection. +- RJ-10 temporary non-action: FAIL. The model treated `erstmal noch nicht` as + abandonment and derived an unsupported rejection. +- RJ-11 explicit rejection with material scope: PASS. Real-plant and + Druckversuch scope were preserved. +- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option + was rejected; the Technikum alternative was not absorbed. + +Configuration: exactly twelve successful sequential `qwen3.5:9B` calls, one +per case, temperature 0, `think=false`, `num_ctx=16384`, +`num_predict=1024`, no retries, no voting and no prompt changes. There were zero +technical failed calls. Aggregate runner time was 15.518 seconds; summed +per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681 +evaluation tokens. + +The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced +three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03). +There were four strict target-field expectation mismatches: three were +consequences of false-positive rejection objects populating otherwise locally +correct antecedents, and one was the missing self-contained RJ-03 target. No +derived positive selected the wrong concrete antecedent. Qualifier-loss count +was zero, positive-alternative absorption count was zero, and no responsibility, +decision, outcome or topic-closure field leaked into model output. + +Conclusion: the experiment is not architecturally successful. Deterministic +structural gates cannot contain a semantically well-formed false-positive +rejection with valid local target and provenance. The model did distinguish +concern, uncertainty, deferral and factual negation, and it handled scoped and +alternative-bearing positives correctly, but it did not reliably separate +explicit rejection from personal preference, advice or temporary non-action. +The current binary recognition `explicit_action_rejection | none` is +insufficient for reliable generalization. +No production integration, generic rejection system, prompt tuning or +cross-pattern reconciliation is justified. + +Artifacts are preserved under +`artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/`. + ## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype Date: 2026-08-11 diff --git a/scripts/run_explicit_rejection_gold_experiment.py b/scripts/run_explicit_rejection_gold_experiment.py new file mode 100644 index 0000000..9220778 --- /dev/null +++ b/scripts/run_explicit_rejection_gold_experiment.py @@ -0,0 +1,16 @@ +#!/usr/bin/env python3 +"""Repository entry point for the explicit-rejection Gold experiment.""" + +import sys +from pathlib import Path + + +REPO_ROOT = Path(__file__).resolve().parents[1] +if str(REPO_ROOT) not in sys.path: + sys.path.insert(0, str(REPO_ROOT)) + +from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import main # noqa: E402 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/meeting_lab/controlled_semantic_derivation/experiment_rejection.py b/src/meeting_lab/controlled_semantic_derivation/experiment_rejection.py new file mode 100644 index 0000000..ccbbfa1 --- /dev/null +++ b/src/meeting_lab/controlled_semantic_derivation/experiment_rejection.py @@ -0,0 +1,347 @@ +#!/usr/bin/env python3 +"""Isolated explicit-action-rejection Gold reliability experiment.""" + +from __future__ import annotations + +import argparse +import json +import time +from pathlib import Path +from typing import Any + +from .experiment_h import ( + DEFAULT_ENDPOINT, + DEFAULT_MODEL, + DerivationValidationError, + OBSERVATION_KEYS, + call_ollama, +) + + +GOLD_SCHEMA_VERSION = "experimental-explicit-rejection-gold-v0" +RECOGNITION_KEYS = { + "rejection_observation_id", "target_observation_id", "rejection_form", + "normalized_rejected_action_text", +} +REJECTION_FORMS = {"explicit_action_rejection", "none"} +FORBIDDEN_LLM_KEYS = { + "decision", "decision_status", "outcome", "topic_status", "closed", + "agreement", "responsible_person", "responsibility", "responsibility_scope", + "owner", "ownership", "assignee", "requested_actor", "status", + "explicitly_rejected", "action_item", "protocol", "protocol_category", + "confidence", "relation", "relations", "graph", "unresolved_issue", +} + +PROMPT_TEMPLATE = """Recognize only whether the candidate rejection observation explicitly rejects a concrete action, option, proposal, or future course of action in this small local set of V3-style observations. + +Answer only: +1. Does the candidate rejection observation explicitly reject, abandon, discontinue, or rule out a concrete action, option, proposal, or future course of action? +2. If yes, which supplied observation identifies the rejected target? +3. What is the concise normalized meaning of the rejected action or option? + +The candidate rejection observation is {rejection_observation_id}. + +Use explicit_action_rejection only for an asserted rejection, abandonment, discontinuation, or non-pursuit with a concrete locally resolvable target. Personal preference is not meeting-level explicit rejection. Concern or objection without refusal is not rejection. Uncertainty is not rejection. Negative recommendation or advice is not established rejection. Deferral is not rejection. "Not yet" or temporary non-action is not abandonment. Factual negation is not action rejection. Lack of commitment is not rejection. + +The rejected target may be self-contained in the candidate observation or introduced by one earlier supplied observation. Choose only among supplied observation IDs. If the target is ambiguous or unresolved, return rejection_form none. Preserve material scope limitations in normalized_rejected_action_text. Ignore a separate positive alternative when describing the rejected target. Keep normalized text in the observation language. + +Do not infer responsibility, ownership, decision status, final outcome, topic closure, protocol status, confidence, relations, graphs, or unresolved issues. Do not answer whether this was finally decided, what the meeting outcome was, who is responsible, or whether the topic is closed. + +Return exactly this JSON shape and no additional fields: +{{ + "rejection_observation_id": "{rejection_observation_id}", + "target_observation_id": "supplied observation ID" | null, + "rejection_form": "explicit_action_rejection | none", + "normalized_rejected_action_text": "concise rejected target" | null +}} + +For rejection_form none, target_observation_id and normalized_rejected_action_text must both be null. + +V3-style observations: +{observations_json} +""" + + +def _exact_keys(value: dict[str, Any], required: set[str], location: str) -> None: + missing = required - value.keys() + unknown = value.keys() - required + if missing: + raise DerivationValidationError(f"{location} missing required keys: {sorted(missing)}") + if unknown: + raise DerivationValidationError(f"{location} has unknown keys: {sorted(unknown)}") + + +def _nonempty_text(value: Any, location: str) -> str: + if not isinstance(value, str) or not value.strip(): + raise DerivationValidationError(f"{location} must be a non-empty string") + return value.strip() + + +def _validate_observations(observations: Any) -> None: + if not isinstance(observations, list) or not observations: + raise DerivationValidationError("observations must be a non-empty list") + seen_observations: set[str] = set() + seen_evidence: set[str] = set() + for index, observation in enumerate(observations): + location = f"observations[{index}]" + if not isinstance(observation, dict): + raise DerivationValidationError(f"{location} must be an object") + _exact_keys(observation, OBSERVATION_KEYS, location) + observation_id = _nonempty_text(observation["observation_id"], f"{location}.observation_id") + evidence_id = _nonempty_text(observation["evidence_id"], f"{location}.evidence_id") + if observation_id in seen_observations: + raise DerivationValidationError("observation IDs must be unique") + if evidence_id in seen_evidence: + raise DerivationValidationError("evidence provenance must be unique and consistent") + seen_observations.add(observation_id) + seen_evidence.add(evidence_id) + _nonempty_text(observation["content"], f"{location}.content") + _nonempty_text(observation["speaker"], f"{location}.speaker") + for field in ("named_person", "addressee"): + if observation[field] is not None: + _nonempty_text(observation[field], f"{location}.{field}") + + +def load_gold_cases(path: Path) -> list[dict[str, Any]]: + data = json.loads(path.read_text(encoding="utf-8-sig")) + if not isinstance(data, dict): + raise DerivationValidationError("Gold fixture must be an object") + _exact_keys(data, {"schema_version", "cases"}, "Gold fixture") + if data["schema_version"] != GOLD_SCHEMA_VERSION: + raise DerivationValidationError("unexpected Gold fixture schema_version") + cases = data["cases"] + if not isinstance(cases, list) or not cases: + raise DerivationValidationError("Gold fixture cases must be a non-empty list") + seen: set[str] = set() + for case in cases: + _exact_keys(case, {"case_id", "description", "observations", "expected_recognition", "expected_result"}, "Gold case") + case_id = _nonempty_text(case["case_id"], "Gold case.case_id") + if case_id in seen: + raise DerivationValidationError(f"duplicate case ID: {case_id}") + seen.add(case_id) + _validate_observations(case["observations"]) + if len(case["observations"]) not in (1, 2): + raise DerivationValidationError("rejection Gold cases require one or two observations") + return cases + + +def build_prompt(case: dict[str, Any]) -> str: + observations = case["observations"] + _validate_observations(observations) + rejection_observation_id = observations[-1]["observation_id"] + return PROMPT_TEMPLATE.format( + rejection_observation_id=rejection_observation_id, + observations_json=json.dumps(observations, ensure_ascii=False, indent=2), + ) + + +def parse_model_json(raw_text: str) -> dict[str, Any]: + data = json.loads(raw_text) + if not isinstance(data, dict): + raise DerivationValidationError("semantic recognition must be an object") + return data + + +def _reject_forbidden_keys(value: Any, location: str = "output") -> None: + if isinstance(value, dict): + forbidden = FORBIDDEN_LLM_KEYS.intersection(value) + if forbidden: + raise DerivationValidationError(f"{location} contains forbidden semantic keys: {sorted(forbidden)}") + for key, item in value.items(): + _reject_forbidden_keys(item, f"{location}.{key}") + elif isinstance(value, list): + for index, item in enumerate(value): + _reject_forbidden_keys(item, f"{location}[{index}]") + + +def validate_recognition(data: Any, observations: list[dict[str, Any]]) -> dict[str, Any]: + _validate_observations(observations) + if not isinstance(data, dict): + raise DerivationValidationError("semantic recognition must be an object") + _reject_forbidden_keys(data) + _exact_keys(data, RECOGNITION_KEYS, "output") + rejection_id = _nonempty_text(data["rejection_observation_id"], "output.rejection_observation_id") + known_ids = {item["observation_id"] for item in observations} + if rejection_id not in known_ids: + raise DerivationValidationError("unknown rejection observation ID") + form = data["rejection_form"] + if form not in REJECTION_FORMS: + raise DerivationValidationError("rejection_form has an unsupported value") + target_id = data["target_observation_id"] + action_text = data["normalized_rejected_action_text"] + if form == "none": + if target_id is not None: + raise DerivationValidationError("none rejection must have null target_observation_id") + if action_text is not None: + raise DerivationValidationError("none rejection must have null normalized_rejected_action_text") + else: + target_id = _nonempty_text(target_id, "output.target_observation_id") + if target_id not in known_ids: + raise DerivationValidationError("unknown target observation ID") + _nonempty_text(action_text, "output.normalized_rejected_action_text") + return data + + +def derive_rejection( + observations: list[dict[str, Any]], recognition: dict[str, Any] +) -> tuple[dict[str, bool], dict[str, Any] | None]: + validate_recognition(recognition, observations) + by_id = {item["observation_id"]: item for item in observations} + positions = {item["observation_id"]: index for index, item in enumerate(observations)} + rejection = by_id.get(recognition["rejection_observation_id"]) + target_id = recognition["target_observation_id"] + target = by_id.get(target_id) if target_id is not None else None + gates = { + "recognition_schema_valid": True, + "explicit_action_rejection": recognition["rejection_form"] == "explicit_action_rejection", + "rejection_observation_exists": rejection is not None, + "target_observation_exists": target is not None, + "observation_ids_valid_and_unique": len(by_id) == len(observations), + "evidence_provenance_valid_unique_consistent": len({item["evidence_id"] for item in observations}) == len(observations), + "target_same_or_before_rejection": target is not None and rejection is not None and positions[target["observation_id"]] <= positions[rejection["observation_id"]], + "normalized_rejected_action_present": isinstance(recognition["normalized_rejected_action_text"], str) and bool(recognition["normalized_rejected_action_text"].strip()), + "target_local_to_case": target_id in by_id if target_id is not None else False, + "schema_state_consistent": recognition["rejection_form"] == "explicit_action_rejection" and target_id is not None, + "referenced_provenance_available": target is not None and rejection is not None and bool(target["evidence_id"]) and bool(rejection["evidence_id"]), + } + if not all(gates.values()): + return gates, None + return gates, { + "rejection_id": "rejection_1", + "content": recognition["normalized_rejected_action_text"].strip(), + "status": "explicitly_rejected", + "support": { + "target": {"observation_id": target["observation_id"], "evidence_id": target["evidence_id"]}, + "rejection": {"observation_id": rejection["observation_id"], "evidence_id": rejection["evidence_id"]}, + }, + } + + +def _concepts_present(text: str | None, concepts: list[list[str]]) -> bool: + if not concepts: + return True + if not isinstance(text, str): + return False + folded = text.casefold() + return all(any(alias.casefold() in folded for alias in alternatives) for alternatives in concepts) + + +def _contains_forbidden_concept(text: str | None, concepts: list[str]) -> bool: + return isinstance(text, str) and any(concept.casefold() in text.casefold() for concept in concepts) + + +def evaluate_case(case: dict[str, Any], recognition: dict[str, Any]) -> dict[str, Any]: + validate_recognition(recognition, case["observations"]) + gates, result = derive_rejection(case["observations"], recognition) + expected = case["expected_recognition"] + expected_result = case["expected_result"] + form_correct = recognition["rejection_form"] == expected["rejection_form"] + rejection_observation_correct = recognition["rejection_observation_id"] == expected["rejection_observation_id"] + target_correct = recognition["target_observation_id"] == expected["target_observation_id"] + action_correct = _concepts_present(recognition["normalized_rejected_action_text"], expected["action_concepts"]) + qualifier_preserved = _concepts_present(recognition["normalized_rejected_action_text"], expected["qualifier_concepts"]) + alternative_absorbed = _contains_forbidden_concept(recognition["normalized_rejected_action_text"], expected["forbidden_action_concepts"]) + derived = result is not None + final_correct = derived == expected_result["explicitly_rejected"] + if result is not None: + final_correct = final_correct and result["status"] == "explicitly_rejected" + semantic_correct = form_correct and rejection_observation_correct and target_correct and action_correct and qualifier_preserved and not alternative_absorbed + automatic_failure = (derived and not expected_result["explicitly_rejected"]) or (derived and not target_correct) or (derived and not qualifier_preserved) or alternative_absorbed + classification = "FAIL" if automatic_failure or not final_correct else ("PASS" if semantic_correct else "PARTIAL") + return { + "case_id": case["case_id"], "classification": classification, + "rejection_form_correct": form_correct, + "rejection_observation_correct": rejection_observation_correct, + "target_observation_correct": target_correct, + "normalized_rejected_action_correct": action_correct, + "material_qualifiers_preserved": qualifier_preserved, + "positive_alternative_absorbed": alternative_absorbed, + "deterministic_gates_correct": final_correct, + "final_result_correct": final_correct, + "unsupported_semantic_strengthening": recognition["rejection_form"] == "explicit_action_rejection" and expected["rejection_form"] == "none", + "normative_leakage": False, + "gates": gates, "result": result, + } + + +def _write_json(path: Path, value: Any) -> None: + path.write_text(json.dumps(value, ensure_ascii=False, indent=2) + "\n", encoding="utf-8") + + +def run_gold(args: argparse.Namespace) -> dict[str, Any]: + cases = load_gold_cases(args.cases) + args.output.mkdir(parents=True, exist_ok=False) + _write_json(args.output / "gold_cases.json", {"schema_version": GOLD_SCHEMA_VERSION, "cases": cases}) + evaluations: list[dict[str, Any]] = [] + successful_calls = 0 + technical_failures = 0 + started = time.perf_counter() + for case in cases: + case_dir = args.output / case["case_id"].lower() + case_dir.mkdir() + observations = case["observations"] + _write_json(case_dir / "v3_style_input_observations.json", observations) + prompt = build_prompt(case) + (case_dir / "prompt.txt").write_text(prompt, encoding="utf-8") + try: + raw, metadata = call_ollama(args.endpoint, args.model, prompt, args.timeout, args.num_ctx, args.num_predict) + successful_calls += 1 + except Exception as exc: # one recorded attempt; never retry + technical_failures += 1 + failure = {"case_id": case["case_id"], "classification": "FAIL", "technical_failure": True, "error_type": type(exc).__name__, "error": str(exc)} + _write_json(case_dir / "ollama_metadata.json", {"model": args.model, "configuration": {"temperature": 0, "think": False, "num_ctx": args.num_ctx, "num_predict": args.num_predict, "retries": 0}, "technical_failure": failure}) + _write_json(case_dir / "structural_validation.json", {"valid": False, "error": str(exc)}) + _write_json(case_dir / "deterministic_gate_results.json", {}) + _write_json(case_dir / "final_derived_result.json", None) + _write_json(case_dir / "evaluation.json", failure) + evaluations.append(failure) + continue + (case_dir / "raw_model_response.txt").write_text(raw + "\n", encoding="utf-8") + _write_json(case_dir / "ollama_metadata.json", metadata) + try: + parsed = parse_model_json(raw) + _write_json(case_dir / "parsed_semantic_recognition.json", parsed) + evaluation = evaluate_case(case, parsed) + validation = {"valid": True, "error": None} + gates, result = derive_rejection(observations, parsed) + except (DerivationValidationError, json.JSONDecodeError) as exc: + validation = {"valid": False, "error_type": type(exc).__name__, "error": str(exc)} + evaluation = {"case_id": case["case_id"], "classification": "FAIL", "error": str(exc), "normative_leakage": "forbidden" in str(exc)} + gates, result = {}, None + _write_json(case_dir / "structural_validation.json", validation) + _write_json(case_dir / "deterministic_gate_results.json", gates) + _write_json(case_dir / "final_derived_result.json", result) + _write_json(case_dir / "evaluation.json", evaluation) + evaluations.append(evaluation) + summary = { + "experiment": "explicit_rejection_gold_v0", "model": args.model, + "successful_llm_call_count": successful_calls, + "technical_failed_call_count": technical_failures, + "runtime_seconds": round(time.perf_counter() - started, 3), + "counts": {label: sum(item["classification"] == label for item in evaluations) for label in ("PASS", "PARTIAL", "FAIL")}, + "evaluations": evaluations, + } + _write_json(args.output / "summary.json", summary) + return summary + + +def parse_args() -> argparse.Namespace: + parser = argparse.ArgumentParser(description="Run isolated explicit-rejection Gold experiment") + parser.add_argument("cases", type=Path) + parser.add_argument("-o", "--output", type=Path, required=True) + parser.add_argument("--model", default=DEFAULT_MODEL) + parser.add_argument("--endpoint", default=DEFAULT_ENDPOINT) + parser.add_argument("--timeout", type=int, default=300) + parser.add_argument("--num-ctx", type=int, default=16384) + parser.add_argument("--num-predict", type=int, default=1024) + return parser.parse_args() + + +def main() -> int: + summary = run_gold(parse_args()) + print(json.dumps(summary, ensure_ascii=False, indent=2)) + return 0 if summary["counts"]["FAIL"] == 0 and summary["technical_failed_call_count"] == 0 else 1 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/tests/gold/explicit_rejection_v0/cases.json b/tests/gold/explicit_rejection_v0/cases.json new file mode 100644 index 0000000..63c0da4 --- /dev/null +++ b/tests/gold/explicit_rejection_v0/cases.json @@ -0,0 +1,111 @@ +{ + "schema_version": "experimental-explicit-rejection-gold-v0", + "cases": [ + { + "case_id": "RJ-01", "description": "Explicit collective rejection with local target", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Nein, das machen wir nicht.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"], ["versuch", "trial", "test"]], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": true} + }, + { + "case_id": "RJ-02", "description": "Explicit non-pursuit with paired target", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die externe Lösung weiterzuverfolgen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das verfolgen wir nicht weiter.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["extern"], ["lösung", "solution"], ["weiter", "pursu", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": true} + }, + { + "case_id": "RJ-03", "description": "Self-contained collaboration rejection", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Mit Dr. Schlummer arbeiten wir nicht weiter.", "speaker": "Martin", "named_person": "Dr. Schlummer", "addressee": null} + ], + "expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_1", "target_observation_id": "obs_1", "action_concepts": [["schlummer"], ["arbeit", "collabor"], ["weiter", "fortsetz", "continu"]], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": true} + }, + { + "case_id": "RJ-04", "description": "Personal preference", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage für den Versuch nutzen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde das nicht machen.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-05", "description": "Concern", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten das neue Material einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das wäre kritisch.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-06", "description": "Uncertainty", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Eine Möglichkeit wäre, die Waschstufe einzubauen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich weiß nicht, ob das sinnvoll ist.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-07", "description": "Negative recommendation", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die reale Anlage verwenden.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Ich würde eher davon abraten.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-08", "description": "Deferral", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die externe Lösung einsetzen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das entscheiden wir nächste Woche.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-09", "description": "Factual negation", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Das Material ist nicht verfügbar.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_1", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-10", "description": "Temporary non-action", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten die Waschstufe einbauen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir erstmal noch nicht.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "none", "rejection_observation_id": "obs_2", "target_observation_id": null, "action_concepts": [], "qualifier_concepts": [], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": false} + }, + { + "case_id": "RJ-11", "description": "Explicit rejection with material scope", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Für den Druckversuch steht die reale Anlage zur Diskussion.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Die reale Anlage nutzen wir dafür nicht.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["real"], ["anlage", "plant"]], "qualifier_concepts": [["druckversuch", "dafür", "pressure test"]], "forbidden_action_concepts": []}, + "expected_result": {"explicitly_rejected": true} + }, + { + "case_id": "RJ-12", "description": "Rejection plus positive alternative", + "observations": [ + {"observation_id": "obs_1", "evidence_id": "e1", "content": "Martin: Wir könnten den Versuch in der realen Anlage durchführen.", "speaker": "Martin", "named_person": null, "addressee": null}, + {"observation_id": "obs_2", "evidence_id": "e2", "content": "Martin: Das machen wir nicht; wir testen stattdessen im Technikum.", "speaker": "Martin", "named_person": null, "addressee": null} + ], + "expected_recognition": {"rejection_form": "explicit_action_rejection", "rejection_observation_id": "obs_2", "target_observation_id": "obs_1", "action_concepts": [["versuch", "trial", "test"], ["real"], ["anlage", "plant"]], "qualifier_concepts": [["real"], ["anlage", "plant"]], "forbidden_action_concepts": ["technikum", "technical facility", "technical center", "technical centre"]}, + "expected_result": {"explicitly_rejected": true} + } + ] +} diff --git a/tests/test_explicit_rejection_gold_experiment.py b/tests/test_explicit_rejection_gold_experiment.py new file mode 100644 index 0000000..2ca8194 --- /dev/null +++ b/tests/test_explicit_rejection_gold_experiment.py @@ -0,0 +1,210 @@ +import json +import tempfile +import unittest +from copy import deepcopy +from pathlib import Path + +from src.meeting_lab.controlled_semantic_derivation.experiment_rejection import ( + DerivationValidationError, + build_prompt, + derive_rejection, + evaluate_case, + load_gold_cases, + validate_recognition, +) + + +GOLD_PATH = Path("tests/gold/explicit_rejection_v0/cases.json") + + +POSITIVE_TEXT = { + "RJ-01": "reale Anlage für den Versuch nutzen", + "RJ-02": "externe Lösung weiterverfolgen", + "RJ-03": "Zusammenarbeit mit Dr. Schlummer fortsetzen", + "RJ-11": "reale Anlage für den Druckversuch nutzen", + "RJ-12": "Versuch in der realen Anlage durchführen", +} + + +def recognition_for(case): + expected = case["expected_recognition"] + positive = expected["rejection_form"] == "explicit_action_rejection" + return { + "rejection_observation_id": expected["rejection_observation_id"], + "target_observation_id": expected["target_observation_id"] if positive else None, + "rejection_form": expected["rejection_form"], + "normalized_rejected_action_text": POSITIVE_TEXT.get(case["case_id"]) if positive else None, + } + + +class ExplicitRejectionGoldExperimentTests(unittest.TestCase): + @classmethod + def setUpClass(cls): + cls.cases = load_gold_cases(GOLD_PATH) + cls.by_id = {case["case_id"]: case for case in cls.cases} + + def test_fixture_contains_exactly_rj_01_through_rj_12(self): + self.assertEqual(list(self.by_id), [f"RJ-{number:02d}" for number in range(1, 13)]) + + def test_cases_use_only_minimal_v3_style_observations(self): + keys = {"observation_id", "evidence_id", "content", "speaker", "named_person", "addressee"} + for case in self.cases: + with self.subTest(case=case["case_id"]): + self.assertIn(len(case["observations"]), (1, 2)) + self.assertTrue(all(set(item) == keys for item in case["observations"])) + + def test_rj_01_derives_target_and_both_provenance_paths(self): + case = self.by_id["RJ-01"] + gates, result = derive_rejection(case["observations"], recognition_for(case)) + self.assertTrue(all(gates.values())) + self.assertEqual(result["status"], "explicitly_rejected") + self.assertEqual(result["support"]["target"], {"observation_id": "obs_1", "evidence_id": "e1"}) + self.assertEqual(result["support"]["rejection"], {"observation_id": "obs_2", "evidence_id": "e2"}) + + def test_rj_02_requires_paired_target_and_derives_abandonment(self): + case = self.by_id["RJ-02"] + _, result = derive_rejection(case["observations"], recognition_for(case)) + self.assertIn("externe Lösung", result["content"]) + with self.assertRaisesRegex(DerivationValidationError, "unknown target"): + derive_rejection(case["observations"][1:], recognition_for(case)) + + def test_rj_03_supports_same_observation_target_and_rejection(self): + case = self.by_id["RJ-03"] + _, result = derive_rejection(case["observations"], recognition_for(case)) + self.assertEqual(result["support"]["target"], result["support"]["rejection"]) + self.assertIn("Dr. Schlummer", result["content"]) + + def test_all_required_negative_cases_remain_non_rejections(self): + for case_id in ("RJ-04", "RJ-05", "RJ-06", "RJ-07", "RJ-08", "RJ-09", "RJ-10"): + case = self.by_id[case_id] + gates, result = derive_rejection(case["observations"], recognition_for(case)) + with self.subTest(case=case_id): + self.assertFalse(gates["explicit_action_rejection"]) + self.assertIsNone(result) + + def test_rj_11_derives_and_preserves_location_and_purpose_scope(self): + case = self.by_id["RJ-11"] + _, result = derive_rejection(case["observations"], recognition_for(case)) + self.assertIsNotNone(result) + self.assertIn("reale Anlage", result["content"]) + self.assertIn("Druckversuch", result["content"]) + + def test_rj_12_rejects_only_real_plant_action_and_not_alternative(self): + case = self.by_id["RJ-12"] + _, result = derive_rejection(case["observations"], recognition_for(case)) + self.assertIn("realen Anlage", result["content"]) + self.assertNotIn("Technikum", result["content"]) + + def test_separate_target_cannot_follow_rejection(self): + case = deepcopy(self.by_id["RJ-01"]) + case["observations"].reverse() + gates, result = derive_rejection(case["observations"], recognition_for(case)) + self.assertFalse(gates["target_same_or_before_rejection"]) + self.assertIsNone(result) + + def test_unknown_target_observation_id_is_rejected(self): + case = self.by_id["RJ-01"] + recognition = recognition_for(case) + recognition["target_observation_id"] = "obs_99" + with self.assertRaisesRegex(DerivationValidationError, "unknown target"): + validate_recognition(recognition, case["observations"]) + + def test_unknown_rejection_observation_id_is_rejected(self): + case = self.by_id["RJ-01"] + recognition = recognition_for(case) + recognition["rejection_observation_id"] = "obs_99" + with self.assertRaisesRegex(DerivationValidationError, "unknown rejection"): + validate_recognition(recognition, case["observations"]) + + def test_duplicate_observation_ids_are_rejected(self): + fixture = json.loads(GOLD_PATH.read_text()) + fixture["cases"][0]["observations"][1]["observation_id"] = "obs_1" + self._assert_bad_fixture(fixture, "observation IDs must be unique") + + def test_inconsistent_evidence_provenance_is_rejected(self): + fixture = json.loads(GOLD_PATH.read_text()) + fixture["cases"][0]["observations"][1]["evidence_id"] = "e1" + self._assert_bad_fixture(fixture, "evidence provenance must be unique") + + def test_none_rejects_populated_target_or_action(self): + case = self.by_id["RJ-04"] + for field, value, message in ( + ("target_observation_id", "obs_1", "null target"), + ("normalized_rejected_action_text", "Anlage nutzen", "null normalized"), + ): + recognition = recognition_for(case) + recognition[field] = value + with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, message): + validate_recognition(recognition, case["observations"]) + + def test_explicit_rejection_requires_normalized_target_text(self): + case = self.by_id["RJ-01"] + for value in (None, ""): + recognition = recognition_for(case) + recognition["normalized_rejected_action_text"] = value + with self.subTest(value=value), self.assertRaises(DerivationValidationError): + validate_recognition(recognition, case["observations"]) + + def test_unknown_schema_fields_are_rejected(self): + case = self.by_id["RJ-01"] + recognition = recognition_for(case) + recognition["explanation"] = "extra" + with self.assertRaisesRegex(DerivationValidationError, "unknown keys"): + validate_recognition(recognition, case["observations"]) + + def test_forbidden_normative_fields_are_rejected_recursively(self): + case = self.by_id["RJ-01"] + fields = ( + "decision", "decision_status", "outcome", "topic_status", "closed", + "agreement", "responsible_person", "responsibility", "responsibility_scope", + "owner", "ownership", "assignee", "requested_actor", "status", + "explicitly_rejected", "action_item", "protocol_category", "confidence", + "relation", "relations", "graph", "unresolved_issue", + ) + for field in fields: + recognition = recognition_for(case) + recognition["wrapper"] = {field: "forbidden"} + with self.subTest(field=field), self.assertRaisesRegex(DerivationValidationError, "forbidden semantic keys"): + validate_recognition(recognition, case["observations"]) + + def test_speaker_identity_creates_no_ownership_or_responsibility(self): + case = deepcopy(self.by_id["RJ-03"]) + for speaker in ("Martin", "Clara", "Antonius"): + case["observations"][0]["speaker"] = speaker + _, result = derive_rejection(case["observations"], recognition_for(case)) + with self.subTest(speaker=speaker): + self.assertNotIn("responsible_person", result) + self.assertNotIn("owner", result) + + def test_all_expected_recognitions_evaluate_as_pass(self): + for case in self.cases: + evaluation = evaluate_case(case, recognition_for(case)) + with self.subTest(case=case["case_id"]): + self.assertEqual(evaluation["classification"], "PASS") + + def test_rj_11_qualifier_loss_and_rj_12_alternative_absorption_fail(self): + rj11 = self.by_id["RJ-11"] + recognition = recognition_for(rj11) + recognition["normalized_rejected_action_text"] = "reale Anlage nutzen" + self.assertEqual(evaluate_case(rj11, recognition)["classification"], "FAIL") + rj12 = self.by_id["RJ-12"] + recognition = recognition_for(rj12) + recognition["normalized_rejected_action_text"] += "; stattdessen im Technikum testen" + self.assertEqual(evaluate_case(rj12, recognition)["classification"], "FAIL") + + def test_prompt_is_fixed_narrow_and_does_not_expose_gold_expectation(self): + prompt = build_prompt(self.by_id["RJ-01"]) + self.assertIn("candidate rejection observation is obs_2", prompt) + self.assertNotIn("expected_result", prompt) + self.assertNotIn("Who is responsible", prompt) + + def _assert_bad_fixture(self, fixture, message): + with tempfile.TemporaryDirectory() as temporary: + path = Path(temporary) / "cases.json" + path.write_text(json.dumps(fixture), encoding="utf-8") + with self.assertRaisesRegex(DerivationValidationError, message): + load_gold_cases(path) + + +if __name__ == "__main__": + unittest.main()