Document failed explicit rejection experiment
This commit is contained in:
@@ -1811,6 +1811,85 @@ production integration, group identity inference or another semantic category.
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/collective_commitment_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0034 — Explicit Rejection Gold V0
|
||||
|
||||
Status: Failed architecturally
|
||||
|
||||
Date: 2026-08-20
|
||||
|
||||
This isolated Stage-2 experiment tested the narrow evidence fact that a
|
||||
concrete action, option, proposal or future course was explicitly rejected,
|
||||
abandoned, discontinued or ruled out. It used twelve synthetic cases containing
|
||||
one self-contained observation or one local target/rejection pair. Evidence
|
||||
Observation V3 was not called or changed. The accepted Request/Acceptance and
|
||||
Collective Commitment paths remained unchanged and were not invoked.
|
||||
|
||||
The strict semantic schema contains exactly `rejection_observation_id`,
|
||||
`target_observation_id`, `rejection_form` and
|
||||
`normalized_rejected_action_text`. `rejection_form` is closed to
|
||||
`explicit_action_rejection` and `none`. A positive recognition requires a
|
||||
known local target and non-empty normalized target; `none` requires both target
|
||||
and normalized text to be null. Decision, outcome, topic-closure,
|
||||
responsibility, ownership, protocol, confidence and graph fields are forbidden.
|
||||
Target resolution is limited to the same observation or one earlier supplied
|
||||
observation. Deterministic code validates schema, IDs, ordering and complete
|
||||
provenance before emitting the narrow status `explicitly_rejected`.
|
||||
|
||||
`explicitly_rejected` means rejected by the cited evidence only. It is not yet
|
||||
a final meeting decision or final topic outcome, does not close a topic, and
|
||||
does not supersede an earlier commitment.
|
||||
|
||||
Gold results:
|
||||
|
||||
- RJ-01 explicit collective rejection with local target: PASS.
|
||||
- RJ-02 explicit non-pursuit with paired target: PASS.
|
||||
- RJ-03 self-contained collaboration rejection: FAIL. The model returned
|
||||
`none`, producing one recognition false negative.
|
||||
- RJ-04 personal preference: FAIL. The model promoted the preference to an
|
||||
explicit rejection and derived an unsupported rejection.
|
||||
- RJ-05 concern: PASS; remained a non-rejection.
|
||||
- RJ-06 uncertainty: PASS; remained a non-rejection.
|
||||
- RJ-07 negative recommendation: FAIL. The model promoted advice to an
|
||||
explicit rejection and derived an unsupported rejection.
|
||||
- RJ-08 deferral: PASS; remained a non-rejection.
|
||||
- RJ-09 factual negation: PASS; remained a non-rejection.
|
||||
- RJ-10 temporary non-action: FAIL. The model treated `erstmal noch nicht` as
|
||||
abandonment and derived an unsupported rejection.
|
||||
- RJ-11 explicit rejection with material scope: PASS. Real-plant and
|
||||
Druckversuch scope were preserved.
|
||||
- RJ-12 rejection plus positive alternative: PASS. Only the real-plant option
|
||||
was rejected; the Technikum alternative was not absorbed.
|
||||
|
||||
Configuration: exactly twelve successful sequential `qwen3.5:9B` calls, one
|
||||
per case, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=1024`, no retries, no voting and no prompt changes. There were zero
|
||||
technical failed calls. Aggregate runner time was 15.518 seconds; summed
|
||||
per-call time was 15.493 seconds, with 6,972 prompt-evaluation tokens and 681
|
||||
evaluation tokens.
|
||||
|
||||
The outcome was eight PASS, zero PARTIAL and four FAIL. Recognition produced
|
||||
three false positives (RJ-04, RJ-07 and RJ-10) and one false negative (RJ-03).
|
||||
There were four strict target-field expectation mismatches: three were
|
||||
consequences of false-positive rejection objects populating otherwise locally
|
||||
correct antecedents, and one was the missing self-contained RJ-03 target. No
|
||||
derived positive selected the wrong concrete antecedent. Qualifier-loss count
|
||||
was zero, positive-alternative absorption count was zero, and no responsibility,
|
||||
decision, outcome or topic-closure field leaked into model output.
|
||||
|
||||
Conclusion: the experiment is not architecturally successful. Deterministic
|
||||
structural gates cannot contain a semantically well-formed false-positive
|
||||
rejection with valid local target and provenance. The model did distinguish
|
||||
concern, uncertainty, deferral and factual negation, and it handled scoped and
|
||||
alternative-bearing positives correctly, but it did not reliably separate
|
||||
explicit rejection from personal preference, advice or temporary non-action.
|
||||
The current binary recognition `explicit_action_rejection | none` is
|
||||
insufficient for reliable generalization.
|
||||
No production integration, generic rejection system, prompt tuning or
|
||||
cross-pattern reconciliation is justified.
|
||||
|
||||
Artifacts are preserved under
|
||||
`artifacts/experiments/explicit_rejection_gold_v0/20260820_qwen35_9b_single_run/`.
|
||||
|
||||
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
|
||||
|
||||
Date: 2026-08-11
|
||||
|
||||
Reference in New Issue
Block a user