Add evidence-near semantic architecture experiments

Record the V1-V3 experiments and accept the minimal semantic-preservation first stage.
This commit is contained in:
2026-08-19 15:46:22 +02:00
parent bcb197a908
commit 18beb3385f
29 changed files with 4542 additions and 0 deletions
+290
View File
@@ -1350,3 +1350,293 @@ accordance with the Gold Standard methodology.
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
remains Open; no phrase-specific deterministic filter was introduced.
## EXP-0027 — Evidence-near observation extraction
Date: 2026-08-18
Hypothesis: `qwen3.5:9B` can more reliably extract evidence-near linguistic and
semantic properties than directly synthesize protocol-level events, outcomes,
actions and unresolved issues. This isolated experiment stops before semantic
interpretation and does not connect to the production pipeline.
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
source evidence, explicit targets, a five-value relation vocabulary, modality,
temporality, evaluation, agreement, responsibility/person, uncertainty,
clarification need and free-text scope. It contains no protocol-level category
field. The validator requires sequential observation IDs, known evidence IDs,
backward-only valid observation targets, closed categorical vocabularies,
consistent responsibility/person pairs, and one canonical absence form: JSON
null for person and `absent` for scope.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries. All nine cases ran exactly once, for nine LLM
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
responses, parsed observations, validation results, Ollama metadata and
comparisons are preserved under
`/tmp/meeting-lab-evidence-observations-v1-20260818/`.
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
Five cases failed structure because the model represented a single target as a
one-element list, usually `["discussion_subject"]`; the accepted schema permits
a list only for two or more jointly referenced observations. Several responses
also copied the relation label `limits_scope` into the free-text scope field.
These were systematic model-output errors, not transport or parser failures.
The prompt and run were not retried or tuned.
Human semantic review of the preserved raw responses:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override
strict structural failures; it separates useful semantic signal from schema
compliance.
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving
closer to evidence reduced some direct promotion behavior: the geometry mention
did not become work, the washing disadvantage did not become an unresolved
issue, both alternatives in B remained uncommitted, and the publication query
did not become an assignment. However, the important promotion errors did not
disappear. C acquired unsupported acceptance, and G still promoted a personal
preference into rejection. Positive cases were only partly preserved: explicit
rejection was recognized but incorrectly linked; trial-only language was kept
but responsibility was invented; Nina's request and commitment were recognized
but responsibility states were wrong; and the publication issue was recognized
but its uncertainty was contradicted by rejection.
Result: **B — evidence-near extraction is promising, but specific observation
dimensions remain unreliable.** Target/relation selection, scope attachment,
responsibility state/person attribution, and agreement versus uncertainty are
not reliable enough to justify designing the later interpretation stage yet.
No production integration or later interpretation stage was implemented.
## EXP-0028 — Evidence-Near Observation Extraction V2
Date: 2026-08-19
V2 tested whether `qwen3.5:9B` preserves the evidence needed by a later
controlled interpretation stage when direct responsibility, agreement and
semantic graph relations are removed. Responsibility was replaced by explicit
participant/discourse facts (`speaker`, `named_person`, `addressee`, singular
self-reference, collective `we`, and impersonal person reference). Agreement
was replaced by explicit affirmation, explicit negation and determination
statement signals. Graph relations were reduced to nullable scalar
`refers_to`; scope became free-text `qualifier` plus nullable scalar
`limits_target`. No later derivation stage was implemented.
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
human interpretations. It contains no responsibility, agreement, action,
decision, open-question, accepted-trial, rejected-alternative or protocol
eligibility fields. Validation enforces known evidence IDs, sequential unique
observation IDs, backward-only scalar references, closed vocabularies, boolean
participant flags, JSON-nullable participant/qualifier/reference fields and no
string `"null"`.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries or voting. A launch-path defect was corrected
before the live run; the failed launch made zero model calls. A sandbox-blocked
localhost attempt also made zero model calls. The completed run called the
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
Persistent prompts, Gold input and expectations, raw and parsed model output,
validation, automatic comparison, Ollama metadata and human evaluation are in
`artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/`.
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
were schema-invalid. The dominant serialization pattern was use of `present`
instead of the specified `explicit` for affirmation/negation; E additionally
used `none` instead of `absent` for a determination signal, while D emitted the
separate uncertainty concept as an invalid modality. These errors are
contract violations, although most `present`/`explicit` differences are
deterministically normalizable without changing meaning. A and C were valid
JSON/schema outputs but had critical semantic mismatches.
Human semantic review:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; `present` caused schema failure. |
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced
V1 promotion errors: collective speech and speaker identity no longer became
individual responsibility; personal preference no longer became a group-level
rejection field; an information question did not become work; and explicit
negation/affirmation survived as separate evidence. Useful participant evidence
also survived strongly in H and collective-speech evidence in F.
Simplification did not make all evidence-near dimensions reliable. Scalar
references and `limits_target` were almost entirely omitted, qualifiers were
usually omitted, committed modality was missed in E, F and H, and C/G repeated
important modality, uncertainty and participant-feature errors. Some positive
semantic information therefore survived only in free-text `content`, not in
the structural signals a controlled derivation stage would need.
Result: **B — V2 is materially better, but specific evidence-near dimensions
still require refinement.** Direct responsibility, agreement and graph-relation
classification should remain excluded. Before designing the derivation stage,
the next work should examine the minimal reliable representation of explicit
reference/scope limitation, commitment modality and participant deixis. No
production integration, Progeo run or derivation implementation was performed.
## EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
Date: 2026-08-19
Hypothesis: `qwen3.5:9B` is substantially more reliable when the first semantic
stage preserves meeting meaning as atomic natural-language observations with
provenance and only simple participant information, without classifying or
deriving higher-level meeting semantics.
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
observation contains exactly `observation_id`, `evidence_id`, `content`,
`speaker`, nullable `named_person`, and nullable `addressee`. It contains no
modality, temporality, evaluation, affirmation, negation, determination,
uncertainty, clarification, responsibility, agreement, relation, reference,
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
Instead, the prompt asks for conservative atomic content that retains hedges,
conditions, personal/collective/impersonal language, requests, acceptances,
rejections, quantities, deadlines and boundaries in natural language.
Structural validation is intentionally small: exact schema keys, non-empty
observations/content, unique `obs_N` identifiers, known evidence IDs, speaker
matching its evidence, explicit named people/addressees, and no string
`"null"`. Human semantic preservation against per-case requirements is the
primary evaluation; wording differences do not fail a case.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`, no retries, voting or per-case tuning. One sandbox-blocked
localhost launch made zero model calls. The completed run made exactly nine
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
structural validation. Persistent source evidence, semantic requirements,
exact prompts, raw and parsed responses, validation, Ollama metadata and human
evaluation are stored under
`artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/`.
Human semantic preservation results:
| Case | Verdict | Main result |
| --- | --- | --- |
| A | PASS | Preserved `kann`, `vielleicht`, tentative follow-up, and explicit `Dann` sequence without commitment. |
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, `wenn überhaupt`, impersonal checking need and no decision, but dropped collective `wir` from who would receive contaminated material. |
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
“that we receive contaminated material back” into an impersonal passive phrase;
the risk itself remained hypothetical. H translated `Freitag` to `Friday`, a
harmless wording difference. I retained two compound observations rather than
splitting every proposition, but all required semantic boundaries and
dependencies remained explicit.
Compared with V2, categorical-field removal improved content preservation in
A, G and I: A retained `Dann`; G retained `wenn überhaupt`, personal `Ich` and
impersonal `Man`; I retained publication purpose and all boundaries. It also
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
strengthening into responsibility, group rejection, established work or
assigned clarification. F and H remain sufficiently complete in natural
language for a later interpretation experiment. No useful meaning was shown to
depend on the removed fields; the V3 content retained the useful signals that
V2's fields had attempted to encode.
Result: **A — MINIMAL FIRST STAGE ACCEPTED.** On A-I, minimal atomic content
with evidence provenance and simple participants is sufficiently reliable to
be the candidate first semantic stage. A later bounded experiment may examine
controlled semantic interpretation, but no derivation stage, production
integration or Progeo run was implemented here.
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
Date: 2026-08-11
Hypothesis: the primary protocol should be a topic-oriented reconstruction of
the meeting rather than a category-oriented list of extracted information.
This first isolated prototype does not replace or connect to the production
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
transcript excerpts to `qwen3.5:9B` and requests Discussion Subjects. Each
subject may contain supported discourse events, an outcome with mandatory
scope, resulting actions and unresolved issues. Optional structures must be
omitted when absent. Every semantic object must reference known evidence IDs.
The strict experimental schema validates:
- non-empty subjects and globally unique semantic identifiers;
- a closed discourse-event vocabulary;
- non-empty, known and non-duplicated evidence references;
- outcome text, scope, certainty and evidence;
- action text, JSON-nullable responsibility/deadline and evidence;
- unresolved-issue text and evidence;
- omission rather than null or empty optional structures.
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only,
multiple options, unaccepted proposal, proposal with objection, rejected
alternative, trial-scoped acceptance, no-decision discussion, resulting Action
Item, and outcome plus unresolved issue. Evaluation targets semantic identity,
development, outcome scope, actions, unresolved issues, traceability and
absence of invented commitments rather than exact wording.
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
`num_predict=4096`. Each case received exactly one model call; there were no
model retries or prompt iterations. The nine completed calls took 59.251
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
failure artifacts were preserved under
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/` and
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/`. Two earlier launch
attempts made zero LLM calls: one failed on the script import path and one was
blocked by sandbox networking.
Human-reviewed results after correcting two objectively wrong Gold assumptions
without another model call:
| Case | Verdict | Reason |
| --- | --- | --- |
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled `considered_option` rather than `introduced_idea`. |
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case
F: the model distinguished acceptance for a bounded trial from acceptance as a
final solution. It also handled the explicit action in case H well. However,
the experiment failed systematically on sparse structured output, subject
grouping and restraint around absent outcomes/actions/unresolved issues. The
model frequently mirrored optional schema fields as empty/null values, treated
alternatives as outcomes, split one discussion into multiple subjects, or
invented open issues from mere non-selection.
The focused experiment is not promising enough to justify a real Progeo chunk
sanity check. No such run was performed, and no architecture is accepted on
the basis of this prototype. Further work should first analyze whether the
failure comes from the schema/prompt representation, the model's sparse-output
reliability, or the boundary between subject grouping and semantic synthesis.
It should not proceed through repeated prompt tuning against these nine cases.