Add evidence-near semantic architecture experiments
Record the V1-V3 experiments and accept the minimal semantic-preservation first stage.
This commit is contained in:
@@ -1350,3 +1350,293 @@ accordance with the Gold Standard methodology.
|
||||
|
||||
Result: partially improved, not accepted as a complete BUG-015 fix. BUG-015
|
||||
remains Open; no phrase-specific deterministic filter was introduced.
|
||||
|
||||
## EXP-0027 — Evidence-near observation extraction
|
||||
|
||||
Date: 2026-08-18
|
||||
|
||||
Hypothesis: `qwen3.5:9B` can more reliably extract evidence-near linguistic and
|
||||
semantic properties than directly synthesize protocol-level events, outcomes,
|
||||
actions and unresolved issues. This isolated experiment stops before semantic
|
||||
interpretation and does not connect to the production pipeline.
|
||||
|
||||
The fixed Gold fixture reuses the unchanged A-I evidence and fixed Discussion
|
||||
Subjects from Semantic Synthesis Isolation. It defines atomic observations with
|
||||
source evidence, explicit targets, a five-value relation vocabulary, modality,
|
||||
temporality, evaluation, agreement, responsibility/person, uncertainty,
|
||||
clarification need and free-text scope. It contains no protocol-level category
|
||||
field. The validator requires sequential observation IDs, known evidence IDs,
|
||||
backward-only valid observation targets, closed categorical vocabularies,
|
||||
consistent responsibility/person pairs, and one canonical absence form: JSON
|
||||
null for person and `absent` for scope.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`, no retries. All nine cases ran exactly once, for nine LLM
|
||||
calls total. The run took 74.732 seconds and used 7,627 prompt-evaluation tokens
|
||||
and 4,514 evaluation tokens. Exact prompts, Gold input and expectations, raw
|
||||
responses, parsed observations, validation results, Ollama metadata and
|
||||
comparisons are preserved under
|
||||
`/tmp/meeting-lab-evidence-observations-v1-20260818/`.
|
||||
|
||||
Strict automated validation/evaluation produced 0 PASS, 1 PARTIAL and 8 FAIL.
|
||||
Five cases failed structure because the model represented a single target as a
|
||||
one-element list, usually `["discussion_subject"]`; the accepted schema permits
|
||||
a list only for two or more jointly referenced observations. Several responses
|
||||
also copied the relation label `limits_scope` into the free-text scope field.
|
||||
These were systematic model-output errors, not transport or parser failures.
|
||||
The prompt and run were not retried or tuned.
|
||||
|
||||
Human semantic review of the preserved raw responses:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PARTIAL | Kept the geometry change uncommitted but collapsed the follow-up observation and weakened explicit uncertainty. |
|
||||
| B | PARTIAL | Preserved two unselected alternatives without commitment, but used incorrect targets/relations and omitted the joint-reference observation. |
|
||||
| C | FAIL | The tentative contact remained ownerless, but the follow-up was incorrectly marked factual and accepted. |
|
||||
| D | PARTIAL | Preserved the negative energy consequence without creating a clarification need, but omitted the explicit uncertainty about whether washing is worthwhile and misused agreement. |
|
||||
| E | PARTIAL | Preserved explicit rejection and verbal confirmation, but failed to target the confirmation at the rejection and weakened the committed future rejection to a completed fact. |
|
||||
| F | FAIL | Preserved the trial-versus-series wording, but failed atomic scope targeting and incorrectly assigned responsibility to Tim from collective speech. |
|
||||
| G | FAIL | Correctly recognized impersonal necessity, but promoted Martin's preference to rejection and responsibility and weakened risk/availability uncertainty. |
|
||||
| H | PARTIAL | Distinguished request from commitment and captured Nina's acceptance, but named the requester as responsible in the request and failed accepted-responsibility/target encoding. |
|
||||
| I | PARTIAL | Preserved the bounded production facts and the information question without assigning work, but lost scope relations and marked the unresolved permission as rejected and not uncertain. |
|
||||
|
||||
Human-review total: 0 PASS, 6 PARTIAL, 3 FAIL. This review does not override
|
||||
strict structural failures; it separates useful semantic signal from schema
|
||||
compliance.
|
||||
|
||||
Compared with Semantic Synthesis Isolation (1 PASS, 2 PARTIAL, 6 FAIL), moving
|
||||
closer to evidence reduced some direct promotion behavior: the geometry mention
|
||||
did not become work, the washing disadvantage did not become an unresolved
|
||||
issue, both alternatives in B remained uncommitted, and the publication query
|
||||
did not become an assignment. However, the important promotion errors did not
|
||||
disappear. C acquired unsupported acceptance, and G still promoted a personal
|
||||
preference into rejection. Positive cases were only partly preserved: explicit
|
||||
rejection was recognized but incorrectly linked; trial-only language was kept
|
||||
but responsibility was invented; Nina's request and commitment were recognized
|
||||
but responsibility states were wrong; and the publication issue was recognized
|
||||
but its uncertainty was contradicted by rejection.
|
||||
|
||||
Result: **B — evidence-near extraction is promising, but specific observation
|
||||
dimensions remain unreliable.** Target/relation selection, scope attachment,
|
||||
responsibility state/person attribution, and agreement versus uncertainty are
|
||||
not reliable enough to justify designing the later interpretation stage yet.
|
||||
No production integration or later interpretation stage was implemented.
|
||||
|
||||
## EXP-0028 — Evidence-Near Observation Extraction V2
|
||||
|
||||
Date: 2026-08-19
|
||||
|
||||
V2 tested whether `qwen3.5:9B` preserves the evidence needed by a later
|
||||
controlled interpretation stage when direct responsibility, agreement and
|
||||
semantic graph relations are removed. Responsibility was replaced by explicit
|
||||
participant/discourse facts (`speaker`, `named_person`, `addressee`, singular
|
||||
self-reference, collective `we`, and impersonal person reference). Agreement
|
||||
was replaced by explicit affirmation, explicit negation and determination
|
||||
statement signals. Graph relations were reduced to nullable scalar
|
||||
`refers_to`; scope became free-text `qualifier` plus nullable scalar
|
||||
`limits_target`. No later derivation stage was implemented.
|
||||
|
||||
The V2 Gold fixture preserves the unchanged A-I source evidence and intended
|
||||
human interpretations. It contains no responsibility, agreement, action,
|
||||
decision, open-question, accepted-trial, rejected-alternative or protocol
|
||||
eligibility fields. Validation enforces known evidence IDs, sequential unique
|
||||
observation IDs, backward-only scalar references, closed vocabularies, boolean
|
||||
participant flags, JSON-nullable participant/qualifier/reference fields and no
|
||||
string `"null"`.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`, no retries or voting. A launch-path defect was corrected
|
||||
before the live run; the failed launch made zero model calls. A sandbox-blocked
|
||||
localhost attempt also made zero model calls. The completed run called the
|
||||
model exactly once for each of A-I: nine calls total, in 125.645 seconds.
|
||||
Persistent prompts, Gold input and expectations, raw and parsed model output,
|
||||
validation, automatic comparison, Ollama metadata and human evaluation are in
|
||||
`artifacts/experiments/evidence_observations_v2/20260819_v2_single_run/`.
|
||||
|
||||
Strict automated comparison produced 0 PASS, 0 PARTIAL and 9 FAIL. Seven cases
|
||||
were schema-invalid. The dominant serialization pattern was use of `present`
|
||||
instead of the specified `explicit` for affirmation/negation; E additionally
|
||||
used `none` instead of `absent` for a determination signal, while D emitted the
|
||||
separate uncertainty concept as an invalid modality. These errors are
|
||||
contract violations, although most `present`/`explicit` differences are
|
||||
deterministically normalizable without changing meaning. A and C were valid
|
||||
JSON/schema outputs but had critical semantic mismatches.
|
||||
|
||||
Human semantic review:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PARTIAL | Preserved possibility, uncertainty, atomic follow-up and its reference, but classified the initial possibility as suggestion/existing and missed implicit clarification. |
|
||||
| B | PARTIAL | Preserved both alternatives without commitment or ownership, but over-fragmented, omitted references/qualifiers and added a determination signal; `present` caused schema failure. |
|
||||
| C | FAIL | Preserved the initial uncertain suggestion and no ownership, but missed self-reference and again converted the follow-up possibility to a factual existing statement. |
|
||||
| D | PARTIAL | Preserved possibility, explicit uncertainty, process description and negative energy consequence without assignment; references/qualifiers were lost and uncertainty was also emitted as an invalid modality. |
|
||||
| E | PARTIAL | Preserved explicit no, collective speech, explicit yes and a determination statement, but missed future commitment and the confirmation reference and over-fragmented the rejection. |
|
||||
| F | PARTIAL | Preserved collective speech, explicit affirmation, future test, quantity, trial-only boundary and non-adoption as series solution without individual ownership, but missed committed modality and all reference/scope attachments. |
|
||||
| G | FAIL | Avoided responsibility and group-rejection promotion, but weakened risk uncertainty, personal-preference features, conditionality and impersonal necessity. |
|
||||
| H | PARTIAL | Correctly preserved speaker, named addressee, request, response speaker, self-reference, affirmation and future conduct without responsibility, but duplicated the request and missed committed modality, reference and deadline qualifier. |
|
||||
| I | PARTIAL | Preserved production content, information question without assignment, unresolved permission and clarification need, but lost every reference/qualifier/limit and weakened impersonal necessity. |
|
||||
|
||||
Human total: 0 PASS, 7 PARTIAL, 2 FAIL. The reduced schema materially reduced
|
||||
V1 promotion errors: collective speech and speaker identity no longer became
|
||||
individual responsibility; personal preference no longer became a group-level
|
||||
rejection field; an information question did not become work; and explicit
|
||||
negation/affirmation survived as separate evidence. Useful participant evidence
|
||||
also survived strongly in H and collective-speech evidence in F.
|
||||
|
||||
Simplification did not make all evidence-near dimensions reliable. Scalar
|
||||
references and `limits_target` were almost entirely omitted, qualifiers were
|
||||
usually omitted, committed modality was missed in E, F and H, and C/G repeated
|
||||
important modality, uncertainty and participant-feature errors. Some positive
|
||||
semantic information therefore survived only in free-text `content`, not in
|
||||
the structural signals a controlled derivation stage would need.
|
||||
|
||||
Result: **B — V2 is materially better, but specific evidence-near dimensions
|
||||
still require refinement.** Direct responsibility, agreement and graph-relation
|
||||
classification should remain excluded. Before designing the derivation stage,
|
||||
the next work should examine the minimal reliable representation of explicit
|
||||
reference/scope limitation, commitment modality and participant deixis. No
|
||||
production integration, Progeo run or derivation implementation was performed.
|
||||
|
||||
## EXP-0029 — Evidence-Near Observation Extraction V3 — Minimal Semantic Preservation
|
||||
|
||||
Date: 2026-08-19
|
||||
|
||||
Hypothesis: `qwen3.5:9B` is substantially more reliable when the first semantic
|
||||
stage preserves meeting meaning as atomic natural-language observations with
|
||||
provenance and only simple participant information, without classifying or
|
||||
deriving higher-level meeting semantics.
|
||||
|
||||
V3 uses the unchanged A-I evidence and intended meanings from V1/V2. Each
|
||||
observation contains exactly `observation_id`, `evidence_id`, `content`,
|
||||
`speaker`, nullable `named_person`, and nullable `addressee`. It contains no
|
||||
modality, temporality, evaluation, affirmation, negation, determination,
|
||||
uncertainty, clarification, responsibility, agreement, relation, reference,
|
||||
qualifier, scope, limit, protocol-category or protocol-eligibility fields.
|
||||
Instead, the prompt asks for conservative atomic content that retains hedges,
|
||||
conditions, personal/collective/impersonal language, requests, acceptances,
|
||||
rejections, quantities, deadlines and boundaries in natural language.
|
||||
|
||||
Structural validation is intentionally small: exact schema keys, non-empty
|
||||
observations/content, unique `obs_N` identifiers, known evidence IDs, speaker
|
||||
matching its evidence, explicit named people/addressees, and no string
|
||||
`"null"`. Human semantic preservation against per-case requirements is the
|
||||
primary evaluation; wording differences do not fail a case.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`, no retries, voting or per-case tuning. One sandbox-blocked
|
||||
localhost launch made zero model calls. The completed run made exactly nine
|
||||
calls, one for each A-I case, in 40.074 seconds. All nine outputs passed
|
||||
structural validation. Persistent source evidence, semantic requirements,
|
||||
exact prompts, raw and parsed responses, validation, Ollama metadata and human
|
||||
evaluation are stored under
|
||||
`artifacts/experiments/evidence_observations_v3/20260819_v3_single_run/`.
|
||||
|
||||
Human semantic preservation results:
|
||||
|
||||
| Case | Verdict | Main result |
|
||||
| --- | --- | --- |
|
||||
| A | PASS | Preserved `kann`, `vielleicht`, tentative follow-up, and explicit `Dann` sequence without commitment. |
|
||||
| B | PASS | Preserved insufficient strength, both alternatives, their two-approach framing, and non-selection. |
|
||||
| C | PASS | Preserved Tim's tentative personal Textor contact, possible follow-up and absence of established work. |
|
||||
| D | PASS | Preserved washing possibility, explicit uncertainty, process, energy consequence and absence of an invented task. |
|
||||
| E | PASS | Preserved neutral cost, collective explicit rejection/non-pursuit and subsequent confirmation that it is decided. |
|
||||
| F | PASS | Preserved collective possibility and test commitment, small extruder, 20 metres, next trial, trial-only limit and not-yet series adoption without individual ownership. |
|
||||
| G | PARTIAL | Preserved hypothetical risk, Martin's personal stance, `wenn überhaupt`, impersonal checking need and no decision, but dropped collective `wir` from who would receive contaminated material. |
|
||||
| H | PASS | Preserved Antonius's request to Nina, Friday, Nina's explicit acceptance and future first-person commitment without a responsibility field. |
|
||||
| I | PASS | Preserved production/comparison boundaries, upstream effort, publication purpose, unresolved permission and clarification need without assignment. |
|
||||
|
||||
Human total: 8 PASS, 1 PARTIAL, 0 FAIL. G's only material weakening changed
|
||||
“that we receive contaminated material back” into an impersonal passive phrase;
|
||||
the risk itself remained hypothetical. H translated `Freitag` to `Friday`, a
|
||||
harmless wording difference. I retained two compound observations rather than
|
||||
splitting every proposition, but all required semantic boundaries and
|
||||
dependencies remained explicit.
|
||||
|
||||
Compared with V2, categorical-field removal improved content preservation in
|
||||
A, G and I: A retained `Dann`; G retained `wenn überhaupt`, personal `Ich` and
|
||||
impersonal `Man`; I retained publication purpose and all boundaries. It also
|
||||
reduced fragmentation from 38 observations in V2 to 28 in V3, with no semantic
|
||||
strengthening into responsibility, group rejection, established work or
|
||||
assigned clarification. F and H remain sufficiently complete in natural
|
||||
language for a later interpretation experiment. No useful meaning was shown to
|
||||
depend on the removed fields; the V3 content retained the useful signals that
|
||||
V2's fields had attempted to encode.
|
||||
|
||||
Result: **A — MINIMAL FIRST STAGE ACCEPTED.** On A-I, minimal atomic content
|
||||
with evidence provenance and simple participants is sufficiently reliable to
|
||||
be the candidate first semantic stage. A later bounded experiment may examine
|
||||
controlled semantic interpretation, but no derivation stage, production
|
||||
integration or Progeo run was implemented here.
|
||||
|
||||
## EXP-0026 — Topic-oriented Discussion Subject reconstruction V2 prototype
|
||||
|
||||
Date: 2026-08-11
|
||||
|
||||
Hypothesis: the primary protocol should be a topic-oriented reconstruction of
|
||||
the meeting rather than a category-oriented list of extracted information.
|
||||
|
||||
This first isolated prototype does not replace or connect to the production
|
||||
pipeline or Working Protocol renderer. It sends small evidence-ID-tagged
|
||||
transcript excerpts to `qwen3.5:9B` and requests Discussion Subjects. Each
|
||||
subject may contain supported discourse events, an outcome with mandatory
|
||||
scope, resulting actions and unresolved issues. Optional structures must be
|
||||
omitted when absent. Every semantic object must reference known evidence IDs.
|
||||
|
||||
The strict experimental schema validates:
|
||||
|
||||
- non-empty subjects and globally unique semantic identifiers;
|
||||
- a closed discourse-event vocabulary;
|
||||
- non-empty, known and non-duplicated evidence references;
|
||||
- outcome text, scope, certainty and evidence;
|
||||
- action text, JSON-nullable responsibility/deadline and evidence;
|
||||
- unresolved-issue text and evidence;
|
||||
- omission rather than null or empty optional structures.
|
||||
|
||||
Focused Gold material contains nine BUG-015/Progeo-derived cases: idea only,
|
||||
multiple options, unaccepted proposal, proposal with objection, rejected
|
||||
alternative, trial-scoped acceptance, no-decision discussion, resulting Action
|
||||
Item, and outcome plus unresolved issue. Evaluation targets semantic identity,
|
||||
development, outcome scope, actions, unresolved issues, traceability and
|
||||
absence of invented commitments rather than exact wording.
|
||||
|
||||
Configuration: `qwen3.5:9B`, temperature 0, `think=false`, `num_ctx=16384`,
|
||||
`num_predict=4096`. Each case received exactly one model call; there were no
|
||||
model retries or prompt iterations. The nine completed calls took 59.251
|
||||
seconds in aggregate and used 6,680 prompt-evaluation tokens plus 3,274
|
||||
evaluation tokens. Raw model responses, prompts, parsed JSON, metadata and
|
||||
failure artifacts were preserved under
|
||||
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run2/` and
|
||||
`/tmp/meeting-lab-topic-reconstruction-v2-gold-run3/`. Two earlier launch
|
||||
attempts made zero LLM calls: one failed on the script import path and one was
|
||||
blocked by sandbox networking.
|
||||
|
||||
Human-reviewed results after correcting two objectively wrong Gold assumptions
|
||||
without another model call:
|
||||
|
||||
| Case | Verdict | Reason |
|
||||
| --- | --- | --- |
|
||||
| A — idea only | PARTIAL | Correct subject and no invented outcome/action, but the isolated idea was labeled `considered_option` rather than `introduced_idea`. |
|
||||
| B — multiple options | FAIL | Invalid empty optional list; one discussion subject was split into three, and alternatives were promoted to tentative outcomes and invented unresolved issues. |
|
||||
| C — unaccepted proposal | FAIL | Proposal was detected, but output used forbidden null/empty structures and promoted it to an Action Item. |
|
||||
| D — proposal with objection | FAIL | Invalid null/empty structures; the objection was not reconstructed as a discourse event and was converted into an unresolved issue. |
|
||||
| E — rejected alternative | FAIL | Rejection, scope and evidence were semantically correct, but strict validation failed on empty optional lists. |
|
||||
| F — trial-only acceptance | PARTIAL | Crucially preserved the 20-metre trial scope and excluded final-series acceptance; it represented the limitation as state/unresolved context rather than a clarification event. |
|
||||
| G — no decision | FAIL | Invalid empty lists, split a connected subject, represented “no decision” as a tentative outcome and invented a prerequisite outcome. |
|
||||
| H — resulting action | PASS | Correct subject, explicit acceptance, Nina responsibility, Friday deadline, outcome and evidence references. |
|
||||
| I — outcome plus unresolved | FAIL | Captured the production-only outcome scope, but omitted supporting evidence and the unresolved publication question; output also contained an empty optional list. |
|
||||
|
||||
Result: 1 PASS, 2 PARTIAL, 6 FAIL. The most important positive signal was case
|
||||
F: the model distinguished acceptance for a bounded trial from acceptance as a
|
||||
final solution. It also handled the explicit action in case H well. However,
|
||||
the experiment failed systematically on sparse structured output, subject
|
||||
grouping and restraint around absent outcomes/actions/unresolved issues. The
|
||||
model frequently mirrored optional schema fields as empty/null values, treated
|
||||
alternatives as outcomes, split one discussion into multiple subjects, or
|
||||
invented open issues from mere non-selection.
|
||||
|
||||
The focused experiment is not promising enough to justify a real Progeo chunk
|
||||
sanity check. No such run was performed, and no architecture is accepted on
|
||||
the basis of this prototype. Further work should first analyze whether the
|
||||
failure comes from the schema/prompt representation, the model's sparse-output
|
||||
reliability, or the boundary between subject grouping and semantic synthesis.
|
||||
It should not proceed through repeated prompt tuning against these nine cases.
|
||||
|
||||
Reference in New Issue
Block a user