Add RC1 regression benchmark artifacts
Add benchmark artifacts created during the first RC1 robustness evaluation of the Meeting Lab pipeline. Included benchmark sets: - hardware_experimental_qwen35_9b - meeting_context_v1 - progeo_meeting_20260804_083849 - progeo_meeting_context_v1_20260804_110913 - progeo_meeting_rc1_20260804_121502 These artifacts document the evolution of the pipeline during the implementation and verification of BUG-009, BUG-010 and BUG-011. The benchmark data provide reproducible real-life regression cases for future development and allow quality comparisons across pipeline revisions. Current benchmark policy: During the active development phase, representative benchmark artifacts are intentionally versioned to preserve reproducibility and simplify regression analysis. Benchmark artifacts are considered part of the engineering evidence rather than temporary build output. Benchmark retention strategy will be revisited once the Meeting Assistant reaches production maturity.
This commit is contained in:
@@ -0,0 +1,106 @@
|
||||
# Hardware Benchmark Report
|
||||
|
||||
## System
|
||||
|
||||
- OS: Microsoft Windows 11 Pro 10.0.26200, 64-bit
|
||||
- CPU: AMD Ryzen 5 9600X 6-Core Processor, 6 cores / 12 logical processors
|
||||
- RAM: 33,988,124,672 bytes installed (~31.7 GiB)
|
||||
- GPU: AMD Radeon RX 9070
|
||||
- GPU memory: `ollama ps` reported 6.7 GB model resident on GPU; Windows CIM AdapterRAM reported 4,293,918,720 bytes, likely driver-limited reporting
|
||||
- Ollama version: 0.32.5
|
||||
- Model: `qwen3.5:9B` / `qwen3.5:9b`, 9.7B, Q4_K_M, digest `6488c96fa5faab64bb65cbd30d4289e20e6130ef535a93ef9a49f42eda893ea7`
|
||||
- Commit: `63075eaca91088beb92e646b6b0a2539cfe049ca`
|
||||
- Branch: `feature/windowed-segmentation`
|
||||
- Upstream: `origin/feature/windowed-segmentation`
|
||||
|
||||
## Runtime comparison
|
||||
|
||||
| Stage | MacBook Air M4 | Experimental computer | Speedup |
|
||||
|---|---:|---:|---:|
|
||||
| Semantic Consolidator V0 | 390.119 s | 38.451 s | 10.15x |
|
||||
| Working Protocol Renderer V2 | 722.900 s | 37.675 s | 19.19x |
|
||||
| Total | 1113.019 s | 76.126 s | 14.62x |
|
||||
|
||||
## Semantic Consolidator V0
|
||||
|
||||
- Runtime state: cold start; `ollama ps` was empty before the measured run
|
||||
- Wall-clock runtime: 38.451 s
|
||||
- Ollama total_duration: 38.4172797 s
|
||||
- load_duration: 4.0657233 s
|
||||
- prompt_eval_count: 4295
|
||||
- prompt_eval_duration: 1.629666 s
|
||||
- eval_count: 2399
|
||||
- eval_duration: 32.712261 s
|
||||
- Fact count: 33
|
||||
- Merged fact groups: 1
|
||||
- Singleton groups: 31
|
||||
- Validation: passed
|
||||
- Source fact coverage: every source fact occurred exactly once
|
||||
- Missing IDs: none
|
||||
- Duplicate source IDs: none
|
||||
- Non-fact categories unchanged: yes
|
||||
- Accepted duplicate: `fact_0025` + `fact_0031`, F&E project list / biweekly interface stand-up
|
||||
- GPU acceleration: active; `ollama ps` showed `PROCESSOR 100% GPU`, context 32768, model size 6.7 GB while inference was active
|
||||
- Monitoring: see `semantic_consolidator_v0/monitoring_log.txt`
|
||||
|
||||
## Working Protocol Renderer V2 - official Python benchmark
|
||||
|
||||
- Result: valid benchmark run
|
||||
- Execution method: synchronous foreground Python using `requests.post`; no PowerShell background job, no `Invoke-WebRequest`, no PowerShell response redirection
|
||||
- Runtime state: cold start; `ollama ps` was empty before the measured run
|
||||
- Wall-clock runtime: 37.675 s
|
||||
- MacBook Air M4 baseline: 722.900 s
|
||||
- Speedup: 19.19x
|
||||
- total_duration: 37.6437146 s
|
||||
- load_duration: 4.0374398 s
|
||||
- prompt_eval_count: 28349
|
||||
- prompt_eval_duration: 14.694067 s
|
||||
- eval_count: 1268
|
||||
- eval_duration: 18.840999 s
|
||||
- Prompt characters: 112964
|
||||
- Output characters: 5567
|
||||
- Output bytes: 5626
|
||||
- Detected output language: German
|
||||
- Readable Markdown: true
|
||||
- HTTP status: 200
|
||||
- Content-Type: `application/json; charset=utf-8`
|
||||
- Raw response byte length: 148913
|
||||
- Top-level JSON keys: `context, created_at, done, done_reason, eval_count, eval_duration, load_duration, model, prompt_eval_count, prompt_eval_duration, response, total_duration`
|
||||
- done: true
|
||||
- done_reason: stop
|
||||
- Request count: 1, recorded in `working_protocol_renderer_v2_python/request_count_marker.txt`
|
||||
- GPU acceleration: active; `ollama ps` showed `PROCESSOR 100% GPU`, context 32768, model size 6.7 GB after the request
|
||||
- Output path: `working_protocol_renderer_v2_python/working_protocol.md`
|
||||
- Raw response path: `working_protocol_renderer_v2_python/raw_ollama_response.json`
|
||||
- Run report: `working_protocol_renderer_v2_python/report.md`
|
||||
|
||||
## Working Protocol Renderer V2 - original invalid attempt
|
||||
|
||||
- Result: invalid benchmark run; no valid `working_protocol.md` was produced
|
||||
- Runtime state at attempted start: warm; model was still loaded from Run 1
|
||||
- Reason: the PowerShell background job started outside the repository working directory and failed to resolve `prompts/working_protocol.md` and the consolidated input path
|
||||
- Comparability: not valid; no speedup claimed
|
||||
- Monitoring: see `working_protocol_renderer_v2/monitoring_log.txt`
|
||||
|
||||
## Working Protocol Renderer V2 - PowerShell retry invalid attempt
|
||||
|
||||
- Result: invalid benchmark run; no valid `working_protocol.md` was produced
|
||||
- Runtime state at attempted start: cold; `ollama ps` was empty before the request
|
||||
- Request count marker: exactly one request start and one request end were recorded locally in `working_protocol_renderer_v2_retry/request_count_marker.txt`
|
||||
- Failure: the wrapper wrote `$resp.Content`, but `$resp.Content` was null/empty in the background job. `Set-Content -Encoding utf8` therefore created a 3-byte UTF-8 BOM file, and the following `$data.response` access produced the null-reference error.
|
||||
- GPU acceleration during attempt: active; `ollama ps` showed `PROCESSOR 100% GPU`, context 32768, model size 6.7 GB while inference was active
|
||||
- Comparability: not valid; no speedup claimed
|
||||
- Monitoring: see `working_protocol_renderer_v2_retry/monitoring_log.txt`
|
||||
|
||||
## Renderer Diagnosis
|
||||
|
||||
- No committed renderer CLI or module exists in this checkout for Working Protocol Renderer V2.
|
||||
- The Mac benchmark artifact identifies a temporary prompt file at `/private/tmp/meeting_lab_working_protocol_synthesis_prompt.txt`, input adapter required, one LLM call, model `qwen3.5:9B`, and `think=false`.
|
||||
- The successful Python benchmark used the same practical request shape as the successful Mac artifact: `/api/generate`, `stream=false`, `think=false`, no JSON format constraint, `temperature=0.0`, `num_ctx=32768`, `num_predict=4096`, prompt file plus consolidated JSON input.
|
||||
- Response parsing reads the top-level Ollama `response` string and writes it directly as `working_protocol.md`. The complete raw HTTP response JSON is preserved before parsing.
|
||||
|
||||
## Comparability limitations
|
||||
|
||||
- Semantic Consolidator V0 used the requested input, model, quantization, endpoint, prompt, and `think=false` setting.
|
||||
- Working Protocol Renderer V2 used the requested prompt, hardware Semantic Consolidator V0 output, model, quantization, endpoint, and `think=false` setting.
|
||||
- The renderer prompt character count differs from the Mac artifact because this run used the newly generated hardware Semantic Consolidator V0 output, as requested.
|
||||
+28
@@ -0,0 +1,28 @@
|
||||
Fact item count: 33
|
||||
Estimated prompt size chars: 15393
|
||||
Estimated prompt tokens: 3849
|
||||
Expected LLM call count: 1
|
||||
Expected runtime: 5-10 minutes on current local benchmark basis
|
||||
Ollama request start: 2026-08-01T08:42:34
|
||||
Ollama endpoint: http://127.0.0.1:11434/api/generate
|
||||
Ollama model: qwen3.5:9B
|
||||
Ollama stream: False
|
||||
Ollama think: False
|
||||
Ollama timeout seconds: 1800
|
||||
Ollama options: {"num_ctx": 32768, "num_predict": 4096, "temperature": 0.0}
|
||||
Waiting for Ollama response: 30.0 seconds elapsed
|
||||
Ollama response metadata:
|
||||
total_duration: 38417279700
|
||||
load_duration: 4065723300
|
||||
prompt_eval_count: 4295
|
||||
prompt_eval_duration: 1629666000
|
||||
eval_count: 2399
|
||||
eval_duration: 32712261000
|
||||
Runtime seconds: 38.451
|
||||
Merged fact groups: 1
|
||||
Source facts involved in merges: 2
|
||||
Singleton fact groups: 31
|
||||
Validation result: passed
|
||||
Output: samples\benchmarks\hardware_experimental_qwen35_9b\semantic_consolidator_v0\consolidated_extractions.json
|
||||
Report: samples\benchmarks\hardware_experimental_qwen35_9b\semantic_consolidator_v0\report.md
|
||||
Raw model response: samples\benchmarks\hardware_experimental_qwen35_9b\semantic_consolidator_v0\raw_model_response.txt
|
||||
+1723
File diff suppressed because it is too large
Load Diff
+32
@@ -0,0 +1,32 @@
|
||||
SAMPLE 0 2026-08-01T08:42:34.7197665+02:00 FreeMemKB=16797340 OllamaCPU=13.84375 OllamaWS=49446912
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
|
||||
SAMPLE 1 2026-08-01T08:42:39.9119216+02:00 FreeMemKB=15942168 OllamaCPU=14.234375 OllamaWS=69578752
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 2 2026-08-01T08:42:45.0374639+02:00 FreeMemKB=15835780 OllamaCPU=14.25 OllamaWS=69705728
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 3 2026-08-01T08:42:50.1735639+02:00 FreeMemKB=17120500 OllamaCPU=14.296875 OllamaWS=66314240
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 4 2026-08-01T08:42:55.3234153+02:00 FreeMemKB=18431496 OllamaCPU=14.328125 OllamaWS=66527232
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 5 2026-08-01T08:43:00.4718461+02:00 FreeMemKB=18403816 OllamaCPU=14.40625 OllamaWS=66707456
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 6 2026-08-01T08:43:05.6096000+02:00 FreeMemKB=18421668 OllamaCPU=14.421875 OllamaWS=66990080
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 7 2026-08-01T08:43:10.7368162+02:00 FreeMemKB=18445280 OllamaCPU=14.4375 OllamaWS=67121152
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
PROCESS_EXIT 2026-08-01T08:43:15.8756249+02:00 ExitCode= ActiveConfirmed=True
|
||||
+164
@@ -0,0 +1,164 @@
|
||||
{
|
||||
"groups": [
|
||||
{
|
||||
"canonical_text": "Das ist das aktuelle Projektdeckblatt, was wir in der F&E benutzen.",
|
||||
"source_item_ids": ["fact_0001"],
|
||||
"merge_reason": "Singleton: No other items express the specific identity and usage of this document."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Das Ding ist kein Filtersystem.",
|
||||
"source_item_ids": ["fact_0002"],
|
||||
"merge_reason": "Singleton: Specific negation regarding system type, not equivalent to general process descriptions or other project attributes."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Kriterien sind nicht relevant für ein Realprodukt.",
|
||||
"source_item_ids": ["fact_0003"],
|
||||
"merge_reason": "Singleton: Specific statement about criterion relevance for real products, distinct from general process steps or other criteria lists."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Der Prozess ist im Wesentlichen abgestimmt auf Entwicklungsideen und geht um Projekte, die entweder physische oder nicht-physischer Natur sind.",
|
||||
"source_item_ids": ["fact_0004"],
|
||||
"merge_reason": "Singleton: Describes the scope and nature of projects in a specific process context. Distinct from general flow descriptions or responsibility assignments."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Der grundsätzliche Ablauf bleibt gleich, aber andere Fragestellungen und Dokumente werden verwendet.",
|
||||
"source_item_ids": ["fact_0005"],
|
||||
"merge_reason": "Singleton: Describes a variation in process flow and documents. Distinct from specific document identities or general scope definitions."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Projektleitung ist überhaupt nicht abteilungszugeordnet.",
|
||||
"source_item_ids": ["fact_0006"],
|
||||
"merge_reason": "Singleton: Specific organizational attribute of the project leadership. Distinct from general responsibility assignments or process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Variante mit PowerPoint-Folien und Stichpunkten hat sich im letzten Jahr als wirklich sehr gut funktionierend und im Arbeitsalltag wertvoll erwiesen.",
|
||||
"source_item_ids": ["fact_0007"],
|
||||
"merge_reason": "Singleton: Specific evaluation of a presentation format's utility. Distinct from process steps or general criteria."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Bevor ein Projekt entsteht, müssen im Vorfeld andere Kriterien abgehandelt werden.",
|
||||
"source_item_ids": ["fact_0008"],
|
||||
"merge_reason": "Singleton: Describes a prerequisite step before project initiation. Distinct from the specific criteria themselves or general process flow."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Solange die Sachen nicht geklärt sind, brauchen wir noch gar nicht loslegen.",
|
||||
"source_item_ids": ["fact_0009"],
|
||||
"merge_reason": "Singleton: Describes a condition for starting work. Distinct from the specific criteria or general process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Aufgabe im Vorfeld alles zu machen liegt im kaufmännischen Bereich.",
|
||||
"source_item_ids": ["fact_0010"],
|
||||
"merge_reason": "Singleton: Assigns a specific responsibility (pre-project work) to the commercial area. Distinct from general process steps or other responsibilities."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Der Prozess ist ein F&E-Prozess gewesen, einfach historisch.",
|
||||
"source_item_ids": ["fact_0011"],
|
||||
"merge_reason": "Singleton: Historical classification of the process. Distinct from current definitions or specific steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Eine neue Idee muss mindestens folgende Punkte beantworten: Titel, Projektidee und ein Projektziel.",
|
||||
"source_item_ids": ["fact_0012"],
|
||||
"merge_reason": "Singleton: Defines mandatory content for a new idea. Distinct from optional criteria or other validation steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Optional ist eine konkrete Aufzählung der zugrunde liegenden Annahmen, zum Beispiel erwartbares Marktvolumen.",
|
||||
"source_item_ids": ["fact_0013"],
|
||||
"merge_reason": "Singleton: Defines optional content for a new idea. Distinct from mandatory requirements or other validation steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Ideen müssen geprüft werden, ob sie auf Team- oder Abteilungsebene abzuarbeiten sind.",
|
||||
"source_item_ids": ["fact_0014"],
|
||||
"merge_reason": "Singleton: Describes a specific scope validation step. Distinct from other criteria or responsibility assignments."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Ideen müssen abgefragt werden, ob sie abteilungsübergreifend, strategisch relevant sind oder ein Budget notwendig ist.",
|
||||
"source_item_ids": ["fact_0015"],
|
||||
"merge_reason": "Singleton: Describes a specific relevance/budget validation step. Distinct from other criteria or responsibility assignments."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Netzwerkverbindung bei Gemeinden ist sehr langsam.",
|
||||
"source_item_ids": ["fact_0016"],
|
||||
"merge_reason": "Singleton: Specific technical status statement regarding network speed at a specific location. Distinct from process or organizational facts."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Sekretariate sind da sehr abweisend, weil sie sagen, dass sie keine Ahnung davon haben und viel anderes zu tun.",
|
||||
"source_item_ids": ["fact_0017"],
|
||||
"merge_reason": "Singleton: Describes the attitude and reasoning of secretariats. Distinct from process steps or technical facts."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Filter müssen definiert werden und dienen als grobe Filter, nicht bis ins Kleinste.",
|
||||
"source_item_ids": ["fact_0018"],
|
||||
"merge_reason": "Singleton: Describes the nature and necessity of defining filters. Distinct from specific filter criteria or process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Das Schlimmste ist, dass einer kommt und sagt, er wirft nur ein Stichwort rein.",
|
||||
"source_item_ids": ["fact_0019"],
|
||||
"merge_reason": "Singleton: Describes a negative behavior pattern in the process. Distinct from positive criteria or structural facts."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Man sollte nicht im Vorhinein versuchen, alles zu planen.",
|
||||
"source_item_ids": ["fact_0020"],
|
||||
"merge_reason": "Singleton: Expresses an intention/strategy against early planning. Distinct from factual process steps or criteria."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Durchsetzung und Umsetzung ist das Entscheidende am Schluss.",
|
||||
"source_item_ids": ["fact_0021"],
|
||||
"merge_reason": "Singleton: Highlights the importance of implementation at the end. Distinct from planning steps or criteria."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Wir machen keine Teppiche für Autos.",
|
||||
"source_item_ids": ["fact_0022"],
|
||||
"merge_reason": "Singleton: Specific business scope limitation/negation. Distinct from general process descriptions."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Geschäftsführung hat den Überblick über Projekte im Haus zu behalten.",
|
||||
"source_item_ids": ["fact_0023"],
|
||||
"merge_reason": "Singleton: Assigns a specific responsibility (overview) to the Management Board. Distinct from other responsibilities or process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Der F&E-Topf wird gemeinsam geführt.",
|
||||
"source_item_ids": ["fact_0024"],
|
||||
"merge_reason": "Singleton: Describes a specific organizational arrangement for funding/resources. Distinct from general responsibility assignments."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up.",
|
||||
"source_item_ids": ["fact_0025", "fact_0031"],
|
||||
"merge_reason": "Semantic equivalence: Both items state that the Head of R&D leads the project list at the bi-weekly interface stand-up meeting. The difference in wording ('auf den' vs 'auf dem') and minor context does not change the core factual proposition."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Zu diesem Zeitpunkt sind Projekte bereits nach gewissen Kriterien kategorisiert; Vorfilter haben entlaufen.",
|
||||
"source_item_ids": ["fact_0026"],
|
||||
"merge_reason": "Singleton: Describes a state of categorization and filtering. Distinct from the definition of filters or specific criteria."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Die Bearbeitung der Projekte muss und kann nur in den Fachabteilungen passieren.",
|
||||
"source_item_ids": ["fact_0027"],
|
||||
"merge_reason": "Singleton: Defines a strict constraint on where project processing occurs. Distinct from general responsibility assignments or process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Ich kann nicht darüber entscheiden, ob ein Marketingprojekt erfolgreich gelaufen ist.",
|
||||
"source_item_ids": ["fact_0028"],
|
||||
"merge_reason": "Singleton: Expresses a lack of competence/authority regarding marketing project success. Distinct from general responsibility assignments or process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Bestimmte Dinge müssen bereits jetzt im Prozess definiert werden.",
|
||||
"source_item_ids": ["fact_0029"],
|
||||
"merge_reason": "Singleton: Expresses a requirement for early definition. Distinct from specific criteria or process steps."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Jede Abteilung hat diese Kriterien ja im Grunde sowieso.",
|
||||
"source_item_ids": ["fact_0030"],
|
||||
"merge_reason": "Singleton: States a general assumption about departmental possession of criteria. Distinct from the specific content or application of those criteria."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Der erste Filter dient als Entscheidungsvorlage und nicht als endgültige Entscheidung.",
|
||||
"source_item_ids": ["fact_0032"],
|
||||
"merge_reason": "Singleton: Defines the specific nature/function of the first filter. Distinct from other filters or general filtering concepts."
|
||||
},
|
||||
{
|
||||
"canonical_text": "Es gibt zwei Orte zum Filtern: einmal vorher (Vorfilter) und dann das Gremium.",
|
||||
"source_item_ids": ["fact_0033"],
|
||||
"merge_reason": "Singleton: Describes the structure of filtering locations. Distinct from specific criteria or filter definitions."
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
# Semantic Consolidator V0 Report
|
||||
|
||||
- Scope: facts only
|
||||
- Model: `qwen3.5:9B`
|
||||
- LLM call count: 1
|
||||
- Runtime: 38.451 seconds
|
||||
- Fact item count: 33
|
||||
- Prompt characters: 15393
|
||||
- Estimated prompt tokens: 3849
|
||||
- Merged fact groups: 1
|
||||
- Source facts involved in merges: 2
|
||||
- Singleton fact groups: 31
|
||||
- Output path: `samples\benchmarks\hardware_experimental_qwen35_9b\semantic_consolidator_v0\consolidated_extractions.json`
|
||||
|
||||
## Actual Merges
|
||||
|
||||
### Der Leiter F&E führt die Projektliste auf dem zweiwöchentlichen Schnittstellen-Stand-Up.
|
||||
|
||||
- Source fact IDs: fact_0025, fact_0031
|
||||
- Merge reason: Semantic equivalence: Both items state that the Head of R&D leads the project list at the bi-weekly interface stand-up meeting. The difference in wording ('auf den' vs 'auf dem') and minor context does not change the core factual proposition.
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
SAMPLE 0 2026-08-01T08:44:56.1505788+02:00 FreeMemKB=18388280 OllamaCPU=14.4375 OllamaWS=67276800
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 3 minutes from now
|
||||
|
||||
JOB_EXIT 2026-08-01T08:45:01.4585914+02:00 State=Running ActiveConfirmed=True
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
# Working Protocol Renderer V2 Report
|
||||
|
||||
- Result: invalid benchmark run
|
||||
- Valid Working Protocol output: no
|
||||
- Runtime: not measured
|
||||
- Model state at attempted start: warm (`qwen3.5:9b` was loaded from Run 1)
|
||||
- Intended model: `qwen3.5:9B`
|
||||
- Intended thinking setting: disabled (`think=false`)
|
||||
- Intended endpoint: `http://127.0.0.1:11434/api/generate`
|
||||
- Intended prompt: `prompts/working_protocol.md` unchanged
|
||||
- Intended input: `samples/benchmarks/hardware_experimental_qwen35_9b/semantic_consolidator_v0/consolidated_extractions.json`
|
||||
- Failure: PowerShell background job used the wrong working directory and could not resolve repository-relative paths.
|
||||
- Retry policy: no second renderer request was started because exactly-one-request compliance could not be proven after the failed wrapper attempt.
|
||||
+1
File diff suppressed because one or more lines are too long
+43
@@ -0,0 +1,43 @@
|
||||
# Working Protocol Renderer V2 Python Report
|
||||
|
||||
- Result: valid benchmark run
|
||||
- Model: `qwen3.5:9B`
|
||||
- Thinking setting: disabled (`think=false`)
|
||||
- Endpoint: `http://127.0.0.1:11434/api/generate`
|
||||
- Prompt path: `C:\Users\marti\Documents\git-projects\meeting-lab\prompts\working_protocol.md`
|
||||
- Input path: `C:\Users\marti\Documents\git-projects\meeting-lab\samples\benchmarks\hardware_experimental_qwen35_9b\semantic_consolidator_v0\consolidated_extractions.json`
|
||||
- Cold/warm state: cold
|
||||
- Request count: 1
|
||||
- Wall-clock runtime: 37.675 seconds
|
||||
- MacBook Air M4 baseline: 722.900 seconds
|
||||
- Speedup: 19.19x
|
||||
- total_duration: 37643714600
|
||||
- load_duration: 4037439800
|
||||
- prompt_eval_count: 28349
|
||||
- prompt_eval_duration: 14694067000
|
||||
- eval_count: 1268
|
||||
- eval_duration: 18840999000
|
||||
- Prompt characters: 112964
|
||||
- Output characters: 5567
|
||||
- Output bytes: 5626
|
||||
- Detected output language: German
|
||||
- Readable Markdown: True
|
||||
- Output path: `C:\Users\marti\Documents\git-projects\meeting-lab\samples\benchmarks\hardware_experimental_qwen35_9b\working_protocol_renderer_v2_python\working_protocol.md`
|
||||
- Raw response path: `C:\Users\marti\Documents\git-projects\meeting-lab\samples\benchmarks\hardware_experimental_qwen35_9b\working_protocol_renderer_v2_python\raw_ollama_response.json`
|
||||
|
||||
## Response Diagnostics
|
||||
|
||||
- HTTP status code: 200
|
||||
- Content-Type: `application/json; charset=utf-8`
|
||||
- Response byte length: 148913
|
||||
- Top-level JSON keys: `context, created_at, done, done_reason, eval_count, eval_duration, load_duration, model, prompt_eval_count, prompt_eval_duration, response, total_duration`
|
||||
- Response text length: 5567
|
||||
- done: True
|
||||
- done_reason: stop
|
||||
|
||||
## GPU Snapshot
|
||||
|
||||
```text
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
```
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
REQUEST_START 2026-08-01T09:09:01+0200 endpoint=http://127.0.0.1:11434/api/generate
|
||||
REQUEST_END 2026-08-01T09:09:38+0200 status=200
|
||||
+274
@@ -0,0 +1,274 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run the Meeting Lab Working Protocol Renderer V2 hardware benchmark once."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
REPO = Path(r"C:\Users\marti\Documents\git-projects\meeting-lab")
|
||||
PROMPT_PATH = REPO / "prompts" / "working_protocol.md"
|
||||
INPUT_PATH = (
|
||||
REPO
|
||||
/ "samples"
|
||||
/ "benchmarks"
|
||||
/ "hardware_experimental_qwen35_9b"
|
||||
/ "semantic_consolidator_v0"
|
||||
/ "consolidated_extractions.json"
|
||||
)
|
||||
OUTPUT_DIR = (
|
||||
REPO
|
||||
/ "samples"
|
||||
/ "benchmarks"
|
||||
/ "hardware_experimental_qwen35_9b"
|
||||
/ "working_protocol_renderer_v2_python"
|
||||
)
|
||||
RAW_RESPONSE_PATH = OUTPUT_DIR / "raw_ollama_response.json"
|
||||
PROTOCOL_PATH = OUTPUT_DIR / "working_protocol.md"
|
||||
REPORT_PATH = OUTPUT_DIR / "report.md"
|
||||
REQUEST_MARKER_PATH = OUTPUT_DIR / "request_count_marker.txt"
|
||||
|
||||
MODEL = "qwen3.5:9B"
|
||||
ENDPOINT = "http://127.0.0.1:11434/api/generate"
|
||||
NUM_CTX = 32768
|
||||
NUM_PREDICT = 4096
|
||||
TEMPERATURE = 0.0
|
||||
TIMEOUT_SECONDS = 1800
|
||||
MAC_RENDERER_BASELINE_SECONDS = 722.9
|
||||
|
||||
|
||||
def run_ollama_ps() -> str:
|
||||
result = subprocess.run(
|
||||
["ollama", "ps"],
|
||||
check=False,
|
||||
capture_output=True,
|
||||
text=True,
|
||||
encoding="utf-8",
|
||||
errors="replace",
|
||||
)
|
||||
return result.stdout.strip()
|
||||
|
||||
|
||||
def detect_language(text: str) -> str:
|
||||
lower = text.lower()
|
||||
german_markers = [
|
||||
" der ",
|
||||
" die ",
|
||||
" das ",
|
||||
" und ",
|
||||
" nicht ",
|
||||
" verantwortlich",
|
||||
" entscheidung",
|
||||
" offene",
|
||||
" fragen",
|
||||
" maßnahmen",
|
||||
" fuer ",
|
||||
" für ",
|
||||
]
|
||||
english_markers = [
|
||||
" the ",
|
||||
" and ",
|
||||
" not ",
|
||||
" responsible",
|
||||
" decision",
|
||||
" open questions",
|
||||
" action items",
|
||||
]
|
||||
german_score = sum(lower.count(marker) for marker in german_markers)
|
||||
english_score = sum(lower.count(marker) for marker in english_markers)
|
||||
if german_score > english_score:
|
||||
return "German"
|
||||
if english_score > german_score:
|
||||
return "English"
|
||||
return "unknown"
|
||||
|
||||
|
||||
def response_text_from_ollama_data(data: dict[str, Any]) -> str | None:
|
||||
text = data.get("response")
|
||||
if isinstance(text, str):
|
||||
return text
|
||||
message = data.get("message")
|
||||
if isinstance(message, dict):
|
||||
content = message.get("content")
|
||||
if isinstance(content, str):
|
||||
return content
|
||||
return None
|
||||
|
||||
|
||||
def is_readable_markdown(text: str) -> bool:
|
||||
stripped = text.strip()
|
||||
return bool(stripped) and stripped.startswith("#") and "\n" in stripped
|
||||
|
||||
|
||||
def write_report(lines: list[str]) -> None:
|
||||
REPORT_PATH.write_text("\n".join(lines).rstrip() + "\n", encoding="utf-8")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||
if not PROMPT_PATH.exists():
|
||||
print(f"Missing prompt: {PROMPT_PATH}", file=sys.stderr)
|
||||
return 1
|
||||
if not INPUT_PATH.exists():
|
||||
print(f"Missing input: {INPUT_PATH}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
cold_warm = "warm" if MODEL.lower() in run_ollama_ps().lower() else "cold"
|
||||
prompt = PROMPT_PATH.read_text(encoding="utf-8")
|
||||
input_text = INPUT_PATH.read_text(encoding="utf-8-sig")
|
||||
full_prompt = (
|
||||
prompt
|
||||
+ "\n\nCONSOLIDATED MEETING REPRESENTATION:\n"
|
||||
+ input_text
|
||||
)
|
||||
payload = {
|
||||
"model": MODEL,
|
||||
"prompt": full_prompt,
|
||||
"think": False,
|
||||
"stream": False,
|
||||
"options": {
|
||||
"temperature": TEMPERATURE,
|
||||
"num_ctx": NUM_CTX,
|
||||
"num_predict": NUM_PREDICT,
|
||||
},
|
||||
}
|
||||
|
||||
REQUEST_MARKER_PATH.write_text(
|
||||
f"REQUEST_START {time.strftime('%Y-%m-%dT%H:%M:%S%z')} endpoint={ENDPOINT}\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
started = time.perf_counter()
|
||||
response = requests.post(ENDPOINT, json=payload, timeout=TIMEOUT_SECONDS)
|
||||
runtime = time.perf_counter() - started
|
||||
REQUEST_MARKER_PATH.write_text(
|
||||
REQUEST_MARKER_PATH.read_text(encoding="utf-8")
|
||||
+ f"REQUEST_END {time.strftime('%Y-%m-%dT%H:%M:%S%z')} status={response.status_code}\n",
|
||||
encoding="utf-8",
|
||||
)
|
||||
|
||||
RAW_RESPONSE_PATH.write_bytes(response.content)
|
||||
|
||||
print(f"HTTP status code: {response.status_code}")
|
||||
print(f"response Content-Type: {response.headers.get('Content-Type', '')}")
|
||||
print(f"response byte length: {len(response.content)}")
|
||||
|
||||
try:
|
||||
data = response.json()
|
||||
except ValueError as exc:
|
||||
print(f"Invalid JSON response: {exc}", file=sys.stderr)
|
||||
write_report(
|
||||
[
|
||||
"# Working Protocol Renderer V2 Python Report",
|
||||
"",
|
||||
"- Result: invalid benchmark run",
|
||||
f"- HTTP status code: {response.status_code}",
|
||||
f"- Response byte length: {len(response.content)}",
|
||||
f"- Error: invalid JSON response: {exc}",
|
||||
"- Request count: 1",
|
||||
]
|
||||
)
|
||||
return 1
|
||||
|
||||
if not isinstance(data, dict):
|
||||
print("Top-level JSON is not an object.", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
keys = sorted(data.keys())
|
||||
text = response_text_from_ollama_data(data)
|
||||
text_length = len(text) if isinstance(text, str) else 0
|
||||
print(f"top-level JSON keys: {keys}")
|
||||
print(f"response text length: {text_length}")
|
||||
print(f"done flag: {data.get('done')}")
|
||||
print(f"done_reason: {data.get('done_reason')}")
|
||||
|
||||
if response.status_code != 200 or not response.content or not isinstance(text, str):
|
||||
write_report(
|
||||
[
|
||||
"# Working Protocol Renderer V2 Python Report",
|
||||
"",
|
||||
"- Result: invalid benchmark run",
|
||||
f"- HTTP status code: {response.status_code}",
|
||||
f"- Response byte length: {len(response.content)}",
|
||||
f"- Top-level JSON keys: {', '.join(keys)}",
|
||||
f"- Response text length: {text_length}",
|
||||
f"- done: {data.get('done')}",
|
||||
f"- done_reason: {data.get('done_reason')}",
|
||||
"- Request count: 1",
|
||||
]
|
||||
)
|
||||
return 1
|
||||
|
||||
PROTOCOL_PATH.write_text(text, encoding="utf-8")
|
||||
|
||||
language = detect_language(text)
|
||||
markdown_ok = is_readable_markdown(text)
|
||||
speedup = MAC_RENDERER_BASELINE_SECONDS / runtime if runtime > 0 else 0.0
|
||||
output_bytes = len(text.encode("utf-8"))
|
||||
validation_ok = bool(text.strip()) and markdown_ok and language == "German"
|
||||
|
||||
total_duration = data.get("total_duration")
|
||||
load_duration = data.get("load_duration")
|
||||
prompt_eval_count = data.get("prompt_eval_count")
|
||||
prompt_eval_duration = data.get("prompt_eval_duration")
|
||||
eval_count = data.get("eval_count")
|
||||
eval_duration = data.get("eval_duration")
|
||||
gpu_snapshot = run_ollama_ps()
|
||||
|
||||
write_report(
|
||||
[
|
||||
"# Working Protocol Renderer V2 Python Report",
|
||||
"",
|
||||
"- Result: valid benchmark run" if validation_ok else "- Result: invalid benchmark run",
|
||||
f"- Model: `{MODEL}`",
|
||||
"- Thinking setting: disabled (`think=false`)",
|
||||
f"- Endpoint: `{ENDPOINT}`",
|
||||
f"- Prompt path: `{PROMPT_PATH}`",
|
||||
f"- Input path: `{INPUT_PATH}`",
|
||||
f"- Cold/warm state: {cold_warm}",
|
||||
"- Request count: 1",
|
||||
f"- Wall-clock runtime: {runtime:.3f} seconds",
|
||||
f"- MacBook Air M4 baseline: {MAC_RENDERER_BASELINE_SECONDS:.3f} seconds",
|
||||
f"- Speedup: {speedup:.2f}x",
|
||||
f"- total_duration: {total_duration}",
|
||||
f"- load_duration: {load_duration}",
|
||||
f"- prompt_eval_count: {prompt_eval_count}",
|
||||
f"- prompt_eval_duration: {prompt_eval_duration}",
|
||||
f"- eval_count: {eval_count}",
|
||||
f"- eval_duration: {eval_duration}",
|
||||
f"- Prompt characters: {len(full_prompt)}",
|
||||
f"- Output characters: {len(text)}",
|
||||
f"- Output bytes: {output_bytes}",
|
||||
f"- Detected output language: {language}",
|
||||
f"- Readable Markdown: {markdown_ok}",
|
||||
f"- Output path: `{PROTOCOL_PATH}`",
|
||||
f"- Raw response path: `{RAW_RESPONSE_PATH}`",
|
||||
"",
|
||||
"## Response Diagnostics",
|
||||
"",
|
||||
f"- HTTP status code: {response.status_code}",
|
||||
f"- Content-Type: `{response.headers.get('Content-Type', '')}`",
|
||||
f"- Response byte length: {len(response.content)}",
|
||||
f"- Top-level JSON keys: `{', '.join(keys)}`",
|
||||
f"- Response text length: {text_length}",
|
||||
f"- done: {data.get('done')}",
|
||||
f"- done_reason: {data.get('done_reason')}",
|
||||
"",
|
||||
"## GPU Snapshot",
|
||||
"",
|
||||
"```text",
|
||||
gpu_snapshot,
|
||||
"```",
|
||||
]
|
||||
)
|
||||
return 0 if validation_ok else 1
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
+97
@@ -0,0 +1,97 @@
|
||||
# Working Protocol
|
||||
|
||||
## Prozessstruktur und Projekttypen
|
||||
|
||||
### Background
|
||||
|
||||
Der bestehende F&E-Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt. Der grundsätzliche Ablauf bleibt gleich, wobei spezifische Dokumente je nach Projekttyp angepasst werden. Die Bearbeitung der Projekte muss und kann nur in den Fachabteilungen passieren; die Projektleitung ist dabei nicht abteilungszugeordnet.
|
||||
|
||||
### Decisions
|
||||
|
||||
- Der bestehende Prozess wird grundsätzlich auch für Business Development (BD)-Projekte genutzt, wobei die spezifischen Dokumente je nach Projekttyp angepasst werden.
|
||||
- Die Prozessschritte für verschiedene Projekttypen (F&E, PM, Marketing, BD) werden im Swimlane-Diagramm offen gelassen und nicht formal festgelegt.
|
||||
- Es wird vereinbart, dass für BD-Projekte ein Projektdeckblatt und Lastenheft erstellt werden müssen.
|
||||
- Der erste Filter dient als Entscheidungsvorlage und nicht als endgültige Entscheidung.
|
||||
|
||||
### Action Items
|
||||
|
||||
- Giovanna stellt die Kriterien zusammen und erstellt einen ersten Entwurf für den Auswahlkatalog, der EDD-, Marketing- und PM-Kriterien integriert. (Verantwortlich: Giovanna)
|
||||
- Björn definiert andere Kriterien für Business Development Projekte (z.B. digitale Affinität, kulturelle Hürden) und integriert diese in den Filter. (Verantwortlich: Björn)
|
||||
|
||||
### Open Questions
|
||||
|
||||
- Wie werden spezifische Auswahlkriterien für digitale Produkte (z.B. Portal) versus Realprodukte definiert und integriert?
|
||||
- Wie genau sollen die Anforderungen an das Lastenheft für nicht-F&E-Projekte definiert werden?
|
||||
- Sollte die Vorab-Filterung durch eine zentrale Stelle oder dezentral in den Abteilungen erfolgen?
|
||||
- Wer übernimmt die Prozessverantwortlichkeit für die erste Prüfung von Projektideen?
|
||||
|
||||
## Minimalanforderungen und Filterkriterien
|
||||
|
||||
### Background
|
||||
|
||||
Neue Ideen müssen mindestens Titel, Projektidee und ein Projektziel beantworten. Optional ist eine konkrete Aufzählung der zugrunde liegenden Annahmen (z.B. erwartbares Marktvolumen). Vorab-Filter dienen als Entscheidungsvorlage; weitere Kriterien werden je nach Fall definiert. Strategiekonformität ist ein Filterkriterium.
|
||||
|
||||
### Decisions
|
||||
|
||||
- Festlegung der Mindestanforderungen für neue Projektideen (Titel, Idee, Ziel) und optionaler Annahmen.
|
||||
- Etablierung eines Prüfprozesses zur Unterscheidung zwischen abteilungsinternen Aufgaben und strategisch relevanten Ideen.
|
||||
- Zustimmung zur Einbindung von Business Development für die Vorsortierung und das Abfragen der Minimaldaten.
|
||||
|
||||
### Action Items
|
||||
|
||||
- Ideen müssen mit Minimalanforderungen angereichert werden, bevor sie bearbeitet werden können. (Verantwortung offen)
|
||||
- Bei BD-Projekten müssen typische drei bis vier typische BD-Sachen abgefragt werden. (Verantwortung offen)
|
||||
- Definition von Filterkriterien für verschiedene Fälle (neue Produkte/Anwendungen vs. bestehende Märkte). (Verantwortung offen)
|
||||
|
||||
### Open Questions
|
||||
|
||||
- Welche Prüfsteine sind relevant für die Fachabteilung?
|
||||
|
||||
## Projektlaufbahn und Gremiumsarbeit
|
||||
|
||||
### Background
|
||||
|
||||
Der Prozess beginnt mit der Entscheidung, ob ein Projekt in die Strategie passt. Wenn ja, wird entschieden, in welche Abteilung (BD, PM, Marketing) es tropft. Zu diesem Zeitpunkt sind Projekte bereits nach gewissen Kriterien kategorisiert; Vorfilter haben entlaufen. Ab dem Termin, wo das Projekt vorgestellt wird, muss jemand den Hut aufsetzen und prüfen oder einen Projektplan erstellen.
|
||||
|
||||
### Decisions
|
||||
|
||||
- Die Bearbeitung der Projekte muss und kann nur in den Fachabteilungen passieren.
|
||||
- Der Prozess soll nicht zu komplex gemacht werden; es wird ein Gefühl für notwendige Informationen entwickelt.
|
||||
- Der Miro-Link wird in den Chat kopiert und allen relevanten Personen (insb. Henning) zur Verfügung gestellt.
|
||||
|
||||
### Action Items
|
||||
|
||||
- Björn soll mit PM oder Business Development über die Diskussion von Projekten (z.B. Papatikus-Anker) sprechen und diese zu den Fachabteilungen rüberspielen.
|
||||
- Fachabteilungen sollen den Katalog ausarbeiten und zur Entscheidung bringen.
|
||||
- Eine Liste führen, die zeigt, wer sich um was kümmert.
|
||||
- Prozess etablieren, in dem Filter sicher abgefragt werden.
|
||||
- Gatekeeper (z.B. James und David) sollen formal ein paar Sachen abarbeiten und Informationen zusammentragen.
|
||||
- Raik muss Ideen in den Pool tragen lassen (nicht nur Stichworte reinwerfen).
|
||||
- Nachdem eine Idee im Prozess ist (z.B. nach 4 Wochen), Status prüfen: Wird es bearbeitet?
|
||||
- Miro-Link in den Chat kopieren und an alle verteilen. (Verantwortlich: Malte)
|
||||
|
||||
### Open Questions
|
||||
|
||||
- Für welche Projekte soll das übergeordnete Projektmanagement gewünscht werden?
|
||||
- Wer stellt welches Projekt in welchem Gremium vor?
|
||||
|
||||
## Organisationale Rahmenbedingungen und Ressourcen
|
||||
|
||||
### Background
|
||||
|
||||
Die Geschäftsführung hat den Überblick über Projekte im Haus zu behalten. Der F&E-Topf wird gemeinsam geführt. Das Netzwerk bei Gemeinden ist sehr langsam, Sekretariate sind da sehr abweisend. Die Durchsetzung und Umsetzung ist das Entscheidende am Schluss. Man sollte nicht im Vorhinein versuchen, alles zu planen.
|
||||
|
||||
### Decisions
|
||||
|
||||
- Es wurde vereinbart, Filterkriterien für verschiedene Fälle (neue Produkte/Anwendungen vs. bestehende Märkte) zu definieren.
|
||||
- Die Gruppe stimmt der gemeinsamen Meinung und dem vorgeschlagenen Vorgehen zur Prozessanpassung zu.
|
||||
|
||||
### Action Items
|
||||
|
||||
- Ein paar Prüfsteine definieren. (Verantwortung offen)
|
||||
- Giovanna und Björn in die Diskussion über Filterkriterien einbeziehen. (Verantwortung offen)
|
||||
|
||||
### Open Questions
|
||||
|
||||
- Sind wir so davon überzeugt, dass wir die Marktstudie anschieben?
|
||||
- An welcher Stelle entscheidet eine Abteilung einfach so und an welcher Stelle nicht?
|
||||
+32
@@ -0,0 +1,32 @@
|
||||
SAMPLE 0 2026-08-01T08:51:07.5108682+02:00 FreeMemKB=19747780 OllamaCPU=14.6875 OllamaWS=43548672
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
|
||||
SAMPLE 1 2026-08-01T08:51:12.6834762+02:00 FreeMemKB=18754304 OllamaCPU=15.1875 OllamaWS=61841408
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 2 2026-08-01T08:51:17.8264165+02:00 FreeMemKB=18733028 OllamaCPU=15.1875 OllamaWS=61841408
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 3 2026-08-01T08:51:22.9486785+02:00 FreeMemKB=18508580 OllamaCPU=15.1875 OllamaWS=61841408
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 4 2026-08-01T08:51:28.0758752+02:00 FreeMemKB=18453568 OllamaCPU=15.21875 OllamaWS=46608384
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 5 2026-08-01T08:51:33.2017774+02:00 FreeMemKB=18453764 OllamaCPU=15.265625 OllamaWS=46620672
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 6 2026-08-01T08:51:38.3370762+02:00 FreeMemKB=18457940 OllamaCPU=15.28125 OllamaWS=46653440
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
SAMPLE 7 2026-08-01T08:51:43.4802353+02:00 FreeMemKB=18472864 OllamaCPU=15.3125 OllamaWS=46653440
|
||||
NAME ID SIZE PROCESSOR CONTEXT UNTIL
|
||||
qwen3.5:9b 6488c96fa5fa 6.7 GB 100% GPU 32768 4 minutes from now
|
||||
|
||||
JOB_EXIT 2026-08-01T08:51:48.6321784+02:00 State=Running ActiveConfirmed=True
|
||||
+1
@@ -0,0 +1 @@
|
||||
|
||||
+20
@@ -0,0 +1,20 @@
|
||||
# Working Protocol Renderer V2 Retry Report
|
||||
|
||||
- Result: invalid benchmark run
|
||||
- Valid Working Protocol output: no
|
||||
- Runtime: not valid for benchmark comparison
|
||||
- Preflight repository root: `C:\Users\marti\Documents\git-projects\meeting-lab`
|
||||
- Required prompt existed: yes
|
||||
- Required consolidated input existed: yes
|
||||
- Renderer request currently running before start: no visible request in `ollama ps`
|
||||
- Model state before start: cold; `ollama ps` was empty
|
||||
- Intended model: `qwen3.5:9B`
|
||||
- Intended thinking setting: disabled (`think=false`)
|
||||
- Intended endpoint: `http://127.0.0.1:11434/api/generate`
|
||||
- Intended prompt: `prompts/working_protocol.md` unchanged
|
||||
- Intended input: `samples/benchmarks/hardware_experimental_qwen35_9b/semantic_consolidator_v0/consolidated_extractions.json`
|
||||
- Request marker: one request start and one request end recorded in `request_count_marker.txt`
|
||||
- Failure: captured raw response file contains only a UTF-8 BOM; no usable Ollama JSON body was available for renderer output or metrics extraction
|
||||
- GPU confirmation: `ollama ps` showed `qwen3.5:9b` with `PROCESSOR 100% GPU`, context 32768, model size 6.7 GB during the request
|
||||
- Validation: failed; `working_protocol.md` does not exist
|
||||
- Retry policy: no additional renderer request was started
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
REQUEST_START 2026-08-01T08:51:07.6738275+02:00 endpoint=http://127.0.0.1:11434/api/generate
|
||||
REQUEST_END 2026-08-01T08:51:48.6341848+02:00 status=
|
||||
Reference in New Issue
Block a user