Pull requests / #1291

#1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context

open · @NRK-SH · 0 Kommentare · Auf GitHub

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

Beschreibung

This PR adds a community benchmark report under `bench/results/`, following
`docs/COMMUNITY_BENCHMARKS.md`.

### What was measured

Strata 0.1.34 (source build with a local Ivy Bridge compatibility port) serving
the GSQ-RCO IQ3_S quantization of Qwen3.8-Flash-Next - a 125B-parameter MoE
model with 6B active parameters per token - on an HP Z820 workstation:
2x Intel Xeon E5-2697 v2 (AVX only; no AVX2/FMA), 256 GB DDR3, one RTX 3090
24 GB, Windows 11 Pro for Workstations. Hybrid inference: GPU expert cache
(7,237 slots / 13.73 GiB), CPU fallback (23 workers), INT8 KV streaming,
mmap-based n-gram / PLE tables, MTP speculative decoding.

**Keywords:** Qwen3.8-Flash-Next, GSQ-RCO, IQ3_S, Strata, MoE, 125B, 512K
context, AVX-only, Xeon E5-2697 v2, HP Z820, RTX 3090, hybrid inference,
expert cache, CPU fallback, MTP, GGUF, long-context, vision, tool calling,
Windows 11.

### Model

`ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, revision `ed59f920...`, IQ3_S
(54,817,524,224 + 28,800,138,432 bytes; both SHA-256 hashes match the
repository's published LFS hashes). Local pack prepared with
`compat_bf16: false`; BF16 vision projector; MTP draft pack.

### Method

Three to four fresh-prompt runs at approximately 4K / 32K / 128K prompt tokens
and two runs at approximately 508K, in increasing-length order on the same
loaded engine. Greedy decoding, thinking disabled, 320-token output cap, and a
unique nonce per request to prevent prefix reuse (`cache_n = 0` for every run).
Engine timing lines for all runs are attached. Long-context recall was measured
with the unmodified `tools/needle_bench.py` at 32k/128k, depths 10/50/90.

### Results (medians)

| Size | Actual prompt tokens | Prompt tok/s | Decode tok/s | TTFT (s) |
| --- | ---: | ---: | ---: | ---: |
| ~4K | 4,139 | 995.4 | 87.4 | 4.22 |
| ~32K | 32,646 | 1,659.8 | 85.3 | 19.85 |
| ~128K | 130,777 | 1,648.0 | 74.9 | 80.08 |
| ~508K | 508,255 | 1,071.9 | 59.0 | 477.1 |

Recall: 6/6. Additional checks: vision code-word question, tool-call round
trip, and a small objective task set (8/9).

### Attachments

`README.md`, `benchmark.py`, `results.json`, `results.csv`, `summary.json`,
`engine-timings.log`, `needles.json`, `memory-summary.json`,
`strata-iq3s-512k-vision.json`, `pack-conversions.json`,
`provenance-local.json`, `hashes.sha256`, and `assets/` (charts and system
overview).

### Limitations

One machine, a custom pack and a local engine build; synthetic workload;
single-sequence operation. The 508K runs processed 508,255 tokens, slightly
below the full 524,288-token window. No general quality claim is made. The
report is results-only and does not touch engine code.

Happy to adjust the structure or add measurements if that would make the
report more useful.

Mehr auf der Site

Links zu Install, Modellen, Releases.