Pull requests / #1291

#1291 bench: community report - HP Z820, dual Xeon E5-2697 v2 (AVX only) + RTX 3090, Qwen3.8-Flash-Next IQ3_S, 512K context

open · @NRK-SH · 0 コメント · GitHub で見る

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentationWindows

本文

This PR adds a community benchmark report under `bench/results/`, following
`docs/COMMUNITY_BENCHMARKS.md`.

### What was measured

Strata 0.1.34 (source build with a local Ivy Bridge compatibility port) serving
the GSQ-RCO IQ3_S quantization of Qwen3.8-Flash-Next - a 125B-parameter MoE
model with 6B active parameters per token - on an HP Z820 workstation:
2x Intel Xeon E5-2697 v2 (AVX only; no AVX2/FMA), 256 GB DDR3, one RTX 3090
24 GB, Windows 11 Pro for Workstations. Hybrid inference: GPU expert cache
(7,237 slots / 13.73 GiB), CPU fallback (23 workers), INT8 KV streaming,
mmap-based n-gram / PLE tables, MTP speculative decoding.

**Keywords:** Qwen3.8-Flash-Next, GSQ-RCO, IQ3_S, Strata, MoE, 125B, 512K
context, AVX-only, Xeon E5-2697 v2, HP Z820, RTX 3090, hybrid inference,
expert cache, CPU fallback, MTP, GGUF, long-context, vision, tool calling,
Windows 11.

### Model

`ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF`, revision `ed59f920...`, IQ3_S
(54,817,524,224 + 28,800,138,432 bytes; both SHA-256 hashes match the
repository's published LFS hashes). Local pack prepared with
`compat_bf16: false`; BF16 vision projector; MTP draft pack.

### Method

Three to four fresh-prompt runs at approximately 4K / 32K / 128K prompt tokens
and two runs at approximately 508K, in increasing-length order on the same
loaded engine. Greedy decoding, thinking disabled, 320-token output cap, and a
unique nonce per request to prevent prefix reuse (`cache_n = 0` for every run).
Engine timing lines for all runs are attached. Long-context recall was measured
with the unmodified `tools/needle_bench.py` at 32k/128k, depths 10/50/90.

### Results (medians)

| Size | Actual prompt tokens | Prompt tok/s | Decode tok/s | TTFT (s) |
| --- | ---: | ---: | ---: | ---: |
| ~4K | 4,139 | 995.4 | 87.4 | 4.22 |
| ~32K | 32,646 | 1,659.8 | 85.3 | 19.85 |
| ~128K | 130,777 | 1,648.0 | 74.9 | 80.08 |
| ~508K | 508,255 | 1,071.9 | 59.0 | 477.1 |

Recall: 6/6. Additional checks: vision code-word question, tool-call round
trip, and a small objective task set (8/9).

### Attachments

`README.md`, `benchmark.py`, `results.json`, `results.csv`, `summary.json`,
`engine-timings.log`, `needles.json`, `memory-summary.json`,
`strata-iq3s-512k-vision.json`, `pack-conversions.json`,
`provenance-local.json`, `hashes.sha256`, and `assets/` (charts and system
overview).

### Limitations

One machine, a custom pack and a local engine build; synthetic workload;
single-sequence operation. The 508K runs processed 508,255 tokens, slightly
below the full 524,288-token window. No general quality claim is made. The
report is results-only and does not touch engine code.

Happy to adjust the structure or add measurements if that would make the
report more useful.

関連リンク

インストール・モデル・リリースへの站内リンク。