Pull requests / #508
#508 Community benchmark: 2x RTX 5060 Ti (16 GB), Xeon E5-2690 v4, IQ3_XXS, engine 0.1.35
closed · @R6DJO · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentationLinux
描述
Results-only community benchmark report, per docs/COMMUNITY_BENCHMARKS.md. No engine changes. ## Setup - Strata 0.1.35, source build, commit `d9ab843` (current `main`); CUDA 13.4, driver 610.57.04, Linux Mint 22.3. - 2x RTX 5060 Ti 16 GB (layer-split auto, each card holds its weights), Xeon E5-2690 v4 (no AVX-512), 125 GB RAM. - Original Flash-Next GSQ-RCO IQ3_XXS, 262,144-token context, KV int8, MTP on, calibrated settings from a previous `--calibrate` on this machine. ## Results (9 serial fresh-prompt runs, greedy, 256-token cap, `reused: 0` on all) | Prompt | Prompt tok/s (median, min-max) | Decode tok/s (median, min-max) | TTFT s (median) | | ---: | --- | --- | ---: | | 4,096 | 808 (785-810) | 77.5 (66.7-77.9) | 5.11 | | 32,768 | 1,931 (1,928-1,931) | 72.7 (70.9-73.5) | 17.09 | | 128,000 | 2,199 (2,197-2,200) | 72.4 (72.1-75.8) | 58.67 | Prompt and decode tok/s are from the engine's per-request `prompt_ms` / `decode_ms`, not total request time. ## Correctness Needle recall (`tools/needle_bench.py`) at 32K and 128K, depths 10/50/90: **6/6 exact** (needles.json). ## Limitations One quantization (IQ3_XXS); one synthetic code-explanation workload; output capped at 256 tokens; GPU clocks not fixed; the expert cache was still warming on the first measured run (hit rate 48% warm-up -> 87% first run -> 98-99% onward, all values in results.json); RAM/VRAM during inference not sampled; exact GGUF repository revision not recorded (file names and sizes are in model-files.txt). Desktop services ran on the same machine. Local paths in the report are replaced with `<strata-dir>` / `<strata-data>` / `<gguf-dir>` placeholders. --- ## Rerun 2026-10-05 (`bench/results/2026-10-05-community-rtx-5060-ti-x2`) Same machine (hardware clarified: **1 CPU, 1 socket**; RAM is quad-channel **4x 32 GB DDR4-2133 ECC** — the "125 GB" in the first report is the same memory as MemTotal). Engine updated to **0.1.39** with a changed launch configuration (`--remote-expert-opt --conversation-cache-mib 8192 --trim-stage-weights`, learned expert profile, fixed `--layer-split 25`), so the speed difference is not attributable to any single change. | Prompt | Prompt tok/s (median, min-max) | Decode tok/s (median, min-max) | TTFT s (median) | | ---: | --- | --- | ---: | | 4,096 | 911 (911-915) | 84.8 (74.3-88.9) | 4.54 | | 32,768 | 2,087 (2,087-2,099) | 86.5 (80.1-87.2) | 15.82 | | 128,000 | 2,440 (2,438-2,457) | 83.8 (81.4-84.0) | 52.87 | Needle recall at 32K and 128K, depths 10/50/90: **6/6 exact** (`needles.json`).
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。