贡献 / #1157

#1157 Community benchmark: 2x Tesla P100, Flash-Next IQ3_S, 128k context

closed · @lechlna000 · 0 评论 · 去 GitHub 看

BenchmarksMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

说明

Results-only submission, no engine changes.

**Hardware:** 2x Tesla P100 PCIe 16GB (Gen3 x16, layer split CUDA0 layers 0-22 / CUDA1 layers 23-47), Xeon E5-2699 v4, 125.67 GiB RAM, NVMe storage, Fedora 44, driver 580.178.04, CUDA 12.8.

**Software:** Strata `6f32ec0` (v0.1.39), locally built CUDA12 engine.

**Model:** ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S, 128,000-token context, INT8 KV (32,768 resident), expert cache auto (9,858 resident experts, profile-prefilled, no eviction), MTP --spec 4, GPU vision enabled but untested by these requests.

**Workload:** the same benchmark.py as the RTX 5090 community report. One warm-up excluded; three serial runs of a fresh 127,736-token prompt (zero reused tokens) with a 256-token output cap, greedy, reasoning off, on the same loaded engine (loading excluded).

**Results:** prefill 531.4 tok/s [526.6-532.1]; decode 31.2 tok/s [28.2-31.7]; client TTFT 350.9 s [240.4-352.4]; RAM peak 62.3 GiB (46.8 GiB arena), VRAM ~15.9/16.0 GiB per card, no paging/OOM. Needle recall 125k depths 10/50/90: 3/3 found.

**Limitations / notes:**
- Single prompt length only (no 4k/32k sweep, no single-GPU run).
- A nominal 128,000-token prompt + 256 cap is rejected by the server (prompt + max_tokens + 8 > context), so the target is 127,736; documented in the README.
- `needle_bench.py --lengths 128k` is skipped by the tool itself at ctx 128,000 (target 128,450); used 125k (actual 122,640 tokens).
- sm60 paths (no bf16 tensor cores); PLE shard read from NVMe with a warm page cache.

本站相关内容

相关页面的快捷入口。