贡献 / #1062
#1062 bench: community report, 2x RTX 4060 Ti 16 GB + Threadripper PRO 3975WX, UD-IQ4_XS at 131K, layer split
closed · @fioruccione · 0 评论 · 去 GitHub 看
BenchmarksSetup & installServer & APIMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
说明
A community benchmark report, data only: one new folder under `bench/results/` and its line in `docs/COMMUNITY_BENCHMARKS.md`. No engine or setup changes. **Machine:** 2x NVIDIA RTX 4060 Ti 16 GB (AD106, PCIe 4.0 x8 each, no P2P), AMD Threadripper PRO 3975WX (32C, AVX2), 128 GB DDR4, Pop!_OS 24.04, driver 610.57.04, source build for sm_89 with CUDA 12.6. **Model:** Unsloth UD-IQ4_XS at setup's pinned revision (hashes checked by setup), 131,072-token context, layer split `auto` (CUDA0 layers 0-22, CUDA1 23-47), experts partly in VRAM (hit rate 0.80-0.90), `--kv int8 --kv-resident 32768 --spec 4 --mtp --remote-expert-opt --pcie-frac 0.00`, one request at a time. **What was run:** the RTX 5090 report's `benchmark.py` and `monitor.py`, unchanged (three runs each at 4,096 / 32,768 / 128,000 prompt tokens, 256-token greedy outputs, 0 reused tokens), in two configurations that differ only in the MTP draft vocabulary: setup's default (`stock`) and one rebuilt from Italian Wikipedia text (`prod`, what this machine serves day to day). `tools/needle_bench.py` at 32K and 128K, depths 10/50/90, on `prod`. **Medians:** | | prompt tok/s at 4K / 32K / 128K | decode tok/s at 4K / 32K / 128K | TTFT at 128K | | --- | ---: | ---: | ---: | | `prod` (Italian draft vocab) | 880 / 1,917 / 2,221 | 54.4 / 56.6 / 54.7 | 58.0 s | | `stock` (default draft vocab) | 906 / 1,934 / 2,234 | 51.9 / 55.3 / 55.7 | 57.6 s | All 6 needles were found. No request failed or stalled; no swap used; VRAM peaked at ~15.7-15.8 GB per card. On these English/code prompts the Italian draft vocabulary was within run-to-run noise of the default one (slightly faster at 4K). The README states the limits: one machine, three runs per point, no concurrency in this report (`"parallel": 2` on this machine is in the comment on #857), no images, and none of the open prompt-path PRs. Developed with an AI coding assistant; every number above was measured on this machine. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
本站相关内容
相关页面的快捷入口。