Pull requests / #995
#995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)
closed · @lfontanez · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation
描述
## Adds a community benchmark entry per `docs/COMMUNITY_BENCHMARKS.md`
**Hardware:** NVIDIA GeForce RTX 4070 Ti SUPER, 16 GiB VRAM, sm_89, 94 GiB system RAM, Samsung 990 PRO 2 TB NVMe, Ubuntu 24.04.5 LTS, NVIDIA driver 595.91.07 (CUDA 13.2), nvidia-container-toolkit 1.20.1, Docker 29.8.1.
**Engine:** upstream-main commit `6f32ec07`, locally compiled `sm_89` fat-binary (`CUDA_ARCHITECTURES=89`, `BUILD_VISION=0`). Base image `nvidia/cuda:13.0.0-devel-ubuntu24.04`. Single consumer card; layer-split deliberately not exercised on the 16 GiB ceiling.
## Eight configurations measured
| Model | CONTEXT | KV | Cold prefill | Decode avg | Peak | Engine TPS | SSE prefill | Draft | TTFT |
|----------|---------:|------:|-------------:|-----------:|--------:|-----------:|------------:|------:|-------:|
| IQ2_XS | 32768 | int8 | **2 695** | 95.62 | 136.34 | 93.47 | 141.3 | 75.8% | 183 ms |
| IQ2_XS | 65536 | int8 | **2 678** | 91.80 | 123.20 | 89.90 | 140.9 | 74.5% | 271 ms |
| IQ2_XS | 131072 | int8 | **2 670** | 87.95 | 117.52 | 86.67 | 132.6 | 76.8% | 422 ms |
| IQ2_XS | 262144 | int8 | **2 484** | 79.63 | 107.36 | 78.50 | 127.3 | 69.6% | 726 ms |
| IQ3_XXS | 32768 | int8 | **2 747** | 73.32 | 97.96 | 71.73 | 116.5 | 69.8% | 217 ms |
| IQ3_XXS | 65536 | int8 | **2 545** | 69.16 | 98.66 | 68.36 | 107.0 | 70.3% | 295 ms |
| IQ3_XXS | 131072 | q4_0 | **2 478** | 65.51 | 84.42 | 64.10 | 105.2 | 67.7% | 461 ms |
| IQ3_XXS | 262144 | k8v4 | **2 404** | 53.70 | 71.09 | 52.97 | 91.9 | 66.6% | 768 ms |
Each row = 1 cold-prefill probe (`max_tokens=1` non-stream after `POST /unload`) + 3 SSE stream calls (`max_tokens=192`, temperature 0, `reasoning_effort="minimal"`).
## Recall (`tools/needle_bench.py --lengths 32k --depths 10,50,90`)
| model | depth | found | prompt tok | seconds |
|----------|------:|:-----:|----------:|--------:|
| IQ2_XS | 10 % | ✅ | 32 171 | 13 |
| IQ2_XS | 50 % | ✅ | 32 171 | 12 |
| IQ2_XS | 90 % | ✅ | 32 172 | 6 |
| IQ3_XXS | 10 % | ✅ | 32 171 | 12 |
| IQ3_XXS | 50 % | ✅ | 32 171 | 12 |
| IQ3_XXS | 90 % | ✅ | 32 172 | 6 |
**3/3 FOUND** for both IQ2_XS and IQ3_XXS at 32k context.
## Cross-table comparison with the 0.1.26 headline (RTX 5070 12 GB)
| model | 32k cold prefill (this host) | 32k cold prefill (RTX 5070) | delta |
|----------|------------------------------:|----------------------------:|------:|
| IQ2_XS | **2 695 t/s** | **2 092 t/s** | **+29 %** |
| IQ3_XXS | **2 747 t/s** | **1 745 t/s** | **+57 %** |
The 4070 Ti SUPER is comfortably ahead of the 5070 on cold prefill; the host's faster CPU pool + Gen4 host-to-device pipeline let the engine keep the KV-streaming reserve in RAM and run the experts through it without choking decode.
## Canonical `bench/results/2026-09-29-speed-0126/` extension
The upstream `matrix.json` was patched (no upstream row contents changed):
- Added a `kv` schema field. The 25 baseline rows are `int8` per `setup.py` defaults.
- Added a `prefill_cold_tok_s` column (true end-of-prompt throughput). `null` on the 25 baselines; populated for the 8 new rows added here.
- The README grew two sections: **KV per cell and True-Prefill (host extension)** (the 8 new rows) and **Why the two prefill columns disagree** (explains `--prefill auto`'s 8 192-token chunked prefill overlaps decode).
`tools/integrate_speed_0126_shin_blackmamba.py` is the **idempotent** script that produced both changes.
## Files
- **NEW:** `bench/results/2026-10-01-ctx-ladder/` — 8 raw JSON, `matrix.{json,md}`, `README.md`, `run.py`, `aggregate.py`, summary.
- **NEW:** `bench/results/2026-10-05-community-rtx-4070-ti-super/` — community report folder.
- **NEW:** `tools/integrate_speed_0126_shin_blackmamba.py`.
- **MODIFIED:** `bench/results/2026-09-29-speed-0126/matrix.json` (33 rows total), `bench/results/2026-09-29-speed-0126/README.md`.
- **MODIFIED:** `docs/COMMUNITY_BENCHMARKS.md` (this PR's entry added).
## Limitations
- Single consumer NVIDIA card. Two 16 GiB cards in this chassis would allow the layer split path (not exercised here).
- KV choices for IQ3_XXS at 128k/256k were forced to `q4_0` and `k8v4` because IQ3_XXS residents 47 GiB of experts and KV=`int8` would push past the 16 GiB VRAM envelope at those contexts.
- The host shares the GPU with no other workload during this test, but other containers were running CPU-bound, so the OS file-cache and RAM pressure were not zero. A truly isolated bench environment might give decode TPS another ±2-3 % improvement.
- Recall matrix is `3 of 3 FOUND` for both models at the single 32k context tested; deeper × wider recall sweeps would have required another `git reset --hard` cycle and were scoped out for this PR.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。