Pull requests / #995

#995 Community benchmark: RTX 4070 Ti SUPER 16 GiB (Shin-BlackMamba, upstream 6f32ec07)

closed · @lfontanez · 0 Kommentare · Auf GitHub

BenchmarksSetup & installMulti-GPUNVIDIA / CUDAModels & quantsDocumentation

Beschreibung

## Adds a community benchmark entry per `docs/COMMUNITY_BENCHMARKS.md`

**Hardware:** NVIDIA GeForce RTX 4070 Ti SUPER, 16 GiB VRAM, sm_89, 94 GiB system RAM, Samsung 990 PRO 2 TB NVMe, Ubuntu 24.04.5 LTS, NVIDIA driver 595.91.07 (CUDA 13.2), nvidia-container-toolkit 1.20.1, Docker 29.8.1.

**Engine:** upstream-main commit `6f32ec07`, locally compiled `sm_89` fat-binary (`CUDA_ARCHITECTURES=89`, `BUILD_VISION=0`). Base image `nvidia/cuda:13.0.0-devel-ubuntu24.04`. Single consumer card; layer-split deliberately not exercised on the 16 GiB ceiling.

## Eight configurations measured

| Model    | CONTEXT  | KV    | Cold prefill | Decode avg | Peak    | Engine TPS | SSE prefill | Draft | TTFT   |
|----------|---------:|------:|-------------:|-----------:|--------:|-----------:|------------:|------:|-------:|
| IQ2_XS   |    32768 | int8  |   **2 695**  |     95.62  |  136.34 |     93.47  |      141.3  | 75.8% | 183 ms |
| IQ2_XS   |    65536 | int8  |   **2 678**  |     91.80  |  123.20 |     89.90  |      140.9  | 74.5% | 271 ms |
| IQ2_XS   |   131072 | int8  |   **2 670**  |     87.95  |  117.52 |     86.67  |      132.6  | 76.8% | 422 ms |
| IQ2_XS   |   262144 | int8  |   **2 484**  |     79.63  |  107.36 |     78.50  |      127.3  | 69.6% | 726 ms |
| IQ3_XXS  |    32768 | int8  |   **2 747**  |     73.32  |   97.96 |     71.73  |      116.5  | 69.8% | 217 ms |
| IQ3_XXS  |    65536 | int8  |   **2 545**  |     69.16  |   98.66 |     68.36  |      107.0  | 70.3% | 295 ms |
| IQ3_XXS  |   131072 | q4_0  |   **2 478**  |     65.51  |   84.42 |     64.10  |      105.2  | 67.7% | 461 ms |
| IQ3_XXS  |   262144 | k8v4  |   **2 404**  |     53.70  |   71.09 |     52.97  |       91.9  | 66.6% | 768 ms |

Each row = 1 cold-prefill probe (`max_tokens=1` non-stream after `POST /unload`) + 3 SSE stream calls (`max_tokens=192`, temperature 0, `reasoning_effort="minimal"`).

## Recall (`tools/needle_bench.py --lengths 32k --depths 10,50,90`)

| model    | depth | found | prompt tok | seconds |
|----------|------:|:-----:|----------:|--------:|
| IQ2_XS   |  10 % | ✅    |    32 171 |   13    |
| IQ2_XS   |  50 % | ✅    |    32 171 |   12    |
| IQ2_XS   |  90 % | ✅    |    32 172 |    6    |
| IQ3_XXS  |  10 % | ✅    |    32 171 |   12    |
| IQ3_XXS  |  50 % | ✅    |    32 171 |   12    |
| IQ3_XXS  |  90 % | ✅    |    32 172 |    6    |

**3/3 FOUND** for both IQ2_XS and IQ3_XXS at 32k context.

## Cross-table comparison with the 0.1.26 headline (RTX 5070 12 GB)

| model    | 32k cold prefill (this host) | 32k cold prefill (RTX 5070) | delta |
|----------|------------------------------:|----------------------------:|------:|
| IQ2_XS   |              **2 695 t/s**     |              **2 092 t/s**  | **+29 %** |
| IQ3_XXS  |              **2 747 t/s**     |              **1 745 t/s**  | **+57 %** |

The 4070 Ti SUPER is comfortably ahead of the 5070 on cold prefill; the host's faster CPU pool + Gen4 host-to-device pipeline let the engine keep the KV-streaming reserve in RAM and run the experts through it without choking decode.

## Canonical `bench/results/2026-09-29-speed-0126/` extension

The upstream `matrix.json` was patched (no upstream row contents changed):

- Added a `kv` schema field. The 25 baseline rows are `int8` per `setup.py` defaults.
- Added a `prefill_cold_tok_s` column (true end-of-prompt throughput). `null` on the 25 baselines; populated for the 8 new rows added here.
- The README grew two sections: **KV per cell and True-Prefill (host extension)** (the 8 new rows) and **Why the two prefill columns disagree** (explains `--prefill auto`'s 8 192-token chunked prefill overlaps decode).

`tools/integrate_speed_0126_shin_blackmamba.py` is the **idempotent** script that produced both changes.

## Files

- **NEW:** `bench/results/2026-10-01-ctx-ladder/` — 8 raw JSON, `matrix.{json,md}`, `README.md`, `run.py`, `aggregate.py`, summary.
- **NEW:** `bench/results/2026-10-05-community-rtx-4070-ti-super/` — community report folder.
- **NEW:** `tools/integrate_speed_0126_shin_blackmamba.py`.
- **MODIFIED:** `bench/results/2026-09-29-speed-0126/matrix.json` (33 rows total), `bench/results/2026-09-29-speed-0126/README.md`.
- **MODIFIED:** `docs/COMMUNITY_BENCHMARKS.md` (this PR's entry added).

## Limitations

- Single consumer NVIDIA card. Two 16 GiB cards in this chassis would allow the layer split path (not exercised here).
- KV choices for IQ3_XXS at 128k/256k were forced to `q4_0` and `k8v4` because IQ3_XXS residents 47 GiB of experts and KV=`int8` would push past the 16 GiB VRAM envelope at those contexts.
- The host shares the GPU with no other workload during this test, but other containers were running CPU-bound, so the OS file-cache and RAM pressure were not zero. A truly isolated bench environment might give decode TPS another ±2-3 % improvement.
- Recall matrix is `3 of 3 FOUND` for both models at the single 32k context tested; deeper × wider recall sweeps would have required another `git reset --hard` cycle and were scoped out for this PR.

Mehr auf der Site

Links zu Install, Modellen, Releases.