Pull requests / #1592

#1592 bench: RTX 5090 1M-agent follow-up — 0.1.40.2 vs 0.1.41, int8 vs k8v4 (interleaved protocol)

open · @gravitomagnetic · 0 コメント · GitHub で見る

BenchmarksServer & APINVIDIA / CUDAModels & quantsDocumentation

本文

Follow-up to our merged 5090 report (`bench/results/2026-10-02-community-rtx5090-1m-agent/`, #466/#1134 lineage), same box and config.

Two A/Bs run today with a protocol change worth noting: our first blocked runs (A A A then B B B) showed the last leg of each block decoding ~30% faster than its own first two — warm-up drift, not a version effect. We re-ran everything as matched alternating pairs (A B A B A B); the drift vanished (cold-prefill spread 0.1 s, greedy decode identical run to run). The README explains it and we'd suggest interleaving for small-n A/Bs on expert-streaming rigs.

Results:
- **0.1.41: no regression** on our regime (prefill 58.3→58.2 s, decode tied, all gates clean). We keep it as the daily engine.
- **k8v4: correct but not worth it here.** The #1264 fix verifies on sm_120 at 200K context (3/3 needles, zero repetition, three for three). The 2.8 GiB pinned-RAM saving buys +146 expert slots, but at ~90% cache hit those slots add nothing measurable to decode, while the streaming path costs a consistent ~3% on 200K prefill. int8 stays.

Raw data: six version JSONs (one per leg); KV per-leg numbers from the engine's serve log (extract attached — the client JSONs share a timestamp and the last leg of each condition overwrote the earlier ones).

Standing offer: happy to re-run the exact interleaved legs on any candidate build — KV, cache, or prefill changes.

関連リンク

インストール・モデル・リリースへの站内リンク。