Pull requests / #1494

#1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMA

open · draft · @k93k2J-glitch · 0 commentaires · Sur GitHub

BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindows

Description

## Summary

Results-only community report for a heterogeneous RTX 5070 Ti 16GB + RTX 5060 Ti 16GB pair on Windows: experimental PP32/16, native 262K int8 KV capacity, KV-grow/pipeline2, and optional pinned adaptive DMA staging.

Every changed path is under `bench/results/2026-10-08-community-5070ti-5060ti-kvg-dma/`. This PR does not apply engine changes to upstream; the prototype patches are reproduction artifacts in that folder.

## Recorded evidence

- Custom fork based on source tag v0.1.40.1, native engine string `0.1.40`: 10 interleaved paired rounds at actual 65,536 / 131,072 / 261,000 input tokens and actual 1,024 generated tokens.
- 60 steady measurements plus 60 first-request diagnostics; all selected outputs passed four functional cases, with prefix reuse disabled and native token counters checked.
- Full-ten median/ranges and independent eight-sample trimmed means are both retained, including slow samples and one disclosed validator-timeout partial-visit retry.
- Separate v0.1.40.3 prototype validation: eight functional outputs and per-device long-to-short KV/expert observations. This is not a repeated v0.1.40.3 speed comparison.
- Synthetic frozen messages, portable opt-in runner, sanitized numeric CSV/JSON, and clean-tag prototype patches are included.

## Main limits

The measured engine is a custom prototype, not stock upstream. The workload completes a Python function then generates comments to maintain the output cap; it does not establish prose speed, long-context recall, perplexity or general coding quality. At 128K, the full-ten median decode difference is about +0.93%, whereas the trimmed-mean difference is +7.49%; slow samples materially affect the latter. There is no general DMA speedup or causal paging claim.

Clocks/background load were not fully controlled, and full model/expert-profile hashes were not frozen. Public stripped-patch builds and inference reruns remain outstanding; the local full prototype previously built and ran.

## Publication preparation

Clean official-tag patch application and core-file/DONE-parser equivalence to the previously tested source were checked offline. Python syntax, synthetic fixture grammar/message hashes, default no-run behavior and an explicit outbound file allowlist were checked. No model request, GPU test or service restart was performed during publication preparation.

Machine usernames, personal email, host/network identifiers, absolute personal paths, credential values/locations, private prompt content, raw service/Codex logs, Windows launchers, auth bypass and local Git history are excluded. Git commit metadata uses the submitting account's GitHub noreply identity.

This is a **draft** for maintainer review of the report and reproduction packaging. Related work: #1223 (VMM/peer ownership), #1231 (no-loan safety), #765 (budgeting), #807 (batched-upload design). The layer-split feature discussion can be coordinated separately rather than proposing the integration branch wholesale.

Feature discussion: #1495.

Sur le site

Liens install, modèles, releases.