贡献 / #1494
#1494 bench: heterogeneous RTX 5070 Ti + 5060 Ti, experimental PP32 KV-grow and adaptive DMA
open · draft · @k93k2J-glitch · 0 评论 · 去 GitHub 看
BenchmarksMulti-GPUAMD / HIPNVIDIA / CUDAModels & quantsSecurityWindows
说明
## Summary Results-only community report for a heterogeneous RTX 5070 Ti 16GB + RTX 5060 Ti 16GB pair on Windows: experimental PP32/16, native 262K int8 KV capacity, KV-grow/pipeline2, and optional pinned adaptive DMA staging. Every changed path is under `bench/results/2026-10-08-community-5070ti-5060ti-kvg-dma/`. This PR does not apply engine changes to upstream; the prototype patches are reproduction artifacts in that folder. ## Recorded evidence - Custom fork based on source tag v0.1.40.1, native engine string `0.1.40`: 10 interleaved paired rounds at actual 65,536 / 131,072 / 261,000 input tokens and actual 1,024 generated tokens. - 60 steady measurements plus 60 first-request diagnostics; all selected outputs passed four functional cases, with prefix reuse disabled and native token counters checked. - Full-ten median/ranges and independent eight-sample trimmed means are both retained, including slow samples and one disclosed validator-timeout partial-visit retry. - Separate v0.1.40.3 prototype validation: eight functional outputs and per-device long-to-short KV/expert observations. This is not a repeated v0.1.40.3 speed comparison. - Synthetic frozen messages, portable opt-in runner, sanitized numeric CSV/JSON, and clean-tag prototype patches are included. ## Main limits The measured engine is a custom prototype, not stock upstream. The workload completes a Python function then generates comments to maintain the output cap; it does not establish prose speed, long-context recall, perplexity or general coding quality. At 128K, the full-ten median decode difference is about +0.93%, whereas the trimmed-mean difference is +7.49%; slow samples materially affect the latter. There is no general DMA speedup or causal paging claim. Clocks/background load were not fully controlled, and full model/expert-profile hashes were not frozen. Public stripped-patch builds and inference reruns remain outstanding; the local full prototype previously built and ran. ## Publication preparation Clean official-tag patch application and core-file/DONE-parser equivalence to the previously tested source were checked offline. Python syntax, synthetic fixture grammar/message hashes, default no-run behavior and an explicit outbound file allowlist were checked. No model request, GPU test or service restart was performed during publication preparation. Machine usernames, personal email, host/network identifiers, absolute personal paths, credential values/locations, private prompt content, raw service/Codex logs, Windows launchers, auth bypass and local Git history are excluded. Git commit metadata uses the submitting account's GitHub noreply identity. This is a **draft** for maintainer review of the report and reproduction packaging. Related work: #1223 (VMM/peer ownership), #1231 (no-loan safety), #765 (budgeting), #807 (batched-upload design). The layer-split feature discussion can be coordinated separately rather than proposing the integration branch wholesale. Feature discussion: #1495.
本站相关内容
相关页面的快捷入口。