Pull requests / #336
#336 cuda: add experimental Tesla P4 8GB (sm_61) support
closed · draft · @CC-David-CC · 0 コメント · GitHub で見る
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
本文
Adds experimental Tesla P4 8GB support using CUDA 12.x, with a bounded FP32 prefill path, numerical tests, configurations, and benchmark results. Tested IQ3_S with Q8 KV, 8192 input tokens, up to 512 output, and 9216 allocated context. Tuned MTP generation: 22.09 tok/s counting, 20.24 coding, 14.94 writing. MTP increased total fresh-request time because prefill was slower. Full timings and limitations are in docs/DETAILS.md. Separate from AMD PR #323. Includes shared non-MTP serving and benchmark support, but excludes the RX 5500 XT GPU port. Preserves upstream README. Draft: measurements come from tested revision 9d393df. This isolated contribution branch has not yet been rebuilt and retested. Credit for Strata, expert caching, and MTP belongs to upstream.
関連リンク
インストール・モデル・リリースへの站内リンク。