Pull requests / #336

#336 cuda: add experimental Tesla P4 8GB (sm_61) support

closed · draft · @CC-David-CC · 0 コメント · GitHub で見る

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

本文

Adds experimental Tesla P4 8GB support using CUDA 12.x, with a bounded FP32 prefill path, numerical tests, configurations, and benchmark results.

Tested IQ3_S with Q8 KV, 8192 input tokens, up to 512 output, and 9216 allocated context. Tuned MTP generation: 22.09 tok/s counting, 20.24 coding, 14.94 writing. MTP increased total fresh-request time because prefill was slower. Full timings and limitations are in docs/DETAILS.md.

Separate from AMD PR #323. Includes shared non-MTP serving and benchmark support, but excludes the RX 5500 XT GPU port. Preserves upstream README.

Draft: measurements come from tested revision 9d393df. This isolated contribution branch has not yet been rebuilt and retested. Credit for Strata, expert caching, and MTP belongs to upstream.

関連リンク

インストール・モデル・リリースへの站内リンク。