Pull requests / #336
#336 cuda: add experimental Tesla P4 8GB (sm_61) support
closed · draft · @CC-David-CC · 0 commentaires · Sur GitHub
BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation
Description
Adds experimental Tesla P4 8GB support using CUDA 12.x, with a bounded FP32 prefill path, numerical tests, configurations, and benchmark results. Tested IQ3_S with Q8 KV, 8192 input tokens, up to 512 output, and 9216 allocated context. Tuned MTP generation: 22.09 tok/s counting, 20.24 coding, 14.94 writing. MTP increased total fresh-request time because prefill was slower. Full timings and limitations are in docs/DETAILS.md. Separate from AMD PR #323. Includes shared non-MTP serving and benchmark support, but excludes the RX 5500 XT GPU port. Preserves upstream README. Draft: measurements come from tested revision 9d393df. This isolated contribution branch has not yet been rebuilt and retested. Credit for Strata, expert caching, and MTP belongs to upstream.
Sur le site
Liens install, modèles, releases.