Pull requests / #336

#336 cuda: add experimental Tesla P4 8GB (sm_61) support

closed · draft · @CC-David-CC · 0 comentários · No GitHub

BenchmarksAMD / HIPNVIDIA / CUDAModels & quantsDocumentation

Descrição

Adds experimental Tesla P4 8GB support using CUDA 12.x, with a bounded FP32 prefill path, numerical tests, configurations, and benchmark results.

Tested IQ3_S with Q8 KV, 8192 input tokens, up to 512 output, and 9216 allocated context. Tuned MTP generation: 22.09 tok/s counting, 20.24 coding, 14.94 writing. MTP increased total fresh-request time because prefill was slower. Full timings and limitations are in docs/DETAILS.md.

Separate from AMD PR #323. Includes shared non-MTP serving and benchmark support, but excludes the RX 5500 XT GPU port. Preserves upstream README.

Draft: measurements come from tested revision 9d393df. This isolated contribution branch has not yet been rebuilt and retested. Credit for Strata, expert caching, and MTP belongs to upstream.

No site

Links install, modelos, releases.