Pull requests / #240

#240 gemm: name the cuBLAS/CUDA call that failed at init, and flag resource shortages

closed · @lukmanfauzie · 0 评论 · 在 GitHub 查看

Setup & installAMD / HIPNVIDIA / CUDAModels & quants

描述

The independent one @Niko1221 asked to split out part of #151, onto v0.1.29 (`d6708a4`).  No behaviour change: `Gemm::init` collapsed five distinct failures into three bare strings ("cublasCreate failed", "workspace", "dequant scratch of N MiB"), so an init
failure on a tight machine could not be told apart from a driver or install problem without attaching a debugger.

`Gemm::init` collapsed five distinct failures into three bare strings ("cublasCreate failed", "workspace", "dequant scratch of N MiB"), so an init failure on a tight machine could not be told apart from a driver or install problem without a debugger.

Each call is now checked through two small helpers that report the operation and the real status - `cuBLAS status N` or `cudaGetErrorString` - and optionally set `resource_failure` when the failure means the machine is out of room (`CUBLAS_STATUS_ALLOC_FAILED`, `cudaErrorMemoryAllocation`).  A caller can then retry with a smaller cache or prompt instead of treating a shortage as fatal.

The parameter is defaulted to nullptr, so every existing caller is source-compatible; the upstream HIP/hipblaslt block is untouched.

站内延伸阅读

链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。