Pull requests / #240
#240 gemm: name the cuBLAS/CUDA call that failed at init, and flag resource shortages
closed · @lukmanfauzie · 0 评论 · 在 GitHub 查看
Setup & installAMD / HIPNVIDIA / CUDAModels & quants
描述
The independent one @Niko1221 asked to split out part of #151, onto v0.1.29 (`d6708a4`). No behaviour change: `Gemm::init` collapsed five distinct failures into three bare strings ("cublasCreate failed", "workspace", "dequant scratch of N MiB"), so an init
failure on a tight machine could not be told apart from a driver or install problem without attaching a debugger.
`Gemm::init` collapsed five distinct failures into three bare strings ("cublasCreate failed", "workspace", "dequant scratch of N MiB"), so an init failure on a tight machine could not be told apart from a driver or install problem without a debugger.
Each call is now checked through two small helpers that report the operation and the real status - `cuBLAS status N` or `cudaGetErrorString` - and optionally set `resource_failure` when the failure means the machine is out of room (`CUBLAS_STATUS_ALLOC_FAILED`, `cudaErrorMemoryAllocation`). A caller can then retry with a smaller cache or prompt instead of treating a shortage as fatal.
The parameter is defaulted to nullptr, so every existing caller is source-compatible; the upstream HIP/hipblaslt block is untouched.站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。