Issues / #1337
#1337 setup --calibrate loses the measurements it already has when a later engine start fails
open · @aly8246 · 0 评论 · 在 GitHub 查看
BenchmarksSetup & installNVIDIA / CUDA
描述
`tools/calibrate.py` measures every candidate through one engine and then restarts it per value in step 4 for the CPU worker count. If one of those starts fails, `start_engine` raises, `measure()` never returns, and `setup --calibrate` reports a failed run while the config keeps its defaults. The winners the earlier steps already measured and printed are lost with it. That is what happened to the run in #1245. On an RTX 5090 Laptop the MTP draft head (212.9 MiB, CJK vocab) no longer fitted during the 23-worker restart, with 213-274 MiB free, so the run died there. The sweep winners printed before it (`PCIe 0.55`, floor `0.70`) were applied by hand. It costs more than one setting when the sweep has found something worth keeping. On this machine (Core Ultra 9 290HX Plus, RTX 5090 Laptop, PCIe 5.0 x16) the PCIe share at 1.0 is worth 39% over the engine default of 0.55, 51.1 to 71.3 tok/s interleaved per request, and #1332 adds 1.0 to the sweep so that `--calibrate` can find it. A run that aborts before the confirm step, or before the worker step, gives that away. The same shape can fire the other way: a candidate can fail to start for a reason that has nothing to do with the setting, for instance VRAM taken by something else at that moment, and the run then loses the candidates it had already measured. Suggestion: treat a candidate whose engine does not start as a losing candidate rather than a failed run. Step 4 already has the loop for it, and it could record the value as a loss and carry on to `pick()`, so the settings from steps 1-3 survive. If every restart fails the run is a failure as it is now. `engine_error()` already returns the failed start's own line, which could be named in the report so the reason is not lost. I am not attached to the shape of the fix. The point is that a run should not lose measurements it already has because a later start failed. Adjacent, not the same thing: #1197 is about the calibration loading the model more than once and leaving the server running.
站内延伸阅读
链到安装、模型与版本说明,便于 SEO/GEO,非官方 issue 正文。