Pull requests / #1491
#1491 bench: completed Flash Next tool-use results with charts and examples
open · @CC-David-CC · 0 commentaires · Sur GitHub
BenchmarksNVIDIA / CUDAModels & quants
Description
## Summary Replaces closed #1456 with the completed twelve-variant evaluation and visual results. All twelve requested model variants now have completed Tool-Eval-Bench results. This update adds Gyro-M and the three AP variants, completes the matrix, and includes two figures with short, paraphrased test examples. Benchmark created by **[SeraphimSerapis](https://github.com/SeraphimSerapis)**: **[Tool-Eval-Bench](https://github.com/SeraphimSerapis/tool-eval-bench)**. Scenarios and scoring are unchanged.   ## Added results | Variant | Short /100 | Standard /100 | Scored cases | |---|---:|---:|---:| | Gyro-M | 93 | 86 | 68 | | AP-Q4_K_XL | 97 | 85 | 69 | | AP-IQ3_XXS | 90 | 89 | 69 | | AP-IQ2_S | 93 | 88 | 69 | ## What changed Completed report, traces, portable deployment settings, resource samples, conversion records and reproducible PNG/SVG figures under `bench/results/tool-eval-20261007`. No engine or benchmark scoring changes in this PR. ## Interpretation - Six ISTA/Swift runs share the mainline build and RTX PRO 6000. AP uses the same GPU and settings, with the existing Q6 build option enabled for its embedding. - Gyro uses patched agentionai Strata rc1; TC-45 is excluded because its endpoint does not enforce required tool use. These are not results from the newer Gyro merge branch. - Q8 uses a separate three-P4 engine; its short score is extracted from the first 15 standard cases. - Single trials; Swift/AP are modified variants. Structured-output API restrictions, stale mock timestamps and background transfers are documented. Examples are paraphrases, not model quotes; all tool actions are mocks. This benchmark report is independent of the Responses/cache integration in #1489 and the Responses implementation in #1446.
Sur le site
Liens install, modèles, releases.