Issues / #973

#973 Golden drift: all 46 test_enter_for_every_question subtests fail on main (tools/test_setup_golden.py)

closed · @AmjedMVP · 2 comments · View on GitHub

Setup & installServer & APIModels & quantsLinux

Description

## Summary

On a fresh clone of `main` (6f32ec0), `tools/test_setup_golden.py` fails **46 of 46** subtest cases of `test_enter_for_every_question`: the configs captured from `setup.py` drift from `tools/test_setup_golden.json`.

## Environment

- Ubuntu 26.04 LTS (WSL2), x86_64, Python 3.14.4
- Fresh `git clone --depth 1` of main @ `6f32ec0`, `.venv` + `pip install -r requirements.txt`
- GPU-free paths only (these tests don't touch the engine)

## Reproduce

```sh
python3 -m unittest tools/test_setup_golden.py -v
```

Result: `Ran 4 tests ... FAILED (failures=46)` — all in `test_enter_for_every_question`.

Representative diff (from a `[64GB-1x32GB qwen recommended]` subtest):

```
AssertionError: {'exe': '<T>/engine/<EXE>', 'args': ['--pack', '<T>/data/packs/iq3_xxs... (645 chars) ...True} != {'args': ['--pack', '<T>/data/packs/iq3_xxs... (646 chars) ...zer'}
```

Affected subtest labels seen in the log: `[64GB-1x32GB qwen recommended|Q2_0|IQ3_XXS|IQ3_S|unsloth UD-Q4_K_XL]`, `[32GB-2x24GB qwen recommended|Q2_0|IQ3_XXS]` and the rest of the matrix (46 total).

Full run log available on request (all 30 `tools/` + `serve/` suites: 526 tests, 476 passed, 49 failed, 1 error on this commit — golden is the bulk of the failures).

## Interpretation

Either the golden JSON is stale after a recent arg-generation change, or a regression landed in the arg builder. Since every matrix cell fails, a systematic change (flag rename/reorder/addition) rather than per-model drift looks likely. Happy to run with a golden refresh locally to confirm which side moved if that helps.

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.