Issues / #941
#941 UPD: Successfully run on RTX 3090 24gb + Z590E + 64gb ddr4 3200 2ch
closed · @gamalsaad7234 · 0 Kommentare · Auf GitHub
BenchmarksServer & APINVIDIA / CUDAModels & quants
Beschreibung
Errors and debug log: Latest version from github. And after re-run: [strata] still starting (24 s) - please wait ... [strata] experts loaded: 46.84 GiB at 2.91 GiB/s (29 s so far) [strata] filling the GPU's expert cache (7153 experts, 13.57 GiB of VRAM) ... [strata] almost ready ... [strata] the engine is running again [strata] reading the prompt: 346 of 351 tokens, 2 s so far [strata] the engine stopped unexpectedly (exit code 3221226505). The engine exited (code 3221226505). Its last log line: strata serve: 305 MiB of VRAM free with everything loaded - if that does not explain it, please report it at github.com/Niko1221/Strata/issues with the log. The next request starts the engine again. Its log: D:\Strata-main\strata-iq3_s.log [strata] done: 0 tokens in 9 s (0.0 tok/s) (error, cancel=False) Sadly! Any ways to fix that? UPD: Deleting json config hangs the strata and explorer.exe UPD: After non normal closing or crashing is something holds venv (Needs to be fixed) UPD: No calibration, no speed projection --> Succesfully run on ctx 65k, generation speed is 60-77.9 tok/s and prompt 2k processed in 2 seconds. My logic test on LLMs passed 👍 Design and code test passed 👍 Vision is working 👍 So question is, for bigger context i need more RAM?
Mehr auf der Site
Links zu Install, Modellen, Releases.