Issues / #941

#941 UPD: Successfully run on RTX 3090 24gb + Z590E + 64gb ddr4 3200 2ch

closed · @gamalsaad7234 · 0 comments · View on GitHub

BenchmarksServer & APINVIDIA / CUDAModels & quants

Description

Errors and debug log:
Latest version from github.
And after re-run:

[strata] still starting (24 s) - please wait ...
[strata] experts loaded: 46.84 GiB at 2.91 GiB/s (29 s so far)
[strata] filling the GPU's expert cache (7153 experts, 13.57 GiB of VRAM) ...
[strata] almost ready ...
[strata] the engine is running again
[strata] reading the prompt: 346 of 351 tokens, 2 s so far
[strata] the engine stopped unexpectedly (exit code 3221226505). The engine exited (code 3221226505). Its last log line: strata serve: 305 MiB of VRAM free with everything loaded - if that does not explain it, please report it at github.com/Niko1221/Strata/issues with the log. The next request starts the engine again. Its log: D:\Strata-main\strata-iq3_s.log
[strata] done: 0 tokens in 9 s (0.0 tok/s) (error, cancel=False)

Sadly! Any ways to fix that?

UPD:
Deleting json config hangs the strata and explorer.exe

UPD:
After non normal closing or crashing is something holds venv (Needs to be fixed)

UPD: 
No calibration, no speed projection -->
Succesfully run on ctx 65k, generation speed is 60-77.9 tok/s and prompt 2k processed in 2 seconds.
My logic test on LLMs passed 👍
Design and code test passed 👍 
Vision is working 👍

So question is, for bigger context i need more RAM?

Related on strata.com

Editorial links to help you install, pick models, or read release notes — not part of the upstream thread.