● FitLLM — fit receipt · Q4_K_M · 8K tokens ctx · KV F16 · engine 2.8.1
WON'T FIT ✗
| weights | KV cache | overhead | reserve | total | short by |
|---|---|---|---|---|---|
| 17.1 GB | 0.4 GB | 2.4 GB | 2.0 GB | 21.9 / 12 GB | 9.9 GB |
→ RTX 4090 (24GB) would fit.
Replay: npx fitllm-engine "GLM-4.7-Flash" --gpu "RTX 4070" · JSON · interactive
— embed:

Computed by the open fitllm-engine (MIT) from official config.json values — estimates; runtime varies. Ran it for real? Challenge this prediction — measured reports calibrate the engine for everyone.