● FitLLM — fit receipt · Q4_K_M · 8K tokens ctx · KV F16 · engine 2.8.1
FITS ✓
| weights | KV cache | linear state | overhead | reserve | total | free |
|---|---|---|---|---|---|---|
| 15.8 GB | 0.5 GB | 0.1 GB | 2.2 GB | 2.0 GB | 20.7 / 24 GB | 3.3 GB |
Max context at this quant: ~29K tokens
Replay: npx fitllm-engine "Qwen 3.8 27B" --gpu "RTX 3090" · JSON · interactive
— embed:

Computed by the open fitllm-engine (MIT) from official config.json values — estimates; runtime varies. Ran it for real? Challenge this prediction — measured reports calibrate the engine for everyone.