● FitLLM — fit receipt · 8-bit · 8K tokens ctx · KV F16 · engine 2.8.1
FITS ✓
| weights | KV cache | overhead | reserve | total | free |
|---|---|---|---|---|---|
| 27.9 GB | 0.4 GB | 3.7 GB | 8.0 GB | 40.0 / 64 GB | 24.0 GB |
Max context at this quant: ~135K tokens
Replay: npx fitllm-engine "GLM-4.7-Flash" --mac 64 · JSON · interactive
— embed:

Computed by the open fitllm-engine (MIT) from official config.json values — estimates; runtime varies. Ran it for real? Challenge this prediction — measured reports calibrate the engine for everyone.