● FitLLM — fit receipt · 8-bit · 8K tokens ctx · KV F16 · engine 2.8.1

GLM-4.7-Flash on Mac 32GB

WON'T FIT ✗

weightsKV cacheoverheadreservetotalshort by
27.9 GB0.4 GB3.7 GB8.5 GB40.5 / 32 GB8.5 GB

Quantize to 4-bit and it fits (small quality cost).

Replay: npx fitllm-engine "GLM-4.7-Flash" --mac 32 · JSON · interactive

live fit badge — embed: ![fits](https://img.shields.io/endpoint?url=https%3A%2F%2Ffitllm.run%2Fapi%2Fbadge%3Fmodel%3DGLM-4.7-Flash%26ram%3D32%26quant%3D8%26ctx%3D8192%26kv%3D16)

Computed by the open fitllm-engine (MIT) from official config.json values — estimates; runtime varies. Ran it for real? Challenge this prediction — measured reports calibrate the engine for everyone.