● FitLLM — fit receipt · Q4_K_M · 8K tokens ctx · KV F16 · engine 2.8.1

gpt-oss-120b on RTX 5080

WON'T FIT ✗

weightsKV cacheoverheadreservetotalshort by
66.7 GB0.3 GB8.3 GB2.0 GB77.2 / 16 GB61.2 GB

A100 80GB would fit.

Replay: npx fitllm-engine "gpt-oss-120b" --gpu "RTX 5080" · JSON · interactive

live fit badge — embed: ![fits](https://img.shields.io/endpoint?url=https%3A%2F%2Ffitllm.run%2Fapi%2Fbadge%3Fmodel%3Dgpt-oss-120b%26gpu%3DRTX%2B5080%26quant%3DQ4_K_M%26ctx%3D8192%26kv%3D16)

Computed by the open fitllm-engine (MIT) from official config.json values — estimates; runtime varies. Ran it for real? Challenge this prediction — measured reports calibrate the engine for everyone.