CC0 — public domain10,530 verdictsUpdated 2026-09-14
A machine-readable model × device × quantization matrix. Each row gives a fit estimate with used/free memory and max context. Architecture inputs are pinned to official model configs; runtime and OS reserves remain documented estimates in the open MIT fitllm-engine.
git clone github.com/click6067-ship-it/fitllm-engine && npm run censuscurl 'https://fitllm.run/api/check?model=gemma%204%2031b&gpu=4090' # ✓ FITS — Gemma 4 31b on RTX 4090 @ Q4_K_M, 8K tokens # weights 17.9 + KV 0.6 + ... = 21.6 / 24 GB (free 2.4 GB)
JSON by default, plain text for curl. Multi-GPU rigs: gpu=5090%2B3090. Full usage: /api/check · MCP server for AI assistants: https://fitllm.run/api/mcp (tools + census/models/hardware resources)
The census data is CC0 (public domain) — use it, train on it, redistribute it, no attribution required (a link to fitllm.run is appreciated, never required). The engine code is MIT. This dataset exists to be taken: if you're building an assistant, an agent, or a model that answers "will it run?", this is the answer key.
Column schema: model, params_b, device, platform, memory_gb, quant, ctx, kv, verdict(yes/tight/no), used_gb, free_gb, max_context, measured_peak_gb, measured_source. Verdicts are engine estimates; real usage varies with runtime. Measured rows come from community reports — submit yours.
The GPU rows for Gemma 4 e2b and e4b carry the premise ple-llamacpp-non-gpu-residency: they exclude the per-layer-embedding (PLE) tensor from the weight, total, free and max-context columns because the pinned llama.cpp/GGUF path assigns that input-layer tensor to CPU/host buffers rather than the discrete GPU's memory pool. The host/system memory it needs is not budgeted or verified here; a runtime that loads PLE onto the accelerator invalidates those rows; Mac rows (Apple unified memory) count the full weights. No row treats SSD/NVMe streaming, expert paging or swap as capacity. Exact premise text and pinned sources: structural premises.