DeepSeek R1
DeepSeek R1
Full reasoning model, not a distilled checkpoint; 37B parameters are activated per token but all 671B weights must be stored.
- Parameters
- 671B
- Context
- 131K
Models / Open weights
Creator-published model specifications from Hugging Face, paired with transparent memory estimates for common accelerator capacities.
Creator model cards
10 of 10 models
DeepSeek R1
Full reasoning model, not a distilled checkpoint; 37B parameters are activated per token but all 671B weights must be stored.
DeepSeek V3
Mixture-of-experts model with 37B parameters activated per token; local weight memory reflects 671B total parameters.
Llama 3.2
Image-plus-text is officially supported only in English; custom license restrictions apply.
Llama 3.2
Custom commercial license with gated access and acceptable-use policy; actual architecture count is approximately 1.23B.
Llama 3.3
Custom commercial license, gated access, and acceptable-use policy apply.
Mistral Small
Supports image understanding, not image generation; requires Mistral-specific serving and tokenizer settings.
Qwen3
Repository name uses 0.6B; Hugging Face safetensors metadata reports approximately 752M parameters.
Qwen3
Mixture-of-experts model with 3.3B parameters activated per token; weight memory reflects 30.5B total parameters. Validated to 131,072 tokens with YaRN.
Estimated local deployment
Estimated weights plus 20% loading overhead. Runtime, KV cache, context length, and software can require substantially more memory.
Method
The estimate uses parameters × planning bits ÷ 8 × 1.20: 5 planning bits for Q4 and 9 for Q8 to allow for scales, metadata, and mixed-precision tensors. “Comfortable” reserves at least 20% of device memory. “Tight” means estimated weights fit but leave less headroom.
This is a capacity screen—not a speed benchmark or guarantee that a particular runtime, context length, multimodal projector, or GPU configuration will work.
| Model | Parameters | Q4 estimate | Q8 estimate | 8 GB | 12 GB | 16 GB | 24 GB | 32 GB | 48 GB | 80 GB | 141 GB | 192 GB | 256 GB | 512 GB |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen3 0.6B | 0.752B | 0.6 GB | 1.0 GB | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Qwen3 30B-A3B | 30.5B | 22.9 GB | 41.2 GB | No | No | No | Tight | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Gemma 3 1B IT | 1B | 0.8 GB | 1.3 GB | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Gemma 3 4B IT | 4B | 3.0 GB | 5.4 GB | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Mistral Small 3.1 24B | 24B | 18.0 GB | 32.4 GB | No | No | No | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Llama 3.2 1B Instruct | 1.23B | 0.9 GB | 1.7 GB | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Llama 3.2 11B Vision | 10.6B | 7.9 GB | 14.3 GB | Tight | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| Llama 3.3 70B Instruct | 70B | 52.5 GB | 94.5 GB | No | No | No | No | No | No | Comfortable | Comfortable | Comfortable | Comfortable | Comfortable |
| DeepSeek V3 | 671B | 503.3 GB | 905.9 GB | No | No | No | No | No | No | No | No | No | No | Tight |
| DeepSeek R1 | 671B | 503.3 GB | 905.9 GB | No | No | No | No | No | No | No | No | No | No | Tight |
What comes next
This release answers whether model weights are likely to fit. A useful throughput comparison must additionally pin the model revision, quantization, prompt and generation lengths, batch size, runtime, kernels, driver, and exact hardware configuration.
The next hardware dataset will add source-grounded tokens-per-second measurements without mixing incomparable benchmark setups.