Which local LLMs can the Apple M5 run?
Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Specs
- Memory options
- 16 GB · 24 GB · 32 GB
- Memory bandwidth
- 153.6 GB/s
- FP16 compute
- 10.0 TFLOPS
- Launch year
- 2025
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- 16 GBRuns great2 modelsRuns well8 modelsRuns slowly3 modelsWon't run16 models3 more models run at a lower quant.
- 24 GBRuns great2 modelsRuns well10 modelsRuns slowly7 modelsWon't run10 models3 more models run at a lower quant.
- 32 GBRuns great2 modelsRuns well12 modelsRuns slowly9 modelsWon't run6 models1 more model runs at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the Apple M5
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 70.2 tok/s56.1–84.2calibrated estimate ±20% | Q4_K_M | 1.9 / 22.4 GB | Unified memory | up to 32k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 50.8 tok/s40.6–60.9calibrated estimate ±20% | Q4_K_M | 2.4 / 22.4 GB | Unified memory | up to 16k | — |
| gpt-oss-20b | Runs well | est. 20.1 tok/s16.1–24.2calibrated estimate ±20% | MXFP4 | 12.9 / 22.4 GB | Unified memory | up to 64k |
|
| Qwen3 30B-A3B (2507) | Runs well | est. 16.4 tok/s13.1–19.7calibrated estimate ±20% | Q4_K_M | 19.9 / 22.4 GB | Unified memory | up to 16k |
|
| Gemma 4 26B-A4B | Runs well | est. 15.9 tok/s12.8–19.1calibrated estimate ±20% | Q4_K_M | 17.9 / 22.4 GB | Unified memory | up to 64k |
|
| DeepSeek R1 Distill Llama 8B | Runs well | est. 15.4 tok/s12.3–18.5calibrated estimate ±20% | Q4_K_M | 6.6 / 22.4 GB | Unified memory | up to 16k |
|
| Kanana 1.5 8B | Runs well | est. 15.4 tok/s12.3–18.5calibrated estimate ±20% | Q4_K_M | 6.6 / 22.4 GB | Unified memory | up to 32k |
|
| Llama 3.1 8B | Runs well | est. 15.4 tok/s12.3–18.5calibrated estimate ±20% | Q4_K_M | 6.6 / 22.4 GB | Unified memory | up to 32k |
|
| Kanana 1.5 15.7B-A3B | Runs well | est. 15.0 tok/s12.0–18.0calibrated estimate ±20% | Q4_K_M | 12.1 / 22.4 GB | Unified memory | up to 16k | — |
| Qwen3.5 9B | Runs well | est. 15.0 tok/s12.0–18.0calibrated estimate ±20% | Q4_K_M | 6.7 / 22.4 GB | Unified memory | up to 128k | — |
| Qwen3 8B | Runs well | est. 14.8 tok/s11.8–17.7calibrated estimate ±20% | Q4_K_M | 6.8 / 22.4 GB | Unified memory | up to 32k |
|
| Gemma 4 12B | Runs well | est. 12.0 tok/s9.6–14.4calibrated estimate ±20% | Q4_K_M | 8.2 / 22.4 GB | Unified memory | up to 64k | — |
| Gemma 3 12B | Runs well | est. 11.8 tok/s9.4–14.1calibrated estimate ±20% | Q4_K_M | 8.4 / 22.4 GB | Unified memory | up to 32k | — |
| Qwen3 14B | Runs well | est. 8.9 tok/s7.1–10.7calibrated estimate ±20% | Q4_K_M | 10.9 / 22.4 GB | Unified memory | up to 8k | — |
| Qwen3.5 35B-A3B | Runs slowly | est. 22.4 tok/s15.7–29.1theoretical estimate ±30% | Q4_K_M | 22.8 GB RAM | CPU only | up to 256k |
|
| HyperCLOVA X SEED Think 14B | Runs slowly | est. 9.0 tok/s6.3–11.8theoretical estimate ±30% | Q4_K_M | 10.8 / 22.4 GB | Unified memory | up to 64k |
|
| Phi-4 | Runs slowly | est. 8.6 tok/s6.9–10.3calibrated estimate ±20% | Q4_K_M | 11.3 / 22.4 GB | Unified memory | up to 16k |
|
| Mistral Small 3.2 24B | Runs slowly | est. 5.9 tok/s4.7–7.1calibrated estimate ±20% | Q4_K_M | 16.2 / 22.4 GB | Unified memory | up to 32k |
|
| Gemma 3 27B | Runs slowly | est. 5.4 tok/s4.3–6.4calibrated estimate ±20% | Q4_K_M | 17.8 / 22.4 GB | Unified memory | up to 32k |
|
| Qwen3.5 27B | Runs slowly | est. 5.2 tok/s4.2–6.3calibrated estimate ±20% | Q4_K_M | 18.2 / 22.4 GB | Unified memory | up to 32k | — |
| EXAONE 4.0 32B | Runs slowly | est. 4.6 tok/s3.7–5.6calibrated estimate ±20% | Q4_K_M | 20.5 / 22.4 GB | Unified memory | up to 32k |
|
| EXAONE 4.5 33B | Runs slowly | est. 4.5 tok/s3.6–5.4calibrated estimate ±20% | Q4_K_M | 21.2 / 22.4 GB | Unified memory | up to 16k |
|
| Qwen3 32B | Runs slowly | est. 2.4 tok/s1.7–3.2theoretical estimate ±30% | Q4_K_M | 21.9 GB RAM | CPU only | up to 16k | — |
| DeepSeek R1 Distill Qwen 32B | Won't runTry Q2_K (heavy quality loss): Runs slowly | est. 6.4 tok/s5.1–7.6calibrated estimate ±20% | Q2_K | 15.0 / 22.4 GB | Unified memory | up to 32k |
|
| gpt-oss-120b | Won't run | — | MXFP4 | needs 64.3 GB | — | — |
|
| Llama 3.3 70B | Won't run | — | Q4_K_M | needs 45.8 GB | — | — |
|
| Qwen3.5 122B-A10B | Won't run | — | Q4_K_M | needs 79.0 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Verdict by memory size
| Model | 16 GB | 24 GB | 32 GB |
|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | Runs great | Runs great |
| HyperCLOVA X SEED 1.5B | Runs great | Runs great | Runs great |
| gpt-oss-20b | Won't run | Runs well | Runs well |
| Qwen3 30B-A3B (2507) | Won't run | Runs slowly | Runs well |
| Gemma 4 26B-A4B | Won't run | Runs slowly | Runs well |
| DeepSeek R1 Distill Llama 8B | Runs well | Runs well | Runs well |
| Kanana 1.5 8B | Runs well | Runs well | Runs well |
| Llama 3.1 8B | Runs well | Runs well | Runs well |
| Kanana 1.5 15.7B-A3B | Runs slowly | Runs well | Runs well |
| Qwen3.5 9B | Runs well | Runs well | Runs well |
| Qwen3 8B | Runs well | Runs well | Runs well |
| Gemma 4 12B | Runs well | Runs well | Runs well |
| Gemma 3 12B | Runs well | Runs well | Runs well |
| Qwen3 14B | Runs well | Runs well | Runs well |
| Qwen3.5 35B-A3B | Won't run | Won't run | Runs slowly |
| HyperCLOVA X SEED Think 14B | Runs slowly | Runs slowly | Runs slowly |
| Phi-4 | Runs slowly | Runs slowly | Runs slowly |
| Mistral Small 3.2 24B | Won't run | Runs slowly | Runs slowly |
| Gemma 3 27B | Won't run | Runs slowly | Runs slowly |
| Qwen3.5 27B | Won't run | Runs slowly | Runs slowly |
| EXAONE 4.0 32B | Won't run | Won't run | Runs slowly |
| EXAONE 4.5 33B | Won't run | Won't run | Runs slowly |
| Qwen3 32B | Won't run | Won't run | Runs slowly |
| DeepSeek R1 Distill Qwen 32B | Won't run | Won't run | Won't run |
| gpt-oss-120b | Won't run | Won't run | Won't run |
| Llama 3.3 70B | Won't run | Won't run | Won't run |
| Qwen3.5 122B-A10B | Won't run | Won't run | Won't run |
| Solar Open 100B | Won't run | Won't run | Won't run |
| Solar Open 2 250B | Won't run | Won't run | Won't run |
Measured results on the Apple M5
No public measurements for this device yet.
Frequently asked questions
What is the largest model that runs entirely on the Apple M5?
Qwen3.5 35B-A3B at IQ4_XS (a 18.17 GB file) fits entirely in 32 GB of unified memory with 8k context, at est. 27.4 tok/s (21.9–32.8, calibrated estimate ±20%).
How many local LLMs run well on the Apple M5?
At Q4_K_M with 8k context, with 32 GB of memory, out of 29 tracked models: 2 run great, 12 run well, 9 run slowly and 6 won't run.
Can the Apple M5 run a 70B model like Llama 3.3 70B?
No. Llama 3.3 70B at Q4_K_M needs about 45.8 GB, while this setup offers 22.4 GB of GPU memory and 28.0 GB of free system RAM.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.