Which local LLMs can the Apple M4 run?
Assumes 8k context and an f16 KV cache. Verdicts are shown for each memory size.
Specs
- Memory options
- 16 GB · 24 GB · 32 GB
- Memory bandwidth
- 120.0 GB/s
- FP16 compute
- 8.5 TFLOPS
- Launch year
- 2024
Verdicts at a glance
At Q4_K_M (or the closest available quant) with 8k context.
- 16 GBRuns great2 modelsRuns well7 modelsRuns slowly4 modelsWon't run16 models3 more models run at a lower quant.
- 24 GBRuns great2 modelsRuns well9 modelsRuns slowly8 modelsWon't run10 models3 more models run at a lower quant.
- 32 GBRuns great2 modelsRuns well11 modelsRuns slowly9 modelsWon't run7 models2 more models run at a lower quant.
How to read the verdicts
- Runs great
- Fully on the GPU at 20 tok/s or more
- Runs well
- Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
- Runs slowly
- 2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
- Won't run
- Does not fit, or under 2 tok/s
Every model on the Apple M4
Scroll sideways to see every column.
| Model | Verdict | Speed | Quant | Memory | Runs as | Context | Notes |
|---|---|---|---|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | est. 54.8 tok/s43.9–65.8calibrated estimate ±20% | Q4_K_M | 1.9 / 22.4 GB | Unified memory | up to 32k | — |
| HyperCLOVA X SEED 1.5B | Runs great | est. 39.7 tok/s31.7–47.6calibrated estimate ±20% | Q4_K_M | 2.4 / 22.4 GB | Unified memory | up to 16k | — |
| gpt-oss-20b | Runs wellDetails | est. 15.7 tok/s12.6–18.9calibrated estimate ±20% | MXFP4 | 12.9 / 22.4 GB | Unified memory | up to 32k |
|
| Qwen3 30B-A3B (2507) | Runs wellDetails | est. 12.8 tok/s10.2–15.4calibrated estimate ±20% | Q4_K_M | 19.9 / 22.4 GB | Unified memory | up to 16k | — |
| Gemma 4 26B-A4B | Runs well | est. 12.5 tok/s10.0–14.9calibrated estimate ±20% | Q4_K_M | 17.9 / 22.4 GB | Unified memory | up to 32k | — |
| DeepSeek R1 Distill Llama 8B | Runs well | est. 12.0 tok/s9.6–14.4calibrated estimate ±20% | Q4_K_M | 6.6 / 22.4 GB | Unified memory | up to 8k |
|
| Kanana 1.5 8B | Runs well | est. 12.0 tok/s9.6–14.4calibrated estimate ±20% | Q4_K_M | 6.6 / 22.4 GB | Unified memory | up to 16k | — |
| Llama 3.1 8B | Runs wellDetails | est. 12.0 tok/s9.6–14.4calibrated estimate ±20% | Q4_K_M | 6.6 / 22.4 GB | Unified memory | up to 16k | — |
| Kanana 1.5 15.7B-A3B | Runs well | est. 11.7 tok/s9.4–14.1calibrated estimate ±20% | Q4_K_M | 12.1 / 22.4 GB | Unified memory | up to 16k | — |
| Qwen3.5 9B | Runs well | est. 11.7 tok/s9.4–14.1calibrated estimate ±20% | Q4_K_M | 6.7 / 22.4 GB | Unified memory | up to 64k | — |
| Qwen3 8B | Runs well | est. 11.5 tok/s9.2–13.9calibrated estimate ±20% | Q4_K_M | 6.8 / 22.4 GB | Unified memory | up to 16k | — |
| Gemma 4 12B | Runs well | est. 9.4 tok/s7.5–11.3calibrated estimate ±20% | Q4_K_M | 8.2 / 22.4 GB | Unified memory | up to 16k | — |
| Gemma 3 12B | Runs wellDetails | est. 9.2 tok/s7.3–11.0calibrated estimate ±20% | Q4_K_M | 8.4 / 22.4 GB | Unified memory | up to 16k | — |
| Qwen3.5 35B-A3B | Runs slowly | est. 17.5 tok/s12.3–22.8theoretical estimate ±30% | Q4_K_M | 22.8 GB RAM | CPU only | up to 256k |
|
| HyperCLOVA X SEED Think 14B | Runs slowly | est. 7.1 tok/s4.9–9.2theoretical estimate ±30% | Q4_K_M | 10.8 / 22.4 GB | Unified memory | up to 64k |
|
| Qwen3 14B | Runs slowlyDetails | est. 7.0 tok/s5.6–8.4calibrated estimate ±20% | Q4_K_M | 10.9 / 22.4 GB | Unified memory | up to 32k |
|
| Phi-4 | Runs slowly | est. 6.7 tok/s5.4–8.1calibrated estimate ±20% | Q4_K_M | 11.3 / 22.4 GB | Unified memory | up to 16k |
|
| Mistral Small 3.2 24B | Runs slowly | est. 4.6 tok/s3.7–5.5calibrated estimate ±20% | Q4_K_M | 16.2 / 22.4 GB | Unified memory | up to 32k | — |
| Gemma 3 27B | Runs slowly | est. 4.2 tok/s3.3–5.0calibrated estimate ±20% | Q4_K_M | 17.8 / 22.4 GB | Unified memory | up to 32k | — |
| Qwen3.5 27B | Runs slowly | est. 4.1 tok/s3.3–4.9calibrated estimate ±20% | Q4_K_M | 18.2 / 22.4 GB | Unified memory | up to 32k | — |
| EXAONE 4.0 32B | Runs slowly | est. 3.6 tok/s2.9–4.3calibrated estimate ±20% | Q4_K_M | 20.5 / 22.4 GB | Unified memory | up to 32k |
|
| EXAONE 4.5 33B | Runs slowly | est. 3.5 tok/s2.8–4.2calibrated estimate ±20% | Q4_K_M | 21.2 / 22.4 GB | Unified memory | up to 16k |
|
| DeepSeek R1 Distill Qwen 32B | Won't runTry Q2_K (heavy quality loss): Runs slowly | est. 5.0 tok/s4.0–6.0calibrated estimate ±20% | Q2_K | 15.0 / 22.4 GB | Unified memory | up to 32k |
|
| Qwen3 32B | Won't runTry IQ4_XS: Runs slowly | est. 3.6 tok/s2.9–4.4calibrated estimate ±20% | IQ4_XS | 20.4 / 22.4 GB | Unified memory | up to 8k | — |
| gpt-oss-120b | Won't runDetails | — | MXFP4 | needs 64.3 GB | — | — |
|
| Llama 3.3 70B | Won't run | — | Q4_K_M | needs 45.8 GB | — | — |
|
| Qwen3.5 122B-A10B | Won't run | — | Q4_K_M | needs 79.0 GB | — | — |
|
| Solar Open 100B | Won't run | — | Q4_K_M | needs 64.4 GB | — | — |
|
| Solar Open 2 250B | Won't run | — | IQ4_XS | needs 137.2 GB | — | — |
|
Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).
Verdict by memory size
| Model | 16 GB | 24 GB | 32 GB |
|---|---|---|---|
| EXAONE 4.0 1.2B | Runs great | Runs great | Runs great |
| HyperCLOVA X SEED 1.5B | Runs great | Runs great | Runs great |
| gpt-oss-20b | Won't run | Runs well | Runs well |
| Qwen3 30B-A3B (2507) | Won't run | Runs slowly | Runs well |
| Gemma 4 26B-A4B | Won't run | Runs slowly | Runs well |
| DeepSeek R1 Distill Llama 8B | Runs well | Runs well | Runs well |
| Kanana 1.5 8B | Runs well | Runs well | Runs well |
| Llama 3.1 8B | Runs well | Runs well | Runs well |
| Kanana 1.5 15.7B-A3B | Runs slowly | Runs well | Runs well |
| Qwen3.5 9B | Runs well | Runs well | Runs well |
| Qwen3 8B | Runs well | Runs well | Runs well |
| Gemma 4 12B | Runs well | Runs well | Runs well |
| Gemma 3 12B | Runs well | Runs well | Runs well |
| Qwen3.5 35B-A3B | Won't run | Won't run | Runs slowly |
| HyperCLOVA X SEED Think 14B | Runs slowly | Runs slowly | Runs slowly |
| Qwen3 14B | Runs slowly | Runs slowly | Runs slowly |
| Phi-4 | Runs slowly | Runs slowly | Runs slowly |
| Mistral Small 3.2 24B | Won't run | Runs slowly | Runs slowly |
| Gemma 3 27B | Won't run | Runs slowly | Runs slowly |
| Qwen3.5 27B | Won't run | Runs slowly | Runs slowly |
| EXAONE 4.0 32B | Won't run | Won't run | Runs slowly |
| EXAONE 4.5 33B | Won't run | Won't run | Runs slowly |
| DeepSeek R1 Distill Qwen 32B | Won't run | Won't run | Won't run |
| Qwen3 32B | Won't run | Won't run | Won't run |
| gpt-oss-120b | Won't run | Won't run | Won't run |
| Llama 3.3 70B | Won't run | Won't run | Won't run |
| Qwen3.5 122B-A10B | Won't run | Won't run | Won't run |
| Solar Open 100B | Won't run | Won't run | Won't run |
| Solar Open 2 250B | Won't run | Won't run | Won't run |
Measured results on the Apple M4
Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.
| Model | Quant | Backend | Context | Prompt (tok/s) | Generation (tok/s) | Flags | Source | Measured |
|---|---|---|---|---|---|---|---|---|
| llama-2-7b | Q4_0 | llama.cpp | 512 | 221.0 | 24.1 | Metal | github.com | 2024-11-15 |
Frequently asked questions
What is the largest model that runs entirely on the Apple M4?
Qwen3.5 35B-A3B at IQ4_XS (a 18.17 GB file) fits entirely in 32 GB of unified memory with 8k context, at est. 21.4 tok/s (17.1–25.7, calibrated estimate ±20%).
How many local LLMs run well on the Apple M4?
At Q4_K_M with 8k context, with 32 GB of memory, out of 29 tracked models: 2 run great, 11 run well, 9 run slowly and 7 won't run.
Can the Apple M4 run a 70B model like Llama 3.3 70B?
No. Llama 3.3 70B at Q4_K_M needs about 45.8 GB, while this setup offers 22.4 GB of GPU memory and 28.0 GB of free system RAM.
Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.