Which local LLMs can the DDR5-5600 dual-channel CPU run?

No GPU — the whole model runs from system RAM

CPU-only inference with 4 GB of RAM kept for the OS, 8k context and an f16 KV cache. Verdicts are shown for each RAM size.

Specs

Memory options
16 GB · 32 GB · 64 GB · 96 GB
Memory bandwidth
89.6 GB/s
Launch year
2021

Verdicts at a glance

At Q4_K_M (or the closest available quant) with 8k context.

  • 16 GBRuns great0 modelsRuns well0 modelsRuns slowly13 modelsWon't run16 models3 more models run at a lower quant.
  • 32 GBRuns great0 modelsRuns well0 modelsRuns slowly21 modelsWon't run8 models1 more model runs at a lower quant.
  • 64 GBRuns great0 modelsRuns well0 modelsRuns slowly21 modelsWon't run8 models3 more models run at a lower quant.
  • 96 GBRuns great0 modelsRuns well0 modelsRuns slowly24 modelsWon't run5 models1 more model runs at a lower quant.

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Every model on the DDR5-5600 dual-channel CPU

Scroll sideways to see every column.

29 models sorted by verdict and speed, with 96 GB of memory
ModelVerdictSpeedQuantMemoryRuns asContextNotes
EXAONE 4.0 1.2B
Runs slowly
est. 34.1 tok/s27.3–40.9calibrated estimate ±20%
Q4_K_M1.3 GB RAMCPU onlyup to 64k
  • CPU inference — capped at Runs slowly
HyperCLOVA X SEED 1.5B
Runs slowly
est. 24.7 tok/s19.7–29.6calibrated estimate ±20%
Q4_K_M1.8 GB RAMCPU onlyup to 16k
  • CPU inference — capped at Runs slowly
Qwen3 30B-A3B (2507)
Runs slowly
19.0 tok/s8k estimate 13.7 tok/s (11.0–16.4, calibrated estimate ±20%)measured (1 run, 2k context)
Q4_K_M19.4 GB RAMCPU onlyup to 128k
  • CPU inference — capped at Runs slowly
Qwen3.5 35B-A3B
Runs slowly
est. 18.7 tok/s15.0–22.5calibrated estimate ±20%
Q4_K_M22.8 GB RAMCPU onlyup to 256k
  • CPU inference — capped at Runs slowly
gpt-oss-20b
Runs slowly
est. 16.8 tok/s13.5–20.2calibrated estimate ±20%
MXFP412.3 GB RAMCPU onlyup to 128k
  • CPU inference — capped at Runs slowly
  • reasoning model
Gemma 4 26B-A4B
Runs slowly
est. 13.3 tok/s10.7–16.0calibrated estimate ±20%
Q4_K_M17.3 GB RAMCPU onlyup to 256k
  • CPU inference — capped at Runs slowly
Kanana 1.5 15.7B-A3B
Runs slowly
est. 12.6 tok/s10.0–15.1calibrated estimate ±20%
Q4_K_M11.5 GB RAMCPU onlyup to 32k
  • CPU inference — capped at Runs slowly
gpt-oss-120b
Runs slowly
est. 12.6 tok/s10.0–15.1calibrated estimate ±20%
MXFP463.7 GB RAMCPU onlyup to 128k
  • CPU inference — capped at Runs slowly
  • reasoning model
DeepSeek R1 Distill Llama 8B
Runs slowly
est. 7.5 tok/s6.0–9.0calibrated estimate ±20%
Q4_K_M6.0 GB RAMCPU onlyup to 64k
  • reasoning model
Kanana 1.5 8B
Runs slowly
est. 7.5 tok/s6.0–9.0calibrated estimate ±20%
Q4_K_M6.0 GB RAMCPU onlyup to 32k—
Llama 3.1 8B
Runs slowly
est. 7.5 tok/s6.0–9.0calibrated estimate ±20%
Q4_K_M6.0 GB RAMCPU onlyup to 128k—
Qwen3.5 9B
Runs slowly
est. 7.3 tok/s5.8–8.8calibrated estimate ±20%
Q4_K_M6.1 GB RAMCPU onlyup to 256k—
Qwen3 8B
Runs slowly
est. 7.2 tok/s5.7–8.6calibrated estimate ±20%
Q4_K_M6.2 GB RAMCPU onlyup to 32k—
Qwen3.5 122B-A10B
Runs slowly
est. 6.0 tok/s4.8–7.2calibrated estimate ±20%
Q4_K_M78.5 GB RAMCPU onlyup to 256k—
Gemma 4 12B
Runs slowly
est. 5.9 tok/s4.7–7.0calibrated estimate ±20%
Q4_K_M7.7 GB RAMCPU onlyup to 128k—
Gemma 3 12B
Runs slowly
est. 5.7 tok/s4.6–6.9calibrated estimate ±20%
Q4_K_M7.8 GB RAMCPU onlyup to 128k—
HyperCLOVA X SEED Think 14B
Runs slowly
est. 4.4 tok/s3.1–5.7theoretical estimate ±30%
Q4_K_M10.2 GB RAMCPU onlyup to 32k
  • reasoning model
Solar Open 100B
Runs slowly
est. 4.3 tok/s3.5–5.2calibrated estimate ±20%
Q4_K_M63.9 GB RAMCPU onlyup to 32k—
Qwen3 14B
Runs slowly
est. 4.3 tok/s3.5–5.2calibrated estimate ±20%
Q4_K_M10.3 GB RAMCPU onlyup to 32k—
Phi-4
Runs slowly
est. 4.2 tok/s3.3–5.0calibrated estimate ±20%
Q4_K_M10.7 GB RAMCPU onlyup to 16k
  • reasoning model
Mistral Small 3.2 24B
Runs slowly
est. 2.9 tok/s2.3–3.4calibrated estimate ±20%
Q4_K_M15.7 GB RAMCPU onlyup to 32k—
Gemma 3 27B
Runs slowly
est. 2.6 tok/s2.1–3.1calibrated estimate ±20%
Q4_K_M17.2 GB RAMCPU onlyup to 64k—
Qwen3.5 27B
Runs slowly
est. 2.5 tok/s2.0–3.0calibrated estimate ±20%
Q4_K_M17.6 GB RAMCPU onlyup to 64k—
Qwen3 32B
Runs slowly
est. 2.0 tok/s1.6–2.5calibrated estimate ±20%
Q4_K_M21.9 GB RAMCPU onlyup to 8k—
DeepSeek R1 Distill Qwen 32B
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 3.1 tok/s2.5–3.7calibrated estimate ±20%
Q2_K14.5 GB RAMCPU onlyup to 8k
  • reasoning model
EXAONE 4.0 32B
Won't run
est. 2.3 tok/s1.8–2.7calibrated estimate ±20%
Q4_K_M19.9 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
EXAONE 4.5 33B
Won't run
est. 2.2 tok/s1.7–2.6calibrated estimate ±20%
Q4_K_M20.6 GB RAMCPU only—
  • Fits, but too slow to use
  • reasoning model
Llama 3.3 70B
Won't run
est. 1.0 tok/s0.8–1.2calibrated estimate ±20%
Q4_K_M45.2 GB RAMCPU only—
  • Fits, but too slow to use
Solar Open 2 250B
Won't run
—IQ4_XSneeds 136.6 GB——
  • Needs about 136.6 GB; 92.0 GB of RAM is free

Reasoning models spend extra tokens thinking, so their speed thresholds are 1.5× stricter (30 / 12 / 3 tok/s).

Verdict by memory size

Q4_K_M verdict for each memory configuration — open the model page for speeds
Model16 GB32 GB64 GB96 GB
EXAONE 4.0 1.2BRuns slowlyRuns slowlyRuns slowlyRuns slowly
HyperCLOVA X SEED 1.5BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Qwen3 30B-A3B (2507)Won't runRuns slowlyRuns slowlyRuns slowly
Qwen3.5 35B-A3BWon't runRuns slowlyRuns slowlyRuns slowly
gpt-oss-20bWon't runRuns slowlyRuns slowlyRuns slowly
Gemma 4 26B-A4BWon't runRuns slowlyRuns slowlyRuns slowly
Kanana 1.5 15.7B-A3BRuns slowlyRuns slowlyRuns slowlyRuns slowly
gpt-oss-120bWon't runWon't runWon't runRuns slowly
DeepSeek R1 Distill Llama 8BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Kanana 1.5 8BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Llama 3.1 8BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Qwen3.5 9BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Qwen3 8BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Qwen3.5 122B-A10BWon't runWon't runWon't runRuns slowly
Gemma 4 12BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Gemma 3 12BRuns slowlyRuns slowlyRuns slowlyRuns slowly
HyperCLOVA X SEED Think 14BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Solar Open 100BWon't runWon't runWon't runRuns slowly
Qwen3 14BRuns slowlyRuns slowlyRuns slowlyRuns slowly
Phi-4Runs slowlyRuns slowlyRuns slowlyRuns slowly
Mistral Small 3.2 24BWon't runRuns slowlyRuns slowlyRuns slowly
Gemma 3 27BWon't runRuns slowlyRuns slowlyRuns slowly
Qwen3.5 27BWon't runRuns slowlyRuns slowlyRuns slowly
Qwen3 32BWon't runRuns slowlyRuns slowlyRuns slowly
DeepSeek R1 Distill Qwen 32BWon't runWon't runWon't runWon't run
EXAONE 4.0 32BWon't runWon't runWon't runWon't run
EXAONE 4.5 33BWon't runWon't runWon't runWon't run
Llama 3.3 70BWon't runWon't runWon't runWon't run
Solar Open 2 250BWon't runWon't runWon't runWon't run

Measured results on the DDR5-5600 dual-channel CPU

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured results on the DDR5-5600 dual-channel CPU
ModelQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
mistral-7bQ4_K_Mllamafile512—11.2—github.com2024-06-15
Qwen3 30B-A3B (2507)Q4_K_Mllama.cpp2k—19.0—hardware-corner.net2026-08-09

Frequently asked questions

What is the largest model this CPU setup can run?

Qwen3.5 122B-A10B at Q4_K_M (a 78.26 GB file) runs from 96 GB of RAM at est. 6.0 tok/s (4.8–7.2, calibrated estimate ±20%).

How many local LLMs run well on the DDR5-5600 dual-channel CPU?

At Q4_K_M with 8k context, with 96 GB of memory, out of 29 tracked models: 0 run great, 0 run well, 24 run slowly and 5 won't run.

Can the DDR5-5600 dual-channel CPU run a 70B model like Llama 3.3 70B?

It fits in memory but is too slow to use: est. 1.0 tok/s (0.8–1.2, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.