Llama 3.3 70B hardware requirements

About the model

Parameters
70.6B
Attention
Full attention
Max context
131,072 tokens
Released
2024-12-06

GGUF files

Tracked quantizations of Llama 3.3 70B, smallest first
QuantFile sizeQualityPublished by
IQ2_M24.12 GBHeavy lossCommunity GGUF · bartowski
Q2_K26.38 GBHeavy lossCommunity GGUF · bartowski
IQ4_XS37.90 GBGood balanceCommunity GGUF · bartowski
Q4_K_Mbaseline42.52 GBGood balanceCommunity GGUF · bartowski
Q8_074.98 GBNear-losslessCommunity GGUF · bartowski

Memory needed by quant and context

Weights + f16 KV cache + compute buffer, in GB. Add about 0.6 GB if the GPU also drives your display on Windows.
Quant4k8k16k32k64k128k
IQ2_M26.027.430.135.746.768.8
Q2_K28.329.632.437.949.071.0
IQ4_XS39.841.243.949.460.582.5
Q4_K_M44.445.848.554.165.187.2
Q8_076.978.281.086.597.6119.6

Which GPUs can run Llama 3.3 70B?

How to read the verdicts

Runs great
Fully on the GPU at 20 tok/s or more
Runs well
Fully on the GPU at 8–20 tok/s; MoE experts in system RAM at 20 tok/s or more; or at least 90% on the GPU at 8 tok/s or more
Runs slowly
2–8 tok/s; CPU-only; less than 90% on the GPU; or MoE experts in system RAM below 20 tok/s
Won't run
Does not fit, or under 2 tok/s

Dense models split between GPU and CPU slow down sharply — the CPU side sets the pace. MoE models that keep only their experts in system RAM degrade far more gently.

Scroll sideways to see every column.

54 GPUs, Macs and CPU setups at Q4_K_M with 8k context, sorted by verdict and speed
HardwareVerdictSpeedQuantMemoryRuns asContextNotesPrice
Apple M5 Max (40-core GPU) · 128 GB
Runs well
est. 8.1 tok/s6.5–9.8calibrated estimate ±20%
Q4_K_M45.8 / 89.6 GBUnified memoryup to 8k——
Apple M3 Ultra · 96 GB
Runs slowly
est. 8.0 tok/s6.4–9.6calibrated estimate ±20%
Q4_K_M45.8 / 67.2 GBUnified memoryup to 64k
  • Try IQ4_XS: Runs well
—
Apple M2 Ultra · 64 GB
Runs slowly
est. 6.0 tok/s4.2–7.8theoretical estimate ±30%
Q4_K_M45.2 GB RAMCPU onlyup to 32k
  • Try IQ4_XS: Runs well
—
Apple M4 Max (40-core GPU) · 64 GB
Runs slowly
est. 4.2 tok/s2.9–5.5theoretical estimate ±30%
Q4_K_M45.2 GB RAMCPU onlyup to 32k
  • Try IQ4_XS: Runs well
—
Ryzen AI Max+ 395 (Strix Halo) · 64 GB
Runs slowly
est. 2.8 tok/s2.0–3.7theoretical estimate ±30%
Q4_K_M45.2 GB RAMCPU onlyup to 32k—$1,999 launch MSRP
GeForce RTX 5090
Runs slowly
est. 2.7 tok/s2.2–3.3calibrated estimate ±20%
Q4_K_M32.0 / 32.0 GB + 14.4 GB RAMPartial offloadup to 16k
  • Q2_K (heavy quality loss): Runs great
$4,200 street (as of 2026-08-09)
NVIDIA DGX Spark · 128 GB
Runs slowly
est. 2.4 tok/s1.9–2.9calibrated estimate ±20%
Q4_K_M45.8 / 89.6 GBUnified memoryup to 32k—$3,999 launch MSRP
Apple M5 Pro · 64 GB
Runs slowly
est. 2.4 tok/s1.7–3.1theoretical estimate ±30%
Q4_K_M45.2 GB RAMCPU onlyup to 32k——
Apple M4 Pro · 64 GB
Runs slowly
est. 2.1 tok/s1.5–2.7theoretical estimate ±30%
Q4_K_M45.2 GB RAMCPU onlyup to 8k——
Apple M5 Max (32-core GPU) · 48 GB
Won't runTry IQ4_XS: Runs slowly
est. 4.0 tok/s2.8–5.1theoretical estimate ±30%
IQ4_XS40.6 GB RAMCPU onlyup to 16k
  • Q2_K (heavy quality loss): Runs well
—
Radeon RX 7900 XT
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 3.6 tok/s2.9–4.3calibrated estimate ±20%
Q2_K20.0 / 20.0 GB + 10.2 GB RAMPartial offloadup to 16k—$899 launch MSRP
Apple M4 Max (32-core GPU) · 48 GB
Won't runTry IQ4_XS: Runs slowly
est. 3.5 tok/s2.5–4.6theoretical estimate ±30%
IQ4_XS40.6 GB RAMCPU onlyup to 16k
  • Q2_K (heavy quality loss): Runs well
—
GeForce RTX 5080
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.2–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$1,256 street (as of 2026-08-09)
GeForce RTX 5070 Ti
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.2–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$949 street (as of 2026-08-09)
GeForce RTX 5080 Laptop
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.2–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k——
GeForce RTX 4080 Super
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.1–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$999 launch MSRP
Radeon RX 9070
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.1–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$639 street (as of 2026-08-09)
Radeon RX 9070 XT
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.1–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$689 street (as of 2026-08-09)
GeForce RTX 4080
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.7 tok/s2.1–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$1,199 launch MSRP
Radeon RX 7800 XT
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.6 tok/s2.1–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$499 launch MSRP
GeForce RTX 4070 Ti Super
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.6 tok/s2.1–3.2calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$799 launch MSRP
GeForce RTX 4090 Laptop
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.6 tok/s2.1–3.1calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k——
GeForce RTX 5060 Ti 16GB
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.5 tok/s2.0–3.1calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 16k—$569 street (as of 2026-08-09)
Radeon RX 9060 XT 16GB
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.5 tok/s2.0–3.0calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 8k—$449 street (as of 2026-08-09)
GeForce RTX 4060 Ti 16GB
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.4 tok/s1.9–2.9calibrated estimate ±20%
Q2_K16.0 / 16.0 GB + 14.2 GB RAMPartial offloadup to 8k—$499 launch MSRP
Radeon RX 7900 XTX
Won't runTry IQ4_XS: Runs slowly
est. 2.2 tok/s1.8–2.6calibrated estimate ±20%
IQ4_XS24.0 / 24.0 GB + 17.8 GB RAMPartial offloadup to 8k—$999 launch MSRP
GeForce RTX 3090 Ti
Won't runTry IQ4_XS: Runs slowly
est. 2.2 tok/s1.8–2.6calibrated estimate ±20%
IQ4_XS24.0 / 24.0 GB + 17.8 GB RAMPartial offloadup to 8k—$1,999 launch MSRP
GeForce RTX 4090
Won't runTry IQ4_XS: Runs slowly
est. 2.2 tok/s1.8–2.6calibrated estimate ±20%
IQ4_XS24.0 / 24.0 GB + 17.8 GB RAMPartial offloadup to 8k—$2,755 street (as of 2026-08-09)
GeForce RTX 3090
Won't runTry IQ4_XS: Runs slowly
est. 2.2 tok/s1.7–2.6calibrated estimate ±20%
IQ4_XS24.0 / 24.0 GB + 17.8 GB RAMPartial offloadup to 8k—$1,050 street (as of 2026-08-09)
GeForce RTX 3080 10GB
Won't runTry IQ2_M (heavy quality loss): Runs slowly
est. 2.2 tok/s1.7–2.6calibrated estimate ±20%
IQ2_M10.0 / 10.0 GB + 18.0 GB RAMPartial offloadup to 8k—$699 launch MSRP
GeForce RTX 5090 Laptop
Won't runTry IQ4_XS: Runs slowly
est. 2.2 tok/s1.7–2.6calibrated estimate ±20%
IQ4_XS24.0 / 24.0 GB + 17.8 GB RAMPartial offloadup to 8k——
GeForce RTX 3080 12GB
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.2 tok/s1.7–2.6calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k—$799 launch MSRP
GeForce RTX 5070
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.6calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k—$629 street (as of 2026-08-09)
GeForce RTX 5070 Ti Laptop
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.6calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k——
GeForce RTX 4070
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.5calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k—$599 launch MSRP
GeForce RTX 4070 Super
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.5calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k—$599 launch MSRP
GeForce RTX 4080 Laptop
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.5calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k——
Intel Arc B580
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.5calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k—$290 street (as of 2026-08-09)
GeForce RTX 3060 12GB
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.1 tok/s1.7–2.5calibrated estimate ±20%
Q2_K12.0 / 12.0 GB + 18.2 GB RAMPartial offloadup to 8k—$250 street (as of 2026-08-09)
GeForce RTX 2080 Ti
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.0 tok/s1.6–2.4calibrated estimate ±20%
Q2_K11.0 / 11.0 GB + 19.2 GB RAMPartial offloadup to 8k—$999 launch MSRP
GeForce GTX 1080 Ti
Won't runTry Q2_K (heavy quality loss): Runs slowly
est. 2.0 tok/s1.4–2.6theoretical estimate ±30%
Q2_K11.0 / 11.0 GB + 19.2 GB RAMPartial offloadup to 8k—$699 launch MSRP
DDR5-6000 dual-channel CPU · 96 GB
Won't run
est. 1.1 tok/s0.8–1.3calibrated estimate ±20%
Q4_K_M45.2 GB RAMCPU only—
  • Fits, but too slow to use
—
DDR5-5600 dual-channel CPU · 96 GB
Won't run
est. 1.0 tok/s0.8–1.2calibrated estimate ±20%
Q4_K_M45.2 GB RAMCPU only—
  • Fits, but too slow to use
—
DDR4-3200 dual-channel CPU · 64 GB
Won't run
est. 0.6 tok/s0.5–0.7calibrated estimate ±20%
Q4_K_M45.2 GB RAMCPU only—
  • Fits, but too slow to use
—
Apple M4 · 32 GB
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
—
Apple M5 · 32 GB
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 28.0 GB of memory is free (22.4 GB usable by the GPU)
—
GeForce RTX 3050 8GB
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
$249 launch MSRP
GeForce RTX 3070
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
$499 launch MSRP
GeForce RTX 4060
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
$299 launch MSRP
GeForce RTX 4060 Laptop
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
—
GeForce RTX 4060 Ti 8GB
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
$399 launch MSRP
GeForce RTX 4070 Laptop
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
—
GeForce RTX 5060
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
$339 street (as of 2026-08-09)
GeForce RTX 5060 Ti 8GB
Won't run
—Q4_K_Mneeds 45.8 GB——
  • Needs about 45.8 GB; 7.4 GB of VRAM and 28.0 GB of RAM are free
$429 street (as of 2026-08-09)

Cheapest GPUs that run it

No GPU with a current street price reaches Runs great or Runs well at Q4_K_M. The table above lists Macs, unified-memory PCs and CPU setups.

Measured results for Llama 3.3 70B

Public benchmarks we calibrate against. Their conditions (context, backend, flags) can differ from the estimates above.

Measured results for Llama 3.3 70B
HardwareQuantBackendContextPrompt (tok/s)Generation (tok/s)FlagsSourceMeasured
Apple M3 Max 128GBQ4_K_Mollama2k—9.8—heyuan110.com2026-04-14

Frequently asked questions

How much VRAM does Llama 3.3 70B need?

At Q4_K_M the weights are 42.52 GB; with 8k context the total is about 45.8 GB (KV cache 2.7 GB, compute buffer 0.6 GB). The smallest tracked file (IQ2_M (heavy quality loss)) needs about 27.4 GB.

What is the cheapest GPU that runs Llama 3.3 70B well?

No GPU with a current street price reaches Runs great or Runs well at Q4_K_M. The best tracked option is Apple M5 Max (40-core GPU) · 128 GB: Runs well.

Can Llama 3.3 70B run on an 8 GB, 16 GB or 24 GB GPU?

At Q4_K_M with 8k context: GeForce RTX 4060 — Won't run; GeForce RTX 5060 Ti 16GB — Won't run; GeForce RTX 3090 — Won't run, est. 1.8 tok/s (1.4–2.1, calibrated estimate ±20%).

Speeds are estimates from memory bandwidth, calibrated against public benchmarks, and each one comes with an error band and a confidence label. Real results vary with drivers, backend, context length and thermals.