Can I Run AI locally?

Find out which AI models your machine can actually run.

GPU | VRAM | BW | RAM | CORES 2,560

Estimates based on browser APIs. Actual specs may vary. WebGPU

12 Runs great 7 Runs well 4 Decent 9 Tight fit 24 Barely runs 21 Too heavy

Llama 3.1 8B

1 year ago

Meta · 8B · Llama 3.1 Community

Meta's versatile 8B — great quality/speed ratio

4.6 GB 77% · 128K ctx · ~29 tok/s
Tight fit 51/100

Qwen 3.5 9B

2mo ago

Alibaba · 9B · Apache 2.0

Multimodal Qwen 3.5 mid-size

5.1 GB 85% · 32K ctx · ~26 tok/s
Tight fit 46/100

Phi-4 14B

1 year ago

Microsoft · 14B · MIT

Microsoft's reasoning-focused model

7.7 GB 128% ↗ RAM · 16K ctx · ~9 tok/s
Barely runs 15/100

Mistral Small 3.1 24B

1 year ago

Mistral AI · 24B · Apache 2.0

Multimodal Mistral with vision support

12.8 GB 213% ↗ RAM · 128K ctx · ~4 tok/s
Barely runs 10/100

GPT-OSS 20B

8mo ago

OpenAI · 21B · Apache 2.0

OpenAI's open-weight MoE with configurable reasoning

11.3 GB 188% ↗ RAM · 128K ctx · ~4 tok/s
Barely runs 9/100

Gemma 3 27B

1 year ago

Google · 27B · Gemma

Google's flagship Gemma 3 model

14.3 GB 238% ↗ RAM · 128K ctx · ~3 tok/s
Barely runs 9/100

Qwen 2.5 Coder 32B

1 year ago

Alibaba · 32B · Apache 2.0

Best open-source coding model at release

16.9 GB 282% ↗ RAM · 128K ctx · ~2 tok/s
Barely runs 8/100

Qwen 3 32B

1 year ago

Alibaba · 32B · Apache 2.0

Qwen 3 flagship dense model

16.9 GB 282% ↗ RAM · 128K ctx · ~2 tok/s
Barely runs 8/100

DeepSeek R1 Distill 32B

1 year ago

DeepSeek · 32B · MIT

R1 reasoning distilled into Qwen 32B — sweet spot

16.9 GB 282% ↗ RAM · 64K ctx · ~2 tok/s
Barely runs 8/100

Llama 3.3 70B

1 year ago

Meta · 70B · Llama 3.3 Community

Best open model at 70B class

36.4 GB 607% · 128K ctx · 0 tok/s
Too heavy 0/100

Llama 4 Scout 17B

1 year ago

Meta · 109B · Llama 4 Community

MoE with 16 experts, 17B active params

56.3 GB 938% · 128K ctx · 0 tok/s
Too heavy 0/100

GPT-OSS 120B

8mo ago

OpenAI · 117B · Apache 2.0

OpenAI's flagship open-weight MoE — 52.6% SWE-bench

60.4 GB 1007% · 128K ctx · 0 tok/s
Too heavy 0/100

Devstral 2 123B

4mo ago

Mistral AI · 123B · MRL

Dense 123B coding model — 72.2% SWE-bench Verified

63.5 GB 1058% · 256K ctx · 0 tok/s
Too heavy 0/100

DeepSeek R1

1 year ago

DeepSeek · 671B · MIT

Massive MoE reasoning model — 37B active

344.2 GB 5737% · 64K ctx · 0 tok/s
Too heavy 0/100

DeepSeek V3.2

4mo ago

DeepSeek · 685B · MIT

State-of-the-art MoE — 37B active params

351.4 GB 5857% · 128K ctx · 0 tok/s
Too heavy 0/100

Kimi K2

9mo ago

Moonshot AI · 1T · Kimi

1T-param MoE with 384 experts — 32B active, strong agentic coding

512.7 GB 8545% · 128K ctx · 0 tok/s
Too heavy 0/100
All models

Qwen 3.5 0.8B

2mo ago

Alibaba · 0.8B · Apache 2.0

Ultra-tiny model for embedded and edge

0.9 GB 15% · 32K ctx · ~149 tok/s
Runs great 92/100

Llama 3.2 1B

1 year ago

Meta · 1B · Llama 3.2 Community

Meta's smallest Llama for edge devices

1 GB 17% · 128K ctx · ~134 tok/s
Runs great 92/100

Gemma 3 1B

1 year ago

Google · 1B · Gemma

Google's tiny Gemma for on-device

1 GB 17% · 32K ctx · ~134 tok/s
Runs great 92/100

TinyLlama 1.1B

2y ago

Community · 1.1B · Apache 2.0

Ultralight model for constrained devices

1.1 GB 18% · 2K ctx · ~122 tok/s
Runs great 92/100

Qwen 2.5 Coder 1.5B

1 year ago

Alibaba · 1.5B · Apache 2.0

Ultra-lightweight coding model

1.3 GB 22% · 32K ctx · ~103 tok/s
Runs great 92/100

DeepSeek R1 1.5B

1 year ago

DeepSeek · 1.5B · MIT

Tiny reasoning model distilled from R1

1.3 GB 22% · 64K ctx · ~103 tok/s
Runs great 92/100

Qwen 3 1.7B

1 year ago

Alibaba · 1.7B · Apache 2.0

Compact multilingual Qwen 3

1.4 GB 23% · 32K ctx · ~96 tok/s
Runs great 92/100

Qwen 3 0.6B

1 year ago

Alibaba · 0.6B · Apache 2.0

Ultra-light Qwen 3 model for constrained devices

0.8 GB 13% · 32K ctx · ~168 tok/s
Runs great 91/100

Qwen 3.5 2B

2mo ago

Alibaba · 2B · Apache 2.0

Small multimodal Qwen 3.5

1.5 GB 25% · 32K ctx · ~90 tok/s
Runs great 91/100

Gemma 2 2B

1 year ago

Google · 2B · Gemma

Google's compact open model

1.5 GB 25% · 8K ctx · ~90 tok/s
Runs great 91/100

Llama 3.2 3B

1 year ago

Meta · 3B · Llama 3.2 Community

Lightweight Llama for mobile and edge

2 GB 33% · 128K ctx · ~67 tok/s
Runs great 85/100

SmolLM3 3B

9mo ago

HuggingFace · 3B · Apache 2.0

Lightweight multilingual reasoning

2 GB 33% · 128K ctx · ~67 tok/s
Runs great 85/100

Phi-3.5 Mini

1 year ago

Microsoft · 3.8B · MIT

Microsoft's efficient small model with long context

2.4 GB 40% · 128K ctx · ~56 tok/s
Runs well 79/100

Phi-4 Mini Reasoning

1 year ago

Microsoft · 3.8B · MIT

Lightweight reasoning model

2.4 GB 40% · 16K ctx · ~56 tok/s
Runs well 79/100

Qwen 3 4B

1 year ago

Alibaba · 4B · Apache 2.0

Compact Qwen 3 for general tasks

2.5 GB 42% · 32K ctx · ~54 tok/s
Runs well 78/100

Gemma 3 4B

1 year ago

Google · 4B · Gemma

Multimodal Gemma with 128K context

2.5 GB 42% · 128K ctx · ~54 tok/s
Runs well 78/100

Qwen 3.5 4B

2mo ago

Alibaba · 4B · Apache 2.0

Small multimodal Qwen 3.5

2.5 GB 42% · 32K ctx · ~54 tok/s
Runs well 78/100

Gemma 4 E2B IT

5d ago

Google · 5B · Gemma

Gemma 4 efficient instruct model (official)

3.1 GB 52% · 256K ctx · ~43 tok/s
Runs well 70/100

Gemma 4 E2B

5d ago

Google · 5B · Gemma

Gemma 4 efficient base model (official)

3.1 GB 52% · 256K ctx · ~43 tok/s
Runs well 70/100

Mistral 7B v0.3

1 year ago

Mistral AI · 7B · Apache 2.0

High-quality 7B with sliding window attention

4.1 GB 68% · 32K ctx · ~33 tok/s
Decent 57/100

Qwen 2.5 7B

1 year ago

Alibaba · 7B · Apache 2.0

Strong multilingual and coding capabilities

4.1 GB 68% · 128K ctx · ~33 tok/s
Decent 57/100

Qwen 2.5 Coder 7B

1 year ago

Alibaba · 7B · Apache 2.0

Dedicated coding model

4.1 GB 68% · 128K ctx · ~33 tok/s
Decent 57/100

DeepSeek R1 Distill 7B

1 year ago

DeepSeek · 7B · MIT

R1 reasoning distilled into Qwen 7B

4.1 GB 68% · 64K ctx · ~33 tok/s
Decent 57/100

Gemma 4 E4B IT

5d ago

Google · 8B · Gemma

Gemma 4 balanced instruct model (official)

4.6 GB 77% · 256K ctx · ~29 tok/s
Tight fit 51/100

Gemma 4 E4B

5d ago

Google · 8B · Gemma

Gemma 4 balanced base model (official)

4.6 GB 77% · 256K ctx · ~29 tok/s
Tight fit 51/100

Qwen 3 8B

1 year ago

Alibaba · 8B · Apache 2.0

Qwen 3 with thinking mode support

4.6 GB 77% · 128K ctx · ~29 tok/s
Tight fit 51/100

Ministral 8B

1 year ago

Mistral AI · 8B · MRL

Mistral's efficient 8B model

4.6 GB 77% · 32K ctx · ~29 tok/s
Tight fit 51/100

Gemma 2 9B

1 year ago

Google · 9B · Gemma

Google's best mid-size open model

5.1 GB 85% · 8K ctx · ~26 tok/s
Tight fit 46/100

GLM-4 9B

1 year ago

Zhipu AI · 9B · GLM-4

Multilingual model supporting 26 languages with 128K context

5.1 GB 85% · 128K ctx · ~26 tok/s
Tight fit 46/100

Nemotron Nano 9B v2

10mo ago

NVIDIA · 9B · NVIDIA Open

Hybrid Mamba2 architecture for reasoning

5.1 GB 85% · 128K ctx · ~26 tok/s
Tight fit 46/100

Llama 3.2 11B Vision

1 year ago

Meta · 11B · Llama 3.2 Community

Multimodal vision and text model

6.1 GB 102% · 128K ctx · ~18 tok/s
Barely runs 26/100

Gemma 3 12B

1 year ago

Google · 12B · Gemma

Multimodal Gemma with 128K context

6.6 GB 110% · 128K ctx · ~14 tok/s
Barely runs 23/100

Mistral Nemo 12B

1 year ago

Mistral AI · 12B · Apache 2.0

Multilingual 12B with 128K context

6.6 GB 110% · 128K ctx · ~14 tok/s
Barely runs 23/100

Qwen 2.5 14B

1 year ago

Alibaba · 14B · Apache 2.0

Excellent quality for its size class

7.7 GB 128% ↗ RAM · 128K ctx · ~9 tok/s
Barely runs 15/100

Qwen 3 14B

1 year ago

Alibaba · 14B · Apache 2.0

Strong all-rounder with thinking mode

7.7 GB 128% ↗ RAM · 128K ctx · ~9 tok/s
Barely runs 15/100

DeepSeek R1 Distill 14B

1 year ago

DeepSeek · 14B · MIT

R1 reasoning distilled into Qwen 14B

7.7 GB 128% ↗ RAM · 64K ctx · ~9 tok/s
Barely runs 15/100

LFM2 24B

5mo ago

Liquid AI · 24B · Liquid AI

Hybrid MoE with convolution+attention layers — 2.3B active

12.8 GB 213% ↗ RAM · 32K ctx · ~4 tok/s
Barely runs 10/100

Devstral Small 2 24B

4mo ago

Mistral AI · 24B · Apache 2.0

Coding-focused model with 256K context — 68% SWE-bench

12.8 GB 213% ↗ RAM · 256K ctx · ~4 tok/s
Barely runs 10/100

Gemma 2 27B

1 year ago

Google · 27B · Gemma

Google's largest Gemma 2 model

14.3 GB 238% ↗ RAM · 8K ctx · ~3 tok/s
Barely runs 9/100

Gemma 4 26B-A4B IT

5d ago

Google · 27B · Gemma

Gemma 4 MoE instruct model (official)

14.3 GB 238% ↗ RAM · 256K ctx · ~3 tok/s
Barely runs 9/100

Gemma 4 26B-A4B

5d ago

Google · 27B · Gemma

Gemma 4 MoE base model (official)

14.3 GB 238% ↗ RAM · 256K ctx · ~3 tok/s
Barely runs 9/100

Qwen 3.5 27B

2mo ago

Alibaba · 27.8B · Apache 2.0

Flagship native multimodal Qwen 3.5

14.7 GB 245% ↗ RAM · 256K ctx · ~3 tok/s
Barely runs 9/100

Qwen 3 30B-A3B

1 year ago

Alibaba · 30B · Apache 2.0

MoE with only 3.3B active — extremely efficient

15.9 GB 265% ↗ RAM · 128K ctx · ~3 tok/s
Barely runs 9/100

Nemotron 3 Nano 30B

10mo ago

NVIDIA · 30B · NVIDIA Open

MoE with 1M context and 3B active

15.9 GB 265% ↗ RAM · 1024K ctx · ~3 tok/s
Barely runs 9/100

Qwen 2.5 32B

1 year ago

Alibaba · 32B · Apache 2.0

High-quality reasoning and multilingual

16.9 GB 282% ↗ RAM · 128K ctx · ~2 tok/s
Barely runs 8/100

EXAONE 4.0 32B

9mo ago

LG AI · 32B · EXAONE AI

Hybrid reasoning, multilingual

16.9 GB 282% ↗ RAM · 128K ctx · ~2 tok/s
Barely runs 8/100

OLMo 2 32B

1 year ago

Allen AI · 32B · Apache 2.0

Fully open research model by Allen AI

16.9 GB 282% ↗ RAM · 4K ctx · ~2 tok/s
Barely runs 8/100

Gemma 4 31B IT

5d ago

Google · 33B · Gemma

Gemma 4 flagship instruct model (official)

17.4 GB 290% · 256K ctx · 0 tok/s
Too heavy 0/100

Gemma 4 31B

5d ago

Google · 33B · Gemma

Gemma 4 flagship base model (official)

17.4 GB 290% · 256K ctx · 0 tok/s
Too heavy 0/100

Command R 35B

2y ago

Cohere · 35B · CC BY-NC 4.0

Optimized for retrieval-augmented generation

18.4 GB 307% · 128K ctx · 0 tok/s
Too heavy 0/100

Qwen 3.5 35B-A3B

2mo ago

Alibaba · 35B · Apache 2.0

Efficient multimodal MoE with 3B active

18.4 GB 307% · 256K ctx · 0 tok/s
Too heavy 0/100

Mixtral 8x7B

2y ago

Mistral AI · 47B · Apache 2.0

MoE with 12.9B active params

24.6 GB 410% · 32K ctx · 0 tok/s
Too heavy 0/100

Qwen 2.5 72B

1 year ago

Alibaba · 72B · Qwen

Alibaba's flagship open model

37.4 GB 623% · 128K ctx · 0 tok/s
Too heavy 0/100

Qwen 3.5 122B-A10B

2mo ago

Alibaba · 122B · Apache 2.0

Large multimodal MoE

63 GB 1050% · 256K ctx · 0 tok/s
Too heavy 0/100

Mixtral 8x22B

2y ago

Mistral AI · 141B · Apache 2.0

Large MoE with 39B active params

72.7 GB 1212% · 64K ctx · 0 tok/s
Too heavy 0/100

Qwen 3 235B-A22B

1 year ago

Alibaba · 235B · Apache 2.0

Massive MoE with 22B active — frontier quality

120.9 GB 2015% · 128K ctx · 0 tok/s
Too heavy 0/100

Qwen 3.5 397B-A17B

2mo ago

Alibaba · 397B · Apache 2.0

Largest multimodal Qwen 3.5 MoE

203.9 GB 3398% · 256K ctx · 0 tok/s
Too heavy 0/100

Llama 4 Maverick 17B-128E

1 year ago

Meta · 400B · Llama 4 Community

Multimodal MoE with 128 experts — 17B active, 1M context

205.4 GB 3423% · 1024K ctx · 0 tok/s
Too heavy 0/100

Llama 3.1 405B

1 year ago

Meta · 405B · Llama 3.1 Community

Largest open-weight dense model by Meta

208 GB 3467% · 128K ctx · 0 tok/s
Too heavy 0/100

Qwen 3 Coder 480B

9mo ago

Alibaba · 480B · Apache 2.0

Largest open coding MoE — 35B active

246.4 GB 4107% · 256K ctx · 0 tok/s
Too heavy 0/100

DeepSeek V3.1

8mo ago

DeepSeek · 671B · MIT

Improved V3 with hybrid thinking and tool use

344.2 GB 5737% · 128K ctx · 0 tok/s
Too heavy 0/100
6 GB ✱