Explore AI
No GPU at all
Possible, and slower than you would like — but not useless.
Possible, and slower than you would like — but not useless.
- Platform
- Any
- Memory
- 16 GB ram
llama.cpp on CPU can run a small model at a few tokens per second. Reading speed is roughly 5–10 tokens per second, so it feels like waiting for someone to type.
Runs comfortably
- Qwen3 1.7B 3 GB at Q4 Apache-2.0
- Llama 3.2 1B Instruct 2 GB at Q4 Llama 3.2 Community
- Qwen3 4B 5 GB at Q4 Apache-2.0
Will not fit
- Qwen3 8Bneeds about 8 GB
- Qwen3 32Bneeds about 24 GB
Worth knowing
Good enough for background jobs — transcribing, summarising a folder overnight. Not good enough for conversation.