Run Llama 3 in your browser
Meta's Llama 3.2 1B is one of the models you can pick in the chat on this site. It downloads once, about 880 MB, and then answers entirely on your own machine — no server, no API key, no account.
What your device needs
Llama 3.2 1B runs on your graphics chip through WebGPU, the browser standard that gives a web page access to graphics hardware. That support is what decides whether this model is available to you.
What to expect
Download
Speed
Quality
Privacy
Offline
How to try it
- Open the chat and pass the short human check.
- Choose "Llama 3.2 1B Instruct" in the model picker.
- Wait for the one-time download — on a fast connection this is a couple of minutes.
- Ask it something. Everything after that happens on your device.
The picker also offers Qwen2.5 0.5B for the fastest start, Qwen2.5 1.5B as a stronger all-rounder, and Phi-4 mini for the most capable answers on a recent desktop.