Compare the models
The chat on this site offers four models that run on your graphics chip. They differ mainly in how much you download once, how much memory they need, and how quickly the answers come out.
Qwen2.5 0.5B Instruct
One-time download
Memory needed
Speed
Best for
Llama 3.2 1B Instruct
One-time download
Memory needed
Speed
Best for
Qwen2.5 1.5B Instruct
One-time download
Memory needed
Speed
Best for
Phi-4 mini Instruct
One-time download
Memory needed
Speed
Best for
There is no honest single number here: the same model can be several times faster on a desktop graphics card than on a tablet. The ordering above follows model size, which is what actually decides the difference on any one device. The one figure measured on this site is for processor-only mode, where the smallest model produces roughly two and a half words a second.
How to choose
Older iPad or laptop
An everyday laptop from the last few years
Desktop with a good graphics card
No graphics acceleration?
If your browser cannot reach the graphics chip, the chat quietly switches to a much smaller model that runs on the processor. Those two are separate from the four above:
SmolLM2 360M Instruct (processor only)
Qwen2.5 0.5B Instruct (processor only)