The difference in one table
Every AI writing tool has to run a language model somewhere. Either it runs on a server, and your text is sent there, or it runs on your computer, and your text stays put. Everything else follows from that one choice.
| On-device | Cloud | |
|---|---|---|
| Where your text is processed | On your computer | On the provider's servers |
| Works offline | Yes, once the model is downloaded | No |
| Model size | Small: roughly 0.5–7 billion parameters on a laptop | Very large; often undisclosed |
| Best at | Continuing text, short predictions | Open-ended writing, reasoning, long drafts |
| Speed depends on | Your chip and memory | Your connection and the provider's load |
| Cost model | Usually a one-off or flat price | Subscription or per-use, to cover servers |
| Uses | Your disk (the model file) and memory | Nothing locally beyond the app |
| Account needed | Often not | Usually |
What “on-device” actually means
An on-device tool ships, or downloads, the model’s weights: a file of learned numbers, typically several hundred megabytes to a few gigabytes. When you type, software on your computer loads that file into memory and runs the model locally. On Apple Silicon Macs this runs on the GPU, which shares one pool of unified memory with the rest of the system, so a laptop can run a model that would need a dedicated graphics card elsewhere.
WeType is an example. It uses llama.cpp, an open-source engine for running language models on ordinary hardware, and models in the GGUF file format that you download inside the app. Predictions happen in the app’s own process; nothing is sent anywhere to produce one.
What you give up, and what you do not
You give up raw capability. A model that fits in a few gigabytes of laptop memory is far smaller than the models behind large cloud assistants. Ask it to plan a project or summarise a report and the difference is obvious.
For autocomplete, you give up much less. Predicting the next few words of a sentence you have already started is a narrow task that small models do well, especially base models trained purely to continue text. And the constraints of autocomplete, a suggestion that must arrive before your next keystroke, favour a model sitting next to the keyboard over one on the far side of a network connection.
You gain predictability. No rate limits, no outage when a provider has a bad day, no change in behaviour because the provider swapped models, and a price that does not depend on how much you type.
How to verify a privacy claim yourself
You do not have to take any app’s word for it, including ours.
Turn the Wi-Fi off and keep typing
If suggestions keep coming with no connection, the model is local. This is the quickest test and it is conclusive for that question.Watch its network traffic
A per-app firewall such as Little Snitch or the free LuLu shows every connection an app makes, and to where. Activity Monitor's Network tab shows how much each process sends.Read the privacy policy for specifics
Look for named files, named events and stated retention periods. “We may collect usage data to improve our services” tells you nothing; a table of exactly what is sent tells you something you can test.
Which to use when
- Drafting from scratch, research, summarising: a cloud assistant is the stronger tool. Avoid pasting in anything you would not send to a third party.
- Finishing your own sentences as you type, all day, everywhere: on-device autocomplete. It sees everything you write, which is exactly why it should not be sending it anywhere.
- Confidential work such as legal, medical or client correspondence: prefer on-device tools, and check your organisation’s policy either way.
Many people use both, and that is sensible. For what each kind of tool is actually for, see AI autocomplete vs AI writing assistants; for picking a model to run locally, see choosing a local model.