Skip to content
WeType

On-device vs cloud AI writing tools: privacy, speed and cost

“Private AI” is on every product page now. The useful question is narrower: when you type, does the text leave your machine? Here is what on-device actually means, what it costs you, and how to verify it rather than trust it.

By Published 7 min read

Short answer

An on-device AI writing tool runs its language model on your own computer, so the text you type is processed locally and never uploaded. A cloud tool sends your text to a server, where a much larger model generates the result. Cloud models are more capable at open-ended writing; on-device models are smaller but work offline, have no per-use cost, and keep sensitive text on your machine. For inline autocomplete, which only needs to predict a few words, a small local model is usually enough.

The difference in one table

Every AI writing tool has to run a language model somewhere. Either it runs on a server, and your text is sent there, or it runs on your computer, and your text stays put. Everything else follows from that one choice.

On-device and cloud AI writing tools, compared
On-deviceCloud
Where your text is processedOn your computerOn the provider's servers
Works offlineYes, once the model is downloadedNo
Model sizeSmall: roughly 0.5–7 billion parameters on a laptopVery large; often undisclosed
Best atContinuing text, short predictionsOpen-ended writing, reasoning, long drafts
Speed depends onYour chip and memoryYour connection and the provider's load
Cost modelUsually a one-off or flat priceSubscription or per-use, to cover servers
UsesYour disk (the model file) and memoryNothing locally beyond the app
Account neededOften notUsually

What “on-device” actually means

An on-device tool ships, or downloads, the model’s weights: a file of learned numbers, typically several hundred megabytes to a few gigabytes. When you type, software on your computer loads that file into memory and runs the model locally. On Apple Silicon Macs this runs on the GPU, which shares one pool of unified memory with the rest of the system, so a laptop can run a model that would need a dedicated graphics card elsewhere.

WeType is an example. It uses llama.cpp, an open-source engine for running language models on ordinary hardware, and models in the GGUF file format that you download inside the app. Predictions happen in the app’s own process; nothing is sent anywhere to produce one.

What you give up, and what you do not

You give up raw capability. A model that fits in a few gigabytes of laptop memory is far smaller than the models behind large cloud assistants. Ask it to plan a project or summarise a report and the difference is obvious.

For autocomplete, you give up much less. Predicting the next few words of a sentence you have already started is a narrow task that small models do well, especially base models trained purely to continue text. And the constraints of autocomplete, a suggestion that must arrive before your next keystroke, favour a model sitting next to the keyboard over one on the far side of a network connection.

You gain predictability. No rate limits, no outage when a provider has a bad day, no change in behaviour because the provider swapped models, and a price that does not depend on how much you type.

How to verify a privacy claim yourself

You do not have to take any app’s word for it, including ours.

  1. Turn the Wi-Fi off and keep typing

    If suggestions keep coming with no connection, the model is local. This is the quickest test and it is conclusive for that question.
  2. Watch its network traffic

    A per-app firewall such as Little Snitch or the free LuLu shows every connection an app makes, and to where. Activity Monitor's Network tab shows how much each process sends.
  3. Read the privacy policy for specifics

    Look for named files, named events and stated retention periods. “We may collect usage data to improve our services” tells you nothing; a table of exactly what is sent tells you something you can test.

Which to use when

  • Drafting from scratch, research, summarising: a cloud assistant is the stronger tool. Avoid pasting in anything you would not send to a third party.
  • Finishing your own sentences as you type, all day, everywhere: on-device autocomplete. It sees everything you write, which is exactly why it should not be sending it anywhere.
  • Confidential work such as legal, medical or client correspondence: prefer on-device tools, and check your organisation’s policy either way.

Many people use both, and that is sensible. For what each kind of tool is actually for, see AI autocomplete vs AI writing assistants; for picking a model to run locally, see choosing a local model.

Questions

Is on-device AI completely private?
The text you type is, if the model really runs locally: nothing needs to be sent for a prediction. An app can still make other network requests, such as downloading the model, checking for updates or validating a licence, so read what an app says it sends and check it with a network monitor.
Are on-device models as good as ChatGPT?
Not for open-ended writing. A model small enough to run on a laptop has a fraction of the parameters of a large cloud model. For predicting the next few words of a sentence you are already writing, the gap matters much less, which is why on-device autocomplete is practical.
How much memory does a local AI model need on a Mac?
It depends on the model. In WeType, the smallest models run comfortably on an 8 GB Mac, the recommended 3B model is best on 16 GB, and the largest code model wants 32 GB or more. Apple Silicon's unified memory lets the GPU use system RAM directly.

Keep reading

Try AI autocomplete in every app on your Mac.

5 days free · 100 words a day · no card, no account

Try WeType free