The quick answer, by Mac
| Your Mac | For writing | For code |
|---|---|---|
| 8 GB | Llama 3.2 1B (770 MB), or Qwen 2.5 0.5B (379 MB) on a busy machine | Qwen 2.5 Coder 1.5B (940 MB) |
| 16 GB | Qwen 2.5 3B (1.9 GB), the default | Qwen 2.5 Coder 1.5B |
| 24–32 GB | Qwen 2.5 3B, or Gemma 3 4B for the most natural phrasing | Qwen 2.5 Coder 1.5B |
| 32 GB or more, Pro or Max chip | Any of the above | Qwen 2.5 Coder 7B (4.5 GB) |
Check your memory in Apple menu › About This Mac. Sizes are the download sizes of the models in WeType.
Memory comes first
A local model has to fit in memory alongside everything else you have open: the browser, the chat app, the editor. On Apple Silicon the model shares unified memory with the system, which is what makes running one on a laptop possible, but it also means a model that is too large for your Mac competes with your other apps and slows everything down, including its own suggestions.
As a rule of thumb, a model uses a little more memory than its file size while it runs. That is why WeType’s model picker labels each model with the Mac it is comfortable on, and why a smaller model that answers instantly beats a bigger one that answers after you have already typed the word.
Base models vs chat models
The models behind chat assistants are instruction-tuned: after learning language, they were trained further to follow requests and answer questions. That extra training is exactly wrong for autocomplete. Given “Thanks for the update, I’ll”, a chat model is inclined to respond to you; a base model, trained only to continue text, simply continues it.
Base models are also better at the fiddliest part of autocomplete: finishing a word you are halfway through typing. That is why seven of the nine models in WeType are base or code-base models, and why the two Gemma 3 models, which are only published in instruction-tuned form, say so in the picker.
Why size decides whether it reads the conversation
The most useful thing an autocomplete model can do is read the message you are replying to and suggest something that answers it. Small models cannot do this well. In WeType’s testing, models under about 1.5 GB tend to copy phrases from the surrounding text into your sentence instead of responding to it, which makes suggestions worse rather than better.
So WeType only offers surrounding context to models of 1.5 GB or more: Gemma 2B, the two 3B models, Gemma 3 4B and Qwen 2.5 Coder 7B. The smaller models still continue your own sentence well; they just do not look beyond it.
Every model WeType offers
| Model | Type | Download | Comfortable on | Reads context | Notes |
|---|---|---|---|---|---|
| Qwen 2.5 3Bdefault | base | 1.9 GB | 16 GB | Yes | The default. Big enough to read the conversation around the field, which is what makes a suggestion fit rather than merely read well. |
| Llama 3.2 3B | base | 1.9 GB | 16 GB | Yes | Same tier as the default, if you would rather run Meta's model. |
| Gemma 2B | base | 1.6 GB | 16 GB | Yes | Google's base model. A solid middle tier. |
| Llama 3.2 1B | base | 770 MB | 8 GB | No | Half the memory and the quickest to finish. Too small to use surrounding context. |
| Qwen 2.5 0.5B | base | 379 MB | 8 GB | No | The lightest option, for older or low-memory Macs. |
| Gemma 3 1B | instruction-tuned | 778 MB | 8 GB | No | The newest small Gemma. Instruction-tuned only, so mid-word completion is less precise, and too small to use surrounding context. |
| Gemma 3 4B | instruction-tuned | 3.2 GB | 16–32 GB | Yes | The sharpest prose predictions, though instruction-tuned, so mid-word completion is less precise. |
| Qwen 2.5 Coder 1.5B | code | 940 MB | 8 GB | No | Code-tuned: understands syntax and fills in the middle of code. |
| Qwen 2.5 Coder 7B | code | 4.5 GB | 32 GB+ | Yes | A much stronger code model, for powerful Macs. |
Models are downloaded inside the app from the Ollama registry and run with llama.cpp. Code models understand syntax and can fill in the middle of code. In code editors, WeType also switches to a code-completion mode.
Trying and switching models
Open Settings › Models
Each model shows its size, the Mac it suits, whether it reads context, and its pros and cons.Download one and make it active
Downloads run in the background. You can keep several installed and switch between them.Judge it on your own writing for a day
Quality differences show up in real replies, not in a single test sentence. If suggestions arrive late, step down a size; if they feel generic, step up.Delete the ones you do not use
Models are the large part of WeType on disk. The uninstall page shows where they live.
Not sure your Mac qualifies at all? WeType needs Apple Silicon and macOS 14 or later; the compatibility page has the details.