Skip to content
WeType

Choosing a local AI model for autocomplete on your Mac

The model decides how good suggestions are, how quickly they appear and how much memory they use. Here is how to choose one for the Mac you actually have, using the models WeType offers as the worked example.

By Published 7 min read

Short answer

For AI autocomplete on a Mac, choose by memory first. On an 8 GB Mac, run a model under about 1 GB, such as Llama 3.2 1B or Qwen 2.5 0.5B. On 16 GB, a 3B base model such as Qwen 2.5 3B is the best balance and can use the conversation around the field. For code, use a code-tuned model: Qwen 2.5 Coder 1.5B on most Macs, or Coder 7B on a Mac with 32 GB or more. Prefer base models over chat (instruction-tuned) models for completion.

The quick answer, by Mac

A starting point for AI autocomplete, by how much memory your Mac has
Your MacFor writingFor code
8 GBLlama 3.2 1B (770 MB), or Qwen 2.5 0.5B (379 MB) on a busy machineQwen 2.5 Coder 1.5B (940 MB)
16 GBQwen 2.5 3B (1.9 GB), the defaultQwen 2.5 Coder 1.5B
24–32 GBQwen 2.5 3B, or Gemma 3 4B for the most natural phrasingQwen 2.5 Coder 1.5B
32 GB or more, Pro or Max chipAny of the aboveQwen 2.5 Coder 7B (4.5 GB)

Check your memory in Apple menu › About This Mac. Sizes are the download sizes of the models in WeType.

Memory comes first

A local model has to fit in memory alongside everything else you have open: the browser, the chat app, the editor. On Apple Silicon the model shares unified memory with the system, which is what makes running one on a laptop possible, but it also means a model that is too large for your Mac competes with your other apps and slows everything down, including its own suggestions.

As a rule of thumb, a model uses a little more memory than its file size while it runs. That is why WeType’s model picker labels each model with the Mac it is comfortable on, and why a smaller model that answers instantly beats a bigger one that answers after you have already typed the word.

Base models vs chat models

The models behind chat assistants are instruction-tuned: after learning language, they were trained further to follow requests and answer questions. That extra training is exactly wrong for autocomplete. Given “Thanks for the update, I’ll”, a chat model is inclined to respond to you; a base model, trained only to continue text, simply continues it.

Base models are also better at the fiddliest part of autocomplete: finishing a word you are halfway through typing. That is why seven of the nine models in WeType are base or code-base models, and why the two Gemma 3 models, which are only published in instruction-tuned form, say so in the picker.

Why size decides whether it reads the conversation

The most useful thing an autocomplete model can do is read the message you are replying to and suggest something that answers it. Small models cannot do this well. In WeType’s testing, models under about 1.5 GB tend to copy phrases from the surrounding text into your sentence instead of responding to it, which makes suggestions worse rather than better.

So WeType only offers surrounding context to models of 1.5 GB or more: Gemma 2B, the two 3B models, Gemma 3 4B and Qwen 2.5 Coder 7B. The smaller models still continue your own sentence well; they just do not look beyond it.

Every model WeType offers

All 9 models in WeType, from the in-app catalogue
ModelTypeDownloadComfortable onReads contextNotes
Qwen 2.5 3Bdefaultbase1.9 GB16 GBYesThe default. Big enough to read the conversation around the field, which is what makes a suggestion fit rather than merely read well.
Llama 3.2 3Bbase1.9 GB16 GBYesSame tier as the default, if you would rather run Meta's model.
Gemma 2Bbase1.6 GB16 GBYesGoogle's base model. A solid middle tier.
Llama 3.2 1Bbase770 MB8 GBNoHalf the memory and the quickest to finish. Too small to use surrounding context.
Qwen 2.5 0.5Bbase379 MB8 GBNoThe lightest option, for older or low-memory Macs.
Gemma 3 1Binstruction-tuned778 MB8 GBNoThe newest small Gemma. Instruction-tuned only, so mid-word completion is less precise, and too small to use surrounding context.
Gemma 3 4Binstruction-tuned3.2 GB16–32 GBYesThe sharpest prose predictions, though instruction-tuned, so mid-word completion is less precise.
Qwen 2.5 Coder 1.5Bcode940 MB8 GBNoCode-tuned: understands syntax and fills in the middle of code.
Qwen 2.5 Coder 7Bcode4.5 GB32 GB+YesA much stronger code model, for powerful Macs.

Models are downloaded inside the app from the Ollama registry and run with llama.cpp. Code models understand syntax and can fill in the middle of code. In code editors, WeType also switches to a code-completion mode.

Trying and switching models

  1. Open Settings › Models

    Each model shows its size, the Mac it suits, whether it reads context, and its pros and cons.
  2. Download one and make it active

    Downloads run in the background. You can keep several installed and switch between them.
  3. Judge it on your own writing for a day

    Quality differences show up in real replies, not in a single test sentence. If suggestions arrive late, step down a size; if they feel generic, step up.
  4. Delete the ones you do not use

    Models are the large part of WeType on disk. The uninstall page shows where they live.

Not sure your Mac qualifies at all? WeType needs Apple Silicon and macOS 14 or later; the compatibility page has the details.

Questions

Can I run AI autocomplete on an 8 GB Mac?
Yes, with a small model. Llama 3.2 1B (770 MB) and Qwen 2.5 0.5B (379 MB) run comfortably on 8 GB. They continue your sentence well but are too small to take the surrounding conversation into account.
Why is a base model better than a chat model for autocomplete?
A base model is trained to continue text, which is exactly what autocomplete asks of it. A chat (instruction-tuned) model is trained to answer, so it tends to respond to your sentence rather than finish it, and it completes half-typed words less precisely.
Which model does WeType use by default?
Qwen 2.5 3B (base), about 1.9 GB. It is the smallest base model that is large enough to use the text around the field, and it is best on a Mac with 16 GB of memory.

Keep reading

Try AI autocomplete in every app on your Mac.

5 days free · 100 words a day · no card, no account

Try WeType free