Learn
AI autocomplete, explained plainly.
Guides to how AI autocomplete and sentence completion work, what running a model on your own Mac involves, and how to type less. Each one starts with the short answer.
New to the idea? Start with what AI autocomplete is. Weighing up apps? The comparisons are separate, and say who wrote them.
What is AI autocomplete, and how does it work?
AI autocomplete is a typing aid that uses a language model to predict the next words of whatever you are writing and offers them inline, usually as faint grey “ghost text” after the cursor. You press a key such as Tab to accept, or keep typing to ignore it. Traditional autocomplete matches your input against a fixed list of words or past entries; AI autocomplete generates a new continuation from the context of what you have already written.
Read the 8-minute answerAI sentence completion: what a sentence completer does, and how to use one
An AI sentence completer is a tool that uses a language model to finish a sentence you have started. Web-based sentence finishers take pasted text and return several possible endings. Inline sentence completers work inside the app you are typing in: they show the likely rest of the sentence as grey text at the cursor, and a key press accepts it. For everyday email, chat and notes, the inline kind is faster because you never leave what you are writing.
Read the 7-minute answerOn-device vs cloud AI writing tools: privacy, speed and cost
An on-device AI writing tool runs its language model on your own computer, so the text you type is processed locally and never uploaded. A cloud tool sends your text to a server, where a much larger model generates the result. Cloud models are more capable at open-ended writing; on-device models are smaller but work offline, have no per-use cost, and keep sensitive text on your machine. For inline autocomplete, which only needs to predict a few words, a small local model is usually enough.
Read the 7-minute answerChoosing a local AI model for autocomplete on your Mac
For AI autocomplete on a Mac, choose by memory first. On an 8 GB Mac, run a model under about 1 GB, such as Llama 3.2 1B or Qwen 2.5 0.5B. On 16 GB, a 3B base model such as Qwen 2.5 3B is the best balance and can use the conversation around the field. For code, use a code-tuned model: Qwen 2.5 Coder 1.5B on most Macs, or Coder 7B on a Mac with 32 GB or more. Prefer base models over chat (instruction-tuned) models for completion.
Read the 7-minute answerHow to type faster on a Mac: built-in tricks, plus AI autocomplete
To type faster on a Mac, cut keystrokes rather than only moving your fingers faster: set up text replacements for phrases you repeat, edit by word with Option-arrow and Option-Delete, jump by line with Command-arrow, dictate long passages, raise the key repeat rate, and use inline predictions or an AI autocomplete app to finish predictable sentences with a single key.
Read the 6-minute answermacOS inline predictive text: how it works, and how to turn it off
macOS inline predictive text shows a suggested word or short phrase in grey after your cursor as you type; press Space to accept it. To turn it off, open System Settings › Keyboard, click Edit next to Input Sources, and switch off “Show inline predictive text”. Turn it back on the same way. Turn it off if it keeps catching your spacebar, or if you use another autocomplete app, so two sets of grey text do not overlap.
Read the 4-minute answerAI autocomplete glossary
This glossary defines the terms used when talking about AI autocomplete and on-device language models, from the everyday (ghost text, predictive text, word-by-word acceptance) to the technical (tokens, base versus instruction-tuned models, context windows, KV caches, quantization, GGUF and llama.cpp).
Read the 8-minute answer
Rather see it than read about it? The trial asks for nothing.
5 days free · 100 words a day · no card, no account