A definition
AI autocomplete is a typing aid that predicts what you are about to write and offers it before you write it. A language model reads the text before your cursor, works out the most likely continuation, and shows it in place, usually as faint grey text after the cursor. Accept it with a key, or ignore it by typing on.
The idea is older than AI. Search boxes, phone keyboards and code editors have completed words for decades. What is new is the length and the fit of the suggestion: a language model can finish the whole sentence, and finish it in a way that follows from the one before.
Can we move Thursday's review to next week?
No problem at all. Next week works better for me too, how about Tuesday afternoon?
Traditional vs AI autocomplete
Traditional autocomplete looks things up. It compares what you have typed with a list, a dictionary, your contacts, past searches or the names in a codebase, and offers matches. It is fast and predictable, and it can only ever suggest something that is already on the list. AI autocomplete generates. It has no list; it has a model of how text tends to continue, so it can produce a phrase nobody has typed before.
| Traditional autocomplete | AI autocomplete | |
|---|---|---|
| Where suggestions come from | A fixed list or your history | A language model generating text |
| Typical length | The rest of one word | A phrase or the rest of the sentence |
| Uses the context | Only the current word | The sentence, and often the text around it |
| Can suggest something new | No | Yes |
| Can be wrong in fluent ways | Rarely | Yes; it predicts what is likely, not what is true |
| Examples | Search suggestions, contact names, IDE symbol lists | Gmail Smart Compose, GitHub Copilot, WeType |
How AI autocomplete works, step by step
Every AI autocomplete tool, whether it runs in the cloud or on your computer, goes through roughly the same loop each time you pause.
It reads the context
It collects the text before the cursor, and depending on the tool, the text after it, the rest of the document, or the message you are replying to. On a Mac, a system-wide tool reads this through the accessibility layer, the same interface screen readers use.It builds a prompt
That text is arranged into the input the model expects. Some tools add instructions (a language, a tone) or examples of your own writing so the continuation sounds like you.The model predicts one token at a time
A language model does not write sentences; it predicts the next token, a word or part of a word, from probabilities learned in training. It appends the most likely one and predicts again, until it reaches a natural stopping point such as the end of the sentence.The suggestion is shown in place
The result is drawn after your cursor in a muted colour so it is clearly not yet part of your text. This is called ghost text.You accept, partly accept, or ignore it
A key such as Tab accepts. Better tools let you take one word at a time, so you can keep the start of a suggestion and write your own ending. Typing something different discards it and the loop starts again.
Speed matters more than anything else here. A suggestion that arrives after you have typed the next word is useless, so autocomplete uses small, fast models and tricks such as caching the part of the prompt that has not changed, so each new keystroke only processes the new text.
Why the kind of model matters
Language models come in two broad flavours, and autocomplete wants the less famous one. A base model is trained only to continue text. An instruction-tuned model, the kind behind chat assistants, is further trained to answer requests. Ask a chat model to continue “I was hoping we could” and it may reply to you rather than finish your sentence; a base model simply continues it. Base models are also better at finishing a half-typed word, because they have not been trained to treat your text as a question.
Size matters too. Very small models continue a sentence well but cannot make use of a long context, such as the email you are answering. Larger models can, at the cost of memory. The guide to choosing a local model covers that trade-off in detail.
Where you have already seen it
- Email. Gmail’s Smart Compose suggests the rest of a sentence in grey as you write.
- Code editors. Tools such as GitHub Copilot complete lines and whole functions inline.
- Phone and Mac keyboards. iOS and macOS offer predictions as you type; since macOS Sonoma they can appear inline in grey (how that works).
- Everywhere at once. System-wide apps put the same idea into every text field on a computer, rather than inside one product.
Is it good for your writing?
It depends on what you are writing. AI autocomplete is at its best on text that is predictable: replies, confirmations, status updates, the second half of a sentence whose first half already says where it is going. It saves the least on original thinking, because a model cannot know what you have not decided yet.
Getting AI autocomplete on a Mac
You have three routes, and they are not exclusive.
- macOS’s built-in inline predictions. Free, already there, a word or short phrase at a time.
- Features inside individual apps, such as an email service’s compose suggestions or a coding assistant in your editor.
- A system-wide AI autocomplete app that works in every text field. WeType is one: it runs the language model on your Mac, shows the rest of your sentence as ghost text, and Tab takes it a word at a time. Here is how the options compare.