Skip to content
WeType

AI autocomplete glossary

The vocabulary of AI autocomplete, defined in a sentence or two each, in plain English. Each term has its own link, so you can point someone straight at one.

By Published 8 min read

Short answer

This glossary defines the terms used when talking about AI autocomplete and on-device language models, from the everyday (ghost text, predictive text, word-by-word acceptance) to the technical (tokens, base versus instruction-tuned models, context windows, KV caches, quantization, GGUF and llama.cpp).

Using autocomplete

Autocomplete
A feature that suggests how to finish what you are typing, so you can accept the suggestion instead of typing it out. Classic autocomplete completes a word from a list; AI autocomplete generates a continuation.
AI autocomplete
Autocomplete powered by a language model, which predicts a phrase or the rest of a sentence from the text around the cursor rather than looking words up in a list. How AI autocomplete works.
Predictive text
The general name for keyboard features that guess your next word. It predates language models; AI autocomplete is a more capable form of it.
Sentence completion
Finishing a partly written sentence. An AI sentence completer does this with a language model, either on a web page or inline as you type. AI sentence completion.
Ghost text
A suggestion drawn in faint grey directly after the cursor, on the same line as your text. It is not part of the document until you accept it.
Inline completion
A suggestion shown inside the text you are editing, as ghost text, rather than in a pop-up list or a separate panel.
Word-by-word acceptance
Accepting a suggestion one word at a time, so you can keep its beginning and write your own ending. In WeType, Tab takes one word.
Caret
The blinking text cursor that marks where typing will appear. Autocomplete suggestions are anchored to it.
Text replacement (text expansion)
A fixed shortcut that expands into a saved phrase, such as ;addr becoming your address. Unlike autocomplete, it never guesses; it only does what you set up. Typing faster on a Mac.

Models

Language model (LLM)
A program trained on large amounts of text to predict what text comes next. “Large language model” usually refers to the biggest ones, though the same idea scales down to models that fit on a laptop.
Small language model
A language model small enough to run on a personal computer or phone, typically under about 8 billion parameters. Less capable than large cloud models, but fast and private.
Parameters
The learned numbers inside a model. “3B” means about three billion of them. More parameters generally means more capability and more memory.
Token
The unit a model reads and writes: a whole word, part of a word, or punctuation. Models predict one token at a time.
Next-token prediction
What a language model fundamentally does: given some text, estimate which token is most likely to come next. Repeating it produces a continuation.
Base model
A model trained only to continue text. It is the best kind for autocomplete, because continuing your text is exactly the task.
Instruction-tuned model
A base model trained further to follow requests and hold conversations, as chat assistants do. It tends to answer your text rather than continue it. Base vs chat models.
Code model
A model trained heavily on source code, so it understands syntax and programming conventions. Better for code, usually weaker for prose.
Fill-in-the-middle (FIM)
A way of prompting a model with the text both before and after the cursor, so it can fill a gap rather than only continue from the end.
Context window
The maximum amount of text a model can take into account at once, measured in tokens.
Surrounding context
Text near the field you are typing in, such as the message you are replying to, given to the model so its suggestion fits the conversation.
Token healing
A technique that backs up to the start of a half-typed word and lets the model regenerate it, so completing “recomm” yields “recommend” rather than an awkward split.

Running a model locally

On-device inference
Running a model on your own computer instead of a server. The text it processes does not need to leave the machine. On-device vs cloud AI.
Model weights
The file containing a model's parameters. For a local model, this is the large download.
Quantization
Storing a model's numbers with fewer bits so the file is smaller and runs faster, at a small cost in quality. Q4_K_M is a common four-bit scheme.
GGUF
A file format for storing (often quantized) model weights, designed for llama.cpp and widely used for running models locally.
llama.cpp
An open-source engine for running language models efficiently on ordinary hardware, including Apple Silicon Macs. WeType uses it.
KV cache (prompt cache)
Saved intermediate results for text the model has already processed. Reusing it means each new keystroke only processes the new text, which keeps suggestions fast.
Latency
How long you wait between pausing and seeing a suggestion. For autocomplete it matters more than raw model quality.

On a Mac

Apple Silicon
Apple's own Mac processors (M1 and later). Their GPU and unified memory make running local models practical on a laptop.
Unified memory
Memory shared by the CPU and GPU on Apple Silicon, so a model loaded once is available to both without copying.
Metal
Apple's graphics and compute framework, which llama.cpp uses to run models on the Mac's GPU.
Accessibility API
The macOS interface that lets assistive software read and edit text in other apps. System-wide autocomplete uses it to read the focused field and insert what you accept.
Input Monitoring
A macOS privacy permission that lets an app observe keystrokes. Autocomplete needs it to notice your typing and the accept key.
Inline predictive text (macOS)
macOS's built-in grey predictions, introduced in Sonoma, accepted with Space. Turning it on or off.

Keep reading

Try AI autocomplete in every app on your Mac.

5 days free · 100 words a day · no card, no account

Try WeType free