Using autocomplete
- Autocomplete
- A feature that suggests how to finish what you are typing, so you can accept the suggestion instead of typing it out. Classic autocomplete completes a word from a list; AI autocomplete generates a continuation.
- AI autocomplete
- Autocomplete powered by a language model, which predicts a phrase or the rest of a sentence from the text around the cursor rather than looking words up in a list. How AI autocomplete works.
- Predictive text
- The general name for keyboard features that guess your next word. It predates language models; AI autocomplete is a more capable form of it.
- Sentence completion
- Finishing a partly written sentence. An AI sentence completer does this with a language model, either on a web page or inline as you type. AI sentence completion.
- Ghost text
- A suggestion drawn in faint grey directly after the cursor, on the same line as your text. It is not part of the document until you accept it.
- Inline completion
- A suggestion shown inside the text you are editing, as ghost text, rather than in a pop-up list or a separate panel.
- Word-by-word acceptance
- Accepting a suggestion one word at a time, so you can keep its beginning and write your own ending. In WeType, Tab takes one word.
- Caret
- The blinking text cursor that marks where typing will appear. Autocomplete suggestions are anchored to it.
- Text replacement (text expansion)
- A fixed shortcut that expands into a saved phrase, such as ;addr becoming your address. Unlike autocomplete, it never guesses; it only does what you set up. Typing faster on a Mac.
Models
- Language model (LLM)
- A program trained on large amounts of text to predict what text comes next. “Large language model” usually refers to the biggest ones, though the same idea scales down to models that fit on a laptop.
- Small language model
- A language model small enough to run on a personal computer or phone, typically under about 8 billion parameters. Less capable than large cloud models, but fast and private.
- Parameters
- The learned numbers inside a model. “3B” means about three billion of them. More parameters generally means more capability and more memory.
- Token
- The unit a model reads and writes: a whole word, part of a word, or punctuation. Models predict one token at a time.
- Next-token prediction
- What a language model fundamentally does: given some text, estimate which token is most likely to come next. Repeating it produces a continuation.
- Base model
- A model trained only to continue text. It is the best kind for autocomplete, because continuing your text is exactly the task.
- Instruction-tuned model
- A base model trained further to follow requests and hold conversations, as chat assistants do. It tends to answer your text rather than continue it. Base vs chat models.
- Code model
- A model trained heavily on source code, so it understands syntax and programming conventions. Better for code, usually weaker for prose.
- Fill-in-the-middle (FIM)
- A way of prompting a model with the text both before and after the cursor, so it can fill a gap rather than only continue from the end.
- Context window
- The maximum amount of text a model can take into account at once, measured in tokens.
- Surrounding context
- Text near the field you are typing in, such as the message you are replying to, given to the model so its suggestion fits the conversation.
- Token healing
- A technique that backs up to the start of a half-typed word and lets the model regenerate it, so completing “recomm” yields “recommend” rather than an awkward split.
Running a model locally
- On-device inference
- Running a model on your own computer instead of a server. The text it processes does not need to leave the machine. On-device vs cloud AI.
- Model weights
- The file containing a model's parameters. For a local model, this is the large download.
- Quantization
- Storing a model's numbers with fewer bits so the file is smaller and runs faster, at a small cost in quality. Q4_K_M is a common four-bit scheme.
- GGUF
- A file format for storing (often quantized) model weights, designed for llama.cpp and widely used for running models locally.
- llama.cpp
- An open-source engine for running language models efficiently on ordinary hardware, including Apple Silicon Macs. WeType uses it.
- KV cache (prompt cache)
- Saved intermediate results for text the model has already processed. Reusing it means each new keystroke only processes the new text, which keeps suggestions fast.
- Latency
- How long you wait between pausing and seeing a suggestion. For autocomplete it matters more than raw model quality.
On a Mac
- Apple Silicon
- Apple's own Mac processors (M1 and later). Their GPU and unified memory make running local models practical on a laptop.
- Unified memory
- Memory shared by the CPU and GPU on Apple Silicon, so a model loaded once is available to both without copying.
- Metal
- Apple's graphics and compute framework, which llama.cpp uses to run models on the Mac's GPU.
- Accessibility API
- The macOS interface that lets assistive software read and edit text in other apps. System-wide autocomplete uses it to read the focused field and insert what you accept.
- Input Monitoring
- A macOS privacy permission that lets an app observe keystrokes. Autocomplete needs it to notice your typing and the accept key.
- Inline predictive text (macOS)
- macOS's built-in grey predictions, introduced in Sonoma, accepted with Space. Turning it on or off.