The AI glossary: 60+ terms explained in plain English

AI has a vocabulary problem. Every launch comes wrapped in terms like agentic, context window, open weights and RAG, usually explained by someone who assumes you already know. This glossary defines 62 of the terms you will meet most often, in plain English, with an example where one helps.

It is written for people who use AI at work rather than build it, though developers should find it a useful refresher. Terms link to our longer guides where we have one, and to the daily digest issue where something happened. Jump to a letter below, or search the page with Ctrl+F (Cmd+F on a Mac).

Agent
An AI system that does not just answer but acts: it plans steps, uses tools such as a browser, files or other apps, checks the results and keeps going until a task is done. Booking travel, fixing a bug across a codebase or reconciling a spreadsheet are agent jobs. Meta's Muse and OpenAI's always-on Dots are consumer examples.See also: Agentic AI, Tool use, Subagent.
Agentic AI
The broad shift from AI that talks to AI that does. When a product is described as agentic, it means the model is trusted to take a series of actions with some independence, usually with a person approving the riskier steps.
AGI (artificial general intelligence)
A hypothetical AI that can do almost any intellectual task a person can, at least as well. There is no agreed test for it, so claims are partly marketing. OpenAI used the phrase openly when it launched GPT-6 Astra in September 2026.
Alignment
The work of making AI systems do what their makers and users actually intend, including refusing harmful requests and being honest about uncertainty. A model can be very capable and poorly aligned, which is the combination safety researchers worry about.
API (application programming interface)
The way software talks to other software. AI companies sell access to their models through APIs, usually priced per million tokens, so developers can build the model into their own products.
Attention
The mechanism inside a transformer that lets a model weigh which earlier words matter most when working out the next one. It is why modern models can follow long, complicated text, and why very long prompts get expensive.
Benchmark
A standard test used to compare models, such as SWE-bench for coding or GPQA for expert-level science questions. Useful for rough comparisons, but labs choose which benchmarks to publish, and models can be tuned to score well on them. Your own tasks are the benchmark that matters.
Chain of thought
A model working through a problem step by step before giving its answer, either because the prompt asked it to or because it was trained to. It noticeably improves maths, logic and planning. Reasoning models do this automatically.
Chatbot
A conversational interface to an AI model, such as ChatGPT, Claude or Gemini. The word undersells what these products now do, which increasingly includes searching, coding, editing files and running agents.
Computer use
A model operating a computer the way a person does: looking at the screen, moving the pointer, clicking and typing. It lets AI work in apps that have no API. Anthropic brought it to the Mac in March 2026, and OpenAI built GPT-6 Astra around it.
Context engineering
The practice of deciding exactly what goes into a model's context for each step of a task: which documents, tool results, memories and instructions, and in what order. As agents run longer, choosing what the model sees has become as important as how you word the prompt.
Context window
How much text a model can consider at once, measured in tokens. It includes your instructions, any documents, the conversation so far and the model's reply. Several 2026 models handle around a million tokens, roughly a few thousand pages.See also: Token, RAG.
Deep research
A mode in ChatGPT, Gemini, Claude, Perplexity and others where the AI spends several minutes searching dozens of sources and writes a cited report, rather than answering instantly. Good for a first pass on a topic; always check the sources it cites.
Diffusion model
A type of model that creates images, video or audio by starting from random noise and refining it step by step into a result. Most image and video generators use it. Google has also tested diffusion for text, with DiffusionGemma reaching 1,000 tokens a second.
Distillation
Training a smaller, cheaper model to imitate a larger one, so it keeps much of the quality at a fraction of the running cost. Many of the fast, cheap models released in 2026 are distilled from bigger siblings.
Embedding
A list of numbers that represents the meaning of a piece of text, image or audio. Things with similar meanings get similar numbers, which lets software search by meaning rather than exact words. Embeddings are the backbone of RAG.See also: Vector database.
Evals (evaluations)
Tests you build to check whether an AI system does your job well, such as 50 real customer questions with known good answers. Running them after every change is the difference between improving a system and guessing.
Fine-tuning
Training an existing model further on your own examples to change how it behaves: its tone, format or skill at a narrow task. It is the wrong tool for teaching facts. Our guide on how to fine-tune an LLM explains when it is worth it.See also: LoRA, RAG.
Foundation model
A large, general model trained on broad data that other products and specialised models are built on top of. GPT, Claude, Gemini, Llama and Qwen are all families of foundation models.
Frontier model
A model at or near the leading edge of capability at the time, usually from Anthropic, OpenAI, Google or another top lab. The frontier moves every few weeks. Our 2026 model timeline tracks where it is now.
Generative AI
AI that creates new content, such as text, images, code, audio or video, rather than only classifying or predicting. Most of what people mean by AI today is generative AI.
GPU (graphics processing unit)
The chip that trains and runs most AI models, because it can do huge numbers of calculations in parallel. NVIDIA dominates the market, and shortages of GPUs and the power to run them shape what AI companies can offer and at what price.
Grounding
Tying a model's answer to a specific source, such as a search result or a company document, so it can be checked. A grounded answer comes with citations. RAG is the most common way to ground answers.
Guardrails
Rules and filters around a model that block unwanted inputs or outputs, such as personal data, harmful instructions or off-topic requests. They sit outside the model and can be adjusted without retraining it.
Hallucination
When a model states something false with confidence, such as a made-up citation, statistic or policy. It happens because models generate plausible text rather than looking up facts. Grounding, citations and checking reduce it. Nothing yet eliminates it.
Inference
Running a trained model to get an answer. Training happens once; inference happens every time anyone uses the model, which is why the cost and speed of inference drive AI pricing.
Jailbreak
A prompt designed to trick a model into ignoring its safety rules. Labs patch known jailbreaks continuously, and new ones keep appearing.See also: Prompt injection.
Knowledge cut-off
The date after which a model has no training data. It will not know about events after it unless it can search the web or is given documents. Always check whether an answer about recent events came from a search.
Large language model (LLM)
A model trained on enormous amounts of text to predict the next word, which turns out to be enough to write, summarise, translate, reason and code. The technology behind ChatGPT, Claude and Gemini.
Latency
How long you wait for a response. Fast modes and smaller models cut latency, which matters for voice assistants and coding tools where a few seconds feels like an age.
LoRA (low-rank adaptation)
A cheap way to fine-tune a model by training a small add-on rather than changing all of its weights. It makes custom models affordable on a single rented GPU.
MCP (Model Context Protocol)
An open standard that lets AI assistants connect to tools and data, such as your files, calendar, CRM or design software, through one shared format. Think of it as USB-C for AI. Our full MCP explainer covers how it works and the security risks.
Mixture of experts (MoE)
A model design where only a small part of the model switches on for each word. It lets a model be enormous but cheap to run. Moonshot's Kimi K3 has 2.8 trillion parameters but uses only 16 of its 896 experts per token.
Model card
A document a lab publishes alongside a model, describing what it was trained for, how it scored on tests, known limits and safety findings. A launch with no model card is a launch you cannot properly assess.
Multimodal
Able to work with more than one kind of input or output, such as text, images, audio and video. Most frontier models are now multimodal: you can show them a screenshot or a chart and ask about it.
Neural network
A system of connected mathematical units, loosely inspired by the brain, that learns patterns from data by adjusting the strength of its connections. Every model on this page is a neural network.
Open weights
A model whose trained weights are published so anyone can download, run and modify it, usually under a licence. It is not quite the same as open source, because the training data and code are often withheld. In 2026 the strongest open-weight models come mostly from Chinese labs such as DeepSeek, Moonshot and Alibaba.
Parameters
The adjustable numbers inside a model that are set during training. More parameters usually means more capability and higher running costs, though mixture-of-experts designs break that link. Frontier models now have hundreds of billions to trillions.See also: Weights.
Pre-training
The first, most expensive stage of building a model: learning general patterns from a vast amount of text and other data. Post-training then shapes it into a helpful, safe assistant.
Prompt
The input you give a model: the question, the instructions, any examples and any documents. The quality of the prompt still decides much of the quality of the answer.
Prompt caching
Storing the processed form of a prompt you send repeatedly, such as a long system prompt or document, so later requests are faster and much cheaper. In 2026 cheaper cache reads drove some of the biggest real-world price cuts, as with Opus 5.5 and GPT-6.
Prompt engineering
The craft of writing prompts that get reliable, useful results: clear goals, context, examples, format instructions and constraints. Our prompt engineering basics guide covers the techniques that matter.
Prompt injection
An attack where instructions are hidden in content the AI reads, such as a web page, email or document, so the model treats them as orders. It is the biggest security risk for agents that read untrusted content and can take actions.
Quantisation
Storing a model's numbers at lower precision so it needs less memory and runs faster, with a small loss of quality. It is what lets capable models run on laptops and phones: Google's Gemma 4 fits in under 1GB this way.
RAG (retrieval-augmented generation)
Looking up relevant passages from your own documents and giving them to the model with the question, so it answers from your material rather than memory, with sources. Our RAG explainer covers how to build one.
Reasoning model
A model trained to think through a problem at length before answering, trading speed and cost for accuracy on maths, coding, planning and analysis. Most 2026 models let you choose how much effort they spend.
Red teaming
Deliberately attacking an AI system to find its weaknesses before others do, from jailbreaks to dangerous capabilities. Labs run red teams before major releases and publish summaries in the model card.
RLHF (reinforcement learning from human feedback)
A training method where people rate a model's answers and the model learns to produce the kind they prefer. It is a large part of why chatbots are helpful and polite rather than merely fluent.
Small language model (SLM)
A compact model, typically a few billion parameters, designed to run cheaply or on a device such as a phone or laptop. Less capable than a frontier model, but fast, private and often good enough for a focused task.
Subagent
An agent started by another agent to handle part of a job in parallel, such as researching one file while the main agent works on another. Coding tools like Claude Code use them to split up large tasks.
Sycophancy
A model's tendency to tell you what you want to hear: agreeing with a flawed plan, praising weak work or backing down when you push. Labs train against it, and asking the model to argue the other side is a useful check.
Synthetic data
Training data generated by AI rather than collected from people. Labs increasingly use it to fill gaps, especially for maths and coding, where answers can be checked automatically.
System prompt
Standing instructions given to a model before any conversation, setting its role, tone, rules and knowledge. Every AI product has one, even if you never see it.
Temperature
A setting that controls how predictable a model's output is. Low temperature gives consistent, focused answers. High temperature gives more varied, creative ones, and more mistakes.
Token
The unit models read and write in, roughly three-quarters of an English word. Context windows, usage limits and API prices are all counted in tokens, so a price of $2 per million input tokens is about $2 for 750,000 words read.
Tool use (function calling)
A model's ability to call external tools, such as a search engine, calculator, database or another app, and use the results. It is what turns a model into an agent.See also: MCP.
Training
The process of building a model by showing it vast amounts of data and adjusting its parameters until it predicts well. Frontier training runs cost hundreds of millions of dollars and months of GPU time.
Transformer
The neural network architecture, introduced by Google researchers in 2017, that underpins almost every modern language model. The T in GPT stands for transformer.
Vector database
A database built to store embeddings and find the closest matches by meaning very quickly. It is the search engine inside most RAG systems. Examples include pgvector, Pinecone, Weaviate, Qdrant and Chroma.
Vibe coding
Building software by describing what you want in plain language and letting AI write the code, often without reading it closely. Great for prototypes, risky for anything that handles real data or money. Our guide to the best AI coding tools covers the options.
Weights
The learned values inside a trained model; in practice, the model itself. Releasing the weights means releasing the model for others to run.See also: Open weights, Parameters.
Zero-shot and few-shot
Zero-shot means asking a model to do a task with no examples. Few-shot means including a handful of examples in the prompt. Showing two or three examples of what good looks like is one of the most reliable ways to improve results.

Missing a term? We add to this glossary as new words take hold. If there is one you keep seeing and cannot pin down, tell us and we will add it.

Frequently asked questions

What is the difference between AI, machine learning and generative AI?

AI is the broad field of making computers do tasks that normally need human intelligence. Machine learning is the main way that is done today, by learning patterns from data instead of following hand-written rules. Generative AI is the branch of machine learning that creates new content such as text, images and code, which is what tools like ChatGPT and Claude do.

What is an AI agent in simple terms?

An AI agent is an AI that can take actions to complete a task, not just answer questions. It plans the steps, uses tools such as a browser, files or other apps, checks the results and keeps going until the job is done, usually asking a person to approve anything risky.

What is a token in AI?

A token is the unit of text an AI model reads and writes, roughly three-quarters of an English word. Context windows, usage limits and API prices are all measured in tokens.

What does open weights mean?

An open-weights model is one whose trained parameters are published so anyone can download and run it, usually under a licence that sets the terms. It differs from fully open source, because the training data and code are often not released.

What is the difference between an LLM and a chatbot?

An LLM (large language model) is the underlying model that generates text. A chatbot is a product built on top of one, adding a conversation interface, a system prompt, safety rules and often tools such as web search. ChatGPT is a chatbot; GPT-6 is the model inside it.

How to use this glossary

You do not need all of these words to use AI well. The handful that matter most day to day are prompt, context window, token, hallucination and agent, because they explain what AI tools can see, what they cost and where they go wrong. Learn those five and the rest of the jargon becomes much easier to decode.

The jargon changes weekly. We translate it.

Every weekday morning before 8am we explain what launched in AI and what it means for your work, in plain English. Free, and you can unsubscribe any time.

You're in. First issue lands tomorrow morning.