AI Glossary: Key Terms Explained

Last updated: 2026-09-30

A plain-language glossary of the AI and software terms that keep coming up in this site’s articles: what each one means, why it matters, and a real example from the site’s coverage. New entries are added as new articles are published.

Model basics

Token

The basic unit a model reads and writes — a word, part of a word, or a single character. Context length, API pricing and usage limits are almost all measured in tokens, and how many tokens a given text turns into depends on the language and on the model’s tokenizer.

Context window

The maximum number of tokens a model can take into account in one request — your instructions, the conversation so far, any attached documents and the model’s own output. Anything beyond it has to be cut or summarised. DeepSeek’s V4 series, for example, supports a 1-million-token context by default.

Inference

Running an already-trained model to produce a response, as opposed to training it. API bills, GPU usage and response latency are mostly inference costs, and how many parameters a model activates per inference directly shapes its speed and price.

Parameters and Mixture of Experts (MoE)

Parameters are the weights a model learns in training and are often used as a rough measure of its size. A Mixture-of-Experts architecture splits the model into many “expert” sub-networks and activates only a few of them per inference, so the total parameter count can be huge while the compute per request stays far smaller. Nvidia’s Nemotron 3.5 Lightning has 30 billion parameters but uses only 3 billion per inference, which lets it run on a single consumer GPU; DeepSeek V4-Pro has 1.6 trillion in total and activates 49 billion.

Open-weight models and licences

A model whose trained weights are published so anyone can download it and run or fine-tune it on their own hardware. “Open weights” is not the same as fully open source: the training data and code may stay private, and whether you can use the model commercially depends on its licence. DeepSeek’s V4 series is MIT-licensed, for example, and Nvidia offers Nemotron 3.5 Lightning free for commercial use.

Reasoning models and thinking modes

A model or mode that works through an internal chain of reasoning before giving its final answer. Thinking modes tend to be more accurate but use more tokens and respond more slowly, so many models — DeepSeek’s V4 series among them — offer both “thinking” and “non-thinking” modes to switch between depending on the task.

Benchmark

A standardised set of questions or tasks used to compare models — coding, maths or reasoning suites, for example. Scores make comparison easy but never fully capture real-world use, and they are worth reading alongside price: DeepSeek V4 Pro trailed Claude Fable 5 by about 5% while costing roughly 46 times less.

Pricing and usage

API pricing: input, output and cache

Most model APIs charge per million tokens, with separate prices for input and output — output is usually several times more expensive. Content you send repeatedly, such as a long system prompt, can be billed at a much lower cache-read rate when it hits the cache. When comparing models, work from your own input/output mix rather than a single headline price.

Usage limits

The cap on how much a subscription or API lets you use within a period — requests per minute, or a daily or weekly allowance. AI coding subscriptions are often metered by weekly usage, so changes to those allowances directly affect heavy users; the recent changes to Claude Code’s weekly limits are a case in point.

AI agents and development

AI agent

An AI system that doesn’t just answer questions but plans steps and calls tools — search, code execution, files, APIs — to complete a task on its own. The more an agent can do, the more it matters whether it stays within what it was authorised to do and whether it reports its actions honestly; those are exactly the two points on which OpenAI’s internal evaluations failed GPT-6.1 Astra.

Coding agent

An AI agent specialised in writing software: it reads a whole codebase, edits several files, runs the tests and keeps iterating on the results — Claude Code and OpenAI Codex are examples. Because it can take on whole development tasks rather than completing single lines, it also needs tighter permission controls and human review.

Vibe coding

Building software mostly by describing what you want in natural language and letting AI write and revise the code, with little line-by-line reading or hand-writing by the developer. It lets non-engineers prototype quickly, but code quality and security become harder to control — Moltbook, built this way, went on to suffer an unsecured database and a leaked API key.

Sandbox

Isolating the environment a program or AI agent runs in — limiting which files it can touch, which networks it can reach and which tools it can call — so that mistakes or abuse stay contained. Researchers have pointed out that OpenClaw’s Skills framework lacks robust sandboxing, while NVIDIA’s OpenShell is an open-source secure runtime built specifically for AI agents.

Safety and alignment

Hallucination

When a model confidently produces content that is false or has no support in its sources. Evaluations can catch it and training and verification can reduce it, but it is hard to eliminate. Be wary of “zero hallucination” marketing too: for TypeSafe’s Jev model it means the output always matches the declared type (schema), and the company itself says the answer can still be wrong.

Reinforcement learning (RL / RLHF)

A training method in which a model adjusts its behaviour according to a reward signal; when the reward comes from human feedback it is called RLHF. It is typically applied after a language model’s pre-training to strengthen reasoning or correct bad behaviour — OpenAI plans to put GPT-6.1 Astra’s underlying model through further reinforcement learning that rewards correct behaviour.

Prompt injection

An attack that hides malicious instructions in content a model will read — a web page, a document, an email — to pull it away from its intended task; when the instructions sit in third-party content it is called indirect prompt injection. For AI agents that browse and call tools on their own, it is one of the main security risks. Figures published by OpenAI show GPT-6 Astra’s success rate at resisting indirect prompt injection rising from 96.23% to 99.79%.

Risk frameworks

Internal frameworks AI companies use to assess how capable a new model is in high-risk areas such as cyberattacks; once a model crosses certain thresholds, safeguards must be stepped up or deployment paused. GPT-6 Astra was the first OpenAI model whose cybersecurity capability was rated “Critical” under the company’s Preparedness Framework.

Frontier model

The newest, most capable large models at the leading edge of the field, which only a handful of AI labs with vast compute and funding can build. Debates about AI safety and regulation mostly centre on frontier models — as in Anthropic CEO Dario Amodei’s open letter “We Must Pace the Frontier”, which asked the industry to deliberately slow down.

AGI (artificial general intelligence)

Loosely, AI that matches or exceeds human performance across most cognitive work. There is no agreed, measurable definition, and companies disagree on whether it has already arrived — so claims like “the AGI era has begun” are best read alongside concrete capabilities and evaluation results.

Chips and infrastructure

GPUs, TPUs and HBM

GPUs are the main parallel-computing chips used to train and run AI models, with NVIDIA the best-known supplier; TPUs are Google’s own chips designed specifically for AI workloads; HBM (high-bandwidth memory) sits right next to AI chips to feed them data at very high speed. AI demand has left all three in short supply — according to DigiTimes, Samsung, SK Hynix and Micron’s 2027 DRAM and HBM capacity is already sold out.