Large language model (LLM)
A neural network trained on very large amounts of text that generates output by repeatedly predicting the next token. ChatGPT, Claude, Gemini and DeepSeek are all built on large language models, and most recent ones also accept images, audio and other inputs.
Related articles: Opus 5.5 Is Here, and Your Free Limit Reset Expires Oct 22 · GPT-6 Astra: OpenAI's First 'Critical Risk' AI Model
Token
The basic unit a model reads and writes — a word, part of a word, or a single character. Context length, API pricing and usage limits are almost all measured in tokens, and how many tokens a given text turns into depends on the language and on the model’s tokenizer.
Related articles: DeepSeek Undercuts Astra on Output by 83x
Context window
The maximum number of tokens a model can take into account in one request — your instructions, the conversation so far, any attached documents and the model’s own output. Anything beyond it has to be cut or summarised. DeepSeek’s V4 series, for example, supports a 1-million-token context by default.
Related articles: DeepSeek V4 Pro Ships, Price Hike Already Flagged
Inference
Running an already-trained model to produce a response, as opposed to training it. API bills, GPU usage and response latency are mostly inference costs, and how many parameters a model activates per inference directly shapes its speed and price.
Related articles: Nvidia's New Open-Source AI Model Runs Free on a Single GPU
Parameters and Mixture of Experts (MoE)
Parameters are the weights a model learns in training and are often used as a rough measure of its size. A Mixture-of-Experts architecture splits the model into many “expert” sub-networks and activates only a few of them per inference, so the total parameter count can be huge while the compute per request stays far smaller. Nvidia’s Nemotron 3.5 Lightning has 30 billion parameters but uses only 3 billion per inference, which lets it run on a single consumer GPU; DeepSeek V4-Pro has 1.6 trillion in total and activates 49 billion.
Related articles: Nvidia's New Open-Source AI Model Runs Free on a Single GPU · DeepSeek V4 Pro Ships, Price Hike Already Flagged
Open-weight models and licences
A model whose trained weights are published so anyone can download it and run or fine-tune it on their own hardware. “Open weights” is not the same as fully open source: the training data and code may stay private, and whether you can use the model commercially depends on its licence. DeepSeek’s V4 series is MIT-licensed, for example, and Nvidia offers Nemotron 3.5 Lightning free for commercial use.
Related articles: DeepSeek V4 Pro Ships, Price Hike Already Flagged · Nvidia's New Open-Source AI Model Runs Free on a Single GPU
Distillation
Training a smaller “student” model on the outputs of a larger “teacher” model so it picks up similar abilities at a fraction of the cost. Nemotron 3.5 Lightning was distilled from the larger Nemotron 3 Ultra — and because distillation can borrow a rival’s capabilities, Anthropic added anti-distillation measures alongside Claude Fable 5.1.
Related articles: Nvidia's New Open-Source AI Model Runs Free on a Single GPU · Claude Fable 5.1: A Hedge Fund's 5-Year-Old Bug, Solved
Multimodal
Able to understand or generate more than one kind of data — text, images, audio, video. ByteDance’s Seedance 2.5, for instance, generates 30-second videos with synchronised audio and accepts up to 50 multimodal reference inputs.
Related articles: Seedance 2.5: 30-Second Video With Synced Audio · Zero API Cost: A Local AI Animation Workflow
Reasoning models and thinking modes
A model or mode that works through an internal chain of reasoning before giving its final answer. Thinking modes tend to be more accurate but use more tokens and respond more slowly, so many models — DeepSeek’s V4 series among them — offer both “thinking” and “non-thinking” modes to switch between depending on the task.
Related articles: DeepSeek V4 Pro Ships, Price Hike Already Flagged
Benchmark
A standardised set of questions or tasks used to compare models — coding, maths or reasoning suites, for example. Scores make comparison easy but never fully capture real-world use, and they are worth reading alongside price: DeepSeek V4 Pro trailed Claude Fable 5 by about 5% while costing roughly 46 times less.
Related articles: DeepSeek Quietly Ships V4 Pro, Targets Fable 5