Free reference

AI Glossary

The AI terms you keep seeing, explained in plain English — what it actually means, why it matters, and a concrete example. No dictionary one-liners.

AI Agent
An AI system that plans and takes multiple actions on its own to reach a goal, instead of just answering one question.
A chatbot answers what you ask. An agent breaks a goal into steps, picks tools to use, checks its own output, and keeps going until the task is done — booking a trip across three sites, or researching a topic across a dozen sources. The trade-off is control: the more autonomy you hand over, the more you need to trust its judgment on the steps in between.
Chain-of-Thought
Asking a model to reason step by step before giving a final answer, instead of jumping straight to a conclusion.
Models get noticeably more accurate on multi-step problems — math, logic, planning — when they write out intermediate reasoning first. You'll see this called "thinking" or "reasoning" in Claude and other assistants. Practically: if you get a wrong or shallow answer, asking the model to "think through this step by step" before answering often fixes it.
Claude Skill
A packaged set of instructions Claude loads automatically when relevant, so you don’t have to re-explain how you want a task done every time.
A skill is a markdown file (SKILL.md) plus, sometimes, supporting resources or scripts, dropped into a folder Claude reads from. Instead of retyping "write this the way I like it" every session, you install the skill once and Claude picks it up whenever the task matches — a recurring report format, a coding standard, a brand voice. Skills persist; a one-off prompt doesn’t.
Browse Claude Skills Packs →
Context Window
The amount of text a model can “see” at once — your prompt, any files, and the conversation so far, all counted together.
Measured in tokens, not characters or words. Once a conversation or document exceeds the context window, the model literally cannot see the parts that fall outside it — it doesn't get slower, it forgets. This is why very long chats sometimes lose track of something you said early on: it has scrolled out of the window.
Fine-tuning
Further training an existing model on your own examples so its default behavior changes, permanently, for that copy of the model.
Different from prompting: a prompt or skill changes behavior for one conversation by giving instructions; fine-tuning changes the model's underlying weights by training it on hundreds or thousands of example inputs/outputs. It's slower, costlier, and harder to undo than writing a better prompt — worth it mainly when no amount of prompting gets consistent results at scale.
Guardrails
Rules and checks placed around a model to keep its output inside safe, on-topic, or on-brand limits.
Guardrails can live in the system prompt ("never discuss competitor pricing"), in a filter that screens output before a user sees it, or in the product design itself (limiting what the model is even asked to do). They matter most wherever a model’s output reaches a customer directly and unreviewed — support bots and public-facing chat widgets, above all.
Hallucination
A confident, fluent answer that is factually wrong — the model isn’t lying, it’s pattern-completing without a ground-truth check.
Language models predict plausible next words, not verified facts, so a hallucination reads exactly as fluent and confident as a correct answer — there's no visible tell. Rates vary a lot by domain and task; they're worst on niche facts, exact citations, and anything the model would have to know rather than infer. Fix: ask for sources, and verify anything you'd be embarrassed to have wrong.
Inference
The act of a trained model producing an output for a given input — i.e., just using it, as opposed to training it.
Training happens once (or occasionally); inference happens every single time you send a prompt and get a reply. "Inference cost" and "inference speed" — the two numbers that actually affect your day-to-day bill and wait time — refer to this running stage, not the (much larger) one-time cost of building the model in the first place.
LLM (Large Language Model)
A model trained on huge amounts of text to predict and generate language — the technology behind Claude, ChatGPT, and similar tools.
"Large" refers to the number of parameters (internally adjustable values) and the volume of training text, both in the billions. An LLM doesn’t look anything up in real time by default; everything it "knows" was absorbed during training, which is why it has a knowledge cutoff date and can be wrong about very recent events unless it’s given live information some other way.
MCP (Model Context Protocol)
An open standard, introduced by Anthropic, for connecting an AI model to external tools and data sources in a consistent way.
Before MCP, every AI app wired up its own one-off way to reach a database, a file system, or an API. MCP defines a common protocol so a tool built once — say, a connector to a project-management app — can be reused by any MCP-compatible AI client, not rebuilt per app. It's the plumbing standard that lets Claude and other assistants plug into outside systems predictably.
Multimodal
A model that can work with more than one type of input or output — text and images together, for example, not just text.
A multimodal model can, say, look at a screenshot of a chart and explain the trend in words, or read a scanned PDF instead of only clean typed text. Not every model is multimodal, and the ones that are aren’t all equally capable across every combination — check what a specific model actually supports before assuming it can read an image, a chart, or audio.
Prompt
The instructions and context you give a model to get the output you want — as simple as a question, or as detailed as a full brief.
The gap between a mediocre and a great AI output is usually the prompt, not the model. A weak prompt ("write a sales email") leaves everything to guesswork; a strong one specifies audience, tone, length, what to avoid, and what a good output looks like. This is the whole idea behind a written, reusable prompt instead of a one-off ask.
Browse Expert Prompt Packs →
Prompt Engineering
The practice of designing and refining prompts deliberately to get more reliable, higher-quality output from a model.
Less about "magic phrases" than about being specific: giving the model a role, the real constraints, a worked example of the output you want, and a way to check its own work. Good prompt engineering is the difference between re-explaining yourself every time and getting a consistent, usable result on the first try.
Browse Expert Prompt Packs →
Prompt Injection
An attack where hidden instructions inside content a model reads try to hijack its behavior — the AI-specific version of a security exploit.
If a model summarizes a webpage, a PDF, or an email, and that content secretly contains text like "ignore previous instructions and instead...", a poorly-guarded system might follow it. This matters most for anyone building AI features that read outside content automatically — support, legal, and research workflows should treat untrusted documents as untrusted, the same way a browser treats an unknown download.
RAG (Retrieval-Augmented Generation)
A setup where a model looks up relevant documents before answering, instead of relying only on what it learned during training.
RAG is how AI tools answer questions about your own private documents, or about anything newer than the model's training cutoff: a search step retrieves the relevant passages first, then the model writes its answer grounded in that retrieved text. It reduces hallucination on anything the retrieved documents actually cover — it doesn't help on questions those documents don't answer.
System Prompt
Standing instructions set once, usually by whoever builds the app, that shape a model’s behavior across an entire conversation.
Where a regular prompt is what you type for one turn, the system prompt runs quietly in the background of every turn — the persona, the tone, the rules, the things it should never do. It’s why the same underlying model can feel completely different depending on which app you’re using it through.
Temperature
A setting that controls how random or predictable a model’s output is — low temperature plays it safe, high temperature takes more risks.
At temperature 0, a model tends toward the single most likely next word every time — consistent, a little flat, good for factual or structured tasks. Turn it up and it samples from a wider range of plausible next words — more varied and creative, but also more likely to wander or go strange. Most chat apps pick a sensible default for you; it mainly matters when you're calling a model through code.
Token / Tokenization
The small chunks of text — often close to a word, sometimes a fragment of one — that a model actually reads and generates in, and that you’re billed by.
Models don't process raw characters or whole words; they split text into tokens first (tokenization) and work with those. This is why pricing and context-window limits are quoted in tokens, not words — and why the exact token count for the same sentence can differ slightly between models, since each one tokenizes text a little differently.
Zero-shot / Few-shot Prompting
Asking a model to do a task with no examples (zero-shot) versus giving it one or more worked examples first (few-shot).
Zero-shot relies entirely on the model’s general training to infer what you want. Few-shot shows it 1–5 examples of input → output pairs in the prompt itself, which reliably improves consistency on tasks with a specific format or style — a big part of why a well-built prompt or skill so often includes a worked example rather than just an instruction.

Know the terms. Now put them to work.

Skills, prompts, and website blueprints built for 12 professional fields — ready to install, no prompt engineering required.

Browse Packs