AI Glossary
The AI terms you keep seeing, explained in plain English — what it actually means, why it matters, and a concrete example. No dictionary one-liners.
- AI Agent
- An AI system that plans and takes multiple actions on its own to reach a goal, instead of just answering one question.
- A chatbot answers what you ask. An agent breaks a goal into steps, picks tools to use, checks its own output, and keeps going until the task is done — booking a trip across three sites, or researching a topic across a dozen sources. The trade-off is control: the more autonomy you hand over, the more you need to trust its judgment on the steps in between.
- Chain-of-Thought
- Asking a model to reason step by step before giving a final answer, instead of jumping straight to a conclusion.
- Models get noticeably more accurate on multi-step problems — math, logic, planning — when they write out intermediate reasoning first. You'll see this called "thinking" or "reasoning" in Claude and other assistants. Practically: if you get a wrong or shallow answer, asking the model to "think through this step by step" before answering often fixes it.
- Claude Skill
- A packaged set of instructions Claude loads automatically when relevant, so you don’t have to re-explain how you want a task done every time.
- A skill is a markdown file (SKILL.md) plus, sometimes, supporting resources or scripts, dropped into a folder Claude reads from. Instead of retyping "write this the way I like it" every session, you install the skill once and Claude picks it up whenever the task matches — a recurring report format, a coding standard, a brand voice. Skills persist; a one-off prompt doesn’t. Browse Claude Skills Packs →
- Context Window
- The amount of text a model can “see” at once — your prompt, any files, and the conversation so far, all counted together.
- Measured in tokens, not characters or words. Once a conversation or document exceeds the context window, the model literally cannot see the parts that fall outside it — it doesn't get slower, it forgets. This is why very long chats sometimes lose track of something you said early on: it has scrolled out of the window.
- Fine-tuning
- Further training an existing model on your own examples so its default behavior changes, permanently, for that copy of the model.
- Different from prompting: a prompt or skill changes behavior for one conversation by giving instructions; fine-tuning changes the model's underlying weights by training it on hundreds or thousands of example inputs/outputs. It's slower, costlier, and harder to undo than writing a better prompt — worth it mainly when no amount of prompting gets consistent results at scale.
- Guardrails
- Rules and checks placed around a model to keep its output inside safe, on-topic, or on-brand limits.
- Guardrails can live in the system prompt ("never discuss competitor pricing"), in a filter that screens output before a user sees it, or in the product design itself (limiting what the model is even asked to do). They matter most wherever a model’s output reaches a customer directly and unreviewed — support bots and public-facing chat widgets, above all.
- Hallucination
- A confident, fluent answer that is factually wrong — the model isn’t lying, it’s pattern-completing without a ground-truth check.
- Language models predict plausible next words, not verified facts, so a hallucination reads exactly as fluent and confident as a correct answer — there's no visible tell. Rates vary a lot by domain and task; they're worst on niche facts, exact citations, and anything the model would have to know rather than infer. Fix: ask for sources, and verify anything you'd be embarrassed to have wrong.
- Inference
- The act of a trained model producing an output for a given input — i.e., just using it, as opposed to training it.
- Training happens once (or occasionally); inference happens every single time you send a prompt and get a reply. "Inference cost" and "inference speed" — the two numbers that actually affect your day-to-day bill and wait time — refer to this running stage, not the (much larger) one-time cost of building the model in the first place.
- LLM (Large Language Model)
- A model trained on huge amounts of text to predict and generate language — the technology behind Claude, ChatGPT, and similar tools.
- "Large" refers to the number of parameters (internally adjustable values) and the volume of training text, both in the billions. An LLM doesn’t look anything up in real time by default; everything it "knows" was absorbed during training, which is why it has a knowledge cutoff date and can be wrong about very recent events unless it’s given live information some other way.
- MCP (Model Context Protocol)
- An open standard, introduced by Anthropic, for connecting an AI model to external tools and data sources in a consistent way.
- Before MCP, every AI app wired up its own one-off way to reach a database, a file system, or an API. MCP defines a common protocol so a tool built once — say, a connector to a project-management app — can be reused by any MCP-compatible AI client, not rebuilt per app. It's the plumbing standard that lets Claude and other assistants plug into outside systems predictably.
- Multimodal
- A model that can work with more than one type of input or output — text and images together, for example, not just text.
- A multimodal model can, say, look at a screenshot of a chart and explain the trend in words, or read a scanned PDF instead of only clean typed text. Not every model is multimodal, and the ones that are aren’t all equally capable across every combination — check what a specific model actually supports before assuming it can read an image, a chart, or audio.
- Prompt
- The instructions and context you give a model to get the output you want — as simple as a question, or as detailed as a full brief.
- The gap between a mediocre and a great AI output is usually the prompt, not the model. A weak prompt ("write a sales email") leaves everything to guesswork; a strong one specifies audience, tone, length, what to avoid, and what a good output looks like. This is the whole idea behind a written, reusable prompt instead of a one-off ask. Browse Expert Prompt Packs →
- Prompt Engineering
- The practice of designing and refining prompts deliberately to get more reliable, higher-quality output from a model.
- Less about "magic phrases" than about being specific: giving the model a role, the real constraints, a worked example of the output you want, and a way to check its own work. Good prompt engineering is the difference between re-explaining yourself every time and getting a consistent, usable result on the first try. Browse Expert Prompt Packs →
- Prompt Injection
- An attack where hidden instructions inside content a model reads try to hijack its behavior — the AI-specific version of a security exploit.
- If a model summarizes a webpage, a PDF, or an email, and that content secretly contains text like "ignore previous instructions and instead...", a poorly-guarded system might follow it. This matters most for anyone building AI features that read outside content automatically — support, legal, and research workflows should treat untrusted documents as untrusted, the same way a browser treats an unknown download.
- RAG (Retrieval-Augmented Generation)
- A setup where a model looks up relevant documents before answering, instead of relying only on what it learned during training.
- RAG is how AI tools answer questions about your own private documents, or about anything newer than the model's training cutoff: a search step retrieves the relevant passages first, then the model writes its answer grounded in that retrieved text. It reduces hallucination on anything the retrieved documents actually cover — it doesn't help on questions those documents don't answer.
- System Prompt
- Standing instructions set once, usually by whoever builds the app, that shape a model’s behavior across an entire conversation.
- Where a regular prompt is what you type for one turn, the system prompt runs quietly in the background of every turn — the persona, the tone, the rules, the things it should never do. It’s why the same underlying model can feel completely different depending on which app you’re using it through.
- Temperature
- A setting that controls how random or predictable a model’s output is — low temperature plays it safe, high temperature takes more risks.
- At temperature 0, a model tends toward the single most likely next word every time — consistent, a little flat, good for factual or structured tasks. Turn it up and it samples from a wider range of plausible next words — more varied and creative, but also more likely to wander or go strange. Most chat apps pick a sensible default for you; it mainly matters when you're calling a model through code.
- Token / Tokenization
- The small chunks of text — often close to a word, sometimes a fragment of one — that a model actually reads and generates in, and that you’re billed by.
- Models don't process raw characters or whole words; they split text into tokens first (tokenization) and work with those. This is why pricing and context-window limits are quoted in tokens, not words — and why the exact token count for the same sentence can differ slightly between models, since each one tokenizes text a little differently.
- Zero-shot / Few-shot Prompting
- Asking a model to do a task with no examples (zero-shot) versus giving it one or more worked examples first (few-shot).
- Zero-shot relies entirely on the model’s general training to infer what you want. Few-shot shows it 1–5 examples of input → output pairs in the prompt itself, which reliably improves consistency on tasks with a specific format or style — a big part of why a well-built prompt or skill so often includes a worked example rather than just an instruction.
Know the terms. Now put them to work.
Skills, prompts, and website blueprints built for 12 professional fields — ready to install, no prompt engineering required.
Browse Packs