How AI Actually Works: Claude, ChatGPT and LLMs Explained Simply
No jargon, no hype. How these systems read your question, find context, call tools, and write an answer. Plus the real safety picture and why the doomsday headlines miss the point.
This article started at a barbershop. Small town, midwest, the kind of place where the conversation covers everything from local news to the end of the world. Someone brought up AI and the room had questions: robots taking over, bots walking the streets, the whole picture painted by whatever came across their feed that week. I pulled out my phone, showed them what I was actually doing with it, and tried to explain how it works in plain English. This is that explanation, written down.
Here is the simple version first. When you ask Claude or ChatGPT something, the system reads your question, breaks it into smaller pieces, checks patterns it learned during training, may pull in outside information or tools, and then writes the answer back one small piece at a time. Modern models also handle images, audio, and video, not just text.
A few quick definitions before we start. An LLM means large language model. It is software trained mainly on text to predict what comes next. Many products built around LLMs are also multimodal, meaning they can work with images, audio, or video too. A token is a small piece of text. A tool is an extra helper like web search, a calculator, a code runner, or a camera feed.
Open the sections below if you want the step by step version. If you are a barber in a small midwest town or someone who just wants to understand what is actually going on, this is written for you.
The basics: how an LLM writes text
Tokens and embeddings
Your text gets split into tokens. A token is a small piece of text, not always a full word. "Hello world" might become ["Hello", " world"]. Longer or unusual words often get broken into smaller parts.
Try this simplified illustration. Real token splits vary by model and are more complicated than this display:
Each token then gets turned into numbers. In AI, those numbers are often called a vector or an embedding. Think of that as the model's internal way of storing meaning so related words and ideas end up closer together.
Attention: how the model keeps track of meaning
Attention is the part that helps the model decide what matters most in a sentence. Each token checks the other tokens around it to figure out which words are important for meaning.
When the model reads "The cat sat on the mat", the word "sat" is linked closely to "cat" and "mat" because those words help explain who did the action and where it happened.
Large models do this many times in parallel using attention heads. An attention head is just one small pattern detector focused on a certain kind of relationship:
- →One head may track who did the action
- →Another may track what the action applies to
- →Another may connect names, pronouns, and references
- →Many others learn patterns around order, grammar, tone, and structure
How the answer is written one token at a time
The last step is prediction. The model scores many possible next tokens and samples one, usually favoring the higher-scoring options. Those scores are called probabilities, which just means how likely each option is.
After it picks the next token, that token gets added to the running answer. Then the model does the same thing again. That loop is why replies appear word by word even though a lot of math is happening underneath.
Tool calling: how Claude or ChatGPT reaches outside the model
The model by itself does not magically know live prices, your private documents, or what is inside a spreadsheet. For that it uses tools. A tool is a connected helper like web search, a calculator, a code runner, or a company database. You may also hear this called function calling. An API is simply a formal way for one software system to ask another software system for data or an action.
Common assistant tools
- Web search
- Code or data analysis in a sandbox
- Image generation
- File upload and analysis
- Connections to approved apps or data sources
What the product owner controls
- Which tools are available
- What data each tool can access
- Whether an action needs approval
- Network, file, time, and cost limits
- Logging and safety checks
Search and retrieval: how RAG finds the right information
When Claude or ChatGPT needs facts from documents, websites, or a company knowledge base, it usually does not rely only on what the model learned during training. It searches for relevant text first. One common setup is called RAG, short for Retrieval Augmented Generation. In plain English, that means find useful source material first, then write the answer using that material.
How vector search works in plain English
Code execution: how the AI can test things safely
Some AI products can run code, inspect files, or make charts. They do this inside a sandbox. A sandbox is an isolated workspace that keeps the run separate from the rest of the system so mistakes or risky code do less damage.
Hosted assistant sandbox
- Runs selected code in an isolated environment
- Works with files and data the user provides
- Applies network and resource restrictions
- Has time, storage, and package limits
Developer-agent workspace
- Can read, write, and test approved project files
- May run terminal commands or install dependencies
- Uses the environment and permissions it is given
- Should be reviewed before it makes consequential changes
Exact capabilities and limits vary by product, plan, configuration, and date. Check the product's current documentation before relying on a specific permission or limit.
Watch AI Build a Simple App From Scratch
When you ask AI to build software, it does not just guess. It considers options, checks its tools, thinks through the logic step-by-step, and then brings the app to life.
Infrastructure: the systems behind the answer
All of this sits on backend infrastructure, which is just the hidden machinery behind the app. That includes GPUs for the heavy math, search systems for finding context, sandboxes for running code, and monitoring tools that watch speed, failures, and cost. You may hear the word inference here. Inference simply means the model generating an answer. Latency just means how long you wait for that answer.
What contributes to request cost
There is no universal cost per request. Providers publish current model and tool pricing, while subscription plans bundle changing usage limits. Estimate with the actual model, token volume, and tools your product uses.
AI agents: how the system handles multi-step work
Sometimes you do not want one answer. You want the system to handle a job with several steps, like checking an order, applying a policy, sending an email, and updating a record. That is where AI agents come in. An agent is simply a model plus tools plus a step-by-step loop for deciding what to do next.
The Three Components: Agent, Tools, Workflow
🧠 Agent (the coordinator)
This is the part that reads the goal, picks the next step, and decides when the job is done. Think of it like a coordinator, not a magic brain.
🔧 Tools (the helpers)
These are the connected helpers the agent can use: search a knowledge base, check a database, send an email, update a system, or write a file. Each tool has rules about what input it expects and what it returns.
🔄 Workflow (the step-by-step loop)
The workflow is the repeatable loop. Read the goal. Pick a step. Use a tool. Look at the result. Decide the next step. Stop when the task is done or hand it to a person if it gets stuck.
Demo only: this simulated workflow skips a human-approval step. A production system should pause for approval before a refund, other money movement, deletion, or a policy exception.
Action: call lookup_order(order_id="8472")
- →Tool fails: Agent sees error message, tries alternative approach
- →Ambiguous result: Agent asks clarifying question or makes best-effort decision
- →Max iterations hit: Agent summarizes progress, flags for human review
How companies put agents into real products
How the system keeps its place
These systems need to remember where they were. Teams save the current state, which means the conversation, tool results, and next step, in a database. If the process stops or times out, it can restart from the last saved point instead of beginning again.
Where a person still has to approve
For important actions, a human still needs to sign off. The agent prepares the action, pauses, and waits for approval. This is common for money movement, risky customer messages, deletions, or policy exceptions.
Safety rules around tools
Not every tool is safe to let loose automatically. Real systems put clear limits around what the agent can do:
- →Rate limits per tool (max 10 DB queries per workflow)
- →Permission checks (agent can read orders, not delete them)
- →Cost limits (max $5 in API calls per workflow)
- →Audit logs (every tool call logged with timestamp, inputs, outputs)
How teams watch what the agent is doing
If you cannot see what the agent is doing, you cannot trust it. Teams track simple things like success rate, number of steps, common failures, and cost:
Tracing tools follow each workflow from start to finish. Dashboards show slow steps, repeated loops, and failures so teams can fix the system before it causes trouble.
Common software used to build these systems
LangGraph (LangChain)
- State machines for agent workflows
- Built-in persistence and checkpointing
- Python-native, integrates with LangChain tools
- Good for complex multi-agent systems
CrewAI
- Multiple agents working together
- Role-based agents (researcher, writer, reviewer)
- Sequential and hierarchical workflows
- Good for content generation pipelines
AutoGen (Microsoft)
- Conversational agent framework
- Agents that talk to each other
- Built-in code execution and debugging
- Good for research and analysis tasks
Custom (Roll Your Own)
- Direct LLM API + tool calling
- Full control over how the workflow runs
- Minimal dependencies, max flexibility
- Good for production at scale
Going deeper: how tokenization actually works
The intro above uses "tokens" as shorthand. This optional section is for people building on an LLM API or writing prompts professionally. The design decisions here affect model behavior, cost, and context-window math.
The model never sees words. It sees integers. Every piece of input, your query, the system prompt, retrieved context, conversation history, gets converted to a flat sequence of token IDs before a single matrix multiply happens. How that conversion works matters more than most people realize.
Three ways to slice text: characters, words, and subwords
A vocabulary is the complete set of units a model can recognize and produce. How you define those units involves a hard tradeoff with no perfect answer.
Split on whitespace and punctuation. Intuitive, but vocabulary bloat kills you fast. English surface forms explode once you account for plurals, conjugations, compounds, proper nouns, typos, and domain jargon. Anything outside the fixed vocabulary at inference time becomes an unknown token, breaking the prediction chain. Pre-2018 NLP systems lived and died by this problem.
Every Unicode codepoint is its own token. Vocabulary stays tiny, you never hit an unknown token, and you can handle any script. The cost: sequence length explodes. English words average around 4-5 characters plus spaces, so a 512-token word-level context becomes roughly 2,500 tokens at the character level for the same passage. Learning long-range dependencies across that expanded length is significantly harder to train.
High-frequency strings get their own token. Low-frequency strings get decomposed into smaller pieces that individually appear more often. You get vocabulary coverage close to character models without the sequence length penalty. Most mainstream chat models use a subword scheme, although their exact algorithms and vocabularies differ.
Byte Pair Encoding: building a vocabulary from frequency
BPE is a bottom-up compression algorithm adapted for vocabulary construction. The loop is simple in principle:
The result: common
strings like function
or return
get a single token. Rare words get decomposed into pieces, often slicing right through what a linguist
would call morpheme boundaries. That is intentional: BPE optimizes for corpus frequency, not linguistic
structure, and those two things do not align.
Where tokenization creates surprising model behavior
If you have spent time prompting models at scale, you have probably run into behaviors that trace back to the tokenization layer rather than the model weights.
Character counting failures
Ask a model to count the letter "s" in "Mississippi" and it often gets it wrong. The model does not operate on characters directly. If "iss" merges into a single token, the individual letters inside it are not immediately recoverable from the model's internal representation. It has to reconstruct character-level information from the embedding, which is lossy.
Whitespace-sensitive tokenization
The token for
"python" and " python" with a leading space are
usually different IDs with different embeddings. Stripping whitespace from prompts can produce subtly
different outputs. This matters if you are doing token-level log probability analysis or building
evals that compare raw completions.
Rhyme and phonetics
Models struggle with rhyming more than their language fluency suggests they should. Two words can rhyme while sharing almost no subword tokens. The model has to infer phonetic similarity from spelling patterns, which are encoded unevenly across the vocabulary depending on how each word was tokenized.
Multilingual token inflation
A vocabulary built on an English-heavy corpus tokenizes English efficiently and everything else less so. A 100-token English passage might cost 300-500 tokens in a morphologically rich language with less training representation. Your effective context window shrinks for non-Latin input, and you burn more tokens per API call for the same semantic content.
What the model "knows" about its own tokens
Each token ID maps to a learned embedding vector. That vector is not engineered by hand. It starts random and gets updated continuously during training as the model learns to predict the next token across trillions of examples. By the end of training, the embedding for a given token implicitly encodes everything the model observed that token doing across the corpus.
This raises a real question: does the embedding for a token like deploy
"know" it is spelled d-e-p-l-o-y? Formally, no. The token is just an integer. But research shows
character-level information does bleed into embeddings. The most likely reason: the same root word
surfaces in many tokenized forms across a large corpus (singular, plural, past tense, mid-sentence,
sentence-initial, after punctuation, inside a code block) and the model is incentivized to build
representations that generalize across those surface variations.
Practical implication: models handle spelling and character-level tasks better for common words (which appear in many tokenized surface forms during training) and worse for rare words or niche proper nouns that were almost always tokenized the same way. If you are building a prompt that requires character-level precision, explicitly spell out the string in your prompt rather than assuming the model can decompose it cleanly.
Tokenizers are versioned separately from models. OpenAI models, for example, may use different encodings such as cl100k_base or o200k_base; other providers use different vocabularies and rules. Never estimate token count by dividing word count by a fixed ratio when cost or context-window headroom matters. Run the tokenizer for the specific model you are calling.
Is AI actually safe?
Short answer: the most immediate risks are not sci-fi robot takeovers. They are misuse and failure in the systems people build: disinformation, autonomous weapons, surveillance, fraud, and bias in high-impact decisions. Safety is an engineering discipline with several layers, but no layer is perfect. Here is what those layers can and cannot do.
Training alignment: teaching the model what not to do
Before a model ever reaches you, it goes through a process called alignment training. This is where the model is trained not just on what is factually correct but on what is helpful, harmless, and honest.
Common techniques include RLHF (Reinforcement Learning from Human Feedback), where human reviewers rate model responses, and newer preference-optimization methods. The model is then updated to produce more helpful responses and fewer unsafe ones. Over many rounds, this shapes behavior before anyone outside the lab sees it.
- →The model learns to decline requests that could cause harm
- →It learns to flag uncertainty instead of making things up confidently
- →It learns to avoid helping with things that are clearly illegal or dangerous
- →It learns tone, nuance, and when to push back on a bad idea
This is not perfect. Models can still be pushed in wrong directions with the right prompts. But the baseline behavior is trained, not bolted on as an afterthought.
System prompts and operator guardrails
Many AI products add system instructions and other guardrails. Before the model sees your message, the company or developer may give it instructions about its role, available tools, and safety rules.
Think of it as an employer giving a contractor rules. Those rules influence behavior, but they are not a security boundary by themselves. Reliable systems combine them with permissions, validation, and monitoring.
On top of that, most platforms run content filters that check both input and output:
- →Input filters catch known harmful patterns before they reach the model
- →Output filters check the model's response before it reaches you and block or rewrite anything that crosses a line
- →Rate limits and monitoring flag unusual usage patterns and slow down or stop automated abuse
What the model can and cannot do on its own
A base model cannot take actions outside its response. It can process the input it is given and generate text or other output, but it cannot reach the internet, execute code, send emails, move money, or control systems unless the surrounding application connects a tool and allows that action.
When a model does have tools, those tools are scoped. A customer support agent with access to your order data does not also have access to your bank account. The access is defined by whoever built the product, not decided by the model.
This is important because it makes the risk surface reviewable. An agent that can only read your calendar and draft meeting summaries is limited by those permissions. The question becomes what the people who built it connected, what data it can reach, and which actions still need approval.
Red teaming and ongoing testing
Many major model releases go through red-team exercises. Red teaming is where people try to break the system: elicit harmful content, bypass safeguards, or trigger unexpected behavior.
Findings can inform mitigations before release, and problems found afterward need patches and monitoring. This is similar to security testing: evaluate, ship carefully, monitor, and improve.
Major labs publish different levels of safety documentation, and some releases include independent evaluations. The depth of reporting and third-party review varies by organization and model release.
The doom question
Every few months someone writes a piece about AI ending humanity. Here is a more grounded take on what the actual risk picture looks like today.
🚫 Overhyped / not how it works
- A model spontaneously "waking up" with goals
- AI deciding on its own to take over systems
- Terminator-style self-preservation instincts
- Models conspiring without human instruction
🔴 Already happening right now
- AI-generated disinformation at election scale
- Autonomous targeting in active warzones
- Deepfakes used for fraud and harassment
- Biased models used in hiring and sentencing
⚠️ Real and getting harder to ignore
- Multi-agent systems with minimal human oversight
- Concentration of capability in 3-4 companies
- Economic disruption faster than retraining allows
- Alignment gaps at much larger model scales
✅ Already working reasonably well
- Refusing clearly harmful requests
- Transparency about being an AI
- Scoped tool access in production systems
- Public safety research and third-party audits
The practical reality: The models you use every day are tools. Very capable, sometimes surprising tools, but tools. They do not have goals or desires of their own. Some products offer optional memory or saved-context features; availability, defaults, and controls depend on the product, plan, and settings. The risks worth taking seriously are the same risks you take seriously with any powerful technology: who controls it, what it can access, how it is audited, and whether the people building it are honest about what it can and cannot do. Warfare, surveillance, fraud, and discriminatory use deserve more attention than sensational claims about sentient machines.
That is the simple version of how Claude and ChatGPT work. Not magic. A stack of models, data, tools, and guardrails.
Behind one clean reply, the system may be predicting text, searching for context, calling tools, checking results, and managing cost and safety. Once you can see the pieces, the black box starts to look a lot more understandable.