AI Agents, Explained: How AI Went From Answering to Acting
The version of AI that doesn't just answer you — it goes and does the task.

Ask a chatbot "how do I plan a trip to Tokyo?" and it gives you a tidy list of steps. Helpful — but you still have to open the tabs, compare the flights, check the dates, and book everything yourself.
Now imagine telling it "plan my Tokyo trip" — and it actually searches flights, compares hotels, checks your calendar, and comes back with a booked itinerary. That second thing is an AI agent. And the leap between the two is the biggest shift in AI right now.
Across this series we've kept bumping into the word "agent." In Part 1 it was the frontier; in Part 2 we saw that an AI can only act when it's given tools; in Part 3 we learned to brief it like a new intern. This article puts it all together. By the end you'll understand exactly what an agent is, how it works, what it can and can't do in 2026, and how to use one without getting burned. No hype — just the real picture.
So what is an AI agent? An AI agent is an AI system that doesn't just answer questions — it takes actions to accomplish a goal, using tools and working in steps. A regular chatbot is an advisor: it tells you what to do. An agent is a doer: it goes and does it, then reports back.
The intern analogy from Part 3 makes it click. Last time, you briefed a brilliant but deskbound intern who handed you a draft. An agent is that same intern — except now they can get up from the desk, open a browser, send emails, run code, and work through a multi-step task on their own before coming back to you with it finished.
A chatbot answers. An agent acts. That one-word difference is the whole story. The one idea: a brain, tools, and a loop Strip away the buzzwords and every agent is three simple things working together.
How an agent works: think, act, observe, repeat until the goal is met [ Upload c4_loop.png here ]
A brain. At the centre sits a large language model — the same next-word-predicting engine from Part 2 — doing the reasoning: figuring out what to do next.
Tools. On its own, remember, an LLM can only produce text; it can't actually open a website or run a calculation. Tools are what change that. Give the brain access to a web browser, a code runner, your email, your files, a calendar, or any software with an API, and suddenly it can affect the real world.
A loop. This is the secret sauce. Instead of answering in one shot, an agent works the way you would: it thinks about the next step, acts by using a tool, observes what happened, and then loops back to think again — over and over until the goal is met.
Think → act → observe → repeat. That little loop is what turns a clever text generator into something that gets things done. And here's the mindset shift it demands of you: with a chatbot you give instructions ("write me X"). With an agent you give an outcome ("achieve X") and let it work out the steps. You're no longer typing commands — you're delegating.
The leap from chatbot to agent To feel the difference, watch the same request handled both ways.
A chatbot advises; an AI agent acts [ Upload c4_chatbot_vs_agent.png here ]
Ask a chatbot to "analyze our top competitor" and it explains how you might do a competitor analysis. Useful, but the work is still yours. Hand the same goal to an agent and it goes and does it: searches the web, opens the competitor's site, pulls their pricing and product pages, maybe checks reviews, and hands you a finished summary. The underlying model is the same. The difference is that the agent has tools and a loop — so it can take action instead of just describing it.
Watch one work, step by step Let's trace the loop on a real goal: "Find our three biggest competitors and summarize their pricing." Here's roughly what runs inside the agent:
Think: "First I need to know who the competitors are." → Act: search the web. → Observe: gets a shortlist of names. Think: "Now I need each one's pricing page." → Act: open competitor A's site and find pricing. → Observe: captures the plans and prices. Loop again for B and C — and when one site hides its pricing, the agent adapts: it tries the site's search, or checks a review page, instead of giving up. Think: "I have all three; time to write it up." → Act: draft the comparison. → Observe: goal met. → Done. That ability to adapt mid-task — to hit a dead end and try another route — is exactly what separates an agent from an old-fashioned automation script. A script follows fixed rules and breaks the moment reality doesn't match them. An agent reasons its way around the surprise. In a supervised setup, you'd simply approve each action as it goes; in a more autonomous one, it runs the whole loop and hands you the finished summary.
What agents can actually do today This is where it gets real — and where I'll keep it honest. In 2026, agents are genuinely useful for well-defined, repeatable work. The standout categories:
Coding agents. The most mature use by far. They read a codebase, write and run code, fix bugs, and open pull requests — working inside real developer workflows rather than just chatting. Research & browser agents. Give a question and they'll visit dozens of pages, extract what matters, and compile a sourced briefing — the kind of legwork that used to eat an afternoon. Customer support. They resolve common, well-scoped tickets end to end — looking up an order, processing a return — and escalate the messy ones to a human with full context. Workflow & ops automation. Routing invoices, updating records across systems, drafting and filing reports — the repetitive multi-step glue between business tools. Computer use. The newest frontier: agents that operate a screen like a person — clicking, typing, filling forms — to drive software that has no neat API. Notice the common thread, captured well by one 2026 industry overview: agents shine at systems that execute workflows, coordinate tools, and make bounded decisions — not fully autonomous operators that handle every possible edge case. They're best when the task is clear and the boundaries are drawn.
The autonomy dial "Agent" isn't one thing — it's a spectrum of how much you let it do without checking in. Picture a dial.
The autonomy dial: supervised, semi-autonomous, highly autonomous, fully autonomous [ Upload c4_autonomy.png here ]
At the low end, supervised agents propose each action and wait for your approval. A step up, semi-autonomous agents handle routine work themselves and only escalate the exceptions. Highly autonomous agents run within set limits, working through long tasks and checking in only when they hit a boundary. And fully autonomous agents — ones that set their own goals and run indefinitely without oversight — are largely still a sci-fi headline, not a product you can buy.
The reality check that cuts through the hype: in 2026, the vast majority of real deployments sit at the supervised or semi-autonomous end, with a human in the loop. That's not a failure of the technology — it's the sensible default when actions have consequences.
Where agents fall down (the honest part) Every guide that only shows the magic is selling you something. Here's where agents genuinely struggle.
Small errors compound. This is the big one. A single chatbot reply can be 95% reliable and you'd barely notice the misses. But an agent chains many steps, and the errors multiply. If each of 10 steps is 95% reliable, the odds the whole chain succeeds are 0.9510 — about 60%. Add more steps and it drops fast. Long, unsupervised agent runs fail in ways a single answer never would.
A wrong sentence is a typo. A wrong action — a deleted file, a sent email, a placed order — is a problem you can't unsend. No real judgment or common sense. As we saw in Part 2, the model matches patterns; it doesn't truly understand intent. It can confidently take a technically-correct action that makes no sense in your actual situation.
Actions carry real risk. A hallucination (Part 2) is annoying when it's just text. When an agent can act on a hallucination — spending money, changing data, messaging a client — the stakes jump. That's why permissions matter: an agent should only be able to touch what it genuinely needs.
They forget, and they cost. Most agents don't remember you between sessions, and long autonomous runs can be slow and expensive in compute. Useful to know before you point one at an open-ended task.
The rule of thumb: the more steps an agent takes unsupervised, and the more consequential each action, the more a human needs to stay in the loop. Reserve "let it run" for cheap, reversible, well-bounded tasks.
What about "teams" of agents? You'll hear about multi-agent systems — instead of one agent doing everything, several specialized agents split the work: one researches, one writes, one checks the output, with a "manager" agent coordinating. The analogy is exactly what it sounds like: a small team with defined roles instead of a single generalist.
It can be powerful for complex jobs, but be aware of the trade-off — every extra agent adds coordination overhead and another place for things to go wrong. More agents isn't automatically better; it's better only when the task genuinely has separable roles.
How to use agents well Everything you learned about prompting in Part 3 still applies — you're just delegating an outcome now instead of requesting a draft. A few habits that follow directly from how agents work:
Start with bounded, repeatable tasks. The sweet spot is well-defined work with clear success criteria — not vague, open-ended missions. Give a clear goal, the constraints, and the access. Spell out what "done" looks like, what it must and mustn't do, and only the tools and permissions it actually needs. Keep a human in the loop for anything consequential. Approve actions that spend money, send messages, or change important data — at least until you trust the agent on that task. Verify the output. Same golden rule as always: confident completion isn't proof of correctness. Check the work. Prefer reversible over irreversible. Let agents draft, stage, and propose; keep the final irreversible click for yourself when it matters. Why this matters For three articles we've talked about AI that talks. Agents are the moment AI starts to do — and that changes the most valuable human skill alongside it. It's no longer just knowing how to ask; it's knowing how to delegate: how to scope a task, set the guardrails, and supervise the result, exactly as you would with a capable new hire.
The realistic near future isn't an army of fully autonomous robots running your life. It's a growing set of bounded, supervised agents quietly handling the repetitive multi-step work — the research, the routing, the first drafts of action — while humans set the goals and keep judgment where it belongs. The people who learn to direct that well will get an enormous amount of leverage. The ones who handed an agent the keys and walked away will learn why the human-in-the-loop existed.
"Will they take my job?" The honest, evidence-based answer mirrors Part 1: agents automate tasks, not whole jobs, and they're best at the repetitive, well-defined parts. That genuinely reshapes roles — the routine multi-step grunt work is most exposed — but it also raises the value of the things agents are worst at: judgment, taste, knowing which task is worth doing, and supervising the result. The safest career move isn't to out-type an agent; it's to become the person who directs them well.
Key Takeaways A chatbot answers; an agent acts. An agent takes actions to reach a goal, using tools, in steps. Every agent is a brain + tools + a loop: an LLM that thinks, acts with a tool, observes, and repeats until done. You delegate outcomes, not instructions — give it a goal and let it work out the steps. Today's real wins are bounded and repeatable: coding, research, support, workflow automation, and computer use. Autonomy is a dial, and almost all real 2026 use sits at the supervised / semi-autonomous end, with a human in the loop. Errors compound across steps, and actions carry real consequences — so oversight, limited permissions, and verification matter more than ever. Conclusion: from talking to doing That completes the arc. Part 1 told you what AI is. Part 2 opened up how it works. Part 3 taught you to talk to it. And now you can see the next step clearly: agents are simply that same technology, handed tools and a loop, so it can act on the world instead of just describing it.
Strip away the breathless headlines and an agent is wonderfully understandable — a reasoning engine, some tools, and a think-act-observe loop chasing a goal. Knowing that, you can spot what's real, ignore what's hype, and start handing off the boring multi-step work while keeping your hand on the wheel.
The future of AI isn't a smarter answer. It's a finished task — with you deciding which tasks, and how far to trust it. Try this: pick one small, repeatable task you do every week — sorting an inbox, gathering links for a report, drafting routine replies — and imagine briefing an agent to do it. What goal, what limits, what would you check before trusting it? That's the new core skill. Tell me the task you'd hand off first in the responses.





