AI · Automation · Engineering

What I'd Learn in an Agentic AI Course (Shipping Agents)

By Lazar MilicevicOctober 10, 20269 min read
Lines of code on a dark monitor, representing building and shipping agentic AI systems in production

I run multi-agent systems that research, write, optimize and publish content with nobody watching. When I look at what agentic AI courses teach, I recognize about half of what I had to learn. The other half I learned the expensive way, at 3 a.m., reading a log of an agent that had politely looped itself into a bill.

This is a breakdown of what courses cover well, what they skip, and the self-study path I'd follow if I were starting today with production in mind.

What the good courses actually get right

The best courses teach the vocabulary and the core patterns, and that matters more than it sounds. Andrew Ng's Agentic AI course on DeepLearning.AI covers four design patterns: reflection, tool use, planning and multi-agent collaboration, and it builds a deep research agent as the worked example. Those four patterns are the real skeleton of almost every agent I have shipped. If you cannot name which of the four a piece of your system is using, you cannot debug it.

The Hugging Face AI Agents Course takes a different route: it teaches established libraries (smolagents, LlamaIndex and LangGraph), and its certification is free with no deadline. Its final project is worth noting for a specific reason. You are scored on a subset of the GAIA benchmark, and you need 30% or higher to earn the Certificate of Completion (source). That is the first time many learners face an uncomfortable truth: agents fail a lot, and you measure the failure rate or you do not know it.

Both are legitimate starting points. Neither will tell you what happens in week six of production.

The one thing I'd weight above everything else: evals

Andrew Ng says the most important part of his course is a disciplined process for evals and error analysis, and that it is the biggest predictor of whether a team executes well (source). I agree with this more than with anything else in the curriculum, and it matches my own experience.

In my content system, the most useful thing I built was not a smarter agent. It was the measure-learn-target loop: the system looks at what real search performance says about published pages, feeds that back into what it targets next, and generates again. That loop is an eval harness pointed at reality instead of a test set. Before I had it, I was tuning prompts by vibes. After, every change had to survive contact with actual impressions and clicks.

The practical version for anyone building an agent:

  1. Write the failure list first. Before you touch the prompt, run 30 to 50 real inputs and write down, in plain words, every way the output was wrong.
  2. Categorize the failures. Wrong tool, bad retrieval, ignored instruction, format break, hallucinated fact. The distribution tells you where to work.
  3. Fix the largest bucket only. Re-run the same inputs. Did that bucket shrink without growing another?
  4. Keep the set. It becomes your regression suite. Every production incident adds a case.

Courses teach you to do error analysis. What they cannot teach is the discipline of doing it when the demo already looks fine.

Planning loops: the loop must be able to stop

Anthropic's Building effective agents draws the line between workflows, which give you predictability and consistency on well-defined tasks, and agents, which are better when you need flexibility and model-driven decisions at scale. Its advice is blunt: find the simplest solution possible and add complexity only when needed, which may mean not building an agentic system at all (analysis).

That is the most undersold lesson in agent education. A lot of what I have shipped that runs reliably is a workflow with one or two narrow agentic steps, not a free-roaming agent. The pipeline stages (research, draft, optimize, publish) are fixed. The model makes decisions inside a stage, not about which stage comes next.

Where I do use an open loop, two properties are non-negotiable, and both appear in Anthropic's pattern: the loop ends on completion or a stopping condition such as a maximum number of iterations, and the agent gets ground truth from the environment (tool results, code execution) at each step (source). The OpenAI Agents SDK bakes the same idea in: the loop re-invokes the model after tool calls until there is a final output or a max-turn limit, which doubles as a runaway-cost guard.

A minimal version of the discipline, framework-agnostic:

MAX_STEPS = 8

def run_agent(task, tools, llm):
    state = {"task": task, "history": []}
    for step in range(MAX_STEPS):
        action = llm.decide(state)
        if action.is_final:
            return action.output
        result = tools[action.tool](**action.args)   # ground truth from the world
        state["history"].append((action, result))
    # Do NOT silently return a half answer. Fail loudly and log it.
    raise AgentStalled(task, state["history"])

The last line is the lesson. A budget that expires into a confident-sounding partial answer is worse than a crash, because nobody notices it.

Tool use and the protocol layer is moving

Tool use is taught as "define a function, let the model call it." In production the hard parts are narrower. Tool descriptions are prompts, and vague ones cause wrong calls. Tool errors must come back as readable text the model can act on, not stack traces. And every tool with side effects needs a permission model, which is why I sandbox anything high-risk.

One thing worth watching if you invest in MCP: the protocol is changing underneath you. The 2026-07-28 MCP specification removes the initialize/initialized handshake and the Mcp-Session-Id header, making the protocol core stateless, and it shipped with updated Tier 1 SDKs for TypeScript, Python, Go and C# (Cloudflare). If your course material predates that, treat its MCP chapter as concepts, not copy-paste code. Check the official changelog for deprecations before you build on a specific feature.

A guardrail detail that bit people I know: in the OpenAI Agents SDK, tool guardrails do not apply to the handoff call itself, because handoffs run through a separate pipeline from function tools (docs). If you assumed your input checks covered agent-to-agent transfers, they do not. Read the guardrail docs of whatever framework you adopt for exactly this kind of gap.

Memory: what courses call it vs what you build

Courses present memory as a feature. In my systems it is three separate decisions that I keep apart on purpose:

Kind What it holds Where I keep it Failure mode
Working context The current run's steps and tool results The prompt, trimmed Context bloat, lost instructions
Run state What this job has done so far A database row, not the prompt Agent re-does or contradicts itself
Learned knowledge What worked in past runs (e.g. search performance) Postgres, retrieved on demand Stale or misleading lessons

The shortcut that fails is stuffing everything into the prompt and hoping. The pattern that works is treating the database as the agent's real memory and the prompt as a small, curated view of it. This is where my RAG and hybrid search background pays off directly: retrieving the right three facts beats carrying thirty.

Cost control: the module nobody takes

This is the biggest gap I see between courses and production. Unattended systems have a bill that nobody is watching, so cost has to be designed in.

Three levers I actually use, with numbers from Anthropic's pricing documentation. Check the pricing docs for current base prices, since models and rates change and third-party pages disagree about which are current:

  • Prompt caching. A cache hit costs 5% of the standard input price, listed at $0.20/MTok for Claude Opus 5.5 and $0.10/MTok for Claude Sonnet 5.5 (source). Writes cost more: 1.25x base input for the 5-minute TTL and 2x for the 1-hour TTL (source). The practical rule: put the large, stable material (system prompt, style guide, tool definitions) first and the variable part last, so repeated calls hit the cache.
  • Batch processing. The Claude Batch API halves both input and output token rates (source). For any pipeline stage that does not need a real-time answer, this is close to free money. If a stage runs on a schedule and nobody is waiting on it, latency is irrelevant, so the discount costs you nothing.
  • Step budgets. The max-iteration cap above is also a cost cap. An illustrative example with round numbers: if one step costs a cent, an uncapped loop that spirals to 500 steps is a five-dollar mistake per run, multiplied by every run per day. The cap makes the worst case a known number.

I deliberately do not quote an "agents cost N times a chat" multiplier. The figures that circulate are not measured data. Measure your own system with tracing instead.

Observability: turn tracing on before you need it

The OpenAI Agents SDK records LLM generations, tool calls, handoffs, guardrails and custom events by default, and you can disable it with OPENAI_AGENTS_DISABLE_TRACING=1. Whatever stack you use, the lesson is the same: you will debug agents by reading traces, so make tracing the default and opt out deliberately. An agent failure is rarely in the last step. It is a bad tool result three steps earlier that the model trusted.

What I'd do

If I were starting from zero and wanted production skills, in this order:

  1. Take one structured course for vocabulary. Ng's Agentic AI course for the four patterns and the error-analysis habit, or the Hugging Face course if you prefer free and library-based. Do not collect certificates. Verify current details (length, plan, certificate) on the official page first, because secondary sources conflict.
  2. Read Anthropic's "Building effective agents" twice. Once now, once after you ship something. It reads differently the second time.
  3. Build one small unattended agent with a real consequence. A scheduled job that does something you would notice if it broke. A toy chatbot teaches you nothing about stopping conditions or cost.
  4. Write the eval set before the second prompt revision. Even 30 cases.
  5. Add the step cap, the cost log and tracing on day one, not after the first surprise bill.
  6. Prefer a workflow with one agentic step over a free agent until you can prove you need the freedom.

Courses give you the map. The terrain is the failure log of something running while you sleep.

If you are working through this and want to compare notes on a production agent, you can reach me at lazar-milicevic.com/#contact, or keep reading the other write-ups on the blog.

Frequently asked questions

What do agentic AI courses teach, and what do they leave out?

The good ones teach the vocabulary and core patterns: reflection, tool use, planning and multi-agent collaboration. Courses like Andrew Ng's on DeepLearning.AI and the Hugging Face AI Agents Course cover these well. What they tend to skip is what happens in week six of production, such as runaway loops, cost control and keeping evals alive after the demo looks fine. I recognize about half of what courses teach, and the other half I learned the expensive way.

What are the four agentic design patterns?

The four patterns taught in Andrew Ng's Agentic AI course are reflection, tool use, planning and multi-agent collaboration. They form the skeleton of almost every agent I have shipped. If you cannot name which of the four a given part of your system is using, you cannot debug it. Knowing the patterns gives you a shared vocabulary for diagnosing failures.

How do I evaluate and debug an AI agent?

Start by running 30 to 50 real inputs and writing down every way the output was wrong before touching the prompt. Then categorize the failures (wrong tool, bad retrieval, ignored instruction, format break, hallucinated fact) and fix only the largest bucket, re-running the same inputs to check you did not grow another. Keep that set as a regression suite and add a case for every production incident. In my own system, a loop that feeds real search performance back into targeting works as an eval harness pointed at reality.

When should I build a workflow instead of an autonomous agent?

Workflows give predictability and consistency on well-defined tasks, while agents suit cases that need flexibility and model-driven decisions. Anthropic's guidance is to find the simplest solution possible and add complexity only when needed, which may mean no agentic system at all. Much of what I run reliably is a fixed pipeline (research, draft, optimize, publish) with one or two narrow agentic steps. The model decides inside a stage, not which stage comes next.

How do I stop an AI agent from looping forever or running up costs?

Give every open loop a stopping condition, such as a maximum number of iterations, and make the agent take ground truth from the environment at each step, like tool results or code execution. The OpenAI Agents SDK does the same by re-invoking the model until there is a final output or a max-turn limit, which doubles as a runaway-cost guard. When the budget runs out, fail loudly and log the history instead of returning a confident-sounding partial answer. A silent partial answer is worse than a crash because nobody notices it.

Lazar Milicevic

Lazar Milićević

Senior Technical Engineer. I build AI automation, GenAI/LLM systems and cloud architecture — autonomous systems that run while you sleep. Founder of BizFlowAI.

Building something hard with AI or automation? I am open to talk.

Get in touch

← All posts