edidiong umana · writing
home
Engineering notes5 min read

What an agent actually is: a loop with a model, tools, memory and a budget

Strip away the hype and an agent is a short loop in your own code. You get that loop in runnable Python with hard step and cost limits, a test for when a plain workflow is the better build, and the rules for tools and MCP that keep it honest.

You've been asked to "add an agent" to a product. The demos look like magic, the frameworks have long feature lists, and nobody can tell you exactly what you're building. So it's hard to estimate, hard to test and hard to say no to.

Here is the plain version. An agent is a loop in your code. On each turn a model reads some context and either answers or asks for a tool. Your code runs the tool, adds the result to the context, and goes round again. Four parts make it work: a model that decides, tools that act, memory that your code builds, and a budget that makes it stop.

Anthropic's guide to building agents puts it in one line: agents are "typically just LLMs using tools based on environmental feedback in a loop."

First, ask whether you need one

There are three shapes of LLM app. A chatbot takes a message and returns text. A workflow runs the model through steps you fixed in code: classify the message, fill a template, send it. An agent lets the model choose the steps.

Workflows are faster, cheaper and easier to test, because the path never changes. Anthropic advises reaching for multi-step agent systems only once simpler approaches have proved not good enough. Agents earn their cost on open-ended tasks: ones where nobody can say in advance how many steps they will take, or write the path down as code.

My test is short. If you can draw the steps on a whiteboard and they're the same every time, write a workflow. An order confirmation that always does the same three things is a workflow. "Find out why this customer's parcel is late and what we can do about it" might need an agent, because the next step depends on what the last one found.

The loop, in about 60 lines

The model never runs anything itself. It only asks. Your code decides whether to do what it asks, and that gap is where every safety rule lives. Below is a complete loop you can run with plain Python 3 and no API key. A stub function stands in for the model, so you can watch the loop and its limits without spending anything.

import json

# Tools: plain functions the model can ask for. Fake data, so it runs offline.
ORDERS = {1042: {"items": "2 phone cases", "total_usd": 12, "paid_usd": 12}}

def get_order(order_id):
    order = ORDERS.get(order_id)
    if order is None:
        return {"error": f"order {order_id} not found; check the number"}
    return {"order_id": order_id, "items": order["items"], "total_usd": order["total_usd"]}

def check_payment(order_id):
    order = ORDERS.get(order_id)
    if order is None:
        return {"error": f"order {order_id} not found; check the number"}
    return {"order_id": order_id, "paid_in_full": order["paid_usd"] >= order["total_usd"]}

TOOLS = {"get_order": get_order, "check_payment": check_payment}  # the closed list
PRICE_PER_1K_TOKENS = 0.01  # made up for the demo: use your provider's real price

def stub_model(messages):
    """Stands in for a real model API. Returns (reply, tokens_used)."""
    results = sum(1 for m in messages if m["role"] == "tool")
    if results == 0:
        reply = {"tool": "get_order", "args": {"order_id": 1042}}
    elif results == 1:
        reply = {"tool": "check_payment", "args": {"order_id": 1042}}
    else:
        reply = {"answer": "Order 1042 is paid in full (12 USD). You can ship it today."}
    tokens = sum(len(m["content"]) for m in messages) // 4 + 30  # about 4 characters per token
    return reply, tokens

def stuck_model(messages):
    """A model that never finishes, to show the limits working."""
    return {"tool": "check_payment", "args": {"order_id": 1042}}, 450

def run_agent(goal, model, max_steps=6, max_cost_usd=0.02, max_errors=2):
    messages = [{"role": "system", "content": "You help a shop owner. Use the tools. Never guess."},
                {"role": "user", "content": goal}]
    spent, errors = 0.0, 0
    for step in range(1, max_steps + 1):
        reply, tokens = model(messages)                        # decide
        spent += tokens / 1000 * PRICE_PER_1K_TOKENS
        if "answer" in reply:                                  # done: no tool requested
            return f"{reply['answer']} [{step} steps, ${spent:.4f}]"
        if spent >= max_cost_usd:                              # budget
            return f"Stopped: budget used up (${spent:.4f}) after {step} steps."
        name, args = reply["tool"], reply["args"]
        try:                                                   # act: your code runs the tool
            result = TOOLS[name](**args)
        except Exception as err:                               # unknown tool or bad arguments
            result = {"error": f"{type(err).__name__}: {err}"}
        if "error" in result:
            errors += 1
            if errors >= max_errors:                           # error limit
                return "Stopped: tools kept failing. Handing over to a person."
        print(f"step {step}: {name}({args}) -> {result}")
        messages.append({"role": "assistant", "content": json.dumps(reply)})
        messages.append({"role": "tool", "content": json.dumps(result)})  # observe
    return f"Stopped: step limit ({max_steps}) reached without an answer."

print(run_agent("Did order 1042 pay, and can I ship today?", stub_model))
print(run_agent("Did order 1042 pay, and can I ship today?", stuck_model))
print(run_agent("Did order 1042 pay, and can I ship today?", stuck_model, max_steps=3))

Run it and you get three runs: one that finishes, one stopped by the budget and one stopped by the step limit.

step 1: get_order({'order_id': 1042}) -> {'order_id': 1042, 'items': '2 phone cases', 'total_usd': 12}
step 2: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
Order 1042 is paid in full (12 USD). You can ship it today. [3 steps, $0.0023]
step 1: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
step 2: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
step 3: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
step 4: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
Stopped: budget used up ($0.0225) after 5 steps.
step 1: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
step 2: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
step 3: check_payment({'order_id': 1042}) -> {'order_id': 1042, 'paid_in_full': True}
Stopped: step limit (3) reached without an answer.

To use a real model, replace stub_model with an API call and read the token counts from the response instead of estimating them. OpenAI-format responses report usage.prompt_tokens and usage.completion_tokens, and Ollama's compatible endpoint does too. The loop itself doesn't change.

Two details matter. The try block turns an unknown tool or bad arguments into an error the model can read, instead of a crash. And the budget check runs after each call, because you only know the cost once the call returns. That means a run can overshoot the cap by one call, so set the cap with that margin in mind.

Stopping is a feature

A loop without an exit burns money. Every agent needs several ways to stop, and the loop above has four of them:

  • Done: the model answers without asking for a tool. The normal finish.
  • Step limit: a maximum number of turns, such as 6 or 10. The most important safety net you'll write.
  • Budget: stop when a run has used too many tokens, too much money or too many seconds.
  • Error limit: stop after repeated tool failures instead of retrying forever.

The fifth is a human checkpoint: pause before anything risky, such as sending money or messages, until a person approves. When any limit fires, say so plainly and log it. "I couldn't finish this" beats a confident, made-up answer.

Put a second limit outside your code too, because code limits have bugs. OpenAI projects can have a monthly spend limit enforced as a hard cap. Use one key per agent, so you can revoke one without breaking the others.

Tools are an interface for a reader who guesses

The model picks tools by reading their names and descriptions, so write them as you would for a new colleague. The ideas below come from Anthropic's guide to writing tools for agents:

  • Say when to use it, not just what it does: "Check whether an order is paid. Use before confirming shipping."
  • Tight arguments. Types, required fields and enums make wrong calls hard to make.
  • Small results. Return the few fields the model needs, not a database dump. Results stay in the context for every later turn.
  • Helpful errors. "Order 1042 not found; check the number" lets the model recover. A stack trace doesn't.
  • Few, distinct tools. Every definition is sent with every call, and overlapping tools confuse the model.

Validate every call before you run it. OpenAI's own SDK warns that arguments may not be valid JSON and may include parameters you never defined. In the loop above, TOOLS is a closed list: the model can only reach what you put there.

MCP, in plain language

Without a standard, every AI app needs its own integration for every tool. The Model Context Protocol (MCP) is an open standard for connecting AI applications to tools and data. Anthropic introduced it in November 2024 and donated it to the Linux Foundation's Agentic AI Foundation in December 2025.

Three roles. The host is the app a person uses, such as Claude Desktop, VS Code or your own agent. Inside it, each client holds one connection to one server. A server offers tools (actions), resources (data to read) and prompts (reusable templates). Local servers talk over stdio, remote ones over Streamable HTTP, and the messages are JSON-RPC 2.0.

Build a server once and any MCP host can use it. For a single agent, plain function tools like the ones above are fine. MCP pays off when the same tools should work in several apps.

One warning. An MCP server runs with your permissions, and its tool descriptions go straight into the model's context. Invariant Labs showed "tool poisoning", where a malicious server hides instructions inside a description. Install servers only from sources you trust, and read what each tool can do.

Memory is something your code builds

The model forgets everything between calls. A chat feels continuous only because the app re-sends the history every time. So memory is a design choice, and there are three common kinds: files the agent reads and writes, like Claude Code's CLAUDE.md or OpenClaw's MEMORY.md; retrieval, which searches a large store and inserts only the matching pieces; and summaries that replace old turns with a short recap.

More context isn't better. Chroma's 2025 study of 18 models found performance grew increasingly unreliable as input length grew, even on simple tasks. Send what the next step needs and little else. The self-hosted harnesses I run on a small VPS, OpenClaw, Hermes Agent and NanoClaw, are versions of this same loop with more tools and memory around it.

Do this todayOpen whatever agent you're building and search for the loop. If it has no step limit and no budget check, add both before you change anything else.

Learn it properly

This post is a slice of Track 1 of the AI Study Group, Agentic AI engineer roadmap. It's free and self-paced, and it goes from tokens and prompts to a tool-using agent with evals, running around the clock on a small server.

Sources