edidiong umana · writing
home
Career playbooks5 min read

Before the agent: scoping an AI deployment

Eight questions to answer before you write code when an organisation asks for "an AI agent", drawn from Anthropic's and OpenAI's published guidance, with a one-page template you can copy.

A manager tells you, "We want an AI agent for customer support." It sounds like a spec. It isn't. It doesn't say which requests, what the agent may touch, what counts as success, or what happens when it gets something wrong. Start coding from that sentence and you'll build a demo that handles the easy questions well, and nobody on the team will trust it with the hard ones.

Scoping turns that sentence into answers. Below are eight questions, in order, plus one before them all. The two labs that publish most about agents agree on that first one: you may not need an agent at all.

Question zero: does this need an agent?

Anthropic's Building effective agents advises finding the simplest solution that works and adding complexity only when it's needed, which might mean no agentic system at all. It separates workflows, where LLMs and tools follow code paths you wrote in advance, from agents, where the model directs its own process and tool use. Agentic systems usually trade latency and cost for better task performance, and for many applications one well-built LLM call with retrieval and examples is enough.

OpenAI's A practical guide to building agents gives a test. Agents earn their place in workflows that have resisted automation, of three kinds:

  • complex decisions full of exceptions and judgement, such as refund approval;
  • rules too intricate to maintain, such as vendor security reviews;
  • heavy unstructured data, such as reading documents to process a home insurance claim.

If the use case doesn't clearly fit one of these, the guide says a deterministic solution may be enough. Write down which kind applies, and why plain rules won't do the job.

1. The job to be done

Name the job as an outcome for a person, not as a feature. "A support agent" is a feature. "Customers with a damaged order get a correct refund decision within an hour" is a job. Scope the first version to one request type, and write down what is out of scope. That list will save you more arguments than anything else on the page.

2. The current process and its cost

Map how the job is done today: who does it, in which systems, how long it takes and where it goes wrong. Then put a number on it. Here is a worked example with made-up inputs:

  • The team handles 1,200 refund requests a month at about 9 minutes each: 10,800 minutes, or 180 staff hours.
  • Suppose 30% are simple cases an assistant could draft for a person to approve in 2 minutes. That's 360 requests saving 7 minutes each: 2,520 minutes, or 42 hours a month.

Those 42 hours are the most version one can return. Compare them with what the system costs to build and run. The same number becomes your baseline later.

3. Data access

List the data the system must read and the systems it must act on. For each, note where it lives, who owns it, how you'll get access and what rules cover it, especially personal data. OpenAI's guide notes that where old systems have no API, agents can work through their screens with computer-use models. If that's your only route, you want to know in week one, not week six. Ask for read-only access first. And ask for a sample of real, anonymised cases: they become your eval set.

4. Risk and irreversibility

List every action the system could take. OpenAI suggests rating each tool low, medium or high risk by whether it reads or writes, whether its effects can be undone, which account permissions it needs, and its financial impact. High-risk tools should pause for extra checks or go to a person. Anthropic adds that autonomy brings higher costs and the potential for compounding errors, and recommends extensive testing in sandboxed environments with guardrails.

ActionReads or writesUndo?Money moves?Rating
Look up order statusReadsNothing to undoNoLow
Draft a reply for staffWrites a draftYesNoLow
Send the reply to the customerWritesNoNoMedium
Issue a refundWritesHardYesHigh

5. Success metric

Pick one primary metric tied to the job, with a target, plus one or two guard metrics that must not get worse. For the refund case: share of simple requests resolved correctly without rework, with zero wrong refunds and no rise in complaints as guards. Anthropic observes that agents add the most value where tasks have clear success criteria, feedback loops and meaningful human oversight. If you can't write the metric down, that's a scoping finding, not a detail to settle later.

6. Eval plan

Decide how you'll know it works before you build it. OpenAI recommends setting up evals to establish a baseline, prototyping with the most capable model, and then trying smaller models where they still meet the target. Anthropic's Demystifying evals for AI agents says 20 to 50 simple tasks drawn from real failures is a great start, and that a good task is one where two domain experts would independently reach the same pass or fail verdict. Your anonymised cases from question 3 are the raw material.

7. Human in the loop

OpenAI names two triggers for handing control to a person. One is exceeding failure thresholds: set limits on retries or actions, and escalate when the agent goes past them. The other is high-risk actions, meaning sensitive, irreversible or high-stakes ones such as cancelling orders, authorising large refunds or making payments, which should get human oversight until confidence in the agent grows. Anthropic describes agents pausing for human feedback at checkpoints, with stopping conditions such as a maximum number of iterations.

Name the person, not "a human": who approves, in which tool, and how quickly.

8. Rollout

OpenAI reports that customers typically do better with an incremental approach than by building a fully autonomous system straight away, and its advice is to start small, validate with real users and grow capabilities over time. One sequence that fits that advice: the system drafts while staff still do the work, then staff approve each action, then low-risk actions run on their own. Agree in advance what triggers a pause, and who can pull it.

The one-page template

Copy this into a doc and fill it in with the people who do the work today. If a line stays blank, that's your next meeting.

AI DEPLOYMENT SCOPE: [name]          Owner: [name]   Date: [date]

0. Agent or not?   Kind: complex decisions / intricate rules /
                   unstructured data. Why rules alone won't do:
1. Job to be done: [person] gets [outcome] within [time].
   In scope (v1):            Out of scope:
2. Current process: steps, people, systems, time per case.
   Volume/month:     Minutes/case:     Staff hours/month:
   Expected hours saved (v1):          Cost to build and run:
3. Data access: source | owner | access route | personal data? | read/write
4. Actions: action | reads/writes | undo? | money? | low/med/high
5. Primary metric + target:          Guard metrics:
6. Evals: 20-50 real cases, pass/fail rule two experts agree on,
   baseline before build, rerun on every change.
7. Humans: failure threshold that escalates:
   High-risk actions needing approval:     Approver + tool + speed:
8. Rollout: drafts only -> approve each action -> low-risk alone.
   Pause trigger:                      Who can pause:
Do this todayTake the AI request in front of you and answer question zero in three written sentences before you open an editor.

Learn it properly

Question 6 is where most projects go thin. Track 8 of the AI Study Group, Evals and observability, is free and walks through error analysis, test sets and graders around one worked example.

Sources