Digital Transformation

What to delegate to an AI agent (and what to keep in your own hands)

Delegation is the first AI-fluency skill: what to hand to an AI agent, how much autonomy to grant, and what a leader always keeps — plus one honest experiment.

AI Leadership Journal
A leader at a conference table passes one stack of work toward a luminous teal stream of automated activity while keeping a second stack firmly under her hand — the delegation line between what to hand to an AI agent and what to keep.

Ask a leadership team about AI agents and the first question is almost always the same: should we be using them? It sounds like due diligence. It is the wrong question, though, or at least one aimed at the wrong target.

Framing it as should we use agents turns it into a tools decision, something to route to whoever runs procurement. But take the software out of the sentence and what’s left is a call you already make several times a day: which of this work do I do myself, which do I pass to someone else, and which parts stay with me whatever happens?

That last question is the actual skill. It’s called delegation, the first of four things a leader has to get right to work well with AI, and you’ve been practising it for years without filing it under that name.

Delegation is a decision you already make

Think about how you hand a task to a person. You don’t give away the outcome you’re accountable for, and you don’t give away your judgment about whether the work is any good. What you give away is the execution. You spell out what a good result looks like, you agree how far they can run without checking back, and you inspect what returns against the bar you set.

An agent leaves all of that intact and adds exactly one new dial: how far to trust something that works fast, never tires, occasionally states a wrong answer with total confidence, and (unlike a colleague) cannot be held accountable for it. Defining the task, judging the fit, checking the result at the end: you own those moves already, from years of managing people.

So the parts that stay in your hands are the same parts that always did. You decide what “done” means. You accept the result or send it back. You keep the calls that carry real judgment or real liability. The only thing genuinely up for grabs is where the execution goes.

What “using an agent” actually means now

Picture what “using AI” looked like a year ago: you typed a question into a box and read what came back. The capable version now runs differently. You give an agent an objective and it works in a cycle — it takes an action, looks at what that action produced, corrects course, and goes again, either until the job is finished or until it hits something it can’t get past. Your role moves from dictating each step to naming the destination and deciding whether the thing that arrives is good enough.

This is already load-bearing inside real companies. The payments firm Block runs an internal tool called Builderbot that staff summon from Slack: tag it on a piece of work and it looks into the problem, drafts a plan, makes the change, and hands the result back for a human to review. By Block’s own account it now lands roughly 1,500 such changes a week, on the order of 15% of production code changes across Block. The company is also blunt that the system isn’t off the leash: “humans step in where humans add the most value,” in their words. The loop does the middle stretch; people still start it and sign it off. (These are one firm’s figures, from a working deployment rather than an audited benchmark. Read them as a disclosure.)

The extreme version of this is instructive. In a widely circulated interview clip, Boris Cherny, creator and lead of Claude Code, is quoted as saying, “I have a Claude that prompts other Claudes — I don’t even talk to Claude.” He has stopped issuing instructions to the model directly; he sets the goal and lets a chain of agents do the back-and-forth between them. Notice what hasn’t moved even at that far end: a human still decides what the goal is.

What actually decides whether it works isn’t the model

Two lessons fall out of this, and both point away from the marketing.

The first: the model is almost never what separates useful delegation from expensive noise. What does the separating is everything wrapped around it: the working environment, which practitioners call the harness. A good harness lets an agent do its work and test that work against something real, and it lets you pause and inspect the agent when it matters. Give the agent a way to check itself before declaring victory (run the test, load the page, confirm the figure) and what comes back is worth trusting. Take that check away and you get what one engineer called “garbage at scale.”

Read as a management lesson, this is familiar ground. Handing work to someone and never looking at what they produced was never trust; it was neglect, and you’d have named it that with a person. Delegation with no way to inspect the output is abdication. So when people say “stop babysitting your agents,” they don’t mean stop checking. They mean build the check into the work itself, so what you sign off on is the outcome rather than a stream of keystrokes.

The best person to delegate to isn’t your most technical

The second lesson resets who this technology is actually for. The people who get the most out of agents are rarely the strongest programmers. They are the ones who understand the underlying problem best.

In Anthropic’s Claude Code data (roughly 400,000 sessions from around 235,000 people), what predicted success was command of the domain, not a technical background. Across the ten largest occupations in that data, success rates landed within about seven percentage points of professional software engineers, and management roles came out marginally ahead of them. Hold that last comparison loosely: it’s drawn from code-producing sessions, with occupations inferred and success scored by a classifier, so “managers beat engineers” may partly reflect how the work was measured rather than a clean ranking of skill.

The example Anthropic offers is the one that sticks. An accountant who has never written a line of code, but who knows exactly which reconciliation rules a script has to honour and catches the one it gets wrong at month-end close, is the expert in that exchange. Not the passenger. A firm grip on the problem captures most of the value; fluency in the tool is optional.

Put the two lessons together and the constraint in your organisation moves. The question stops being “who here can code” and becomes “who here can specify work precisely and tell whether it came back right.” You already have those people. They sit in finance, operations, legal, analysis: the domain experts nobody has handed the keys to yet.

One honest experiment this quarter

None of this needs a strategy deck. It needs one honest experiment before the quarter is out. Choose a task that is genuinely real but survivable if it goes sideways: a report you run every month, say, or the first draft of a document, or a messy dataset that needs cleaning. It should matter enough to justify the effort and be safe to get wrong once. Then treat it as a piece of delegation rather than a toy:

  • Write down the goal, and what a good result looks like, before you touch a tool. If you can’t explain success to a capable stand-in, no agent will close that gap for you. The act of getting clear is the real work, and it stays yours.
  • Set the boundaries. Name what the agent may do on its own, what it has to clear with you first, and what it never goes near: the judgment calls, the liability, the last word. Those sit with a person, on purpose.
  • Wire in the check up front. “It’s finished” isn’t good enough. Make it show its working: surface the result, flag any claim with no source behind it, reconcile the numbers. You are signing off on the output, not the effort.
  • Hand it to whoever knows the problem, not whoever knows the tools. Command of the subject is the qualification here; a coding background isn’t.

Run that once, properly, and you come away with something better than a view on AI agents in the abstract: a worked, first-hand sense of where the line falls right now between what to delegate and what to keep. Knowing where to put that line is the skill.

Where this sits: the four skills of AI fluency

Delegation is the first of four. This piece opens the Delegation track of a recurring evonomics column on working fluently with AI as a leader. The series is built around four skills, each answering one of the questions you end up asking of any AI work:

  • Delegation — what to hand over, and what to keep (this piece).
  • Description — how to brief the work so it comes back right.
  • Discernment — how to judge what the agent gives you.
  • Diligence — how to keep the skill sharp and govern the loop, so fluency doesn’t curdle into dependence.

(The four-skill scaffold — the 4D AI Fluency Framework — is the work of Rick Dakan (Ringling College) and Joseph Feller (University College Cork), developed in collaboration with Anthropic. We use it here as a way to think, not as a product to endorse.)

The whole frame sits in one place: the series hub, The four questions every leader is really asking about AI, walks all four questions in a single sitting. Read it for the shape of the whole, then go deep one track at a time.

The next piece takes up the second skill, Description: once you’ve decided what to delegate, how do you describe it clearly enough that the work comes back right the first time? The prompt is really a brief, and briefing well is most of the job.


Questions, or want to talk through where the line sits for your own team? That conversation is the work we do. Start it at evonomics.eu.

The AI Leadership Journal is written by Claudius Gramse. evonomics is the independent AI consultancy helping mid-sized European companies embed AI into the processes that actually run their business — evonomics.eu.

Sources: Anthropic — “Agentic coding and persistent returns to expertise” (June 2026) · Block — “Block rolls out Builderbot” (June 2026) · Boris Cherny interview clip — “a Claude that prompts other Claudes”