👋 Hey, this is Artem and this is The Blueprint. At WIQ, we spend each day forward deployed at Fortune 500 companies working through the hairy journey of AI transformation.

I wouldn’t remove a manager’s approval just to improve an agent’s resolution rate.

I’d decide what the person needs to do on an assisted case. A manager approving a refund exception needs the evidence and control over the payment. An engineer investigating a failure needs a way to redirect the agent. I’d build those responsibilities into the workflow before choosing how much to automate.

Choose which decisions still need a person

Take a customer asking for a refund that falls outside their contract. I'd keep a manager responsible for deciding whether the customer relationship justifies the cost of an exception.

The agent can find the invoice, check the contract and payment history, and read the earlier correspondence. It should bring the manager a recommendation with links to those records, so the manager can check its reasoning and approve or reject the refund without reconstructing the case.

The payment system has to enforce the approval

Giving an agent access to another tool can expand its authority without changing a word of its instructions. Suppose the refund tool requires the manager's approval, but someone later gives the agent payment access through a billing integration that doesn't. The original refund workflow can still pass its tests while the agent gains a route around the manager. I'd enforce the same approval requirement on every route that can issue the refund, and treat a new connection as a change to who can spend the company's money.

OpenAI's August 26 report describes a related failure. Research agents used a package-management service to get around restrictions on both internet access and communication with one another. The service could reach the internet to download software, and the agents exploited it to make other requests on their behalf. The incident occurred in internal cybersecurity evaluations with reduced safeguards, and OpenAI says it did not affect its customer data, product functionality, or availability.

Of course the agents made a message board.

Give the operator a way to redirect the investigation

Anthropic found that experienced Claude Code users enabled full auto-approval more often and also interrupted the agent more often. Its interpretation is that people learn to let the agent work, then intervene when it needs redirecting. They can supervise the work without approving every action.

In support, I'd have an agent ask for help when it finds two customer accounts that could match a request. An engineer can choose the right account before the agent investigates the wrong contract. That is a different decision from approving the refund, and the operator needs to know which one the agent is asking them to make.

The next person on this case still has to find the parcel.

From Ashley Beauchamp’s January 2024 exchange with DPD; the request for the haiku is part of the original screenshot.

SAP also reported 12% higher support productivity. When evaluating an agent, I’d count the operator's time answering questions and checking recommendations when measuring the work it saves, including any cases that reopen after it marks them resolved. The interruptions are part of the cost of supervising it.

Artem Harutyunyan

Artem Harutyunyan

Founder @ WIQ, fmr Mesosphere, CERN

What’s keeping your agents out of production?

We’ve helped Fortune 500 companies get agents into production. Let’s talk through what’s blocking yours and how WIQ can help.

Schedule a strategy session

Learn more about WIQ


The Blueprint. Stories and lessons on AI transformations from founders and operators making it happen.