An agent can give a convincing answer in a demo while leaving most of the production problem untouched.
An agent can give a convincing answer in a demo while leaving most of the production problem untouched.
Consider a hypothetical refund workflow. The agent reads a customer's message, finds the order, and recommends a refund. Before it can issue one, the business needs to establish whether the order record is current, whether a refund has already been paid, and whether this request requires a manager's approval. If the payment system times out, the workflow must determine what happened before trying again.
Those questions are easy to leave outside a pilot. They become unavoidable when the agent can move money.
Our argument is that teams need to test the complete workflow earlier. Useful reasoning is one requirement. The workflow also has to use reliable information, act within its authority, and produce an outcome someone can verify. Until those requirements have been tested, expanding autonomy means expanding uncertainty.
MIT Technology Review Insights surveyed 300 senior data and technology executives at companies with at least $500 million in annual revenue. Respondents reported that, on average, 34% of their agentic AI projects reached production.
That is a self-reported average across respondents. It does not establish a universal failure rate or explain why any particular project stalled. The barriers respondents cited are more useful for deciding what to test: integration, security and privacy, missing context, auditability, hallucination, and missing or inaccurate data.

One possible explanation is that teams encounter more security and reliability questions as they get closer to letting agents act. The survey cannot prove that sequence. It does give us a reason to include those questions in the first workflow evaluation, rather than waiting for integration to be finished.
In the refund example, finding an order is only the beginning. The support system might show the original purchase while the payment system shows a partial refund. A policy document might have changed since the customer placed the order. A previous conversation might contain a commitment that affects how the case should be handled.
These are three different kinds of information: business rules, the history of the case, and the procedure for resolving it. The report makes the same distinction. A workflow needs all three, with a way to resolve conflicts between them.
A shared connection to the payment system can reduce duplicated integration work. It cannot make an incomplete payment record accurate. The workflow still needs checks for missing values, stale records and conflicting transactions, along with a route to a person when the conflict cannot be resolved. This is what testing the data means in practice.
The report, produced with graph database company Neo4j, recommends a knowledge layer built on a knowledge graph. That can help connect business facts and their relationships. In our example, it could help identify which refund policy applies. The team still needs to check that the policy and transaction history are current, and then control the action that follows.
A correct recommendation can still lead to an unauthorized payment. The refund workflow needs rules for which accounts it can access, how much it can refund, and when an expert must sign off. Those checks belong before the payment operation, while it is still possible to stop it.
The model may be useful for interpreting an ambiguous customer request. Once the order, amount and authorization are established, code can perform the payment operation directly. Asking a model to reinterpret those settled inputs introduces another opportunity for error.
Use model judgment where the work requires it, and test the decisions it makes. A policy check can stop an excessive refund. It may not catch a plausible explanation based on the wrong order. Both cases belong in the evaluation, because controlling an action and establishing that a decision is correct are separate requirements.
Suppose the payment request times out. The workflow cannot safely assume that the refund failed and send another request. It needs a way to check the transaction's status and prevent a duplicate payment. If it cannot establish what happened, that uncertainty must reach a person before another attempt.
For this example, completion means the correct customer received the authorized amount once, the order history reflects it, and the customer was informed. A successful tool response is one piece of evidence. The resulting business state is what the workflow needs to verify.
The record should preserve the inputs used, the policy version applied, any approval, the payment reference and the verified result. That lets a reviewer reconstruct what happened. Keeping the record does not itself prove that the decision was justified; a reviewer must still be able to compare it with the applicable rules and evidence. We describe that review in how to audit an AI agent.
A pilot that tests only straightforward refunds will tell you little about disputed orders, conflicting records or interrupted payments. Evaluate those cases deliberately. Measure correct completion, duplicate or incorrect actions, user corrections, escalations and the time people spend resolving them. Frequent escalation may keep the workflow safe while leaving it too expensive to be useful.
Begin with expert review where the consequences warrant it. Reduce that review for a defined class of cases when representative evaluations and supervised runs provide enough evidence. A record of success on small, routine refunds does not establish readiness for disputed or high-value ones. Keep a way to stop the workflow, and reassess it when policies or connected systems change.
Established engineering practices help with connections, permissions, records and recovery. Model judgment still introduces uncertainty. The practical work is deciding where that uncertainty is acceptable, how to detect mistakes, and when to hand the case to someone who can resolve it.
We sell AIOS, and this is the design approach behind it. Blueprint captures the workflow for your team to review, including the business rules and the steps that need an expert's sign-off. Shared connectors give workflows a common route to the systems they use. Known operations run directly, with reasoning reserved for steps that need judgment.
Every action is policy-gated: AIOS checks it against your policies before it reaches a system of record, and you add rules where your process needs them. Approval gates can sit on the steps that carry risk. The execution ledger records actions and approvals, and each run is evaluated against its workflow's goal. Exceptions go to a person with the case context.
In the refund example, those mechanisms would support checking the request, obtaining any required approval, executing the payment and collecting evidence of the result. The actual policy, transaction checks and completion criteria must be configured and tested for the business. A list of platform capabilities cannot establish that a particular workflow is ready for production.
Take one workflow from your own backlog and follow an ordinary case all the way through. Then try it with a stale record, a missing permission and a timeout after an action was sent. For each case, ask what the workflow can establish, what it is allowed to do next, and what evidence would show that the work is complete. The unanswered questions are the work your next evaluation needs to address.
Sources
Keep reading
We embed until it works, then you pay for what worked. Bring the process you would most like to stop staffing.