Sunday, 9 Aug 2026
|
Most AI agent metrics are designed to be reported, not to be acted on. Deflection rate, containment, assist rate, interactions handled — each measures something real, and each can be arranged to look impressive without the underlying operation changing at all.
There is one metric that resists that: auto-resolve rate, defined strictly as the percentage of inbound requests fully completed without any human touch. It is unforgiving, it maps directly to labour cost, and it is the number a CFO will recognize as meaning something.
Getting the definition right before deployment matters more than it sounds, because the definition determines what the team optimizes for. Vague definitions produce agents that look successful while the work quietly moves somewhere else.
The distinction that matters most is between no human touched it and the request was actually completed correctly. An agent that sends a confident wrong answer scores well on every metric except the last one — until the customer emails back, which starts a new interaction that most systems count separately.
That re-contact loop is where inflated numbers get exposed. Any honest definition has to account for it.
Auto-resolve rate should be defined as: the percentage of inbound requests where the agent completed the requested outcome, no human edited or sent any part of the response, and the requester did not re-contact about the same issue within a defined window.
Three clauses, each closing a specific loophole:
"Completed the requested outcome" excludes acknowledgements. An agent replying "we're looking into that" has not resolved anything, though it will happily count itself as having responded.
"No human edited or sent any part" excludes draft-mode work. Draft mode is a genuinely useful deployment stage — it is how you accumulate accuracy data safely — but a human approving every message is not automation, and blending the two makes the number meaningless.
"No re-contact within a window" — typically 48 or 72 hours — closes the wrong-answer loophole. This is the clause most vendors omit, and the one that makes the metric trustworthy.
A single blended auto-resolve rate across an entire inbox is not actionable, because the workflows underneath it have wildly different ceilings.
Status and tracking requests are highly automatable: the information exists in a system, the question is unambiguous, the answer is verifiable. Quote requests with complete data are strong. Quote requests missing dimensions are gated on an intake exchange first. Rate negotiation, claims adjudication and anything involving a commercial concession have a low ceiling by design — and should.
A blended 40% could be an excellent result or a poor one depending entirely on the mix. Reported per workflow, the same data tells you where to invest next and where you have already hit the sensible limit.
This is also the honest answer to "what's a good auto-resolve rate?" — the question is underspecified until you name the workflow.
Auto-resolve rate alone can be improved in ways that damage the business. Three companions prevent that:
Together with auto-resolve rate, those four give a complete picture. Any one of them alone can be moved without improving anything.
The most common reason AI deployments fail to demonstrate value is that nobody recorded what the process cost beforehand. Once the agent is live, the pre-deployment state is unrecoverable — and "it feels faster" does not survive a budget review.
Two weeks of baseline measurement before go-live, capturing volume by workflow, median response time, human minutes per item and re-contact rate, is the cheapest insurance available. It is also the step teams skip most often, and a large part of why so many pilots never reach production.
Once live, the same numbers need to stay visible continuously rather than being pulled for quarterly reviews — which is a question of instrumenting the agent properly rather than of reporting discipline.
What is a good auto-resolve rate for a logistics AI agent? The question needs a workflow attached. Status and tracking requests support a high rate because the information is retrievable and the question is unambiguous. Rate negotiation supports a low one by design. A blended number across mixed workflows is not comparable between companies.
Should draft-mode approvals count toward auto-resolve rate? No. Track them separately as approval rate — it is the leading indicator that tells you when a workflow is ready for autonomy — but never blend it into auto-resolve, or the metric stops meaning anything.
How long should the re-contact window be? 48 to 72 hours covers most logistics workflows. Longer windows are more conservative; the important thing is choosing one and never changing it, because the metric's value is in its comparability over time.
What if our auto-resolve rate goes down after we expand to a new workflow? That is expected and healthy. It is why segmentation matters — a blended number falling because you took on harder work looks like regression and is actually expansion.
Deflection, containment and assist rate can all be arranged to look good while the work moves somewhere else. Auto-resolve rate, defined strictly, cannot.
Define it as completed outcome, zero human touch, no re-contact in 72 hours. Report it per workflow, never blended. Pair it with accuracy, escalation quality and human minutes per transaction. And measure the baseline for two weeks before go-live, because you cannot reconstruct it afterward.
Debales deploys AI agents for freight quoting, order processing, ETA updates, and multi-channel customer communication — reported per workflow with auto-resolve, accuracy and escalation quality visible from day one. Book a demo.

Wednesday, 2 Sep 2026
Gartner projects agentic supply chain software spend reaching $53 billion by 2030 and 40% of enterprise applications embedding agents by the end of 2026. Here's what that means concretely for a broker next year.

Tuesday, 1 Sep 2026
USPS cut its DIM divisor in July, peak surcharges are up as much as 23%, and NMFC reclassification changed LTL pricing. The crossover point between parcel and LTL shifted on both sides at once.