Skip to content
Aqvil

Analysis1 min readPublished Aug 21, 2026

Analysis 1 min read

Why we still keep a human in the loop for refunds over $200

Our agent could technically approve these on its own. The data on why we still do not is more interesting than the policy.

Published · Updated

Our agent handles refund requests under $200 end to end and has for over a year, with an error rate we are comfortable publishing: 0.9% require a manual reversal, almost all of them policy edge cases rather than mistakes. We have the technical capability to raise that threshold. We have not, and this is the actual reasoning, not the marketing version.

The data that almost changed our minds

When we modeled raising the threshold to $500, the projected error rate barely moved, to about 1.1%. On pure accuracy, the case for raising it was reasonable.

What the accuracy number does not capture

We pulled every refund request over $200 from the last year and read the ticket text, not just the outcome. A disproportionate number of them mention something the model's accuracy score cannot see: a customer describing a broader problem with the product, a pattern across their account, or language suggesting they are deciding whether to churn entirely.

A correctly-approved refund does not capture any of that. A human reviewing the same ticket does, and about a third of the time follows up with something the automated flow never would have: a discount on renewal, an apology call, or a product team ticket.

The actual line we draw

We do not automate decisions where the dollar amount is the least important number in the ticket. Under $200, the request is almost always exactly what it says. Above it, the request is often a symptom, and treating it as a symptom requires a person who can see the pattern, not just the case.

Continue exploring

Explore this topic

AI Automation

All AI Automation content

Related experts

Related businesses