Analysis1 min readPublished Aug 21, 2026
Why we still keep a human in the loop for refunds over $200
Our agent could technically approve these on its own. The data on why we still do not is more interesting than the policy.
Our agent handles refund requests under $200 end to end and has for over a year, with an error rate we are comfortable publishing: 0.9% require a manual reversal, almost all of them policy edge cases rather than mistakes. We have the technical capability to raise that threshold. We have not, and this is the actual reasoning, not the marketing version.
The data that almost changed our minds
When we modeled raising the threshold to $500, the projected error rate barely moved, to about 1.1%. On pure accuracy, the case for raising it was reasonable.
What the accuracy number does not capture
We pulled every refund request over $200 from the last year and read the ticket text, not just the outcome. A disproportionate number of them mention something the model's accuracy score cannot see: a customer describing a broader problem with the product, a pattern across their account, or language suggesting they are deciding whether to churn entirely.
A correctly-approved refund does not capture any of that. A human reviewing the same ticket does, and about a third of the time follows up with something the automated flow never would have: a discount on renewal, an apology call, or a product team ticket.
The actual line we draw
We do not automate decisions where the dollar amount is the least important number in the ticket. Under $200, the request is almost always exactly what it says. Above it, the request is often a symptom, and treating it as a symptom requires a person who can see the pattern, not just the case.
Continue exploring
The support ticket router that took us three tries to get right
Our first two attempts at automatic ticket routing made things worse. The third one shipped because we stopped optimizing for accuracy.
What we tell customers before turning on an AI agent
Six sentences we make every customer read and confirm before we enable autonomous replies. Most of the value is in what we exclude.
How we cut LLM inference costs by 61% without touching model quality
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
The Job Search Has Changed. And We're Optimizing for the Wrong Thing
You can have 10+ years in your field, a strong CV, the right certifications, and a track record of actually delivering results and still never get a chance to speak to a human.
Explore this topic
AI Automation
Related experts
Bogdan Dan
It doesn't matter how many times you fall, the important thing is not to break the bottle!
Getronics
1 article
Related businesses
Nodesin
Software · Budapest, Hungary
Cloud Workflow Automation with AI Agents
1 article
