How we cut LLM inference costs by 61% without touching model quality
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
Press passNo. 0004Co-founder & CTO
Co-founder at Northline AI. I write about running LLM systems in production and the bills that come with them.
Amsterdam, NetherlandsMember since Apr 2026
aqvil.com/@mara-holt
In Mara’s words
I spent eight years building data pipelines before starting Northline with two colleagues in 2023. We build document and support automation for mid-sized companies. Most of what I publish here is what we measured, including the parts that did not work.
On Aqvil since Apr 26, 2026 · Amsterdam, Netherlands
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
Our first two attempts at automatic ticket routing made things worse. The third one shipped because we stopped optimizing for accuracy.
Six sentences we make every customer read and confirm before we enable autonomous replies. Most of the value is in what we exclude.
Most of our extraction tasks have no single correct answer to grade against. This is the scoring approach we landed on after two that failed.
Our agent could technically approve these on its own. The data on why we still do not is more interesting than the policy.