How we cut LLM inference costs by 61% without touching model quality
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
Articles
Newest first. Filter by the kind of article you are looking for.
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
Eighteen months in, the costs and benefits side by side. Not a recommendation either way.
The survey reports that get cited for years are built differently from the ones that get read once and forgotten.
We built a lead scoring model, sold it as a core feature, and shut it off eight months later. Here is what replaced it.
We have audited over 60 companies that believed MFA was fully enforced. Fewer than ten actually were.
Our agent could technically approve these on its own. The data on why we still do not is more interesting than the policy.
We assumed our incident response was fine because nobody complained. The data said otherwise.