How we cut LLM inference costs by 61% without touching model quality
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
6 articles
Topic
Replacing manual processes with models, and what it takes to make that stick.
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
Our agent could technically approve these on its own. The data on why we still do not is more interesting than the policy.
Our first two attempts at automatic ticket routing made things worse. The third one shipped because we stopped optimizing for accuracy.
Six sentences we make every customer read and confirm before we enable autonomous replies. Most of the value is in what we exclude.
You can have 10+ years in your field, a strong CV, the right certifications, and a track record of actually delivering results and still never get a chance to speak to a human.
A walkthrough of building your first automation in Nodesin, and why every step, its inputs and outputs, and the bill remain visible.