How we cut LLM inference costs by 61% without touching model quality
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
0 articles7 subtopics
Topic
How software, infrastructure and data actually get built and run.
Prompt caching, request batching and a routing layer did more than any model swap. Here is what we measured, in order.
Our agent could technically approve these on its own. The data on why we still do not is more interesting than the policy.
A step-by-step walkthrough of the two-hour exercise we run with clients, including the scenario we use and where teams usually get stuck.
Our first two attempts at automatic ticket routing made things worse. The third one shipped because we stopped optimizing for accuracy.
Discord bots have revolutionized the way communities interact and engage with each other on the popular communication platform, Discord.
You can have 10+ years in your field, a strong CV, the right certifications, and a track record of actually delivering results and still never get a chance to speak to a human.
A walkthrough of building your first automation in Nodesin, and why every step, its inputs and outputs, and the bill remain visible.
Most phishing that gets through is not sophisticated. It is fast. A practical layered setup for a 50-person company.
Ranked by the incidents we have actually responded to, not by how they are usually marketed.
Six sentences we make every customer read and confirm before we enable autonomous replies. Most of the value is in what we exclude.
We put off adding a dedicated search service for two years. A tutorial on getting real search out of Postgres, and the signals that told us to move on.
We have audited over 60 companies that believed MFA was fully enforced. Fewer than ten actually were.
Most of our extraction tasks have no single correct answer to grade against. This is the scoring approach we landed on after two that failed.