Why our own products run on our inference stack first
Deep Variance does inference optimization for open models: video, image, audio and text. We make them faster and cheaper to run on the GPUs you already have. We also hold ourselves to one rule. Nothing we offer a customer runs on their hardware before it has run in our own production.
This post explains that rule and why we think it matters to anyone who runs open models.
Benchmarks are not production
Every inference change looks good somewhere. A new kernel wins on one input shape. A quantized model holds up on a standard evaluation. A scheduler change clears a synthetic load test.
Production is less tidy. Real traffic arrives in bursts. Requests differ in length, resolution and duration. Some users are patient and some cancel. Models get updated. A change that helped on Tuesday’s benchmark can hurt on Friday’s traffic, and the first sign is often a slower product or a bill that did not drop.
We would rather find that out on our own products than on yours.
What “ours first” means
We run our own APIs for our own products. Botsona.io, our AI testing product for websites, is the first of them. When we change how a model is served, from how requests are batched to how the GPU kernels run, that change goes live for our products before anyone else sees it.
That gives us three things a lab result cannot:
- Real traffic. Our products see the same messy load a customer’s product sees.
- Real cost. We pay for our own GPUs, so a change that saves nothing gets noticed.
- Real users. If quality slips, our own product gets worse, and we hear about it.
Only a change that survives all three is one we will bring to a customer.
Why this matters if you run open models
Open models give you control: your weights, your data, your hardware. The cost of that control is that serving them well is now your problem. The stack between a request and a GPU has many layers, and each one can waste time and money.
Most inference companies solve this by asking you to move your traffic onto their cloud. We do the opposite. We optimize the stack itself and bring it to the GPUs you already run, on your own or rented hardware. You keep your models and your data.
Doing that well takes evidence. “Ours first” is how we get it.
What comes next
We will write here about what we measure, what we change and what we get wrong. If you run open models in production and the GPU bill decides your margin, tell us about your workload.