Blog
Notes on verifying and tuning AI infrastructure, and on what we build.
Why our own products run on our inference stack first
Deep Variance optimizes inference for open models. Every change runs in our own production before it reaches a customer's GPUs. Here is why.
Where the cost goes when you serve an open video model
Open video models cost far more to serve than text models. A plain-English walk through where the GPU time goes, and which levers move it.