Blog

Notes on verifying and tuning AI infrastructure, and on what we build.

  1. Why our own products run on our inference stack first

    Deep Variance optimizes inference for open models. Every change runs in our own production before it reaches a customer's GPUs. Here is why.

  2. Where the cost goes when you serve an open video model

    Open video models cost far more to serve than text models. A plain-English walk through where the GPU time goes, and which levers move it.