Where the cost goes when you serve an open video model

Open video models have become good enough to ship in real products. Serving them is another matter. A request to a video model can keep a GPU busy far longer than a request to a text model, and a team that budgeted for a chatbot is often surprised by the bill.

This post walks through where that GPU time goes. It is written for teams that run open video models themselves and want to know where to look first.

One clip is many full passes

Most open video models today are diffusion models. They do not write a clip in one go. They start from noise and refine it over many steps, and every step is a full pass through a large model over the whole clip.

So the cost of one clip is roughly the cost of one pass, times the number of steps. Anything that makes a pass cheaper, or removes steps, pays off many times over.

Video is a lot of tokens

Inside the model, a clip is cut into small patches across space and time. Each patch becomes a token. A few seconds of video at a useful resolution turns into a very long sequence.

Attention, the part of the model that relates every token to the others, grows faster than the sequence does. Double the frames or the resolution and the attention work grows by more than double. That is why longer and sharper clips get expensive quickly, and why attention kernels matter so much for video.

Guidance can double the work

Many models use a technique called classifier-free guidance to follow the prompt more closely. In its common form it runs the model twice per step: once with the prompt and once without. That can double the cost of every step for a quality gain that some products need and others do not.

Turning it back into video

At the end, a separate decoder turns the model’s compressed output back into frames. For long or high-resolution clips this stage is not free, and it uses a lot of GPU memory at the moment the main model may still be loaded.

The GPU is often waiting

All of the above is about work the GPU does. A lot of cost comes from time the GPU spends doing nothing useful:

  • Requests do not batch neatly. Clips differ in length, resolution and step count, so it is hard to group them, and a GPU running one request at a time is often underused.
  • Queues are uneven. Traffic comes in bursts. Provision for the peak and GPUs sit idle in the quiet hours; provision for the average and users wait.
  • Memory decides concurrency. Long video sequences use a lot of memory, which caps how many requests one GPU can hold at once.

The levers

None of these levers is a secret. The hard part is knowing which ones help a given workload, and checking that quality holds.

  • Fewer steps. Distilled and faster-sampling versions of a model can reach similar quality in fewer steps.
  • Reuse across steps. Neighbouring steps often compute similar values, and some of that work can be cached.
  • Faster attention. Kernels written for long sequences and for the exact GPU in use make every step cheaper.
  • Lower precision. Running parts of the model in lower precision can save time and memory when quality is checked carefully.
  • Smarter batching and scheduling. Grouping requests by shape, and deciding what runs when, keeps GPUs busy without making users wait.
  • Matching the product. Not every product needs the longest clip, the highest resolution or the strongest guidance. Serving what the product needs, and no more, is often the cheapest win.

Each lever trades something. Fewer steps can soften detail. Lower precision can drift. Caching can blur motion. That is why we measure every change on real traffic, in our own production first, before we recommend it to anyone.

Where to start

If you serve an open video model and want to cut what it costs, measure before you change anything: time per step, steps per clip, how full your GPUs really are, and how long requests wait. The answer is usually in one of the sections above.

If you would like a second pair of eyes on your numbers, tell us about your workload.