Longitudinal Study Reveals One-Year Evolution of LLM Serving Workloads
August 17, 2026
Analysis of a one-year production trace from Chutes characterizes LLM serving workloads across aggregate, temporal, and model-specific perspectives. The study provides visibility into how user-model interactions and long-tail model usage shape production traffic over time.
HOW THIS AFFECTS YOU
●
builderYou can use these longitudinal insights to better design auto-scaling and caching strategies for production.
●
researcherThis provides a more realistic benchmark for testing serving system efficiency against real-world volatility.