Autonomous agents function very differently compared to standard inference traffic. A single request might initiate an embedding step, a planning step, and a generation step. Since each of these actions fires unpredictably and relies on distinct hardware, the majority of existing GPU infrastructure was simply not constructed to accommodate them.