- US - English
- China - 简体中文
- India - English
- Japan - 日本語
- Malaysia - English
- Singapore - English
- Taiwan – 繁體中文
I use XPU to mean the full class of AI accelerators, including GPUs, TPUs, and custom AI silicon. Whatever the architecture, utilization is one of the primary metrics leaders monitor. That makes sense. Accelerators are expensive, power-hungry, and often the scarcest asset in the data center.
Here is the problem. Utilization tells us whether silicon is active. It does not tell us whether that activity produces useful, customer-valued AI output. An XPU can look full while it waits on data, moves data, rebuilds evicted context, recomputes prior work, absorbs overhead, or generates work that never reaches the user.
The industry needs a better measure I call Productive Utilization: the share of accelerator time that contributes to useful output delivered within the workload’s quality, latency, and service objectives.
A busy XPU is not necessarily productive
Productive Utilization is a management framework, not an established industry standard or a universal equation. It is meant to connect system design to infrastructure economics.
The central question is not simply, “How busy are my accelerators?” It is, “How much useful output am I getting from every accelerator-hour, watt, and infrastructure dollar?”
The distinction matters because utilization can blur three very different categories. Useful output reaches the user and meets the workload’s quality and service objectives. Necessary overhead keeps the system reliable. Avoidable waste is work that does not improve the delivered result, or work that a better design could eliminate without compromising quality or reliability.
Avoidable waste shows up in two ways. Idle waste is obvious: the XPU waits for memory, storage, communication, or scheduling. Active waste is harder. The utilization chart looks healthy, but the accelerator is rebuilding state, repeating prior computation, or producing output that never contributes to the result.
KV cache shows why activity can mislead
The clearest example is the key-value cache, usually called the KV cache. When a model processes a prompt and prior conversation, it builds internal state that helps generate the next answer. That first pass is often called prefill. If the state stays accessible, the system can reuse it. If it is evicted or discarded, the system may have to rebuild it.
Prefill and decode stress the system differently. Prefill builds the state from the prompt and context. Decode generates the answer token by token. Productive Utilization depends on both: avoiding repeated prefill and feeding decode efficiently.
Rebuilding that state keeps the XPU active. It may also do nothing to improve the answer. The user gets the same result they would have received if the state had been reused, but the operator pays again in accelerator time, power, and capacity.
This is active waste. It often looks like a compute problem. In practice, the root cause is often memory capacity, bandwidth, locality, cache behavior, or data movement. The XPU is doing work because the system could not keep the right state in the right place.
In suitable workloads, a well-designed memory and storage hierarchy can retrieve reusable state more efficiently than recomputing it. That does not make fetching universally better. The answer depends on the model, context length, hardware, topology, workload pattern, and service objective.
There is also a difference between logical reuse and physical reuse. A trace may show that a prompt or context should be reusable. But if the underlying cache is evicted, fragmented, or moved out of reach, the system still incurs the recomputation cost. Productive Utilization must measure what the infrastructure delivered, not what the workload theoretically could have reused.
The point is simple. Eviction can turn accelerator activity into a tax. Reuse can turn the same infrastructure into more useful output.
Memory converts capacity into output
Memory and storage are not side issues in AI infrastructure. They are the mechanism that determines how much accelerator capacity becomes useful work.
HBM supplies bandwidth for the hottest data and active execution. High-capacity main memory and memory expansion keep more context and working state accessible. High-performance NVMe storage can preserve reusable state at scale and, in appropriate workloads, reduce unnecessary recomputation.
The hierarchy matters because AI work is not uniform. Some data must be immediately next to the accelerator. Some states must remain available across a longer interaction. Some reusable work has value beyond the next token. Treating all data as either hot or cold misses out on how the workload behaves.
Decode is often constrained by how quickly the system can move weights, KV state, and intermediate data, rather than by peak compute alone. That is why bandwidth and locality matter so much for productive output.
Capacity matters as well. When KV state consumes available memory, the scheduler may reduce concurrency. Demand may still be there, and the accelerator may still be capable, but the system cannot keep enough working state resident to serve more users efficiently.
This is where Micron has a useful perspective. We build HBM, DRAM, memory expansion, and NVMe storage. That gives us a system-level view of where productive utilization is won or lost across the flow of data and state.
The result is straightforward. Low productive utilization may show up as a compute issue, but the root cause is often a lack of memory bandwidth, IOPS, or capacity. An accelerator cannot produce useful output if the system cannot feed it, keep state close, preserve reusable work, and move data efficiently.
Measure output per accelerator-hour, watt, and dollar
Leaders should keep tracking utilization. It is useful. It can reveal idle resources, scheduling problems, and basic orchestration issues. But it should not be treated as a measure of economic return.
The better operating questions are sharper. How much useful output are we producing per accelerator-hour? How much per watt? How much per infrastructure dollar? How often are we rebuilding state that could have been reused? How much work is necessary overhead, and how much is avoidable?
Goodput is an important step in this direction. NVIDIA AIPerf defines Goodput as completed requests per second that meet specified service-level objectives, such as time to first token and inter-token latency. That connects performance to user experience. Productive Utilization asks the economic follow-up question: what share of paid accelerator time created useful output that met those objectives?
That answer should shape design choices: how leaders size memory, evaluate storage tiers, tune workloads, expand capacity, and decide when more accelerators are truly needed.
The next generation of AI infrastructure will not be judged by how busy its XPUs appear, but by how much useful intelligence each accelerator-hour, watt, and dollar produces. A strong memory and storage hierarchy is what converts activity into output.
The practical question is not simply whether to buy more accelerators. It is about whether the system has the right memory and storage performance, capacity, and tiering to turn expensive accelerator cycles into productive utilization and, ultimately, useful output.