Arrow pointing towards left
All articles
AI

Mapping Data Movement Before Buying Hardware Keeps Inference Spend On The Real Constraint

The Data Wire - News Team

|

October 8, 2026

Nagarjun Rajendran, Cloud and Technology Solutions Manager at Bosch Digital, shares why GPU utilization alone cannot tell a team what is limiting its production inference.

Credit: The Data Wire
Quote Icon
Most infrastructure decisions get made assuming the previous bottleneck is still the current one, but that's not the case. The bottleneck keeps moving.

Nagarjun Rajendran

Cloud & Technology Solutions Manager (Principal Architect)
Bosch Digital

A GPU benchmark reports raw compute, and production inference rarely behaves the way that number predicts. A live serving path can be held up inside the chip, by how fast it reads its own memory. The same path can stall between chips, on how quickly two cards exchange data or how long a hop between nodes takes, and fixing whichever one binds typically exposes the next bottleneck. Software efficiency is worth examining before buying more capacity, since the bottleneck a team measured last quarter has usually been replaced by a different one.

Nagarjun Rajendran is Cloud and Technology Solutions Manager and Principal Architect at Bosch Digital, where he owns end-to-end architecture for a cloud platform that a dozen or more engineering teams build on. Rajendran joined Bosch in 2017 as an IoT consultant across the ASEAN region, then spent three years as a senior cloud and IoT architect on its industrial platforms. His earlier work runs from a microelectronics foundation at the National University of Singapore through embedded automotive software to multi-region cloud design. He holds professional or expert-level certification across AWS, Azure, and Google Cloud, and writes about system design and architecture tradeoffs.

"Most infrastructure decisions get made assuming the previous bottleneck is still the current one, but that's not the case. The bottleneck keeps moving," says Rajendran. A team that assumes otherwise spends money on hardware for a bottleneck that's already gone. Rajendran wants a team to work out how data moves through its system before it buys anything, because fixing the movement can remove the need for more hardware.

Compute, bandwidth, and data movement

Adding GPU power is the first move most teams make, and it helps while compute is what's running short. The next limit is how fast the chip can pull data out of its own memory, and buying more processing power doesn't change it. "If your inference workload is slow, you just start adding a lot of compute. Basically you improve the GPU power," says Rajendran. "But then the team starts to discover there is a memory bandwidth bottleneck that surfaces. Memory bandwidth is nothing but how the memory is accessed within a chip."

One request needs both, at different moments and in different amounts. Reading the prompt pushes all of it through the model at once and loads the processing cores. Producing each token after that means fetching what the model stored about every earlier token, and that traffic runs through memory on the same card. "The LLM serving is having two aspects. One is first the prompt gets understood by the LLM, and then it starts generating the first token. That is actually compute bound," explains Rajendran. "But once you start generating tokens from the second token onwards, every time it has to look back to the previous token, how it has been generated. That aspect is actually memory bound."

The bottlenecks tend to arrive in that order. How soon a particular team runs into them is a different question, and Rajendran ties it to how fast the system is growing, so growth rate is the thing to watch. "It purely depends on how their roadmap looks like, how much they want to sell, and what their customer pipeline looks like," he notes. "If the team is small and they are growing at a slower pace, maybe all of this bottleneck surfaces once they reach a particular milestone."

Three numbers worth monitoring

GPU utilization is the number most teams watch, and Rajendran says they lean on it almost entirely. A consumption dashboard hides the cost of output that has to be generated twice, and a utilization figure hides a card sitting idle while another component in the path saturates. "The majority of people always look at compute in the dashboard," Rajendran adds. "They tend to use only GPU power as a monitoring mechanism for their entire inference."

A utilization figure describes one card. Two of the three measurements Rajendran wants describe what happens between cards. "One is the compute power, flops. Then second is the network connectivity, how fast and how efficiently you are connecting two hardware when you serve for a larger audience," he says. "Then third is how efficiently you are moving data from one hardware to another."

Rajendran is most insistent about the order. A team maps how data moves before it buys hardware and before it stands up an inference server, because what the map shows changes what's worth buying. "I would say start with the data movement, how exactly the data is moving. That is where people have to start before they do anything on the inference," he says. "If they solve that, they don't even need to add computation power."

Splitting cache from compute

The KV cache holds the keys and values for every token a model has already processed, the prompt included, and each new token consults it. Running the cache and the generation together puts two jobs on one piece of silicon that can only be tuned for one of them. "Often they oversee that both KV cache as well as inference can be put into the same GPU. It shouldn't be the case," explains Rajendran. "They have to design the cache and the computation in two different hardware, optimize both for those specific purposes, and link them with a high-efficiency networking protocol."

Published work from the vLLM team on paged attention took the same waste to how GPU memory gets handed out. Reserving one contiguous block per request leaves the unused portion of that block stranded inside it. Handing out small pages keeps that memory available for other work. "Instead of addressing that KV cache in the traditional way, they have to come up with an out-of-the-box solution, where they switch it to work the way an operating system like Windows or Linux handles memory," notes Rajendran. "Whenever they need cache they use the entire memory available and keep adding pages instead of reserving a full block. If you split and use it page by page, you free up a lot of GPU capacity."

Kubernetes assumes a workload it can destroy and recreate. Rajendran says that's fine for general compute, while an inference platform is assembled from particular hardware in a particular arrangement. Where a workload runs depends on which hardware sits where, and he expects orchestrators to start reading that before they provision anything. "Maybe in the next two years all of these players would understand the full picture of how to host inference in a more efficient way, and they would have solutions to address all of this off the shelf," he concludes. "But for now, the engineering team has to put in the effort to understand the finer details before they serve inferencing with maximum utilization of the hardware."

Related Stories