Matching Infrastructure to Agency Needs
One of the key benefits of hybrid infrastructure is that it allows IT teams to place workloads where they best meet agency requirements. Public cloud environments are often well suited for AI experimentation or workloads requiring burst capacity, while on-premises infrastructure may be preferable for applications involving sensitive citizen or public safety data. Likewise, on-premises GPU clusters can support predictable, large-scale inferencing workloads.
READ MORE: Here is a guide to AI governance for state and local agencies.
“Whether an AI workload is running in the cloud or on-premises is going to depend on the organization’s specific needs for artificial intelligence,” says Mariano Carro, principal field solution architect for Microsoft hybrid infrastructure at CDW. “Most of the time, we are going to need some resources in the cloud to do the training and for the high level of compute that we need. But once we get that training complete, we may move a workload on-premises to improve performance or protect data privacy.”
Hybrid infrastructure helps agencies control costs and address GPU scarcity while maintaining data sovereignty and supporting compliance with regulations governing sensitive public-sector information. Keeping data close to processing resources can also reduce latency for AI applications.
Workload Placement: Putting AI Close to Government Data
Workload placement decisions are rarely static. As agencies mature in their use of AI, they often revisit where applications run based on changing cost structures, performance requirements and governance policies.
“Workload placement is going to take into consideration things like latency as well as accessibility,” says Eryn Brodsky, server and storage practice lead at CDW. “Businesses need to think about what AI outcomes they want to leverage the data. You need to have quick access to it, but you also have to ensure that you have proper access to it.”
That balancing act frequently leads to movement between environments over time. Many organizations begin AI initiatives in the cloud because of its flexibility and access to large-scale compute resources. As workloads mature and costs become more predictable, some choose to move those workloads back on-premises.
“It’s not uncommon for us to see customers starting in the cloud, as the majority of our customers are, then bringing those applications back on-premises. We call it repatriation,” Brodsky says. “They repatriate those workloads on-prem because they have to consider the cost of maintaining those applications and those workloads in the cloud versus the cost of being able to invest in the architecture to support them.”
LEARN MORE: IT infrastructure supports AI use cases.
Cost optimization is only one piece of the equation. Performance and efficiency become increasingly important as AI initiatives expand.
“What we’re looking for is lowest cost per function or lowest cost per token,” Gutierrez says. “Leveraging all of the available optimizations — whether it’s through power and cooling, accelerated infrastructure, better data pipelines, better data posture — is really important to achieving the desired outcomes, reducing latency and getting users the information that they need when they need it.”
Those optimizations increasingly depend on matching compute resources to AI workloads. The relationship between CPUs and GPUs has become a critical consideration for agencies modernizing their data centers.
“Advancements in accelerated compute are really important when thinking about the data center of the future,” Gutierrez says. “There are certain applications and certain data center processes that simply work better on a GPU. So, understanding the GPU-CPU relationship in your data center and what workloads or what jobs you can assign to each of those is really important in optimizing your infrastructure.”
