AI Demand and Infrastructure Pressures
As enterprises increasingly adopt AI, the infrastructure needed, spanning compute, storage, networking and energy, has become a central concern. This covers everything from data centres and chip selection to cooling and energy sourcing. Unlike conventional IT stacks, an AI strategy cannot be uniformly deployed; each function within a business may require tailored approaches. AI has become a business-wide decision, not confined to technology departments.
Start with the Problem, Not the Tech
Effective AI initiatives begin by identifying the specific operational challenges or efficiencies sought, allowing clearer ROI calculations. Avoiding AI is typically no longer an option, but indiscriminate adoption does not guarantee transformation.
AI compute and scalability The concept of “AI compute” encompasses the resources required for AI workloads. Requirements depend on data volume and algorithmic complexity. Infrastructure must scale with demand: a few individuals using a chatbot have minimal impact, but entire departments can strain networks and compute resources.Data Centres: On-Premises, Co-location, Modular, Cloud, or Hybrid
- On-premises offers control and security but is capital intensive and energy heavy.
- Co-location lets companies install dedicated hardware in shared facilities, ideal for custom GPU clusters.
- Modular data centres (pre-fabricated units) allow rapid deployment close to operations, cutting latency and easing expansion. The global modular market is growing rapidly, from $30bn in 2024 to a projected $81bn by 2031. (Source)
- Cloud provides flexible, scalable access to compute (GPUs, storage, networking), though dependence on large providers raises portability risks.
- Hybrid approaches (mixing owned equipment and cloud) enable control while benefiting from off-site scale.
Geopolitics and Data Sovereignty
Where data is physically stored matters, governed by data residency and sovereignty laws. Companies increasingly weigh local regulations, geopolitical considerations, and proximity to users to meet compliance and latency needs.
Edge Computing for Speed and Resilience
By processing data near its source, edge computing reduces latency and bandwidth demands. This is essential for time-critical applications (autonomous vehicles, industrial automation). It also provides redundancy if cloud connectivity is lost. Edge typically complements cloud, which remains suited to large-scale model training.
Chips: GPUs, CPUs, TPUs, and Beyond
- GPUs dominate training and deep learning due to parallel processing strengths.
- CPUs still handle lighter inference or edge workloads.
- TPUs and other accelerators are optimized for high-speed parallel tasks, often more efficient for specific workloads.
- Emerging designs such as neuromorphic chips and high-bandwidth memory aim to improve speed and reduce energy use.
Portability and Flexibility
Avoiding lock-in is critical. Businesses need hardware and contracts that keep options open, given rapid innovation and regulatory scrutiny of dominant cloud players. Ensuring data and service portability is key for continuity and competitiveness.
Sustainability and Power Concerns
AI data centres draw up to 10 times more power than standard IT racks. Grid constraints could delay up to 20% of planned capacity by 2030. This has triggered efforts in liquid cooling, waste heat reuse, and even proposals for private small modular nuclear reactors.
Smaller, More Efficient Models
Recent advances (for instance in China with DeepSeek) demonstrate that smaller models can achieve strong performance with dramatically reduced compute needs. This challenges the assumption that only the largest models justify investment and may broaden AI accessibility.