3 min read
Training a single large language model can eat up as much electricity as 100 US homes use in a year. That should tell you something about the hardware underneath.
Every ChatGPT prompt or Midjourney render pulls from a stack of specialized chips, networking, storage, and orchestration software. And the constraints on all of it shape what teams actually build.
NVIDIA GPUs still run most AI training, though Google's TPUs, Amazon's Trainium, and AMD's MI300X are gaining ground. Each one balances memory bandwidth, matrix throughput, and power draw differently, so the "best" chip really depends on what you're training.
One NVIDIA H100 goes for about $30,000 and pulls 700 watts under load. Meta reportedly bought around 350,000 of them in 2024. That's roughly $10 billion just in silicon, and it explains why smaller teams rent GPU time from AWS, CoreWeave, or Lambda instead.
The interesting stuff is happening at the edges. Cerebras builds wafer-scale chips that fit a whole model on one piece of silicon, and Groq's LPUs run transformer inference maybe 10x faster than GPUs for certain workloads. Neither will kill NVIDIA, but they'll win specific fights.
TSMC fabricates most of the leading-edge AI silicon. Their CoWoS packaging capacity is still the bottleneck for the whole industry, and nobody expects that to loosen up before late 2026.
A rack full of GPUs isn't worth much without fast interconnects between them. NVLink handles chip-to-chip talk inside a single server. InfiniBand or 400GbE Ethernet stitches servers into training clusters, sometimes pushing 3.2 terabits per second per node.
Moving data around burns more energy than the math itself. Training sets need to sit close to compute, and building those sets means pulling web-scale content from geographically diverse IPs. That's why plenty of engineers buy datacenter proxies at MarsProxies.com when they need the throughput and IP diversity that residential options just can't hit at scale.
Topology matters as much as raw speed here. Fat-tree, dragonfly, rail-optimized layouts, each one suits a different workload, and getting it wrong wastes serious money on idle capacity.Wikipedia's data center entry covers the standard designs if you want the deeper reading.
A 70-billion-parameter model needs terabytes of training data pumped through it fast. HBM3e memory handles the on-chip stuff, but bulk data sits on NVMe SSD arrays or object storage. Streaming it fast enough to keep 700-watt chips from sitting idle is its own engineering problem.
Checkpoints are the other headache nobody talks about. One checkpoint for a frontier-scale model can be over a terabyte, and slow writes burn expensive GPU-hours.IEEE Spectrum has covered how storage throughput has quietly become as much of a limit as compute itself.
Vector databases matter more on the inference side. Pinecone, Weaviate, Milvus, these store the embeddings that power retrieval-augmented generation, and they need access patterns nothing like a traditional database.
An AI datacenter today feels more like a factory than an office server room. Direct-to-chip liquid cooling isn't optional anymore because you can't push air fast enough to keep a 700-watt accelerator from cooking itself.
Grid capacity is the newer worry. Northern Virginia, Dublin, and other classic hubs have basically maxed out on easy power, and Harvard Business Review reported that hyperscalers are locking in nuclear PPAs to guarantee supply through the 2030s.
Site selection has turned into its own strategy problem. Meta, Microsoft, OpenAI, they're picking spots based on grid access and permitting timelines as much as network latency. Some deals reserve power years before ground even breaks.
Water use gets less press but really shouldn't. A single hyperscale campus can drink millions of gallons a day for evaporative cooling, which is pushing builders toward closed-loop systems and drier locations.
Custom silicon will keep chipping away at NVIDIA's lead. And inference is already moving toward the edge as models get smaller and cheaper to run. The teams that win won't be the ones buying the most GPUs; they'll be the ones who figured out how to keep those chips fed.
If you're building anything serious in AI, you have to think about the whole stack: silicon, network, storage, power, data. Any weak link caps everything else, no matter how much you spend on the shiny parts. Getting the plumbing right isn't fun work, but it's the difference between a shipped model and a project stuck at 40% GPU utilization.
Start with one task and clear approval rules. We handle hosting, saved memory, restarts, and messaging connections.
Plans start at $29/month. Cancel anytime.
Hosted agent
OpenClaw or Hermes