Quest 3 of 15
Behind the Screen: Data Centers & Power
AI feels invisible but runs on football-field-sized buildings, GPUs, and enormous electricity use. Learn the physical backbone so you understand latency, environmental impact, and why AI is not “free magic in the cloud.” For example, you will trace a prompt from your phone through networks to a remote data center.
Start here
A data center is just a building full of computers — but AI needs special, power-hungry machines. We explain every piece from scratch.
Big idea
Cloud AI runs on real machines in massive data centers. Specialized chips (GPUs/TPUs) do parallel math; they need enormous power and cooling — which affects cost, speed, and where services can be built.

Learn one idea at a time
Read, explore, then mark each idea when you can explain it.
Idea 1 of 8
When you use AI on your phone, your words do not stay on your phone. The message travels over mobile data or Wi-Fi, across the internet's cables and exchanges — often crossing borders and sometimes oceans — to a data center: a large, secure building packed with racks of computers running 24 hours a day. There, powerful machines process your request, generate the reply, and send it back, typically within a second or two. The phone in your hand is just the window; the intelligence lives in a warehouse of humming servers possibly thousands of kilometres away. This is what 'the cloud' physically means — someone else's computers, in a building you will never see.
Live interactive diagrams
Tap nodes, stages, or cards to explore — these diagrams match this module’s ideas.
Path of a chat request
Tap each hop — distance and load add delay.
App sends your prompt over mobile or Wi-Fi.
Choose a deep dive
Open the topics you want to explore. The detail stays folded until you need it.
Deep dive 1Why GPUs suit AI work
Training a model involves many similar calculations on large grids of numbers. A GPU can perform many of those calculations at once, while a CPU is designed to handle a wider variety of jobs. This is why a laptop can use some AI tools but huge model training usually happens in specialised data centres.
Deep dive 2The physical cost of a digital answer
Every AI reply uses servers, networking equipment, electricity, and cooling. Data centres must also stay reliable during power cuts and heavy demand. Efficient prompts and smaller models can reduce waste without asking learners to stop using useful tools.
Deep dive 3The journey of one chatbot question
When you press send, your question travels from your phone over mobile data to a local tower, through national fibre links, often under the sea, to a data centre that may be on another continent. There, load balancers route it to a server rack where GPUs run the model and generate your reply token by token, and the text streams back along the same path. The total round trip often takes a second or two — most of it computation, not travel. Understanding this chain explains real behaviour you see: replies slow down at peak times, fail when fibre is cut, and cost providers real money per message.
Deep dive 4Why some AI runs on your phone instead
Not every model needs a data centre. Small models can run directly on a phone — keyboard prediction, camera scene detection, offline translation. On-device AI answers faster, works without signal, and keeps your data on your phone, but it must be small enough to fit in limited memory and battery. Cloud AI can be enormous and always up to date, but it needs connectivity and raises privacy questions. Engineers choose between them per feature, and many products mix both: your phone transcribes your voice locally, then sends the text to a cloud model for the clever answer.
Deep dive 5Where chat answers live
A data center is a building full of servers. GPUs and cooling make large AI possible — and expensive.

One-minute challenge
Connect this lesson to real life
Name one situation where this idea could help, and one thing a person should still check.
Explore a real-world example
Use the arrows to connect the idea to a visible situation.
Photo example
Example: racks and cables
Latency (waiting time) grows when your phone is far from the servers or the network is congested.

Key terms
Tap a term to flip and read the definition.
Optional further learningFree textbooks and trusted online resources
These sources informed the course structure. Use them to revisit a concept or study it in more depth.
Ready check
Tick each idea only when you could explain it without looking back.
Ready for practice? Draw the journey, interview the “GPU” in a role-play challenge, and rank the path from phone to server.
Extra context (audience, logistics, curriculum notes)
Built for: Connects global tech infrastructure to local realities—power grids, climate, and digital access in Southern Africa.
Formats: Diagram walkthrough · AI hardware Q&A lab · Class debate on regional infrastructure
AI Architect Module 3 — data centers, GPUs vs CPUs, cooling and energy (AI_Architect_10_Part_Course.md).
Next up
Ready for the next part?
When you've finished the reading, inline exercises, and knowledge check for this part, check the box to continue.