Computing · · 6 min
GPU computing, explained
Why graphics processors became central to AI, and what determines how effectively they are used.
GPUs were designed to perform the same operation across many pixels at once. That same parallelism suits the matrix arithmetic at the heart of neural networks.
Throughput over latency
A CPU core is optimised to finish one task quickly. A GPU is optimised to finish many similar tasks in parallel. Workloads that can be expressed that way benefit enormously.
Memory is often the constraint
Model size, batch size, and activation memory all compete for on-device memory. Many sizing decisions come down to fitting a workload into available memory efficiently.
Utilisation
Raw capability means little if accelerators are idle. Data pipelines, scheduling, and job design determine how much useful work is actually done.