A-Level · Computer Systems

CPU Architecture Concepts

Three separate design decisions that all trade complexity for speed, or speed for simplicity, in genuinely different ways. Every comparison below is backed by real, worked-through numbers, not just labelled diagrams.

Section 2

Von Neumann vs Harvard architecture

Every simulation in the Fetch-Execute Cycle lesson used a Von Neumann machine without ever naming it: one shared memory, one bus, for both instructions and data. Run the same 2-instruction program on both architectures and watch what that shared pathway actually costs.

Architecture

Controls

Cycles used so far

0

Program: LDA 4 (loads memory[4]), then LDA 6 (loads memory[6])

Exam tips

  • This shared-bus limitation is called the von Neumann bottleneck: instruction fetch and data access permanently compete for the exact same pathway, no matter how fast the CPU itself becomes.
  • Harvard architecture is genuinely faster for this reason, but needs two separate physical memory systems, more hardware, more cost, which is why it's common in small embedded microcontrollers (predictable, small programs) rather than general-purpose PCs (huge, unpredictable memory needs).
Section 3

CISC vs RISC

Same task, multiply two numbers already in memory and store the result. Step through both and watch them do the identical underlying work, packaged completely differently.

CISC: one complex instruction

MULT 2:3, 5:2

RISC: several simple instructions

LOAD A, 2:3 LOAD B, 5:2 PROD A, B STORE 2:3, A

Controls

Instructions the programmer had to write

–

Exam tips

  • Don't just write "RISC is faster", that's an oversimplification. RISC's real advantage is that its simple, uniform instructions pipeline extremely well, CISC's variable-length, variable-duration instructions make pipelining considerably harder to implement cleanly.
  • CISC's advantage is code density: fewer, richer instructions mean smaller compiled programs, which mattered enormously when memory was scarce and expensive, historically why CISC came first.
Section 4

Multicore and GPU processing

Pipelining overlaps stages of a single instruction stream. This is different: genuinely separate processing units, each capable of running a completely independent instruction stream of its own, at the same time. Step through time itself and watch cores pick up work in waves.

8 independent items, step through time

Number of cores

Controls

Time elapsed

0

GPU-style: what happens when items need different operations

Exam tips

  • A GPU has a very large number of simple cores, all genuinely fast, but only when they can all run the exact same instruction on different pieces of data at once (this pattern is called SIMD, Single Instruction Multiple Data), exactly why GPUs excel at image processing and matrix maths, the same operation really does apply to every pixel or every cell.
  • When items genuinely need different logic (branching), a GPU can't run both branches on the same core simultaneously, it has to run all the cores through branch A first, then all of them through branch B, idling half the cores each pass, called branch divergence, a real, measurable performance cost.
  • Multicore CPUs don't have this restriction to nearly the same extent, each core can run a fully independent, differently-branching program, which is why general-purpose computing still relies on CPU cores rather than GPU cores for most everyday tasks.
Section 5

Check your understanding