Three separate design decisions that all trade complexity for speed, or speed for simplicity, in genuinely different ways. Every comparison below is backed by real, worked-through numbers, not just labelled diagrams.
Every simulation in the Fetch-Execute Cycle lesson used a Von Neumann machine without ever naming it: one shared memory, one bus, for both instructions and data. Run the same 2-instruction program on both architectures and watch what that shared pathway actually costs.
Same task, multiply two numbers already in memory and store the result. Step through both and watch them do the identical underlying work, packaged completely differently.
Pipelining overlaps stages of a single instruction stream. This is different: genuinely separate processing units, each capable of running a completely independent instruction stream of its own, at the same time. Step through time itself and watch cores pick up work in waves.