A-Level · Software

Translators and Compilation

A CPU only ever executes binary machine code. Every program you've ever written in a higher-level language had to pass through one of three kinds of translator first. This page builds every stage of that pipeline for real, live, on data you type in yourself.

TranslatorWhat it doesOutput
CompilerTranslates the entire program before any of it runsStandalone machine code executable
InterpreterTranslates and executes one line at a timeNo executable, source needed every run
AssemblerTranslates assembly mnemonics into machine codeOne mnemonic → one machine instruction
Section 2

Stage 1, lexical analysis: turning text into tokens

The very first thing any translator does is stop treating your code as plain text. Type a line below and watch it split into tokens live, exactly as a real lexer would.

keyword identifier number string operator punctuation

Try one

Section 3

Stage 2, syntax analysis: from tokens to a real tree

The lexer above only produces a flat list of tokens, it knows nothing about structure. The parser's job is to build a genuine tree from that list, one that correctly captures precedence, this is the actual Abstract Syntax Tree a compiler would build, not just a description of one.

Try one

Exam tips

  • Notice 3 * 4 sits deeper in the tree than the +, that's precedence made structural: the parser evaluates the deepest branches first, so multiplication naturally happens before addition without needing any special-case rule at evaluation time.
  • If the tokens are in the wrong order, a closing bracket missing for example, the parser cannot build a valid tree at all, that's precisely what a syntax error is.
Section 4

Stage 3, semantic analysis: the symbol table

Syntactically valid code can still be meaningless. "Check the variable was declared before use" isn't just a rule to memorise, here's the actual lookup that enforces it, live.

Controls

Symbol table

empty
Section 5

Compiler vs interpreter: the same bug, two different outcomes

Same 4 lines, same bug on line 3 (q was never defined). Run both and watch the single most commonly examined difference between them play out for real.

Compiler

Interpreter

Exam tips

  • The compiler found this error without running a single line, semantic analysis catches it during translation. The interpreter had no idea anything was wrong until it actually tried to execute line 3.
  • This is exactly why compiled languages report all errors at once (sometimes overwhelming), while interpreted languages report exactly one error at a time, wherever execution happened to stop.
Section 6

Stage 4, optimisation: constant folding

"The compiler optimises the code" is easy to state and easy to leave vague. Here's a genuine optimisation, verified, with a real instruction count before and after.

Without optimisation

With constant folding

Exam tips

  • The compiler can see that 2 and 3 are both fixed, known values at compile time, so it computes 2+3 itself, once, and simply writes 5 into the code. That addition never happens at runtime, not even once.
  • This only works because 2 and 3 are literal constants. If one of them were a variable whose value isn't known until the program runs, the compiler couldn't fold it, the addition would have to stay in the generated code.
Section 7

Why an interpreted loop costs more than it looks like

A naive interpreter doesn't just execute a line, it re-translates that exact same line every single time it's reached. A compiler translates it once, no matter how many times it later runs.

for i in range(N): x = x + 1

Loop iterations (N)

Section 8

Assemblers: why forward references need two passes

This program jumps to a label, END, that hasn't been defined yet when the jump instruction is written. A single top-to-bottom pass genuinely cannot resolve that, here's why two passes fixes it, using the exact LDA/ADD/STA encoding from the Fetch-Execute Cycle lesson.

Controls

Symbol table (labels → addresses)

empty

Exam tips

  • Pass 1 reads every line just to record where each label actually ends up, no machine code is generated yet. Pass 2 reads the source again, now able to look up any label, including ones defined further down the program than where they're used.
  • A backward reference (jumping to an earlier label) doesn't strictly need two passes, its address is already known. Forward references are the reason two passes are necessary in general.
Section 9

Check your understanding