A-Level · Data Representation

Data Compression

Every compression technique below is genuinely running on real, visible data, watch a grid actually shrink into runs, a dictionary actually grow entry by entry, and an image actually lose real, measurable detail as compression gets more aggressive.

Section 2

Run-length encoding: when it wins, and when it genuinely doesn't

RLE replaces a run of identical values with just the value and a count. Try a pattern with long runs, then try one without, and watch what actually happens to the size.

Run breakdown (this is the actual encoded output)

-
original pixels
-
encoded values
-
result

Exam tips

  • RLE's effectiveness depends entirely on the data having long runs of identical values. Verified directly above: Blocks and Stripes compress genuinely well, Checkerboard and Random noise don't, and can make the file larger than the original, every run needs at least a value and a count, even a run of length 1.
  • This is precisely why RLE suits simple graphics (icons, line art, large flat colour areas) and is a poor choice for photographs or naturally noisy data.
Section 3

Dictionary-based compression

Rather than counting runs, build a dictionary of substrings seen so far, and replace repeated substrings with a short reference into it. Watch the dictionary actually grow, and the output actually shrink, symbol by symbol.

Dictionary being built

CodeEntry

Encoded output so far

Exam tips

  • The 7-character string "ABABABA" compresses to just 4 output codes, verified directly above, because the dictionary starts recognising "AB" and then "ABA" as whole units after seeing them once.
  • The dictionary is built from the data itself as it's processed, both the compressor and decompressor build the identical dictionary in the identical order, so only the codes need to be transmitted or stored, never the dictionary itself.
Section 4

Lossy compression: real detail, genuinely discarded

Reduce the number of distinct brightness levels an image is allowed to use. Watch the actual picture degrade, and see exactly how many pixels changed value, this is lossy compression in the most literal sense: information that cannot be recovered.

8 bits
per pixel
0%
size reduction
0
pixels genuinely changed

Exam tips

  • This is called quantisation, values are rounded down into a smaller number of allowed "buckets". Fewer bits per pixel genuinely means fewer possible values, and verified directly above, that genuinely means real pixels changing to a value they weren't originally.
  • Unlike RLE or dictionary coding, this loss is permanent, there is no decoding step that recovers the original values, only an approximation, that's precisely what makes this lossy rather than lossless.
Section 5

Check your understanding