AArch64 Playground
4.2 · System architecture: what is inside the box

System architecture: what is inside the box

A program is a list of instructions stored in memory as numbers. The processor reads them one at a time and does what each one says, billions of times a second. Assembly is written close to that level, so it pays to know the parts involved. This lesson names them, follows one instruction through the processor, and then steps a real program to watch it happen.

The parts of a computer

Every general-purpose computer is built from the same few parts:

  • The CPU (central processing unit), or processor, runs the instructions.
  • Main memory, also called RAM, holds a program while it runs and the data it works on. It is fast but volatile: it loses everything when the power goes off.
  • Secondary storage, such as an SSD (solid-state drive) or a hard disk, keeps files while the power is off. It is much slower than RAM, so a program is copied from storage into memory before it starts.
  • Peripherals are the devices around the edge: the keyboard, the screen, the network card.
  • The bus is the set of wires that connects these parts and carries values between them.
+-------------------------------+|              CPU              ||   control unit     ALU        ||   registers: pc, x0 to x30    |+---------------+---------------+                |  system bus: address, data, control                |     +----------+-----------+     |          |           |  memory     storage    peripherals  (RAM)     (SSD, disk) (keyboard,                         screen)

Inside the CPU

The CPU has three main parts:

  • The control unit reads each instruction, works out what it asks for, and tells the other parts what to do.
  • The ALU (arithmetic logic unit) does the arithmetic and the logic: add, subtract, compare, and the bit operations.
  • The registers are a small set of storage slots inside the CPU itself. They are the fastest storage in the machine, and the ALU works on them directly. An AArch64 processor has 31 general registers, x0 to x30, each 64 bits wide.

A few registers have fixed jobs. The program counter, pc, holds the address of the instruction being run. Inside the control unit, the instruction register holds a copy of that instruction while it is worked on. A status register holds flags: single bits that describe the last result, such as whether it was zero or negative. The branching lessons use the flags to make decisions.

Memory is a long row of numbered bytes

Main memory is a long row of bytes, and each byte has its own number, called its address. A program names a byte by its address. A value wider than one byte, such as a 4-byte int, sits in bytes next to each other and is named by the address of its first byte. AArch64 calls a 4-byte value a word.

On AArch64 an address is 64 bits wide, so it fits in one x register. The program and its data share this one memory: the instructions are numbers in memory too, sitting beside the data. A machine built this way is called a von Neumann machine. A Harvard machine, by contrast, keeps instructions and data in separate memories.

The system bus

The CPU and memory talk over the system bus, which has three parts:

  • the address bus carries the address the CPU wants,
  • the data bus carries the value being read or written,
  • the control bus carries signals such as "read" or "write", and says when a value is ready.

To read a word, the CPU puts the address on the address bus and a read signal on the control bus, and memory answers with the value on the data bus.

The clock

The clock keeps every part in step. It is a signal that flips between low and high at a fixed rate, a square wave, and each tick is one cycle. A 3 GHz processor's clock ticks 3 billion times a second. A simple instruction takes one cycle or a few; a trip out to main memory takes many more.

Fetch, decode, execute

The processor runs every program with the same three-step loop, over and over:

  1. Fetch. Read the instruction at the address in pc from memory into the instruction register.
  2. Decode. The control unit splits the instruction's bits into fields (which operation, which registers, which constant) and sets up the parts that will do the work.
  3. Execute. The ALU, or the part that talks to memory, does the work, and the result goes to its destination register, or to memory for a store. Then pc moves on to the next instruction.

Every AArch64 instruction is exactly 4 bytes long, so the next instruction is normally at pc + 4. The exception is a branch, an instruction whose whole job is to put a different address into pc; that is how loops and if statements work. Here is one instruction from the program below, add w12, w10, w11, going round the loop:

fetch    read the 4 bytes at the address in pc into the instruction registerdecode   operation: add    destination: w12    sources: w10 and w11execute  the ALU adds w10 and w11, and the sum goes into w12next     pc = pc + 4

Watch it happen

The program below keeps two numbers in memory, loads them into registers, adds them, and stores the sum into a third memory word. It prints that word before and after, so you can see the one change to memory:

sum in memory before: 0
sum in memory after:  42

Run it, then press step to run one instruction at a time and watch the registers:

  • pc grows by 4 on every step, because every instruction is 4 bytes. The two bl printf lines are the exception: bl is a branch, so pc jumps away into the library and comes back to the line after the call a few steps later.
  • The two ldr steps that load values fill x10 and x11, and the add fills x12. (w10 is the low 32 bits of x10; the next lessons explain the two names.) The add does not touch memory; only the str writes the 42 back.
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 1try it: run it, or step one instruction at a timeOpen in playground

note

ldr addr_r, =first puts the address of first into a register, and ldr first_r, [addr_r] then loads the value stored at that address. The square brackets mean "the memory at this address". This program keeps its values in w10 to w12 because it is done with them before each call to printf, which is free to change those registers.

Load/store machines and accumulator machines

The add above never touches memory. AArch64 is a load/store design: only load and store instructions reach memory, and every other instruction works on registers. To add two numbers held in memory, a program loads both, adds the registers, and stores the result.

Many older and smaller processors had a single working register, the accumulator. On an accumulator machine, an instruction like ADD b means "add the value in memory at b to the accumulator", and every result lands in the accumulator. With only one register to keep results in, a longer calculation has to park each partial result in memory and fetch it back later. Here are both designs working out (a + b) - (c + d):

load/store machine         accumulator machine(ARMv8)load  r1 <- a              LOAD  aload  r2 <- b              ADD   bload  r3 <- c              STORE temp1load  r4 <- d              LOAD  cadd   r1 <- r1 + r2        ADD   dadd   r3 <- r3 + r4        STORE temp2sub   r1 <- r1 - r3        LOAD  temp1store r1 -> result         SUB   temp2                           STORE resultmemory trips: 5            memory trips: 9

The program below plays an accumulator machine on AArch64. acc_r is the only register allowed to keep a result. AArch64 cannot add straight from memory, so each accumulator ADD becomes a load into in_r, which carries the value for one instruction only, followed by the add. Every partial result is stored before the next one starts. It prints (a + b) - (c + d) = 28.

loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 2try it: run it, or step one instruction at a timeOpen in playground

RISC and CISC

Processor designs fall into two broad families:

  • RISC (reduced instruction set computer) designs, such as ARMv8, have fewer and simpler instructions, all the same size, and are load/store machines. Simple instructions of one size are quick to decode.
  • CISC (complex instruction set computer) designs, such as the x86 chips in most Windows PCs, have many more instructions. Some do several steps at once, including arithmetic straight on memory, and some take many cycles. Their instructions are anywhere from 1 to 15 bytes long, so the processor has to work out where one ends before it can decode the next.

A RISC program usually needs more instructions to do the same job, but each one is simpler.

The operating system and exception levels

Your program never talks to the keyboard, the screen or the disk itself. The operating system (Linux, on the machines this site imitates) sits between programs and the hardware. ARMv8 enforces this with exception levels, privilege levels numbered EL0 to EL3:

  • EL0: ordinary programs, including yours, with the least privilege.
  • EL1: the kernel, the core of the operating system.
  • EL2: a hypervisor, which runs several operating systems side by side.
  • EL3: the secure monitor, the most trusted code on the chip.

When printf has text to show, it asks the kernel to write it, and the kernel, running at EL1, does the part that touches the hardware. A later lesson makes that request directly.

Check yourself

  1. What does the ALU do, and what does the control unit do?
  2. Which keeps its contents when the power goes off: RAM or an SSD?
  3. pc holds 0x400100 and the instruction there is an add. What does pc hold after the add runs?
  4. On a load/store machine, which instructions may read or write memory?
  5. What does the data bus carry?

answers

show answers
  1. The ALU does arithmetic and logic; the control unit decodes each instruction and directs the other parts.
  2. The SSD; RAM is volatile.
  3. 0x400104.
  4. Only the loads and the stores.
  5. The value being read from memory or written to it.

Practice