AArch64 Playground
4.31 · Every instruction is 32 bits

Every instruction is 32 bits

Every instruction in your program ends up as a number. The assembler's last job is to replace a line such as add x9, x9, x12 with the pattern of bits the processor fetches and obeys. That pattern is the instruction's machine code, and on AArch64 it is always exactly 32 bits (four bytes), whichever instruction it is. The fixed size is part of what makes AArch64 a RISC (reduced instruction set) design: the processor always knows where the next instruction starts, which is why pc, the register holding the address of the running instruction, moves on by 4 after every instruction that does not branch.

This lesson reads those 32 bits. By the end you can take an instruction word written in hex, cut it into its fields, and name the instruction inside, and you can build the word for a simple instruction by hand. The two programs on this page read their own machine code out of memory, so you can check every number here by running them.

Reading a word by position

The bits of a word are numbered from 31 on the left (the most significant bit) down to 0 on the right. A field is a run of neighboring bits with one job, such as naming a register or holding a constant. A field is named by its end positions: "bits 9 to 5" is a five-bit field.

Hex is the easiest way to write a word down, because each hex digit stands for exactly four bits. To split a word, write each of its eight hex digits as four binary digits, then ignore the groups of four and cut the 32 bits again where the fields begin and end. To go the other way, line the fields up, regroup the bits in fours from the left, and read off the hex.

Three facts hold in every format below that names a register:

  • A register is named by a 5-bit number from 0 to 31: x0 is 00000, x9 is 01001, x29 (fp) is 11101, and x30 (lr) is 11110. The number 31 (11111) means the zero register or sp, depending on the instruction.
  • The x and w views of a register share one number. In the arithmetic formats a separate bit picks the width: the size bit sf, bit 31, is 1 for 64-bit x registers and 0 for 32-bit w registers.
  • The register being written (or, for a load or store, the register being transferred) sits at the right end, in bits 4 to 0. The first source register comes next, in bits 9 to 5.

Five formats

The top bits of a word, the opcode, say which operation it holds and therefore how to read the rest of it. Instructions that need the same kinds of operands share one layout, called a format. Textbooks name five of them with short letters:

FormatUsed byEach field and the bits it occupies
R, registeradd, sub, adds, subs with two source registersopcode 31-21, Rm 20-16, imm6 15-10, Rn 9-5, Rd 4-0
I, immediateadd, sub, adds, subs with a constantopcode 31-22, imm12 21-10, Rn 9-5, Rd 4-0
D, load and storeldr, str with pre-index or post-indexopcode 31-21, imm9 20-12, op2 11-10, Rn 9-5, Rt 4-0
B, branchb, blopcode 31-26, imm26 25-0
CB, conditional branchb.eq, b.ne, b.lt and the restopcode 31-24, imm19 23-5, bit 4 always 0, cond 3-0

Each row covers all 32 bits with no gaps. Rd is the destination, Rn and Rm are the first and second sources, and Rt is the register a load writes or a store reads. A field whose name starts with imm holds an immediate, a constant stored inside the instruction itself, and the number in its name is its width in bits: imm12 is 12 bits.

The ARM manual cuts the opcode into smaller named parts and has many more layouts (logical instructions, movz, loads with a plain offset, floating point). These five cover the instructions you write most often, and every other layout is read the same way.

The register format

add x9, x9, x12, from the second program on this page, assembles to 0x8b0c0129. Taken apart:

hex        8    b    0    c    0    1    2    9binary     1000 1011 0000 1100 0000 0001 0010 1001field      opcode        Rm      imm6     Rn      Rdbits       31-21         20-16   15-10    9-5     4-0binary     10001011000   01100   000000   01001   01001meaning    add, x regs   x12     shift 0  x9      x9

The 11-bit opcode packs several choices together. Bit 31 is sf. Bit 30, op, is 0 for add and 1 for subtract. Bit 29, S, is 1 for the versions that set the flags, adds and subs. Bits 23 and 22 choose the kind of shift applied to Rm (00 is lsl), and imm6 holds how far to shift it. With no shift written, imm6 is 0. The four 64-bit opcodes differ only in bits 30 and 29:

Instructionopcodeas hex
add100010110000x458
adds101010110000x558
sub110010110000x658
subs111010110000x758

On w registers bit 31 is 0, which takes 0x400 off each value: add on w registers is 0x058.

Some familiar mnemonics are other instructions under a second name, called aliases. cmp x9, x12 is really subs xzr, x9, x12: a subtraction that sets the flags and sends its result to register 31, where it is thrown away. Its word, 0xeb0c013f, ends in 11111 for that reason.

The program below does this split in code. The line labelled specimen is the instruction it inspects: ldr x9, =specimen gets that line's address, exactly as ldr x0, =fmt gets a string's, and ldr word_r, [x9] reads its four bytes as one number. Then ubfx (unsigned bitfield extract, from the shifts and bitfields lesson) copies each field into the low bits of an argument register for printf.

Running it prints:

word = 0x8b0f09cdsf = 1  opcode = 0x458  Rm = 15  imm6 = 2  Rn = 14  Rd = 13
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 1try it: run it, or step one instruction at a timeOpen in playground

Try a few edits to the specimen line and run again:

  • add w13, w14, w15: sf becomes 0, the opcode reads 0x058, imm6 reads 0, and the word is 0x0b0f01cd.
  • sub x13, x14, x15, lsl 2: only bit 30 changes, so the opcode reads 0x658 and the word is 0xcb0f09cd.
  • adds x13, x14, x15, lsl 2: bit 29 turns on, and the opcode reads 0x558.

The immediate format

When the second operand is a constant, the constant goes inside the instruction. sub w11, w11, 1 assembles to 0x5100056b:

binary     0101 0001 0000 0000 0000 0101 0110 1011field      opcode        imm12          Rn      Rdbits       31-22         21-10          9-5     4-0binary     0101000100    000000000001   01011   01011meaning    sub, w regs   1              w11     w11

The opcode starts with the same three choices as the register format, sf, op and S: here 0 (32-bit), 1 (subtract) and 0 (no flags). The next six bits, 100010, mark the add and subtract immediate family. The last opcode bit is sh.

imm12 is a 12-bit unsigned number, so one add or sub can carry a constant from 0 to 4095. When sh is 1, the constant is shifted left 12 places before use, which reaches multiples of 4096: add x10, x11, 8192 is stored as imm12 = 2 with sh = 1. A constant that fits neither way cannot be an immediate. It has to be put in a register first, the way the registers and immediates lesson builds one with movz and movk.

Encoding one by hand

Going from an instruction to its word is the same work in the other direction:

  1. Pick the format from the instruction's operands.
  2. Write each field in binary at its full width, with leading zeros.
  3. Put the fields side by side, opcode first, to make 32 bits.
  4. Cut those 32 bits into groups of four from the left and read each group as a hex digit.

For add x10, x11, 300:

format     immediate, since the second operand is a constantopcode     sf=1 op=0 S=0, then 100010, then sh=0   ->  1001000100imm12      300 = 256 + 32 + 8 + 4                  ->  000100101100Rn         x11                                     ->  01011Rd         x10                                     ->  0101032 bits    1001000100 000100101100 01011 01010by fours   1001 0001 0000 0100 1011 0001 0110 1010hex        9    1    0    4    b    1    6    a      =  0x9104b16a

To check a hand encoding, make it the specimen line in the program above and read the word line. Ignore the second line for an immediate: it splits the word the register-format way, which does not apply here.

The load and store format

The D format covers ldr and str with pre-index or post-index addressing. ldr x12, [x10], 8 assembles to 0xf840854c:

binary     1111 1000 0100 0000 1000 0101 0100 1100field      opcode         imm9        op2    Rn      Rtbits       31-21          20-12       11-10  9-5     4-0binary     11111000010    000001000   01     01010   01100meaning    ldr, 8 bytes   +8          post   x10     x12
  • The opcode's first two bits give the access size: 11 for 8 bytes (an x register) and 10 for 4 bytes (a w register). Its second-to-last bit is 1 for a load and 0 for a store, so str x12, [x10], 8 starts 11111000000.
  • imm9 is the step in bytes. It is signed, so it runs from -256 to 255, and a negative step is stored in two's complement, the same way a negative number is stored in a register.
  • op2 picks the addressing mode: 01 is post-index (use the address in Rn, then add the step to Rn) and 11 is pre-index (add the step first, then use the new address; written with !).
  • Rn is the base register holding the address, and Rt is the register loaded or stored.

The plain offset form, ldr x0, [x1, 16], uses a different layout with a 12-bit unsigned offset counted in units of the access size. Its bit diagram is on the ldr reference entry.

Branches count in instructions

A branch does not store the address it jumps to. It stores an offset: how far the target is from the branch instruction itself, counted in instructions, not bytes. Every instruction is 4 bytes, so a byte distance always ends in two 0 bits; the format leaves them out and multiplies by 4 when the branch runs, which gives the same field four times the reach. A backward jump has a negative offset, stored in two's complement.

  • The B format, for b and bl, has a 6-bit opcode (000101 for b, 100101 for bl) and a 26-bit offset, imm26. That reaches 2^25 instructions, 128 MB, in either direction.
  • The CB format, for b.cond, has the opcode 01010100, a 19-bit offset imm19 (1 MB either way), a 0 bit, and a 4-bit condition code.

The condition codes you will meet most:

condbitscondbits
eq0000ge1010
ne0001lt1011
hs0010gt1100
lo0011le1101
hi1000ls1001

Two branches from the loop in the program below. b sum_test jumps four instructions ahead to the loop's test, and b.gt sum_top jumps four instructions back to the loop's first instruction:

b sum_test     0x14000004binary         0001 0100 0000 0000 0000 0000 0000 0100opcode         000101                        ->  bimm26          00000000000000000000000100    ->  +4 instructionstarget         16 bytes after the bb.gt sum_top   0x54ffff8cbinary         0101 0100 1111 1111 1111 1111 1000 1100opcode         01010100                      ->  b.condimm19          1111111111111111100           ->  leading 1, so negativeflip bits      0000000000000000011add 1          0000000000000000100           ->  4, so imm19 is -4bit 4          0cond           1100                          ->  gttarget         16 bytes before the b.gt

Because the offset is relative, a branch word does not change when the whole program is loaded at a different address. It changes only when the distance to the target changes.

A program that reads its own machine code

The program below adds four numbers in a loop, then prints the words of five of the loop's instructions, one from each format. The loop has the test at the bottom: b sum_test jumps straight to the test, so a list of zero numbers would skip the body entirely. Each inspected instruction has a label (enc_b, enc_ldr and so on) so the program can find it with ldr =label, the same way it finds a string. The text beside each word is an ordinary string showing the line after m4 has replaced the register names.

Running it prints:

sum = 124b       sum_test           0x14000004ldr     x12, [x10], 8      0xf840854cadd     x9, x9, x12        0x8b0c0129sub     w11, w11, 1        0x5100056bb.gt    sum_top            0x54ffff8c
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 2try it: run it, or step one instruction at a timeOpen in playground

Two edits worth making:

  • Add a line holding nop (an instruction that does nothing) just above enc_sub:. The sum stays 124, but both branches now cross one more instruction: b becomes 0x14000005 (+5) and b.gt becomes 0x54ffff6c (-5). No other word changes.
  • Change the sub to subtract 2. The loop runs twice, so the sum becomes 17, and the sub word becomes 0x5100096b because imm12 now holds 2. The text beside it stays as it was, since it is only a string.

pitfall

The usual slips when encoding by hand: counting a branch offset in bytes instead of instructions; measuring it from the next instruction instead of from the branch itself; writing a register field as one hex digit instead of five binary digits; and swapping Rn and Rd (the destination is always the rightmost field). When a hand encoding and the machine disagree, split both words into fields and compare them one field at a time.

Check yourself

  1. Decode 0xcb0a0128. Which instruction is it?
  2. Encode a b.ne whose target is three instructions before the branch.
  3. What is the largest constant imm12 can hold, and what does setting sh do to it?

answers

show answers
  1. The opcode is 11001011000, a 64-bit sub; Rm is 10, imm6 is 0, Rn is 9 and Rd is 8, so it is sub x8, x9, x10.
  2. The offset is -3, which is 1111111111111111101 in 19 bits, and ne is 0001, so the word is 0x54ffffa1.
  3. 4095; setting sh shifts the constant left 12 bits, which multiplies it by 4096.

Practice