AArch64 Playground
4.29 · Vectors: one instruction, many numbers

Vectors: one instruction, many numbers

The floating-point lesson used d0 and s0, the low 64 and low 32 bits of registers named v0 to v31. Each v register is 128 bits wide: room for sixteen bytes, or four ints, at once. The vector instructions read those bits as a row of separate numbers, called lanes, and do the same operation to every lane in one instruction. The name for this is SIMD, single instruction, multiple data. Programs that edit images or mix sound use it heavily, because they repeat one small step on millions of values.

Course programs keep to s and d. The playground runs the vector instructions as well, so this lesson shows what the rest of the register is for.

One register, many views

The scalar names read one value, not a row of them, from the low bits of a v register:

namebitsholds
b08one byte
h016one halfword
s032one single, or one int
d064one double, or one long
q0128the whole register as one value

Written as v0 with an arrangement after a dot, the same 128 bits become lanes:

arrangementlanesbits per lane
v0.16b168
v0.8h816
v0.4s432
v0.2d264

v0.8b, v0.4h and v0.2s use only the low 64 bits. One lane is named by its number in square brackets, counting from 0 at the low end: v0.s[2] is the third 32-bit lane.

Loading, adding, and copying lanes out

  • ld1 {v0.4s}, [x9] loads 16 bytes from the address in x9 into the four lanes, the word at the lowest address into lane 0. st1 {v0.4s}, [x9] stores the lanes back the same way.
  • add v2.4s, v0.4s, v1.4s adds lane 0 to lane 0, lane 1 to lane 1, and so on: four separate adds in one instruction. mul works lane by lane in the same way. All three operands use the same arrangement.
  • dup v3.4s, w9 copies w9 into every lane, and movi v1.8b, 40 puts the constant 40 in every lane.
  • addv s5, v2.4s adds the lanes of one register together and leaves the total in lane 0 of v5, which is s5.
  • printf reads only general registers for %d, so mov w1, v2.s[0] copies one lane out to w1 first.

Each of these has a reference entry with an example to run, such as ld1, dup and addv, and the reference's Calling convention tab covers the jobs of the v registers.

The program below adds two lists of four ints, multiplies the four sums by 3, and totals the sums. It prints:

sums:    11 22 33 44times 3: 33 66 99 132total:   110
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 1try it: run it, or step one instruction at a timeOpen in playground

Vectors and calls

The rule that a function must hand d8 to d15 back unchanged covers only the low 64 bits of v8 to v15. A function you call may change their upper 64 bits, and every bit of v0 to v7 and v16 to v31. So no v register is sure to keep a whole 128-bit value across bl printf.

That is why the program copies the total into w19 and stores the products in memory with st1 before its first printf, then loads them back with ld1 afterward. .align 4 before saved puts it on a 16-byte boundary, the size of one v register.

Each lane stands alone

A carry never crosses from one lane into the next. Add 40 to a byte lane that holds 250, and the true answer, 290, does not fit in 8 bits, so the lane keeps 290 - 256 = 34, while the lane beside it never sees the carry. For the brightness of a pixel that wraparound is wrong: a bright pixel turns nearly black. uqadd (unsigned saturating add) stops at the largest value instead, so 250 + 40 gives 255.

The program below brightens eight pixels both ways and prints each row through a small subroutine. It prints:

in      10  60 120 180 200 215 230 250add     50 100 160 220 240 255  14  34uqadd   50 100 160 220 240 255 255 255
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 2try it: run it, or step one instruction at a timeOpen in playground

The pixel that held 215 lands on 255 exactly, so both rows agree there. Only the last two went past 255, and only those two differ between the rows.

Check yourself

  1. How many lanes does v7.8h have, and how many bits is each one?
  2. v0.4s holds 1, 2, 3 and 4, and v1.4s holds 10 in every lane. What does v2.4s hold after add v2.4s, v0.4s, v1.4s, and what does addv s3, v2.4s then leave in s3?
  3. A byte lane holds 200. What does it hold after an add of 100, and what after a uqadd of 100?
  4. A function keeps a 128-bit value in v9 and calls printf. Which part of v9 can it count on afterward?

answers

show answers
  1. Eight lanes of 16 bits each.
  2. 11, 12, 13 and 14 in the four lanes, and 50 in s3.
  3. 44 after add, because 300 - 256 = 44, and 255 after uqadd.
  4. Only the low 64 bits, d9. To keep all 128 bits, store the register in memory before the call.

Practice

  • Basic quiz: floating point: the register file that the vectors share.
  • Four lanes, one add: add one bonus to four scores with ld1, dup and a vector add, and total them with addv.
  • Basic quiz: vectors: lanes and arrangements, uqadd, dup, copying a lane out, and which v registers a call may change.
  • Predict: vectors (challenge): what lane-by-lane adds, a saturating subtract and addv leave in a register.