Vectors: one instruction, many numbers
The floating-point lesson used d0 and s0, the low 64 and low 32 bits of registers named v0 to v31. Each v register is 128 bits wide: room for sixteen bytes, or four ints, at once. The vector instructions read those bits as a row of separate numbers, called lanes, and do the same operation to every lane in one instruction. The name for this is SIMD, single instruction, multiple data. Programs that edit images or mix sound use it heavily, because they repeat one small step on millions of values.
Course programs keep to s and d. The playground runs the vector instructions as well, so this lesson shows what the rest of the register is for.
One register, many views
The scalar names read one value, not a row of them, from the low bits of a v register:
| name | bits | holds |
|---|---|---|
b0 | 8 | one byte |
h0 | 16 | one halfword |
s0 | 32 | one single, or one int |
d0 | 64 | one double, or one long |
q0 | 128 | the whole register as one value |
Written as v0 with an arrangement after a dot, the same 128 bits become lanes:
| arrangement | lanes | bits per lane |
|---|---|---|
v0.16b | 16 | 8 |
v0.8h | 8 | 16 |
v0.4s | 4 | 32 |
v0.2d | 2 | 64 |
v0.8b, v0.4h and v0.2s use only the low 64 bits. One lane is named by its number in square brackets, counting from 0 at the low end: v0.s[2] is the third 32-bit lane.
Loading, adding, and copying lanes out
ld1 {v0.4s}, [x9]loads 16 bytes from the address inx9into the four lanes, the word at the lowest address into lane 0.st1 {v0.4s}, [x9]stores the lanes back the same way.add v2.4s, v0.4s, v1.4sadds lane 0 to lane 0, lane 1 to lane 1, and so on: four separate adds in one instruction.mulworks lane by lane in the same way. All three operands use the same arrangement.dup v3.4s, w9copiesw9into every lane, andmovi v1.8b, 40puts the constant 40 in every lane.addv s5, v2.4sadds the lanes of one register together and leaves the total in lane 0 ofv5, which iss5.printfreads only general registers for%d, somov w1, v2.s[0]copies one lane out tow1first.
Each of these has a reference entry with an example to run, such as ld1, dup and addv, and the reference's Calling convention tab covers the jobs of the v registers.
The program below adds two lists of four ints, multiplies the four sums by 3, and totals the sums. It prints:
sums: 11 22 33 44times 3: 33 66 99 132total: 110Vectors and calls
The rule that a function must hand d8 to d15 back unchanged covers only the low 64 bits of v8 to v15. A function you call may change their upper 64 bits, and every bit of v0 to v7 and v16 to v31. So no v register is sure to keep a whole 128-bit value across bl printf.
That is why the program copies the total into w19 and stores the products in memory with st1 before its first printf, then loads them back with ld1 afterward. .align 4 before saved puts it on a 16-byte boundary, the size of one v register.
Each lane stands alone
A carry never crosses from one lane into the next. Add 40 to a byte lane that holds 250, and the true answer, 290, does not fit in 8 bits, so the lane keeps 290 - 256 = 34, while the lane beside it never sees the carry. For the brightness of a pixel that wraparound is wrong: a bright pixel turns nearly black. uqadd (unsigned saturating add) stops at the largest value instead, so 250 + 40 gives 255.
The program below brightens eight pixels both ways and prints each row through a small subroutine. It prints:
in 10 60 120 180 200 215 230 250add 50 100 160 220 240 255 14 34uqadd 50 100 160 220 240 255 255 255The pixel that held 215 lands on 255 exactly, so both rows agree there. Only the last two went past 255, and only those two differ between the rows.
Check yourself
- How many lanes does
v7.8hhave, and how many bits is each one? v0.4sholds 1, 2, 3 and 4, andv1.4sholds 10 in every lane. What doesv2.4shold afteradd v2.4s, v0.4s, v1.4s, and what doesaddv s3, v2.4sthen leave ins3?- A byte lane holds 200. What does it hold after an
addof 100, and what after auqaddof 100? - A function keeps a 128-bit value in
v9and callsprintf. Which part ofv9can it count on afterward?
answers
show answers
- Eight lanes of 16 bits each.
- 11, 12, 13 and 14 in the four lanes, and 50 in
s3. - 44 after
add, because 300 - 256 = 44, and 255 afteruqadd. - Only the low 64 bits,
d9. To keep all 128 bits, store the register in memory before the call.
Practice
- Basic quiz: floating point: the register file that the vectors share.
- Four lanes, one add: add one bonus to four scores with
ld1,dupand a vectoradd, and total them withaddv. - Basic quiz: vectors: lanes and arrangements,
uqadd,dup, copying a lane out, and whichvregisters a call may change. - Predict: vectors (challenge): what lane-by-lane adds, a saturating subtract and
addvleave in a register.