Floating point: the other register file
prerequisite
An integer register cannot hold 12.5 or 0.045. For numbers with a fraction the processor has floating-point numbers: a value stored as a sign, a power of two, and a fixed number of significant digits, the way scientific notation writes 6.02 * 10^23, but in binary. How those bits are laid out is the subject of Inside a float: IEEE 754; this lesson is about using them.
Floating-point values live in their own set of 32 registers, their own register file, apart from x0 to x30, and have their own instructions, nearly all of them starting with f.
The s and d registers
Each of the 32 floating-point registers can be read at more than one width. d0 to d31 are 64 bits wide and hold a double (double precision, C's double, good for about 15 significant decimal digits). s0 to s31 are the low 32 bits of the same registers and hold a single (single precision, C's float, about 7 digits). d3 and s3 are one register seen two ways, just as x3 and w3 are, and writing s3 clears the rest of it. The full 128-bit register is called v3; the instructions that work on several values at once use that name, and this lesson needs only s and d.
An instruction works at one width. fadd d0, d1, d2 adds doubles and fadd s0, s1, s2 adds singles. fadd d0, s1, d2 does not assemble: convert one operand first, as shown further down.
Constants
A floating-point constant goes in .data with .double (8 bytes) or .float (4 bytes). The number is written with a 0r prefix, which marks it as a real number rather than an integer: rate: .double 0r0.045. Load it the way you load any variable, its address first and then the value, but into a d or s register: ldr x9, =rate and then ldr d9, [x9].
fmov can put a constant straight into a register, but only from a small set: a whole number of sixteenths times a power of two, from 0.125 up to 31, such as 0.25, 0.5, 1.0, 2.5, and 10.0. fmov d1, 0.1 does not assemble, because no such product equals 0.1 exactly; put 0.1 in .data instead. fmov also copies one register to another (fmov d0, d8) and copies raw bits from a general register: fmov d8, xzr makes 0.0, because a double whose 64 bits are all zero is 0.0.
Arithmetic
fadd, fsub, fmul, and fdiv take a destination and two sources, like their integer cousins. A few more have no integer twin:
fmadd d0, d1, d2, d3computesd3 + d1 * d2, andfmsub d0, d1, d2, d3computesd3 - d1 * d2. Each rounds once, at the end, instead of after the multiply and again after the add.fnmul d0, d1, d2is-(d1 * d2).fabsandfnegtake one source and give its absolute value and its negation.
Dividing by zero does not stop the program. 1.0 / 0.0 gives infinity, and 0.0 / 0.0 gives NaN, a special value meaning "not a number"; printf shows them as inf and nan.
Printing with %f
printf reads each %f from the floating-point registers: the first from d0, the next from d1, and so on. The format string still goes in x0 and each %d still comes from w1, w2, and on, because the two kinds of argument are counted separately. The format "year %2d: interest %6.2f, balance %8.2f\n" takes the year from w1, the interest from d0, and the balance from d1. %.2f prints two digits after the point, and %8.2f also pads the number to 8 characters so the columns line up.
The program below grows a balance of 1000.00 by 4.5 percent a year. Each pass multiplies the balance by the rate to get the year's interest, adds the interest on, and prints both. The balance and the rate have to survive every call to printf, so they live in d8 and d9, two registers that a call must leave as it found them; the rules are in the last section.
It prints ten lines, from year 1: interest 45.00, balance 1045.00 to year 10: interest 66.87, balance 1552.97.
note
Look at year 2: 1045.00 plus 47.02 is 1092.02, yet the line says 1092.03. The machine kept every digit it could. The interest was a hair under 47.025 and the new balance a hair over 1092.025, and printf rounded each one on its own as it printed it, the first down and the second up. Rounding happens only in what you print, never in what the register holds.
Printing a single
%f always means a double. printf is variadic: it takes any number of arguments after the format string. C widens every float it passes to a variadic function into a double, so printf never receives a single at all. Widen one yourself with fcvt d0, s0 before the call. Skip that step and printf reads the 32 bits of s0 as if they were a 64-bit double, a number so tiny it prints as 0.00.
Converting between integers and floating point
scvtf d1, w9turns a signed integer into a double, so 100 becomes 100.0 (ucvtfdoes the same for an unsigned integer).fcvtzs w2, d0turns a double into a signed integer by dropping the fraction, as C's(int)xdoes: 23.7 becomes 23 and -23.7 becomes -23.fcvtns w1, d0rounds to the nearest integer instead: 23.7 becomes 24. A value exactly halfway goes to the even neighbor, so 2.5 becomes 2 and 3.5 becomes 4 (fcvtnurounds the same way to an unsigned result).fcvt d0, s0widens a single to a double, andfcvt s0, d0narrows a double to a single, rounding off the digits that do not fit.
Comparing
fcmp d1, d2 compares two values of the same width and sets the flags, and the signed conditions read them just as they do after cmp: b.gt, b.ge, b.lt, b.le, b.eq, and b.ne. Most results carry a little rounding, so compare against a limit with b.gt or b.le rather than testing for an exact match with b.eq.
The program below shows each conversion. It divides 100 by 8 twice, once with sdiv and once with fdiv after scvtf, then uses fcmp to pick the warmer of two single-precision sensor readings, rounds it both ways, and widens it with fcvt so printf can print it.
It prints three lines: sdiv answers 12, fdiv answers 12.50, and the warmer reading, 23.70, becomes 24 with fcvtns and 23 with fcvtzs.
Floating point and subroutines
The calling convention (the rules every function follows) has a floating-point half that mirrors the integer half:
- Arguments go in
d0tod7(s0tos7for singles), counted separately from the integer arguments inx0tox7. A functionscale(int count, double factor)receivescountinw0andfactorind0. - A floating-point result comes back in
d0(ors0). d8tod15are callee-saved: a function that changes one must put the old value back before it returns, just likex19tox28. Only their low 64 bits count, so saving thedregister is enough.- Every other floating-point register,
d0tod7andd16tod31, may be changed by any call,printfincluded.
The next program turns the formula a + (b - a) * t into a subroutine. This linear interpolation finds the point a fraction t of the way from a to b: t = 0 gives a, t = 1 gives b, and t = 0.5 the point halfway. lerp takes a, b, and t in d0, d1, and d2, does the work with one fsub and one fmadd, and returns the answer in d0. main walks t from 0 to 1 in steps of 0.25 and keeps it in d8 so that it survives the calls. Because d8 is callee-saved, main stores the old d8 in its frame first and loads it back before it returns.
It prints five lines, one for each quarter: 20.0, 60.0, 100.0, 140.0, and 180.0 degrees.
pitfall
The step of 0.25 and the limit of 1.0 go through d16, which any call may change. That is why the loop sets d16 again just before each use, and never counts on it across bl lerp or bl printf. A value that must outlive a call belongs in d8 to d15, saved in the prologue (the function's opening lines), or in memory.
pitfall
Common mistakes from this lesson, each with a broken program and its fix that you can run:
Check yourself
- Which register holds the format string for
printf, and which holds the first%f? - A function takes
(int count, double scale). Where does each argument arrive? - Why does a single need
fcvt d0, s0before it is printed? - A loop keeps a running total in
d5and callsprintfon every pass. What goes wrong, and what is the fix? fmov d1, 0.1does not assemble. What do you write instead?
answers
show answers
x0holds the format string andd0the first%f.countinw0andscaleind0: integer and floating-point arguments are counted separately.printfonly ever receives doubles, because C widens afloatpassed to it, so%freads 64 bits.printfmay changed0tod7, so the total can be lost. Keep it in one ofd8tod15and save that register in the prologue.- Put
.double 0r0.1in.dataunder a label, then load its address andldr d1from it.
Practice
- Average temperature: read doubles into an array and print their mean.
- Newton's square root: close in on a square root with
fdivandfadd, nofsqrt. - Basic quiz: floating point, then core and challenge: the registers, the instructions, and the calling rules.
- Fill in the blank: floating point (core) and Predict: floating point (core): the directives, the conversions, and what each instruction leaves in a register.