AArch64 Playground
4.32 · Inside a float: IEEE 754

Inside a float: IEEE 754

A general register holds whole numbers only. To store 6.75, or 0.1, or a number as large as 10^30, a program needs a format with a fractional part and a very wide range. That format is floating point: each value is stored as a sign, a power of two that sets its scale, and a fixed number of significant bits. Almost every processor, AArch64 included, lays those three parts out the way a standard called IEEE 754 defines, and the s and d registers from the floating-point lesson hold their values in exactly this layout.

This lesson opens those registers up. By the end you can split a single or double into its sign, exponent and fraction, turn a decimal number into its bits and back, and tell zero, subnormal numbers, infinity and NaN apart from the bits alone.

Fractions in binary

Digits after a binary point work like digits after a decimal point, except that each place is worth half of the one to its left: 1/2, 1/4, 1/8, 1/16, and so on. So 110.11 in binary is 4 + 2 + 1/2 + 1/4 = 6.75.

To turn a decimal fraction into binary, keep doubling it. Each doubling pushes one bit across the point: the whole-number part (0 or 1) is the next bit, and only what is left after the point carries on to the next row.

value   doubled   next bit0.75    1.5       1          keep 0.50.5     1.0       1          nothing left, so 0.75 = 0.110.1     0.2       00.2     0.4       00.4     0.8       00.8     1.6       1          keep 0.60.6     1.2       1          keep 0.2, a value already seen                             so 0.1 = 0.000110011001100... without end

The bits of 0.1 never end, in the same way that 1/3 = 0.333... never ends in decimal. A float has room for only so many bits, so 0.1 is stored as the nearest value that fits, not as 0.1 itself. The programs below show exactly how near.

Fixed point, and why floats move the point

One simple way to store fractions is fixed point: agree that the binary point sits at a set place, say 8 bits from the right, and store the number times 2^8 as an ordinary integer. 6.75 becomes 6.75 x 256 = 1728. Integer instructions can add and compare these directly, which is why small devices with no floating-point hardware still use them. The catch is range: with the point fixed, the same 32 bits cannot hold both 0.0001 and ten billion.

Floating point lets the point move. Any number other than zero can be written in binary scientific notation, as 1.something times a power of two:

6.75      = 110.11 in binary    = 1.1011 x 2^20.15625   = 0.00101 in binary   = 1.01   x 2^-3

Moving the point until exactly one 1 sits to its left is called normalizing. After normalizing, that leading digit is always 1, so the format does not spend a bit storing it; it is called the hidden bit. What is left to store is the sign, the power of two (the exponent), and the bits after the point (the fraction).

The single-precision layout

A single-precision float, the 32-bit kind that sits in an s register and that .float places in memory, splits its bits three ways:

bit      31   30 ........ 23   22 ............................. 0field    s    e (8 bits)       f (23 bits)value    1.f  x  2^(e - 127),  negated when s is 1
  • s, the sign bit, is 0 for positive and 1 for negative. Flipping it negates the number and changes nothing else, which is all fneg does.
  • e, the exponent, is stored with a bias: the true power of two plus 127. The bias keeps the field between 0 and 255, so it needs no sign of its own. An e of 127 means 2^0, 130 means 2^3, and 124 means 2^-3.
  • f, the fraction, is the 23 bits after the point of the normalized number, padded with zeros on the right. The hidden 1 goes in front of them.

Encoding -6.75:

sign       negative                          s = 1binary     6.75 = 110.11 = 1.1011 x 2^2exponent   2 + 127 = 129                     e = 10000001fraction   the bits after the point, 1011    f = 1011000000000000000000032 bits    1 10000001 10110000000000000000000by fours   1100 0000 1101 1000 0000 0000 0000 0000hex        c    0    d    8    0    0    0    0     = 0xc0d80000

Decoding a word

Decoding runs the same steps backwards: split the bits 1, 8 and 23; subtract 127 from e; put the hidden 1 in front of f; and move the point by the true exponent. The IEEE 754 view of the base converter shows the three fields of any word you type, which makes it a quick check for a hand decoding. For 0x42a50000:

binary     0100 0010 1010 0101 0000 0000 0000 0000split      0 | 10000101 | 01001010000000000000000s          0, so positivee          10000101 = 133, and 133 - 127 = 61.f        1.0100101move       1.0100101 x 2^6 = 1010010.1value      64 + 16 + 2 + 0.5 = 82.5

Range and precision

e values from 1 to 254 give ordinary numbers, so the true exponent runs from -126 to 127. The largest single is just under 2^128, about 3.4 x 10^38, and the smallest ordinary one is 2^-126, about 1.2 x 10^-38. The two leftover e values, 0 and 255, are kept for special cases, covered further down.

Precision is set by the fraction: 23 stored bits plus the hidden bit give 24 significant bits, about 7 decimal digits. Past that, even whole numbers go missing: 16,777,217 (2^24 + 1) has no single-precision form, and converting it with scvtf gives 16,777,216.

Taking floats apart in a program

The program below prints four single-precision values beside their bits and fields. Its helper, show_float, works in three steps:

  • fmov bits_r, s0 (with bits_r standing for w9) copies the 32 bits of s0 into a general register unchanged. It copies the pattern and does not convert the value: for 1.0 the integer it leaves is 1,065,353,216 (0x3f800000), not 1.
  • ubfx pulls out s (bit 31), e (bits 30 to 23) and f (bits 22 to 0), and sub takes the bias off e.
  • fcvt d0, s0 widens the value to a double before the call, because printf always receives a floating-point argument as a double.

Running it prints:

1            0x3f800000  s = 0  e = 127  e - 127 =    0  f = 0x000000-6.75        0xc0d80000  s = 1  e = 129  e - 127 =    2  f = 0x5800000.100000001  0x3dcccccd  s = 0  e = 123  e - 127 =   -4  f = 0x4ccccdinf          0x7f800000  s = 0  e = 255  e - 127 =  128  f = 0x000000
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 1try it: run it, or step one instruction at a timeOpen in playground

Reading the output:

  • 1.0 is 1.0 x 2^0, so e equals the bias and f is all zeros.
  • -6.75 matches the hand encoding above. f prints as 0x580000 because its 23 bits, 1011 and nineteen zeros, regroup in fours from the right as 101 1000 0000 0000 0000 0000.
  • 0.1 is 1.6 x 2^-4. Its fraction is a pattern of bits that repeats without end (it shows as the run of c digits in f), so the last bit is rounded up, and f ends in d where the pattern alone would give c. Printed to nine significant digits, the stored value shows as 0.100000001.
  • 1.0 / 0.0 does not stop the program. It gives infinity, which has e = 255 and f = 0.

Change 0r-6.75 to 0r82.5 and run again: the second line shows 0x42a50000, the word decoded by hand above.

Zero, subnormals, infinity, and NaN

The two e values left out of ordinary numbers mark the special cases:

efMeaning
00zero; s gives +0 or -0
0not 0subnormal: 0.f x 2^-126, with no hidden 1
1 to 254anynormal: 1.f x 2^(e - 127)
2550infinity; s gives plus or minus
255not 0NaN, short for "not a number"
  • Zero has no leading 1 to hide, so it gets its own pattern: e and f all zeros. That leaves two zeros, +0 and -0, which fcmp reports as equal.
  • A subnormal number is smaller than the smallest normal one. With e = 0 the hidden bit becomes 0, so these values fill the gap between 2^-126 and zero, at the price of fewer significant bits. The smallest, 0x00000001, is 2^-149, about 1.4 x 10^-45.
  • Infinity is the result when the true answer is too large to store, or when a nonzero number is divided by zero.
  • A NaN marks a result with no sensible value, such as 0.0 / 0.0 or infinity minus infinity; 0x7fc00000 is the NaN that AArch64 produces for these. A NaN compares unequal to everything, itself included: after fcmp with a NaN on either side, b.eq is not taken.

The program below reads each pattern's fields and names its kind. main keeps its pointer and its counter in x19 and x20 because they must survive every printf call, so it saves those two registers on entry and restores them before returning, as the calling conventions lesson requires. Replace any word in the .word lists with a hex word of your own, such as one from a practice question, to classify it.

Running it prints:

0x00000000  s = 0  e =   0  f = 0x000000  zero0x80000000  s = 1  e =   0  f = 0x000000  zero0x00000001  s = 0  e =   0  f = 0x000001  subnormal0x00800000  s = 0  e =   1  f = 0x000000  normal0x7f7fffff  s = 0  e = 254  f = 0x7fffff  normal0x7f800000  s = 0  e = 255  f = 0x000000  infinity0xff800000  s = 1  e = 255  f = 0x000000  infinity0x7fc00000  s = 0  e = 255  f = 0x400000  NaN
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 2try it: run it, or step one instruction at a timeOpen in playground

Double precision

A double-precision float, the 64-bit kind in a d register and made by .double, follows the same rules with wider fields:

bit      63   62 ........ 52   51 ............................. 0field    s    e (11 bits)      f (52 bits)value    1.f  x  2^(e - 1023),  negated when s is 1

The bias is 1023, the special cases are e = 0 and e = 2047, and 53 significant bits give about 16 decimal digits, over a range out to about 1.8 x 10^308. -6.75 as a double keeps the same sign and the same fraction bits 1011; only the exponent field is wider, holding the true exponent 2 as 2 + 1023 = 1025:

s e f      1 10000000001 1011 followed by 48 zerosby fours   1100 0000 0001 1011 0000 ... 0000hex        0xc01b000000000000

The program below does the same split for doubles, with x registers and wider ubfx fields. It prints 17 significant digits, which is enough to tell any two doubles apart.

Running it prints:

1  0x3ff0000000000000    s = 0  e = 1023  e - 1023 =     0  f = 0x0000000000000-6.75  0xc01b000000000000    s = 1  e = 1025  e - 1023 =     2  f = 0xb0000000000000.10000000000000001  0x3fb999999999999a    s = 0  e = 1019  e - 1023 =    -4  f = 0x999999999999a0.10000000149011612  0x3fb99999a0000000    s = 0  e = 1019  e - 1023 =    -4  f = 0x99999a0000000
loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 3try it: run it, or step one instruction at a timeOpen in playground

The third and fourth values both started as 0.1. The double 0.1 is off by about 5.6 x 10^-18. The fourth line is the single 0.1 widened with fcvt, and it is off by about 1.5 x 10^-9, the single's own error. Widening only appends zero bits to the fraction (the single's f, 0x4ccccd, moved left 29 places is 0x99999a0000000), so the error comes along unchanged. Converting to a wider type never brings back bits that were already rounded away.

pitfall

The usual slips when decoding by hand: forgetting to subtract the bias (127 for a single, 1023 for a double); forgetting the hidden 1 in front of the fraction (it is there for every normal number and missing for subnormals); reading f as decimal digits, when they are binary digits after the point; and splitting a double 1, 8, 23 instead of 1, 11, 52. The matching slip in code is using scvtf or fcvtzs where fmov was meant: those two convert the value, while fmov between a general register and a floating-point register copies the bits.

Check yourself

  1. Decode 0xc1200000.
  2. Encode 0.375 as a single-precision word.
  3. Name the kind of value each word holds: 0x80000000, 0x7f800001, 0x00400000.

answers

show answers
  1. s is 1, e is 130 so the true exponent is 3, and 1.f is 1.01, so the value is -1.01 x 2^3 = -1010 in binary, which is -10.0.
  2. 0.375 = 0.011 = 1.1 x 2^-2, so s is 0, e is 125 (01111101), and f is a 1 then 22 zeros, giving 0x3ec00000.
  3. Negative zero (s 1, e 0, f 0); NaN (e 255, f not 0); subnormal (e 0, f not 0).

Practice