Inside a float: IEEE 754
prerequisite
A general register holds whole numbers only. To store 6.75, or 0.1, or a number as large as 10^30, a program needs a format with a fractional part and a very wide range. That format is floating point: each value is stored as a sign, a power of two that sets its scale, and a fixed number of significant bits. Almost every processor, AArch64 included, lays those three parts out the way a standard called IEEE 754 defines, and the s and d registers from the floating-point lesson hold their values in exactly this layout.
This lesson opens those registers up. By the end you can split a single or double into its sign, exponent and fraction, turn a decimal number into its bits and back, and tell zero, subnormal numbers, infinity and NaN apart from the bits alone.
Fractions in binary
Digits after a binary point work like digits after a decimal point, except that each place is worth half of the one to its left: 1/2, 1/4, 1/8, 1/16, and so on. So 110.11 in binary is 4 + 2 + 1/2 + 1/4 = 6.75.
To turn a decimal fraction into binary, keep doubling it. Each doubling pushes one bit across the point: the whole-number part (0 or 1) is the next bit, and only what is left after the point carries on to the next row.
value doubled next bit0.75 1.5 1 keep 0.50.5 1.0 1 nothing left, so 0.75 = 0.110.1 0.2 00.2 0.4 00.4 0.8 00.8 1.6 1 keep 0.60.6 1.2 1 keep 0.2, a value already seen so 0.1 = 0.000110011001100... without endThe bits of 0.1 never end, in the same way that 1/3 = 0.333... never ends in decimal. A float has room for only so many bits, so 0.1 is stored as the nearest value that fits, not as 0.1 itself. The programs below show exactly how near.
Fixed point, and why floats move the point
One simple way to store fractions is fixed point: agree that the binary point sits at a set place, say 8 bits from the right, and store the number times 2^8 as an ordinary integer. 6.75 becomes 6.75 x 256 = 1728. Integer instructions can add and compare these directly, which is why small devices with no floating-point hardware still use them. The catch is range: with the point fixed, the same 32 bits cannot hold both 0.0001 and ten billion.
Floating point lets the point move. Any number other than zero can be written in binary scientific notation, as 1.something times a power of two:
6.75 = 110.11 in binary = 1.1011 x 2^20.15625 = 0.00101 in binary = 1.01 x 2^-3Moving the point until exactly one 1 sits to its left is called normalizing. After normalizing, that leading digit is always 1, so the format does not spend a bit storing it; it is called the hidden bit. What is left to store is the sign, the power of two (the exponent), and the bits after the point (the fraction).
The single-precision layout
A single-precision float, the 32-bit kind that sits in an s register and that .float places in memory, splits its bits three ways:
bit 31 30 ........ 23 22 ............................. 0field s e (8 bits) f (23 bits)value 1.f x 2^(e - 127), negated when s is 1s, the sign bit, is 0 for positive and 1 for negative. Flipping it negates the number and changes nothing else, which is allfnegdoes.e, the exponent, is stored with a bias: the true power of two plus 127. The bias keeps the field between 0 and 255, so it needs no sign of its own. Aneof 127 means 2^0, 130 means 2^3, and 124 means 2^-3.f, the fraction, is the 23 bits after the point of the normalized number, padded with zeros on the right. The hidden 1 goes in front of them.
Encoding -6.75:
sign negative s = 1binary 6.75 = 110.11 = 1.1011 x 2^2exponent 2 + 127 = 129 e = 10000001fraction the bits after the point, 1011 f = 1011000000000000000000032 bits 1 10000001 10110000000000000000000by fours 1100 0000 1101 1000 0000 0000 0000 0000hex c 0 d 8 0 0 0 0 = 0xc0d80000Decoding a word
Decoding runs the same steps backwards: split the bits 1, 8 and 23; subtract 127 from e; put the hidden 1 in front of f; and move the point by the true exponent. The IEEE 754 view of the base converter shows the three fields of any word you type, which makes it a quick check for a hand decoding. For 0x42a50000:
binary 0100 0010 1010 0101 0000 0000 0000 0000split 0 | 10000101 | 01001010000000000000000s 0, so positivee 10000101 = 133, and 133 - 127 = 61.f 1.0100101move 1.0100101 x 2^6 = 1010010.1value 64 + 16 + 2 + 0.5 = 82.5Range and precision
e values from 1 to 254 give ordinary numbers, so the true exponent runs from -126 to 127. The largest single is just under 2^128, about 3.4 x 10^38, and the smallest ordinary one is 2^-126, about 1.2 x 10^-38. The two leftover e values, 0 and 255, are kept for special cases, covered further down.
Precision is set by the fraction: 23 stored bits plus the hidden bit give 24 significant bits, about 7 decimal digits. Past that, even whole numbers go missing: 16,777,217 (2^24 + 1) has no single-precision form, and converting it with scvtf gives 16,777,216.
Taking floats apart in a program
The program below prints four single-precision values beside their bits and fields. Its helper, show_float, works in three steps:
fmov bits_r, s0(withbits_rstanding forw9) copies the 32 bits ofs0into a general register unchanged. It copies the pattern and does not convert the value: for 1.0 the integer it leaves is 1,065,353,216 (0x3f800000), not 1.ubfxpulls outs(bit 31),e(bits 30 to 23) andf(bits 22 to 0), andsubtakes the bias offe.fcvt d0, s0widens the value to a double before the call, becauseprintfalways receives a floating-point argument as a double.
Running it prints:
1 0x3f800000 s = 0 e = 127 e - 127 = 0 f = 0x000000-6.75 0xc0d80000 s = 1 e = 129 e - 127 = 2 f = 0x5800000.100000001 0x3dcccccd s = 0 e = 123 e - 127 = -4 f = 0x4ccccdinf 0x7f800000 s = 0 e = 255 e - 127 = 128 f = 0x000000Reading the output:
- 1.0 is 1.0 x 2^0, so
eequals the bias andfis all zeros. - -6.75 matches the hand encoding above.
fprints as0x580000because its 23 bits,1011and nineteen zeros, regroup in fours from the right as101 1000 0000 0000 0000 0000. - 0.1 is 1.6 x 2^-4. Its fraction is a pattern of bits that repeats without end (it shows as the run of
cdigits inf), so the last bit is rounded up, andfends indwhere the pattern alone would givec. Printed to nine significant digits, the stored value shows as 0.100000001. - 1.0 / 0.0 does not stop the program. It gives infinity, which has
e= 255 andf= 0.
Change 0r-6.75 to 0r82.5 and run again: the second line shows 0x42a50000, the word decoded by hand above.
Zero, subnormals, infinity, and NaN
The two e values left out of ordinary numbers mark the special cases:
e | f | Meaning |
|---|---|---|
| 0 | 0 | zero; s gives +0 or -0 |
| 0 | not 0 | subnormal: 0.f x 2^-126, with no hidden 1 |
| 1 to 254 | any | normal: 1.f x 2^(e - 127) |
| 255 | 0 | infinity; s gives plus or minus |
| 255 | not 0 | NaN, short for "not a number" |
- Zero has no leading 1 to hide, so it gets its own pattern:
eandfall zeros. That leaves two zeros, +0 and -0, whichfcmpreports as equal. - A subnormal number is smaller than the smallest normal one. With
e= 0 the hidden bit becomes 0, so these values fill the gap between 2^-126 and zero, at the price of fewer significant bits. The smallest,0x00000001, is 2^-149, about 1.4 x 10^-45. - Infinity is the result when the true answer is too large to store, or when a nonzero number is divided by zero.
- A NaN marks a result with no sensible value, such as 0.0 / 0.0 or infinity minus infinity;
0x7fc00000is the NaN that AArch64 produces for these. A NaN compares unequal to everything, itself included: afterfcmpwith a NaN on either side,b.eqis not taken.
The program below reads each pattern's fields and names its kind. main keeps its pointer and its counter in x19 and x20 because they must survive every printf call, so it saves those two registers on entry and restores them before returning, as the calling conventions lesson requires. Replace any word in the .word lists with a hex word of your own, such as one from a practice question, to classify it.
Running it prints:
0x00000000 s = 0 e = 0 f = 0x000000 zero0x80000000 s = 1 e = 0 f = 0x000000 zero0x00000001 s = 0 e = 0 f = 0x000001 subnormal0x00800000 s = 0 e = 1 f = 0x000000 normal0x7f7fffff s = 0 e = 254 f = 0x7fffff normal0x7f800000 s = 0 e = 255 f = 0x000000 infinity0xff800000 s = 1 e = 255 f = 0x000000 infinity0x7fc00000 s = 0 e = 255 f = 0x400000 NaNDouble precision
A double-precision float, the 64-bit kind in a d register and made by .double, follows the same rules with wider fields:
bit 63 62 ........ 52 51 ............................. 0field s e (11 bits) f (52 bits)value 1.f x 2^(e - 1023), negated when s is 1The bias is 1023, the special cases are e = 0 and e = 2047, and 53 significant bits give about 16 decimal digits, over a range out to about 1.8 x 10^308. -6.75 as a double keeps the same sign and the same fraction bits 1011; only the exponent field is wider, holding the true exponent 2 as 2 + 1023 = 1025:
s e f 1 10000000001 1011 followed by 48 zerosby fours 1100 0000 0001 1011 0000 ... 0000hex 0xc01b000000000000The program below does the same split for doubles, with x registers and wider ubfx fields. It prints 17 significant digits, which is enough to tell any two doubles apart.
Running it prints:
1 0x3ff0000000000000 s = 0 e = 1023 e - 1023 = 0 f = 0x0000000000000-6.75 0xc01b000000000000 s = 1 e = 1025 e - 1023 = 2 f = 0xb0000000000000.10000000000000001 0x3fb999999999999a s = 0 e = 1019 e - 1023 = -4 f = 0x999999999999a0.10000000149011612 0x3fb99999a0000000 s = 0 e = 1019 e - 1023 = -4 f = 0x99999a0000000The third and fourth values both started as 0.1. The double 0.1 is off by about 5.6 x 10^-18. The fourth line is the single 0.1 widened with fcvt, and it is off by about 1.5 x 10^-9, the single's own error. Widening only appends zero bits to the fraction (the single's f, 0x4ccccd, moved left 29 places is 0x99999a0000000), so the error comes along unchanged. Converting to a wider type never brings back bits that were already rounded away.
pitfall
The usual slips when decoding by hand: forgetting to subtract the bias (127 for a single, 1023 for a double); forgetting the hidden 1 in front of the fraction (it is there for every normal number and missing for subnormals); reading f as decimal digits, when they are binary digits after the point; and splitting a double 1, 8, 23 instead of 1, 11, 52. The matching slip in code is using scvtf or fcvtzs where fmov was meant: those two convert the value, while fmov between a general register and a floating-point register copies the bits.
Check yourself
- Decode
0xc1200000. - Encode 0.375 as a single-precision word.
- Name the kind of value each word holds:
0x80000000,0x7f800001,0x00400000.
answers
show answers
sis 1,eis 130 so the true exponent is 3, and 1.f is 1.01, so the value is -1.01 x 2^3 = -1010 in binary, which is -10.0.- 0.375 = 0.011 = 1.1 x 2^-2, so
sis 0,eis 125 (01111101), andfis a 1 then 22 zeros, giving0x3ec00000. - Negative zero (
s1,e0,f0); NaN (e255,fnot 0); subnormal (e0,fnot 0).
Practice
- Float anatomy: split a float into its sign, exponent and fraction, and name what kind of value it is.
- Build a float from parts: the same layout, run backwards.
- Advanced quiz: floating point: how
fmovfits a constant into 8 bits (a sign, 3 exponent bits and 4 fraction bits, this lesson's layout in miniature). - Fill in the blank: floating point (core):
.floatand.double. - Predict: floating point (core):
fnegandfabs, which touch only the sign bit.