AArch64 Playground
4.14 · Binary, hex, and octal: one number, many spellings

Binary, hex, and octal: one number, many spellings

A register holds bits and nothing else. The same 32 bits can be a count, a temperature below zero, or a letter, and the machine does not know which. The meaning comes from the instruction that uses the bits or, when you print them, from the format specifier you give printf.

This lesson reads one bit pattern several ways: as an unsigned number, as a signed number, and in hexadecimal and octal, two shorter ways to write binary. It also gives the range of values each register size can hold.

Bits and place value

A bit is one binary digit, 0 or 1. A binary number works like a decimal one, except that each place is worth twice the place to its right instead of ten times. From the right, the places are worth 1, 2, 4, 8, 16, and so on.

Bits are numbered from the right, starting at 0. Bit 0 is the least significant bit, worth 1. In an 8-bit value, bit 7 is the most significant bit, worth 128. Here is 200 in eight bits:

bit76543210
worth1286432168421
20011001000

To turn a decimal number into binary, take away the largest place value that fits, put a 1 in that place, and repeat with what is left. For 200: 200 - 128 = 72, 72 - 64 = 8, 8 - 8 = 0. Bits 7, 6 and 3 are 1 and the rest are 0, so 200 is 11001000.

Hex and octal: binary in short form

Rows of bits are hard to read and easy to miscount, so people write them in hexadecimal (hex for short, base 16) or in octal (base 8).

Hex works because 16 is 2 to the power 4, so each hex digit stands for exactly four bits. The digits are 0 to 9 and then a to f for 10 to 15. To convert, split the bits into groups of four starting from the right and write one digit per group:

1100 1000 becomes c 8, which is 0xc8.

Octal uses groups of three, because 8 is 2 to the power 3, and its digits run from 0 to 7:

11 001 000 becomes 3 1 0, which is 0310.

Going back is the same step in reverse: each hex digit becomes four bits and each octal digit becomes three. A 32-bit value always takes 8 hex digits. It takes 11 octal digits, and since 32 is not a multiple of 3, the top octal digit covers only two bits and is never more than 3. Octal comes back later for Linux file permissions, where each digit holds three yes-or-no bits. To check a conversion, type the number into the base converter, which shows it in binary, octal, decimal and hex at once.

pitfall

In source code, 0x starts a hex number and 0b starts a binary one, and a number that starts with a plain 0 is octal. The assembler reads 010 as eight, not ten, so never pad a decimal number with leading zeros.

Sizes and unsigned ranges

Data comes in four sizes. A byte is 8 bits, a halfword is 16 bits, a word is 32 bits, and a doubleword is 64 bits. A w register holds a word and an x register holds a doubleword.

C on 64-bit Linux follows the LP64 model: long values and pointers are 64 bits (the L and the P in the name), while int stays at 32.

sizebitsC typeregister view
byte8charlow 8 bits of a register
halfword16shortlow 16 bits of a register
word32intw0 to w30
doubleword64long, every pointerx0 to x30

With n bits there are 2 to the power n different patterns. Read as unsigned numbers, which have no sign and are never negative, they run from 0 up to 2 to the power n, minus 1. A byte holds 0 to 255, and a word holds 0 to 4,294,967,295.

Negative numbers: two's complement

Signed numbers use two's complement: the top bit counts as a negative amount instead of a positive one, and every other place keeps its usual value. In a byte, bit 7 is worth -128 instead of +128, so the pattern 11001000 reads as -128 + 64 + 8 = -56.

The top bit is called the sign bit: when it is 1, the number is negative. So the same eight bits are 200 when read as unsigned and -56 when read as signed. When the sign bit is 1, the two readings always differ by 2 to the power n (256 for a byte); when it is 0, they are the same number.

The signed range is lopsided by one, because zero takes one of the patterns with the sign bit off:

sizeunsigned rangesigned range
byte0 to 255-128 to 127
halfword0 to 65,535-32,768 to 32,767
word0 to 4,294,967,295-2,147,483,648 to 2,147,483,647
doubleword0 to 18,446,744,073,709,551,615-9,223,372,036,854,775,808 to 9,223,372,036,854,775,807

note

All ones is -1 at every size. In a byte, 11111111 is -128 + 64 + 32 + 16 + 8 + 4 + 2 + 1 = -1. That is why -1 printed as an unsigned number shows the largest value the size can hold.

One pattern, four readings

printf has a specifier for each reading of a word: %d prints it as signed, %u as unsigned, %x as hex (without the 0x), and %o as octal (without the leading 0). Putting an l in front, as in %ld, %lu, %lx and %lo, makes printf read a whole x register instead of a w register.

The program below passes the same register value to all four specifiers at once, one row per value. In the header, %% prints a single %, so each column is labeled with its own specifier.

Each value goes into w9 just before the call and is not needed after it. w9 is a scratch register: a register whose value does not have to survive a call, because printf is free to change x0 to x18.

loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 1try it: run it, or step one instruction at a timeOpen in playground

The 200 row reads the same in the first two columns, because the sign bit of the word (bit 31) is 0. The -56 row ends in c8, the same low byte as 200, and its unsigned column shows 4294967240, which is -56 + 4,294,967,296. The -1 row is all ones: 4294967295 as unsigned, ffffffff in hex, and 37777777777 in octal, whose top digit is 3 because only two bits are left for it.

The last four lines hand an x register holding -1 to the l specifiers, and print 18446744073709551615, sixteen fs, and 1777777777777777777777.

Change 200 to 0xc8 or to 0b11001000 and run again. The output does not change, because all three spellings put the same bits into the register.

Negating by hand

To negate a two's complement number, flip every bit and then add 1. Flipping alone gives -n - 1, because a number and its flipped copy add up to all ones, which is -1. The extra 1 turns that into -n.

mvn flips every bit of a register (the name reads as "move not"; Binary logic: masks, flags, and tst covers it with the other logic instructions), and neg does the flip and the add in one step.

The program below negates 13 by hand and prints each step twice: in hex with %08x, which pads to eight digits with zeros, and as a signed number with %d. 13 and its negation are needed after several calls, so they live in w19 and w20: printf may change x0 to x18, but it has to give x19 to x28 back unchanged. This main uses w19 and w20 without saving them first, as the programs in these lessons do. A subroutine you write yourself has to save any of x19 to x28 that it changes; Subroutines: saved registers, pointers, and big arguments shows how.

loading editor...

regfile

N clearZ clearC clearV clear

x0–x30 are the integer registers.

X0arg00x0000000000000000
X1arg10x0000000000000000
X2arg20x0000000000000000
X3arg30x0000000000000000
X4arg40x0000000000000000
X5arg50x0000000000000000
X6arg60x0000000000000000
X7arg70x0000000000000000
X8ind0x0000000000000000
X90x0000000000000000
X100x0000000000000000
X110x0000000000000000
X120x0000000000000000
X130x0000000000000000
X140x0000000000000000
X150x0000000000000000
X16ip00x0000000000000000
X17ip10x0000000000000000
X18pr0x0000000000000000
X190x0000000000000000
X200x0000000000000000
X210x0000000000000000
X220x0000000000000000
X230x0000000000000000
X240x0000000000000000
X250x0000000000000000
X260x0000000000000000
X270x0000000000000000
X280x0000000000000000
X29fp0x0000000000000000
X30lr0x0000000000000000
SP0x0000000080000000
PC0x0000000000400000
console

Output prints here as your program runs.

Press step or run under the editor, or feed stdin from the box below.

not assembled

example 2try it: run it, or step one instruction at a timeOpen in playground

It prints 0000000d = 13, then fffffff2 = -14 after the flip and fffffff3 = -13 after adding one. neg gives the same fffffff3.

The last line adds 13 and -13 and gets 00000000 = 0. The true sum is 0x100000000: the addition produces a carry, a 1 that moves into the next place, out of bit 31. A w register has no bit 32 to keep it in, so it is dropped and zero is left. Dropping that carry is what lets one adder work for signed and unsigned numbers alike: the same bits come out either way, and only the reading differs.

Check yourself

  1. What is the byte 10000001 as an unsigned number, and as a two's complement number?
  2. Write 0x2f in binary and in octal.
  3. w1 holds -2. What does printf print for it with %x, and with %u?
  4. What is the largest signed halfword, in decimal and in hex?

answers

show answers
  1. 129 unsigned; -128 + 1 = -127 signed.
  2. 0010 1111 in binary; regrouped as 00 101 111, that is 057 in octal (47 in decimal).
  3. fffffffe and 4294967294.
  4. 32,767, which is 0x7fff: every bit set except the sign bit.

Practice