Bit, Nibble, Byte

Three words for three sizes, and none of them was inevitable. A byte was not always eight bits, a nibble is a joke that stuck, and a machine's word size decided which bases grouped its bits evenly. Other bases still work, with a shorter group at the left.

One byte, four notations

Word size
    the byte, IBM System/360

      Hex groups them in fours. On a byte that comes out even, and each group is one hex digit. That group is a nibble.

      Octal groups them in threes from the right. On a byte that leaves a ragged two on the end. Change the word size above and watch which base stops fitting.

      Binary
      0100 1010
      Hex
      4A
      Octal
      112
      Decimal
      74
      Signed
      +74
      BCD
      4 ×
      ASCII
      J

      Flip any bit and watch which notations move.

      Where the three words came from

      Bit is the oldest, and the only one of the three coined to shorten something people were already writing out. John Tukey contracted "binary digit" to "bit" in a Bell Labs memo on 9 January 1947. Claude Shannon put it into the literature the following year in A Mathematical Theory of Communication, and credited Tukey for it there.

      Byte was coined by Werner Buchholz in 1956, during the design of the IBM 7030, the machine everyone called Stretch. He spelled it with a y deliberately, so that a typo could not turn it into "bit". On Stretch a byte was not eight bits: it was whatever you addressed, anywhere from one to eight.

      Nibble is a pun. A byte sounds like a bite, so half of one is a nibble. It is credited to David Benson around 1958 and does not show up in print much before the mid-1970s. It survived because it turned out to name something real: four bits is exactly one hexadecimal digit.

      Why this comes after Shannon on the timeline

      The order looks wrong. Bitwise sits at 1937, ten years earlier, and it is about pushing bits around. How can the operations predate the word for the thing they operate on?

      Because they did. Shannon showed in 1937 that a relay circuit is a piece of algebra, and engineers spent the next decade building machines that added, shifted and masked binary digits. They just wrote out binary digit every time, or said "pulse", or "position", or nothing at all.

      Names arrive late, when someone gets tired of the long version or needs to draw a line that was not there before. Tukey was tired of writing binary digit. Buchholz needed a word for however many bits an instruction reached at once, and that only became a question when Stretch let each instruction pick a different number. Nobody wanted a word for four bits until hexadecimal made four bits worth exactly one character.

      So the gap between 1937 and 1947 is not a gap in the engineering. It is the gap between doing something and finding it worth naming.

      Why eight

      Eight was an argument, not a discovery. IBM's SPREAD task force recommended the eight-bit byte in December 1961, and it was settled when System/360 was announced on 7 April 1964. Fred Brooks, its architect, pushed for eight against internal pressure to make it four or six, which would have been cheaper.

      The case for eight was not that ASCII had landed, because it had not: IBM’s SPREAD committee recommended eight in December 1961, and X3.4 approved ASCII in June 1963. The argument Brooks made was about what a character set would need. Six bits cannot hold upper case, lower case, the digits and punctuation at once, and lower case is what the machines of the day mostly did without. Eight can, with room left over, and it makes a byte hold 256 values. Seven would have done for the character set alone; eight is also a power of two, which is what makes the addressing arithmetic a shift.

      Most machines since have inherited that argument, and the ones that did not are the reason the page says most. The PDP-10 kept a 36-bit word and bytes of whatever width you asked for until the 1980s, and some digital signal processors still address memory in units that are not eight bits.

      Why your machine picked a base

      An octal digit is three bits. A hex digit is four. So a machine whose word size divides by three reads naturally in octal, and a machine built out of eight-bit bytes reads naturally in hex. That is the whole rule, and it explains a split that looks like tribal preference from the outside.

      DEC's PDP-8 had a 12-bit word and its programmers wrote octal, though twelve divides evenly into hex too and either base would have fit. The PDP-10's 36-bit word is twelve octal digits exactly. System/360 addressed eight-bit bytes, and two hex digits are exactly one byte, so its programmers wrote hex. Hex won in the end because the byte won.

      The rule is only divisibility, and the switcher above is set up to prove it rather than to flatter it. Six bits and eighteen bits divide by three and not by four, so hex is the ragged one there. Twelve divides by both, so a PDP-8 word reads cleanly either way: four octal digits or three hex digits. Octal did not suit "old machines" and hex "new" ones; each suited the widths it divided.

      The budget this page names is the one Morse, Baudot and the punch card had all been paying for a century without a word for it. Eight is where it settled.

      These ran in this browser when the page loaded. Each claim, whether it held, and the number behind it.

      Each claim, whether it held, and the values behind it
      claimheldmeasured
      the byte divides by 4 and not by 3; 6 and 18 divide by 3 and not by 4; 12 divides by bothyesso the argument is not that word size DICTATES the base -- the PDP-8's twelve bits would take either, and it wrote octal by lineage. The byte has a shorter leading octal group; six and eighteen bits have a shorter leading hex group. Both bases still represent every value
      grouping a byte in threes from the right gives 3 groups, the leftmost holding only 2 bitsyesagainst exactly 2 whole nibbles; you can see which one the byte is shaped for without being told
      all 256 values a byte can hold read back identically from binary, octal and hexyesevery readout is derived from one integer on every change, so the four notations cannot disagree; there is no table here
      setting and clearing each of the 8 bits in turn puts it exactly where the page draws ityesindex 0 is the most significant, so the array reads left to right the way the number is written; an off-by-one here would mirror the display and nothing else would notice
      control codes are named rather than printed, and "not ASCII" is what a byte above 127 getsyesprinting a control code shows nothing and calls it a character; 128 to 255 are above the seven-bit set, which is the point rather than an omission

      What is real here, and what is not

      The arithmetic is real, and there is only one of it

      Every readout is computed from a single integer on every change. There is no lookup table of pre-rendered answers, which is why the four notations cannot disagree with each other. The groupings are drawn from the same width the value uses, so the ragged octal group is a consequence of the arithmetic rather than a picture of one.

      The word sizes are chosen to be awkward

      Six, eight, twelve and eighteen are not a sample of famous machines; they show the three cases these widths can produce. Eight is ragged in octal, six and eighteen are ragged in hex, and twelve is ragged in neither. A width ragged in both, such as ten, is not shown. A widget offering only 8 and 12 would appear to show that octal belongs to older machines, and that is not what the arithmetic says. Thirty-six bits is omitted for space, and it behaves like twelve: even in both.

      BCD is shown as a reading, not as arithmetic

      Each nibble is read as a decimal digit, and 1010 through 1111 are marked with a cross because they are not digits. A real machine doing BCD arithmetic also carries at ten rather than sixteen, which this page does not do; it only decodes. The reason BCD is here at all is that four bits is the storage one BCD digit needs, which is the size the word nibble names. Whether it was coined for that purpose, as a pun on a small bite, is a story the honesty note below treats as the best available rather than as a fact: the attribution rests on one recollection and the printed record starts much later.

      The dates are the contested part

      1947 for "bit" and 1964 for the eight-bit byte are firm. 1956 for "byte" is well attested but sources differ on the month. The 1958 attribution for "nibble" rests on one person's recollection and the printed record only starts in the 1970s, so treat it as the best available story rather than as a settled date.

      ASCII above 127 is not ASCII

      ASCII is a seven-bit set, so any value from 128 up is outside it and the readout says so rather than guessing at a character. What those values mean depends on a code page, and picking one silently is exactly the kind of thing that makes text turn to mojibake.

      Sources