Binary Code: How Computers Understand Text

September 7, 2026

In school computer science class, they show you a string of zeros and ones and say: "this is how a computer sees text." If you sat there thinking that the word "hello" looks like a magical jumble of bytes to the processor — you weren't alone. The mechanics are actually simple, and you can try them hands-on: convert any text to binary in our binary translator and see what you get.

A zero and a one are just a switch

Physically, a computer can only tell two states apart: current flowing or no current. One means yes, zero means no. A single cell like this is called a bit, and on its own it's next to useless: the most one bit can say is "yes" or "no." But take eight bits in a row (a byte) and you already have 256 combinations — enough for a whole alphabet, digits, and punctuation.

From byte to letter: the code table

For a byte to become a letter, there has to be an agreement about which combination stands for which character. That agreement is called an encoding. Classic ASCII declared: 01000001 is a capital Latin A, 01000010 is B, and so on. It had no room for Cyrillic, so extensions appeared, and eventually a single standard — Unicode — where a single character can take up several bytes.

That's where the familiar phrase "a byte per character" comes from. In UTF-8, a Latin letter takes one byte, while a Cyrillic letter takes two. That's why Russian text in binary form is always longer than English: "Privet" (Russian for "hello") is six letters and twelve bytes.

The letter A and the letter П
A → 01000001 (1 byte, ASCII)
П → 11010000 10011111 (2 bytes, UTF-8)

Why some letters are "heavier": Cyrillic and bytes

In UTF-8, a Latin letter takes one byte and a Cyrillic letter takes two. The word "hello" weighs 5 bytes, while "privet" already weighs 12. The reason is that eight bits give you only 256 combinations: enough for the Latin alphabet, digits, and punctuation, but Cyrillic needs a second "page" of codes, and a character from that page is written with two bytes. So Russian text in binary is almost twice as long — and "why does Cyrillic take up more bits" is a question with a precise answer: the letters themselves aren't heavier, their codes simply don't fit in a single byte.

Numbers are simpler: any value from 0 to 255 fits in a single byte. The number 243 in 8-bit binary is written as 11110011 — 128 + 64 + 32 + 16 + 2 + 1. You can check it by hand: write out the powers of two from left to right (128, 64, 32, 16, 8, 4, 2, 1) and add up the ones with a 1 underneath.

How to convert text yourself

  1. take the first character of the text;
  2. look up its code in the encoding — it's a number;
  3. convert the number to binary (divide by 2 keeping the remainders, then read them bottom-up);
  4. repeat for every character and join the results with spaces.
Doing this by hand is interesting exactly once — to understand the mechanics. After that it's faster to open the binary code converter: it translates text to bits and back, in both directions, Cyrillic and emoji included.

The reverse task is just as useful, by the way. If a file or a message made of zeros and ones lands on your desk, knowing the encoding lets you read it. Get the encoding wrong, and you end up with the classic mojibake — the gibberish that pops up when a document is opened in the wrong character set.

How many bits are in a byte, and why eight

One byte is always 8 bits, which gives 2⁸ = 256 possible combinations: from 00000000 to 11111111. Eight wasn't the immediate standard: early text encodings made do with seven bits per letter (128 slots covered Latin letters, digits and punctuation), and the eighth bit was sometimes used for error checking. When Cyrillic and other alphabets needed a home, that eighth bit was carved up — that's how 256-character tables like Windows-1251 appeared. Hence the school rule of thumb: 1 byte = 8 bits, 1 kilobyte = 1024 bytes = 2¹⁰ bytes — the whole system is built on powers of two.

A quick eyeball check: every group of eight 0s and 1s is one character. If there are no spaces between letters, count the eights — it's the easiest way to estimate how many characters a long binary string encodes.

Why this is worth knowing in practice

  • debugging: when data arrives as garbled text, you'll know right away that the problem is the encoding, not the program;
  • sizes: knowing that Cyrillic weighs twice as much as Latin text explains why the same text takes up a different amount of space in different systems;
  • exams and interviews: converting between number systems is a classic question, and the base converter helps you check yourself.

Binary seems like an abstraction right up until the first real-world use case. Convert your own name to bits — and the abstraction turns concrete.

Related articles