ARYXTOOLS

Free Online Tools

ARYXTOOLS

Binary and Base64: How Your Computer Turns Letters Into Numbers

Why computers use binary, how ASCII turns a letter into a number, and why Base64 shows up in email, login tokens, and data URLs. With the math worked out step by step.

·7 min read

Type the letter A into any text box and your computer never really stores an A. It stores the number 65, and it stores that number as eight switches in a row, some on and some off: 01000001. Every letter, every emoji, every password you have ever typed goes through the same translation before it becomes ones and zeros. Most of the time nobody needs to think about it. Then a login token looks like gibberish, an email attachment inflates in size for no obvious reason, or a downloaded file opens as garbled symbols, and suddenly the translation underneath everything is worth understanding.

This is that translation, in the order it happens: why computers settled on two digits instead of ten, how those two digits turn into letters, and why one particular encoding called Base64 shows up everywhere from email to login tokens.

Why two digits instead of ten

A transistor, the tiny switch inside every chip, is built to be reliably read as one of two states: on or off, high voltage or low. Reading ten distinct voltage levels reliably, one for each digit 0 through 9, is a much harder engineering problem than reading two, the same way a light switch is simpler and more reliable than a dimmer. Binary is not a workaround. It is the representation that matches what the hardware is good at.

A single binary digit is a bit. Group eight of them together and you get a byte, which can represent 256 different values, 0 through 255, since two options repeated eight times is 2 to the eighth power. Almost everything a computer stores gets measured in bytes for exactly that reason.

How a letter becomes a number

In the early 1960s, computers needed an agreed-upon way to turn letters into numbers, since two machines that each invented their own mapping could not read each other's text. The result was ASCII: A is assigned the number 65, B is 66, and so on, with 128 characters covering the English alphabet, digits, and common punctuation. Store 65 as an 8-bit byte and you get 01000001. That byte is not a stylized or compressed version of the letter A. It is the letter A, as far as the machine is concerned.

Modern text mostly uses UTF-8 instead of plain ASCII, since ASCII's 128 characters have no room for accented letters, most non-English scripts, or emoji. UTF-8 keeps every ASCII character exactly as it was, one byte each, and reaches beyond that range by using 2, 3, or 4 bytes for anything outside the original 128. Plain English text looks identical whether a program assumes ASCII or UTF-8, which is the compatibility that let the whole web move to UTF-8 without breaking older systems along the way.

The Binary to Decimal Converter turns any binary string into the number it represents, and the number it represents is what a byte like 01000001 means once you stop reading it as a pattern and start reading it as 65.

Where plain binary runs into trouble

Binary data is not always safe to send through a system that expects plain text. Early email is the clearest example. SMTP, the protocol email still runs on, was built in 1982 around 7-bit ASCII text, and it was not designed to carry the full range of an 8-bit byte cleanly. Send a raw image or PDF through a 7-bit-only pipe and bytes get stripped, shifted, or dropped, and the file arrives corrupted.

Base64 exists to solve exactly that problem, and it predates the modern web: an early version showed up in the 1987 PEM email standard, and the format was formalized under the name Base64 in 1992. The idea is to represent any binary data using only 64 characters that every text system can be trusted to pass through untouched: the 26 uppercase letters, 26 lowercase letters, the digits 0 through 9, and two more symbols to round out the set to 64. Nothing in that set is a line break, a null byte, or any other character that a text-only pipe might mishandle.

How Base64 encodes something

Base64 reads input 3 bytes at a time, 24 bits, and repacks those 24 bits into four groups of 6 bits each. Six bits gives 64 possible values, one for each character in the Base64 alphabet, so 3 bytes of input become exactly 4 characters of output. That fixed 3-to-4 ratio is also why a Base64 string is always about 33 percent larger than the data it holds, and why an email with attachments is noticeably bigger than the files inside it.

The word hello, 5 bytes, becomes aGVsbG8= in Base64. The trailing equals sign is padding: 5 is not a clean multiple of 3, so the last group is short 1 byte, and the = marks exactly that. Paste any text into the Base64 Encoder / Decoder and watch the output grow by roughly a third, then decode it back to confirm nothing was lost along the way, since Base64 is fully reversible with no data thrown away.

Base64 is not a lock, and it was never meant to be

Because Base64 output looks unreadable at a glance, it gets mistaken for security constantly. It is not one. There is no key involved anywhere in the process, which means decoding it back to the original bytes takes no more effort than encoding it did in the first place. A JSON Web Token, the kind of login token that keeps you signed into a site, is three Base64 sections joined by dots, and anyone can read the header and payload of one without needing to crack anything. HTTP Basic Authentication sends a username and password as a single Base64 string, which is exactly why that method is only ever safe to use over HTTPS. If a password, an API key, or anything genuinely sensitive needs to stay hidden, actual encryption has to happen before Base64 ever gets involved, not instead of it.

That distinction matters most when a password itself is on the line. A strong password's real defense is not that it looks scrambled. It is the number of possible combinations an attacker would have to try, measured in bits, the same binary digits this entire piece has been about, not whatever encoding happens to wrap around it afterward. Generate one with the Password Generator and the randomness is doing the real work, not any encoding layered on top.

One catch: characters and bytes are not always the same count

In UTF-8, plain English letters and digits cost 1 byte each, same as old ASCII always did. Step outside that range and the cost goes up: most accented letters and non-Latin characters take 2 or 3 bytes, and most emoji take 4. A tool that reports a character count and a system that enforces a byte limit are not always counting the same thing, and the gap only shows up once the text stops being plain English. Run text through the Word Counter and that count is characters, not the byte size the text would take up once UTF-8 encodes it.

Frequently asked questions

No, the opposite. Base64 takes 3 bytes of input and turns them into 4 characters of output, so the encoded version is about a third larger than the original. It trades size for safety: every character in the output is guaranteed to survive a text-only channel intact.

No. Encoding and encryption solve different problems. Base64 has no key and no secret, so anyone can decode a Base64 string back to its original bytes in seconds with nothing more than a few lines of code. If something needs to stay secret, it needs real encryption first, with Base64 only used afterward to make the encrypted bytes text-safe for transport.

Base64 works in groups of 3 input bytes at a time. When the input length is not a clean multiple of 3, the last group gets padded, and the equals signs mark how much padding was added. One = means the last group had 2 real bytes, two == means it had only 1.

A transistor is simplest and most reliable when it only needs to represent two states: on or off, high voltage or low. Building a reliable circuit that distinguishes ten different voltage levels for the digits 0 through 9 is far harder than building one that only needs to tell two states apart, so binary won out for the same reason a light switch is more reliable than a dimmer.

Not in UTF-8, the encoding most of the web uses today. Plain English letters and numbers still take 1 byte each, the same as old ASCII. An accented letter or a symbol from a non-Latin script typically takes 2 or 3 bytes, and an emoji usually takes 4. A word counter that reports character counts is not the same number as a byte count once the text leaves plain English.

A bit is a single binary digit, either 0 or 1. A byte is a group of 8 bits, and it is the unit computers generally use to store or move data, since 8 bits together can represent 256 distinct values, from 0 to 255.

JWT tokens used for web login sessions are three Base64 sections joined by dots. A data URI lets a small image sit directly inside a line of HTML or CSS instead of a separate file. HTTP Basic Authentication sends a username and password as one Base64 string in a request header, which is exactly why that method is only safe over HTTPS.

No. Binary, decimal, and Base64 conversion all happen in your browser using plain JavaScript. Nothing you type into any of these tools is uploaded or logged.

More from the blog