Encoding turns data from one representation into another so that it can travel or be stored somewhere that expects a particular format. It is not encryption: there is no key, anyone can reverse it, and it hides nothing. Base64 is the most common example. It writes any bytes using only 64 safe printable characters, so binary data can pass through systems built for text.
Why Base64 exists
Many older protocols were designed for plain text. Email bodies, JSON strings, URLs and HTML attributes can all break or corrupt arbitrary bytes: a stray null byte, a newline in the wrong place or a character that means something special. Base64 avoids the problem by using only A-Z, a-z, 0-9, + and / (plus = for padding). You will see it in email attachments, data: URLs, JSON Web Tokens and the keys in certificate files.
How it works
Base64 reads the input three bytes (24 bits) at a time and re-cuts those 24 bits into four groups of 6 bits. Each 6-bit group is a number from 0 to 63, and each number picks a character from the alphabet (A is 0, Z is 25, a is 26, z is 51, 0 is 52, 9 is 61, + is 62, / is 63).
Take the three letters Man:
| Character | M | a | n |
|---|---|---|---|
| Byte value | 77 | 97 | 110 |
| 8 bits | 01001101 |
01100001 |
01101110 |
Joined, that is 010011010110000101101110, which splits into four 6-bit groups:
| 6-bit group | Index | Character |
|---|---|---|
010011 |
19 | T |
010110 |
22 | W |
000101 |
5 | F |
101110 |
46 | u |
So Man becomes TWFu. You can check this with Base64 Encode.
Padding and size
Three input bytes always become four characters, so the output is about one third larger than the input. When the input length is not a multiple of three, the last group is padded with = signs so the output length stays a multiple of four.
| Input | Base64 | Input bytes | Output chars |
|---|---|---|---|
Hi |
SGk= |
2 | 4 |
Hello |
SGVsbG8= |
5 | 8 |
Hello, world! |
SGVsbG8sIHdvcmxkIQ== |
13 | 20 |
a |
YQ== |
1 | 4 |
ab |
YWI= |
2 | 4 |
abc |
YWJj |
3 | 4 |
One leftover byte gives two characters and ==; two leftover bytes give three characters and =.
Variants
- URL-safe Base64 replaces
+with-and/with_(and often drops the=), because+,/and=have special meanings in URLs. The bytesfb ff feare+//+in standard Base64 and-__-in the URL-safe form. JSON Web Tokens use this variant. - MIME Base64 (email) inserts a line break every 76 characters.
- Base32 and Base16 (hex) work the same way with smaller alphabets: they are longer but friendlier to case-insensitive systems and to people typing codes by hand.
Common mistakes
- Treating Base64 as security. A password or secret that is only Base64-encoded is stored in plain sight. Anyone can paste it into Base64 Decode. Use real encryption or hashing for secrets.
- Encoding text without choosing a character set. Base64 encodes bytes, not letters. For text it matters whether those bytes are UTF-8 or something else, and a mismatch turns
éinto garbage on the other end. - Double encoding. Encoding already-encoded data inflates it by another third and decodes to unreadable text. If decoded data still looks like Base64, it was encoded twice.
- Forgetting the variant. Standard and URL-safe Base64 are not interchangeable: a string with
-or_fails in a strict standard decoder. - Breaking it with line wrapping or whitespace. Some decoders reject embedded newlines; others ignore them.
Where to try it
Encode with Base64 Encode and decode with Base64 Decode; both run entirely in your browser. For the bytes under the text, see ASCII and UTF-8.