Glossary Data & Encoding

UTF-8

UTF-8 is the dominant Unicode encoding, writing each character as one to four bytes and staying identical to ASCII for the first 128 characters.

Last reviewed:

AdSense Placeholder
Slot: header_reference_page

In more detail

Because it is variable-length, a string's byte count is not its character count: é takes two bytes and most emoji take four. It is the default for the web, JSON and most modern software.

Example

A is 41, é is C3 A9 and € is E2 82 AC. See every character's bytes with Unicode Character Inspector.

Try these tools

See also

  • Glossary ASCII
    ASCII is a 7-bit character encoding that assigns numbers 0 to 127 to English letters, digits, punctuation and control codes.
AdSense Placeholder
Slot: footer_leaderboard