Glossary Data & Encoding

Unicode

Unicode is the universal standard that gives every character in the world's writing systems, plus symbols and emoji, a unique number called a code point.

Last reviewed:

AdSense Placeholder
Slot: header_reference_page

In more detail

Code points are written U+0041 for A or U+1F600 for a smiling face. Unicode defines the characters; encodings such as UTF-8 define how code points are stored as bytes. Look-alike and invisible characters can be abused, for example to disguise a domain name.

Try these tools

See also

  • Glossary UTF-8
    UTF-8 is the dominant Unicode encoding, writing each character as one to four bytes and staying identical to ASCII for the first 128.
  • Glossary ASCII
    ASCII is a 7-bit character encoding that assigns numbers 0 to 127 to English letters, digits, punctuation and control codes.
AdSense Placeholder
Slot: footer_leaderboard