Unicode Character Inspector - Code Points, UTF-8 Bytes and Escapes
Unicode Character Inspector
See exactly which characters a piece of text contains
Character by Character
Code point, category, script, UTF-8 and UTF-16 for every character, with invisible ones highlighted.
Ready-Made Escapes
Copy \u escapes for JavaScript or Python, HTML entities, CSS escapes or percent-encoding.
Four Ways to Count
See why one emoji can be 1 visible character, 2 code points, 4 UTF-16 units and 8 bytes.
Private
Everything runs in your browser; nothing you paste is uploaded.
Code Points, Bytes and What You See
Every Unicode character has a code point such as U+00E9 for é. UTF-8, the encoding of the web, stores it in 1 to 4 bytes; UTF-16, used inside JavaScript, Java and Windows, uses one or two 16-bit units. What looks like a single character can be several code points: é can be one precomposed character or e plus a combining accent, and a waving hand with a skin tone is two code points. That is why string lengths and byte limits often disagree.
Normalization makes such text comparable: NFC composes characters (the usual choice for storage), NFD decomposes them, and the K forms also fold compatibility characters such as the ligature fi into f and i. Invisible format characters (category Cf) like zero-width spaces and direction marks are highlighted, since they often cause baffling bugs in search, passwords and code.
Key Takeaways
- Length is ambiguous: Visible characters, code points, UTF-16 units and bytes differ.
- Normalize before comparing: Use NFC so é and e + accent match.
- Spot invisibles: Format and control characters are flagged in the table.