Unicode Character Inspector - Code Points, UTF-8 Bytes and Escapes

AdSense Placeholder
Slot: header_tool

Unicode Character Inspector

See exactly which characters a piece of text contains

Code points
0
Visible characters
0
UTF-16 units (JS length)
0
UTF-8 bytes
0
CharCode pointCategoryScriptUTF-8UTF-16Escape

Only the first 500 characters are listed.

Normalization forms
NFC
--
NFD
--
NFKC
--
NFKD
--

AdSense Placeholder
Slot: tool_mid_article

Character by Character

Code point, category, script, UTF-8 and UTF-16 for every character, with invisible ones highlighted.

Ready-Made Escapes

Copy \u escapes for JavaScript or Python, HTML entities, CSS escapes or percent-encoding.

Four Ways to Count

See why one emoji can be 1 visible character, 2 code points, 4 UTF-16 units and 8 bytes.

Private

Everything runs in your browser; nothing you paste is uploaded.

Code Points, Bytes and What You See

Every Unicode character has a code point such as U+00E9 for é. UTF-8, the encoding of the web, stores it in 1 to 4 bytes; UTF-16, used inside JavaScript, Java and Windows, uses one or two 16-bit units. What looks like a single character can be several code points: é can be one precomposed character or e plus a combining accent, and a waving hand with a skin tone is two code points. That is why string lengths and byte limits often disagree.

Normalization makes such text comparable: NFC composes characters (the usual choice for storage), NFD decomposes them, and the K forms also fold compatibility characters such as the ligature fi into f and i. Invisible format characters (category Cf) like zero-width spaces and direction marks are highlighted, since they often cause baffling bugs in search, passwords and code.

Key Takeaways

  • Length is ambiguous: Visible characters, code points, UTF-16 units and bytes differ.
  • Normalize before comparing: Use NFC so é and e + accent match.
  • Spot invisibles: Format and control characters are flagged in the table.
AdSense Placeholder
Slot: footer_leaderboard