Unicode to UTF-8 Converter
Code points or text to UTF-8 bytes, and back.
This text to Unicode converter lists the code point of every character, such as U+20AC for €. You can also get JavaScript, CSS, and HTML escapes, or paste code points to turn them back into text.
| Character | Code point | Decimal | UTF-8 bytes | UTF-16 |
|---|
Every character in Unicode has a number, its code point, written as U+ and at least four hex digits. The converter reads text one code point at a time, so an emoji counts as one character even though JavaScript stores it as two UTF-16 units.
€ → U+20AC → \u20AC → € → €
😀 → U+1F600 → \uD83D\uDE00 (UTF-16 pair) → \u{1F600}
A code point is a character's number; an encoding is how that number is stored as bytes. U+20AC is the euro sign in every encoding, but UTF-8 stores it as three bytes (e2 82 ac) and UTF-16 as one 16-bit unit (20ac). For the bytes, use the Unicode to UTF-8 or Text to Hex converters.
The number Unicode gives a character, written as U+ followed by hex digits. A is U+0041.
Use \u{1F600} in modern JavaScript, or the surrogate pair \uD83D\uDE00 in older code and JSON.
Use &#x followed by the hex code point and a semicolon, such as € for €.
Often used together with the Text to Unicode Converter.
Code points or text to UTF-8 bytes, and back.
Text to hexadecimal bytes and back.
Encodes text into HTML entities and decodes named, decimal, and hex entities.