Text to Unicode Converter
Text to U+ code points, escapes, and entities.
This Unicode to UTF-8 converter turns code points like U+20AC, or plain text, into the UTF-8 bytes a file or network stores. It also decodes UTF-8 bytes and flags any invalid sequence.
| Character | Code point | Decimal | UTF-8 bytes | UTF-16 |
|---|
U+20AC U+00E9, or any text.UTF-8 uses one to four bytes per code point, depending on its size:
| Code point range | Bytes | Bit pattern |
|---|---|---|
| U+0000, U+007F | 1 | 0xxxxxxx |
| U+0080, U+07FF | 2 | 110xxxxx 10xxxxxx |
| U+0800, U+FFFF | 3 | 1110xxxx 10xxxxxx 10xxxxxx |
| U+10000, U+10FFFF | 4 | 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx |
The code point's bits fill the x positions from left to right.
UTF-8 stores ASCII text in exactly the same bytes as ASCII, so old files and protocols kept working. It needs no byte-order mark, and a lost byte only damages one character. Today it is used by about 98% of websites and is the default for HTML, JSON, and most programming languages.
One to four. ASCII uses one, most European letters two, most Asian characters three, and emoji four.
e2 82 ac.
The UTF-8 bytes c3 a9 were read as Latin-1, which shows each byte as a separate character. Decode the text as UTF-8 instead.
Often used together with the Unicode to UTF-8 Converter.
Text to U+ code points, escapes, and entities.
Text to hexadecimal bytes and back.
Percent-encodes and decodes text for URLs.