Unicode to UTF-8 Converter

This Unicode to UTF-8 converter turns code points like U+20AC, or plain text, into the UTF-8 bytes a file or network stores. It also decodes UTF-8 bytes and flags any invalid sequence.

Updated
Runs in your browser. Your text is not uploaded.
Direction
Output

        

How to use the Unicode to UTF-8 Converter

  1. Choose Text or code points → UTF-8 bytes, or the reverse.
  2. Type code points such as U+20AC U+00E9, or any text.
  3. Read the bytes and the character table.

How it works

UTF-8 uses one to four bytes per code point, depending on its size:

Code point rangeBytesBit pattern
U+0000, U+007F10xxxxxxx
U+0080, U+07FF2110xxxxx 10xxxxxx
U+0800, U+FFFF31110xxxx 10xxxxxx 10xxxxxx
U+10000, U+10FFFF411110xxx 10xxxxxx 10xxxxxx 10xxxxxx

The code point's bits fill the x positions from left to right.

Examples

  • U+20AC (€) = e2 82 ac.
  • U+00E9 (é) = c3 a9.
  • U+1F600 (😀) = f0 9f 98 80.
  • U+0041 (A) = 41.

Why UTF-8 won

UTF-8 stores ASCII text in exactly the same bytes as ASCII, so old files and protocols kept working. It needs no byte-order mark, and a lost byte only damages one character. Today it is used by about 98% of websites and is the default for HTML, JSON, and most programming languages.

Limitations

  • Bytes must be valid UTF-8; overlong forms and lone surrogates are rejected.
  • Other encodings (Latin-1, UTF-16) are not decoded.

Frequently asked questions

How many bytes does a UTF-8 character use?

One to four. ASCII uses one, most European letters two, most Asian characters three, and emoji four.

What is the UTF-8 encoding of €?

e2 82 ac.

Why do I see é instead of é?

The UTF-8 bytes c3 a9 were read as Latin-1, which shows each byte as a separate character. Decode the text as UTF-8 instead.

Often used together with the Unicode to UTF-8 Converter.