Web & Encoding

ASCII vs UTF-8: Bytes, Unicode and Compatibility

ASCII and UTF-8 agree on bytes 00–7F. The difference becomes visible when your text includes an accented letter, Cyrillic, Hindi or emoji.

Published

ASCII, Unicode and UTF-8 describe different things

ASCII is a 7-bit coded character set, usually stored in bytes with the high bit clear. UTF-8 preserves those byte values. Unicode is not a synonym for UTF-8: the same code point also has UTF-16 and UTF-32 representations.

NameWhat it describesExample
ASCII128 character codes, including control codesA = decimal 65 = hex 41
UnicodeCode points for text from many writing systemsक = U+0915
UTF-8A byte encoding of Unicode scalar valuesक = E0 A4 95

The same ASCII bytes work in UTF-8

Each byte in this example is below 80 hex. Decode the sequence as ASCII or UTF-8 and you get Hello. This compatibility also applies to ASCII control codes such as line feed 0A; it does not make every byte printable.

48 65 6C 6C 6F
Hello
TextCode pointUTF-8 bytesASCII?
AU+004141Yes
éU+00E9C3 A9No
ЯU+042FD0 AFNo
कU+0915E0 A4 95No
😀U+1F600F0 9F 98 80No

Non-ASCII text uses several bytes per code point

Here Я occupies two bytes, the space one, and क three. Reading each byte as an independent character loses the UTF-8 structure. Switch the example below between ASCII and UTF-8 to see the difference.

One visible character can contain several Unicode code points. Hindi syllables and emoji sequences can therefore use more than four bytes even though each individual Unicode scalar value uses at most four UTF-8 bytes.

D0 AF 20 E0 A4 95
Я क

A non-ASCII byte sequence is not automatically UTF-8

C3 requires a continuation byte in the range 80–BF. The next byte, 28, is an opening parenthesis and cannot continue it. The UTF-8 view reports an invalid sequence; a replacement character is an error indicator, not recovered original text.

Check the source encoding before converting. A lone E9 may represent é in a legacy encoding, but E9 alone is not a complete UTF-8 sequence. Trying random encodings can produce plausible-looking text without proving it is correct.

C3 28

Encode text and decode bytes explicitly in Python

Encoding turns text into bytes; decoding interprets bytes as text. Choose the encoding that the file or protocol actually uses. Ignoring decoding errors can discard data, so keep strict decoding while diagnosing a problem.

text = "Я क"
raw = text.encode("utf-8")
print(raw.hex(" "))  # d0 af 20 e0 a4 95
print(bytes.fromhex("d0 af 20 e0 a4 95").decode("utf-8"))
# bytes.fromhex("c3 28").decode("utf-8") raises UnicodeDecodeError

Extended ASCII and UTF-16 are separate cases

Extended ASCII is an informal label for several incompatible 8-bit encodings, not one universal mapping for 80–FF. UTF-8 compatibility covers standard ASCII only.

UTF-16 also represents Unicode, but its bytes differ: A is 41 00 in UTF-16LE and 00 41 in UTF-16BE. A BOM, file metadata or protocol declaration can help identify an encoding; a successful decode alone does not identify the original format.

Check a byte sequence in three steps

  • Start with the original bytes and the source's declared encoding.
  • Use the ASCII view to locate printable characters and control or non-ASCII bytes.
  • Select UTF-8, check validity, and compare the decoded result with the expected text before copying it.

Frequently asked questions

Is every ASCII file valid UTF-8?

A sequence consisting only of standard ASCII bytes 00–7F is valid UTF-8 and represents the same characters. That says nothing about the syntax of the file format.

Can ASCII represent Hindi or Russian?

Standard ASCII does not include Cyrillic or Devanagari letters. Use a suitable Unicode encoding such as UTF-8. A font that draws a different glyph over an ASCII code does not change the stored character.

Does UTF-8 always use more space than ASCII?

For ASCII-only content the byte count is identical. Other characters use multiple UTF-8 bytes; byte length and visible character count are different measurements.

Try the example

Start with ASCII-compatible bytes

Decode Hello, then switch between ASCII and UTF-8.

48 65 6C 6C 6F

Expected result: Both views display Hello; the input contains five bytes.

Try the example

Decode Cyrillic and Devanagari

A two-byte letter, a space and a three-byte letter share one UTF-8 sequence.

D0 AF 20 E0 A4 95

Expected result: UTF-8 displays Я क. The six bytes are not six independent characters.

Try the example

Inspect an invalid UTF-8 sequence

The byte 28 cannot continue the sequence started by C3.

C3 28

Expected result: The UTF-8 view reports invalid input. A replacement character does not recover the missing original value.

Inspect your bytes

Compare ASCII and UTF-8 output

Paste hexadecimal bytes, inspect the ASCII view, then switch to UTF-8 and check whether the sequence is valid.

Open Hex to ASCII Converter →

Documentation and standards