Web & Encoding
ASCII vs UTF-8: Bytes, Unicode and Compatibility
ASCII and UTF-8 agree on bytes 00–7F. The difference becomes visible when your text includes an accented letter, Cyrillic, Hindi or emoji.
Published
ASCII, Unicode and UTF-8 describe different things
ASCII is a 7-bit coded character set, usually stored in bytes with the high bit clear. UTF-8 preserves those byte values. Unicode is not a synonym for UTF-8: the same code point also has UTF-16 and UTF-32 representations.
| Name | What it describes | Example |
|---|---|---|
| ASCII | 128 character codes, including control codes | A = decimal 65 = hex 41 |
| Unicode | Code points for text from many writing systems | क = U+0915 |
| UTF-8 | A byte encoding of Unicode scalar values | क = E0 A4 95 |
The same ASCII bytes work in UTF-8
Each byte in this example is below 80 hex. Decode the sequence as ASCII or UTF-8 and you get Hello. This compatibility also applies to ASCII control codes such as line feed 0A; it does not make every byte printable.
48 65 6C 6C 6F
Hello| Text | Code point | UTF-8 bytes | ASCII? |
|---|---|---|---|
| A | U+0041 | 41 | Yes |
| é | U+00E9 | C3 A9 | No |
| Я | U+042F | D0 AF | No |
| क | U+0915 | E0 A4 95 | No |
| 😀 | U+1F600 | F0 9F 98 80 | No |
Non-ASCII text uses several bytes per code point
Here Я occupies two bytes, the space one, and क three. Reading each byte as an independent character loses the UTF-8 structure. Switch the example below between ASCII and UTF-8 to see the difference.
One visible character can contain several Unicode code points. Hindi syllables and emoji sequences can therefore use more than four bytes even though each individual Unicode scalar value uses at most four UTF-8 bytes.
D0 AF 20 E0 A4 95
Я कA non-ASCII byte sequence is not automatically UTF-8
C3 requires a continuation byte in the range 80–BF. The next byte, 28, is an opening parenthesis and cannot continue it. The UTF-8 view reports an invalid sequence; a replacement character is an error indicator, not recovered original text.
Check the source encoding before converting. A lone E9 may represent é in a legacy encoding, but E9 alone is not a complete UTF-8 sequence. Trying random encodings can produce plausible-looking text without proving it is correct.
C3 28Encode text and decode bytes explicitly in Python
Encoding turns text into bytes; decoding interprets bytes as text. Choose the encoding that the file or protocol actually uses. Ignoring decoding errors can discard data, so keep strict decoding while diagnosing a problem.
text = "Я क"
raw = text.encode("utf-8")
print(raw.hex(" ")) # d0 af 20 e0 a4 95
print(bytes.fromhex("d0 af 20 e0 a4 95").decode("utf-8"))
# bytes.fromhex("c3 28").decode("utf-8") raises UnicodeDecodeErrorExtended ASCII and UTF-16 are separate cases
Extended ASCII is an informal label for several incompatible 8-bit encodings, not one universal mapping for 80–FF. UTF-8 compatibility covers standard ASCII only.
UTF-16 also represents Unicode, but its bytes differ: A is 41 00 in UTF-16LE and 00 41 in UTF-16BE. A BOM, file metadata or protocol declaration can help identify an encoding; a successful decode alone does not identify the original format.
Check a byte sequence in three steps
- Start with the original bytes and the source's declared encoding.
- Use the ASCII view to locate printable characters and control or non-ASCII bytes.
- Select UTF-8, check validity, and compare the decoded result with the expected text before copying it.
Frequently asked questions
Is every ASCII file valid UTF-8?
A sequence consisting only of standard ASCII bytes 00–7F is valid UTF-8 and represents the same characters. That says nothing about the syntax of the file format.
Can ASCII represent Hindi or Russian?
Standard ASCII does not include Cyrillic or Devanagari letters. Use a suitable Unicode encoding such as UTF-8. A font that draws a different glyph over an ASCII code does not change the stored character.
Does UTF-8 always use more space than ASCII?
For ASCII-only content the byte count is identical. Other characters use multiple UTF-8 bytes; byte length and visible character count are different measurements.
Try the example
Start with ASCII-compatible bytes
Decode Hello, then switch between ASCII and UTF-8.
48 65 6C 6C 6FExpected result: Both views display Hello; the input contains five bytes.
Try the example
Decode Cyrillic and Devanagari
A two-byte letter, a space and a three-byte letter share one UTF-8 sequence.
D0 AF 20 E0 A4 95Expected result: UTF-8 displays Я क. The six bytes are not six independent characters.
Try the example
Inspect an invalid UTF-8 sequence
The byte 28 cannot continue the sequence started by C3.
C3 28Expected result: The UTF-8 view reports invalid input. A replacement character does not recover the missing original value.
Inspect your bytes
Compare ASCII and UTF-8 output
Paste hexadecimal bytes, inspect the ASCII view, then switch to UTF-8 and check whether the sequence is valid.