FAQ
Q: What is the relationship between ASCII and Unicode?
ASCII was the earliest character encoding standard, covering only 128 characters — primarily English letters, digits, and common symbols. Unicode is a superset of ASCII: its first 128 code points are identical to ASCII, ensuring full backward compatibility. Unicode extends far beyond ASCII with over 140,000 characters, covering virtually every writing system in the world. In practice, UTF-8 — the most common Unicode encoding — represents ASCII characters with a single byte, just like the original standard.
Q: Why does ASCII only have 128 characters?
ASCII uses 7 bits for encoding, allowing a maximum of 2⁷ = 128 distinct characters (codes 0–127). When ASCII was designed in the 1960s, computer memory was scarce, and 7 bits were sufficient for English letters, digits, punctuation, and essential control characters. Although a byte has 8 bits, the highest bit was typically reserved as a parity bit for error detection during data transmission.
Q: What are control characters used for?
The first 32 ASCII characters (0–31) and character 127 (DEL) are control characters, originally designed to control hardware devices like teletypewriters. Several remain widely used in modern programming: LF (Line Feed, code 10), CR (Carriage Return, code 13), TAB (Horizontal Tab, code 9), and NUL (Null, code 0). Notably, the different line-ending conventions — LF on Unix/macOS versus CRLF on Windows — trace directly back to these two control characters.