UTF-16 Explained
UTF-16 (16-bit Unicode Transformation Format) is a variable-length character encoding for Unicode. See UTF-16 on Wikipedia.
Unicode's encoding space is divided into 17 planes, each containing 65,536 code points. The first plane is called the Basic Multilingual Plane (BMP, Plane 0); the remaining planes are called supplementary planes.
BMP vs Supplementary Planes
| Plane | Number of bytes | First code point | Last code point |
|---|---|---|---|
| Basic Multilingual Plane (BMP) | 2 | U+0000 | U+FFFF |
| Supplementary Planes | 4 | U+10000 | U+10FFFF |
The BMP contains most commonly used characters (including the majority of CJK characters), while the supplementary planes contain emoji and other less common characters.
Encoding Process
- Convert the character to its Unicode code point (e.g.,
str.codePointAt(0)) - If the code point is less than U+10000, it maps directly to a single UTF-16 code unit
- If the code point is U+10000 or greater, it is split into a surrogate pair:
(codePoint - 0x10000) / 1024 + 0xD800(high surrogate) and(codePoint - 0x10000) % 1024 + 0xDC00(low surrogate)
For example, "🍉" has code point 0x1F349. Subtract 0x10000 to get 0x0F349. The high surrogate is 0xD83C and the low surrogate is 0xDF49.