跳到主要内容

Unicode to UTF-16

Quick Try:

Conversion Process

Unicode
Binary
Process
UTF-16

UTF-16 Explained

UTF-16 (16-bit Unicode Transformation Format) is a variable-length character encoding for Unicode. See UTF-16 on Wikipedia.

Unicode's encoding space is divided into 17 planes, each containing 65,536 code points. The first plane is called the Basic Multilingual Plane (BMP, Plane 0); the remaining planes are called supplementary planes.

BMP vs Supplementary Planes

PlaneNumber of bytesFirst code pointLast code point
Basic Multilingual Plane (BMP)2U+0000U+FFFF
Supplementary Planes4U+10000U+10FFFF

The BMP contains most commonly used characters (including the majority of CJK characters), while the supplementary planes contain emoji and other less common characters.

Encoding Process

  1. Convert the character to its Unicode code point (e.g., str.codePointAt(0))
  2. If the code point is less than U+10000, it maps directly to a single UTF-16 code unit
  3. If the code point is U+10000 or greater, it is split into a surrogate pair: (codePoint - 0x10000) / 1024 + 0xD800 (high surrogate) and (codePoint - 0x10000) % 1024 + 0xDC00 (low surrogate)

For example, "🍉" has code point 0x1F349. Subtract 0x10000 to get 0x0F349. The high surrogate is 0xD83C and the low surrogate is 0xDF49.