Skip to content
Hex Calculator

UTF-16 to Hex Converter

Encode text as UTF-16 bytes in hex (big-endian, 2 bytes per code unit) β€” the format JavaScript strings use internally.

UTF-16 hex result
0x004800690021
Swapnil Sanghvi

Built by

Swapnil Sanghvi

Full-Stack Web Developer & WordPress Developer

Swapnil Sanghvi is a full-stack web and WordPress developer, UI designer, and full-time freelancer who builds and maintains Hex Calculator.

UTF-16 to Hex Converter tool preview card from Hex Calculator
Share preview: this is the card that appears when you share the UTF-16 to Hex Converter page on social media.

Further reading: Character encoding β€” Wikipedia

How to Convert Text to UTF-16 Hex

Every character becomes one 16-bit code unit (2 bytes) β€” or, for characters above U+FFFF like most emoji, a pair of code units (4 bytes total). Each unit is written as 4 hex digits, most significant byte first, concatenated in order.

Where UTF-16 Hex Actually Comes Up

Debugging a JavaScript String's Internals

JavaScript strings are UTF-16 under the hood β€” seeing the raw hex helps explain surprising behavior like .length counting an emoji as 2, or .charCodeAt() returning half of a surrogate pair.

πŸ˜€ -> D83D DE00

Reading Windows API String Data

Windows APIs that take wide strings (WCHAR, LPWSTR) use UTF-16 natively β€” converting a known string to hex gives you the exact bytes to expect when inspecting memory or a captured buffer.

Hi! -> 0048 0069 0021

Comparing Java char Array Contents

Java's char type is a 16-bit UTF-16 code unit β€” encoding a string to UTF-16 hex here lets you verify a char array's contents match what the source text should produce.

A -> 0041

Common Mistakes When Converting Text to UTF-16 Hex

  • Expecting the same byte count as UTF-8 β€” UTF-16 is wider for plain ASCII text.
  • Forgetting emoji and other high code points need a 4-byte surrogate pair, not 2 bytes.
  • Assuming byte order matches little-endian systems without checking β€” this output is big-endian.

Why Use This Instead of Doing It by Hand

  • Handles surrogate pairs for emoji and other high code points automatically
  • Runs entirely in your browser β€” nothing you type gets sent anywhere
  • Saves you from manually splitting UTF-16 code units into byte pairs
  • Matches the exact internal representation JavaScript, Java, and Windows APIs use

Limitations

  • Encodes as big-endian, 2 bytes per code unit.

Frequently Asked Questions

How do you convert text to UTF-16 hex?

Each character becomes one or two 16-bit code units (2 bytes each) β€” most characters need one, characters above U+FFFF like emoji need a surrogate pair of two β€” written as hex, most significant byte first.

Why does this output more hex digits than UTF-8 for the same text?

UTF-16's baseline is 2 bytes per code unit even for plain ASCII characters, while UTF-8 uses just 1 byte for those same characters β€” English text roughly doubles in size going from UTF-8 to UTF-16.

What does this converter use for emoji?

A surrogate pair β€” two 2-byte code units (4 bytes total) that together represent the one character, following the same rule JavaScript uses internally for strings containing emoji.

Is the output big-endian or little-endian?

Big-endian β€” each code unit's high byte comes first. Some systems (notably Windows internals) use little-endian UTF-16 instead; swap each byte pair if you need that format.

Why would I need UTF-16 hex specifically?

Debugging raw string data from JavaScript, Java, or Windows APIs, all of which use UTF-16 internally β€” seeing the exact bytes helps track down encoding mismatches.

Is this the reverse of Hex to UTF-16?

Yes β€” that page decodes UTF-16 hex bytes back into readable text.

Is this the same as a "UTF-16LE converter"?

Related but not identical β€” UTF-16LE is the little-endian variant, with each code unit's bytes swapped compared to this converter's big-endian output. Reverse each 2-byte pair if you specifically need UTF-16LE.

Why does JavaScript's string.length lie for emoji?

Because JavaScript strings are UTF-16 internally, and .length counts 16-bit code units, not characters β€” an emoji needing a surrogate pair counts as 2 toward .length even though it's visually one character.