SSerguey Asael Shinder
Java coding notes: the JVM, and writing software that lasts

Serguey Asael Shinder: String.length() counts UTF-16 code units, so an emoji is two characters long

· by Serguey Asael Shinder / Serguey Shinder

This line looks like it enforces a limit of 20 characters:

if (name.length() > 20) {
    name = name.substring(0, 20);
}

It enforces a limit of 20 char values, which is not the same thing, and the difference is documented in the first paragraphs of the String javadoc: a String "represents a string in the UTF-16 format in which supplementary characters are represented by surrogate pairs". And: "Index values refer to char code units, so a supplementary character uses two positions in a String."

length() says the same in one sentence: "The length is equal to the number of Unicode code units in the string." Every character above U+FFFF, which includes most emoji and a number of CJK ideographs, counts twice.

String s = "😀";                           // U+1F600, one emoji
int units = s.length();                              // 2
int chars = s.codePointCount(0, s.length());         // 1

Where it goes wrong. Three places, in my experience.

Serguey Asael Shinder: String.length() counts UTF-16 code units, so an emoji is two characters long
String.length() counts UTF-16 code units, so an emoji is two characters long — Serguey Asael Shinder

Limits. A user types 20 emoji into a field with a 20-character limit and is told it is too long, or a limit enforced in Java disagrees with the same limit enforced by a database that counts code points.

Truncation. substring(0, 20) can cut between the two halves of a surrogate pair. The result is a String ending in an unpaired high surrogate, which is not valid UTF-16. Encoders will typically replace it with ? or U+FFFD, so the damage shows up far from the line that caused it.

Positions. A column number computed from char offsets drifts by one for every supplementary character earlier on the line. This is exactly what Checkstyle's issue #10924 is about: checks that measure spacing with getLine() should count code points instead of characters.

What I use instead. When the rule is about characters, I count them as code points: codePointCount(int, int) "returns the number of Unicode code points in the specified text range", and codePoints() gives a stream in which surrogate pairs are combined. To truncate, I move by code points with offsetByCodePoints(int, int) and cut at the index it returns, so a pair is never split.

int limit = 20;
if (name.codePointCount(0, name.length()) > limit) {
    name = name.substring(0, name.offsetByCodePoints(0, limit));
}

Code points are still not what a reader perceives as one character: a flag or an accented letter written with a combining mark is several code points. If the limit is about what is displayed, the unit is the grapheme cluster, and in the JDK that means BreakIterator.getCharacterInstance(). Most limits in the systems I work on are about storage, though, and there the right question is what the other side counts. Ask it before writing length().

The code in this note is written to compile as shown; the behaviour described is the one stated in the linked javadoc for Java SE 25.