SSerguey Asael Shinder
Java coding notes: the JVM, and writing software that lasts

Serguey Asael Shinder: trim() and strip() disagree on what is space, and both keep U+00A0

· by Serguey Asael Shinder / Serguey Shinder

Java has two ways to remove space from both ends of a string, and they use two different definitions of space. Neither of them covers the character that most often causes trouble in pasted input.

trim() works on code values. The javadoc defines space as "any character whose codepoint is less than or equal to 'U+0020'". That includes the ordinary space, tab and line breaks, but also every control character below it, \u0000 among them. It does not include anything above U+0020, so a Unicode space such as U+3000 (the ideographic space) or U+2003 (em space) is left in place.

strip() works on Unicode. Added in Java 11, its javadoc says it removes "white space", and links to Character.isWhitespace(int). That method counts Unicode space, line and paragraph separators, plus tab, line feed, vertical tab, form feed, carriage return and U+001C to U+001F. So strip() removes U+3000 and U+2003, which trim() keeps. But it does not remove NUL or the other low control characters outside that list, which trim() does remove.

" id ".trim();    // " id " - above U+0020, left alone
" id ".strip();   // "id"
"\u0000id\u0000".trim();    // "id" - NUL is at or below U+0020
"\u0000id\u0000".strip();   // "\u0000id\u0000" - NUL is not Java whitespace
" id ".trim();    // " id "
" id ".strip();   // " id "
Serguey Asael Shinder: trim() and strip() disagree on what is space, and both keep U+00A0
trim() and strip() disagree on what is space, and both keep U+00A0 — Serguey Asael Shinder

The case neither handles. Character.isWhitespace is explicit that a Unicode space character counts only if it "is not also a non-breaking space" and names three: U+00A0, U+2007 and U+202F. U+00A0 is what HTML's   becomes, and it arrives constantly in values copied from web pages, spreadsheets and word processors. Neither method touches it. isBlank() follows the same definition as strip(), so a string made only of non-breaking spaces is not blank either.

What to write instead. Decide what "space" means for the field and say it in code. For identifiers and codes, a strict check is often better than trimming: reject anything outside the allowed characters. If you do want to remove non-breaking spaces too, name them:

String cleaned = input.replaceAll("^[\\s\\u00A0\\u2007\\u202F]+|[\\s\\u00A0\\u2007\\u202F]+$", "");

In a Java regex, \s without flags matches only [ \t\n\x0B\f\r], which is why the non-breaking spaces are listed separately.

The rule I take from it. trim() is about ASCII control and space, strip() is about Unicode whitespace, and neither is about what a user sees as blank. When input comes from outside, pick the definition on purpose.

No JDK runs on the machine this note was written on; the code is written to compile against Java 21 and the behaviour described is the one stated in the linked javadoc.