SSerguey Asael Shinder
Java coding notes: the JVM, and writing software that lasts

Serguey Asael Shinder: Integer.parseInt accepts Arabic-Indic digits; the regex digit class does not

· by Serguey Asael Shinder / Serguey Shinder

A validation rule I have seen in several code bases looks like this: check the input against a pattern of digits, then parse it.

if (input.matches("\\d{1,9}")) {
    int n = Integer.parseInt(input);
}

The two lines use two different definitions of "digit", and they only agree on ASCII.

Serguey Asael Shinder: Integer.parseInt accepts Arabic-Indic digits; the regex digit class does not
Integer.parseInt accepts Arabic-Indic digits; the regex digit class does not — Serguey Asael Shinder

What parseInt accepts. The javadoc of Integer.parseInt(String, int) says the characters must all be digits of the radix "as determined by whether Character.digit(char, int) returns a nonnegative value". The one-argument version delegates to it with radix 10. And Character.digit treats a character as a valid digit when isDigit is true of it and its Unicode decimal value is below the radix. Character.isDigit is true for every character in the Unicode category DECIMAL_DIGIT_NUMBER, and its own javadoc lists some of them: Arabic-Indic digits ٠ to ٩, Extended Arabic-Indic ۰ to ۹, Devanagari ० to ९, fullwidth 0 to 9, and adds that many other ranges contain digits as well.

So by the documented contract, Integer.parseInt("١٢٣") returns 123, and so does the same number typed with fullwidth digits.

What \d accepts. The Pattern javadoc defines \d as "A digit: [0-9] if UNICODE_CHARACTER_CLASS is not set". Without that flag, the pattern above rejects the Arabic-Indic string that parseInt would happily accept.

In the snippet above the regex runs first, so nothing bad happens: non-ASCII digits are refused before the parser sees them. The trouble starts when the order changes or the check disappears. Code that relies on parseInt alone as its validation accepts inputs that the rest of the system, a database constraint, a downstream service, a log search, may treat as different strings. Two IDs that parse to the same number but are stored as different text are a classic source of duplicate records.

What I do instead. I decide which alphabet the field is allowed to use and enforce it in one place. For identifiers and anything that crosses a system boundary that is almost always ASCII, so I check for it explicitly before parsing, with a character loop over '0' to '9' or with the regex, and I keep the check and the parse next to each other. Where users are allowed to type numbers in their own script, I parse deliberately with Character.digit and store the normalised ASCII form, so that the value the system keeps has one spelling.

The code in this note is written to compile as shown; the behaviour described is the one stated in the linked javadoc for Java SE 25.