globalizejs / globalizejs/globalize

Bug: Globalize number formatter is incorrect for numeric digits in supplemental plane

Open
#922 4 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
JavaScript
Stars
4.8k
Forks
585
PR merge metrics
No merged PRs in 30d

Description

Hi there

globalise (v1.7.0) number formatting is incorrect for cldr-data (v36.0.0), when cldr numeric digits are from the UTF-16 [supplemental plane](https://en.wikipedia.org/wiki/UTF-16#Code_points_from_U+010000_to_U+10FFFF) (from U+010000 to U+10FFFF).

Short example, discussed below: 44.56 formatted in ccp locale
* Should be: "𑄺𑄺.𑄻𑄼" = ["1113a", "1113a", "2e", "1113b", "1113c"] (hex codepoints)
* But returned by globalise: "��.��" = [ 'd804', 'd804', '2e', 'dd38', 'd804' ]

Based on the formatted value returned by globalise, I initially suspected that individual characters are somehow being represented in globalize as surrogate pairs (so two 16-bit hex values), but only the first of these hex values is returned. There's a worked example below, except I now have some doubts over this theory: for the 4 numeric digits involved, 3 of the digits returned by globalize seem to be the first half of a surrogate pair, but one isn't.

**Example (no code)**

For the "ccp" locale, digitals 0-9 are "𑄶𑄷𑄸𑄹𑄺𑄻𑄼𑄽𑄾𑄿", which have unicode hex codepoints of ["11136", "11137", "11138", "11139", "1113a", "1113b", "1113c", "1113d", "1113e", "1113f"].

So the number 44.56 formatted in ccp should be "𑄺𑄺.𑄻𑄼" = ["1113a", "1113a", "2e", "1113b", "1113c"]

What is actually returned from globalise is "��.��" = [ 'd804', 'd804', '2e', 'dd38', 'd804' ]

Using the [Surrogate Pair Calculator](http://www.russellcottrell.com/greek/utilities/SurrogatePairCalculator.htm) for the individual characters in "𑄺𑄺.𑄻𑄼" = ["1113a", "1113a", "2e", "1113b", "1113c"]
* 1113a = **D804** + DD3A
* 1113a = **D804** + DD3A
* 2e = **2e** (no pair needed)
* 1113b = **D804** + DD3B **(but globalise actually returns dd38)**
* 1113c = **D804** + DD3C

So maybe globalise is returning the first hex value from each surrogate pair? But dd38 is returned, not D804 (for 1113b)

**Example (code)**

```
// Output hex values for Javascript unicode characters
var asUnicodePoints = function(value) {
return Array.from(value).map(function(codePoint) {
return codePoint.codePointAt(0).toString(16);
});
};

// For us locale, works fine
var result = Globalize('us').numberFormatter()(44.56);
console.log(result);
=> 44.56
console.log(asUnicodePoints(result));
=> [ '34', '34', '2e', '35', '36' ]

// For cpp locale, wrongly returns first hex value from each surrogate pair?
var result = Globalize('ccp').numberFormatter()(44.56);
console.log(result);
=> ��.��
console.log(asUnicodePoints(result));
=> [ 'd804', 'd804', '2e', 'dd38', 'd804' ]

// For ccp locale, the true hex values for formatted 44.56 should be..
console.log(asUnicodePoints("𑄺𑄺.𑄻𑄼"));
=> [ '1113a', '1113a', '2e', '1113b', '1113c' ]
```

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.