nodejs / nodejs/node

`TextEncoder.encodeInto()` underfills the destination for some non-ASCII text

Open Beginner friendly
#65,994 2 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
JavaScript
Stars
122k
Forks
37.3k
Avg merge
4d 2h
Merged PRs (30d)
283

Description

There are two problems I found with TextEncoder.encodeInto(), both of which can cause encoding to stall even when the next char can fit the destination.

  1. A 2-byte char requires a 3-byte destination

    const encoder = new TextEncoder();
    const text = '\u0400'.repeat(33);
    
    console.log(encoder.encodeInto(text, new Uint8Array(2)));  // { read: 0, written: 0 }
    console.log(encoder.encodeInto(text, new Uint8Array(3)));  // { read: 1, written: 2 }
    

    The second call proves that '\u0400' should fit into a 2-byte array.

  2. Appending an unread character changes encoding progress

    const encoder = new TextEncoder();
    const text = 'é'.repeat(33);
    
    console.log(encoder.encodeInto(text, new Uint8Array(2)));  // { read: 0, written: 0 }
    console.log(encoder.encodeInto(text + '☺', new Uint8Array(2)));  // { read: 1, written: 2 }
    

    Appending should not change whether preceding chars can be read into the buffer, but there it is.

The bugs were introduced by the encodeInto() performance change in Node.js v25.4.0. The examples above use length 33 strings to exercise that optimized path(kSmallStringThreshold = 32). Unfortunately the current encodeInto.any.js WPT tests fail to expose the problems because:

  1. all input cases use 7 or fewer code units
  2. even then, the cases don't use chars between U+0400 and U+07FF, and
  3. their cases don't contain a narrow dst capacity to reveal the signed-byte problem.

Proposed fixes

src/encoding_binding.cc

  1. Incorrect cutoff in simpleUtfEncodingLength()

    -  if (c < 0x400) return 2;
    +  if (c < 0x800) return 2;
    

    (very likely a typo, given the comment immediately below it says "Code points < 0x800: 2 bytes")

  2. Signed-byte handling in findBestFit()

    -    size_t extra = simpleUtfEncodingLength(data[pos]);
    +    size_t extra = simpleUtfEncodingLength(UTF16 ? data[pos] : static_cast<uint8_t>(data[pos]));
    

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in src/encoding_binding.cc, especially simpleUtfEncodingLength() and findBestFit(), then inspect the encodeInto.any.js WPT coverage. Reproduce the 33-character examples and add coverage for U+0400–U+07FF and narrow destinations; done means encodeInto reports the expected read/written values and the relevant WPT tests pass.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp, javascript, nodejs
Domain
api, backend
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
84/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.