Devolutions / Devolutions/IronRDP

Fix UTF-16 string length computations

Open
#130 1 comment 0 reactions 0 assignees View on GitHub
kind/technical-debt scope/core
Dominant language
Rust
Stars
3.2k
Forks
275
Avg merge
1d 11h
Merged PRs (30d)
189

Description

Some code in IronRDP is incorrectly computing the size of UTF-16 strings.

Example:
```rust
// This is not the right way to compute the number of bytes for unicode strings encoded in UTF-16.
// This is a time bomb: it will returns the correct result some times (e.g.: when the string is valid ASCII),
// but not always.
fn utf16_len(utf8_str: &str) -> usize {
utf8_str.len() * 2
}
```

Both UTF-8 and UTF-16 are using a variable-length encoding and code points may be encoded using multiple code units. The thing is, UTF-16 uses one or two 16-bit code units and UTF-8 uses between one and four 8-bit code units. It’s really not always the case that a code point in UTF-16 is twice as big as the same code point in UTF-8.

This kind of erroneous code is present at multiple places. One such instance is [`ironrdp_pdu::rdp::client_info::string_len`](https://github.com/Devolutions/IronRDP/blob/8857c3c25ea40bd2a2f4d850f67b16b2306393c9/crates/ironrdp-pdu/src/rdp/client_info.rs#L572).

Instead, something like that must be used:
```rust
utf8_str.encode_utf16().count() * 2 // add 2 if we need to account for a null terminator (0x0000)
```

Refer to `ironrdp_pdu::pcb` module for a correct implementation.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.