github / github/codeql-cli-binaries

Improve `int.toUnicode()` documentation

未關閉
#80 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
CLI
主要語言
沒有語言資料
星號
1k
分支
184
PR 合併指標
30 天內沒有已合併 PR

描述

The documentation for the newly added `int.toUnicode()` predicate [says](https://codeql.github.com/codeql-standard-libraries/java/predicate.int$toUnicode.0.html):
> Returns the unicode character for the receiver seen as a unicode code point

This is slightly misleading because CodeQL strings consist of UTF-16 code points. Therefore supplementary code points (> U+FFFF) will result in two CodeQL string characters (demonstrated by [this query](https://lgtm.com/query/8363238946666626198/)). It might also be good to describe its behavior for invalid code point values. For surrogate code point it does not seem to have a result either, e.g. `55296.toUnicode()`.
Also it should uppercase "Unicode".

I would recommend the following description (or similar):
> Returns the Unicode character for the receiver seen as a Unicode code point. Because CodeQL strings consist of UTF-16 code units, supplementary code points (that is > U+FFFF) result in a CodeQL string of length 2. This predicate has no result if the int receiver does not represent a valid Unicode code point, or represents the code point of a surrogate character.

This requires changes to the built-in documentation (which is why I created the issue here) as well as the language specification.

貢獻指南

開啟貢獻指南

研究方向

Start with the built-in documentation entry for int.toUnicode() and the related language specification, both identified in the issue. Update both to clarify UTF-16 behavior, supplementary and surrogate code points, invalid values, and Unicode capitalization; done when the documentation matches the proposed semantics.

由索引模型根據 Issue 內容生成。

評估

領域
documentation
Issue 類型
文件
難度
3/5
預估耗時
1-2 天
活躍度
停滯
描述清晰度
基本清楚
新手友好度
35/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。