github / github/codeql

False positive: js/incomplete-hostname-regexp treats LinkifyIt.match(text) as a regex call

Open
#22,546 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
CodeQL
Stars
10.1k
Forks
2.1k
Avg merge
2d 15h
Merged PRs (30d)
141

Description

**Description of the false positive**

`js/incomplete-hostname-regexp` treats the argument to `LinkifyIt.match(text)` as a regular expression. The receiver is a `LinkifyIt` instance from `linkify-it@6.1.0`; its argument is document text to scan for links, not a regex pattern. Literal dots in these URLs are therefore correct.

Observed with CodeQL **2.26.4**, JavaScript/TypeScript analysis, `build-mode: none`, and `github/codeql-action@cdf488f595d80d6e07e03d4674febd5ab45fa938` (v4.37.9), using the default query suite without custom queries or exclusions.

Diagnostic on the text literal:

> This string, which is used as a regular expression here, has an unescaped '.' before 'youtube.com/watc', so it might match more hosts than expected.

The linked use is `scanner.match(text)`.

**Code samples or links to source code**

Small excerpt retaining the constructor, configuration, literal and call from the reported test:

```js
import { LinkifyIt } from "linkify-it";

const scanner = new LinkifyIt({ fuzzyLink: false, fuzzyEmail: false })
.add("ftp:", null)
.add("mailto:", null)
.add("//", null);
const text =
"😀 *literal* (https://www.youtube.com/watch?v=tax4e4hBBZc), then https://store.steampowered.com/app/457140/.";
const matches = scanner.match(text);
console.log(matches.map((m) => m.raw));
```

With `linkify-it@6.1.0`, the API returns matches for the two literal URLs. The package's `build/index.d.ts` declares `match(text: string): Match[] | null`; `build/index.mjs` implements the method by scanning that document text with the library's link recognizers.

[Exact reported source at the PR merge revision](https://github.com/MaksymShostak/steam-community-bbcode/blob/34d9cea0dccc2fbf8fae35558333b23b5e049758/test/link-recognition-qualification.test.js#L4-L13).

[Source at the immutable PR head](https://github.com/MaksymShostak/steam-community-bbcode/blob/c1c5de6ec6685b5f2c36aac24470dfb41634f961/test/link-recognition-qualification.test.js#L4-L13).

Validation: `node --test test/link-recognition-qualification.test.js` passes 1/1, including exact URL and offset assertions. The hosted CodeQL analysis reports the finding in the linked full test. The reduced excerpt above has not been separately analyzed with CodeQL; no claim is made that it is the smallest scanner reproducer.

Expected: recognize that this imported library's `match` method consumes document text, so these literals should not be classified as hostname regular expressions. Ordinary `String.match` regex findings should remain enabled.

Related precedent: https://github.com/github/codeql/pull/19854 explicitly models Sinon's `match` calls as non-RegExp. I found no existing `linkify-it` report in the upstream search.

**URL to the alert on GitHub code scanning (optional)**

https://github.com/MaksymShostak/steam-community-bbcode/security/code-scanning/2

[Failing PR check](https://github.com/MaksymShostak/steam-community-bbcode/pull/4/checks?check_run_id=103162822542).

Contributor guide

Open the contributing guide

Research direction

Inspect the js/incomplete-hostname-regexp query and its library models, then compare the Sinon precedent in PR 19854. Re-run the linked CodeQL alert against test/link-recognition-qualification.test.js; done when LinkifyIt.match(text) literals are no longer reported while ordinary String.match regex findings remain enabled.

Written by the indexing model from the issue text.

Assessment

Tech stack
javascript
Domain
security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
55/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.