aboutcode-org / aboutcode-org/scancode-toolkit

Explain when and where comma, whitespace, etc. are ignored

Open
#4,843 1 comment 0 reactions 1 assignee Claimed by @AyanSinhaMahapatra View on GitHub
documentation
Dominant language
Python
Stars
2.6k
Forks
791
Avg merge
1d 12h
Merged PRs (30d)
5

Description

### Description

> A brief description of the Documentation Improvement or New Section request.

I am currently scanning a file with the following license reference:

```
Licensed under GPLv2, see file LICENSE in this source tree.
```

scancode reports a match for rule `gpl-2.0_147.RULE` which has the following:

```
Licensed under GPLv2.
```

The file I am scanning has a comma, while the rule has a dot, yet it matches. This makes me wonder if scancode first preprocesses data and ignores commas, whitespace, and others, and if so: when (at the end of a sentence? mid-sentence?). This could possibly be relevant when writing license rules.

### Link to Documentation Page

> Where the confusion/inconsistency/incomplete documentation is.

### Select Category

- [ ] Inconsistency
- [ ] New Section Request
- [ ] General Improvement
- [ ] Typo/Mistakes
- [X] Other

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.