aboutcode-org / aboutcode-org/scancode-toolkit
Explain when and where comma, whitespace, etc. are ignored
- Dominant language
- Python
- Stars
- 2.6k
- Forks
- 791
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 5
Description
### Description
> A brief description of the Documentation Improvement or New Section request.
I am currently scanning a file with the following license reference:
```
Licensed under GPLv2, see file LICENSE in this source tree.
```
scancode reports a match for rule `gpl-2.0_147.RULE` which has the following:
```
Licensed under GPLv2.
```
The file I am scanning has a comma, while the rule has a dot, yet it matches. This makes me wonder if scancode first preprocesses data and ignores commas, whitespace, and others, and if so: when (at the end of a sentence? mid-sentence?). This could possibly be relevant when writing license rules.
### Link to Documentation Page
> Where the confusion/inconsistency/incomplete documentation is.
### Select Category
- [ ] Inconsistency
- [ ] New Section Request
- [ ] General Improvement
- [ ] Typo/Mistakes
- [X] Other
Contributor guide
Assessment
This issue has not been assessed yet.