aboutcode-org / aboutcode-org/scancode-toolkit
Treat 'comments' and 'actual code' differently while scanning a file.
- 主要语言
- Python
- 星标
- 2.6k
- 派生
- 791
- 平均合并
- 1 天 12 小时
- 30 天内合并 PR
- 5
描述
## Short Description
Examine the following code sample:
```python
def main():
for i in range(20):
print("cc", "by", "nd")
gpl = 20
print(gpl*80)
```
although, they don't really signify those licences, but scancode outputs following:
- cc-by-4.0
- GPL 2.0
- GPL 1.0 or later
It would be better if scancode separates comments/docstrings and actual code in a file before scanning *since the the licence or stuff like that are almost always found in comments/docstrings.*
For this we need to accurately detect programming language (pygments doesn't) and scan accordingly since now we know what character(s) (for that programming language) is used to add comment/docstring.
*Also, this would result in faster scans.*
Similar issues:
https://github.com/nexB/scancode-toolkit/issues/1933
## Possible Labels
- new feature
## Select Category
- Enhancement [x]
- Add License/Copyright []
- Scan Feature [x]
- Packaging []
- Documentation []
- Expand Support []
- Other []
贡献指南
评估
这个 Issue 还没有评估数据。