aboutcode-org / aboutcode-org/scancode-toolkit

Treat 'comments' and 'actual code' differently while scanning a file.

未關閉
#1,995 6 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
new feature
主要語言
Python
星號
2.6k
分支
791
平均合併
1 天 12 小時
30 天內合併 PR
5

描述

## Short Description

Examine the following code sample:

```python
def main():
for i in range(20):
print("cc", "by", "nd")

gpl = 20
print(gpl*80)
```
although, they don't really signify those licences, but scancode outputs following:
- cc-by-4.0
- GPL 2.0
- GPL 1.0 or later

It would be better if scancode separates comments/docstrings and actual code in a file before scanning *since the the licence or stuff like that are almost always found in comments/docstrings.*

For this we need to accurately detect programming language (pygments doesn't) and scan accordingly since now we know what character(s) (for that programming language) is used to add comment/docstring.

*Also, this would result in faster scans.*

Similar issues:
https://github.com/nexB/scancode-toolkit/issues/1933

## Possible Labels

- new feature

## Select Category

- Enhancement [x]
- Add License/Copyright []
- Scan Feature [x]
- Packaging []
- Documentation []
- Expand Support []
- Other []

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。