aboutcode-org / aboutcode-org/scancode-toolkit

Extend copyright output with raw data option

未关闭
#3,743 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
new feature
主要语言
Python
星标
2.6k
派生
791
平均合并
1 天 12 小时
30 天内合并 PR
5

描述

## Short Description

As discussed in the ORT Community Meeting (https://github.com/oss-review-toolkit/ort/wiki/ORT-Community-Meeting#2024-04-11) we propose to add a raw data export option to ScanCode so that post-processing of copyrights becomes possible on the raw and not normalized data which ScanCode usually does.

## Possible Labels

- new feature
- copyright scan

## Select Category

- [x] Enhancement
- [ ] Add License/Copyright
- [ ] Scan Feature
- [ ] Packaging
- [ ] Documentation
- [ ] Expand Support
- [ ] Other

## **Describe the Update**

ScanCode should be able to export raw data on demand. Either as a second field as an addition to the existing copyright or as a configuration option to not normalize copyrights and provide the raw data in the existing result field.

## **How This Feature will help you/your organization**

We have a strong requirement from our legal department that copyrights need to be correct and complete. Some of the normalizations of ScanCode currently omit information which might be required. Therefore raw data would help us to do manual and automated post-processing of copyright statements.

## **Possible Solution/Implementation Details**

@pombredanne suggested to also export the existing raw tokenized data before normalization happens.

## **Example/Links if Any**

## **Can you help with this Feature**

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。