aboutcode-org / aboutcode-org/scancode-toolkit

Bug: CSV output assigns start_line to end_line for copyrights, holders, and authors

未关闭
#4,785 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
bug
主要语言
Python
星标
2.6k
派生
791
平均合并
1 天 12 小时
30 天内合并 PR
5

描述

## Description

In `src/formattedcode/output_csv.py`, the `flatten_scan()` function incorrectly assigns `copyr['start_line']` to `inf['end_line']` for copyrights, holders, and authors entries. This means the `end_line` column in CSV output always contains the same value as `start_line`, losing the actual end line information.

## Affected lines

- **Line 188** (copyrights): `inf['end_line'] = copyr['start_line']`
- **Line 196** (holders): `inf['end_line'] = copyr['start_line']`
- **Line 204** (authors): `inf['end_line'] = copyr['start_line']`

All three should be `copyr['end_line']`.

## Evidence

The underlying data model (`CopyrightDetection`, `HolderDetection`, `AuthorDetection` in `src/cluecode/copyrights.py`) all define both `start_line` and `end_line` as attrs fields. The JSON output correctly includes both values. Only the CSV output has this issue.

Git blame traces this to commit `ef8086d` ("Update CSV output to latest copyright data format", ~2018), suggesting a copy-paste error in the original refactor.

## Expected behavior

`inf['end_line']` should be assigned `copyr['end_line']` so that the CSV output reflects accurate line ranges for copyright, holder, and author detections.

## How to reproduce

1. Run a scan with copyright detection and CSV output: scancode -c --csv output.csv
2. Open `output.csv` and compare the `start_line` and `end_line` columns for copyright/holder/author rows
3. Notice `end_line` always equals `start_line`, even for multi-line copyright statements

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。