Add converter and formatter parameters to csv.reader and csv.writer
還沒有人認領這個 Issue。
- 主要語言
- Python
- 星號
- 77.2k
- 分支
- 36k
- PR 合併指標
- PR 指標待擷取
描述
Feature or enhancement
Proposal:
The conversion between Python values and CSV fields is hard-coded in both directions. The reader converts unquoted fields with float(), and only in the QUOTE_NONNUMERIC and QUOTE_STRINGS modes. The writer converts every non-string value with str().
This is the common cause of several open issues:
- gh-74232 --
boolis written unquoted asTrue, which cannot be read back. - gh-98485 -- the same for
complex;FractionandIntEnumare affected too, andDecimalsilently round-trips throughfloat. - gh-110852 -- there is no way to write floats with a fixed precision, because preformatted strings are quoted in the
QUOTE_NONNUMERICmode. - gh-85002 -- there is no way to reject values which are neither strings nor numbers.
I propose two parameters, mirroring parse_float in json:
csv.reader(f, converter=None)-- called asconverter(index, field)instead offloat().csv.writer(f, formatter=None)-- called asformatter(index, value)instead ofstr(). It must return a string.
index is the 0-based position of the field in the record. Both default to None, which keeps the current behavior. The hooks only replace the existing calls -- what is not passed to float() or str() now is not passed to them either. Quoting is still decided by the original value.
The index goes first, like in enumerate(). This also makes a wrong one-argument callable fail at once: converter=int raises TypeError on the first field instead of taking the index as the base.
The index makes the hooks per-column, which is what the dtype and converters parameters of pandas.read_csv() are used for:
>>> types = [str, int, Decimal, Fraction]
>>> list(csv.reader(['spam,42,1.10,1/2'], quoting=csv.QUOTE_NONNUMERIC,
... converter=lambda i, field: types[i](field)))
[['spam', 42, Decimal('1.10'), Fraction(1, 2)]]
>>> def money(index, value):
... return format(value, '.2f') if index == 2 else str(value)
>>> csv.writer(sys.stdout, formatter=money).writerow(['a', 1, 0.0, 3.14159])
a,1,0.00,3.14159
gh-85002 no longer needs a parameter of its own -- a strict writer is a formatter which refuses everything except numbers.
I have a working prototype (about 90 lines in Modules/_csv.c).
Open question: should these be parameters of the reader and the writer, or attributes of the dialect? A dialect is a portable description of the file syntax -- it is registered under a global name, sniffed, and copied -- so keeping callables out of it seems better.
Linked PRs
- gh-155099
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
研究方向
從 Modules/_csv.c 以及提案中描述的 reader/writer 進入點開始;檢視現有的 float() 和 str() 轉換路徑,以及其中提到的原型。比較參數和方言的替代方案,然後使用連結的 gh-155099 工作來確定已達成共識的 API 和完成標準。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- python
- 領域
- data
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 停滯
- 描述清晰度
- 基本清楚
- 新手友好度
- 25/100