python / python/cpython

Add converter and formatter parameters to csv.reader and csv.writer

Đang mở
#155,097 0 bình luận 1 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

extension-modules type-feature
Ngôn ngữ chính
Python
Star
77.2k
Fork
35.9k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

Feature or enhancement

Proposal:

The conversion between Python values and CSV fields is hard-coded in both directions. The reader converts unquoted fields with float(), and only in the QUOTE_NONNUMERIC and QUOTE_STRINGS modes. The writer converts every non-string value with str().

This is the common cause of several open issues:

  • gh-74232 -- bool is written unquoted as True, which cannot be read back.
  • gh-98485 -- the same for complex; Fraction and IntEnum are affected too, and Decimal silently round-trips through float.
  • gh-110852 -- there is no way to write floats with a fixed precision, because preformatted strings are quoted in the QUOTE_NONNUMERIC mode.
  • gh-85002 -- there is no way to reject values which are neither strings nor numbers.

I propose two parameters, mirroring parse_float in json:

  • csv.reader(f, converter=None) -- called as converter(index, field) instead of float().
  • csv.writer(f, formatter=None) -- called as formatter(index, value) instead of str(). It must return a string.

index is the 0-based position of the field in the record. Both default to None, which keeps the current behavior. The hooks only replace the existing calls -- what is not passed to float() or str() now is not passed to them either. Quoting is still decided by the original value.

The index goes first, like in enumerate(). This also makes a wrong one-argument callable fail at once: converter=int raises TypeError on the first field instead of taking the index as the base.

The index makes the hooks per-column, which is what the dtype and converters parameters of pandas.read_csv() are used for:

>>> types = [str, int, Decimal, Fraction]
>>> list(csv.reader(['spam,42,1.10,1/2'], quoting=csv.QUOTE_NONNUMERIC,
...                 converter=lambda i, field: types[i](field)))
[['spam', 42, Decimal('1.10'), Fraction(1, 2)]]
>>> def money(index, value):
...     return format(value, '.2f') if index == 2 else str(value)
>>> csv.writer(sys.stdout, formatter=money).writerow(['a', 1, 0.0, 3.14159])
a,1,0.00,3.14159

gh-85002 no longer needs a parameter of its own -- a strict writer is a formatter which refuses everything except numbers.

I have a working prototype (about 90 lines in Modules/_csv.c).

Open question: should these be parameters of the reader and the writer, or attributes of the dialect? A dialect is a portable description of the file syntax -- it is registered under a global name, sniffed, and copied -- so keeping callables out of it seems better.

Linked PRs
  • gh-155099

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu với Modules/_csv.c và các điểm vào reader/writer được mô tả trong đề xuất; xem xét các đường chuyển đổi float() và str() hiện có cùng prototype được đề cập trong đó. So sánh các phương án về tham số và dialect, sau đó sử dụng công việc được liên kết gh-155099 để xác định API đã thống nhất và các tiêu chí hoàn thành.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
data
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.