python / python/cpython

Improve `tarfile` streaming mode to handle very large archives

Đang mở
#139,960 3 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

stdlib type-feature
Ngôn ngữ chính
Python
Star
77.2k
Fork
35.9k
Chỉ số merge pull request
Chỉ số pull request đang chờ

Mô tả

Bug report

Bug description:

I am trying to use tarfile to write very large archives, and the process is being killed by the OOM killer. I would expect+want to be able to use it in streaming mode (write an unlimited number of files) like the standard tar command line utility supports by default.

Reproduction case attached.

import gc
import io
import os
import psutil
import tarfile

if __name__ == "__main__":
    t = tarfile.open("a.tar", mode="w") # default compresslevel 9
    for i in range(1,100_000_000):
        if i % 10_000 == 0:
            gc.collect()
            process = psutil.Process(os.getpid())
            mem_info = process.memory_info()
            mem = mem_info.rss
            print(f"Iteration {i}, memory usage: {mem}")

        bs = (" "*1000 + str(i)).encode('utf8')
        with io.BytesIO(bs) as file:
            tarinfo = tarfile.TarInfo(name="cool_files/{i}.txt")
            tarinfo.size = len(bs)
            t.addfile(tarinfo, file)

The memory usage increases without bound because the list of this line in addfile():

self.members.append(tarinfo)

I'm not sure what the use-case is for this line. In write-only mode, it does not seem useful. Maybe mixed read/write? But in general it does not seem correct to assume you can fit all the tarinfo's in memory.

Edit: As a workaround, I'm setting t.members=[] manually.

CPython versions tested on:

3.13

Operating systems tested on:

Linux

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Hướng nghiên cứu

Bắt đầu với phần triển khai tarfile xung quanh TarFile.addfile() và dòng self.members.append(tarinfo), sử dụng bản tái hiện được báo cáo để quan sát mức tăng bộ nhớ. Xác định hành vi dự kiến cho chế độ chỉ ghi hoặc streaming, sau đó xác minh rằng việc ghi các archive rất lớn vẫn giữ mức sử dụng bộ nhớ trong giới hạn mà không làm hỏng các trường hợp sử dụng khác của tarfile.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
python
Lĩnh vực
backend
Loại issue
Lỗi
Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
45/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.