pytorch / pytorch/pytorch

[Bug] ValueError: I/O operation on closed file during torch.save() on Windows (PyTorch Nightly + CUDA 13.2)

Open
#180,458 1 comment 0 reactions 0 assignees View on GitHub
bot-triaged module: multiprocessing module: regression module: serialization module: windows needs reproduction triaged
Dominant language
Python
Stars
103k
Forks
29.8k
PR merge metrics
PR metrics pending

Description

### 🐛 Describe the bug

When attempting to save model checkpoints (via torch.save()) during a training loop on Windows using PyTorch 2.12 Nightly (with CUDA 13.2), the process randomly crashes with a ValueError: I/O operation on closed file inside torch.serialization.py.

The crash happens specifically when zip_file.write_record() attempts to write data to the .pt archive. This seems to be a race condition or a file-lock issue related to the new ZipFile writer implementation in the nightly build, possibly exacerbated by Windows multiprocessing/I/O handling.

**To Reproduce**
Steps to reproduce the behavior:
1. Install the PyTorch Nightly build with CUDA 13.2 support on a Windows machine:
`pip install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/cu132`
2. Start a training loop using the `ultralytics` package (which calls `torch.save()` at the end of each epoch to save `last.pt`).
3. The training loop runs normally for the first epochs, but crashes during the file serialization process.

*Minimal Code Context:*
```python
from ultralytics import YOLO

if __name__ == '__main__':
model = YOLO('yolo26n.pt')
model.train(
data="data.yaml",
epochs=100,
imgsz=640,
device=0,
workers=4, # Issue might be related to Windows DataLoader workers
)
```

**Expected behavior**
`torch.save()` should successfully write the `.pt` archive without throwing an I/O Exception or dropping the file lock prematurely.

**Error logs / Traceback**
```python
Traceback (most recent call last):
File "[...]\.env\Lib\site-packages\torch\serialization.py", line 1004, in save
_save(
obj,
...<3 lines>...
_disable_byteorder_record,
)
File "[...]\.env\Lib\site-packages\torch\serialization.py", line 1260, in _save
zip_file.write_record("data.pkl", data_value, len(data_value))
ValueError: I/O operation on closed file.

During handling of the above exception, another exception occurred:

Traceback (most recent call last):
File "[...]\main.py", line 16, in
model.train(
data="./[...]/data.yaml",
...<12 lines>...
name="RTX3080_Cuda13_2"
)
File "[...]\.env\Lib\site-packages\ultralytics\engine\model.py", line 787, in train
self.trainer.train()
File "[...]\.env\Lib\site-packages\ultralytics\engine\trainer.py", line 246, in train
self._do_train()
File "[...]\.env\Lib\site-packages\ultralytics\engine\trainer.py", line 535, in _do_train
if (self.args.save or final_epoch) and self.save_model():
File "[...]\.env\Lib\site-packages\ultralytics\engine\trainer.py", line 643, in save_model
torch.save(
{
...<21 lines>...
buffer,
)
File "[...]\.env\Lib\site-packages\ultralytics\utils\patches.py", line 197, in torch_save
return _torch_save(*args, **kwargs)
File "[...]\.env\Lib\site-packages\torch\serialization.py", line 1003, in save
with _open_zipfile_writer(f) as opened_zipfile:
File "[...]\.env\Lib\site-packages\torch\serialization.py", line 855, in __exit__
self.file_like.write_end_of_file()
ValueError: I/O operation on closed file.
```

**Additional context **
Setting `workers=0` inside the Ultralytics DataLoader seems to mitigate/bypass the issue, suggesting a potential clash between Windows multi-processing and the new ZipFile writer implementation when acquiring or releasing the file handle.

### Versions

PyTorch version: 2.12.0.dev20260415+cu132
Is debug build: False
CUDA used to build PyTorch: 13.2
ROCM used to build PyTorch: N/A

OS: Microsoft Windows 11 Pro (10.0.26200 64 bit)
GCC version: Could not collect
Clang version: Could not collect
CMake version: Could not collect
Libc version: N/A

Python version: 3.14.3 (tags/v3.14.3:323c59a, Feb 3 2026, 16:04:56) [MSC v.1944 64 bit (AMD64)] (64-bit runtime)
Python platform: Windows-11-10.0.26200-SP0
Is CUDA available: True
CUDA runtime version: 13.2.78
CUDA_MODULE_LOADING set to:
GPU models and configuration: GPU 0: NVIDIA GeForce RTX 3080
Nvidia driver version: 595.97
cuDNN version: Could not collect
Is XPU available: False
HIP runtime version: N/A
MIOpen runtime version: N/A
Is XNNPACK available: True
Caching allocator config: N/A

CPU:
Name: AMD Ryzen 7 5800X 8-Core Processor
Manufacturer: AuthenticAMD
Family: 107
Architecture: 9
ProcessorType: 3
DeviceID: CPU0
CurrentClockSpeed: 3801
MaxClockSpeed: 3801
L2CacheSize: 4096
L2CacheSpeed: None
Revision: 8448

Versions of relevant libraries:
[pip3] numpy==2.4.4
[pip3] torch==2.12.0.dev20260415+cu132
[pip3] torchaudio==2.11.0
[pip3] torchvision==0.27.0.dev20260414+cu132
[conda] Could not collect

cc @peterjc123 @mszhanyi @skyline75489 @nbcsm @iremyux @Blackhex @VitalyFedyunin @albanD @pragupta @ppwwyyxx @mruberry @mikaylagawarecki

Contributor guide

Open the contributing guide

Research direction

Start in torch/serialization.py at save(), _save(), and _open_zipfile_writer.__exit__, then reproduce the failure on Windows with the reported nightly build and workers=4. Compare with workers=0 as a diagnostic. Done means torch.save() completes without the closed-file ValueError during archive writing and cleanup.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.