pytorch / pytorch/vision

`PILToTensor` input shape description is wrong/confusing

Open
#9,221 2 comments 0 reactions 1 assignee View on GitHub

@AntoineSimoulin is already working on this.

Since Sep 16, 2025.

Dominant language
Python
Stars
17.9k
Forks
7.3k
Avg merge
1d 15h
Merged PRs (30d)
13

Description

📚 The doc issue

The PILToTensor documentation pages (for both torchvision.transforms.v2.PILToTensor and torchvision.transforms.PILToTensor) state:

Converts a PIL Image (H x W x C) to a Tensor of shape (C x H x W).

This is confusing, because img.size returns the dimensions of a PIL.Image (img, in this case) in (width, height) format.

Suggest a potential alternative/fix

Change the quoted statement to the following:

Converts a PIL Image (W x H x C) to a Tensor of shape (C x H x W).

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.