pytorch / pytorch/executorch

Ops with included input dim order conversion fail at runtime

Open
#11,086 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

module: exir
Dominant language
Python
Stars
5k
Forks
1.2k
Avg merge
2d 10h
Merged PRs (30d)
581

Description

🐛 Describe the bug

Certain ops in pytorch can internally convert their inputs to contiguous format. Linear, for example, seems to do this. These dim order conversions are reflected in the node metadata (and thus ET dim order), but not in the graph. Our kernels don't include the input dim order conversion and thus fail at runtime.

A slightly contrived example is as follows, though I expect that this is reproable with any op with an ATen kernel that includes an implicit conversion to contiguous, of which there are quite a few.

import torch

class Module(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.linear = torch.nn.Linear(10, 10)

    def forward(self, x):
        return self.linear(x)

inputs = (torch.randn(1,2,3,10).to(memory_format=torch.channels_last),)
module = Module()

ep = torch.export.export(module, inputs)
from executorch.exir import to_edge_transform_and_lower

et_program = to_edge_transform_and_lower(ep).to_executorch()
from executorch.extension.pybindings.portable_lib import _load_for_executorch_from_buffer

et_model = _load_for_executorch_from_buffer(et_program.buffer)

Output:

[tensor_util_portable.cpp:132] Check failed (all_contiguous || all_channels_last): 2 input tensors have different dim orders
[op_expand_copy.cpp:89] Check failed (tensors_have_same_dim_order(self, out)): 
[method.cpp:1322] KernelCall failed at instruction 0:1 in operator aten::expand_copy.out: 0x12
[method.cpp:1328] arg 0 with type id 1
[method.cpp:1328] arg 1 with type id 8
[method.cpp:1328] arg 2 with type id 5
[method.cpp:1328] arg 3 with type id 1
[method.cpp:1328] arg 4 with type id 1

Note that the failure actually occurs in expand_copy due to the weird decomp in this case.

Versions

N/A

cc @JacobSzwejbka @angelayi

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the provided torch.export.export and to_edge_transform_and_lower reproduction, then inspect tensor_util_portable.cpp and op_expand_copy.cpp around the reported failures. Trace how input dim-order metadata reaches the graph and portable kernels. Done means the channels-last Linear example loads and runs without dim-order mismatch errors.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
embedded-iot, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.