ROCm / ROCm/AMDMIGraphX

Incorrect output aliases

Open
#1,252 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C++
Stars
333
Forks
150
Avg merge
4d 19h
Merged PRs (30d)
54

Description

The program fails with an error if the last op is slice.

terminate called after throwing an instance of 'migraphx::version_1::exception'
what(): /AMDMIGraphX/src/program.cpp:270: operator(): Incorrect shape {float_type, {1, 63, 672, 672}, {28901376, 451584, 672, 1}} for parameter: main:#output_1

Code example
#include <iostream>
#include <map>
#include <memory>
#include <string>
#include <vector>

#include <migraphx/argument.hpp>
#include <migraphx/generate.hpp>
#include <migraphx/gpu/hip.hpp>
#include <migraphx/gpu/target.hpp>
#include <migraphx/instruction.hpp>
#include <migraphx/make_op.hpp>
#include <migraphx/program.hpp>
#include <migraphx/ref/target.hpp>
#include <migraphx/target.hpp>

int main() {
  auto target = migraphx::target(migraphx::gpu::target{});
  migraphx::module *network_;
  migraphx::program program_;

  network_ = program_.get_main_module();

  const size_t W = 672;
  const size_t H = 672;
  const size_t K = 64;
  const size_t C = 3;
  const size_t R = 3;
  const size_t S = 3;

  const migraphx::shape input_shape{migraphx::shape::float_type, {1, C, W, H}};
  const migraphx::shape filter_shape{migraphx::shape::float_type, {K, C, R, S}};
  const auto input = network_->add_parameter("data", input_shape);

  std::vector<float> weights(filter_shape.elements(), 1);

  const auto filter =
      network_->add_literal(migraphx::literal(filter_shape, weights));

  auto out = network_->add_instruction(
      migraphx::make_op("convolution", {{"padding", {1, 1}},
                                        {"stride", {1, 1}},
                                        {"dilation", {1, 1}},
                                        {"group", 1}}),
      input, filter);

  auto slice_out_0 = network_->add_instruction(
      migraphx::make_op("slice",
                        {{"axes", {1}}, {"starts", {0}}, {"ends", {1}}}),
      out);

  auto slice_out_1 = network_->add_instruction(
      migraphx::make_op("slice",
                        {{"axes", {1}}, {"starts", {1}}, {"ends", {K}}}),
      out);

  const auto out0 =
      network_->add_parameter("main:#output_0", slice_out_0->get_shape());
  const auto out1 =
      network_->add_parameter("main:#output_1", slice_out_1->get_shape());

  network_->add_return({slice_out_0, slice_out_1});
  program_.compile(target);

  auto param_shapes = program_.get_parameter_shapes();
  for (auto &p : param_shapes) {
    std::cout << p.first << " " << p.second << '\n';
  }

  auto output_shapes = program_.get_output_shapes();
  for (auto &o : output_shapes) {
    std::cout << o << '\n';
  }

  std::vector<float> input_data;
  input_data.resize(input_shape.elements());

  std::vector<float> out0_data;
  std::vector<float> out1_data;
  out0_data.resize(output_shapes[0].elements());
  out1_data.resize(output_shapes[1].elements());

  migraphx::parameter_map pm;
  pm["data"] =
      migraphx::gpu::to_gpu(migraphx::argument(input_shape, input_data.data()));
  pm["main:#output_0"] = migraphx::gpu::to_gpu(
      migraphx::argument(output_shapes[0], out0_data.data()));
  pm["main:#output_1"] = migraphx::gpu::to_gpu(
      migraphx::argument(output_shapes[1], out1_data.data()));

  auto output = program_.eval(pm);
  for (auto &o : output) {
    std::cout << o.get_shape() << " ";
  }
}

Local configuration

  • latest develop (c99be32c013a21c41f6ad31154fc12c03eac1124)
  • gpu target
  • gfx1030

Root cause

instruction.cpp:get_output_alias returns the input tensor as the slice op defines the next output alias:

https://github.com/ROCmSoftwarePlatform/AMDMIGraphX/blob/c99be32c013a21c41f6ad31154fc12c03eac1124/src/include/migraphx/op/slice.hpp#L114

And then replace_allocate.cpp:create_output_names creates incorrect output_names with tensor from the previous layer instead of sliced tensor.

Solution

I tried the same test program with default output_alias for slice operation and it works correctly. I suppose that other operations like reshape transpose broadcast also may throw an error.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with instruction.cpp:get_output_alias and op/slice.hpp, then inspect replace_allocate.cpp:create_output_names. Reproduce the provided C++ program on the GPU target and compare the reported parameter and output shapes. Done means sliced outputs retain the correct aliases and the example evaluates without the incorrect-shape error; check whether the same behavior affects reshape, transpose, or broadcast.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
backend, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.