rust-lang / rust-lang/rust

Logically equivalent code produces larger assembly and slower performance when `Arc` and `Pin` wrapped in a struct

Open
#146,630 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

C-optimization T-compiler
Dominant language
Rust
Stars
119k
Forks
16.1k
PR merge metrics
PR metrics pending

Description

I tried this code:

https://www.godbolt.org/z/Pd5aa9r86

#![feature(arbitrary_self_types)]
#![feature(rustc_attrs)]

use std::{sync::Arc,pin::Pin,};

trait Trait {
fn by_arc(self: Arc<Self>) -> i64;
fn by_pin_box(self: Pin<Box<Self>>) -> i64;
}

impl Trait for i64 {
    fn by_arc(self: Arc<Self>) -> i64 {
        *self
    }
    fn by_pin_box(self: Pin<Box<Self>>) -> i64 {
        *self
    }
}

pub fn main() {
    let arc = Arc::new(2i64) as Arc<dyn Trait>;
    assert_eq!(2, arc.by_arc());

    let pin_box = Into::<Pin<Box<i64>>>::into(Box::new(4i64)) as Pin<Box<dyn Trait>>;
    assert_eq!(4, pin_box.by_pin_box());
}

and:

#![feature(arbitrary_self_types)]
#![feature(rustc_attrs)]

pub struct TempGroup1 {
    pub field1: Arc<dyn Trait>,
    pub field2: Pin<Box<dyn Trait>>,
}

use std::{sync::Arc, pin::Pin};

trait Trait{
    fn by_arc(self: Arc<Self>) -> i64;
    fn by_pin_box(self: Pin<Box<Self>>) -> i64;
}


impl Trait for i64 {
    fn by_arc(self: Arc<Self>) -> i64{
        *self
    }
    fn by_pin_box(self: Pin<Box<Self>>) -> i64{
        *self
    }
}

pub fn main()
{
    let temp_group_1 = TempGroup1 { 
        field1: Arc::new(2i64) as Arc<dyn Trait>, 
        field2: Into::<Pin<Box<i64>>>::into(Box::new(4i64)) as Pin<Box<dyn Trait>> 
    };
    let arc = temp_group_1.field1;
    assert_eq!(2, arc.by_arc());
    let pin_box = temp_group_1.field2;
    assert_eq!(4, pin_box.by_pin_box());
}

I expected to see this happen:

Both code snippets to generate similar optimized assembly and to have similar runtime performance, since they are logically equivalent (both just wrap an i64 in Arc or Pin<Box> and call the same trait methods).

Instead, this happened:

The first version produces about 90 lines of optimized assembly (rustc -C opt-level=3 --emit=asm --crate-type=lib), while the second version produces about 286 lines.
When benchmarked with hyperfine (the main function loop with 10,000,000 iterations), the second version runs about 79.2% slower:

Version A: ≈ 0.1476 s

Version B: ≈ 0.2645 s

This suggests that introducing the TempGroup1 struct changes the generated code in a way that significantly increases overhead, even though the program logic is the same.

Meta

rustc --version --verbose:

rustc 1.89.0-nightly (586ad391f 2025-06-15)
binary: rustc
commit-hash: 586ad391f5ee4519acc7cae340e34673bae762b1
commit-date: 2025-06-15
host: x86_64-unknown-linux-gnu
release: 1.89.0-nightly
LLVM version: 20.1.5
Version A hyperfine.json

{
  "results": [
    {
      "command": "./target/release/init",
      "mean": 0.1476577922789474,
      "stddev": 0.010808251268166473,
      "median": 0.1482381797,
      "user": 0.14682449894736843,
      "system": 0.0010573589473684211,
      "min": 0.1263320227,
      "max": 0.1653164557,
      "times": [
        0.1482381797,
        0.1593237487,
        0.1408860767,
        0.1636745287,
        0.1653164557,
        0.1612710107,
        0.1519671997,
        0.1549548597,
        0.1503653417,
        0.1405962647,
        0.1263320227,
        0.1509238887,
        0.1472579667,
        0.1534259367,
        0.1366232097,
        0.1329117547,
        0.1408619007,
        0.1355554947,
        0.1450122127
      ],
      "exit_codes": [
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0
      ]
    }
  ]
}

Version B hyperfine.json

{
  "results": [
    {
      "command": "./target/release/new",
      "mean": 0.2645699604018182,
      "stddev": 0.009253164823394697,
      "median": 0.26372722522000003,
      "user": 0.2632035254545454,
      "system": 0.0014536563636363636,
      "min": 0.24764987322,
      "max": 0.28167238222,
      "times": [
        0.27004933622000005,
        0.26448221622000007,
        0.25952847222000003,
        0.24764987322,
        0.26372722522000003,
        0.26165526122000005,
        0.25714677222000004,
        0.28167238222,
        0.26352976122000005,
        0.27705794322000005,
        0.26377032122000005
      ],
      "exit_codes": [
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0,
        0
      ]
    }
  ]
}

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the two main reproductions and the linked Godbolt example, then run rustc -C opt-level=3 --emit=asm --crate-type=lib on both versions. Compare the generated assembly and the reported hyperfine timings, narrowing down which compiler stage accounts for the difference. Done means the cause is identified and the equivalent snippets no longer show the reported code-generation or performance gap.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
compilers, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.