rust-lang / rust-lang/rust

Inefficient code generation with target-feature AVX2

Open
#137,335 5 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

A-codegen A-LLVM A-target-feature C-bug I-slow T-compiler
Dominant language
Rust
Stars
119k
Forks
16.1k
PR merge metrics
PR metrics pending

Description

I tried this code:

use std::time::Instant;

fn main() {
  let mut scratch_buf = [25; 4096];
  let color = [30; 16];
  let x = 0;
  let width = 256;

  let start = Instant::now();
  for _ in 0..200000 {
    fill_solid(&mut scratch_buf, &color, x, width);
  }

  println!("Ran for {:?}", start.elapsed());
}

pub(crate) fn fill_solid(
  scratch: &mut [u8; 4096],
  color: &[u8; 16],
  x: usize,
  width: usize,
) {
  let target = &mut scratch[x * 16..][..16 * width];

  let dest = target.chunks_exact_mut(16);

  for cb in dest {
    for i in 0..16 {
      cb[i] = color[i] + ((color[i] as u16 * cb[i] as u16) / 255) as u8;
    }
  }
}

I expected to see this happen: When compiling with RUSTFLAGS="-C target-feature=+avx2", I at least expected the code to not be much slower.

Instead, this happened: The code runs 7x slower than when compiled without this target feature.

RUSTFLAGS="-C target-feature=+avx2" cargo run --release
Ran for 226.4182ms

cargo run --release
Ran for 31.5613ms

I can't tell what exactly is going on, but the AVX code definitely looks much more verbose: https://godbolt.org/z/s48ccT1fn

Meta

rustc --version --verbose:

rustc 1.84.0 (9fc6b4312 2025-01-07)
binary: rustc
commit-hash: 9fc6b43126469e3858e2fe86cafb4f0fd5068869
commit-date: 2025-01-07
host: x86_64-pc-windows-msvc
release: 1.84.0
LLVM version: 19.1.5

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the benchmark with and without RUSTFLAGS="-C target-feature=+avx2" using the Rust snippet and cargo run --release. Compare the generated assembly through the linked Godbolt example and investigate the rustc/LLVM code-generation path; done means identifying why AVX2 is slower and documenting or fixing the regression with a validating test or benchmark.

Written by the indexing model from the issue text.

Assessment

Tech stack
rust
Domain
compilers, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.