bytecodealliance / bytecodealliance/wasmtime

Code generated by `wasmtime` doesn't cache-align loops

Open
#4,883 3 comments 0 reactions 0 assignees View on GitHub
cranelift cranelift:E-easy cranelift:goal:optimize-speed enhancement performance
Dominant language
Rust
Stars
18.6k
Forks
1.8k
Avg merge
1d 18h
Merged PRs (30d)
126

Description

## The problem

Currently `wasmtime`/`cranelift` (unlike e.g. LLVM which doesn't have this problem AFAIK) doesn't cache-align the loops it generates, leading to potentially huge performance regressions if a hot loop ends up accidentally spanning over multiple cache lines.

## Background

Recently we were updating from `wasmtime` 0.38 to 0.40 and we saw a peculiar performance regression when doing so. One of our benchmarks took almost 2x the time to run, with a lot of them taking around ~45% more time. A huge regression. Ultimately it ended up being unrelated to the 0.38 -> 0.40 upgrade. We tracked the problem down to `memset` within the WASM (we're currently not using the bulk memory ops extension) suddenly taking a lot more time to run for no apparent reason. Depending on which exact address `wasmtime` decided to generate the code for `memset` at (which is essentially random, although consistent for the same code with the same flags in the same environment) the benchmarks were either slow, or fast, and it all boiled down to whether the hot loop of the `memset` spanned multiple cache lines or not.

You can find a detailed analysis of the problem in [this comment](https://github.com/paritytech/substrate/pull/12096#issuecomment-1238225600) and [this comment](https://github.com/paritytech/substrate/pull/12096#issuecomment-1239560799) of mine.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.