dimforge / dimforge/nalgebra

Nalgebra seems to run an order of magnitude slower than Numpy, with the same backend bindings

Open
#1,468 5 comments 1 reaction 0 assignees View on GitHub
Dominant language
Rust
Stars
4.8k
Forks
565
PR merge metrics
No merged PRs in 30d

Description

Hi,

I wanted to try out nalgebra and see how it compares to very vanilla operations in numpy. I _think_ I've implemented everything the 'right' way according to the documentation, but I'm really struggling with the performance here:

My rust code here:

```rust
use rand::distributions::{Distribution, Uniform};
use rand::thread_rng;
use rayon::prelude::*;

extern crate nalgebra as na;
extern crate nalgebra_lapack;

use std::time::Instant;

fn main() {

let start = Instant::now();
let rows = 2000;
let cols = 2000;

// Pre-allocate the vector with exact capacity
let mut data = Vec::with_capacity(rows * cols);

// Initialize random number generator and distribution
let uniform = Uniform::new(0.0, 1.0);

// Generate random numbers in parallel
data.par_extend(
(0..rows * cols)
.into_par_iter()
.map(|_|
{
let mut local_rng = thread_rng(); // Create a new RNG for each thread
uniform.sample(&mut local_rng)
})
);

// Create matrix from pre-generated data
let a = na::DMatrix::from_vec(rows, cols, data);

let creation_time = start.elapsed();
println!("Matrix creation: {:?}", creation_time);
println!("Matrix dimensions: {} x {}", a.nrows(), a.ncols());

// Compute A^T A directly without storing the transpose
let mult_start = Instant::now();
let c = &a * &a.transpose(); // Use references to avoid unnecessary copies

println!("Result dimensions: {} x {}", c.nrows(), c.ncols());
let mult_time = mult_start.elapsed();
println!("Matrix multiplication: {:?}", mult_time);

// Use symmetric eigenvalue computation since C = A^T A is symmetric
let eig_start = Instant::now();
let _eigvals = na::linalg::SymmetricEigen::new(c.clone()).eigenvalues;
let eig_time = eig_start.elapsed();
println!("Eigenvalues computation: {:?}", eig_time);

// Use symmetric SVD since C is symmetric positive semidefinite
let svd_start = Instant::now();
let svd = na::linalg::SVD::new(c, true, true);
let _singular_values = svd.singular_values;
let svd_time = svd_start.elapsed();
println!("SVD calculation: {:?}", svd_time);

println!("Total time: {:?}", start.elapsed());
}
```

my cargo.toml:

```rust
[package]
name = "rust"
version = "0.1.0"
edition = "2021"

[dependencies]
nalgebra = "0.33.2"
nalgebra-lapack = { version = "0.25.0", default-features = false, features = ["accelerate"] }
rand = "0.8.5"
rayon = "1.10.0"

and my build.rs

fn main() {
#[cfg(target_os = "macos")]
{
println!("cargo:rustc-link-lib=framework=Accelerate");
}
}
```

and, for comparison, my Python code, a super simple script using numpy:

```python
import numpy as np
import time

t1 = time.time()
a = np.random.rand(2_000, 2_000)
a.shape
print(a)
t2 = time.time()
b = a.T @ a
t3 = time.time()
# calculate the svd of b
result = np.linalg.eigvals(b)
# print(result)
t4 = time.time()
t5 = time.time()
result = np.linalg.svd(b)
t6 = time.time()
# cholesky decomp
result = np.linalg.cholesky(b)
t7 = time.time()
# np.show_config()
print(f"Total time: {(t6 - t1)*1e3:.1f} miliseconds")
print(f"Matrix creation: {(t2 - t1)*1e3:.1f} miliseconds")
print(f"Matrix multiplication: {(t3 - t2)*1e3:.1f} miliseconds")
print(f"Eigenval calculation: {(t4 - t3)*1e3:.1f} miliseconds")
print(f"SVD calculation: {(t6 - t5)*1e3:.1f} miliseconds")
print(f"Cholesky decomposition: {(t7 - t6)*1e3:.1f} miliseconds")
```

I find that when I run both, I get these results:

RUST:

❯ cargo run --release
Finished `release` profile [optimized] target(s) in 0.08s
Running `target/release/rust`
Matrix creation: 9.157083ms
Matrix dimensions: 2000 x 2000
Result dimensions: 2000 x 2000
Matrix multiplication: 364.901042ms
Eigenvalues computation: 4.586041791s
SVD calculation: 14.505883666s
Total time: 19.466049208s

and

PYTHON:

Total time: 3177.9 miliseconds
Matrix creation: 17.5 miliseconds
Matrix multiplication: 22.2 miliseconds
Eigenval calculation: 1377.0 miliseconds
SVD calculation: 1761.1 miliseconds
Cholesky decomposition: 27.5 miliseconds

You can see that Python is more than 10x faster at matmul, ~4x faster at calculating eigenvalues, and an order of magnitude faster at computing the SVD.

Please could someone tell me where I'm going wrong here?

Some further Info:

1. I'm running in 'release' mode.
2. I know my binary is linked to Apple accelerate because:
❯ otool -L target/release/rust
target/release/rust:
/System/Library/Frameworks/Accelerate.framework/Versions/A/Accelerate (compatibility version 1.0.0, current version 4.0.0)
/usr/lib/libiconv.2.dylib (compatibility version 7.0.0, current version 7.0.0)
/usr/lib/libSystem.B.dylib (compatibility version 1.0.0, current version 1351.0.0)
3. I've also tried with the default openblas setting, and performance is even worse.

If someone could please help me out / point me to where I'm going wrong that would be great. I am aware Python is going to be hard to beat here (because this Python code is basically just running C / Fortran optimised for Apple chips), but, I would expect at the very least that the Rust would be on-par with the performance of really-simple Python, given that both are targeting the same backend (Apple accelerate).

Thanks!

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.