Lightning-AI / Lightning-AI/torchmetrics

Calculations nDCG using GPU are 2x slower than CPU

Open
#2,287 12 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug / fix help wanted v1.2.x
Dominant language
Python
Stars
2.5k
Forks
526
Avg merge
6d 11h
Merged PRs (30d)
5

Description

## 🐛 Bug

Hi TorchMetrics Team,

In the following example, nDCG calculation using GPU tensors spent 2 times longer the time using CPU tensors and numpy array.

### To Reproduce

The codes were tested on both Google Colab and a Slurm cluster.

Code sample

```python
import timeit

import numpy as np
import torch
from sklearn.metrics import ndcg_score
from torchmetrics.functional.retrieval import retrieval_normalized_dcg

# p and t are examples given by both sklearn and torchmetrics
p = [.1, .2, .3, 4, 70] * 100
t = [10, 0, 0, 1, 5] * 100

number = int(1e4)

# 1. BENCHMARK: numpy array
preds = np.asarray([p])
target = np.asarray([t])

def a():
return ndcg_score(target, preds)

print(f'numpy array: {timeit.timeit("a()", setup="from __main__ import a", number=number):.4f}')

# 2. cpu tensor
preds_cpu = torch.tensor(p)
target_cpu = torch.tensor(t)

assert preds_cpu.device == torch.device("cpu")

def b():
retrieval_normalized_dcg(preds_cpu, target_cpu)

print(f'CPU tensor: {timeit.timeit(f"b()", setup="from __main__ import b", number=number):.4f}')

# 3. gpu tensor
preds_gpu = torch.tensor(p, device="cuda")
target_gpu = torch.tensor(t, device="cuda")

assert preds_gpu.device == torch.device("cuda:0")

def c():
retrieval_normalized_dcg(preds_gpu, target_gpu)

print(f'GPU tensor: {timeit.timeit("c()", setup="from __main__ import c", number=number):.4f}')
```

Results:
```
# Tesla T4
numpy array: 6.4896
CPU tensor: 5.8501
GPU tensor: 10.4120
```

I also tested the codes on the Slurm Cluster I'm currently using, the GPU here is an A100.

```
numpy array: 3.8700
CPU tensor: 2.9305
GPU tensor: 7.7575
```

### Expected behavior

The performance of calculation using GPU tensors, if not superior, should be at least close to CPU tensors.

### Environment

- TorchMetrics version (and how you installed TM, e.g. `conda`, `pip`, build from source): 1.2.1 (pip)
- Python & PyTorch Version (e.g., 1.0): Python 3.10.12 and 3.10.13, Torch 2.1.0 and 2.1.1
- Any other relevant information such as OS (e.g., Linux): Ubuntu 22.04.3 LTS and Linux 5.4.204-ql-generic-12.0-19 x86_64

### Additional context

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the supplied timeit comparison and the retrieval_normalized_dcg entry point, testing CPU and CUDA tensors under the reported environment. Compare the nDCG results and execution costs; done means GPU calculation preserves the metric result without the reported slowdown relative to CPU.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.