MaartenGr / MaartenGr/BERTopic

Multi-GPU Utilisation

Open
#1,837 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
7.8k
Forks
920
Avg merge
22h 24m
Merged PRs (30d)
5

Description

Hi Maarten,
I'm attempting to execute one of your examples in Google Colab for processing large-scale databases. Here are the specifications of my machine: 8 NVIDIA A100 cards and a 50TB SSD. However, when running the code, it appears to only utilize one of the GPUs. Could you advise on how I can distribute the workload across all 8 cards?

```
import numpy as np
from torch import cuda

# Set device to use all available GPUs
num_gpus = cuda.device_count()
if num_gpus > 0:
device_ids = list(range(num_gpus)) # Assuming GPUs are indexed from 0 to 7
device = f'cuda:{device_ids[0:8]}' # Set the device to the first GPU
print("Available GPUs:", num_gpus)
print("Using GPUs:", device_ids)
else:
device = 'cpu'
print("CUDA is not available. Using CPU.")

print("Device:", device)
##################
from datasets import load_dataset
# Extract 1 millions records
lang = 'en'
data = load_dataset(f"Cohere/wikipedia-22-12", lang, split='train', streaming=True)
docs = [doc["text"] for doc in data if doc["id"] != "1_000_000"];
# Embeddings
from sentence_transformers import SentenceTransformer
# Create embeddings
model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
embeddings = model.encode(docs, show_progress_bar=True)
import collections
from tqdm import tqdm
from sklearn.feature_extraction.text import CountVectorizer

# Extract vocab to be used in BERTopic
vocab = collections.Counter()
tokenizer = CountVectorizer().build_tokenizer()
for doc in tqdm(docs):
vocab.update(tokenizer(doc))
vocab = [word for word, frequency in vocab.items() if frequency >= 15];
len(vocab)

# Train BERTopic
from cuml.manifold import UMAP
from cuml.cluster import HDBSCAN
from bertopic import BERTopic

# Prepare sub-models
embedding_model = SentenceTransformer('all-MiniLM-L6-v2')
umap_model = UMAP(n_components=5, n_neighbors=50, random_state=42, metric="cosine", verbose=True)
hdbscan_model = HDBSCAN(min_samples=20, gen_min_span_tree=True, prediction_data=False, min_cluster_size=20,
verbose=True)
vectorizer_model = CountVectorizer(vocabulary=vocab, stop_words="english")

# Fit BERTopic without actually performing any clustering
topic_model = BERTopic(
embedding_model=embedding_model,
umap_model=umap_model,
hdbscan_model=hdbscan_model,
vectorizer_model=vectorizer_model,
verbose=True
).fit(docs, embeddings=embeddings)

```

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the supplied Python example and inspect the SentenceTransformer.encode and BERTopic.fit calls. The issue names no repository file or test, so determine which project entry point controls device use and define completion as a documented, reproducible multi-GPU path or a clearly stated supported limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.