MaartenGr / MaartenGr/BERTopic

fit_transform tries to access embedding_model if representation_model is not None

Open
#2,189 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Python
Stars
7.8k
Forks
920
Avg merge
22h 24m
Merged PRs (30d)
5

Description

### Have you searched existing issues? 🔎

- [X] I have searched and found no existing issues

### Desribe the bug

I was using BERTopic on a cluster of queries with my own embeddings (computed on a model that is hard to pass as a parameter) and it was working as expected.
After trying to use `representation_model = KeyBERTInspired()` and adding ` representation_model=representation_model` to BERTopic as a parameter. I got this error :

```
AttributeError Traceback (most recent call last)
[1] representation_model = KeyBERTInspired()
[2] topic_model = BERTopic(
[3] calculate_probabilities=True,
[4] min_topic_size=1
[5] embedding_model=None,
[6] representation_model=representation_model,
[7] )
----> [8] topics, probs = topic_model.fit_transform(corpus, np.array(corpus_embeddings))
[9] topic_model.get_topic_info()

File ~/query_analysis/bertopic_env/lib/python3.11/site-packages/bertopic/_bertopic.py:493, in BERTopic.fit_transform(self, documents, embeddings, images, y)
[490] self._save_representative_docs(custom_documents)
[491] else:
[492] # Extract topics by calculating c-TF-IDF
--> [493] self._extract_topics(documents, embeddings=embeddings, verbose=self.verbose)
[495] # Reduce topics
[496] if self.nr_topics:

File ~/query_analysis/bertopic_env/lib/python3.11/site-packages/bertopic/_bertopic.py:3991, in BERTopic._extract_topics(self, documents, embeddings, mappings, verbose)
[3989] documents_per_topic = documents.groupby(["Topic"], as_index=False).agg({"Document": " ".join})
[3990] self.c_tf_idf_, words = self._c_tf_idf(documents_per_topic)
-> [3991] self.topic_representations_ = self._extract_words_per_topic(words, documents)
...
[3680] "Make sure to use an embedding model that can either embed documents"
[3681] "or images depending on which you want to embed."
[3682]

AttributeError: 'NoneType' object has no attribute 'embed_documents'
```

### Reproduction

```python
from query import Query
import json
import numpy as np
from sklearn.cluster import DBSCAN
from bertopic import BERTopic
from openai import OpenAI
from bertopic.representation import KeyBERTInspired

representation_model = KeyBERTInspired()
topic_model = BERTopic(
calculate_probabilities=True,
min_topic_size=15,
embedding_model=None,
representation_model=representation_model,
)
topics, probs = topic_model.fit_transform(corpus, np.array(corpus_embeddings))
topic_model.get_topic_info()

```

### BERTopic Version

0.16.4

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in bertopic/_bertopic.py at fit_transform, then follow _extract_topics into _extract_words_per_topic, where the reported AttributeError occurs with KeyBERTInspired and embedding_model=None. Reproduce the example using supplied corpus_embeddings and confirm that fitting with a representation_model no longer requires access to embedding_model.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.