MaartenGr / MaartenGr/BERTopic
Zero-shot topic modeling with nr_topics parameter results in empty topic_sizes_ and get_topic_info()
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 7.8k
- Forks
- 920
- Avg merge
- 22h 24m
- Merged PRs (30d)
- 5
Description
Have you searched existing issues? 🔎
- I have searched and found no existing issues
Desribe the bug
When using zero-shot topic modeling with the nr_topics parameter set, and all (or most) documents are assigned to zero-shot topics, the topic_sizes_ attribute remains empty and get_topic_info() returns an empty DataFrame.
Root cause:
The issue occurs in the fit_transform() method when all documents are assigned to zero-shot topics. The code has this conditional:
if len(documents) > 0:
# Cluster reduced embeddings
documents, probabilities = self._cluster_embeddings(umap_embeddings, documents, y=y)
if self._is_zeroshot() and len(assigned_documents) > 0:
documents, embeddings = self._combine_zeroshot_topics(
documents, embeddings, assigned_documents, assigned_embeddings
)
else:
# All documents matches zero-shot topics
documents = assigned_documents
embeddings = assigned_embeddings
When all documents are assigned to zero-shot topics, len(documents) == 0, so the else branch is taken. However, this branch never calls _update_topic_size(), which is normally called within _cluster_embeddings() or _combine_zeroshot_topics().
Additionally, when nr_topics is specified, the fallback call to _sort_mappings_by_frequency() (which would call _update_topic_size()) is skipped because if not self.nr_topics: evaluates to False.
Expected behavior:
topic_sizes_should contain the count of documents per topicget_topic_info()should return a populated DataFrame with topic information
Actual behavior:
topic_sizes_is an empty dictionary{}get_topic_info()returns an empty DataFrametopic_representations_works correctly (populated as expected)
Note: The issue does NOT occur when nr_topics is not specified, because _sort_mappings_by_frequency() gets called, which internally calls _update_topic_size().
Reproduction
from bertopic import BERTopic
# Sample documents and zero-shot topics
docs = ["I need help with my voucher", "Gift card not working", "Customer service was poor"] * 50
zeroshot_topics = ["Voucher inquiries", "Gift card issues", "Customer service feedback"]
# BUG: Setting nr_topics causes the issue
model = BERTopic(
zeroshot_topic_list=zeroshot_topics,
zeroshot_min_similarity=-1, # Force all documents to zero-shot assignment
nr_topics=4, # This triggers the bug
)
topics, _ = model.fit_transform(docs)
# Demonstrates the bug
print("topic_sizes_:", model.topic_sizes_) # Expected: Counter with topic counts, Actual: None
print("topic_representations_ works:", bool(model.topic_representations_)) # This works: True
BERTopic Version
0.17.0
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start in fit_transform() and trace the all-zero-shot else branch alongside _cluster_embeddings(), combine_zeroshot_topics(), and sort_mappings_by_frequency(). Use the supplied zero-shot reproduction with nr_topics=4, then verify that topic_sizes contains document counts and get_topic_info() returns a populated DataFrame while topic_representations remains populated.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Clearly specified
- Newbie friendliness
- 48/100