MaartenGr / MaartenGr/BERTopic

Issues with visualizations on loaded models.

Open
#2,032 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
7.8k
Forks
920
Avg merge
22h 24m
Merged PRs (30d)
5

Description

I am storing my model in this manner:
embedding_model = "sentence-transformers/all-MiniLM-L6-v2"
topic_model.save("data/topic_model", serialization="safetensors", save_ctfidf=True, save_embedding_model=embedding_model)

Then, I load it like:
topic_model = BERTopic.load("/tmp/topic_model")

Then I want to use the original set of documents that the model was fitted on to visualize these topics over time.

topics_over_time = main_model.topics_over_time( ticket_training_set.description, ticket_training_set.created_at, nr_bins=30 )
main_model.visualize_topics_over_time(topics_over_time)

I get the following error:

ValueError                                Traceback (most recent call last)
Cell In[35], line 1
----> 1 topics_over_time = main_model.topics_over_time(
      2     ticket_training_set.description,
      3     ticket_training_set.created_at,
      4     nr_bins=30
      5 )
      6 main_model.visualize_topics_over_time(topics_over_time)

File ~/anaconda3/envs/python3/lib/python3.10/site-packages/bertopic/_bertopic.py:820, in BERTopic.topics_over_time(self, docs, timestamps, topics, nr_bins, datetime_format, evolution_tuning, global_tuning)
    818 if global_tuning:
    819     selected_topics = [all_topics_indices[topic] for topic in documents_per_topic.Topic.values]
--> 820     c_tf_idf = (global_c_tf_idf[selected_topics] + c_tf_idf) / 2.0
    822 # Extract the words per topic
    823 words_per_topic = self._extract_words_per_topic(words, selection, c_tf_idf, calculate_aspects=False)

File ~/anaconda3/envs/python3/lib/python3.10/site-packages/scipy/sparse/_index.py:77, in IndexMixin.__getitem__(self, key)
     75         return self._get_arrayXint(row, col)
     76     elif isinstance(col, slice):
---> 77         return self._get_arrayXslice(row, col)
     78 else:  # row.ndim == 2
     79     if isinstance(col, INT_TYPES):

File ~/anaconda3/envs/python3/lib/python3.10/site-packages/scipy/sparse/_csr.py:216, in _csr_base._get_arrayXslice(self, row, col)
    214     col = np.arange(*col.indices(self.shape[1]))
    215     return self._get_arrayXarray(row, col)
--> 216 return self._major_index_fancy(row)._get_submatrix(minor=col)

File ~/anaconda3/envs/python3/lib/python3.10/site-packages/scipy/sparse/_compressed.py:711, in _cs_matrix._major_index_fancy(self, idx)
    708 np.cumsum(row_nnz, out=res_indptr[1:])
    710 nnz = res_indptr[-1]
--> 711 res_indices = np.empty(nnz, dtype=idx_dtype)
    712 res_data = np.empty(nnz, dtype=self.dtype)
    713 csr_row_index(M, indices, self.indptr, self.indices, self.data,
    714               res_indices, res_data)

ValueError: negative dimensions are not allowed

The same error happens on other visualizations, such as topic_model.visualize_hierarchy().

What am I missing here?
Thank you.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the save/load sequence using BERTopic.load and then call topics_over_time and visualize_hierarchy. Start in _bertopic.py around topics_over_time line 820 and compare the loaded model state with the in-memory model. Done means the saved model can be loaded and its visualizations run without the reported ValueError.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-visualization, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.