MaartenGr / MaartenGr/BERTopic

[inhomogeneous shape unresolved] [Colab] ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (2,) + inhomogeneous part.

Open
#1,951 1 comment 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
7.8k
Forks
920
Avg merge
22h 24m
Merged PRs (30d)
5

Description

Using BERTopic for some research task, I am getting:
`ValueError: setting an array element with a sequence. The requested array has an inhomogeneous shape after 1 dimensions. The detected shape was (2,) + inhomogeneous part`

Is there plan for `BERTopic` supporting `numpy >1.23.5` and `numba > 0.56.4`?
This will make it easier to use BERTopic in Colab without having to downgrade persistently manually (`!pip install numba==0.56.4`): especially for those without Colab.pro
Thanks.

[**Noted**]
Related issue: [**closed**][1] | [#1697, #1602, #1309]
Related issue: [**open**][2] | [#1814, #1799, #1684, #1584, #1421, #1269]
**SO**: https://stackoverflow.com/a/76504825

PS: I gather that a walkaround is to downgrade numpy!
in issue #1421: @aaron-imani hinted at `np.average()` in `_guided_topic_modelling()`. @MaartenGR indicated that setting `numba` to `0.56.4` or earlier should ideally fix the issue

Thanks @MaartenGr for your insightful engagement in previous issues and suggestions.

**Environment**: `colab.research.google.com`
NB: (as at 26 April 2024, Colab runs `python: 3.10.12, bertopic: 0.16.1, numpy: 1.25.2, numba: 0.58.1`)

**BERTopic**:
```python
# Install BERTopic
!pip install bertopic
# Load libraries
from bertopic import BERTopic
from bertopic.representation import KeyBERTInspired
from bertopic.vectorizers import ClassTfidfTransformer
from umap import UMAP
from sentence_transformers import SentenceTransformer
```
[Code snippet]
```python
# Train BERTopic

topic_model_03 = BERTopic(
verbose=True,
min_topic_size= 5, #10, #12, #15, ## the higher, the lower the clusters/topics
nr_topics = 4, #5, ## reduce the initial number of topics
#seed_topic_list=seed_topic_list, ##ValueError: ...
n_gram_range = (1,3), ## n-gram range for the CountVectorizer
zeroshot_topic_list=zeroshot_topic_list,
zeroshot_min_similarity=.35, #.55, #.75, #.85,
embedding_model="thenlper/gte-small", ## pass string directly to sbert sentence-transformers models
#umap_model = umap_model, ## dimensionality
ctfidf_model = ctfidf_model,
representation_model=KeyBERTInspired(),
)
```
The seed topics look like this {redacted in part}
```python
### Try with Guided Representation
seed_topic_list = [["sustainability", "...", "sustain", "...", "..."],
["climate change", "climate", "...", "...", "...", "...", "ozone"],
["social justice", "social", "...", "...", "...", "...", "...", "..."],
["net zero", "..."]
]

## fit and transform
#topics_03, probs_03 = topic_model_03.fit_transform(docs)
## visualise documents' topics spread
#topic_model_03.visualize_documents(docs)
```

[1]: https://github.com/MaartenGr/BERTopic/issues?q=is%3Aissue+inhomogeneous+is%3Aclosed
[2]: https://github.com/MaartenGr/BERTopic/issues?q=is%3Aissue+inhomogeneous+is%3Aopen

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the BERTopic.fit_transform example in Colab with the reported Python, BERTopic, NumPy, and numba versions. Read related issues #1421, #1584, #1684, #1799, and #1814; done means identifying a supported dependency combination or a confirmed BERTopic change that removes the need to downgrade NumPy or numba.

Written by the indexing model from the issue text.

Assessment

Tech stack
numpy, python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.