MaartenGr / MaartenGr/BERTopic
Issues about the heatmap generated by BERTopic
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 7.8k
- Forks
- 920
- Avg merge
- 22h 24m
- Merged PRs (30d)
- 5
Description
BERTopic generate topics by clustering semantically similar clusters of documents. However, in the heatmap, some topcis have high simiarty scores (e.g., 0.8 or above). Why are these topics with a high similarity split into separate topics?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the reported heatmap behavior in BERTopic and inspect how the heatmap represents topic similarity. Determine why topics with similarity scores around 0.8 or higher remain separate, then document the explanation or identify the behavior that needs correction.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data-visualization, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100