huggingface / huggingface/datatrove
Exploring documents in different clusters
Open
- Dominant language
- Python
- Stars
- 3.3k
- Forks
- 302
- Avg merge
- 2h 18m
- Merged PRs (30d)
- 2
Description
I see that there is an option to save the cluster ids. But, when I read *.clusters file for the original documents file, I see that the number of documents do not match. I see lesser number of documents. For each document, I want the cluster id so that I can view the documents that fall into the same cluster.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.