Example notebook with Datalab on a text dataset with sota LLM & embeddings model
Open
help wanted
- Dominant language
- Jupyter Notebook
- Stars
- 136
- Forks
- 28
- PR merge metrics
- No merged PRs in 30d
Description
Make a version of this tutorial: https://docs.cleanlab.ai/stable/tutorials/datalab/text.html
but using more modern ML models. `pred_probs` can be produced by a (pretrained) LLM, and `features` produced via a recently popular Embeddings model.
Recommend using models from HuggingFace. Try to select a dataset where the detected issues are interesting, particularly one where the under-performing group issue is present.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.