huggingface / huggingface/datatrove

FastText: newer model

Open
#321 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.3k
Forks
302
Avg merge
2h 18m
Merged PRs (30d)
2

Description

I was curious why the older model of fastText is used by default instead of the NLLB version that covers more languages (https://huggingface.co/facebook/fasttext-language-identification). Is it because of its non-commercial license?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.