Proposal: Open Low-Resource Language Corpus for Culturally Grounded LLM Research
- Lingua principale
- Python
- Stelle
- 622
- Fork
- 45
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hello AI2 team,
I would like to propose a new corpus initiative focused on under-resourced languages and culturally grounded language understanding.
I am building an open, AI-ready corpus that begins with Western Armenian literary and historical texts and will later expand to additional Armenian and Turkish literary sources. Beyond digitization, the project is designed to preserve linguistic diversity, dialectal variation, historical language change, and culturally specific knowledge that is often absent from current large language model training and evaluation resources.
The corpus will include high-quality OCR correction, structured metadata, linguistic annotations, and cultural context annotations, making it suitable for research on low-resource languages, multilingual NLP, and LLM evaluation. The long-term goal is to provide an open resource that helps researchers study how AI systems understand historically and culturally rich texts from underrepresented language communities.
Before preparing a contribution, I would appreciate your feedback on whether this type of resource aligns with AI2's goals for open language resources and whether there are recommended standards or contribution pathways for integrating such datasets into the broader ecosystem.
Thank you for your time and consideration. I look forward to your feedback.
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.