Proposal: Open Low-Resource Language Corpus for Culturally Grounded LLM Research
- Linguagem predominante
- Python
- Estrelas
- 622
- Forks
- 45
- Métricas de merge de PRs
- Nenhum PR com merge em 30d
Descrição
Hello AI2 team,
I would like to propose a new corpus initiative focused on under-resourced languages and culturally grounded language understanding.
I am building an open, AI-ready corpus that begins with Western Armenian literary and historical texts and will later expand to additional Armenian and Turkish literary sources. Beyond digitization, the project is designed to preserve linguistic diversity, dialectal variation, historical language change, and culturally specific knowledge that is often absent from current large language model training and evaluation resources.
The corpus will include high-quality OCR correction, structured metadata, linguistic annotations, and cultural context annotations, making it suitable for research on low-resource languages, multilingual NLP, and LLM evaluation. The long-term goal is to provide an open resource that helps researchers study how AI systems understand historically and culturally rich texts from underrepresented language communities.
Before preparing a contribution, I would appreciate your feedback on whether this type of resource aligns with AI2's goals for open language resources and whether there are recommended standards or contribution pathways for integrating such datasets into the broader ecosystem.
Thank you for your time and consideration. I look forward to your feedback.
Guia de contribuição
Nenhum guia de contribuição indexado para este repositório
Direção de pesquisa
The issue proposes an open corpus beginning with Western Armenian literary and historical texts, with OCR correction, structured metadata, linguistic annotations, and cultural context annotations. No repository files, tests, entry points, contribution format, or acceptance criteria are identified. Start by seeking maintainer feedback on whether this resource belongs in WildDet3D and what standards or contribution pathway would apply; done would require an agreed scope and integration plan.
Escrita pelo modelo de indexação a partir do texto da issue.
Avaliação
- Domínio
- data, machine-learning
- Tipo de issue
- Funcionalidade
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Status de atividade
- Pouca atividade
- Clareza
- Precisa de esclarecimento
- Facilidade para iniciantes
- 18/100