abhijithneilabraham / abhijithneilabraham/Aksharam
Create corpora for Malayalam
未关闭
enhancement
good first issue
help wanted
- 主要语言
- Jupyter Notebook
- 星标
- 1
- 派生
- 2
- PR 合并指标
- 30 天内没有已合并 PR
描述
Hint: Take Nltk corpora(brown,for example) and understand how it is defined
https://www.nltk.org/book/ch02.html
Use the corpora to tokenize malayalam words.
贡献指南
这个仓库没有索引到贡献指南
调研方向
Start with the linked NLTK book chapter 2 and compare how the Brown corpus is structured. Then inspect the Aksharam notebooks or existing corpus/tokenization entry points to see where a Malayalam corpus would fit. Done would mean the project has a Malayalam corpus definition that can be used to tokenize Malayalam words.
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- jupyter-notebook, python
- 领域
- data, internationalization, machine-learning
- Issue 类型
- 功能
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 活跃度
- 停滞
- 描述清晰度
- 需要澄清
- 新手友好度
- 25/100