abhijithneilabraham / abhijithneilabraham/Aksharam

Create corpora for Malayalam

未关闭
#1 0 条评论 0 个 reaction 已指派 2 人 已被 @AugustinJose1221 认领 在 GitHub 查看
enhancement good first issue help wanted
主要语言
Jupyter Notebook
星标
1
派生
2
PR 合并指标
30 天内没有已合并 PR

描述

Hint: Take Nltk corpora(brown,for example) and understand how it is defined

https://www.nltk.org/book/ch02.html

Use the corpora to tokenize malayalam words.

贡献指南

这个仓库没有索引到贡献指南

调研方向

Start with the linked NLTK book chapter 2 and compare how the Brown corpus is structured. Then inspect the Aksharam notebooks or existing corpus/tokenization entry points to see where a Malayalam corpus would fit. Done would mean the project has a Malayalam corpus definition that can be used to tokenize Malayalam words.

由索引模型根据 Issue 内容生成。

评估

技术栈
jupyter-notebook, python
领域
data, internationalization, machine-learning
Issue 类型
功能
难度
4/5
预计耗时
3-5 天
活跃度
停滞
描述清晰度
需要澄清
新手友好度
25/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。