abhijithneilabraham / abhijithneilabraham/Aksharam

Create corpora for Malayalam

Open
#1 0 comments 0 reactions 2 assignees Claimed by @AugustinJose1221 View on GitHub
enhancement good first issue help wanted
Dominant language
Jupyter Notebook
Stars
1
Forks
2
PR merge metrics
No merged PRs in 30d

Description

Hint: Take Nltk corpora(brown,for example) and understand how it is defined

https://www.nltk.org/book/ch02.html

Use the corpora to tokenize malayalam words.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the linked NLTK book chapter 2 and compare how the Brown corpus is structured. Then inspect the Aksharam notebooks or existing corpus/tokenization entry points to see where a Malayalam corpus would fit. Done would mean the project has a Malayalam corpus definition that can be used to tokenize Malayalam words.

Written by the indexing model from the issue text.

Assessment

Tech stack
jupyter-notebook, python
Domain
data, internationalization, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.