adobe / adobe/NLP-Cube

Classical Chinese Model needed

Open
#100 31 comments 0 reactions 0 assignees View on GitHub
enhancement help wanted
Dominant language
HTML
Stars
562
Forks
92
PR merge metrics
No merged PRs in 30d

Description

I've almost finished to build up [UD_Classical_Chinese-Kyoto](https://github.com/UniversalDependencies/UD_Classical_Chinese-Kyoto/tree/dev) Treebank, and now I'm trying to make a Classical Chinese model for NLP-Cube (please check my [diary](https://srad.jp/~yasuoka/journal/629704/)). But in my model sentence_accuracy<35 and I can't sentencize "天平二年正月十三日萃于帥老之宅申宴會也于時初春令月氣淑風和梅披鏡前之粉蘭薰珮後之香加以曙嶺移雲松掛羅而傾盖夕岫結霧鳥封縠而迷林庭舞新蝶空歸故鴈於是盖天促膝飛觴忘言一室之裏開衿煙霞之外淡然自放快然自足若非翰苑何以攄情詩紀落梅之篇古今夫何異矣宜賦園梅聊成短詠" (check gold standard [here](https://srad.jp/~yasuoka/journal/629612/)). How do I tune up sentencization for Classical Chinese?

Contributor guide

Open the contributing guide

Research direction

Start by reading the linked UD_Classical_Chinese-Kyoto treebank and the linked gold-standard example, then inspect how NLP-Cube currently evaluates sentence accuracy for this text. Done means identifying a reproducible way to tune Classical Chinese sentencization so the quoted passage is segmented correctly and sentence_accuracy exceeds 35.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.