UniDic 2.1.2 support for Japanese Tokenizer (Kuromoji) [LUCENE-5169]
Open
legacy-jira-priority:Minor
module:analysis
type:enhancement
- Dominant language
- Java
- Stars
- 3.6k
- Forks
- 1.4k
- Avg merge
- 2d 11h
- Merged PRs (30d)
- 88
Description
I made some amendments to support UniDic 2.1.2 into kuromoji.
The attached patch is against lucene_solr_4_4 branch.
---
Migrated from [LUCENE-5169](https://issues.apache.org/jira/browse/LUCENE-5169) by Adrien Grand (@jpountz), updated Jun 22 2021
Attachments: [unidic.patch](https://apache.github.io/lucene-jira-archive/attachments/LUCENE-5169/unidic.patch)
Contributor guide
Research direction
Start by reviewing the attached unidic.patch and the Kuromoji tokenizer changes on the lucene_solr_4_4 branch. Trace how UniDic data is integrated, then verify that UniDic 2.1.2 is supported without regressing existing Japanese tokenization behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- search
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100