Contributing a deep-learning, BERT-based analyzer
- Dominant language
- Java
- Stars
- 3.6k
- Forks
- 1.4k
- Avg merge
- 2d 11h
- Merged PRs (30d)
- 88
Description
### Description
Hi,
We are building an open-source custom Hebrew/Arabic analyzer (lemmatizer and stopwords), based on a BERT model. We'd like to contribute this to this repository. How can we do that and be accepted? Can we compile it to native code and use JNI or [Panama ](https://openjdk.org/projects/panama/)? If not, what is the best approacch?
https://github.com/apache/lucene/issues/12502#issuecomment-1675084211
@uschindler would be very happy to hear what you think
Contributor guide
Research direction
Start with the linked Lucene issue discussion and determine whether the proposed Hebrew/Arabic BERT-based analyzer has an accepted integration path, including JNI or Panama. Done means a maintainer-approved contribution plan with clear acceptance criteria; this issue names no source files, entry points, or tests.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- search
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100