dice-group / dice-group/gerbil
Each dataset should have Language as a feature
Open
type:enhancement
- Dominant language
- Java
- Stars
- 231
- Forks
- 55
- PR merge metrics
- No merged PRs in 30d
Description
Implement language as a feature of a dataset so we can move forward towards multilingual benchmarking of annotators.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the dataset model and the benchmarking entry point. Read how existing dataset features are represented, then define the expected language value and verify that multilingual annotator benchmarking can consume it; done means a dataset exposes language and the relevant benchmark behavior is covered.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- java
- Domain
- data, internationalization
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 32/100