dice-group / dice-group/gerbil

Each dataset should have Language as a feature

Open
#35 0 comments 0 reactions 0 assignees View on GitHub
type:enhancement
Dominant language
Java
Stars
231
Forks
55
PR merge metrics
No merged PRs in 30d

Description

Implement language as a feature of a dataset so we can move forward towards multilingual benchmarking of annotators.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by locating the dataset model and the benchmarking entry point. Read how existing dataset features are represented, then define the expected language value and verify that multilingual annotator benchmarking can consume it; done means a dataset exposes language and the relevant benchmark behavior is covered.

Written by the indexing model from the issue text.

Assessment

Tech stack
java
Domain
data, internationalization
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
32/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.