是否考虑构建一个“检索+阅读”的中文openQA数据集?
Open
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 306
- PR merge metrics
- No merged PRs in 30d
Description
openQA会根据问题,从知识库(百万量级以上的文本)中检索相关的文本,然后进行“阅读”以抽取出问题的答案。目前openQA的数据集主要都是英文的,如:NaturalQuestions、WebQuestions。
dureader其实可以在现有的基础上,整理出一版针对openQA任务的数据集,构建一个中文 openQA的榜单,这将对中文openQA的发展很有帮助。想问下有这个计划吗?谢谢~
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue proposes creating a Chinese openQA dataset and leaderboard from DuReader, but names no files, tests, or entry points. Start by reviewing the existing DuReader dataset structure and project plans; completion would require a defined dataset and benchmark scope, which the issue does not specify.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100