baidu / baidu/DuReader

是否考虑构建一个“检索+阅读”的中文openQA数据集?

Open
#83 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.2k
Forks
306
PR merge metrics
No merged PRs in 30d

Description

openQA会根据问题,从知识库(百万量级以上的文本)中检索相关的文本,然后进行“阅读”以抽取出问题的答案。目前openQA的数据集主要都是英文的,如:NaturalQuestions、WebQuestions。

dureader其实可以在现有的基础上,整理出一版针对openQA任务的数据集,构建一个中文 openQA的榜单,这将对中文openQA的发展很有帮助。想问下有这个计划吗?谢谢~

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue proposes creating a Chinese openQA dataset and leaderboard from DuReader, but names no files, tests, or entry points. Start by reviewing the existing DuReader dataset structure and project plans; completion would require a defined dataset and benchmark scope, which the issue does not specify.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.