OpenMOSS / OpenMOSS/MOSS

多轮对话数据构造的时候是否会有上下文不一致的问题

Open
#375 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
12.3k
Forks
1.1k
PR merge metrics
No merged PRs in 30d

Description

您好 readme中提到:
moss-moon-003-sft所使用的多轮对话数据,基于MOSS-002内测阶段采集的约10万用户输入数据和gpt-3.5-turbo构造而成

请问是将user prompt输入到chatgpt中一轮一轮来增量构造的吗?
那么是否会存在用户在第二轮提的内容在gpt第一轮中没有出现过,比如下面的示例:

user: 给我写一个快排
gpt: code.....
user: 你的代码里面的quicksort函数是什么意思
gpt:对不起我之前的回答里面并没有提到quicksort这个函数

这种情况是不是上下文的语义不太统一,开源的数据里面考虑过这种问题吗?

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the README passage describing moss-moon-003-sft and the MOSS-002 data collected with gpt-3.5-turbo. Determine how the multi-turn conversations were constructed and whether later user turns remain consistent with earlier assistant responses; done means documenting the finding and any data-quality handling.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.