huggingface / huggingface/transfer-learning-conv-ai

How are the distractors made in the dataset?

Open
#77 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
430
PR merge metrics
No merged PRs in 30d

Description

I want to use my own custom dataset with this project, but I don't understand how the distractors were made in the original dataset to get a grasp on how to do this. Are they randomly sampled from other conversations?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by examining how the original dataset and its distractors are represented in the project. Determine whether distractors are randomly sampled from other conversations or generated another way, then document that process so someone using a custom dataset can reproduce it.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data, machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.