facebookresearch / facebookresearch/fairo
Feedback on writeup for Dashboard in Turk
- Dominant language
- Jupyter Notebook
- Stars
- 929
- Forks
- 123
- PR merge metrics
- No merged PRs in 30d
Description
## Type of Issue
Select the type of issue:
- [ ] Bug report (to report a bug)
- [ ] Feature request (to request an additional feature)
- [ ] Tracker (I am just using this as a tracker)
- [ ] Refactor request
- [ ] Documentation Ask
## Checklist
Few things that come to mind when I look at the task from a Turker's perspective:
- [ ] The description is too big, and not consumable for a Turker who has to work through the tasks quickly.
- [ ] If Turkers aren't marking failures, the first thing that comes to mind is - they aren't even getting to read the "Mark the error" section. Have you tried any experiments with rearranging sections ? Has that helped ?
- [x] We need to really cut down the description - I still see a major part of the old writeup (from first version) here. Have you tried running tasks with smaller list of capabilities allowed for example ?
- [x] I think we really should be running some basic qualification tasks to prune out the noise before doing any goal oriented / targeted experiments. This will also ensure we don't need to narrate the full story for each task and we can cut down on the first part by a lot.
- [x] I see a discrepancy between : ".Click "start" in the bottom right to begin session. Interact with the bot for full 4 minutes." vs "We ask that you interact with the agent for at least 5 minutes." Note that these are small things but can really confuse Turkers from my experience.
- [x] `The bot might be slow to respond due to network lag. You may expect a few seconds between sending a command and getting responses from the agent (we are improving it!).`-> Also mention that the agent will do stuff in the 3-D grid world and it doesn't just respond back for every command
- [x] To avoid overwhelming folks, we could also skip this : On the bottom left you will be able to see a history of the last 5 commands you’ve sent to the assistant in the game.
- [x] This list can be reduced: `You can assume that the bot has the following capabilities:`
- [ ] We need to explain more on what "Marking Errors" means.
- [ ] General guidelines are:
- [ ] Turk task real estate is expensive, we should use it very wisely. If the explanation/ content is getting too big , we need to think twice - should this still be one task or can this be split ?
- [ ] More examples of successful task is way more helpful than plain text of explanation - for example, here we can add some examples of commands that can be given (and ask them to not repeat same commands), and even give some example of commands , how those were errors. We can cover few examples of erroneous commands covering different kinds of errors.
- [ ] Things we really want to emphasize on can be repeated but nothing else.
- [ ] We really need to iterate fast on the description and this should be guided by learnings from the data coming in.
Contributor guide
Assessment
This issue has not been assessed yet.