tensorflow / tensorflow/datasets

[data request] YouTube 8M

Open
#18 11 comments 11 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

dataset request
Dominant language
Python
Stars
4.6k
Forks
1.6k
Avg merge
3h 54m
Merged PRs (30d)
1

Description

  • Name of dataset: YouTube 8M
  • URL of dataset: https://research.google.com/youtube8m/
  • License of dataset: Creative Commons Attribution 4.0 International (CC BY 4.0) license.
  • Short description of dataset and use case(s): YouTube-8M is a large-scale labeled video dataset that consists of millions of YouTube video IDs, with high-quality machine-generated annotations from a diverse vocabulary of 3,800+ visual entities. Can be used for large-scale video understanding, representation learning, noisy data modeling, transfer learning, and domain adaptation approaches for video, and multi-task learning.

Folks who would also like to see this dataset in tensorflow/datasets, please +1/thumbs-up so the developers can know which requests to prioritize.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reviewing the YouTube-8M dataset details at the linked URL and the tensorflow/datasets repository context. Determine the scope and requirements for adding this dataset, including how its video IDs, annotations, and CC BY 4.0 license would be represented; done means the dataset is supported in TFDS.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, tensorflow
Domain
data, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.