LAION-AI / LAION-AI/project-menu

data extraction project ideas

Open
#9 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
No language data
Stars
12
Forks
4
PR merge metrics
No merged PRs in 30d

Description

Ideas:

  • extract text+image+audio link+video link from pages ; could be valuable for future multimodal gpt3-like models
  • extract text content+images from pages : much more similar to image+text dataset but that would already provide some more context for images
  • extract embeddings+metadata from videos from torrents
  • extract embeddings+metadata from audio from torrents

Each of these ideas could be fleshed out in these ways:

  • build a POC of data extraction, estimate the amount of data available
  • imagine in more details models that could benefit
  • estimate the infra needed to do the extraction
  • describe how the distribution of the dataset could be done
  • check any legal issues

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

No files, tests, or entry points are mentioned. Start by selecting one extraction idea, then scope its proof of concept, estimate available data and infrastructure, identify useful models, outline dataset distribution, and check legal issues; done means one idea has a concrete, self-contained proposal.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.