LAION-AI / LAION-AI/project-menu
data extraction project ideas
Open
Nobody has claimed this yet.
- Dominant language
- No language data
- Stars
- 12
- Forks
- 4
- PR merge metrics
- No merged PRs in 30d
Description
Ideas:
- extract text+image+audio link+video link from pages ; could be valuable for future multimodal gpt3-like models
- extract text content+images from pages : much more similar to image+text dataset but that would already provide some more context for images
- extract embeddings+metadata from videos from torrents
- extract embeddings+metadata from audio from torrents
Each of these ideas could be fleshed out in these ways:
- build a POC of data extraction, estimate the amount of data available
- imagine in more details models that could benefit
- estimate the infra needed to do the extraction
- describe how the distribution of the dataset could be done
- check any legal issues
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are mentioned. Start by selecting one extraction idea, then scope its proof of concept, estimate available data and infrastructure, identify useful models, outline dataset distribution, and check legal issues; done means one idea has a concrete, self-contained proposal.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100