LAION-AI / LAION-AI/CLAP

did anyone train on your own dataset and got good performance?

Open
#76 8 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
2.3k
Forks
213
PR merge metrics
No merged PRs in 30d

Description

did anyone train on your own dataset and got good performance?

Hi, I wanna train on my own data but seems there is some dependence on datasets like "batch['url']" in train.py: https://github.com/LAION-AI/CLAP/blob/main/src/laion_clap/training/train.py#L326

how the 'url' can be set in my own data? what is this stand for(FOR class label?)? could u please give me an example?

thanks for your time!

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with src/laion_clap/training/train.py around line 326 and trace how batch['url'] is consumed during training. Check the repository's dataset-loading examples or documentation for the expected custom-data fields. Done means documenting whether url is required, what it represents, and showing a working example for a user-owned dataset.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.