caption text at training
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 2.3k
- Forks
- 213
- PR merge metrics
- No merged PRs in 30d
Description
Hi
I am using music_audioset_epoch_15_esc_90.14.pt as a music classifier. I would like to classify the mood and genre of our music files. I am trying to find the cosine similarity using the text "The mood of this song is (romantic, energetic, etc)" but I only get about 0.4. I think that if I use a text similar to the one you used in your training, the value will be better, so could you please tell me what type of text you used?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reviewing the documentation and training configuration associated with music_audioset_epoch_15_esc_90.14.pt and the CLAP training setup. Confirm whether the text prompt or caption format used during training is documented; done means providing the exact text or clearly stating that it is unavailable.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 30/100