Fixed parameters for audio and text encoders
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 2.3k
- Forks
- 213
- PR merge metrics
- No merged PRs in 30d
Description
While performing contrastive pre-training on the audio-text pairs dataset, did you fix the paramters of both the encoders and optimise the parameters of MLP layer only? Also, I was wondering if you use the pre-trained audio encoder (PANN and HTSAT) or you optimise the parameters from scratch? I don't seem to find the detailed information from the paper. Great work, and thanks in advance.
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Review the paper and the contrastive pre-training setup for the audio-text pairs dataset, focusing on whether the audio and text encoders are fixed or optimized and whether PANN and HTSAT are pretrained. Done means providing a clear, documented answer to these training questions.
Written by the indexing model from the issue text.
Assessment
- Domain
- machine-learning
- Issue type
- Documentation
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100