huggingface / huggingface/audio-transformers-course
Happy New Year greetings and request for new topics in the course
- Dominant language
- MDX
- Stars
- 521
- Forks
- 155
- Avg merge
- 6m
- Merged PRs (30d)
- 1
Description
Good day @sanchit-gandhi , @MKhalusova and the whole course team! Congratulations on the new year 2024!
First of all, I would like to thank you and the entire course team once again for the work done. The course turned out to be very good and allowed many of us to immerse ourselves in the topic of working with sound, some of us even managed to find a job thanks to your course. Especially liked the practical orientation of the course and the good and quite accessible presentation of the theoretical material. On behalf of all the members of the course translation team - "Thank you so much for your work!".
Over the past two months, people who have taken the course in Russian have been leaving feedback on topics they would like to see covered in the course. Together with Sergey (@Lightmourne), we systematized this feedback, which eventually became our request in this issue. If possible, could you further address the following topics in the course?
List of topics:
1. Audio data preparation (broadly defined)
2. Finding partial duplicates (duplication by time-shifting the audio) and full duplicate audio (filtering the dataset before training classification models). A common case of filtering datasets of 1 second duration.
3. increasing the volume of audio data
4. determining the need for class balancing for different tasks and models in the audio domain. Examples of class balancing for audio data, methods and techniques.
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue proposes four new course topics: audio data preparation, duplicate detection, audio augmentation, and class balancing. Review the existing course structure and decide whether these topics fit its scope; done would require an agreed content plan rather than a specific code or test change.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100