CodeYourFuture / CodeYourFuture/Changes
Fine-tune an extractive Question Answer model on our own dataset
- Lingua principale
- JavaScript
- Stelle
- 3
- Fork
- 4
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
## I want to do this:
I want to fine-tune an extractive QA model on a CodeYourFuture dataset pulled from our docs and our Slack conversations.
## Here’s why I want to do it:
I want to provide a SMALL hosted model (probably an autotrain on our Huggingface org to make it as approachable as possible) that trainees can build interfaces to interact with
I want to provide an exciting but _achievable_ final project for trainees
I want to create a "taster" of AI and ML that is relevant to our course and illuminating for trainees, who have expressed a lot of interest in this emerging field
I want to reduce the burden on staff of constantly answering the same 5 questions over and over on Slack
## Here’s how it serves our goals:
All our trainees need lots more practice asking good questions and evaluating the answers; projects that create spaces for dialogic pedagogy are needed
_We Believe in Collective Intelligence_
An exciting final project should improve performance in hiring: _good jobs in tech_
## This is how much time I can put towards this change:
I have [stubbed a dataset](https://docs.google.com/spreadsheets/d/153_f9WHn3KN6X5ok_M2eUMEKMoYYtRAys_qYPnrLZJ0/edit?usp=sharing) to think about this more. I can spend up to 6 1 hour sessions on this
I have reached out to some ML/AI experts I know to get advice. I will spend up to 4 2 hours sessions getting advice
I have created an organisation on [Huggingface](https://huggingface.co/CodeYourFuture) -- please join! https://huggingface.co/CodeYourFuture
I have drafted (SUPER DRAFTY) an e[xample final project](https://docs.google.com/document/d/1Oc_1eoNS0ZIRNpAoaBMjBaj70pJoPo2Edj6WNfcqFOE/edit?usp=sharing) that interacts with this hypothetical model
## This is the help I need from others to get this done (if any):
- ML devs to evaluate and instruct on the preparation of data
- trainees and interns to prepare data (approx 3 hours each)
- Any interested parties to build small proof of concepts to share
- Possibly a sponsor to fund training cycles if we find they are needed
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Direzione di ricerca
Inizia dal foglio di calcolo del dataset provvisorio, dall’organizzazione CodeYourFuture su Hugging Face e dalla bozza del documento del progetto finale. Determina innanzitutto i requisiti per la preparazione dei dati e l’addestramento del modello con gli esperti di ML menzionati; il lavoro sarebbe completato quando fossero disponibili un dataset concordato, un modello di QA estrattivo addestrato e ospitato e una proof of concept da utilizzare per i trainees.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- huggingface, machine-learning
- Ambito
- ai, data, machine-learning
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100