CodeYourFuture / CodeYourFuture/Changes

Fine-tune an extractive Question Answer model on our own dataset

Aperta
#9 8 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Resources needed Skills Tech Ed Time Trainees
Lingua principale
JavaScript
Stelle
3
Fork
4
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

## I want to do this:

I want to fine-tune an extractive QA model on a CodeYourFuture dataset pulled from our docs and our Slack conversations.

## Here’s why I want to do it:

I want to provide a SMALL hosted model (probably an autotrain on our Huggingface org to make it as approachable as possible) that trainees can build interfaces to interact with
I want to provide an exciting but _achievable_ final project for trainees
I want to create a "taster" of AI and ML that is relevant to our course and illuminating for trainees, who have expressed a lot of interest in this emerging field
I want to reduce the burden on staff of constantly answering the same 5 questions over and over on Slack

## Here’s how it serves our goals:

All our trainees need lots more practice asking good questions and evaluating the answers; projects that create spaces for dialogic pedagogy are needed
_We Believe in Collective Intelligence_
An exciting final project should improve performance in hiring: _good jobs in tech_

## This is how much time I can put towards this change:

I have [stubbed a dataset](https://docs.google.com/spreadsheets/d/153_f9WHn3KN6X5ok_M2eUMEKMoYYtRAys_qYPnrLZJ0/edit?usp=sharing) to think about this more. I can spend up to 6 1 hour sessions on this
I have reached out to some ML/AI experts I know to get advice. I will spend up to 4 2 hours sessions getting advice
I have created an organisation on [Huggingface](https://huggingface.co/CodeYourFuture) -- please join! https://huggingface.co/CodeYourFuture
I have drafted (SUPER DRAFTY) an e[xample final project](https://docs.google.com/document/d/1Oc_1eoNS0ZIRNpAoaBMjBaj70pJoPo2Edj6WNfcqFOE/edit?usp=sharing) that interacts with this hypothetical model

## This is the help I need from others to get this done (if any):

- ML devs to evaluate and instruct on the preparation of data
- trainees and interns to prepare data (approx 3 hours each)
- Any interested parties to build small proof of concepts to share
- Possibly a sponsor to fund training cycles if we find they are needed

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

Inizia dal foglio di calcolo del dataset provvisorio, dall’organizzazione CodeYourFuture su Hugging Face e dalla bozza del documento del progetto finale. Determina innanzitutto i requisiti per la preparazione dei dati e l’addestramento del modello con gli esperti di ML menzionati; il lavoro sarebbe completato quando fossero disponibili un dataset concordato, un modello di QA estrattivo addestrato e ospitato e una proof of concept da utilizzare per i trainees.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
huggingface, machine-learning
Ambito
ai, data, machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.