CodeYourFuture / CodeYourFuture/Changes

Fine-tune an extractive Question Answer model on our own dataset

Offen
#9 8 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen
Resources needed Skills Tech Ed Time Trainees
Vorherrschende Sprache
JavaScript
Sterne
3
Forks
4
PR-Merge-Kennzahlen
Keine gemergten PRs in 30 T.

Beschreibung

## I want to do this:

I want to fine-tune an extractive QA model on a CodeYourFuture dataset pulled from our docs and our Slack conversations.

## Here’s why I want to do it:

I want to provide a SMALL hosted model (probably an autotrain on our Huggingface org to make it as approachable as possible) that trainees can build interfaces to interact with
I want to provide an exciting but _achievable_ final project for trainees
I want to create a "taster" of AI and ML that is relevant to our course and illuminating for trainees, who have expressed a lot of interest in this emerging field
I want to reduce the burden on staff of constantly answering the same 5 questions over and over on Slack

## Here’s how it serves our goals:

All our trainees need lots more practice asking good questions and evaluating the answers; projects that create spaces for dialogic pedagogy are needed
_We Believe in Collective Intelligence_
An exciting final project should improve performance in hiring: _good jobs in tech_

## This is how much time I can put towards this change:

I have [stubbed a dataset](https://docs.google.com/spreadsheets/d/153_f9WHn3KN6X5ok_M2eUMEKMoYYtRAys_qYPnrLZJ0/edit?usp=sharing) to think about this more. I can spend up to 6 1 hour sessions on this
I have reached out to some ML/AI experts I know to get advice. I will spend up to 4 2 hours sessions getting advice
I have created an organisation on [Huggingface](https://huggingface.co/CodeYourFuture) -- please join! https://huggingface.co/CodeYourFuture
I have drafted (SUPER DRAFTY) an e[xample final project](https://docs.google.com/document/d/1Oc_1eoNS0ZIRNpAoaBMjBaj70pJoPo2Edj6WNfcqFOE/edit?usp=sharing) that interacts with this hypothetical model

## This is the help I need from others to get this done (if any):

- ML devs to evaluate and instruct on the preparation of data
- trainees and interns to prepare data (approx 3 hours each)
- Any interested parties to build small proof of concepts to share
- Possibly a sponsor to fund training cycles if we find they are needed

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Rechercherichtung

Beginnen Sie mit der Platzhalter-Datensatz-Tabelle, der CodeYourFuture Hugging Face-Organisation und dem Entwurf des Dokuments zum Abschlussprojekt. Bestimmen Sie zunächst gemeinsam mit den genannten ML-Expert:innen die Anforderungen an Datenaufbereitung und Modelltraining; abgeschlossen wäre die Aufgabe, wenn ein vereinbarter Datensatz, ein trainiertes, gehostetes Extractive-QA-Modell und ein Proof of Concept für die Nutzung durch Trainees vorliegen.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
huggingface, machine-learning
Bereich
ai, data, machine-learning
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.