ProjectTech4DevAI / ProjectTech4DevAI/kaapi-backend
Classification: AI peer matching experiment
Personne n'a encore pris cette issue.
- Langage dominant
- Python
- Étoiles
- 18
- Forks
- 10
- Merge moyen
- 2 j 20 h
- PR mergées (30 j)
- 14
Description
Is your feature request related to a problem?
Deodar's Use Case 1 (submission cleanup) is on hold due to low volume. The real issue is Use Case 2: classifying writers for peer matching, as new writers need credible feedback and peer groups of similar skill. The challenge is whether AI can classify 50–100+ writers reliably.
Describe the solution you'd like
- Assemble a dummy set of ~30 short stories (good/middling/bad) with guidelines.
- Experiment with AI by:
- Providing samples and guidelines to the AI for organic bucketing.
- Comparing AI's buckets with Deodar's.
- Asking AI to propose a rubric and provide scoring and feedback.
- Use prompt engineering without model training; iterate the rules for improvement.
- Ensure existing AI Assessments pipeline is utilized for classification tasks.
- Kaapi to assist with prompt structure and initial rounds, and provide access for self-iteration afterwards.
Original issue
Context
Deodar's Use Case 1 (submission cleanup) is parked — volume (~700–800/year) doesn't justify AI. The real problem is Use Case 2: classifying writers for peer matching. New writers need credible feedback and want peer groups at or above their own skill. Deodar can bucket 30–40 stories by hand; the question is whether AI can do this reliably at 50–100+ writers. The AI's job is classification at the entry point only — assign a writer to the right room; everything after is human-to-human.
Consent blocker & workaround
Deodar needs to take permission from writers at submission and the stories are the writers' own product, so real submissions can't be sent. Workaround: Deodar assembles a dummy set of ~30 short stories (good/middling/bad, free to share) plus a written guideline (not a rubric) on what makes writing good/bad and what characterises Indian fiction.
First experiment
- Give the AI the 30 samples + guideline; let it bucket organically into top/middle/bottom.
- Compare its buckets against Deodar's.
- Ask the AI to propose its own rubric; score and give feedback per story; sample-check; iterate.
- No model training — entirely prompt engineering (3–6 page prompts workable). First round will underperform; value is in iterating the rules.
- Platform fit: the existing AI Assessments pipeline works (opinionated toward assessment, but classification uses the same rubric-in/scored-buckets-out mechanism). Kaapi stays involved for 2–3 iterations, then hands Deodar a UI to self-iterate.
Notes
- Product shape (login → upload → AI feedback emailed; gated persona → room assignment) is exploratory, not committed. Platform must disclose AI is the first-level reader.
- Volume assumptions (50–100 simultaneous writers) are aspirational; market viability unvalidated; no internal deadline.
Next steps (Kaapi)
- Help structure the prompt and rubric; run the first rounds jointly; provide self-serve platform access once early rounds show promise.
Guide de contribution
Ouvrir le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Piste de recherche
Commencez par le pipeline existant de AI Assessments et comprenez comment il accepte des directives ou des grilles d’évaluation et produit des buckets notés. Constituez le jeu dummy proposé d’environ 30 histoires courtes ainsi que sa directive d’écriture, puis exécutez le bucketing organique initial et comparez-le aux classifications de Deodar. Le travail sera considéré comme terminé lorsque vous aurez documenté si l’expérience est suffisamment fiable pour justifier une itération sur le prompt et la grille d’évaluation.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- machine-learning, python
- Domaine
- backend, machine-learning
- Type d'issue
- Fonctionnalité
- Difficulté
- 5/5
- Temps estimé
- Plus d'une semaine
- Activité
- Calme
- Clarté
- À clarifier
- Accessibilité débutants
- 25/100