Call-for-Code-for-Racial-Justice / Call-for-Code-for-Racial-Justice/TakeTwo-DataScience
Determining the unit of analysis for the machine learning models
- 主要語言
- Jupyter Notebook
- 星號
- 8
- 分支
- 8
- PR 合併指標
- 30 天內沒有已合併 PR
描述
For the [chrome extension](https://github.com/Call-for-Code-for-Racial-Justice/TakeTwo-Marker-ChromeExtension), we need to decide what is the unit or type of selections trusted contributors can select when identifying racially biased content. This will be used to train the ML models and is important to help users understand why content is potentially racially biased and to offer up alternatives.
Considerations
* If content is considered racially biased, context will be a factor. How can we store more data around the selected word / phrase to provide more context for the ML models?
* At what granularity should the trusted contributors be expected to flag content? (Categorically vs binary?)
* Consider the input as a paragraph/group of text and all kind of tokenizations can be used to finally preprocess
貢獻指南
研究方向
先從 issue 中關於選取粒度、上下文儲存、類別標籤與二元標籤的差異,以及段落層級輸入的問題開始。查看連結的 Chrome 擴充功能,了解受信任的貢獻者如何選取內容。對於訓練 ML 模型所使用的分析單位、需要保留的上下文以及標註方法達成共識,即表示完成。
由索引模型根據 Issue 內容生成。
評估
- 技術堆疊
- machine-learning
- 領域
- data, machine-learning
- Issue 類型
- 功能
- 難度
- 5/5
- 預估耗時
- 一週以上
- 活躍度
- 停滯
- 描述清晰度
- 需要釐清
- 新手友好度
- 25/100