EducationalTestingService / EducationalTestingService/rsmtool
Kappa computation when predicted scores are on a different scale
- 主要言語
- Python
- スター
- 71
- フォーク
- 21
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
In a situation where predicted scores are on a completely different scale from the observed scores, kappa computation fails because the range of possible scores is too large. We saw this recently with SGDRegressor which produced scores in the range 1233372304332.22 to 1723509896207.16 when human scores were 1-5.
Few possible solutions:
(1) Do not compute kappa if there is no overlap in range between predicted and observed scores.
(2) Do not compute kappa if the range is greater than certain threshold.
Thoughts?
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
調査の方向性
ファイル、テスト、エントリーポイントのいずれも指定されていません。まず kappa の計算箇所を特定し、1233372304332.22–1723509896207.16 前後の予測スコアと 1–5 の観測スコアで失敗を再現してください。重複しない、または過度に大きいスコア範囲に対する合意済みのポリシーが実装され、失敗がテストでカバーされれば完了です。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python, scikit-learn
- 領域
- analytics, machine-learning
- issue の種類
- バグ
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 停滞
- 明瞭さ
- 説明が足りない
- 初心者へのやさしさ
- 25/100