allenai / allenai/staged-training

Empirical estimate

オープン
#5 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
Jupyter Notebook
スター
33
フォーク
2
PR マージ指標
30日以内にマージされた PR はありません

説明

Thank you for your great work! It is very enlightening to me, but I still struggled to figure out some details.

The most puzzling part for me is still how to empirically estimate thresholds of `PRE-GROWTH` and `OPTIMALITY`. I probably caught that we should train a small original model and a small target model, identify the three points, and calculate ${\mathcal{T}}\_{G}$ , ${\mathcal{T}}\_{opt}$ , and $\rho$.

But how are these three points located? Is the `PRE-GROWTH` where the slope of the original model (loss curve) is about to fall below that of the target model? However, it does not appear to be the case as depicted in Figures 1 and 5. Similarly, what is the basis for determining the `OPTIMALITY` point?

I would be grateful if you could help clarify.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。