allenai / allenai/staged-training

Empirical estimate

Aberta
#5 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Jupyter Notebook
Estrelas
33
Forks
2
Métricas de merge de PRs
Nenhum PR com merge em 30d

Descrição

Thank you for your great work! It is very enlightening to me, but I still struggled to figure out some details.

The most puzzling part for me is still how to empirically estimate thresholds of `PRE-GROWTH` and `OPTIMALITY`. I probably caught that we should train a small original model and a small target model, identify the three points, and calculate ${\mathcal{T}}\_{G}$ , ${\mathcal{T}}\_{opt}$ , and $\rho$.

But how are these three points located? Is the `PRE-GROWTH` where the slope of the original model (loss curve) is about to fall below that of the target model? However, it does not appear to be the case as depicted in Figures 1 and 5. Similarly, what is the basis for determining the `OPTIMALITY` point?

I would be grateful if you could help clarify.

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.