allenai / allenai/olmes

Release Plan for IFEval-OOD and HREF

Aberta
#4 9 comentários 6 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
395
Forks
105
Métricas de merge de PRs
Nenhum PR com merge em 30d

Descrição

Thanks for the amazing work on both the OLMES and Tulu-3 releases.

As stated in Sections 7.3.1 and 7.3.2 of the Tulu-3 paper, AI2 created two benchmarks in-house, IFEval-OOD and HREF, to test the models, which still show a significant gap in optimal performance. I want to ask if the authors plan to release these two benchmarks to facilitate further research purposes. If yes, what is the potential timeline for that?

Many thanks!

Guia de contribuição

Nenhum guia de contribuição indexado para este repositório

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.