InternLM / InternLM/Agent-FLAN

Question about evaluation datasets

Open
#14 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
362
Forks
10
PR merge metrics
No merged PRs in 30d

Description

Hey,

Great observations and work on disentangling the format following from reasoning! Could we share details on evaluation dataset we used and how we can reproduce the result in the paper? I have fine tuned llama3 on the dataset and achieved worse performance in 30 questions curated from HotpotQA dataset. If you could share some light on this it would be super appreciated! Thanks,
Jason

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.