ByteDance-Seed / ByteDance-Seed/AHN

Reproducing Qwen2.5-3B performance

Open
#5 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
182
Forks
6
PR merge metrics
No merged PRs in 30d

Description

Hi there,

I just tried running the LV-Eval -128k subset on the Qwen2.5-3B-Instruct model(no AHN enabled). The performance is terribly low and can't match the number reported in paper.

So here is the screenshot. I ran the evaluation on a single A6000Ada 48G VRAM with single process mode. Could you please share the model prediction (the json files) so I can verify what goes wrong.

Image

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.