InternLM / InternLM/StarBench

Test results of BAT

Open
#6 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
44
Forks
4
PR merge metrics
No merged PRs in 30d

Description

I found that for BAT, the perception part has valid scores but reasoning part is 0.00

I wonder why? Because BAT is never trained with MCQA, so it can not generate valid answers, therefore it would always be 0.00, how can you get valid numbers for BAT?

Thanks.

Contributor guide

No contributing guide indexed for this repository

Research direction

No file or test entry point is named. Start by locating the BAT evaluation path and the MCQA scoring logic, then reproduce the reported 0.00 reasoning result and determine whether it is expected for an untrained model or indicates an evaluation problem.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, testing
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Needs clarification
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.