Test results of BAT
Open
- Dominant language
- Python
- Stars
- 44
- Forks
- 4
- PR merge metrics
- No merged PRs in 30d
Description
I found that for BAT, the perception part has valid scores but reasoning part is 0.00
I wonder why? Because BAT is never trained with MCQA, so it can not generate valid answers, therefore it would always be 0.00, how can you get valid numbers for BAT?
Thanks.
Contributor guide
No contributing guide indexed for this repository
Research direction
No file or test entry point is named. Start by locating the BAT evaluation path and the MCQA scoring logic, then reproduce the reported 0.00 reasoning result and determine whether it is expected for an untrained model or indicates an evaluation problem.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, testing
- Issue type
- Bug
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Active
- Clarity
- Needs clarification
- Newbie friendliness
- 45/100