alibaba / alibaba/FederatedScope

Smaller test/val loss but lower evaluation accuracy

Open
#750 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.5k
Forks
261
PR merge metrics
No merged PRs in 30d

Description

When I finetune llama-7b on gsm-8k with different finetuning methods. I compared the test loss and evaluation accuracy of different methods and found that one of the method has smaller test/val loss but lower evaluation accuracy. Is it reasonable?

Contributor guide

No contributing guide indexed for this repository

Research direction

No files or tests are named. Start by reproducing the llama-7b fine-tuning comparison on gsm-8k across the reported methods, then compare the test/validation loss with evaluation accuracy. Done means establishing whether the discrepancy is expected or identifying a reproducible implementation problem.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.