google / google/uncertainty-baselines
Reproducibility of Rank-1 results after the OSS release: investigate a potential gap and/or update the README.
- Dominant language
- Python
- Stars
- 1.6k
- Forks
- 224
- Avg merge
- 15h 36m
- Merged PRs (30d)
- 2
Description
The current table in the README presents results for Rank-1 BNNs that date back to April/May which :
* used a different evaluation log-likelihood
* possibly had a slightly different CIFAR data loader
* mostly were 3 to 5-seed averages instead of our current standard of 10-seed averages
The latest results with Gaussian Rank-1 BNNs are slightly underperforming (by an order of 0.05%) on CIFAR-10 while outperforming on CIFAR-100 (by an order of 0.1-0.4%) and require investigating or simply updating the README.
Contributor guide
Research direction
Start with the README table and compare its Rank-1 BNN results with the latest Gaussian Rank-1 results. Check the evaluation log-likelihood, CIFAR data loader, and seed counts mentioned in the issue; done means the discrepancy is explained and the README is updated if needed.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, tensorflow
- Domain
- documentation, machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 30/100