google / google/uncertainty-baselines

Reproducibility of Rank-1 results after the OSS release: investigate a potential gap and/or update the README.

Open
#128 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.6k
Forks
224
Avg merge
15h 36m
Merged PRs (30d)
2

Description

The current table in the README presents results for Rank-1 BNNs that date back to April/May which :
* used a different evaluation log-likelihood
* possibly had a slightly different CIFAR data loader
* mostly were 3 to 5-seed averages instead of our current standard of 10-seed averages

The latest results with Gaussian Rank-1 BNNs are slightly underperforming (by an order of 0.05%) on CIFAR-10 while outperforming on CIFAR-100 (by an order of 0.1-0.4%) and require investigating or simply updating the README.

Contributor guide

Open the contributing guide

Research direction

Start with the README table and compare its Rank-1 BNN results with the latest Gaussian Rank-1 results. Check the evaluation log-likelihood, CIFAR data loader, and seed counts mentioned in the issue; done means the discrepancy is explained and the README is updated if needed.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, tensorflow
Domain
documentation, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.