EleutherAI / EleutherAI/pythia
pythia-2.8b: 86 of 154 step branches (incl. step0 and the entire early grid) serve final-era checkpoint content
- Dominant language
- Jupyter Notebook
- Stars
- 2.9k
- Forks
- 224
- PR merge metrics
- No merged PRs in 30d
Description
While extracting embedding matrices across the full Pythia checkpoint grid (HTTP range reads only — no full downloads), we found that the pythia-2.8b training history is largely not step-specific.
## File-level evidence (LFS sha256 of `model.safetensors`, via the HF API)
- `main` = `ab496f1c3fd79e3c749a9d5414136a2c8e4224f94eecb261970315cdb0f813fe`
- **76 step branches are bit-identical to `main`**: steps 0, 1, 2, 4, 8, 16, 32, 64, 128, 256, 512; every 1000 from 1000 through 53000; plus 55000, 60000, 65000, 70000, 80000, 90000, 100000, 103000, 110000, 120000, 130000, 140000.
- A **second blob** (`462f2b960062159c…`) is served by **10 branches**: 54000, 56000, 57000, 58000, 59000, 61000, 62000, 63000, 64000, and 143000.
- The remaining **68 branches** (non-round steps 66000–142000) each serve a unique file.
Full branch → sha256 table attached (`upstream_2p8b_branch_hashes.csv`).
## Tensor-level evidence (embedding byte ranges read from every branch)
- All **86** branches in the two duplicated groups serve **byte-identical `embed_in` and `embed_out` tensors** (re-fetching `step0` and `step143000` returns identical bytes; e.g. row-0 norm 0.889188, abs-sum 2012294.875 on both). The two blobs differ somewhere outside the embedding ranges.
- The 68 unique files carry genuinely distinct, evolving embeddings consistent with late-training checkpoints of the same run (cosine ≥ 0.86 to the shared snapshot).
## The shared snapshot is end-of-training content, not an init
Its never-used vocabulary rows (padding ids ≥ 50277) have median L2 norm 0.19436, i.e. ~0.308 of the init norm (~0.632). Under decoupled AdamW with the published 2.8b schedule (peak LR 1.6e-4, cosine to 1.6e-5, weight decay 0.1), unused rows decay as exp(−λτ), predicting an end-of-training ratio of 0.284 — an ~8% match. A genuine step-0 branch would sit at ratio 1.0 and genuine early steps at ≥0.97. So the branches labelled step0…step53000 (and all round steps) serve **final-era** content; `step143000`'s embeddings equal `main`'s and may themselves be fine.
## Impact
Any research using pythia-2.8b intermediate checkpoints (learning dynamics, interpretability-over-training) is silently analyzing one end-of-training snapshot for the init, the entire early grid, and every round-numbered step. Genuine history appears to survive only on the 68 non-round branches 66000–142000.
## Reproduction (no downloads required)
1. `GET https://huggingface.co/api/models/EleutherAI/pythia-2.8b/tree/` and compare `lfs.oid` of `model.safetensors` across branches.
2. For tensor identity: HTTP range-read `gpt_neox.embed_in.weight` (offsets from the safetensors header) from any two same-hash branches and compare bytes.
Happy to share the extraction code and the full per-branch statistics.
[upstream_2p8b_branch_hashes.csv](https://github.com/user-attachments/files/31050445/upstream_2p8b_branch_hashes.csv)
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the attached upstream_2p8b_branch_hashes.csv and the Hugging Face model.safetensors tree API, then reproduce the LFS hash comparison and HTTP range reads for gpt_neox.embed_in.weight across the listed branches. Done means confirming the duplicated checkpoint content and documenting the affected branches and an agreed remediation path for researchers.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- huggingface
- Domain
- data, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 35/100