google-research / google-research/circuit_training

CT Tool for Ariane-133: Snapshot generation halt

Open
#76 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.7k
Forks
273
PR merge metrics
No merged PRs in 30d

Description

Hello, I wanted to ask about the Ariane-133 design (from MacroPlacement). I ran the CT tool for 11 days (Intel Xeon, 132GB RAM, and I used 3 collect jobs, no GPUs), however it seemed that the tool stopped generating snapshots after day two. On Tensorboard, I also noticed that the losses plateaued around day two. Is there a reason for this? Additionally, is there a measure in place to know when the tool is done? It seems like it finished generating snapshots but continued to run. Should it stop at a certain point? Lastly, I am curious as to why the checkpoints directory was empty (no checkpoints created, even after 31k steps).

Thank you so much!

snapshot results taken on 01/26/25:
![Image](https://github.com/user-attachments/assets/cadc9bd7-b152-47e9-be50-68b370923a84)

Tensorboard Results:
plateau occurs after 1.782 days:
![Image](https://github.com/user-attachments/assets/86263652-b976-4ff5-a215-28eac54e7970)

all jobs ended manually on day 11:
![Image](https://github.com/user-attachments/assets/8767c448-7bbd-432a-ab1a-898763d852fe)

![Image](https://github.com/user-attachments/assets/5c898429-21e9-41d6-a4c8-8545ca0e147f)

![Image](https://github.com/user-attachments/assets/c2ad669b-7fa3-4340-82c7-3caa95d705b0)

![Image](https://github.com/user-attachments/assets/31ae5d38-6ac4-4038-b054-1d6ba57b06c4)

![Image](https://github.com/user-attachments/assets/7f1206ef-f076-452e-8d4e-2e1f275765c1)

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the Ariane-133 run with the CT tool and reviewing the TensorBoard output, snapshot results, and checkpoints directory described in the report. Trace snapshot generation, checkpoint creation, and the run's stopping condition. Done means explaining the plateau and missing checkpoints, and documenting or fixing how completion is detected.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.