iterative / iterative/cml.dev

When should we update dvc.lock and dvc push in GitHub Actions?

Aperta
#393 5 commenti 1 reazione 0 assegnatari Vedi su GitHub
documentation question
Lingua principale
TypeScript
Stelle
13
Fork
22
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

### Discussed in https://github.com/iterative/dvc/discussions/6542

Originally posted by **Hongbo-Miao** September 6, 2021
Currently, my GitHub Actions workflow looks like this: when I open a pull request (change some model codes / params), CML creates a AWS EC2 instance, and DVC pull the data.

Here is my [current GitHub Actions workflow](https://github.com/Hongbo-Miao/hongbomiao.com/blob/078e097d71213f9a0e24b884c4d65fd78bf0ccfd/.github/workflows/cml.yaml#L53-L129):

Click to expand!

```yaml
cml-cloud-set-up-cloud:
name: CML (Cloud) - Set up cloud
runs-on: ubuntu-20.04
steps:
- name: Cancel previous runs
uses: styfle/cancel-workflow-action@0.9.1
with:
access_token: ${{ github.token }}
- name: Checkout
uses: actions/checkout@v2
- name: Set up CML
uses: iterative/setup-cml@v1
- name: Set up cloud
shell: bash
env:
REPO_TOKEN: ${{ secrets.CML_ACCESS_TOKEN }}
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
run: |
cml-runner \
--cloud=aws \
--cloud-region=us-west-2 \
--cloud-type=t2.small \
--labels=cml-runner
cml-cloud-train:
name: CML (Cloud) - Train
needs: cml-cloud-set-up-cloud
runs-on: [self-hosted, cml-runner]
# container: docker://iterativeai/cml:0-dvc2-base1-gpu
container: docker://iterativeai/cml:0-dvc2-base1
steps:
- name: Cancel previous runs
uses: styfle/cancel-workflow-action@0.9.1
with:
access_token: ${{ github.token }}
- name: Checkout
uses: actions/checkout@v2
- name: Set up Miniconda
uses: conda-incubator/setup-miniconda@v2
with:
miniconda-version: "latest"
activate-environment: hm-cnn
- name: Install requirements
working-directory: convolutional-neural-network
shell: bash -l {0}
run: |
conda install pytorch torchvision torchaudio --channel=pytorch
conda install pandas
conda install tabulate
pip install -r requirements.txt
- name: Pull Data
working-directory: convolutional-neural-network
env:
AWS_ACCESS_KEY_ID: ${{ secrets.AWS_ACCESS_KEY_ID }}
AWS_SECRET_ACCESS_KEY: ${{ secrets.AWS_SECRET_ACCESS_KEY }}
run: |
dvc pull
- name: Train model
working-directory: convolutional-neural-network
shell: bash -l {0}
env:
WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
run: |
dvc repro
- name: Write CML report
working-directory: convolutional-neural-network
shell: bash -l {0}
env:
REPO_TOKEN: ${{ secrets.CML_ACCESS_TOKEN }}
run: |
echo "# CML (Cloud) Report" >> report.md
echo "## Params" >> report.md
cat output/reports/params.txt >> report.md
cml-send-comment report.md
```

My [dvc.yaml](https://github.com/Hongbo-Miao/hongbomiao.com/blob/078e097d71213f9a0e24b884c4d65fd78bf0ccfd/convolutional-neural-network/dvc.yaml) looks like this:

```yaml
stages:
prepare:
cmd: tar -xf data/raw/cifar-10-python.tar.gz --dir=data/processed
deps:
- data/raw/cifar-10-python.tar.gz
outs:
- data/processed/cifar-10-batches-py/
main:
cmd: python main.py
deps:
- data/processed/cifar-10-batches-py/
- evaluate.py
- main.py
- model/
- train.py
params:
- lr
- train.epochs
outs:
- output/models/model.pt
```

After training, if I think the change is good because the performance is better based on the reports,

- the [dvc.lock](https://github.com/Hongbo-Miao/hongbomiao.com/blob/078e097d71213f9a0e24b884c4d65fd78bf0ccfd/convolutional-neural-network/dvc.lock) I feel needs to get update.
- the new model `model.pt` needs to be uploaded to AWS S3 in my case.

My question is, after `dvc repro`, am I supposed to add `dvc push` and then commit? Something like

```yaml
- name: Train model
working-directory: convolutional-neural-network
shell: bash -l {0}
env:
WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
run: |
dvc repro
dvc push # New added
git add . # New added
git commit -m "update dvc.lock, etc." # New added
git push origin current_pr # New added, need somehow get the current pull request name
```

This above method will apply when the pull request is open.
However, I kind of feeling the best moment adding would be when I decide merging because I think this is a good pull request that actually improves the machine learning performance. But at this moment, the EC2 instance has been destroyed.

Any suggestion? Thanks!

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Review the linked GitHub Actions workflow, dvc.yaml, and dvc.lock first. Trace when dvc repro changes the lockfile and when model.pt is produced, then determine which workflow stage should own committing and pushing artifacts. Done means the issue has a documented, unambiguous workflow for pull requests and merges.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
aws, github-actions
Ambito
ci-cd, devops
Tipo di issue
Documentazione
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.