canonical / canonical/microcloud

Sometimes reoccuring pipeline failures

Aperta
#653 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Incomplete
Lingua principale
Go
Stelle
532
Fork
83
Merge medio
1g 14h
PR unite (30g)
36

Descrizione

This issue should track reoccurring pipeline errors which are observed from time to time.
Fixed issues should be checked:

- [x] `Error: System "micro02" failed to join the cluster: Failed to update cluster status of services: Put "https://10.69.231.229:9443/1.0/services": EOF`
Fixed in https://github.com/canonical/microcloud/pull/652
- [x] `lxc exec micro01 -- lxc launch ubuntu-minimal-daily:22.04 c1 -c limits.memory=512MiB -d root,size=2GiB -s remote -n default
Launching c1
Error: Failed instance creation: Fetch project database object: Failed to fetch from "projects" table: Failed to fetch from "projects" table: Failed to fetch from "projects" table: sql: transaction has already been committed or rolled back`
Maybe related (due to resources): https://www.google.de/url?sa=t&source=web&rct=j&opi=89978449&url=https://discuss.linuxcontainers.org/t/database-error-sql-transaction-has-already-been-committed-or-rolled-back/14464/4&ved=2ahUKEwjd5Z6xy_CLAxUogP0HHVx1ChMQFnoECBUQAQ&usg=AOvVaw001EheGLCq7CzTSkLx6j-f
* https://github.com/canonical/microcloud/actions/runs/13437321556/job/37542878760
* https://github.com/canonical/microcloud/actions/runs/13653613796/job/38167679646
* https://github.com/canonical/microcloud/actions/runs/14333304587/job/40174251902
* https://github.com/canonical/microcloud/actions/runs/14930531723/job/42055476170
* https://github.com/canonical/microcloud/actions/runs/14930554727/job/41945583317?pr=774
* https://github.com/canonical/microcloud/actions/runs/16221024008/job/45801534172?pr=865#step:3:6316
* https://github.com/canonical/microcloud/actions/runs/17234618780/job/48897915655?pr=887
* https://github.com/canonical/microcloud/actions/runs/17521340182/job/49765605620
Fixed in https://github.com/canonical/microcloud/pull/1246
- [x] `github.com/canonical/microcloud/microcloud/cmd/tui.(*InputHandler).getAllRows(0xc00031a550)` panics `panic: runtime error: index out of range [4] with length 4`
https://github.com/canonical/microcloud/actions/runs/13499838854/job/37715339968?pr=656
https://github.com/canonical/microcloud/actions/runs/13786409444/job/38555414351
https://github.com/canonical/microcloud/actions/runs/14310219002/job/40103165145?pr=733
https://github.com/canonical/microcloud/actions/runs/14598026893/job/40948866855
https://github.com/canonical/microcloud/actions/runs/14972399336/job/42056209191?pr=782
Fixed in https://github.com/canonical/microcloud/pull/799
- [ ] `Error: System "micro02" failed to join the cluster: Failed to update cluster status of services: Failed to join "LXD" cluster: Failed to configure cluster: Failed to initialize member: Failed to initialize storage pools and networks: Failed to create storage pool "remote": Failed to run: rbd --id admin --cluster ceph --pool lxd_remote info lxd_lxd_remote: signal: killed`
https://github.com/canonical/microcloud/actions/runs/13519224737/job/37774700405?pr=666
https://github.com/canonical/microcloud/actions/runs/16264171520/job/45916182877?pr=865
- [ ] `Error: Failed to check if "lxd_cephfs" CephFS exists: Failed to run: ceph --name client.admin --cluster ceph fs get lxd_cephfs: exit status 1 (Error initializing cluster client: InvalidArgumentError('RADOS invalid argument (error calling conf_read_file)'))`
https://github.com/canonical/microcloud/actions/runs/13792565713/job/38576260053
- [ ] `Error: System "micro02" failed to join the cluster: Failed to update cluster status of services: Failed to join "MicroCeph" cluster: failed to add service mgr: Failed to run: ceph auth get-or-create mgr.micro02 mon allow profile mgr osd allow * mds allow * -o /var/snap/microceph/common/data/mgr/ceph-micro02/keyring: exit status 1 (Error initializing cluster client: InvalidArgumentError('RADOS invalid argument (error calling conf_read_file)'))`
https://github.com/canonical/microcloud/actions/runs/13853126505/job/38764167452?pr=697
- [ ] `Error: Failed start validation for device "eth1": MAC address "00:16:3e:29:5e:ee" already defined on another NIC`
https://github.com/canonical/microcloud/actions/runs/14454972118/job/40536205229?pr=750#step:3:6612
- [x] ` ⨯ Error: Failed to request disk addition: Post "http://control.socket/1.0/services/microceph/1.0/disks?target=micro01": context deadline exceeded`
https://github.com/canonical/microcloud/actions/runs/16267366020/job/45926275850
https://github.com/canonical/microcloud/actions/runs/16517071789/job/46710255729?pr=893#step:3:7936
Fixed in https://github.com/canonical/microcloud/pull/894
- [x] `Error: Failed restoring network: Failed adding OVS chassis "micro01" with priority 1138 to chassis group "lxd-net2": Failed to run: ovn-nbctl --timeout=10 --db ssl:10.233.39.247:6641,ssl:10.233.39.196:6641,ssl:10.233.39.87:6641 -c /proc/self/fd/3 -p /proc/self/fd/4 -C /proc/self/fd/5 --wait=sb ha-chassis-group-add-chassis lxd-net2 micro01 1138: signal: alarm clock (2025-07-25T09:13:24Z|00003|fatal_signal|WARN|terminating with signal 14 (Alarm clock))`
https://github.com/canonical/microcloud/actions/runs/16517982655/job/46713166163?pr=894
-> Issue is already identified https://chat.canonical.com/canonical/pl/hy76o6k8dt8n9nk89rakzd1ysw
- [ ] When running e2e and deploying using Terraform, somehow the LXD cluster started to misbehave which caused `Failed to begin transaction: failed to create dqlite connection: no available dqlite leader server found`
https://github.com/canonical/microcloud/actions/runs/17099508932/job/48492273368?pr=927
- [ ] `⨯ Error: System "micro02" failed to join the cluster: Failed to update cluster status of services: Failed to join "MicroCeph" cluster: failed to add service mon: failed to retrieve mon keyring: Failed to run: ceph auth get mon. -o /tmp/3755778528/mon.keyring: exit status 1 (unable to get monitor info from DNS SRV with service name: ceph-mon`
https://github.com/canonical/microcloud/actions/runs/17466017898/job/49602152088?pr=954#step:3:27015
- [ ] `internal error, please report: running "lxd.lxc" failed: cannot create transient scope: DBus error "org.freedesktop.systemd1.TransactionIsDestructive": [Transaction for snap.lxd.lxc-14eb3434-fa1d-4ecd-bbce-19067088da76.scope/start is destructive (poweroff.target has 'start' job queued, but 'stop' is included in transaction).]`
https://github.com/canonical/microcloud/actions/runs/18712031576/job/53362757438?pr=1039
- [x] Using the new Canonical runners we sometimes see `E: Failed to fetch http://archive.ubuntu.com/ubuntu/dists/noble-updates/universe/binary-amd64/Packages.xz File has unexpected size (1500056 != 1500192). Mirror sync in progress? [IP: 185.125.190.81 80]`
https://github.com/canonical/microcloud/actions/runs/19627629308/job/56205410401
Fixed in https://github.com/canonical/microcloud/pull/1140

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Inizia dalle esecuzioni di GitHub Actions non selezionate e dai relativi log degli errori, in particolare dall’esecuzione end-to-end di Terraform, quindi raggruppa gli errori ricorrenti per sottosistema. Il lavoro è completato quando ogni errore rimanente è stato riprodotto e diagnosticato, oppure ha un fix o un issue collegato; le pull request elencate documentano i casi già risolti.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
github-actions, go, terraform, ubuntu
Ambito
ci-cd, distributed-systems, infrastructure, testing-qa
Tipo di issue
Bug
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
20/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.