channable / channable/opsqueue
CI keeps failing due to timeout
- Lenguaje dominante
- Rust
- Estrellas
- 96
- Forks
- 2
- Merge medio
- 1 h 2 min
- PR fusionados (30 d)
- 2
Descripción
When i open a PR, i often have to run the job a couple of times before it completes in time.
We've had issues before with our integration tests being flaky and becoming deadlocked indefinitely. #6 introduced the 20 second timeout for integration tests, so they fail more quickly once they become deadlocked.
It's also difficult to see what exactly causes the failure. A first step towards resolving this issue could be to see what we can do to improve the output that we get when a test times out on CI, because at time of writing this is just a massive stack trace with mostly callsites originating from pytest plugins and the like.
Either the time-out is just too short for CI, or we're still getting deadlocks. So far, i've not been able to really reproduce any deadlocks by running the tests locally. We could try to relax that timeout a bit more, but not by too much. As the comment above the timeout configuration states, we're already being quite generous with our time limit.
**Examples**
The last four PRs have all seen at least one failure of the integration tests due to it exceeding the time limit:
- #19
- [failed push](https://channable.semaphoreci.com/workflows/56fef772-2c93-4009-a632-fa1b51fe5aa4)
- #18
- [failed push](https://channable.semaphoreci.com/jobs/ef466f0d-4f1c-4e1c-9e8f-883d47d118c5)
- [failed re-run](https://channable.semaphoreci.com/jobs/c496f290-f42b-46ea-9605-13204f6fcae5)
- #17
- [failed push](https://channable.semaphoreci.com/jobs/ed000520-408d-4222-a674-02833521bdf9)
- #16
- [failed merge by Hoff](https://github.com/channable/opsqueue/pull/16#issuecomment-3376018058)
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Comienza con la configuración del timeout de las pruebas de integración mencionada en la issue e inspecciona los trabajos de CI vinculados para los PRs #16–#19. Compara la salida del timeout con las ejecuciones locales para determinar si los fallos indican deadlocks o un límite de CI insuficiente. Se considera terminado cuando se haya identificado la causa y los fallos de timeout produzcan una salida útil para actuar o se haya realizado un ajuste del timeout debidamente justificado.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- rust
- Área
- ci-cd, testing
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 35/100