cockroachdb / cockroachdb/cockroach
roachtest: monitor_failure failed
- Dominant language
- Go
- Stars
- 32.5k
- Forks
- 4.1k
- PR merge metrics
- PR metrics pending
Description
roachtest.monitor_failure [failed](https://teamcity.cockroachdb.com/buildConfiguration/Cockroach_Nightlies_RoachtestNightlyAwsBazel/16794955?buildTab=log) with [artifacts](https://teamcity.cockroachdb.com/buildConfiguration/Cockroach_Nightlies_RoachtestNightlyAwsBazel/16794955?buildTab=artifacts#/restore/tpce/8TB/aws/nodes=10/cpus=8) on release-24.1 @ [74b732729e074694be98190cc41eef6f845c3df3](https://github.com/cockroachdb/cockroach/commits/74b732729e074694be98190cc41eef6f845c3df3):
```
test restore/tpce/8TB/aws/nodes=10/cpus=8 failed: (monitor.go:154).Wait: monitor failure: pq: query execution canceled
EOF [owner=test-eng]
test artifacts and logs in: /artifacts/restore/tpce/8TB/aws/nodes=10/cpus=8/run_1
```
Parameters:
- ROACHTEST_arch=amd64
- ROACHTEST_cloud=aws
- ROACHTEST_coverageBuild=false
- ROACHTEST_cpu=8
- ROACHTEST_encrypted=false
- ROACHTEST_fs=ext4
- ROACHTEST_localSSD=false
- ROACHTEST_metamorphicBuild=false
- ROACHTEST_ssd=0
Help
See: [roachtest README](https://github.com/cockroachdb/cockroach/blob/master/pkg/cmd/roachtest/README.md)
See: [How To Investigate \(internal\)](https://cockroachlabs.atlassian.net/l/c/SSSBr8c7)
_Grafana is not yet available for aws clusters_
Same failure on other branches
- #127634 roachtest: monitor_failure failed [O-roachtest O-robot T-testeng X-infra-flake branch-release-24.1.3-rc]
- #126792 roachtest: monitor_failure failed [O-roachtest O-robot T-testeng X-infra-flake branch-master]
/cc @cockroachdb/test-eng
[This test on roachdash](https://roachdash.crdb.dev/?filter=status:open%20t:.*monitor_failure.*&sort=title+created&display=lastcommented+project) | [Improve this report!](https://github.com/cockroachdb/cockroach/tree/master/pkg/cmd/bazci/githubpost/issues)
Jira issue: CRDB-41979
Contributor guide
Research direction
Start with the roachtest README and the failure site at monitor.go:154, then inspect the logs and artifacts under /artifacts/restore/tpce/8TB/aws/nodes=10/cpus=8/run_1. Compare the related failures in #127634 and #126792 to determine whether this is the same recurring failure. Done means identifying the cause and validating a rerun of the restore test without the monitor failure.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- aws, go
- Domain
- databases, infrastructure, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100