cockroachdb / cockroachdb/cockroach

roachtest: monitor_failure failed

Open
#130,292 6 comments 0 reactions 0 assignees View on GitHub
branch-release-24.1 O-roachtest O-robot T-testeng X-infra-flake
Dominant language
Go
Stars
32.5k
Forks
4.1k
PR merge metrics
PR metrics pending

Description

roachtest.monitor_failure [failed](https://teamcity.cockroachdb.com/buildConfiguration/Cockroach_Nightlies_RoachtestNightlyAwsBazel/16794955?buildTab=log) with [artifacts](https://teamcity.cockroachdb.com/buildConfiguration/Cockroach_Nightlies_RoachtestNightlyAwsBazel/16794955?buildTab=artifacts#/restore/tpce/8TB/aws/nodes=10/cpus=8) on release-24.1 @ [74b732729e074694be98190cc41eef6f845c3df3](https://github.com/cockroachdb/cockroach/commits/74b732729e074694be98190cc41eef6f845c3df3):

```
test restore/tpce/8TB/aws/nodes=10/cpus=8 failed: (monitor.go:154).Wait: monitor failure: pq: query execution canceled
EOF [owner=test-eng]
test artifacts and logs in: /artifacts/restore/tpce/8TB/aws/nodes=10/cpus=8/run_1
```

Parameters:
- ROACHTEST_arch=amd64
- ROACHTEST_cloud=aws
- ROACHTEST_coverageBuild=false
- ROACHTEST_cpu=8
- ROACHTEST_encrypted=false
- ROACHTEST_fs=ext4
- ROACHTEST_localSSD=false
- ROACHTEST_metamorphicBuild=false
- ROACHTEST_ssd=0
Help

See: [roachtest README](https://github.com/cockroachdb/cockroach/blob/master/pkg/cmd/roachtest/README.md)

See: [How To Investigate \(internal\)](https://cockroachlabs.atlassian.net/l/c/SSSBr8c7)

_Grafana is not yet available for aws clusters_

Same failure on other branches

- #127634 roachtest: monitor_failure failed [O-roachtest O-robot T-testeng X-infra-flake branch-release-24.1.3-rc]
- #126792 roachtest: monitor_failure failed [O-roachtest O-robot T-testeng X-infra-flake branch-master]

/cc @cockroachdb/test-eng

[This test on roachdash](https://roachdash.crdb.dev/?filter=status:open%20t:.*monitor_failure.*&sort=title+created&display=lastcommented+project) | [Improve this report!](https://github.com/cockroachdb/cockroach/tree/master/pkg/cmd/bazci/githubpost/issues)

Jira issue: CRDB-41979

Contributor guide

Open the contributing guide

Research direction

Start with the roachtest README and the failure site at monitor.go:154, then inspect the logs and artifacts under /artifacts/restore/tpce/8TB/aws/nodes=10/cpus=8/run_1. Compare the related failures in #127634 and #126792 to determine whether this is the same recurring failure. Done means identifying the cause and validating a rerun of the restore test without the monitor failure.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, go
Domain
databases, infrastructure, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.