cloudfoundry / cloudfoundry/capi-release
cc.api_post_start_healthcheck_timeout_in_seconds doesn't seem to matter
Nobody has claimed this yet.
- Dominant language
- HTML
- Stars
- 24
- Forks
- 110
- Avg merge
- 3d 12h
- Merged PRs (30d)
- 8
Description
Issue
The cc.api_post_start_healthcheck_timeout_in_seconds field doesn't seem to have any effect.
Context
The cc.api_post_start_healthcheck_timeout_in_seconds config field seems to imply that the post-start will fail if CCNG fails report healthy within that time frame:
However, we pushed a changed to CCNG that slept for 10 hours before starting the thin server (in the runner), and this timeout did not get triggered. Instead, 20 minutes passed until both ccng_monit_http_healthcheck and nginx_cc failed.
Task 277 | 22:03:30 | L starting jobs: api/abf448d3-74e1-4086-80c2-130814373e14 (0) (canary) Task 277 | 22:03:57 | Updating instance scheduler: scheduler/7d611e5a-1ada-41a4-b811-49b5dbcb2b2f (0) (canary) (00:02:09)
Task 277 | 22:23:32 | Updating instance api: api/abf448d3-74e1-4086-80c2-130814373e14 (0) (canary) (00:21:44)
L Error: 'api/abf448d3-74e1-4086-80c2-130814373e14 (0)' is not running after update. Review logs for failed jobs: ccng_monit_http_healthcheck, nginx_cc
Task 277 | 22:23:32 | Error: 'api/abf448d3-74e1-4086-80c2-130814373e14 (0)' is not running after update. Review logs for failed jobs: ccng_monit_http_healthcheck, nginx_cc
So it seems this check in the post start script:
doesn't seem to matter anymore.
However, trying this experiment on older versions of CAPI, we observer that the deploy fails in the post-start script after the configured time.
We believe this regression was introduced in this PR: https://github.com/cloudfoundry/capi-release/pull/195/files, which reconfigured the monit dependencies between the processes.
What we are not clear on is if this is a regression that reintroduced the issue which prompted the introduction of that post-start check in the first place: https://github.com/cloudfoundry/capi-release/issues/125.
Steps to Reproduce
Add a long sleep to this line: https://github.com/cloudfoundry/cloud_controller_ng/blob/main/lib%2Fcloud_controller%2Frunner.rb#L88 and deploy
Expected result
The post-start script to fail after the configured cc.api_post_start_healthcheck_timeout_in_seconds value
Current result
Deploy fails after 20 minutes.
Possible Fix
Not sure, is this something we need to address?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with jobs/cloud_controller_ng/templates/post-start.sh.erb around line 71 and the spec entries around lines 349-351; compare the monit dependency changes in PR #195 with issue #125. Reproduce by adding a long sleep in cloud_controller_ng/lib/cloud_controller/runner.rb around line 88 and deploy, then verify whether the configured timeout or the 20-minute healthcheck failure controls the result.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- ruby, shell
- Domain
- devops, infrastructure
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100