actions / actions/actions-runner-controller

Runners don't scale down if there are any `num_terminating_busy` replicas

Open
#2,039 7 comments 1 reaction 0 assignees View on GitHub

Nobody has claimed this yet.

bug
Dominant language
Go
Stars
6.5k
Forks
1.5k
Avg merge
2d 2h
Merged PRs (30d)
27

Description

Checks
Controller Version

0.26.0

Helm Chart Version

0.21.0

CertManager Version

1.8.0

Deployment Method

Helm

cert-manager installation

Checks
  • This isn't a question or user support case (For Q&A and community support, go to Discussions. It might also be a good idea to contract with any of contributors and maintainers if your business is so critical and therefore you need priority support
  • I've read releasenotes before submitting this issue and I'm sure it's not due to any recently-introduced backward-incompatible changes
  • My actions-runner-controller version (v0.x.y) does support the feature
  • I've already upgraded ARC (including the CRDs, see charts/actions-runner-controller/docs/UPGRADING.md for details) to the latest and it didn't fix the issue
  • I've migrated to the workflow job webhook event (if you using webhook driven scaling)
Resource Definitions
apiVersion: apps/v1
kind: Deployment
metadata:
  annotations:
    deployment.kubernetes.io/revision: "11"
    meta.helm.sh/release-name: actions-runner-controller
    meta.helm.sh/release-namespace: actions-runner-system
  labels:
    app.kubernetes.io/instance: actions-runner-controller
    app.kubernetes.io/managed-by: Helm
    app.kubernetes.io/name: actions-runner-controller
    app.kubernetes.io/version: 0.26.0
    helm.sh/chart: actions-runner-controller-0.21.0
  name: actions-runner-controller
  namespace: actions-runner-system
spec:
  progressDeadlineSeconds: 600
  replicas: 1
  revisionHistoryLimit: 10
  selector:
    matchLabels:
      app.kubernetes.io/instance: actions-runner-controller
      app.kubernetes.io/name: actions-runner-controller
  strategy:
    rollingUpdate:
      maxSurge: 25%
      maxUnavailable: 25%
    type: RollingUpdate
  template:
    metadata:
      annotations:
        cluster-autoscaler.kubernetes.io/safe-to-evict: "true"
        kubectl.kubernetes.io/default-logs-container: manager
      labels:
        app.kubernetes.io/instance: actions-runner-controller
        app.kubernetes.io/name: actions-runner-controller
    spec:
      containers:
      - args:
        - --metrics-addr=127.0.0.1:8080
        - --enable-leader-election
        - --port=9443
        - --sync-period=5m
        - --default-scale-down-delay=10m
        - --docker-image=docker:dind
        - --runner-image=summerwind/actions-runner:latest
        command:
        - /manager
        env:
        - name: GITHUB_TOKEN
          valueFrom:
            secretKeyRef:
              key: github_token
              name: controller-manager
              optional: true
        - name: GITHUB_APP_ID
          valueFrom:
            secretKeyRef:
              key: github_app_id
              name: controller-manager
              optional: true
        - name: GITHUB_APP_INSTALLATION_ID
          valueFrom:
            secretKeyRef:
              key: github_app_installation_id
              name: controller-manager
              optional: true
        - name: GITHUB_APP_PRIVATE_KEY
          valueFrom:
            secretKeyRef:
              key: github_app_private_key
              name: controller-manager
              optional: true
        - name: GITHUB_BASICAUTH_PASSWORD
          valueFrom:
            secretKeyRef:
              key: github_basicauth_password
              name: controller-manager
              optional: true
        image: summerwind/actions-runner-controller:v0.26.0
        imagePullPolicy: IfNotPresent
        name: manager
        ports:
        - containerPort: 9443
          name: webhook-server
          protocol: TCP
        resources: {}
        securityContext: {}
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
        volumeMounts:
        - mountPath: /etc/actions-runner-controller
          name: secret
          readOnly: true
        - mountPath: /tmp
          name: tmp
        - mountPath: /tmp/k8s-webhook-server/serving-certs
          name: cert
          readOnly: true
      - args:
        - --secure-listen-address=0.0.0.0:8443
        - --upstream=http://127.0.0.1:8080/
        - --logtostderr=true
        - --v=10
        image: quay.io/brancz/kube-rbac-proxy:v0.13.0
        imagePullPolicy: IfNotPresent
        name: kube-rbac-proxy
        ports:
        - containerPort: 8443
          name: metrics-port
          protocol: TCP
        resources: {}
        securityContext: {}
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
      dnsPolicy: ClusterFirst
      restartPolicy: Always
      schedulerName: default-scheduler
      securityContext: {}
      serviceAccount: actions-runner-controller
      serviceAccountName: actions-runner-controller
      terminationGracePeriodSeconds: 10
      volumes:
      - name: secret
        secret:
          defaultMode: 420
          secretName: controller-manager
      - name: cert
        secret:
          defaultMode: 420
          secretName: actions-runner-controller-serving-cert
      - emptyDir: {}
        name: tmp
--
apiVersion: actions.summerwind.dev/v1alpha1
kind: HorizontalRunnerAutoscaler
metadata:
  annotations:
    kubectl.kubernetes.io/last-applied-configuration: |
      {"apiVersion":"actions.summerwind.dev/v1alpha1","kind":"HorizontalRunnerAutoscaler","metadata":{"annotations":{},"name":"evi-platform-study-dev-runner","namespace":"actions-runner-system"},"spec":{"maxReplicas":5,"metrics":[{"scaleDownAdjustment":1,"scaleDownThreshold":"0.3","scaleUpAdjustment":1,"scaleUpThreshold":"0.75","type":"PercentageRunnersBusy"}],"minReplicas":1,"scaleDownDelaySecondsAfterScaleOut":600,"scaleTargetRef":{"name":"evi-platform-study-dev-runner"}}}
  name: evi-platform-study-dev-runner
  namespace: actions-runner-system
spec:
  maxReplicas: 5
  metrics:
  - scaleDownAdjustment: 1
    scaleDownThreshold: "0.3"
    scaleUpAdjustment: 1
    scaleUpThreshold: "0.75"
    type: PercentageRunnersBusy
  minReplicas: 1
  scaleDownDelaySecondsAfterScaleOut: 600
  scaleTargetRef:
    name: evi-platform-study-dev-runner
To Reproduce
1. Run some sort of ill-behaved job on the runner that will be terminated but end up stuck in `Terminating`.  Sorry, I haven't got to the bottom of what's going on with this part, but I understand from the troubleshooting FAQ that this can happen.
2. Drive jobs such that runner replicas are scaled beyond the minimum.
3. Wait for jobs to complete.
4. Observe that runners don't scale down.
Describe the bug

Although there were no non-terminating busy runners, desired replicas remained at 5. Once the Terminating pods were removed by removing their finalizers, scaledown to the reserved limit occurred in the next cycle.

2022-11-23T01:51:10Z	DEBUG	actions-runner-controller.horizontalrunnerautoscaler	Suggested desired replicas of 5 by PercentageRunnersBusy	{"replicas_desired_before": 5, "replicas_desired": 5, "num_runners": 5, "num_runners_registered": 5, "num_runners_busy": 0, "num_terminating_busy": 2, "namespace": "actions-runner-system", "kind": "runnerdeployment", "name": "evi-platform-study-dev-runner", "horizontal_runner_autoscaler": "evi-platform-study-dev-runner", "enterprise": "evidation-health", "organization": "", "repository": ""}
2022-11-23T01:51:10Z	DEBUG	actions-runner-controller.horizontalrunnerautoscaler	Calculated desired replicas of 5	{"horizontalrunnerautoscaler": "actions-runner-system/evi-platform-study-dev-runner", "suggested": 5, "reserved": 0, "min": 1, "max": 5}
Describe the expected behavior

Even with stuck Terminating pods, idle runners are scaled down.

Whole Controller Logs
https://gist.github.com/oeuftete/f83e1d0ff1efb2197071712e9c17ce6c
Whole Runner Pod Logs
https://gist.github.com/oeuftete/6b92499085d3712b18d25b6d151138a4
Additional Context

No response

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the HorizontalRunnerAutoscaler PercentageRunnersBusy calculation and the controller log fields shown in the report, especially num_terminating_busy and num_runners_busy. Reproduce the case with stuck Terminating pods and verify that idle runners scale down toward the minimum even when terminating busy replicas remain.

Written by the indexing model from the issue text.

Assessment

Tech stack
github-actions, go, kubernetes
Domain
infrastructure
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.