aws / aws/amazon-cloudwatch-agent

address already in use

Open
#1,956 5 comments 3 reactions 0 assignees View on GitHub
Dominant language
Go
Stars
550
Forks
271
Avg merge
1d 21h
Merged PRs (30d)
13

Description

**Describe the bug**
Deployment on EKS (kubernetes) cluster fails (from time to time) with error:
```
failed to serve and listen","error":"listen tcp :4311: bind: address already in use
```
This port is used by fluent bit [kuberntes filter](https://docs.fluentbit.io/manual/data-pipeline/filters/kubernetes) (default value for `aws_pod_association_port` field)

The error comes from here - https://github.com/aws/amazon-cloudwatch-agent/blob/main/extension/server/extension.go#L127
And most likely is caused by this line - https://github.com/aws/amazon-cloudwatch-agent/blob/main/extension/server/extension.go#L110 This does not seem to wait for port to be freed.
This should be replaced by `sever.Shutdown(...)`, but the whole code seem heavy weight to reload certs (shutting and starting server).

**Steps to reproduce**
EKS cluster with cloudwatch agent, there is similar issue (errors coming from aws fluent-bit) here - https://github.com/aws/amazon-cloudwatch-agent-operator/issues/269
```
[filter:kubernetes:kubernetes.1] no upstream connections available to cloudwatch-agent.amazon-cloudwatch:4311

```
**What did you expect to see?**
No errors, if server not reloaded do not just blindly log error in go routine (no error is returned from the method) and pretend everything is ok - https://github.com/aws/amazon-cloudwatch-agent/blob/main/extension/server/extension.go#L127
At least there should be retry if the port is available before starting server in go routine.

**What did you see instead?**
Server restarted when `reloadServer` is called or error returned. None of this currently happens.

**What version did you use?**
latest

**What config did you use?**
default

**Environment**
OS: linux

**Additional context**
Add any other context about the problem here.

Contributor guide

Open the contributing guide

Research direction

Start in extension/server/extension.go around lines 110 and 127, then reproduce the reload on an EKS cluster with CloudWatch Agent and Fluent Bit using the default configuration. Trace how reloadServer starts the listener and handles its goroutine error. Done means reloads do not produce address-in-use failures or silently ignore a failed server start.

Written by the indexing model from the issue text.

Assessment

Tech stack
go, kubernetes, linux
Domain
backend, devops, networking
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
48/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.