aws / aws/amazon-ecs-agent

Discrepancy in network TX/RX rate as reported by AWS ECS insights and task metadata endpoint

Open
#4,505 0 comments 0 reactions 0 assignees View on GitHub
scope/Fargate
Dominant language
Go
Stars
2.2k
Forks
662
Avg merge
3d 22h
Merged PRs (30d)
24

Description

## What we have

We are running our workloads on a Fargate cluster, the OS/architecture is Linux/X86_64, the platform version is 1.4.0, the network mode is `awsvpc`.
We are collecting ECS container insights, and we are persisting them to S3 bucket for long term storage and analytics.
We are also using Datadog for near-realtime monitoring and alerting. Datadog agent is running as a sidecar in all tasks and collects the container runtime data from the ECS task metadata endpoint.
With scaling out, our bill for the ECS insights has increased dramatically, and we started replacing them with an OTEL collector, which also runs as a sidecar and its pipeline uses `awsecscontainermetricsreceiver` to scrape the ECS task metadata endpoint and then `s3exporter` to dump the data into S3 bucket in the same format as ECS container insights.

## The problem

Comparing the data, we realised that although CPU and memory statistics show no significant difference between OTEL and ECS insights, network TX/RX rates are about 3x higher as reported by OTEL as compared to ECS insights. We also used Datadog as the third reference source, and Datadog is more in agreement with OTEL than with ECS insights. Datadog data is more granular, but the TX/RX rate values are of about the same magnitude as OTEL values, and are consistently higher than ECS insights.

## Other observations

I looked through the code of cloudwatch agent, ecs agent, and OTEL awsecscontainermetricsreceiver, in a hope that I'd uncover something suspicious.

We noticed long ago that the network tx/rx rates are the same for different containers within the task. I believe I saw a proof of that in the ecs agent code where it [uses task level metrics](https://github.com/aws/amazon-ecs-agent/blob/0f876b5372c9ecb15228f607f11d2c4be629d364/agent/stats/engine.go#L919-L937) to populate the container level metrics when the networking mode is `awsvpc`.

I also noticed the code that [splits task level network stats](https://github.com/aws/amazon-ecs-agent/blob/0f876b5372c9ecb15228f607f11d2c4be629d364/agent/stats/task_linux.go#L105-L117) equally between the containers. Maybe that explains the 3x difference somewhere (we have 3 containers per task).

I also noticed that the cloudwatch agent uses OTEL receiver's code behind the scenes, so in theory it should be sending the same data to CloudWatch, if that's how it works on Fargate. In reality, the data in CW logs and metrics for ECS insights is still different.

There seems to be a problem with data, whether it's a race condition when multiple components are polling the same metadata endpoint, or different approach to interpreting the data. Since it is the ECS agent that publishes the data on the task metadata endpoint, I see it as the potential culprit. Please investigate, because it adversely affects our analytics and makes us doubt the ECS insights data.

## Sample data

I extracted one hour of insights data for one specific container for comparison. I used queries/selectors to drill down on the specific task family, task id, and container.

### Network TX rate

ECS container insights:

![Image](https://github.com/user-attachments/assets/649b8dd7-68b1-41fe-8e45-5f6db162850c)

OTEL insights:

![Image](https://github.com/user-attachments/assets/771dda83-98c8-4351-ab99-82ffdc209a58)

Datadog insights:

![Image](https://github.com/user-attachments/assets/aa047f17-9c68-45b7-9df8-f7059a80ed9f)

### Network RX rate

ECS container insights:

![Image](https://github.com/user-attachments/assets/858ec5c0-786b-4d47-a739-a1c711d0606d)

OTEL insights:

![Image](https://github.com/user-attachments/assets/4057dfdb-45bb-477f-badc-b144ac6f5237)

Datadog insights:

![Image](https://github.com/user-attachments/assets/67e1131a-b34d-44ee-b92f-268509d4458d)

### Raw data

Extracted from CloudWatch log group for ECS insights:

[network - ecs insights.csv](https://github.com/user-attachments/files/18806933/network.-.ecs.insights.csv)

Extracted from data persisted in S3 for OTEL insights:

[network - otel collector.csv](https://github.com/user-attachments/files/18806936/network.-.otel.collector.csv)

Contributor guide

Open the contributing guide

Research direction

Start with agent/stats/engine.go around lines 919-937 and agent/stats/task_linux.go around lines 105-117, then compare their task-level network metric handling with the attached ECS and OTEL CSV data. Reproduce the discrepancy and determine whether task-level splitting or metadata interpretation accounts for it; done means the cause and an actionable correction are established.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, go
Domain
cloud, observability-sre
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.