aws / aws/containers-roadmap

Support Task execution timeout(maximum-lifetime for a container) in ECS fargate

Open
#572 36 comments 210 reactions 0 assignees View on GitHub
ECS Fargate
Dominant language
Shell
Stars
5.4k
Forks
334
PR merge metrics
No merged PRs in 30d

Description

### Summary

ECS does not currently support a task execution timeout so that when a task exceeds more than certain period of time, the task must be stopped automatically like how AWS Batch has job timeouts. The task definition does not have a parameter to enforce a task/container execution timeout that will automatically trigger the container to stop after the set time.

Use-case example from a customer:
I have a NLP model training job I want to run in a fargate container triggered by a lambda function. At some time, a bug might be introduced in the training code that would cause it to run indefinitely. I don't want to accidentally have those tasks piling up and have 50 tasks running for a couple weeks before we notice. That could have a cost implication. Is there a native way to kill a container if it hasn't exited on its own before a certain time?

Can this be considered as a feature request?

Contributor guide

Open the contributing guide

Research direction

No repository files or tests are named. Start by reviewing the requested ECS/Fargate task execution timeout and the AWS Batch job-timeout behavior cited in the issue; done means establishing how a configured maximum lifetime would automatically stop an overlong task or container in the Lambda-triggered training scenario.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws
Domain
cloud, devops
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.