aws / aws/aws-lambda-roadmap

[Lambda][CloudWatch] Throttled Executions Counted as Success, Error, or Ignored

オープン
#66 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
主要言語
言語のデータがありません
スター
196
フォーク
5
PR マージ指標
30日以内にマージされた PR はありません

説明

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment
**Tell us about your request**
When using AWS Lambda with reserved concurrency, Lambda executions can be throttled. Currently, in CloudWatch, these throttled executions are recorded as successes. We request the ability to configure the default behavior of throttled executions, allowing users to choose whether they are logged as successes (default), errors, or ignored entirely.

**Which service(s) is this request for?**
AWS Lambda / CloudWatch

**Tell us about the problem you're trying to solve. What are you trying to do, and why is it hard?**
When using AWS Lambda with reserved concurrency, throttled executions can occur. In CloudWatch, these throttled executions are currently logged as successes, which can lead to misleading metrics and unreliable alarms.

For example:
- **Input:** A Lambda function with reserved concurrency set to 1, a timeout of 15 minutes, and a rule to invoke the function every 5 minutes.
- **Behavior:** If an issue occurs during execution (e.g., Lambda timeout), the following output is observed:
- 1 execution logged as an error (timeout)
- 2 throttled executions logged as successes

This behavior makes it difficult to build reliable CloudWatch alarms. For instance, if we want to allow occasional failures but not consecutive ones (e.g., intermittent issues), the current logging behavior of throttled executions as successes complicates alarm configuration.

We propose adding the ability to configure how throttled executions are logged in CloudWatch:
- **Success (default)**
- **Error**
- **Ignored (no logs)**

**Are you currently working around this issue?**
To work around this issue, I configure CloudWatch alarms with an evaluation period equal to the Lambda timeout plus the invocation delay (15 + 5 minutes). I then set the alarm to trigger if there are 2 errors out of 3 evaluation periods. However, this workaround is complex and not ideal for all use cases.

**Additional context**
Here is an example of the issue: In the "Error count and success rate" graph, throttled executions are logged as successes, even though all real executions are errors. This behavior leads to misleading metrics and unreliable alarms.

**Attachments**
Attached is an example graph showing the issue, where throttled executions are logged as successes despite all actual executions being errors.
Overall view:
Image
Error count and success rate: some executions count as "not in error".
Image

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

Issue では、リポジトリのファイル、テスト、実装のエントリーポイントは特定されていません。まず、AWS Lambda の throttling が CloudWatch メトリクスでどのように表現されるかを確認し、次に、設定可能な成功、エラー、無視の outcome がどのように動作すべきか、また各オプションをどのように検証するかを定義してください。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws
領域
cloud, observability-sre
issue の種類
機能追加
難易度
5/5
見積もり時間
1週間以上
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
25/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。