aws / aws/aws-lambda-roadmap

[Lambda][CloudWatch] Throttled Executions Counted as Success, Error, or Ignored

Aperta
#66 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Nessun dato sulla lingua
Stelle
196
Fork
5
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment
**Tell us about your request**
When using AWS Lambda with reserved concurrency, Lambda executions can be throttled. Currently, in CloudWatch, these throttled executions are recorded as successes. We request the ability to configure the default behavior of throttled executions, allowing users to choose whether they are logged as successes (default), errors, or ignored entirely.

**Which service(s) is this request for?**
AWS Lambda / CloudWatch

**Tell us about the problem you're trying to solve. What are you trying to do, and why is it hard?**
When using AWS Lambda with reserved concurrency, throttled executions can occur. In CloudWatch, these throttled executions are currently logged as successes, which can lead to misleading metrics and unreliable alarms.

For example:
- **Input:** A Lambda function with reserved concurrency set to 1, a timeout of 15 minutes, and a rule to invoke the function every 5 minutes.
- **Behavior:** If an issue occurs during execution (e.g., Lambda timeout), the following output is observed:
- 1 execution logged as an error (timeout)
- 2 throttled executions logged as successes

This behavior makes it difficult to build reliable CloudWatch alarms. For instance, if we want to allow occasional failures but not consecutive ones (e.g., intermittent issues), the current logging behavior of throttled executions as successes complicates alarm configuration.

We propose adding the ability to configure how throttled executions are logged in CloudWatch:
- **Success (default)**
- **Error**
- **Ignored (no logs)**

**Are you currently working around this issue?**
To work around this issue, I configure CloudWatch alarms with an evaluation period equal to the Lambda timeout plus the invocation delay (15 + 5 minutes). I then set the alarm to trigger if there are 2 errors out of 3 evaluation periods. However, this workaround is complex and not ideal for all use cases.

**Additional context**
Here is an example of the issue: In the "Error count and success rate" graph, throttled executions are logged as successes, even though all real executions are errors. This behavior leads to misleading metrics and unreliable alarms.

**Attachments**
Attached is an example graph showing the issue, where throttled executions are logged as successes despite all actual executions being errors.
Overall view:
Image
Error count and success rate: some executions count as "not in error".
Image

Guida per i contributori

Apri la guida per i contributori

Direzione di ricerca

Nell’issue non sono identificati file del repository, test o punti di ingresso dell’implementazione. Inizia esaminando come il throttling di AWS Lambda è rappresentato nelle metriche CloudWatch, quindi definisci come dovrebbero comportarsi gli esiti configurabili di successo, errore e ignorati e come verrebbe verificata ciascuna opzione.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
aws
Ambito
cloud, observability-sre
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Abbastanza chiara
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.