aws / aws/aws-lambda-roadmap

[Lambda][CloudWatch] Throttled Executions Counted as Success, Error, or Ignored

Abierto
#66 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Lenguaje dominante
Sin datos de lenguaje
Estrellas
196
Forks
5
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

### Community Note

* Please vote on this issue by adding a 👍 [reaction](https://blog.github.com/2016-03-10-add-reactions-to-pull-requests-issues-and-comments/) to the original issue to help the community and maintainers prioritize this request
* Please do not leave "+1" or "me too" comments, they generate extra noise for issue followers and do not help prioritize the request
* If you are interested in working on this issue or have submitted a pull request, please leave a comment
**Tell us about your request**
When using AWS Lambda with reserved concurrency, Lambda executions can be throttled. Currently, in CloudWatch, these throttled executions are recorded as successes. We request the ability to configure the default behavior of throttled executions, allowing users to choose whether they are logged as successes (default), errors, or ignored entirely.

**Which service(s) is this request for?**
AWS Lambda / CloudWatch

**Tell us about the problem you're trying to solve. What are you trying to do, and why is it hard?**
When using AWS Lambda with reserved concurrency, throttled executions can occur. In CloudWatch, these throttled executions are currently logged as successes, which can lead to misleading metrics and unreliable alarms.

For example:
- **Input:** A Lambda function with reserved concurrency set to 1, a timeout of 15 minutes, and a rule to invoke the function every 5 minutes.
- **Behavior:** If an issue occurs during execution (e.g., Lambda timeout), the following output is observed:
- 1 execution logged as an error (timeout)
- 2 throttled executions logged as successes

This behavior makes it difficult to build reliable CloudWatch alarms. For instance, if we want to allow occasional failures but not consecutive ones (e.g., intermittent issues), the current logging behavior of throttled executions as successes complicates alarm configuration.

We propose adding the ability to configure how throttled executions are logged in CloudWatch:
- **Success (default)**
- **Error**
- **Ignored (no logs)**

**Are you currently working around this issue?**
To work around this issue, I configure CloudWatch alarms with an evaluation period equal to the Lambda timeout plus the invocation delay (15 + 5 minutes). I then set the alarm to trigger if there are 2 errors out of 3 evaluation periods. However, this workaround is complex and not ideal for all use cases.

**Additional context**
Here is an example of the issue: In the "Error count and success rate" graph, throttled executions are logged as successes, even though all real executions are errors. This behavior leads to misleading metrics and unreliable alarms.

**Attachments**
Attached is an example graph showing the issue, where throttled executions are logged as successes despite all actual executions being errors.
Overall view:
Image
Error count and success rate: some executions count as "not in error".
Image

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

En el issue no se identifican archivos del repositorio, pruebas ni puntos de entrada de implementación. Empieza revisando cómo se representa el throttling de AWS Lambda en las métricas de CloudWatch y, a continuación, define cómo deberían comportarse los resultados configurables de éxito, error e ignorados, y cómo se verificaría cada opción.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
aws
Área
cloud, observability-sre
Tipo de issue
Nueva funcionalidad
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
25/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.