open-telemetry / open-telemetry/opentelemetry-lambda

Improving lambda Cold start

Open
#727 10 comments 33 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Go
Stars
432
Forks
246
Avg merge
3d 10h
Merged PRs (30d)
47

Description

Is your feature request related to a problem? Please describe.

A Lambda cold start happens when a new instance of a Lambda function must be created and initialized. The cold start refers to the delay between invocation and runtime created by the initialization process.
New instances needs to be initialized whenever other instances have expired due to inactivity or when there are more invocations than active instances. Cold starts are an inherent problem with Lambda functions because it is not possible to keep lambda initialized forever.

The OpenTelemetry SDK was not created with Lambda functions in mind. If you use OpenTelemetry inside a lambda function, the overhead of initializing the SDK and optionally auto instrumenting the application code adds up in the cold start time. This is specially painful for users because this inserts high latency in their application and increases the cost of running the lambdas.

Describe the solution you'd like

This proposal will tackle the cold start time of the OpenTelemetry lambda layers with the following plan:

Plan:

  • Continuously measure the cold start time of the layers with each release. This will help catching regressions in performance and also show trends and where we should invest our time.
  • Profile each layer to identify where all this time is spent on in the code.
  • Propose optimizations in the initialization of the SDK in each layer: Using the profiling information from the previous step, look for low hanging fruits and also more complex refactoring that will improve the performance.

Methodology for measuring the cold startup:

  • Measure the cold start time for lambdas with and without the layers for each supported layer.
    • Create a sample application and deploy to a lambda function
    • Generate load for this sample application
    • Vary a parameter in the lambda function that will force the lambda to be recreated.
    • Parse the logs of the lambda function with the following query:
filter @type="REPORT"
| filter ispresent(@initDuration)
| stats count(@initDuration) as coldStartCount, pct(@initDuration, 50) as p50Init, pct(@initDuration, 90) as p90Init, pct(@initDuration, 99) as p99Init group by @log 

Methodology for profiling the lambda functions:

  • TBD - We will need to

Additional context
References

https://github.com/open-telemetry/opentelemetry-lambda/issues/263

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the Lambda layer release process and the cold-start measurement methodology described in the issue. Build sample applications with and without each supported layer, generate load, vary the parameter that recreates the function, and analyze REPORT logs with the provided query. Done means release-to-release measurements, profiling results, and documented initialization optimizations for the affected layers.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws
Domain
cloud, observability, performance
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Active
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.