open-telemetry / open-telemetry/opentelemetry-lambda

Metric flush cannot be triggered automatically

Open
#1,949 1 comment 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

bug javascript
Dominant language
Go
Stars
432
Forks
246
Avg merge
3d 10h
Merged PRs (30d)
47

Description

Describe the bug
I noticed that automatic flushing is triggered before lambda receives shutdown signal. However, it makes the http exporter fail with http timeout.

Forcing the metric flush or triggering shutdown sequence during lambda invocation works. In this case, I've observed same metric being emitted to grafana 5 times, each 1 minute apart. This part of the bug may not be relevant to the project; but I'm unable to find the root cause.

Steps to reproduce
I've created this repo that simulates the behavior. You can find the details logs of the test results there as well.

What did you expect to see?
Basic config of otel exporter with http exproter successfully exports metrics.

What did you see instead?

ERROR	{"stack":"Error: PeriodicExportingMetricReader: metrics export failed (error Error: Request Timeout)\n    at d._doRun (/opt/773.wrapper.js:1:3565)\n    at processTicksAndRejections (node:internal/process/task_queues:95:5)\n    at runNextTicks (node:internal/process/task_queues:64:3)\n    at process.processTimers (node:internal/timers:516:9)\n    at async d._runOnce (/opt/773.wrapper.js:1:2894)\n    at async d.onForceFlush (/opt/773.wrapper.js:1:3784)\n    at async d.forceFlush (/opt/773.wrapper.js:1:2000)\n    at async ee.forceFlush (/opt/773.wrapper.js:1:21040)\n    at async Promise.all (index 0)\n    at async ne.forceFlush (/opt/773.wrapper.js:1:24412)","message":"PeriodicExportingMetricReader: metrics export failed (error Error: Request Timeout)","name":"Error"}

What version of collector/language SDK version did you use?
node20
layers:

  • opentelemetry-nodejs-0_16_0:1
  • opentelemetry-collector-amd64-0_17_0:1

What language layer did you use?
nodejs

Additional context

Reproducible version can be found in https://github.com/kaskavalci/otlp-lambda-example/tree/main

It's weird that when shutdown or forceflush is triggered, same metric appears 5 times in grafana. I appreaciate your comment on this as well.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the linked otlp-lambda-example reproduction and its docs/test-results.md automatic-flush section, using the Node.js 20 and listed layer versions. Trace the automatic flush and shutdown behavior around the HTTP exporter timeout and repeated Grafana emissions. Done means the root cause is identified and the expected metric export completes without timeout or duplicate emissions.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, nodejs
Domain
cloud, observability-sre
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.