On Elastic Cloud, would like a troubleshooting assistance when APM signals stop flowing to the APM GUI
- Dominant language
- Go
- Stars
- 1.3k
- Forks
- 543
- Avg merge
- 1d 18h
- Merged PRs (30d)
- 109
Description
# Problem Statement
When using Elastic through Elastic Cloud (aka ESS), the apm signals may stop flowing to the the APM GUI.
Causes reside in the Elastic cluster or be external and there may be no error reported by the apm agents. An upgrade problem is an example of such a problem.
When this problem happen, I would like to have assistance in Elastic Cloud GUI to troubleshoot.
This assistance could be composed of
* Status and version check: confirm that APM Server is running, verify that the version of APM Server and of the "APM integration" meet the requirements
* Dashboard of the key metrics of APM server
* Ingest rate broken down by ingestion status
* Success
* Failure
* Rejected by APM Server
* Accepted by APM Server but failed to be persisted in Elasticsearch
* Maybe ingest rate broken down by the signal type
* Indications of connection failures like authentication failures
* Timestamps of the last processed data point and of the last successfully ingested data point
* Audit trail of the last log messages produced by APM Server, at least of the last warning & error log messages
# Elastic Cloud screens related to APM Server health



# Workflow to get more troubleshooting details enabling "Logs and metrics"



Contributor guide
Assessment
This issue has not been assessed yet.