influxdata / influxdata/influxdb

sporadic-tls-bad-record-mac-error

Open
#24,368 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
31.7k
Forks
3.7k
Avg merge
13h 37m
Merged PRs (30d)
8

Description

__Steps to reproduce:__
Start influxdb with the following configuration:
```
[http]
auth-enabled = true
pprof-enabled = false
flux-enabled = true
https-enabled = true
https-certificate = "client.crt"
https-private-key = "client.key"
[tls]
min-version = "tls1.2"
max-version = "tls1.3"
# https://wiki.mozilla.org/Security/Server_Side_TLS#Intermediate_compatibility_.28recommended.29
#
# Can this be configured more cleanly?
# strict-ciphers didn't work / or not sure on where to configure it
ciphers = [ "TLS_AES_128_GCM_SHA256",
"TLS_AES_256_GCM_SHA384",
"TLS_CHACHA20_POLY1305_SHA256",
"TLS_ECDHE_ECDSA_WITH_AES_128_GCM_SHA256",
"TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256",
"TLS_ECDHE_ECDSA_WITH_AES_256_GCM_SHA384",
"TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384",
"TLS_ECDHE_ECDSA_WITH_CHACHA20_POLY1305",
"TLS_ECDHE_RSA_WITH_CHACHA20_POLY1305",
"TLS_ECDHE_RSA_WITH_AES_128_CBC_SHA",
"TLS_ECDHE_RSA_WITH_AES_256_CBC_SHA",
"TLS_RSA_WITH_AES_128_CBC_SHA",
"TLS_RSA_WITH_AES_256_CBC_SHA"
]

```
Anyway the configuration seemd to running fine for a while and then suddenly yesterday our influxdb became in accessible. Grafana started throwing 502 errors and trying to do curl commands:

```
curl --fail --silent --show-error -k -u grafana_user: -G "https://10.0.67.1:8086/query?db=metrics" --data-urlencode "q=select LAST(value) from /^some.metric*/ where time > now() - 1m"
```
we get the error `Failed with curl: (7) Failed to connect to 10.0.67.1 port 8086: Connection refused`

On restarting the VM ofcourse everything worked back again. On checking the logs the error stated on influxdb was
```
http: TLS handshake error from 10.0.67.6:38084: local error: tls: bad record MAC
```
Our initial assumption was it might be because of manually listing the ciphers, so we changed the configuration as

```
[http]
auth-enabled = true
log-enabled = <%= p('influxdb.enable_http_log') %>
pprof-enabled = false
flux-enabled = true
https-enabled = true
https-certificate = "client.crt"
https-private-key = "client.key"
[tls]
min-version = "tls1.3"

[[graphite]]
enabled = true
database = "<%= p('influxdb.metrics_database_name') %>"
```
After the new update with just 1.3, we noticed a reduction in the issue, however, it occasionally creeps up. As it just happens sporadically haven’t identified a trigger for it as such, as it just suddenly starts in an entirely healthy production setup. The issue fixes itself after a really long time or when we restart influxdb.

The biggest issue is that the memory spikes up and the VM starts becoming unresponsive (even to ssh )
image

What could be causing the issue and how do we debug this?
Currently using influx 1.8.10 OSS

Contributor guide

Open the contributing guide

Research direction

Start with the supplied HTTPS configuration and curl request, then compare the InfluxDB TLS handshake error with the reported memory spike and connection refusal. Use the logs and the cipher/min-version changes described in the issue to establish a reproducible trigger or diagnostic path; done means the cause and a verified mitigation are documented.

Written by the indexing model from the issue text.

Assessment

Tech stack
grafana
Domain
backend, databases, security
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.