influxdata / influxdata/influxdb

Influx not writing to disk

Open
#25,296 0 comments 2 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
31.7k
Forks
3.7k
Avg merge
13h 37m
Merged PRs (30d)
8

Description

We have a Docker Swarm cluster running InfluxDB v2.7.8 on a single node, where we experience issues every Sunday evening at 00 AM, with a recurrence interval of 2 to 3 weeks between occurrences.

We’ve noticed that memory usage suddenly begins to increase steadily, and InfluxDB struggles to write data to disk. While some data is still being written, most of it is not. I’ve reviewed the logs, including Docker logs, syslog, and docker.service logs, but haven’t found any relevant information. There are no error entries suggesting that Influx is having some sort of problem. It is compacting and doing its normal tasks.

Our first attempt to resolve this issue was upgrading InfluxDB from version 2.7.5 to 2.7.8, as there was a changelog entry addressing an infinite write loop bug. Unfortunately, this didn’t resolve the issue, as we experienced the same problem again today.

We have Telegraf running on the server which gathers metrics from the Influx, where we've seen a spike in the queue active:

![Skjermbilde fra 2024-09-09 10-04-34](https://github.com/user-attachments/assets/02acd069-b114-4724-8bdc-6d0f8bc3ff2b)

Below is also the graph from the memory usage of the server:

![image](https://github.com/user-attachments/assets/50503c70-d179-4d78-a97c-372b17263fa9)

The problem is solved by restarting the docker container, but we have lost all data that Influx had in memory.

__Environment info:__

uname -srm:
Linux 5.15.0-113-generic x86_64

Docker version:
Docker version 24.0.7, build afdd53b

Influxdb-docker image version 2.7.8

The server is a VM running in VMware.

__Logs:__
The only log error I can find which might be relevant is that Telegraf failed to send metrics to our off-site influx at 00:00 AM and 00:15 AM, even though our off-site influx still received some data from the server.

Contributor guide

Open the contributing guide

Research direction

Start with the Sunday 00:00 recurrence described in the issue, reviewing the Docker, syslog, docker.service, and Telegraf logs alongside the queue-active and memory graphs. Compare the reported environment—InfluxDB 2.7.8, Docker 24.0.7, Linux 5.15.0-113-generic, and a VMware VM—with the write failure period. Done means identifying a reproducible cause or the missing diagnostic evidence needed to isolate it.

Written by the indexing model from the issue text.

Assessment

Tech stack
docker, linux
Domain
databases, devops, observability-sre
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.