aws / aws/efs-utils

/usr/bin/amazon-efs-mount-watchdog - OSError: [Errno 28] No space left on device

Open
#154 7 comments 0 reactions 0 assignees View on GitHub
bug
Dominant language
Rust
Stars
361
Forks
238
Avg merge
2d 13h
Merged PRs (30d)
3

Description

On our servers it happens regularly that the servers crash and are inaccessible via SSM / SSH. The only solution is to stop the server (sometimes it restarts normally, sometimes we have to destroy the server)

After investigation I found these elements that correspond with the unavailability of the servers
> Jan 16 13:31:11 ip-XX-XX-XX-142 dhclient[3600]: XMT: Solicit on eth0, interval 120860ms.
Jan 16 13:33:12 ip-XX-XX-XX-142 dhclient[3600]: XMT: Solicit on eth0, interval 115990ms.
Jan 16 13:35:08 ip-XX-XX-XX-142 dhclient[3600]: XMT: Solicit on eth0, interval 129620ms.
Jan 16 13:37:18 ip-XX-XX-XX-142 dhclient[3600]: XMT: Solicit on eth0, interval 108240ms.
Jan 16 13:37:46 ip-XX-XX-XX-142 env: OSError: [Errno 28] No space left on device
Jan 16 13:37:46 ip-XX-XX-XX-142 env: During handling of the above exception, another exception occurred:
Jan 16 13:37:46 ip-XX-XX-XX-142 env: Traceback (most recent call last):
Jan 16 13:37:46 ip-XX-XX-XX-142 env: File "amazon-efs-mount-watchdog", line 2014, in
Jan 16 13:37:46 ip-XX-XX-XX-142 env: main()
Jan 16 13:37:46 ip-XX-XX-XX-142 env: File "/usr/bin/amazon-efs-mount-watchdog", line 2004, in main
Jan 16 13:37:47 ip-XX-XX-XX-142 env: unmount_count_for_consistency,
Jan 16 13:37:47 ip-XX-XX-XX-142 env: File "/usr/bin/amazon-efs-mount-watchdog", line 1005, in check_efs_mounts
Jan 16 13:37:47 ip-XX-XX-XX-142 env: rewrite_state_file(state, state_file_dir, state_file)
Jan 16 13:37:47 ip-XX-XX-XX-142 env: File "/usr/bin/amazon-efs-mount-watchdog", line 921, in rewrite_state_file
Jan 16 13:37:47 ip-XX-XX-XX-142 env: json.dump(state, f)
Jan 16 13:37:47 ip-XX-XX-XX-142 env: OSError: [Errno 28] No space left on device
Jan 16 13:37:47 ip-XX-XX-XX-142 systemd-udevd: fork of child failed: Cannot allocate memory
Jan 16 13:37:47 ip-XX-XX-XX-142 systemd: amazon-efs-mount-watchdog.service: main process exited, code=exited, status=1/FAILURE
Jan 16 13:37:47 ip-XX-XX-XX-142 systemd: Unit amazon-efs-mount-watchdog.service entered failed state.
Jan 16 13:37:47 ip-XX-XX-XX-142 systemd: amazon-efs-mount-watchdog.service failed.
Jan 16 13:38:27 ip-XX-XX-XX-142 systemd: amazon-efs-mount-watchdog.service holdoff time over, scheduling restart.
Jan 16 13:38:45 ip-XX-XX-XX-142 systemd: Stopped amazon-efs-mount-watchdog.
Jan 16 13:38:52 ip-XX-XX-XX-142 systemd: Started amazon-efs-mount-watchdog.
Jan 16 13:39:46 ip-XX-XX-XX-142 dhclient[3600]: XMT: Solicit on eth0, interval 126470ms.
Jan 16 13:41:45 ip-XX-XX-XX-142 dhclient[3600]: XMT: Solicit on eth0, interval 126770ms.
Jan 16 13:42:21 ip-XX-XX-XX-142 amazon-ssm-agent: runtime/cgo: pthread_create failed: Resource temporarily unavailable
Jan 16 13:42:51 ip-XX-XX-XX-142 amazon-ssm-agent: SIGABRT: abort
Jan 16 13:43:09 ip-XX-XX-XX-142 amazon-ssm-agent: PC=0x7fad8a3b4051 m=2 sigcode=18446744073709551610
Jan 16 13:43:16 ip-XX-XX-XX-142 amazon-ssm-agent: goroutine 0 [idle]:
Jan 16 13:43:49 ip-XX-XX-XX-142 amazon-ssm-agent: runtime: unknown pc 0x7fad8a3b4051
Jan 16 13:44:24 ip-XX-XX-XX-142 amazon-ssm-agent: stack: frame={sp:0x7fad63126bc0, fp:0x0} stack=[0x7fad62727678,0x7fad63127278)
Jan 16 13:43:49 ip-XX-XX-XX-142 amazon-ssm-agent: runtime: unknown pc 0x7fad8a3b4051
Jan 16 13:44:24 ip-XX-XX-XX-142 amazon-ssm-agent: stack: frame={sp:0x7fad63126bc0, fp:0x0} stack=[0x7fad62727678,0x7fad63127278)
Jan 16 13:50:40 ip-XX-XX-XX-142 journal: Runtime journal is using 8.0M (max allowed 1.5G, trying to leave 2.3G free of 15.5G available → current limit 1.5G).
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: Linux version 4.14.301-224.520.amzn2.x86_64 (mockbuild@ip-10-0-47-71) (gcc version 7.3.1 20180712 (Red Hat 7.3.1-15) (GCC)) #1 SMP Fri Dec 9 09:57:03 UTC 2022
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: Command line: BOOT_IMAGE=/boot/vmlinuz-4.14.301-224.520.amzn2.x86_64 root=UUID=a482dce8-a78a-42c8-931e-7a3bbdd3eb43 ro console=tty0 console=ttyS0,115200n8 net.ifnames=0 biosdevname=0 nvme_core.io_timeout=4294967295 rd.emergency=poweroff rd.shell=0
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: x86/fpu: Supporting XSAVE feature 0x001: 'x87 floating point registers'
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: x86/fpu: Supporting XSAVE feature 0x002: 'SSE registers'
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: x86/fpu: Supporting XSAVE feature 0x004: 'AVX registers'
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: x86/fpu: xstate_offset[2]: 576, xstate_sizes[2]: 256
Jan 16 13:50:40 ip-XX-XX-XX-142 kernel: x86/fpu: Enabled xstate features 0x7, context size is 832 bytes, using 'standard' format.

Storage :
>[ssm-user@ip-xxx-xxx-xxx-142 bin]$ df -H
Filesystem Size Used Avail Use% Mounted on
devtmpfs 17G 0 17G 0% /dev
tmpfs 17G 0 17G 0% /dev/shm
tmpfs 17G 521k 17G 1% /run
tmpfs 17G 0 17G 0% /sys/fs/cgroup
/dev/nvme0n1p1 275G 17G 259G 6% /
127.0.0.1:/ 9.3E 56G 9.3E 1% /experiments
tmpfs 3.4G 0 3.4G 0% /run/user/1000

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.