hashicorp / hashicorp/nomad

(very) old allocation still existing and breaking new Topology view

Open
#9,617 2 comments 1 reaction 0 assignees View on GitHub
stage/needs-investigation theme/core type/bug
Dominant language
Go
Stars
17k
Forks
2.1k
Avg merge
1d 9h
Merged PRs (30d)
105

Description

### Nomad version
Nomad v1.0.0 (cfca6405ad9b5f66dffc8843e3d16f92f3bedb43)

### Operating system and Environment details
Linux, AWS

### Issue
After upgrading our clusters to Nomad v1.0.0 we can not see the Topology view.
After a quick investigation we found out that the UI is producing following error:
```
Uncaught (in promise) Error: Node b37eb7a6-ff6b-af73-e756-e019f2cabf57 for alloc 05bf3c66-9dbb-b06a-4cf3-3216e6e922e4 not in index.
```

After checking the allocation id we get this:
```
nomad alloc status 05bf3c66-9dbb-b06a-4cf3-3216e6e922e4
ID = 05bf3c66-9dbb-b06a-4cf3-3216e6e922e4
Eval ID = f373dfb7
Name = authentication.hermes[0]
Node ID = b37eb7a6
Node Name = ip-10-169-141-162.eu-central-1.compute.internal
Job ID = authentication
Job Version = 86
Client Status = complete
Client Description = All tasks have completed
Desired Status = stop
Desired Description = alloc not needed due to job update
Created = 1y1mo ago
Modified = 11mo28d ago
Deployment ID = 7eb18ddf
Deployment Health = unset
Canary = true

Couldn't retrieve stats: Unexpected response code: 500 (Unknown node b37eb7a6-ff6b-af73-e756-e019f2cabf57)

Task "authentication" is "dead"
Task Resources
CPU Memory Disk Addresses
1000 MHz 768 MiB 300 MiB api: 10.169.141.162:22564

Task Events:
Started At = N/A
Finished At = 2019-11-01T10:17:37Z
Total Restarts = 0
Last Restart = N/A

Recent Events:
Time Type Description
2019-12-19T08:25:28+01:00 Received Task received by client
2019-12-19T07:10:22+01:00 Killing Sent interrupt. Waiting 5s before force killing
2019-12-19T06:54:40+01:00 Received Task received by client
2019-12-19T06:49:22+01:00 Killing Sent interrupt. Waiting 5s before force killing
2019-12-19T06:28:49+01:00 Received Task received by client
2019-12-19T06:18:38+01:00 Killing Sent interrupt. Waiting 5s before force killing
2019-12-19T05:57:56+01:00 Received Task received by client
2019-12-19T05:52:38+01:00 Killing Sent interrupt. Waiting 5s before force killing
2019-12-19T05:42:32+01:00 Received Task received by client
2019-12-19T05:27:09+01:00 Killing Sent interrupt. Waiting 5s before force killing
```

The job itself:
```
nomad job status authentication
ID = authentication
Name = authentication
Submit Date = 2020-12-11T10:01:57+01:00
Type = service
Priority = 50
Datacenters = dc1
Namespace = default
Status = running
Periodic = false
Parameterized = false

Summary
Task Group Queued Starting Running Failed Complete Lost
hermes 0 0 2 0 3 0

Latest Deployment
ID = dfbb1013
Status = successful
Description = Deployment completed successfully

Deployed
Task Group Auto Revert Promoted Desired Canaries Placed Healthy Unhealthy Progress Deadline
hermes true true 2 1 2 2 0 2020-12-11T09:08:05Z

Allocations
ID Node ID Task Group Version Desired Status Created Modified
c7ee4423 bd478068 hermes 173 run running 2m58s ago 2m17s ago
074106b4 72c99dc4 hermes 173 run running 3m25s ago 2m59s ago
8e0e3a95 a9e0ff4a hermes 172 stop complete 16h43m ago 2m42s ago
2f7934ad bd478068 hermes 172 stop complete 16h45m ago 2m42s ago
05bf3c66 b37eb7a6 hermes 86 stop complete 1y1mo ago 11mo28d ago
```

So it seems like Nomad keeps this allocation, even after about a whole year...
The client node is long gone..

We even tried to stop and purge the job (and recreated it, with a fresh job version counter) but this allocation is still there.

Neither
```
nomad system gc
```
or
```
nomad system reconcile summaries
```
changed anything.

I even tried to stop the allocation, but it yields the same output and result every time, without changing anything.
```
nomad alloc stop -verbose 05bf3c66-9dbb-b06a-4cf3-3216e6e922e4
==> Monitoring evaluation "1fbae262-dc1b-1dd3-ef24-faa5272c2763"
Evaluation triggered by job "authentication"
==> Monitoring evaluation "1fbae262-dc1b-1dd3-ef24-faa5272c2763"
Evaluation within deployment: "dfbb1013-14f8-b248-7742-5a8b99c5e6dd"
Evaluation status changed: "pending" -> "complete"
==> Evaluation "1fbae262-dc1b-1dd3-ef24-faa5272c2763" finished with status "complete"
```

### Reproduction steps

I don't know how to reproduce this.
But we have a few more allocations like this. Also pointing to different (long gone) nodes.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by investigating the stale allocation reported by `nomad alloc status` and the behavior of `nomad system gc`, `nomad system reconcile summaries`, and `nomad alloc stop`. The report provides no source files or reproducible steps; done should mean allocations referencing gone nodes no longer break the Topology view and can be removed or handled safely.

Written by the indexing model from the issue text.

Assessment

Tech stack
aws, linux
Domain
devops, infrastructure
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.