opencontainers / opencontainers/runc

`runc state` reports a CRIU-restored container as `running` after its process has exited

Open Beginner friendly
#5,370 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

llm-generated
Dominant language
Go
Stars
13.5k
Forks
2.3k
Avg merge
2d 8h
Merged PRs (30d)
30

Description

Summary

After a container is restored with runc restore, if its process later exits, runc state (and any consumer of the container status) continues to report the container as
running instead of stopped. A container started normally with runc run/runc start reports stopped correctly in the same situation.

The root cause is that restoredState.transition() does not update the container's
current state, unlike every other state's transition().

Reproduction (state-machine unit test — confirmed)

The bug is in the state-machine logic and reproduces directly against runc's own
libcontainer package. Add this test to libcontainer/ and run
go test ./libcontainer/ -run TestRUNC2 -v:

func TestRUNC2_RestoredExitedReportsStopped(t *testing.T) {
    c := &Container{}
    c.state = &restoredState{c: c}
    // what refreshState() does when the process is gone (!hasInit()):
    if err := c.state.transition(&stoppedState{c: c}); err != nil {
        t.Fatalf("transition returned error: %v", err)
    }
    if got := c.state.status(); got != Stopped {
        t.Fatalf("RUNC-2 BUG: restored+exited reports %v, want Stopped", got)
    }
}

Result on the current code:

--- FAIL: TestRUNC2_RestoredExitedReportsStopped
    RUNC-2 BUG: a restored container whose process exited reports running, want Stopped
--- PASS: TestRUNC2_RunningExitedReportsStopped   (same test with runningState)

The identical test with runningState passes, confirming the asymmetry. Applying the
suggested fix below makes both pass, with no regression in the existing state tests
(go test ./libcontainer/ -run State). (Verified on commit fc89fbd.)

Steps to reproduce (end-to-end, via CRIU)

  1. Create and run a container, then checkpoint it:
    runc run -d mycontainer
    runc checkpoint --image-path /tmp/img mycontainer
    
  2. Restore it:
    runc restore -d --image-path /tmp/img mycontainer
    
  3. Let the container's process exit (or kill it):
    runc kill mycontainer KILL     # or wait for the workload to finish
    
  4. Query the status:
    runc state mycontainer
    

Expected behavior

runc state reports "status": "stopped" (the process is gone), consistent with a
container started via runc run.

Actual behavior

runc state reports "status": "running" for the dead container. runc list shows it
as running as well.

Root cause

Container.currentStatus() calls refreshState(), which re-derives the real status
from reality and applies the corresponding transition
(libcontainer/container_linux.go, refreshState):

if !c.hasInit() {
    return c.state.transition(&stoppedState{c: c})   // process is gone -> should become stopped
}

Every state's transition() sets c.state = s on an accepted transition — except
restoredState.transition()
, which returns nil without updating c.state
(libcontainer/state_linux.go:207-213):

func (r *restoredState) transition(s containerState) error {
    switch s.(type) {
    case *stoppedState, *runningState:
        return nil                 // <-- returns without `r.c.state = s`
    }
    return newStateTransitionError(r, s)
}

Compare runningState.transition() (state_linux.go:117-131), which does
r.c.state = s on the *stoppedState case:

case *stoppedState:
    if r.c.hasInit() {
        return ErrRunning
    }
    r.c.state = s                  // <-- running correctly updates
    return nil

Because restoredState.status() returns Running (state_linux.go:203-205), a
restored container whose process has exited stays in restoredState after
refreshState, so currentStatus() returns Running. A normally-run container takes
the runningState path and correctly becomes stopped.

(This is why the inconsistency is persistent rather than self-correcting: for other
states, refreshState's re-derivation fixes any transient mismatch on the next status
query, but restoredState never leaves that state via transition.)

Suggested fix

Have restoredState.transition() update c.state like the other states, while
preserving the checkpoint-cleanup behavior that restoredState.destroy() provides —
e.g.:

func (r *restoredState) transition(s containerState) error {
    switch s.(type) {
    case *stoppedState, *runningState:
        r.c.state = s
        return nil
    }
    return newStateTransitionError(r, s)
}

(If the intent of keeping restoredState was to retain imageDir for
destroy()-time checkpoint cleanup, an alternative is to make restoredState.status()
re-derive liveness, or to carry the checkpoint-cleanup responsibility into the target
state.)

Version

  • runc 1.5.0-rc.1+dev, commit fc89fbd9ebec617475d7e7a7a38f4e4bf277cf54
  • Please confirm against latest main and check for existing/duplicate reports before
    filing.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in libcontainer/state_linux.go at restoredState.transition and compare it with runningState.transition. Run go test ./libcontainer/ -run TestRUNC2 -v and the existing State tests; done means a restored container whose process exits transitions to stopped and runc state reports stopped without regressions.

Written by the indexing model from the issue text.

Assessment

Tech stack
go
Domain
cli, infrastructure
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
90/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.