NVIDIA-NeMo / NVIDIA-NeMo/nemo-platform

[entities] workspace_cleanup controller goes permanently unhealthy: async engine singleton binds pool to the wrong event loop (startup race)

Open
#1,230 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
78
Forks
23
Avg merge
1d 14h
Merged PRs (30d)
578

Description

Summary

The entities service's workspace_cleanup controller permanently reports unhealthy (and its cleanup loop stops processing pending workspace deletions) whenever the process-global SQLAlchemy async engine singleton is first initialized on a different event loop than the controller's private loop and the connection pool is subsequently exhausted by concurrent load. The controller's connection checkout then waits on an asyncio.Queue bound to the other loop and raises:

RuntimeError: <Queue ... maxsize=5> is bound to a different event loop

Because controllers.healthy aggregates this flag into /status, one stuck background controller falsely degrades the whole deployment's health signal. In a full regression run this single flag caused 25/198 agent test cases to error at setup (health gate refused entry) while every API those tests exercise was healthy.

The failure is a startup race: whether a deployment is affected depends on which loop touches the DB singleton first, so identical builds can come up healthy or permanently degraded.

Source analysis

Three sites interact:

  1. Engine singleton ignores loop affinityservices/core/entities/src/nmp/core/entities/app/repository/__init__.py:
_async_engine: AsyncEngine | None = None

async def initialize_async_engine(config: EntitiesConfig) -> None:
    global _async_engine, _async_session_maker
    if _async_engine is not None:
        return  # Already initialized  ← no check WHICH loop created it
  1. The uvicorn lifespan initializes the same singleton on the server loopservices/core/entities/src/nmp/core/entities/service.py (await initialize_async_engine(cfg) inside the startup/retry path). Controllers run as daemon threads in the same process (packages/nmp_platform_runner/src/nmp/platform_runner/server.py).

  2. The controller thread believes it owns the poolservices/core/entities/src/nmp/core/entities/controllers/main.py:

# Create a single event loop that will be shared for DB init and the cleanup controller,
# so SQLAlchemy's async pool is bound to the same loop that later runs queries.
loop = asyncio.new_event_loop()
loop.run_until_complete(initialize_async_engine(entities_config))  # ← NO-OP if the server won

The comment documents the intended invariant, but the singleton's early return silently breaks it whenever the server lifespan initializes first. workspace_cleanup.step() then drives queries via run_until_complete on the controller loop against a pool whose waiter Queue belongs to the uvicorn loop, and its exception handler pins _is_healthy = False.

Note: while the pool has idle connections the cross-loop checkout happens to succeed (no Queue wait), which is why lightly loaded deployments usually look fine. The defect surfaces exactly under pool contention.

Reproduction (standalone, deterministic — no running platform needed)

Using the product modules in a venv:

  1. Loop A (simulating the uvicorn lifespan): initialize_async_engine(cfg), check out all pool_size + max_overflow connections and hold them (simulating concurrent API load), and leave one waiter queued so the pool Queue's futures bind to loop A.
  2. Loop B (simulating the controller thread): call initialize_async_engine(cfg) again (returns early — singleton reused), then run one repository query via loop_b.run_until_complete(...) exactly like workspace_cleanup.step() does.

Result:

RuntimeError: <Queue at 0x... maxsize=5> is bound to a different event loop

Observed identically in production logs under concurrent regression load (three suites), with the pool queue showing _getters[28] tasks=172977 and /status reporting:

{"healthy": false, "status": {"job_scheduler": true, "job_reconciler": true,
 "models_controller": true, "workspace_cleanup": false}}

After a clean restart (controller loop wins the race) the same suite passes with workspace_cleanup: true throughout — confirming the race dependence.

Expected behavior
  • Background controllers keep a valid DB path regardless of initialization order.
  • workspace_cleanup stays healthy under concurrent API load.
  • A single stuck background janitor should not flip the deployment-wide controllers.healthy signal that external monitors gate on.
Suggested fix directions
  • Make initialize_async_engine loop-aware: record the owning loop and either create per-loop engines/session-makers or raise on foreign-loop reuse so the misconfiguration is visible at startup; or
  • Run the entities controller's DB work on the loop that owns the engine (asyncio.run_coroutine_threadsafe onto the server loop), which is what the controller comment already intends; and
  • Consider reporting per-controller health separately from the hard aggregate.
Environment
  • main @ 0c4dc810cc519ed2b266152e79b4378f7f07de7f (also reproduced from source at that revision)
  • Single nemo services run process, SQLite entities DB, AsyncAdaptedQueuePool size=5 max_overflow=10
  • Internal tracking: NVBUG 6588975

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with services/core/entities/src/nmp/core/entities/app/repository/init.py and service.py to trace engine initialization, then inspect controllers/main.py and packages/nmp_platform_runner/src/nmp/platform_runner/server.py for loop ownership. Run the standalone two-loop reproduction under pool contention. Done means initialization order no longer leaves workspace_cleanup using a foreign-loop pool, and the controller remains healthy during concurrent load.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, sqlalchemy, sqlite
Domain
backend, databases, observability
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.