NVIDIA / NVIDIA/cloudai

No single source of truth for "agent-driven" run; --single-sbatch conflates scheduling with search strategy

Open
#944 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
99
Forks
62
Avg merge
6d 12h
Merged PRs (30d)
17

Description

Summary

env_params are only resolved in the agent loop (CloudAIGymEnv.step), so their validity depends on one fact: will this run be agent-driven? Today that fact has no single source of truth — it's scattered across tr.is_dse_job (config), agent.samples_env_params (config), and args.single_sbatch (CLI), reconciled nowhere.

Symptom

A config with env_params + an env-aware (RL) agent passes validate_dse_env_params (it's is_dse_job and the agent samples). But if run with --single-sbatch, dispatch routes to the grid-unroll path, which calls apply_params_set(combination) with no env_params. The env_params are silently dropped (and the field is left as an unresolved list heading into command-gen).

# src/cloudai/cli/handlers.py  (mode decision)
has_dse = any(tr.is_dse_job for tr in test_scenario.test_runs)
if args.single_sbatch or not has_dse:   # <-- single_sbatch forces grid unroll
    handle_non_dse_job(runner, args)

Root cause (two coupled flaws)

  1. Missing abstraction: there is no single is_agent_driven concept. env_params validity (and dispatch) should gate on it, computed once.
  2. --single-sbatch overloads scheduling with search strategy. It is a scheduling/packaging concern (cram cases into one sbatch) but currently forces the grid strategy, overriding whatever agent the config declared. Scheduling should be orthogonal to env / action-space / search strategy. In future --single-sbatch should support agent-driven runs too (e.g. a genetic algorithm launching multiple evaluations in parallel), where env_params are perfectly valid.

Why not a quick guard

Rejecting env_params when args.single_sbatch (the obvious patch) bakes in the exact coupling we want to remove: the day --single-sbatch supports agent-driven runs, env_params would be valid there, yet the guard would still reject them. The fix belongs in the model, not the run handler.

Direction

  • Introduce a single source of truth for agent-driven execution (config/agent capability); gate both dispatch and env_params validation on it.
  • Decouple --single-sbatch from search strategy so it only affects scheduling/packaging and composes with both grid and agent-driven runs.

Pointers

  • src/cloudai/cli/handlers.pyhandle_dry_run_and_run mode decision.
  • src/cloudai/configurator/env_params.pyvalidate_dse_env_params.
  • src/cloudai/systems/slurm/single_sbatch_runner.py — grid unroll calling apply_params_set without env_params.

Surfaced by the env_params work in #901; the underlying scheduling/strategy coupling predates it.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Read handle_dry_run_and_run in src/cloudai/cli/handlers.py, then compare validate_dse_env_params in src/cloudai/configurator/env_params.py with the grid-unroll path in src/cloudai/systems/slurm/single_sbatch_runner.py. Trace how agent capability, dispatch, and env_params are currently decided. Done means agent-driven execution has one source of truth and --single-sbatch controls scheduling without forcing the search strategy.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
cli, infrastructure
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.