E3SM-Project / E3SM-Project/E3SM
Possible hang in ATM init with ne120 test on pm-gpu `ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu.eamxx-L128--eamxx-prod`
- Dominant language
- Fortran
- Stars
- 440
- Forks
- 481
- Avg merge
- 4d 6h
- Merged PRs (30d)
- 36
Description
With the test `[ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu.eamxx-L128--eamxx-prod](https://my.cdash.org/tests/425251873)`, we see it fail a couple of times recently, timing out after 45 minutes.
While it may be some issue on machine-side, it looks like it could also be a hang on our side in ATM init.
e3sm.log
```
0: number of MPI processes per node: min,max= 4 4
0:
0: application HM_COARSE with ID = 258769486 and external id: 120 is registered now
0: application HM_FINE with ID = 719764396 and external id: 119 is registered now
0: application HM_PGX with ID = 35914825 and external id: 5 is registered now
0: application with ID: 719764396 global id: 119 name: HM_FINE has 1 already defined tags out of 1 tags
0: application with ID: 258769486 global id: 120 name: HM_COARSE has 1 already defined tags out of 1 tags
0: application with ID: 35914825 global id: 5 name: HM_PGX has 1 already defined tags out of 1 tags
0: Note: nsplit=-1, while nsplit must be >=1. We know SCREAM does not know nsplit until runtime, so this is fine.
0: Make sure nsplit is set to a valid value before calling prim_advance_subcycle!
0: application ATM_PHYS_SCREAM with ID = 997430475 and external id: 205 is registered now
0: moab_atm_phys_scream:: register MOAB app:ATM_PHYS_SCREAM mphaid= 997430475
0: application with ID: 997430475 global id: 205 name: ATM_PHYS_SCREAM has 1 already defined tags out of 1 tags
srun: Job step aborted: Waiting up to 32 seconds for job step to finish.
0: [2026-07-15T08:02:01.367] error: *** STEP 55924314.0 ON nid001309 CANCELLED AT 2026-07-15T08:02:01 DUE TO TIME LIMIT ***
```
cpl.log:
```
(cpl7moab_driver) : Initialize each component: atm, lnd, rof, ocn, ice, glc, wav, esp, iac
(component_init_cc:moab) : Initialize component iac
(component_init_cc:moab) : Initialize component atm
```
Here we see it has run 26 time sin last 4 weeks with 2 fails that were timeouts, but all the others with 18-20 min (well except one that was 28min)
```
perlmutter-login36% sacct --start=now-4weeks -X --format=End,JobID,JobName%60,NNodes,Elapsed,Timelimit,Account,State -a -u e3smtest | grep ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1
2026-06-19T00:32:37 54690959 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:12 00:45:00 e3sm_g COMPLETED
2026-06-19T22:38:35 54736130 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:04 00:44:00 e3sm_g COMPLETED
2026-06-20T21:15:10 54774518 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:22 00:46:00 e3sm_g COMPLETED
2026-06-22T11:37:43 54816061 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:45 00:44:00 e3sm_g COMPLETED
2026-06-23T03:15:26 54862694 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:17 00:45:00 e3sm_g COMPLETED
2026-06-23T21:21:20 54932797 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:13 00:46:00 e3sm_g COMPLETED
2026-06-25T00:24:01 54993171 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:13 00:44:00 e3sm_g COMPLETED
2026-06-25T22:13:16 55053021 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:16 00:44:00 e3sm_g COMPLETED
2026-06-27T04:58:28 55114087 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:14 00:44:00 e3sm_g COMPLETED
2026-06-27T21:38:48 55161861 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:13 00:44:00 e3sm_g COMPLETED
2026-06-28T21:58:38 55208591 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:15 00:44:00 e3sm_g COMPLETED
2026-06-29T21:59:21 55289245 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:28:37 00:44:00 e3sm_g COMPLETED
2026-07-01T21:43:50 55375747 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:19 01:00:00 e3sm_g COMPLETED
2026-07-02T21:24:30 55416083 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:34 00:46:00 e3sm_g COMPLETED
2026-07-04T00:45:34 55461627 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:50 00:45:00 e3sm_g COMPLETED
2026-07-04T22:28:25 55511135 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:20:11 00:47:00 e3sm_g COMPLETED
2026-07-06T00:40:51 55550923 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:20:22 00:47:00 e3sm_g COMPLETED
2026-07-07T00:24:18 55617424 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:36 00:47:00 e3sm_g COMPLETED
2026-07-08T03:23:28 55667863 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:35 00:45:00 e3sm_g COMPLETED
2026-07-09T10:02:48 55704672 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:48 00:45:00 e3sm_g COMPLETED
2026-07-09T23:39:17 55736674 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:24 00:45:00 e3sm_g COMPLETED
2026-07-10T22:05:05 55784446 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:45:16 00:45:00 e3sm_g TIMEOUT
2026-07-12T00:28:42 55809325 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:32 00:45:00 e3sm_g COMPLETED
2026-07-12T21:26:06 55841525 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:32 00:45:00 e3sm_g COMPLETED
2026-07-13T23:34:16 55882692 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:31 00:45:00 e3sm_g COMPLETED
2026-07-15T01:02:01 55924314 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:45:01 00:45:00 e3sm_g TIMEOUT
```
So the hang is localized to **EAMxx's `set_initial_conditions` step**, which is where SCORPIO/PIO reads the initial condition (and SPA/topo) files and does the interpolation/remap onto the run grid. Since this happens intermittently (~1/15 runs) at `ne120` scale with 32 MPI ranks across 8 GPU nodes, this smells like a collective I/O or MPI communication race in that init path.
Contributor guide
Research direction
Start by reproducing the CDASH test ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu.eamxx-L128--eamxx-prod and inspect e3sm.log and cpl.log around EAMxx's set_initial_conditions step. Trace the SCORPIO/PIO initial-condition, SPA, and topo reads and the associated interpolation/remap work at 32 MPI ranks. Done means the test no longer intermittently times out during ATM initialization.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- fortran
- Domain
- distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Quiet
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100