E3SM-Project / E3SM-Project/E3SM

Possible hang in ATM init with ne120 test on pm-gpu `ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu.eamxx-L128--eamxx-prod`

Open
#8,563 2 comments 0 reactions 0 assignees View on GitHub
EAMxx pm-gpu
Dominant language
Fortran
Stars
440
Forks
481
Avg merge
4d 6h
Merged PRs (30d)
36

Description

With the test `[ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu.eamxx-L128--eamxx-prod](https://my.cdash.org/tests/425251873)`, we see it fail a couple of times recently, timing out after 45 minutes.

While it may be some issue on machine-side, it looks like it could also be a hang on our side in ATM init.

e3sm.log
```
0: number of MPI processes per node: min,max= 4 4
0:
0: application HM_COARSE with ID = 258769486 and external id: 120 is registered now
0: application HM_FINE with ID = 719764396 and external id: 119 is registered now
0: application HM_PGX with ID = 35914825 and external id: 5 is registered now
0: application with ID: 719764396 global id: 119 name: HM_FINE has 1 already defined tags out of 1 tags
0: application with ID: 258769486 global id: 120 name: HM_COARSE has 1 already defined tags out of 1 tags
0: application with ID: 35914825 global id: 5 name: HM_PGX has 1 already defined tags out of 1 tags
0: Note: nsplit=-1, while nsplit must be >=1. We know SCREAM does not know nsplit until runtime, so this is fine.
0: Make sure nsplit is set to a valid value before calling prim_advance_subcycle!
0: application ATM_PHYS_SCREAM with ID = 997430475 and external id: 205 is registered now
0: moab_atm_phys_scream:: register MOAB app:ATM_PHYS_SCREAM mphaid= 997430475
0: application with ID: 997430475 global id: 205 name: ATM_PHYS_SCREAM has 1 already defined tags out of 1 tags
srun: Job step aborted: Waiting up to 32 seconds for job step to finish.
0: [2026-07-15T08:02:01.367] error: *** STEP 55924314.0 ON nid001309 CANCELLED AT 2026-07-15T08:02:01 DUE TO TIME LIMIT ***
```

cpl.log:
```
(cpl7moab_driver) : Initialize each component: atm, lnd, rof, ocn, ice, glc, wav, esp, iac
(component_init_cc:moab) : Initialize component iac
(component_init_cc:moab) : Initialize component atm
```

Here we see it has run 26 time sin last 4 weeks with 2 fails that were timeouts, but all the others with 18-20 min (well except one that was 28min)
```
perlmutter-login36% sacct --start=now-4weeks -X --format=End,JobID,JobName%60,NNodes,Elapsed,Timelimit,Account,State -a -u e3smtest | grep ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1
2026-06-19T00:32:37 54690959 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:12 00:45:00 e3sm_g COMPLETED
2026-06-19T22:38:35 54736130 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:04 00:44:00 e3sm_g COMPLETED
2026-06-20T21:15:10 54774518 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:22 00:46:00 e3sm_g COMPLETED
2026-06-22T11:37:43 54816061 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:45 00:44:00 e3sm_g COMPLETED
2026-06-23T03:15:26 54862694 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:17 00:45:00 e3sm_g COMPLETED
2026-06-23T21:21:20 54932797 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:13 00:46:00 e3sm_g COMPLETED
2026-06-25T00:24:01 54993171 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:13 00:44:00 e3sm_g COMPLETED
2026-06-25T22:13:16 55053021 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:16 00:44:00 e3sm_g COMPLETED
2026-06-27T04:58:28 55114087 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:14 00:44:00 e3sm_g COMPLETED
2026-06-27T21:38:48 55161861 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:13 00:44:00 e3sm_g COMPLETED
2026-06-28T21:58:38 55208591 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:15 00:44:00 e3sm_g COMPLETED
2026-06-29T21:59:21 55289245 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:28:37 00:44:00 e3sm_g COMPLETED
2026-07-01T21:43:50 55375747 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:19 01:00:00 e3sm_g COMPLETED
2026-07-02T21:24:30 55416083 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:34 00:46:00 e3sm_g COMPLETED
2026-07-04T00:45:34 55461627 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:19:50 00:45:00 e3sm_g COMPLETED
2026-07-04T22:28:25 55511135 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:20:11 00:47:00 e3sm_g COMPLETED
2026-07-06T00:40:51 55550923 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:20:22 00:47:00 e3sm_g COMPLETED
2026-07-07T00:24:18 55617424 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:36 00:47:00 e3sm_g COMPLETED
2026-07-08T03:23:28 55667863 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:35 00:45:00 e3sm_g COMPLETED
2026-07-09T10:02:48 55704672 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:48 00:45:00 e3sm_g COMPLETED
2026-07-09T23:39:17 55736674 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:24 00:45:00 e3sm_g COMPLETED
2026-07-10T22:05:05 55784446 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:45:16 00:45:00 e3sm_g TIMEOUT
2026-07-12T00:28:42 55809325 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:32 00:45:00 e3sm_g COMPLETED
2026-07-12T21:26:06 55841525 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:32 00:45:00 e3sm_g COMPLETED
2026-07-13T23:34:16 55882692 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:18:31 00:45:00 e3sm_g COMPLETED
2026-07-15T01:02:01 55924314 test.ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu+ 8 00:45:01 00:45:00 e3sm_g TIMEOUT

```

So the hang is localized to **EAMxx's `set_initial_conditions` step**, which is where SCORPIO/PIO reads the initial condition (and SPA/topo) files and does the interpolation/remap onto the run grid. Since this happens intermittently (~1/15 runs) at `ne120` scale with 32 MPI ranks across 8 GPU nodes, this smells like a collective I/O or MPI communication race in that init path.

Contributor guide

Open the contributing guide

Research direction

Start by reproducing the CDASH test ERS_Lh6.ne120pg2_ne120pg2.F2010-SCREAMv1.pm-gpu_gnugpu.eamxx-L128--eamxx-prod and inspect e3sm.log and cpl.log around EAMxx's set_initial_conditions step. Trace the SCORPIO/PIO initial-condition, SPA, and topo reads and the associated interpolation/remap work at 32 MPI ranks. Done means the test no longer intermittently times out during ATM initialization.

Written by the indexing model from the issue text.

Assessment

Tech stack
fortran
Domain
distributed-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.