llnl / llnl/mpiP

Possible data corruption for Fortran in the presence of MPI_IN_PLACE

Open
#46 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
C
Stars
112
Forks
35
PR merge metrics
No merged PRs in 30d

Description

A range of MPI operations allow to reuse send buffers as receive buffers by setting the send buffer to a special constant `MPI_IN_PLACE`. With Fortran applications this can lead to data corruption if executed with mpiP.

The corruption can be demonstrated with this simple code:
```
PROGRAM sample_allreduce
USE mpi
IMPLICIT NONE

INTEGER :: ierr
INTEGER :: rank, rank_in_place
INTEGER :: rank_sum

CALL MPI_Init(ierr)
CALL MPI_Comm_rank(MPI_COMM_WORLD, rank, ierr)

rank_in_place = rank

PRINT *, 'Rank: ', rank
CALL MPI_Allreduce(rank, rank_sum, 1, MPI_INT, MPI_SUM, MPI_COMM_WORLD, ierr)
CALL MPI_Allreduce(MPI_IN_PLACE, rank_in_place, 1, MPI_INT, MPI_SUM, MPI_COMM_WORLD, ierr)
PRINT *, 'Sum: ', rank_sum, ' - ', rank_in_place

CALL MPI_Finalize(ierr)
END PROGRAM sample_allreduce
```

Executing without mpiP instrumentation leads to the expected output
```
$ mpirun -np 3 ./a.out
Rank: 0
Sum: 3 - 3
Rank: 1
Sum: 3 - 3
Rank: 2
Sum: 3 - 3
```
while executing with mpiP corrupts the data as
```
$ mpirun -np 3 env LD_PRELOAD=$HLRS_MPIP_ROOT/lib/libmpiP.so ./a.out
mpiP:
mpiP: mpiP V3.5.0 (Build Mar 16 2023/14:16:24)
mpiP:
Rank: 0
Sum: 3 - 0
Rank: 1
Sum: 3 - 0
Rank: 2
Sum: 3 - 0
mpiP:
mpiP: Storing mpiP output in [./a.out.3.1905013.1.mpiP].
mpiP:
```
Note, that the second column (which used `MPI_IN_PLACE`) is "0" while it should be "3".

I guess that the underlying problem is missing or incorrect treatment of constants such as `MPI_IN_PLACE` in the transition from Fortran to C PMPI interfaces. A similar problem has been observed in other projects / tools using PMPI such as [here](https://github.com/PRUNERS/ReMPI/issues/8). In fact, the code above is taken from that issue.

I have observed this behavior for mpiP v3.4.1 and v3.5 using GCC v10.2 with either OpenMPI v4.1.4 or HPE's MPI implementation MPT 2.26.

Also note, that the code runs correctly when replacing `use mpi` with `use mpi_f08`, at least for OpenMPI (but not for MPT).

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the provided sample_allreduce Fortran reproducer and compare runs with and without mpiP instrumentation. Trace the MPI_IN_PLACE call through the Fortran-to-C PMPI transition, using the linked ReMPI issue as context. Done means the MPI_IN_PLACE reduction reports 3 in the second column under mpiP for the stated MPI implementations.

Written by the indexing model from the issue text.

Assessment

Tech stack
c, fortran
Domain
distributed-systems, performance
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.