MagicStack / MagicStack/uvloop

`connect_read_pipe` double-closes the underlying fd, possibly closing an unrelated fd

Open
#763 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Cython
Stars
11.9k
Forks
616
PR merge metrics
No merged PRs in 30d

Description

In our production use of uvloop, we found sporadic cases of file descriptors being randomly closed. (For more details of our specific case, see skypilot-org/skypilot#10681).
We traced this back to loop.connect_read_pipe double-closing its file descriptor. When the fd was re-allocated between the first and second close, this caused an unrelated thread to have its fd closed out from under.

I used this script to check the double-close behavior:

import asyncio, io, os, sys

class BufferedReaderWithProbe(io.BufferedReader):
    def close(self):
        # Probe before status
        fd = self.fileno()
        print(f"closing {self} (fd {fd})")
        try:
            os.fstat(fd); fd_state = "still OPEN"
        except OSError as e:
            fd_state = f"already closed ({e.strerror})"
        # Open a new pipe to see what fds we get
        # Kernel should reuse lowest free number
        newfd_reader, newfd_writer = os.pipe()
        print(f"  is fd {fd} already closed? {fd_state}")
        print(f"  fresh os.pipe() got fds {newfd_reader},{newfd_writer}")

        # Do normal BufferedReader close
        try:
            print("calling BufferedReader.close")
            super().close()
        finally:

            # Probe after status
            print("Status after close:")
            for fd in (newfd_reader, newfd_writer):
                try:
                    os.fstat(fd); print(f"  fresh fd {fd}: alive")
                except OSError as e:
                    print(f"  fresh fd {fd}: DEAD ({e.strerror})  <-- stolen")

async def main():
    loop = asyncio.get_running_loop()
    pipe_read_fd, pipe_write_fd = os.pipe()
    probe_reader = BufferedReaderWithProbe(io.FileIO(pipe_read_fd, "rb"))
    transport, _ = await loop.connect_read_pipe(asyncio.Protocol, probe_reader)
    print(f"transport {type(transport).__name__} on fd {pipe_read_fd}")
    transport.close()

if __name__ == "__main__":
    if sys.argv[1:] == ["uvloop"]:
        import uvloop; uvloop.install(); print(f"== uvloop {uvloop.__version__}")
    else:
        print(f"== stdlib asyncio {sys.version.split()[0]}")
    asyncio.run(main())
$ python repro.py
== stdlib asyncio 3.10.16
transport _UnixReadPipeTransport on fd 6
closing <BufferedReaderWithProbe name=6> (fd 6)
  is fd 6 already closed? still OPEN
  fresh os.pipe() got fds 8,9
calling BufferedReader.close
Status after close:
  fresh fd 8: alive
  fresh fd 9: alive
$ python repro.py uvloop
== uvloop 0.22.1
transport ReadUnixTransport on fd 13
closing <BufferedReaderWithProbe name=13> (fd 13)
  is fd 13 already closed? already closed (Bad file descriptor)
  fresh os.pipe() got fds 13,16
calling BufferedReader.close
Status after close:
  fresh fd 13: DEAD (Bad file descriptor)  <-- stolen
  fresh fd 16: alive

Here you can see that a asyncio event loop correctly leaves fd 6 open, allowing the underlying file object (in this case, our custom BufferedReader) to close the fd.
However, uvloop has already closed the fd, so when the file object is closed, it re-closes fd 13 which is now an unrelated os.pipe() fd.

This specific case was already fixed for sockets: https://github.com/MagicStack/uvloop/commit/d5195d7c10fbae81ef1dcb4609cee50f8aa746fa. It was assumed that the EBADF for non-sockets was benign, but it can actually be a problem if the underlying fd is reassigned before the _fileobj.close() can run.

You can observe the same behavior if the transport is not explicitly closed but is instead GC'd - although in this case it seems that the file object is closed before the uv_close call double-closes the fd. (The order is inverted.) This is actually the case we saw in prod.

It also looks like connect_write_pipe is similarly affected.

AI disclosure: This investigation was heavily assisted by AI tools, but this report was entirely human written without assistance except for the original repro script. The original repro script was then modified by me for clarity and manually validated.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start at the connect_read_pipe and connect_write_pipe transport entry points and compare their file-descriptor cleanup with the socket fix in commit d5195d7c10fbae81ef1dcb4609cee50f8aa746fa. Run the supplied repro with uvloop and the standard asyncio loop. Done means closing or garbage-collecting either pipe transport does not close an unrelated file descriptor.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
networking
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Active
Clarity
Clearly specified
Newbie friendliness
68/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.