pytest-dev / pytest-dev/execnet
Unhandled error in sys.excepthook
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 102
- Forks
- 47
- PR merge metrics
- No merged PRs in 30d
Description
- Bitbucket: https://bitbucket.org/hpk42/execnet/issue/30
- Originally reported by: @alfredodeza
- Originally created at: 2014-02-10T16:05:34.775
After closing the connections we are seeing a few thread errors that seem impossible to get rid of. At the point where they are sent to stderr there is nothing left running and the code has closed the gateway.
Unhandled exception in thread started by
Error in sys.excepthook:
Original exception was:
Running execnet with debug show this:
[53488] gw0 [receiver-thread] RECEIVERTHREAD: starting to run
[53488] gw0 sent <Message.CHANNEL_EXEC channelid=1 len=6354>
[53488] gw0 sent <Message.CHANNEL_DATA channelid=1 'platform_information()'>
[53488] gw0 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 ('Ubuntu', '12.04', 'precise')>
[53488] gw0 sent <Message.CHANNEL_DATA channelid=1 'machine_type()'>
[53488] gw0 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 'x86_64'>
[53488] gw0 sent <Message.CHANNEL_DATA channelid=1 'shortname()'>
[53488] gw0 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 'node1'>
[53488] gw0 sent <Message.CHANNEL_DATA channelid=1 len=324>
[53488] gw0 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 None>
[53488] gw0 sent <Message.CHANNEL_DATA channelid=1 len=127>
[53488] gw0 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 None>
[53488] gw0 gateway.exit() called
[53488] gw0 --> sending GATEWAY_TERMINATE
[53488] gw0 --> io.close_write
[53488] gw0 [receiver-thread] EOF without prior gateway termination message
[53488] gw0 [receiver-thread] entering finalization
[53488] gw0 finished receiving
[53488] gw0 [receiver-thread] terminating execution
[53488] gw0 [receiver-thread] closing read
[53488] gw0 [receiver-thread] closing write
[53488] gw0 [receiver-thread] leaving finalization
[53488] gw1 [receiver-thread] RECEIVERTHREAD: starting to run
[53488] gw1 sent <Message.CHANNEL_EXEC channelid=1 len=6354>
[53488] gw1 sent <Message.CHANNEL_DATA channelid=1 'platform_information()'>
[53488] gw1 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 ('CentOS', '6.4', 'Final')>
[53488] gw1 sent <Message.CHANNEL_DATA channelid=1 'machine_type()'>
[53488] gw1 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 'x86_64'>
[53488] gw1 sent <Message.CHANNEL_DATA channelid=1 'shortname()'>
[53488] gw1 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 'node2'>
[53488] gw1 sent <Message.CHANNEL_DATA channelid=1 len=324>
[53488] gw1 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 None>
[53488] gw1 sent <Message.CHANNEL_DATA channelid=1 len=127>
[53488] gw1 [receiver-thread] received <Message.CHANNEL_DATA channelid=1 None>
[53488] gw1 gateway.exit() called
[53488] gw1 --> sending GATEWAY_TERMINATE
[53488] gw1 --> io.close_write
[53488] === atexit cleanup <Group []> ===
[53488] gw0 1 channel.__del__
[53488] gw1 1 channel.__del__
Unhandled exception in thread started by
Error in sys.excepthook:
Original exception was:
It looks as though this is happening in the threading module in Python itself. I am not sure something is possible to eat up these messages.
There is a ticket open in the Python bug tracker that mentions this problem is closed:
http://bugs.python.org/issue1722344
One of the patches that one commenter added looks like what we would need to get rid of them (http://bugs.python.org/file9356/1722344_squelch_exception.patch)
--- /usr/lib/python2.5/threading.py.orig 2008-02-04 13:58:18.000000000 -0600
+++ ./threading.py 2008-02-05 11:55:33.000000000 -0600
@@ -472,34 +472,18 @@
_sys.stderr.write("Exception in thread %s:\n%s\n" %
(self.getName(), _format_exc()))
else:
- # Do the best job possible w/o a huge amt. of code to
- # approximate a traceback (code ideas from
- # Lib/traceback.py)
- exc_type, exc_value, exc_tb = self.__exc_info()
- try:
- print>>self.__stderr, (
- "Exception in thread " + self.getName() +
- " (most likely raised during interpreter shutdown):")
- print>>self.__stderr, (
- "Traceback (most recent call last):")
- while exc_tb:
- print>>self.__stderr, (
- ' File "%s", line %s, in %s' %
- (exc_tb.tb_frame.f_code.co_filename,
- exc_tb.tb_lineno,
- exc_tb.tb_frame.f_code.co_name))
- exc_tb = exc_tb.tb_next
- print>>self.__stderr, ("%s: %s" % (exc_type, exc_value))
- # Make sure that exc_tb gets deleted since it is a memory
- # hog; deleting everything else is just for thoroughness
- finally:
- del exc_type, exc_value, exc_tb
+ # If _sys is missing, then the interpreter is shutting
+ # down and the thread should no longer exist. If this
+ # happens, ignore the error and exit gracefully.
+ pass
else:
if __debug__:
self._note("%s.__bootstrap(): normal return", self)
finally:
- self.__stop()
+ # Exceptions will also be raised during stop/delete if the
+ # interpreter is shutting down. Ignore these as well.
try:
+ self.__stop()
self.__delete()
except:
pass
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the shutdown sequence with execnet debug output, focusing on gateway.exit(), receiver-thread finalization, and the atexit cleanup shown in the report. Compare the resulting errors with Python's threading shutdown behavior and the linked threading.py patch. Done means interpreter-shutdown cleanup no longer emits the unhandled thread and sys.excepthook errors.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100