the program still deadlocks on python3 but it works perfectly on python2, anyone know what changed between implementations that could be triggering this issue?
the program still deadlocks on python3 but it works perfectly on python2, anyone know what changed between implementations that could be triggering this issue?
For some reason on OP's computer that probably causes an appearance of the program hanging, but really, it will just crash after a while exhausting some resource.
Here's the gdb stack trace:
#0 __futex_abstimed_wait_common64 (private=<optimized out>, cancel=true, abstime=0x0, op=393, expected=0, futex_word=0x17bc7c0) at ./nptl/futex-internal.c:57
#1 __futex_abstimed_wait_common (cancel=true, private=<optimized out>, abstime=0x0, clockid=0, expected=0, futex_word=0x17bc7c0) at ./nptl/futex-internal.c:87
#2 __GI___futex_abstimed_wait_cancelable64 (futex_word=futex_word@entry=0x17bc7c0, expected=expected@entry=0, clockid=clockid@entry=0, abstime=abstime@entry=0x0,
private=<optimized out>) at ./nptl/futex-internal.c:139
#3 0x00007ff326c9cc5f in do_futex_wait (sem=sem@entry=0x17bc7c0, abstime=0x0, clockid=0) at ./nptl/sem_waitcommon.c:111
#4 0x00007ff326c9ccf8 in __new_sem_wait_slow64 (sem=0x17bc7c0, abstime=0x0, clockid=0) at ./nptl/sem_waitcommon.c:183
#5 0x00007ff326c9cd71 in __new_sem_wait (sem=<optimized out>) at ./nptl/sem_wait.c:42
#6 0x000000000042765b in PyThread_acquire_lock_timed (lock=0x17bc7c0, microseconds=-1, intr_flag=0) at ../Python/thread_pthread.h:483
#7 0x0000000000625813 in _enter_buffered_busy (self=0x7ff326f330f0) at ../Modules/_io/bufferedio.c:281
#8 0x000000000045f9b3 in buffered_flush (self=0x7ff326f330f0, args=<optimized out>) at ../Modules/_io/bufferedio.c:825
#9 0x0000000000524185 in method_vectorcall_NOARGS (func=func@entry=<method_descriptor at remote 0x7ff326f611d0>, args=args@entry=0x7fffa8bc7a38, nargsf=<optimized out>,
kwnames=kwnames@entry=0x0) at ../Objects/descrobject.c:436
#10 0x000000000053bd5a in _PyObject_VectorcallTstate (kwnames=0x0, nargsf=<optimized out>, args=0x7fffa8bc7a38, callable=<method_descriptor at remote 0x7ff326f611d0>,
tstate=0x17bf410) at ../Include/cpython/abstract.h:118
#11 PyObject_VectorcallMethod (name=<optimized out>, args=0x7fffa8bc7a38, nargsf=<optimized out>, kwnames=0x0) at ../Objects/call.c:828
#12 0x000000000060283a in _PyObject_CallMethodIdNoArgs (name=0x8f6120 <PyId_flush.lto_priv.2>, self=<optimized out>) at ../Include/cpython/abstract.h:243
#13 _io_TextIOWrapper_flush_impl (self=0x7ff326f40040) at ../Modules/_io/textio.c:3038
#14 _io_TextIOWrapper_flush (self=0x7ff326f40040, _unused_ignored=<optimized out>) at ../Modules/_io/clinic/textio.c.h:685Anyways, another common pitfall in multiprocessing is attempting to serialize multithreading / multiprocessing primitives s.a. locks, variables or mutexes. My memory may fail me, but, I think, it may result in deadlock too. I think, multiprocessing code tries to guard against it, but there are some weird rules for when it's OK for serialized objects to have those primitives (I think, initialization in __init__ is fine, but not so much otherwise or something like that), but the check isn't very good / just a heuristic... But, really, I don't remember this part well.
1. Writing to stderr grabs a lock.
2. Part of the multiprocessing code (perhaps not present in Python 2) also grabs this lock.
3. If you fork at the right moment (which is quite likely with the loop) the lock is held by a thread that is now dead, and so now you're waiting for a lock to release that will never be released.