[FnApiRunner]multi-process runner does not terminate cleanly upon receiving SIGINT
- Dominant language
- Java
- Stars
- 8.7k
- Forks
- 4.7k
- Avg merge
- 2d 5h
- Merged PRs (30d)
- 204
Description
The multi-process runner does not handle SIGINT gracefully. To reproduce, run wordcount.py using the "Run with multiprocessing mode" instructions from the first comment in BEAM-3645 (in Python 3).
Expected: wordcount terminates gracefully when Ctrl-C is pressed during pipeline execution (similarly to default direct runner)
Actual: wordcount hangs forever after printing the following once per worker:
```
Exception in thread run_worker:
Traceback (most recent call last):
File "/usr/lib/python3.6/threading.py",
line 916, in _bootstrap_inner
self.run()
File "/usr/lib/python3.6/threading.py", line 864, in
run
self._target(*self._args, **self._kwargs)
File "/usr/local/google/home/yifanmai/venv/wordcount/lib/python3.6/site-packages/apache_beam/runners/portability/local_job_service.py",
line 216, in run
'Worker subprocess exited with return code %s' % p.returncode)
RuntimeError:
Worker subprocess exited with return code 1
```
Imported from Jira [BEAM-8149](https://issues.apache.org/jira/browse/BEAM-8149). Original Jira may contain additional context.
Reported by: hannahjiang.
Contributor guide
Research direction
Reproduce the hang with wordcount.py using the multiprocessing instructions from BEAM-3645, then inspect local_job_service.py around the run_worker traceback and worker subprocess handling. Done means Ctrl-C during pipeline execution exits cleanly without hanging, matching the direct runner behavior.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- backend, distributed-systems
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100