facebookresearch / facebookresearch/sam2
Backend worker did not respond in given time
- Dominant language
- Jupyter Notebook
- Stars
- 19.9k
- Forks
- 2.5k
- PR merge metrics
- No merged PRs in 30d
Description
I encountered an issue while attempting to segment multiple images in a loop. The error I am receiving is:
"Backend worker did not respond in the given time."
org.pytorch.serve.wlm.WorkerThread - 9000 Worker disconnected. WORKER_MODEL_LOADED
It appears that GPU memory is not being cleared properly, and the available variable space is diminishing with each image processed.
Interestingly, the code was functioning correctly before the latest commits made on December 12, 2024, and December 15, 2024. These recent changes might have introduced the problem.
2025-01-15T10:16:06.329-08:002025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Number or consecutive unsuccessful inference 1 | 2025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Number or consecutive unsuccessful inference 1 | AllTraffic/i-098f59f854f0cab5b
-- | -- | --
| 2025-01-15T10:16:06.329-08:002025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Backend worker error | 2025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Backend worker error | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00org.pytorch.serve.wlm.WorkerInitializationException: Backend worker did not respond in given time | org.pytorch.serve.wlm.WorkerInitializationException: Backend worker did not respond in given time | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00#011at org.pytorch.serve.wlm.WorkerThread.run(WorkerThread.java:247) [model-server.jar:?] | #011at org.pytorch.serve.wlm.WorkerThread.run(WorkerThread.java:247) [model-server.jar:?] | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00#011at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) [?:?] | #011at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) [?:?] | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00#011at java.util.concurrent.FutureTask.run(FutureTask.java:264) [?:?] | #011at java.util.concurrent.FutureTask.run(FutureTask.java:264) [?:?] | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00#011at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) [?:?] | #011at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) [?:?] | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00#011at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) [?:?] | #011at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) [?:?] | AllTraffic/i-098f59f854f0cab5b
| 2025-01-15T10:16:06.329-08:00#011at java.lang.Thread.run(Thread.java:840) [?:?]
2025-01-15T10:16:06.329-08:00
2025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Number or consecutive unsuccessful inference 1
2025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Number or consecutive unsuccessful inference 1
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581186)
2025-01-15T10:16:06.329-08:00
2025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Backend worker error
2025-01-15T18:16:06,223 [ERROR] W-9000-model_1.0 org.pytorch.serve.wlm.WorkerThread - Backend worker error
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581187)
2025-01-15T10:16:06.329-08:00
org.pytorch.serve.wlm.WorkerInitializationException: Backend worker did not respond in given time
org.pytorch.serve.wlm.WorkerInitializationException: Backend worker did not respond in given time
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581188)
2025-01-15T10:16:06.329-08:00
#011at org.pytorch.serve.wlm.WorkerThread.run(WorkerThread.java:247) [model-server.jar:?]
#011at org.pytorch.serve.wlm.WorkerThread.run(WorkerThread.java:247) [model-server.jar:?]
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581189)
2025-01-15T10:16:06.329-08:00
#011at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) [?:?]
#011at java.util.concurrent.Executors$RunnableAdapter.call(Executors.java:539) [?:?]
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581190)
2025-01-15T10:16:06.329-08:00
#011at java.util.concurrent.FutureTask.run(FutureTask.java:264) [?:?]
#011at java.util.concurrent.FutureTask.run(FutureTask.java:264) [?:?]
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581191)
2025-01-15T10:16:06.329-08:00
#011at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) [?:?]
#011at java.util.concurrent.ThreadPoolExecutor.runWorker(ThreadPoolExecutor.java:1136) [?:?]
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581192)
2025-01-15T10:16:06.329-08:00
#011at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) [?:?]
#011at java.util.concurrent.ThreadPoolExecutor$Worker.run(ThreadPoolExecutor.java:635) [?:?]
[AllTraffic/i-098f59f854f0cab5b](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#logsV2:log-groups/log-group/$252Faws$252Fsagemaker$252FEndpoints$252Foats2/log-events/AllTraffic$252Fi-098f59f854f0cab5b$3Fstart$3D1736964966229$26refEventId$3D38735613132877352247412838889996017434895850674447581193)
2025-01-15T10:16:06.329-08:00
#011at java.lang.Thread.run(Thread.java:840) [?:?]
Contributor guide
Assessment
This issue has not been assessed yet.