autogluon / autogluon/autogluon-assistant

Arbitrary Code Execution without Sandboxing in ExecuterAgent

Open
#247 6 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
306
Forks
56
PR merge metrics
No merged PRs in 30d

Description

I noticed that the `ExecuterAgent` executes LLM-generated Python and Bash code directly on the host machine using `subprocess.Popen`.

This is a significant security risk. Beyond the danger of a buggy generation causing accidental damage, this opens a direct attack vector for bad actors. An attacker could manipulate the LLM (e.g., through indirect prompt injection) to intentionally generate malicious code. This makes any system or agent built using this repository vulnerable to being hijacked. An attack could lead to severe consequences like:

* Stealing sensitive data from `~/.ssh/` or `~/.aws/`.
* Deleting user files (`rm -rf /`).
* Installing malware on the host system.

Considering that agentic systems may become more capable, I suggest to add a warning in the the `README.md`, so users understand the risk before running the code. Something like this would be great:

```markdown
---
## ⚠️ Security Warning ⚠️

This tool allows Language Models (LLMs) to execute arbitrary code directly on your machine. This is inherently dangerous and can be exploited by bad actors. It is strongly recommended to always run this code in a sandboxed environment, such as a Docker container or a dedicated VM, to protect your system and data.
---
```

A more robust solution could be to make sandboxing the default execution method, for example using Docker. The `execute_code` function could be modified to spin up a minimal, isolated Docker container for each execution.

Thanks for the great work on this repository!

Contributor guide

Open the contributing guide

Research direction

README.md and the ExecuterAgent/execute_code path using subprocess.Popen are the named starting points. First inspect how Python and Bash commands are launched, then update README.md to warn that generated code runs unsandboxed and recommend isolation such as Docker or a VM. Done means the warning accurately reflects current behavior; the optional sandbox redesign is a separate, larger change.

Written by the indexing model from the issue text.

Assessment

Tech stack
bash, python
Domain
documentation, security
Issue type
Documentation
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
38/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.