feat(environment): stream lines in ReadFileTool to prevent OOM on large files
- Lingua principale
- Python
- Stelle
- 21.5k
- Fork
- 4k
- Merge medio
- 1g 14h
- PR unite (30g)
- 37
Descrizione
## Required Information
### Is your feature request related to a specific problem?
Currently, `ReadFileTool` loads an entire file into memory as bytes via `await self._environment.read_file(path)` and then splits lines across the whole payload before slicing the requested range (`start_line` to `end_line`).
For very large files (hundreds of megabytes or gigabytes), reading the entire file into memory risks high memory usage or Out-Of-Memory (OOM) crashes in the agent environment, even when the caller only requests a small range of lines. There is an explicit `TODO` in `src/google/adk/tools/environment/_read_file_tool.py`:
`# TODO: Avoid loading the entire file into memory to prevent OOM on large files.`
### Describe the Solution You'd Like
1. Add a `read_file_lines(path, start_line, end_line)` method to `BaseEnvironment` with a backward-compatible default fallback that calls `read_file`.
2. Implement an optimized, streaming `read_file_lines` on `LocalEnvironment` that iterates through the file lazily line-by-line in $O(1)$ memory, only collecting the requested range.
3. Update `ReadFileTool.run_async` to call `read_file_lines`, eliminating the need to buffer the entire file into memory.
### Impact on your work
Improves stability and performance of file inspection tools in agent environments handling large log files, datasets, and source files.
### Willingness to contribute
Yes, I have an implementation ready with unit tests and pre-commit checks passing.
---
## Recommended Information
### Proposed API / Implementation
```python
# BaseEnvironment
async def read_file_lines(
self,
path: str | Path,
start_line: int = 1,
end_line: int | None = None,
) -> tuple[list[bytes], int]: ...
# LocalEnvironment
@staticmethod
def _sync_read_lines(
path: Path, start_line: int, end_line: int | None = None
) -> tuple[list[bytes], int]:
selected_lines: list[bytes] = []
total_lines = 0
with open(path, 'rb') as f:
for line in f:
total_lines += 1
if start_line <= total_lines and (
end_line is None or total_lines <= end_line
):
selected_lines.append(line)
return selected_lines, total_lines
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.