aws / aws/sagemaker-python-sdk

Clean temporary files after loading query to dataframe

未关闭
#5,100 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
component: feature store type: bug
主要语言
Python
星标
2.3k
派生
1.3k
平均合并
1 天 22 小时
30 天内合并 PR
35

描述

**Describe the bug**
Loading data from feature group using `as_dataframe()` doesn't clean temporary `.csv files`

**To reproduce**
After creating feature group and ingesting data, load athena query and then use `as_dataframe()` method

**Expected behavior**
This methods loads .csv query file and returns the data inside, but it doesn't clean the .csv file itself.
So after running multiple queries it's really easy to just fill your local disk space up to maximum, because of multiple .csv files with query results

**Screenshots or logs**
`

def as_dataframe(self, **kwargs) -> DataFrame:
"""Download the result of the current query and load it into a DataFrame.

Args:
**kwargs (object): key arguments used for the method pandas.read_csv to be able to
have a better tuning on data. For more info read:
https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.read_csv.html

Returns:
A pandas DataFrame contains the query result.
"""
query_state = self.get_query_execution().get("QueryExecution").get("Status").get("State")
if query_state != "SUCCEEDED":
if query_state in ("QUEUED", "RUNNING"):
raise RuntimeError(
f"Current query {self._current_query_execution_id} is still being executed."
)
raise RuntimeError(f"Failed to execute query {self._current_query_execution_id}")

output_filename = os.path.join(
tempfile.gettempdir(), f"{self._current_query_execution_id}.csv"
)
self.sagemaker_session.download_athena_query_result(
bucket=self._result_bucket,
prefix=self._result_file_prefix,
query_execution_id=self._current_query_execution_id,
filename=output_filename,
)

kwargs.pop("delimiter", None)
return pd.read_csv(filepath_or_buffer=output_filename, delimiter=",", **kwargs)

`

贡献指南

打开贡献指南

调研方向

Start from the as_dataframe() entry point shown in the issue and trace the temporary CSV path through download_athena_query_result and pandas.read_csv. Verify that the query result still loads successfully and that the temporary CSV is removed after the method finishes.

由索引模型根据 Issue 内容生成。

评估

技术栈
aws, pandas, python
领域
cloud, data
Issue 类型
缺陷
难度
2/5
预计耗时
1-3 小时
活跃度
停滞
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。