aws / aws/sagemaker-python-sdk

Enable passing column type to SHAPConfig in combination with ClarifyCheckStep

オープン
#5,131 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る
component: pipelines type: feature request
主要言語
Python
スター
2.3k
フォーク
1.3k
平均マージ
1日 22時間
マージ済み PR(30日)
35

説明

**Describe the feature you'd like**
Add a parameter to SHAPConfig from sagemaker.workflow.clarify_checkstep which lets the user specify the types of the dataset used to create a baseline for the SHAP analysis (e.g. float, int, category, etc..).
Alternatively, make it possible to run ClarifyCheckStep when an S3 URI has been passed as baseline to SHAPConfig.

**How would this feature be used? Please describe.**
When using the ClarifyCheckStep and SHAPConfig from sagemaker.workflow.clarify_checkstep, I am currently unable to specify my dataset's column types (e.g. some columns should be numerical while others should be categorical).

When running the ClarifyCheckStep as part of a SageMaker pipeline, Clarify calculates a baseline which is erroneous due to not having taken the column types into account, so e.g. some columns that should be categorical gets the mean of the column as baseline, where preferrably they should get the mode of the column or something else more appropriate.

I know that I can pass my own baseline to SHAPConfig, but I don't want this hard coded in my SageMaker pipeline definition - I want it to be computed at runtime, based on previous steps in my SageMaker pipeline.
An alternative solution would be to pass to SHAPConfig the S3 URI to a baseline dataset I create in a previous step, however this doesn't seem to work with how ClarifyCheckStep is currently implemented.

**Describe alternatives you've considered**
Make it possible to run ClarifyCheckStep when an S3 URI has been passed as baseline to SHAPConfig.

**Additional context**
```
from sagemaker.workflow.clarify_check_step import ClarifyCheckStep, ModelExplainabilityCheckConfig, SHAPConfig

shap_config = SHAPConfig(seed=123, num_samples=100, num_clusters=5)

model_explainability_check_config = ModelExplainabilityCheckConfig(
data_config=model_explainability_data_config,
model_config=model_config,
explainability_config=shap_config,
)

step_model_explainability_check = ClarifyCheckStep(
name="ModelExplainabilityCheckStep",
display_name="Model Explainability Check",
clarify_check_config=model_explainability_check_config,
check_job_config=check_job_config_clarify,
skip_check=skipCheckModelExplainabilityParam,
register_new_baseline=registerNewBaselineModelExplainabilityParam,
supplied_baseline_constraints=suppliedBaselineConstraintsModelExplainabilityParam,
model_package_group_name=model_package_group_name,
)
```

コントリビューションガイド

コントリビューションガイドを開く

調査の方向性

sagemaker.workflow.clarify_check_step.py から始めて SHAPConfig と ClarifyCheckStep を読み、baseline 値と S3 URI がどのように扱われているかに注目します。提案されている 2 つのアプローチ、つまり列型パラメーターと実行時の S3 baseline サポートを比較し、ClarifyCheckStep がパイプライン内で型を考慮した SHAP baseline を計算または利用できることを完了条件とします。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
aws, python
領域
backend-api-design, machine-learning
issue の種類
機能追加
難易度
4/5
見積もり時間
3〜5日
活発さ
停滞
明瞭さ
おおむね明確
初心者へのやさしさ
35/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。