aws / aws/sagemaker-python-sdk

`ModelBuilder` missing `resource_requirements` and `inference_component_name` fields for `CustomOrchestrator` IC deployment

Đang mở
#6,198 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
component: model builder type: bug
Ngôn ngữ chính
Python
Star
2.3k
Fork
1.3k
Merge trung bình
1 ngày 22 giờ
Pull request đã merge (30 ngày)
35

Mô tả

**PySDK Version**
- [ ] PySDK V2 (2.x)
- [x] PySDK V3 (3.x)

**Describe the bug**

When deploying a `CustomOrchestrator` as an Inference Component — as documented in the [Build and deploy AI inference workflows with new enhancements to the Amazon SageMaker Python SDK](https://aws.amazon.com/blogs/machine-learning/build-and-deploy-ai-inference-workflows-with-new-enhancements-to-the-amazon-sagemaker-python-sdk/) blog post — the `ModelBuilder` class internally reads `self.resource_requirements` and `self.inference_component_name` during `build()` to determine:

1. Whether to deploy the orchestrator as an IC or a standalone Endpoint ([model_builder.py L4507](https://github.com/aws/sagemaker-python-sdk/blob/44410e2c55202f80d980e1dfd4533e66b9d44495/sagemaker-serve/src/sagemaker/serve/model_builder.py#L4507))
2. The IC name and resource allocation to include in the deployable spec ([model_builder.py L4534-4536](https://github.com/aws/sagemaker-python-sdk/blob/44410e2c55202f80d980e1dfd4533e66b9d44495/sagemaker-serve/src/sagemaker/serve/model_builder.py#L4534))

However, **neither `resource_requirements` nor `inference_component_name` are defined as dataclass fields** on `ModelBuilder`. They cannot be passed through the constructor, despite being required for the documented `CustomOrchestrator` workflow.

The `ModelBuilder` dataclass fields are:

```
model, model_path, inference_spec, schema_builder, modelbuilder_list, role_arn,
sagemaker_session, image_uri, s3_model_data_url, source_code, env_vars, model_server,
model_metadata, log_level, content_type, accept_type, compute, network, instance_type,
mode, shared_libs, dependencies, image_config
```

The existing `compute` field is a `Compute(ResourceConfig)` class designed for training jobs (volume sizes, spot training, instance groups) — not inference component resource allocation.

**To reproduce**

Try the following script:

```python
from sagemaker.serve.model_builder import ModelBuilder, SchemaBuilder
from sagemaker.serve.spec.inference_base import CustomOrchestrator
from sagemaker.core.inference_config import ResourceRequirements
from sagemaker.core.helper.session_helper import Session, get_execution_role

class MyOrchestrator(CustomOrchestrator):
def __init__(self, endpoint_name, component_names):
super().__init__()
self.endpoint_name = endpoint_name
self.component_names = component_names

def handle(self, data, context=None):
import json
response = self.client.invoke_endpoint(
EndpointName=self.endpoint_name,
InferenceComponentName=self.component_names[0],
Body=data if isinstance(data, (str, bytes)) else json.dumps(data),
ContentType="application/json"
)
return json.loads(response["Body"].read())

role = get_execution_role()
sess = Session()

# This is the expected usage pattern per the reference notebook:
orchestrator = ModelBuilder(
inference_spec=MyOrchestrator(
endpoint_name="my-endpoint",
component_names=["base-ic", "adapter-ic"],
),
dependencies={"auto": False, "custom": ["cloudpickle"]},
sagemaker_session=sess,
role_arn=role,
schema_builder=SchemaBuilder(sample_input="Test", sample_output={"generated_text": "test"}),
resource_requirements=ResourceRequirements(
requests={"memory": 4096, "num_accelerators": 1, "copies": 1, "num_cpus": 2}
),
inference_component_name="my-orchestrator-ic",
)
```

**Expected behavior**

`ModelBuilder` accepts `resource_requirements` and `inference_component_name` as constructor parameters, consistent with the documented usage in the [Llama3.1-Mistral inference workflow reference notebook](https://github.com/aws-samples/sagemaker-genai-hosting-examples/blob/main/Llama3.1-Mistral-workflow/Llama3.1-8B-Mistral-7B-inference-orchestrator.ipynb) (Cell 15):

```python
orchestrator = ModelBuilder(
inference_spec=PythonCustomInferenceEntryPoint(...),
dependencies={"auto": False, "custom": ["cloudpickle", "graphene"]},
sagemaker_session=Session(),
role_arn=role,
resource_requirements=ResourceRequirements(
requests={"memory": 4096, "num_accelerators": 1, "copies": 1, "num_cpus": 2}
),
inference_component_name=custom_orchestrator_name,
schema_builder=SchemaBuilder(sample_input="Test", sample_output={"generated_text": "test"}),
modelbuilder_list=[llama_model_builder, mistral_mb]
)
```

**Screenshots or logs**
```
╭─────────────────────────────── Traceback (most recent call last) ────────────────────────────────╮
│ in :4 │
│ │
│ 1 from sagemaker.serve.model_builder import ModelBuilder, SchemaBuilder │
│ 2 from sagemaker.core.inference_config import ResourceRequirements │
│ 3 │
│ ❱ 4 orchestrator = ModelBuilder( │
│ 5 │ inference_spec=SequentialWorkflow( │
│ 6 │ │ region_name=region, │
│ 7 │ │ endpoint_name=endpoint_name, │
╰──────────────────────────────────────────────────────────────────────────────────────────────────╯
TypeError: ModelBuilder.__init__() got an unexpected keyword argument 'resource_requirements'
```

```
╭─────────────────────────────── Traceback (most recent call last) ────────────────────────────────╮
│ in :4 │
│ │
│ 1 from sagemaker.serve.model_builder import ModelBuilder, SchemaBuilder │
│ 2 from sagemaker.core.inference_config import ResourceRequirements │
│ 3 │
│ ❱ 4 orchestrator = ModelBuilder( │
│ 5 │ inference_spec=SequentialWorkflow( │
│ 6 │ │ region_name=region, │
│ 7 │ │ endpoint_name=endpoint_name, │
╰──────────────────────────────────────────────────────────────────────────────────────────────────╯
TypeError: ModelBuilder.__init__() got an unexpected keyword argument 'inference_component_name'
```

**Workaround**

````
Setting attributes directly on the instance after construction:

```python
orchestrator = ModelBuilder(
inference_spec=MyOrchestrator(...),
dependencies={"auto": False, "custom": ["cloudpickle"]},
sagemaker_session=sess,
role_arn=role,
schema_builder=SchemaBuilder(sample_input="Test", sample_output={"generated_text": "test"}),
)

# Workaround: set attributes post-construction
orchestrator.resource_requirements = ResourceRequirements(
requests={"memory": 4096, "num_accelerators": 1, "copies": 1, "num_cpus": 2}
)
orchestrator.inference_component_name = "my-orchestrator-ic"

orchestrator.build()
```

This works because `build()` reads `self.resource_requirements` and `self.inference_component_name` via attribute access, but it is undocumented and inconsistent with the published reference sample.

**System information**
A description of your system. Please provide:
- **SageMaker Python SDK version**: `sagemaker-serve 1.20.0 (SDK v3)`
- **Framework name (eg. PyTorch) or algorithm (eg. KMeans)**: sagemaker-distribution-prod:3.2.0-cpu
- **Framework version**:
- **Python version**: 3.12
- **CPU or GPU**: GPU
- **Custom Docker image (Y/N)**: N

**Additional context**
Add any other context about the problem here.

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Hướng nghiên cứu

Bắt đầu trong sagemaker-serve/src/sagemaker/serve/model_builder.py, kiểm tra các trường của dataclass ModelBuilder và các tham chiếu đến build() quanh các dòng 4507 và 4534-4536. Thêm hỗ trợ cho hai tham số đã được tài liệu hóa vào constructor và xác minh rằng ví dụ CustomOrchestrator được build với ResourceRequirements và tên component suy luận.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
aws, python
Lĩnh vực
cloud, machine-learning
Loại issue
Lỗi
Độ khó
3/5
Thời gian dự kiến
1-2 ngày
Mức độ hoạt động
Sôi nổi
Độ rõ ràng
Đặc tả rõ ràng
Mức phù hợp với người mới
76/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.