aws / aws/sagemaker-python-sdk

load_feature_definitions_from_dataframe() doesn't recognize pandas nullable dtypes (Float64, Int64)

Closed
#5,675 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
2.3k
Forks
1.3k
Avg merge
1d 22h
Merged PRs (30d)
35

Description

**PySDK Version**
PySDK 3.6.0

**Describe the bug**
load_feature_definitions_from_dataframe() in sagemaker.mlops.feature_store only recognizes numpy dtypes (float64, int64, etc.) but not pandas nullable dtypes (Float64, Int64, string). When a DataFrame uses nullable dtypes (common after calling pd.DataFrame.convert_dtypes()), all numeric columns are incorrectly mapped to StringFeatureDefinition.

**To reproduce**
```
import pandas as pd
from sagemaker.mlops.feature_store import load_feature_definitions_from_dataframe

# Create a DataFrame with numpy dtypes (works correctly)
df_numpy = pd.DataFrame({
"id": [1, 2, 3],
"price": [1.1, 2.2, 3.3],
"name": ["a", "b", "c"],
})
print("numpy dtypes:", {c: str(df_numpy[c].dtype) for c in df_numpy.columns})
# {'id': 'int64', 'price': 'float64', 'name': 'object'}

defs = load_feature_definitions_from_dataframe(df_numpy)
for d in defs:
print(f" {d.feature_name}: {d.feature_type}")

# Now convert to pandas nullable dtypes (common pattern)
df_nullable = df_numpy.convert_dtypes()
print("\nnullable dtypes:", {c: str(df_nullable[c].dtype) for c in df_nullable.columns})
# {'id': 'Int64', 'price': 'Float64', 'name': 'string'}

defs = load_feature_definitions_from_dataframe(df_nullable)
for d in defs:
print(f" {d.feature_name}: {d.feature_type}")
```

Root cause

In sagemaker/mlops/feature_store/feature_utils.py, _INTEGER_TYPES and _FLOAT_TYPES only contain lowercase numpy dtype names:

_INTEGER_TYPES = {'int8', 'int16', 'int32', 'int64', 'int_', 'uint8', 'uint16', 'uint32', 'uint64'}
_FLOAT_TYPES = {'float16', 'float32', 'float64', 'float_'}

Pandas nullable dtypes are capitalized (Int64, Float64, etc.) and are not matched.

Suggested fix

Add nullable dtype names to the type sets:

_INTEGER_TYPES = {'int8', 'int16', 'int32', 'int64', 'int_',
'Int8', 'Int16', 'Int32', 'Int64',
'uint8', 'uint16', 'uint32', 'uint64',
'UInt8', 'UInt16', 'UInt32', 'UInt64'}
_FLOAT_TYPES = {'float16', 'float32', 'float64', 'float_',
'Float16', 'Float32', 'Float64'}

Or use case-insensitive comparison in _generate_feature_definition().

**Expected behavior**
Panda nullable types should get properly converted.

**System information**
A description of your system. Please provide:
- **SageMaker Python SDK version**: 3.6.0

I think this got fixed/addressed before.. but maybe that 2.x code didn't carray over to 3.x
https://github.com/aws/sagemaker-python-sdk/pull/3740/changes

Contributor guide

Open the contributing guide

Research direction

Start in sagemaker/mlops/feature_store/feature_utils.py, focusing on _generate_feature_definition() and the _INTEGER_TYPES and _FLOAT_TYPES sets. Run the provided DataFrame reproduction and verify that nullable Int64 and Float64 columns are mapped to numeric feature definitions rather than StringFeatureDefinition, while existing numpy dtype behavior remains unchanged.

Written by the indexing model from the issue text.

Assessment

Tech stack
pandas, python
Domain
machine-learning
Issue type
Bug
Difficulty
2/5
Estimated time
1-3 hours
Activity status
Quiet
Clarity
Clearly specified
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.