googleapis / googleapis/google-cloud-python

SpannerDialect does not accept json_serializer / json_deserializer kwargs from create_engine()

Abierto
#15,672 1 comentario 0 reacciones 1 asignado Reclamado por @olavloite Ver en GitHub
api: spanner
Lenguaje dominante
Python
Estrellas
5.4k
Forks
1.8k
Merge medio
3 d 4 h
PR fusionados (30 d)
122

Descripción

## Problem

SQLAlchemy's `create_engine()` supports `json_serializer` and `json_deserializer` parameters to customize how JSON columns are serialized/deserialized. This is [documented behavior](https://docs.sqlalchemy.org/en/20/core/type_basics.html#sqlalchemy.types.JSON) and works with PostgreSQL, MySQL, SQLite, and other dialects.

However, passing these to the Spanner dialect raises `TypeError`:

```python
engine = create_engine(
"spanner:///projects/p/instances/i/databases/d",
json_serializer=lambda obj: json.dumps(obj, cls=MyEncoder),
)
# TypeError: Invalid argument(s) 'json_serializer' sent to create_engine(),
# using configuration SpannerDialect/QueuePool/Engine.
```

### Root cause

SQLAlchemy's `create_engine()` uses `util.get_cls_kwargs()` to determine which kwargs the dialect accepts. Since `SpannerDialect.__init__` does not declare `json_serializer` or `json_deserializer`, they are rejected.

Other dialects (e.g., `PGDialect`) accept these in their `__init__` and store them on the instance, where `DefaultDialect` exposes them as `_json_serializer` / `_json_deserializer` for use by `JSON.bind_processor()`.

### Additional complexity

Even if the dialect accepted the kwargs, there's a pipeline mismatch. SQLAlchemy expects `_json_serializer` to be a `json.dumps`-like callable (returns a string), but the Spanner dialect sets `_json_serializer = JsonObject` — a class constructor that produces a `JsonObject` instance. The actual string serialization happens later in `_helpers._make_param_value_pb` when it calls `obj.serialize()`.

A proposed fix uses a **serialize-then-wrap** strategy: pre-serialize with the user's function, then `JsonObject.from_str()` the result. This requires no changes to `JsonObject` or the core `google-cloud-spanner` library.

## Expected behavior

```python
engine = create_engine(
"spanner:///...",
json_serializer=lambda obj: json.dumps(obj, default=my_handler),
)
```

should work, allowing custom types in JSON columns to be serialized correctly through the DML path.

## Environment

- `sqlalchemy-spanner`: 1.7.0
- `sqlalchemy`: 2.0.x
- `google-cloud-spanner`: 3.x

Guía de contribución

Abrir la guía de contribución

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.