gleanwork / gleanwork/api-client-python
Model serialization drops keys
还没有人认领这个 Issue。
- 主要语言
- Python
- 星标
- 20
- 派生
- 10
- 平均合并
- 1 天 6 小时
- 30 天内合并 PR
- 17
描述
There may be an issue affecting the serialize_model methods of the Pydantic models in this library.
Taking the DocumentContent model as an example, we see:
class DocumentContent(BaseModel):
full_text_list: Annotated[
Optional[List[str]], pydantic.Field(alias="fullTextList")
] = None
r"""The plaintext content of the document."""
@model_serializer(mode="wrap")
def serialize_model(self, handler):
optional_fields = set(["fullTextList"])
serialized = handler(self)
m = {}
for n, f in type(self).model_fields.items():
k = f.alias or n
val = serialized.get(k)
if val != UNSET_SENTINEL:
if val is not None or k not in optional_fields:
m[k] = val
return m
This model uses a field alias that, when constructing the Pydantic object from an API response, will map the fullTextList field of the JSON object to the full_text_list field of the Pydantic object.
However, the model serializer uses:
...
k = f.alias or n
val = serialized.get(k)
...
which means that the field alias (fullTextList) will be used to extract the value rather than the Pydantic field name. This results in value being None and in missing keys in the returned dictionary m when the field name and its alias are different.
To support this claim, please find attached a documents.json file that contains an anonymized response collected from the Glean API (/rest/api/v1/getdocuments endpoint).
And below is a simple debug.py script to run alongside it:
import pathlib
from glean.api_client import models
from glean.api_client.utils.unmarshal_json_response import unmarshal_json_response
class DummyHttpResponse:
def __init__(self, text):
self.status_code = 200
self.text = text
with pathlib.Path("documents.json").open("r") as f:
http_res = DummyHttpResponse(
text=f.read(),
)
documents_response = unmarshal_json_response(models.GetDocumentsResponse, http_res)
assert isinstance(documents_response, models.GetDocumentsResponse)
assert documents_response.documents is not None
assert isinstance(
documents_response.documents["https://company.com/Test"].content,
models.DocumentContent,
)
assert (
documents_response.documents["https://company.com/Test"].content.full_text_list[0]
== "This is a test document."
)
serialized_document_response = documents_response.model_dump()
assert isinstance(serialized_document_response, dict)
assert serialized_document_response["documents"] is not None
# Here's the problem: no `full_text_list` or `fullTextList` in the serialized response!
assert (
len(serialized_document_response["documents"]["https://company.com/Test"]["content"])
> 0
)
Running it yields:
$ ls
debug.py documents.json
$ python debug.py
Traceback (most recent call last):
File "/workspace/app/debug/debug.py", line 40, in <module>
len(serialized_document_response["documents"]["https://company.com/Test"]["content"])
> 0
AssertionError
Note that I'm using:
pydantic_core==2.41.5
pydantic==2.12.5
glean-api-client==0.11.27
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
调研方向
从 src/glean/api_client/models/documentcontent.py 开始,检查 serialize_model,然后通过 unmarshal_json_response 使用 documents.json 运行 debug.py。当序列化后的 DocumentContent content 非空且包含受影响的字段时即表示完成,这一点由 debug.py 中的 assertions 检查。
由索引模型根据 Issue 内容生成。
评估
- 技术栈
- python
- 领域
- api
- Issue 类型
- 缺陷
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 活跃度
- 停滞
- 描述清晰度
- 描述清楚
- 新手友好度
- 45/100