googleapis / googleapis/google-api-python-client

Option to skip per-method docstring generation in build() (memory / high-concurrency)

未關閉
#2,779 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
8.9k
分支
2.6k
平均合併
2 天 28 分鐘
30 天內合併 PR
17

描述

## Feature request: an option to skip per-method docstring generation in `build()`

### Problem

`build()` (and lazy sub-resource construction) generates a fully-expanded, recursively
pretty-printed prototype of each method's **response schema** and attaches it as
`method.__doc__` (`discovery.createMethod` → `schema.Schemas.prettyPrintSchema` /
`prettyPrintByName`). For APIs with large, deeply-nested schemas this is very expensive, and it
is paid **every time a service/resource is constructed**.

Concrete numbers from profiling Sheets v4 (`google-api-python-client==2.198.0`, Python 3.13),
measured with RSS (no tracemalloc, to avoid its overhead):

- Building the service, then touching one sub-resource (`service.spreadsheets()`, **no API
call**): **~66 MB**.
- Of that, ~99.9% is the docstring schema expansion — no-oping `prettyPrintSchema`/
`prettyPrintByName` drops it to **~1 MB**. The `.spreadsheets()` methods themselves are ~24 KB.
- The docstrings are only useful for interactive `help()`; in a server they are never read.

### Impact

In a concurrent server (a fresh service built per request, common with per-user credentials),
these allocations are **not shared** across in-flight requests. 8 concurrent Sheets requests
each build ~66 MB of docstrings simultaneously ≈ **530 MB peak**, which OOM-kills a
memory-limited container. This is the concurrent-peak sibling of the long-standing
reference-cycle memory issue in #535 (whose recommended fix — build/reuse a single service — is
not always feasible when credentials differ per request).

### Request

A supported way to skip docstring generation at build time, e.g.:

```python
build("sheets", "v4", credentials=creds, generate_docstrings=False)
# or a module/env toggle
```

Today the only options are to monkeypatch `Schemas.prettyPrintSchema`/`prettyPrintByName`
(fragile across versions) or fork. A first-class flag would let memory-constrained / high-
concurrency deployments opt out of documentation strings they never use.

### Environment

- `google-api-python-client==2.198.0`, Python 3.13
- Reproly: build any large-schema API (Sheets v4), touch a sub-resource, measure RSS; repeat
concurrently to see the multiplier.

Related: #535 (memory from repeated `build()` / reference cycles).

貢獻指南

開啟貢獻指南

研究方向

從 build() 開始,追蹤 discovery.createMethod、schema.Schemas.prettyPrintSchema 和 prettyPrintByName,查看在建構 service 和 lazy sub-resource 期間 response-schema docstring 是在哪裡產生的。重現 issue 中描述的 Sheets v4 記憶體設定;當受支援的 build-time 選項可以略過該產生程序,且啟用後不影響正常建構時,即視為完成。

由索引模型根據 Issue 內容生成。

評估

技術堆疊
python
領域
api, backend
Issue 類型
功能
難度
4/5
預估耗時
3-5 天
活躍度
冷清
描述清晰度
基本清楚
新手友好度
52/100

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。