googleapis / googleapis/google-api-python-client

Option to skip per-method docstring generation in build() (memory / high-concurrency)

Open
#2,779 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
8.9k
Forks
2.6k
Avg merge
2d 28m
Merged PRs (30d)
17

Description

## Feature request: an option to skip per-method docstring generation in `build()`

### Problem

`build()` (and lazy sub-resource construction) generates a fully-expanded, recursively
pretty-printed prototype of each method's **response schema** and attaches it as
`method.__doc__` (`discovery.createMethod` → `schema.Schemas.prettyPrintSchema` /
`prettyPrintByName`). For APIs with large, deeply-nested schemas this is very expensive, and it
is paid **every time a service/resource is constructed**.

Concrete numbers from profiling Sheets v4 (`google-api-python-client==2.198.0`, Python 3.13),
measured with RSS (no tracemalloc, to avoid its overhead):

- Building the service, then touching one sub-resource (`service.spreadsheets()`, **no API
call**): **~66 MB**.
- Of that, ~99.9% is the docstring schema expansion — no-oping `prettyPrintSchema`/
`prettyPrintByName` drops it to **~1 MB**. The `.spreadsheets()` methods themselves are ~24 KB.
- The docstrings are only useful for interactive `help()`; in a server they are never read.

### Impact

In a concurrent server (a fresh service built per request, common with per-user credentials),
these allocations are **not shared** across in-flight requests. 8 concurrent Sheets requests
each build ~66 MB of docstrings simultaneously ≈ **530 MB peak**, which OOM-kills a
memory-limited container. This is the concurrent-peak sibling of the long-standing
reference-cycle memory issue in #535 (whose recommended fix — build/reuse a single service — is
not always feasible when credentials differ per request).

### Request

A supported way to skip docstring generation at build time, e.g.:

```python
build("sheets", "v4", credentials=creds, generate_docstrings=False)
# or a module/env toggle
```

Today the only options are to monkeypatch `Schemas.prettyPrintSchema`/`prettyPrintByName`
(fragile across versions) or fork. A first-class flag would let memory-constrained / high-
concurrency deployments opt out of documentation strings they never use.

### Environment

- `google-api-python-client==2.198.0`, Python 3.13
- Reproly: build any large-schema API (Sheets v4), touch a sub-resource, measure RSS; repeat
concurrently to see the multiplier.

Related: #535 (memory from repeated `build()` / reference cycles).

Contributor guide

Open the contributing guide

Research direction

Start at build() and trace discovery.createMethod, schema.Schemas.prettyPrintSchema, and prettyPrintByName to see where response-schema docstrings are generated during service and lazy sub-resource construction. Reproduce the Sheets v4 memory profile described in the issue; done means a supported build-time option skips that generation without affecting normal construction when enabled.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
api, backend
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Quiet
Clarity
Mostly clear
Newbie friendliness
52/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.