getsentry / getsentry/sentry-python
`_generate_sample_rand` seeds a Mersenne Twister per Transaction, eagerly, even when unsampled (6.4 µs)
- 主要言語
- Python
- スター
- 2.2k
- フォーク
- 669
- 平均マージ
- 1日 40分
- マージ済み PR(30日)
- 212
説明
### Summary
`Transaction.__init__` unconditionally computes `_generate_sample_rand(self.trace_id)`, and `_generate_sample_rand` seeds a Mersenne Twister to produce a single float. That is **6.4 µs per call** on CPython 3.14 / Apple M2, paid on every request through the ASGI integrations even when tracing is disabled and the value can never be used.
Two independent problems:
**1. It is eager.** `sentry_sdk/tracing.py`, `Transaction.__init__`:
```python
baggage_sample_rand = None if self._baggage is None else self._baggage._sample_rand()
if baggage_sample_rand is not None:
self._sample_rand = baggage_sample_rand
else:
self._sample_rand = _generate_sample_rand(self.trace_id)
```
`_sample_rand` is only read when a sampling decision is actually made. With `traces_sample_rate` unset the transaction is never sampled, so this is pure waste. Making it a lazy property costs nothing.
**2. It is expensive.** `sentry_sdk/tracing_utils.py`:
```python
def _generate_sample_rand(trace_id, *, interval=(0.0, 1.0)):
...
rng = Random(trace_id)
sample_rand_scaled = rng.randrange(lower_scaled, upper_scaled)
return sample_rand_scaled / 1_000_000
```
`Random(seed)` runs the full MT19937 `init_by_array` over a 625-word state. Measured with `timeit`, 50k iterations, best of 5:
| | µs |
|---|---:|
| `Random(trace_id)` (32-char hex string) | 6.39 |
| `Random(int(trace_id, 16))` | 5.90 |
| `Random(12345)` | 5.89 |
| `int(trace_id, 16) / 2**128` | **0.23** |
The cost is the Mersenne Twister initialisation, not the string hashing - seeding with a small int is just as slow. Deriving a uniformly distributed value in `[0, 1)` arithmetically from the same trace id is **27x cheaper** and just as deterministic.
### Impact
On a do-nothing FastAPI endpoint with tracing disabled, making `_generate_sample_rand` cheap moves the SDK's per-request overhead from +61.3 µs to +53.3 µs (in-process measurement, baseline 15.6 µs/req) - about 13% of the SDK's cost, for a value that is discarded.
### Questions
1. Is the fix to `Transaction.__init__` simply making `_sample_rand` lazy? Happy to open a PR.
2. Is the exact output of `_generate_sample_rand` required to be bit-compatible across SDKs, or only to be deterministic-from-`trace_id` and uniformly distributed? If the latter, `int(trace_id, 16) / 2**128` (scaled into the requested interval) would be a drop-in replacement. If the former, the laziness fix alone still helps.
### Repro
```python
import timeit, uuid
from random import Random
tid = uuid.uuid4().hex
n = 50000
for label, fn in [
("Random(hex str)", lambda: Random(tid)),
("Random(int)", lambda: Random(int(tid, 16))),
("Random(12345)", lambda: Random(12345)),
("int(tid,16)/2**128", lambda: int(tid, 16) / 2**128),
]:
print(f"{label:<20} {min(timeit.repeat(fn, number=n, repeat=5)) / n * 1e6:.2f} us")
```
Environment: CPython 3.14.7, sentry-sdk 2.67.1, Apple M2.
Context: this was found while measuring 7400, where the discarded `Transaction` is the larger half of the same problem.
コントリビューションガイド
調査の方向性
Start in sentry_sdk/tracing.py at Transaction.__init__ and sentry_sdk/tracing_utils.py at _generate_sample_rand; inspect how _sample_rand is read and whether output compatibility is required. Run the issue's benchmark to compare the alternatives, then verify that unsampled transactions avoid unnecessary work while deterministic sampling behavior is preserved.
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- python
- 領域
- performance
- issue の種類
- リファクタリング
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 活発さ
- 活発
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 64/100