hardbyte / hardbyte/postgresql-job-queue-benchmarking
awa: investigate non-monotonic throughput dip at 256 workers
- Lenguaje dominante
- HTML
- Estrellas
- 2
- Forks
- 0
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
The 2026-05-01 awa-extended-scaling run
([`results/2026-05-01-awa-extended-scaling/`](https://github.com/hardbyte/postgresql-job-queue-benchmarking/tree/main/results/2026-05-01-awa-extended-scaling))
shows a real non-monotonic dip:
| workers | throughput |
|---:|---:|
| 128 | 6,481 |
| **256** | **4,800** |
| 512 | 5,344 |
| 1,024 | 7,837 |
Each point is a 75 s clean phase median over ~15 samples — well above
measurement noise. The dip is reproducible.
Two plausible causes:
- **Connection-pool saturation around 256.** awa's claim path holds
a connection per active worker; postgres `max_connections=400` is
only ~1.5× the worker count, so contention on connection acquire
goes up sharply right around here.
- **`LWLock` / `Lock` contention** on the queue ring metadata at
some specific concurrency, transitioning back to no-contention
once workers fan out past the cliff.
Both are testable with the wait-event sampler that's about to land
(see #?). Re-running with that instrumentation will surface whether
256 sits on a `LWLock` cliff or a connection-acquire wait, and
whether bumping `max_connections` or pool size shifts the dip
location.
Not a blocker — peak throughput at 1,024 is still the highest in the
run — but worth understanding before publishing "awa scales linearly"
as a claim.
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Start by inspecting results/2026-05-01-awa-extended-scaling/ and the benchmark configuration for the 256-worker run. Re-run the workload with the wait-event sampler mentioned in the issue, then compare connection-acquire and LWLock/Lock waits across worker counts. Done means identifying the likely cause and testing whether changing max_connections or pool size shifts the dip.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- postgresql
- Área
- databases, performance
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Tranquilo
- Claridad
- Bastante claro
- Aptitud para principiantes
- 48/100