pulp / pulp/pulpcore

Measure Pulp's ability to scale to high #s of client requests

Open
#1,902 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Performance Task
Dominant language
Python
Stars
598
Forks
168
Avg merge
1d 4h
Merged PRs (30d)
86

Description

Author: @dralley (dalley)

Redmine Issue: 6928, https://pulp.plan.io/issues/6928


Pulp 3 has a more complicated architecture on the externally-facing side than Pulp 2 does. Whereas Pulp 2 wrote directories of symlinks and relied on a web server (Apache) to serve them as a static directly, Pulp 2 has a custom app which services incoming requests by matching their paths against paths stored in the database.

This means that as the # of clients being simultaneously served scales up, so too does the load on the database. We should measure this impact to gain an understanding of how Pulp 3 is likely to behave in a real-world scenario where many thousands of clients may be requesting packages from a Pulp installation and, concurrently, Pulp administrative tasks such as sync and publish may be loading the database as well.

One possible way of testing this would be to create a Kubernetes cluster of "package consumer" agents which can be scaled as desired for the testing.

We would also need some kind of monitoring - but I do not have any suggestions on how this should be accomplished.

This testing should occur in advance of any dev-freeze date, so that we have time to make adjustments if they are found to be necessary.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The issue names no source files, tests, or entry points to start from. First define a repeatable load-testing setup for Pulp 3's externally facing request path, including scalable package-consumer agents and database monitoring. Done means measuring client-request scaling alongside concurrent sync and publish activity and documenting the observed impact.

Written by the indexing model from the issue text.

Assessment

Tech stack
kubernetes, python
Domain
backend, databases, performance, testing-qa
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.