Worker profile limited to a short timespan
- Langage dominant
- Python
- Étoiles
- 1.7k
- Forks
- 778
- Merge moyen
- 2 h 50 min
- PR mergées (30 j)
- 3
Description
**Describe the issue**:
The worker profile has a limited span and older data seems to be lost. For example, with the minimal example below, the total CPU time is 24 hours, but the profile never contains more than 4 h 26 min. At one point during the run, the profile looks like this:

30 minutes later, it looks like this:

The previous data is completely gone, as can be seen by the "activity over time" graph at the bottom.
This issue has been occurring for several months, most recently with dask and distributed 2024.4.2. You can look at the full discussion on Discourse: https://dask.discourse.group/t/measuring-the-overall-profile-of-long-runs/1859/11
**Minimal Complete Verifiable Example**:
```python
import time
import logging
from dask.distributed import (
Client,
LocalCluster,
get_client,
as_completed,
performance_report,
)
NUM_PROCESSES = 16
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s %(levelname)-8s %(message)s",
)
class DummyManager:
def run(self):
logging.info("Starting the manager")
jobs = list(range(1, 97))
client = get_client()
futs = []
for j in jobs:
futs.append(client.submit(self.job, j))
asc = as_completed(futs, with_results=True)
for fut, ret in asc:
logging.info(f"Processing future {str(fut)}: ret={str(ret)}")
if ret > 0:
logging.info(f"Launching a subjob with time {ret}")
asc.add(client.submit(self.job, ret))
fut.release()
def job(self, n):
time.sleep(60 * 15)
return 0
if __name__ == "__main__":
cluster = LocalCluster(
n_workers=1,
threads_per_worker=NUM_PROCESSES,
processes=False,
)
client = Client(cluster)
with performance_report(filename=f"dask-performance_{time.time():.0f}.html"):
try:
manager = DummyManager()
manager.run()
except KeyboardInterrupt:
logging.info("Stopping the job...")
cluster.close()
exit(0)
client.close()
cluster.close()
```
**Environment**:
- Dask version: 2024.4.2
- Python version: 3.10.12
- Operating System: Linux Mint 21.2
- Install method (conda, pip, source): pip
Guide de contribution
Ouvrir le guide de contribution
Piste de recherche
Exécutez l’exemple minimal Python fourni avec Client, LocalCluster et performance_report, puis comparez le profil du worker et la vue de l’activité au fil du temps pendant l’exécution longue. Suivez la manière dont les données de profilage sont collectées et conservées, en utilisant la discussion Discourse liée comme contexte. Le travail est considéré comme terminé lorsque l’activité plus ancienne reste disponible et que le profil couvre toute l’exécution, au lieu de s’arrêter vers 4 h 26 min.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Évaluation
- Stack technique
- python
- Domaine
- distributed-systems, observability-sre
- Type d'issue
- Bug
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Activité
- À l'abandon
- Clarté
- À clarifier
- Accessibilité débutants
- 35/100