huggingface / huggingface/cosmopedia

python deduplicate_dataset.py

Open
#12 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
574
Forks
49
PR merge metrics
No merged PRs in 30d

Description

https://github.com/huggingface/cosmopedia/blob/main/deduplication/deduplicate_dataset.py

```
2024-02-22 14:17:57.759 | INFO | datatrove.executor.slurm:launch_job:216 - Launching dependency job "mh3"
2024-02-22 14:17:57.759 | INFO | datatrove.executor.slurm:launch_job:216 - Launching dependency job "mh2"
2024-02-22 14:17:57.759 | INFO | datatrove.executor.slurm:launch_job:216 - Launching dependency job "mh1"
2024-02-22 14:17:57.763 | INFO | datatrove.executor.slurm:launch_job:249 - Launching Slurm job mh1 (120 tasks) with launch script "/home/wzp/code/LLMData/open_source/datatrove/data/minhash_logs/signatures/launch_script.slurm"
Traceback (most recent call last):
File "/home/wzp/code/LLMData/open_source/datatrove/demo.py", line 110, in
stage4.run()
File "/home/wzp/code/LLMData/open_source/datatrove/src/datatrove/executor/slurm.py", line 169, in run
self.launch_job()
File "/home/wzp/code/LLMData/open_source/datatrove/src/datatrove/executor/slurm.py", line 217, in launch_job
self.depends.launch_job()
File "/home/wzp/code/LLMData/open_source/datatrove/src/datatrove/executor/slurm.py", line 217, in launch_job
self.depends.launch_job()
File "/home/wzp/code/LLMData/open_source/datatrove/src/datatrove/executor/slurm.py", line 217, in launch_job
self.depends.launch_job()
File "/home/wzp/code/LLMData/open_source/datatrove/src/datatrove/executor/slurm.py", line 262, in launch_job
self.job_id = launch_slurm_job(launch_file_contents, *args)
File "/home/wzp/code/LLMData/open_source/datatrove/src/datatrove/executor/slurm.py", line 349, in launch_slurm_job
return subprocess.check_output(["sbatch", *args, f.name]).decode("utf-8").split()[-1]
File "/home/wzp/anaconda3/envs/3.10/lib/python3.10/subprocess.py", line 421, in check_output
return run(*popenargs, stdout=PIPE, timeout=timeout, check=True,
File "/home/wzp/anaconda3/envs/3.10/lib/python3.10/subprocess.py", line 503, in run
with Popen(*popenargs, **kwargs) as process:
File "/home/wzp/anaconda3/envs/3.10/lib/python3.10/subprocess.py", line 971, in __init__
self._execute_child(args, executable, preexec_fn, close_fds,
File "/home/wzp/anaconda3/envs/3.10/lib/python3.10/subprocess.py", line 1863, in _execute_child
raise child_exception_type(errno_num, err_msg, err_filename)
```

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with deduplication/deduplicate_dataset.py and trace the failure through datatrove executor/slurm.py to the subprocess sbatch call shown in the traceback. Reproduce the command failure if possible, then clarify the expected behavior and document or address the root cause before considering the issue done.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.