[Documentation] How does `SingleNodeExecutor` touch the file system?
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 77
- Forks
- 7
- Avg merge
- 10h 32m
- Merged PRs (30d)
- 12
Description
I'm trying to understand where SingleNodeExecutor with a cache_directory gets its instructions to actually write it's result to file.
For instance, consider the following:
from executorlib import SingleNodeExecutor
def foo(x):
return x + 1
with SingleNodeExecutor() as exe:
f = exe.submit(foo, 1, resource_dict={"cache_key": "my_key", "cache_directory": "my_dir"})
print("Result", f.result())
My understanding of the path of events is
Initialization:
SingleNodeExecutor.__init__triggersBaseExecutor.__init__with aDependencyTaskScheduleras its underlying_task_scheduler- That
DependencyTaskScheduleris initialized with aOneProcessTaskScheduleras itsexecutorarg, which gets stored in_process_kwargs["executor"]
Submission:
SingleNodeExecutor.submitis inherited directly fromBaseExecutor.submitBaseExecutor.submitpasses everything (including theresource_dictas a single kwarg) to the_task_scheduler.submit- Since
DependencyTaskScheduler._generate_dependency_graphhas fed through toFalsewith all the default values, this generates the future by a plainsuper()call toTaskSchedulerBase.submit TaskSchedulerBase.submitsends our information (function, args, kwargs, resource dict, empty future) toself._future_queue.put
Here I get out of my depth, but it seems to me like putting stuff on the future queues is activating the associated Thread, which all the TaskSchedulerBase children initialize in _set_process using a Thread taking some function and the _process_kwargs (which includes the _future_queue!). On that assumption, that means that the self._future_queue.put call we got to from TaskSchedulerBase.submit would route back to the Thread set in the DependencyTaskScheduler._set_process invocation -- i.e. _execute_tasks_with_dependencies
Continuing submission:
_execute_tasks_with_dependenciesindeed takes anexecutor, which isDependencyTaskScheduler._process_kwargs["executor"]i.e. ourOneProcessTaskScheduler; there's lots going on, but...- I don't see any reference to the cache, so I don't think the file system interaction is happening here
- It looks like we ultimately do a
executor_queue.putonto the underlyingOneProcessTaskScheduler._future_queue - I.e. we move to
_execute_task_in_separate_process
_execute_task_in_separate_processis in turn re-directing to_wrap_execute_task_in_separate_processand both of these are now taking aspawner: type[BaseSpawner]argument- But I'm at the end of the line, I don't see anything other than the default
MpiExecSpawnerbeing leveraged, and I never find any references to thecache_directoryor any of thetask_scheduler.filemodule tools
With the SlurmClusterExecutor we sometimes route through create_file_executor, in which case the file system connection is obvious, but in the other case we're still going through DependencyTaskScheduler -- this time with an SrunSpawner instead of a MpiExecSpawner. In this later case I also don't see the connection to file system tools, so I feel like I must be missing something at the diverging point: DependencyTaskScheduler.
What am I missing here? When does the SingleNodeExecutor figure out it needs to leverage the "cache_directory" field in the resource_dict?
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with SingleNodeExecutor, BaseExecutor, DependencyTaskScheduler, OneProcessTaskScheduler, and the task_scheduler.file tools named in the issue. Trace cache_directory from resource_dict through submission and task execution, then document the point where file-backed caching is selected and how it differs from create_file_executor. Done means the execution path and the missing connection are clearly explained.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- documentation, hpc
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100