huggingface / huggingface/accelerate
[Community Contributions] examples on distributed inference using 🤗 Accelerate
- Dominant language
- Python
- Stars
- 9.9k
- Forks
- 1.5k
- Avg merge
- 5d 2h
- Merged PRs (30d)
- 27
Description
The [`inference/distributed` directory](https://github.com/huggingface/accelerate/tree/main/examples/inference/distributed) houses examples on running distributed inference with `accelerate`:
* Phi2 for language generation
* Stable Diffusion for image generation
The strategy followed there is to load an entire model onto each GPU and sending chunks of a batch through each GPU’s model copy at a time. Synthetic data generation has become an essential toolkit for every ML Engineer. So, it'd be beneficial to extend these examples to include some more use cases:
* Image captioning
* Speech data generation
Some nice to haves:
* Include artifact serialization as done in [this](https://gist.github.com/sayakpaul/cfaebd221820d7b43fae638b4dfa01ba)
* Keep the artifact serialization code under a thread to not block GPU execution
## How can you help?
You could help us contribute an example on any of the above-mentioned use cases or you can come up with your own 🤗 Help us make the art of synthetic data generation scalable, easy, and accessible.
Contributor guide
Assessment
This issue has not been assessed yet.