huggingface / huggingface/accelerate

[Community Contributions] examples on distributed inference using 🤗 Accelerate

Open
#3,078 11 comments 2 reactions 0 assignees View on GitHub
contributions-welcome good first issue wip
Dominant language
Python
Stars
9.9k
Forks
1.5k
Avg merge
5d 2h
Merged PRs (30d)
27

Description

The [`inference/distributed` directory](https://github.com/huggingface/accelerate/tree/main/examples/inference/distributed) houses examples on running distributed inference with `accelerate`:

* Phi2 for language generation
* Stable Diffusion for image generation

The strategy followed there is to load an entire model onto each GPU and sending chunks of a batch through each GPU’s model copy at a time. Synthetic data generation has become an essential toolkit for every ML Engineer. So, it'd be beneficial to extend these examples to include some more use cases:

* Image captioning
* Speech data generation

Some nice to haves:

* Include artifact serialization as done in [this](https://gist.github.com/sayakpaul/cfaebd221820d7b43fae638b4dfa01ba)
* Keep the artifact serialization code under a thread to not block GPU execution

## How can you help?

You could help us contribute an example on any of the above-mentioned use cases or you can come up with your own 🤗 Help us make the art of synthetic data generation scalable, easy, and accessible.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.