AI-Hypercomputer / AI-Hypercomputer/maxtext

Create a user friendly inference demo

オープン
#532 コメント 0 件 リアクション 0 件 担当者 1 名 @vipannalla が担当を希望しています GitHub で見る
inference
主要言語
Python
スター
2.4k
フォーク
607
平均マージ
2日 19時間
マージ済み PR(30日)
158

説明

This is a feature request.

I like `maxtext` because it is very customizable and efficient for training.
The main issue I’m having is hacking away an inference function. The code is quite complex so not straightforward to do.
The simple `decode.py` works but it seems mainly experimental development for streaming.

I think streaming will be really cool, but we would also benefit from an easy `model.generate(input_ids, attention_mask, params)` function:
* it should allow prefill based on the length of `input_ids` (user responsibility to try to supply not too many shapes to avoid recompilation)
* it should allow batch input, with left padding to support different input length
* should be compilable with `jit`/`pjit`
* allow a few common sampling strategy: greedy, sample (with temperature, top k, top p), beam search
* allow being used without a separate engine/service in case we want to make it part of a larger function that includes multiple models

This PR looked interesting: https://github.com/google/maxtext/pull/402
I think that it was mainly for benchmarking though as it didn’t stop when the entire batch was eos but had a nice prefill functionality.

コントリビューションガイド

コントリビューションガイドを開く

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。