NVIDIA / NVIDIA/TensorRT-LLM

[Feature]: Support multiple input types

Open
#15,885 0 comments 0 reactions 1 assignee View on GitHub

@laikhtewari is already working on this.

Since Jul 2, 2026.

feature request
Dominant language
Python
Stars
14.7k
Forks
2.8k
Avg merge
2d 23h
Merged PRs (30d)
489

Description

🚀 The feature, motivation and pitch

My LLM is trained as follows: First, it obtains the speech_embedding from an external encoder and adaptor. Then, it manually projects prompt_1 into prompt_embedding_1, and prompt_2 into prompt_embedding_2. Finally, it concatenates these three embeddings and feeds them to the LLM: [prompt_embedding_1, speech_embedding, prompt_embedding_2]. Therefore, I wanted the TensorRT-LLM Executor API to support multiple input types, including tokens and embeddings.

Alternatives

I tried using PromptTuningConfig to achieve this, but without success.

const int32_t max_token_id = 151668;

int64_t length = 20;
int64_t hidden_size = 2560;
const __nv_bfloat16 *host_bf16_ptr = reinterpret_cast<const __nv_bfloat16 *>(data_bytes.data());
size_t num_values = length * hidden_size;
size_t num_bytes = num_values * sizeof(__nv_bfloat16);
__nv_bfloat16 *device_bf16_ptr = nullptr;
cudaMalloc(reinterpret_cast<void **>(&device_bf16_ptr), num_bytes);
cudaMemcpy(device_bf16_ptr, host_bf16_ptr, num_bytes, cudaMemcpyHostToDevice);

tle::Shape tensor_shape{length, hidden_size};
tle::Tensor embedding_tensor = tle::Tensor::of(
    tle::DataType::kBF16,
    static_cast<void *>(device_bf16_ptr),
    tensor_shape
);

std::vector<int32_t> prompt_1 = {151644, 8948, 198, 151645, 198, 151644, 872, 14880, 44063, 87752, 43815, 105395, 12857, 104811, 28311};
std::vector<int32_t> prompt_2 = {151645, 198, 151644, 77091, 198, 151667, 271, 151668, 271};

std::vector<int32_t> placeholder_ids(length);
for (size_t i = 0; i < length; ++i) {
    placeholder_ids[i] = max_token_id + i;
}

std::vector<int32_t> input_token_ids;
input_token_ids.insert(input_token_ids.end(), prompt_1.begin(), prompt_1.end());
input_token_ids.insert(input_token_ids.end(), placeholder_ids.begin(), placeholder_ids.end());
input_token_ids.insert(input_token_ids.end(), prompt_2.begin(), prompt_2.end());

auto trtllm_request = tle::Request(input_token_ids, max_new_tokens, request->is_streaming);
tle::PromptTuningConfig ptuning_config(embedding_tensor);
trtllm_request.setPromptTuningConfig(ptuning_config);

auto request_id = trtllm_executor_->enqueueRequest(trtllm_request);
Additional context

I also tried tle::Request::setEmbeddingBias() and tle::Request::setMultimodalEmbedding(), but neither worked.

Before submitting a new issue...
  • Make sure you already searched for relevant issues, and checked the documentation and examples for answers to frequently asked questions.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.