Kong / Kong/developer.konghq.com

Support for NVIDIA Triton inference Server

Open
#3,364 2 comments 0 reactions 0 assignees View on GitHub
internal priority: high product:ai-gateway release-docs
Dominant language
Ruby
Stars
28
Forks
121
Avg merge
1d 4h
Merged PRs (30d)
313

Description

## Description

Add documentation for supporting NVIDIA Triton inference server models and embeddings. This includes referencing Triton as a provider option and showing how to use Triton endpoints with AI Gateway for both model inference and embedding generation.

Reference: https://developer.nvidia.com/dynamo

https://konghq.atlassian.net/browse/AG-269

## Definition of done

* Add Triton to the AI Providers reference page
* Add a subsection describing Triton usage (similar to other providers)
* Provide example requests for:
* Model inference
* Embedding generation
* Add config examples for AI Proxy / AI Proxy Advanced

## Additional information

Person of contact: Jack Tysoe

## Size

M

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the AI Providers reference page and the existing provider subsections, then review the AI Proxy and AI Proxy Advanced documentation. The work is complete when Triton is listed as a provider, usage is documented for model inference and embeddings, and both proxy configuration examples and example requests are included.

Written by the indexing model from the issue text.

Assessment

Domain
documentation
Issue type
Documentation
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
45/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.