Kong / Kong/developer.konghq.com
Support for NVIDIA Triton inference Server
- Dominant language
- Ruby
- Stars
- 28
- Forks
- 121
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 313
Description
## Description
Add documentation for supporting NVIDIA Triton inference server models and embeddings. This includes referencing Triton as a provider option and showing how to use Triton endpoints with AI Gateway for both model inference and embedding generation.
Reference: https://developer.nvidia.com/dynamo
https://konghq.atlassian.net/browse/AG-269
## Definition of done
* Add Triton to the AI Providers reference page
* Add a subsection describing Triton usage (similar to other providers)
* Provide example requests for:
* Model inference
* Embedding generation
* Add config examples for AI Proxy / AI Proxy Advanced
## Additional information
Person of contact: Jack Tysoe
## Size
M
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the AI Providers reference page and the existing provider subsections, then review the AI Proxy and AI Proxy Advanced documentation. The work is complete when Triton is listed as a provider, usage is documented for model inference and embeddings, and both proxy configuration examples and example requests are included.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation
- Issue type
- Documentation
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 45/100