NVIDIA-NeMo / NVIDIA-NeMo/Guardrails

Error configuring NeMo Guardrails with a TensorRT LLM serving on TRT-LLM server

Open
#657 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
7.2k
Forks
842
Avg merge
3d 1h
Merged PRs (30d)
25

Description

Has anyone successfully integrated TensorRT LLM with NeMo Guardrails? There is a lack of documentation on utilizing TensorRT LLM for Nemo Guardrails so I am hoping for some guidance here.

I am getting the following error when trying to execute self check input:
Error while execution self_check_input: [StatusCode.UNAVAILABLE] DNS resolution failed for http://triton-models:8000/v2/models/ensemble/generate: C-ares status is not ARES_SUCCESS qtype=AAAA name=http://triton-models:8000/v2/models/ensemble/generate is_balancer=0: Could not contact DNS servers

My config.yml:

models:
 - type: main
   engine: trt_llm
   parameters:
    server_url: http://triton-models:8000/v2/models/ensemble/generate

# These instructions configure the bot to answer questions about the employee handbook and the company's policies.
instructions:
  - type: general
    content: |
      Below is a conversation between a user and a bot called X Virtual Assistant.
      The bot is talkative and provides lots of specific details from its context.
      If the bot does not know the answer to a question, it truthfully says it does not know.

rails:
  # Input rails are invoked when a new message from the user is received.
  input:
    flows:
      - self check input

  # Output rails are triggered after a bot message has been generated.
  output:
    flows:
      - self check output

langchain integration:

# load NeMo Guardrails
config = RailsConfig.from_path("./guardrails_config")
guardrails = RunnableRails(config, passthrough=False)

class ChatChain(LLMChain):
    def __init__(self, llm, examples):
        super().__init__(llm)
        prompt = chat_template.format_chat_prompt(examples)
        custom_chain = (
            prompt
            | llm
            | StrOutputParser()
        )

        mem_chain = RunnableWithMessageHistory(
            custom_chain,
            get_message_history,
            input_messages_key="input",
            history_messages_key="history",
            output_messages_key="output"
        )

        self.mem_chain_with_guardrails = guardrails | mem_chain
    
    def get_chain(self):
        return self.mem_chain_with_guardrails

Pip version:
nemoguardrails==0.9.0
tritonclient[all]

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the config.yml model settings and the LangChain integration shown in the issue, then trace how the trt_llm server_url is used during self check input. Reproduce or inspect the reported DNS-resolution failure involving the Triton endpoint. Done should provide documented TensorRT LLM integration guidance for NeMo Guardrails, including the relevant configuration and the observed connection requirements.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
ai
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.