AlibabaResearch / AlibabaResearch/DAMO-ConvAI

Inference script for the level-3 data

未关闭
#102 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
api-bank
主要语言
Python
星标
1.6k
派生
250
PR 合并指标
30 天内没有已合并 PR

描述

Hello,

I am trying to run evaluation of the Lynx model on level-3 data.
However, I did not find the inference script and unsure of how to reproduce it.

My question is:
Did the model generate all steps of tool calling during the lvl-3 evaluation and received feedback from tool after each step? How the errors were handled? How the errors were parsed and added into the tool output?
How the errors were incorporated in the tool output if the model hallucinated and generated something that couldn't be parsed?
Is it possible to provide the full script of running the Lynx model on lvl-1, lvl-2 and lvl-3 data?

Thanks in advance
Gregory

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。