[AI Evaluation] Explore setting up a pipeline for integration tests that can actually invoke the LLM
Open
area-ai-eval
area-infrastructure
- Dominant language
- C#
- Stars
- 3.2k
- Forks
- 894
- Avg merge
- 1d 12h
- Merged PRs (30d)
- 23
Description
This would essentially help us orchestrate and dogfood how tests would run in the end user's environments. The idea is to have a pipeline that can run the tests against a (Azure Storage based) cache which expires periodically and gets refreshed with updated LLM evaluation responses as the tests are run over time.
Contributor guide
Assessment
This issue has not been assessed yet.