finos / finos/community

Proposed: AI Evaluation Framework - APPROVED

Open
#438 2 comments 0 reactions 2 assignees Claimed by @TheJuanAndOnly99 View on GitHub
contribution labs
Dominant language
JavaScript
Stars
75
Forks
40
Avg merge
1m
Merged PRs (30d)
5

Description

### Contribution Prerequisites

- [x] I have read [FINOS Contribution Requirements](https://community.finos.org/docs/governance/Software-Projects/contribution-compliance-requirements).
- [x] I confirm I am contributing on behalf of a FINOS member OR I will seek a maintainer from a member during socialization.
- [x] I have reviewed the [FINOS Contribution Process](https://community.finos.org/docs/governance/Software-Projects/contribution/).
- [x] I have reviewed the [FINOS Project Lifecycle](https://community.finos.org/docs/governance/project-lifecycle) stages.
- [x] I am familiar with [FINOS Governance](https://community.finos.org/docs/governance) and [Maintainer Responsibilities](https://community.finos.org/docs/finos-maintainers-cheatsheet/).

### Target Lifecycle Stage

Labs

### FINOS Member Organization Name

Red Hat

### Name

AI Evaluation Framework

### Slug

ai-evaluation-framework

### Business Problem & Solution

**Business Problem**

- General-purpose AI benchmarks often fall short for the unique requirements of the finance sector.
- Financial tasks rarely yield a single "correct" answer and require precise alignment with financial ontologies and regulatory standards.
- Current focus on model-level evaluations do not fully capture the complexity of real-world AI systems, which combine AI and non-AI components.
- Enterprises increasingly face a "Day 2" Reality Gap when moving AI agents into production, struggling with information overload, runaway costs, and a lack of governance and trust.
- There is a significant risk of "safety-washing," where general capability improvements in models are misinterpreted as advancements in actual system safety.

**Solution**

- Establish a common, open, and transparent evaluation suite for Generative AI and agentic applications in financial services, by utilizing a "Taxonomy-First" approach, mapping specific financial use cases (e.g., credit risk analysis) directly to material risks, and subsequently to measurable metrics.
- The framework bridges the trust gap by transitioning from traditional MLOps to AgentOps, enabling continuous, outcome-oriented evaluation and behavioral auditing for complex AI workflows. It acts as the critical measurement layer within the FINOS Governance-as-Code pipeline, allowing organizations to continuously evaluate their models, agents, and controls against FINOS AI Governance Framework (AIGF) risks.

### Mission Statement

To establish a common, open, and transparent evaluation suite for Generative AI in financial services that bridges the AI trust gap by transitioning from model-level scoring to continuous, system-level AgentOps evaluations, mapping real-world financial use cases directly to material risks, regulatory requirements, and measurable metrics.

**Strategic Objectives**

- System-Level Evaluation: Move beyond generic model leaderboards to rigorously evaluate end-to-end workflows, multi-agent systems, and retrieval-augmented generation (RAG) operating in complex financial contexts.
- Taxonomy-First Framework: Systematically align financial use cases to operational and compliance risks (such as hallucinations, bias, and non-determinism) to define clear, quantitative industry benchmarks for trustworthy AI.
- Governance-as-Code Integration: Function as the open measurement layer in the FINOS Governance-as-Code pipeline, offering continuous auditability, metacognitive guardrails, and runtime evaluation across the system lifecycle.
- Mutualized Innovation: Accelerate enterprise-grade AI adoption across the financial sector through open synthetic datasets, repeatable test cases, and shared reference architectures.

### Existing Materials

Core Framework Repo: [LINK](https://github.com/finos-labs/ai-evals-framework)
Reference Implementation: [LINK](https://github.com/finos-labs/finsight-agent) (The FinSight AI Agent, a multi-agent, metacognitive system designed for earnings call analysis).
Leaderboard Repo: [LINK](https://github.com/finos-labs/Open-Financial-LLMs-Leaderboard/) (A framework for evaluating LLM performance across diverse financial tasks).
OSFF Toronto Training Workshop: [LINK](https://github.com/finos-labs/ai-evals-framework/blob/main/OSFF_Toronto_2026_Workshop_Beyond%20the%20Black%20Box_%20Operationalizing%20AgentOps%20and%20_Glass-Box_%20Evaluations%20for%20Financial%20AI%20(1).pdf)

### Additional Information (Optional)

This will be a multi-repo project with ai-eval-framework holding the framework itself and the various use-case implementation contributed in their own repository as software.

### Maintainer Team

| Name | Affiliation | Email | GitHub Username |
| :--- | :--- | :--- | :--- |
| Vincent Caldeira | Red Hat | vincent.caldeira@redhat.com | @caldeirav |
| Jamie Macdonald | ScottLogic | jmacdonald@scottlogic.com | @JamieWhitMac |

### Compliance & Requirements Agreement

- [x] I understand that before being able to contribute all maintainers and contributors will need to be covered by a [Contributor License Agreement](https://community.finos.org/docs/governance/Software-Projects/contribution-compliance-requirements#contributor-license-agreement) or commits must satisfy the [Developer Certificate of Origin](https://developercertificate.org/) (DCO) requirements.
- [x] I confirm that software will be covered by the [Apache 2.0 License](https://community.finos.org/docs/governance/Software-Projects/contribution-compliance-requirements#license-information) and any [third-party code is compatible](https://community.finos.org/docs/governance/Software-Projects/contribution-compliance-requirements#third-party-code-compliance).
- [x] I have selected the appropriate [lifecycle stage](https://community.finos.org/docs/governance/project-lifecycle) based on that stage's requirements. I agree and commit to the Acceptance & Maintenance Requirements for my target stage.

### Administrative Consent & Transfer Authorization

- [x] I agree to grant 'finos-admin' Admin/Owner access to the repository for security auditing and setup.
- [x] I authorize the transfer of this code/repository to the FINOS GitHub organization upon project acceptance.
- [x] I confirm I have the legal authority (individual or corporate) to grant these permissions.

### Socialization Acknowledgement

- [x] I understand I may need to socialize this proposal to community@finos.org before a TOC review to gauge community interest.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.