aws / aws/fmeval

[Feature] Image fields for multi-modal models

Open
#247 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
291
Forks
60
PR merge metrics
No merged PRs in 30d

Description

I'm trying to evaluate Claude v3's performance for some document understanding tasks, with a workflow that includes passing the image of the page in as one of the inputs.

Is fmeval considering native handling for image/multi-modal fields in input datasets?

Contributor guide

Open the contributing guide

Research direction

The issue does not name any files, tests, or entry points. Start by locating how input datasets and model inputs are represented, then trace the Claude v3 evaluation path to determine where image or multi-modal fields would fit. Done would mean a defined, tested approach for passing image inputs through an evaluation workflow.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.