InternLM / InternLM/xtuner

支持自定义视觉编码器么(llava-llama3)?

Open
#668 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
5.2k
Forks
448
Avg merge
3d 15h
Merged PRs (30d)
26

Description

支持自定义视觉编码器么(llava-llama3)?
例如将clip换成siglip?
该如何实现?哪些代码需要修改?

Contributor guide

Open the contributing guide

Research direction

Start by searching the repository for the llava-llama3 implementation and its CLIP integration points. Determine whether replacing CLIP with SigLIP is supported, which configuration or model components would need changes, and what validation would demonstrate that the custom visual encoder works.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.