bytedance / bytedance/MoMA

what is the CLIP for supervising the generated image embedding in Eq(4)

Open
#4 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
234
Forks
17
PR merge metrics
No merged PRs in 30d

Description

Thanks for the excellent work, and I am trying to reimplement moma. Since moma has adopted the pre-trained weights from IP-Adapter for initilaization, the CLIP version is OpenCLIP-ViT-H-14, is that correct? Or the CLIP is the one mentioned in config.json ``openai/clip-vit-large-patch14-336''? Thanks!

Contributor guide

No contributing guide indexed for this repository

Research direction

Read the CLIP references in config.json and compare them with the OpenCLIP-ViT-H-14 and openai/clip-vit-large-patch14-336 names cited in the issue. Confirm which model supervises the generated image embedding in Eq. (4), then document the definitive choice and resolve the question.

Written by the indexing model from the issue text.

Assessment

Domain
machine-learning
Issue type
Documentation
Difficulty
1/5
Estimated time
Under an hour
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.