google-research / google-research/m2svid

Feature Request regarding Decoder

Open
#4 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
37
Forks
8
PR merge metrics
No merged PRs in 30d

Description

Hi,

I am came across something in other reasearch article i.e. guided vae decoder for svd for high quality stereo movie conversion i.e.
High Quality Details
Epipolar Cross-Attention (Structured Skip Connections) in the VAE Decoder recovers details and mitigates binocular rivalry. Our Guided Decoder produces better results than the Stable Diffusion VAE Decoder.

Paper reference: https://elastic3d.github.io/

Therefore I like to request if vae decoder for M2SVid project can be retrained which can be used to improve the inpainting quality.

Contributor guide

Open the contributing guide

Research direction

The issue names no files, tests, or entry points. Start by comparing the referenced Elastic3D paper with M2SVid’s decoder and training setup, then define the retraining data, objective, and inpainting-quality evaluation. Done would require an agreed decoder change and demonstrated quality improvement.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Quiet
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.