google-research / google-research/m2svid
Feature Request regarding Decoder
- Dominant language
- Python
- Stars
- 37
- Forks
- 8
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
I am came across something in other reasearch article i.e. guided vae decoder for svd for high quality stereo movie conversion i.e.
High Quality Details
Epipolar Cross-Attention (Structured Skip Connections) in the VAE Decoder recovers details and mitigates binocular rivalry. Our Guided Decoder produces better results than the Stable Diffusion VAE Decoder.
Paper reference: https://elastic3d.github.io/
Therefore I like to request if vae decoder for M2SVid project can be retrained which can be used to improve the inpainting quality.
Contributor guide
Research direction
The issue names no files, tests, or entry points. Start by comparing the referenced Elastic3D paper with M2SVid’s decoder and training setup, then define the retraining data, objective, and inpainting-quality evaluation. Done would require an agreed decoder change and demonstrated quality improvement.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Quiet
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100