hzwer / hzwer/Practical-RIFE

Request to incorporate InterpAny-Clearer's technology in Practical-RIFE

Open
#69 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
1k
Forks
131
PR merge metrics
No merged PRs in 30d

Description

I would like to thank you very much for continually adding new, practical RIFE models. For this last one I am particularly grateful to you. Firstly because you are open to the work of others that can improve the already awesome RIFE. Secondly, because I was very interested in adding just this loss function to RIFE, which I requested from the developer InterpAny-Clearer here: Request for [D,R] RIFE trained on Style loss, also called Gram matrix loss (the best perceptual loss function)

As you have applied gram matrix loss to RIFE v4.17 training, I would like to ask you to conduct x128 frame interpolation tests of RIFE v4.15 and v4.17 for the files I0_0.png and I0_1.png and also I1_0.png and I1_1.png available here: https://github.com/zzh-tech/InterpAny-Clearer/tree/main/demo and posting their results as GIF files in the Practical-RIFE repository or here. There will be a total of 4 GIF files.

I think this test will be of interest to any Practical-RIFE enthusiast and will answer an important question: can gram matrix loss really improve the performance of practical RIFE models? And, in particular, whether it eliminates or at least mitigates to some extent the undesirable distortions occurring with VGG loss, which can be seen particularly clearly in the lower example for the column: [D,R] RIFE-vgg (Ours) in the table available at this link: https://github.com/zzh-tech/InterpAny-Clearer?tab=readme-ov-file#time-indexing-vs-distance-indexing

This test will not only compare the differences in the loss function for version 4.15 and version 4.17 of RIFE, but will also give an answer to the question of whether the application of the InterpAny-Clearer enhancement will further improve the clarity of frame interpolation for practical RIFE models, just as it improves it for the base RIFE model, as shown in the table to which I have provided a link above.

We will have a total comparison of 5 different versions of RIFE - 3 already available:

[T] RIFE - RIFE base model
[D,R] RIFE - RIFE base model with full InterpAny-Clearer upgrade
[D,R] RIFE-vgg - RIFE base model with VGG loss and full InterpAny-Clearer upgrade

and 2 versions of the practical RIFE models, which I ask you to compare here:

RIFE v4.15 with standard perceptual loss
RIFE v4.17 with gram matrix loss

This comparison will hopefully give evidence to justify the use of gram matrix loss in all future RIFE practical models. Also, I hope it will demonstrate the need for the use of InterpAny-Clearer in the next RIFE v4.18 model.

I would also like to use these GIF files of 5 different versions of RIFE (of course with links to the sources of the comparisons) in the introduction to my rankings Video Frame Interpolation Rankings and Video Deblurring Rankings, where I would like to encourage other researchers to use the best and above all proven solutions for practical applications when developing new models.

I hope that the next GIF file will already be for RIFE v4.18 with gram matrix loss and the InterpAny-Clearer enhancement. The clarity of the frame interpolation is more important than the super resolution, especially when new monitors already reach 500Hz and the original frames can be seen on the screen very briefly and most of the time the interpolated frames are visible.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with the demo inputs I0_0.png, I0_1.png, I1_0.png, and I1_1.png, then identify how RIFE v4.15 and v4.17 are run for x128 frame interpolation. Done means producing and posting four GIF comparisons, with the versions and loss-function differences identified clearly.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.