allenai / allenai/hyperdecoders
Using Separate Hypernetworks for Each Transformer Layer
Aperta
- Lingua principale
- Python
- Stelle
- 14
- Fork
- 3
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hello,
I noticed in the paper that a single hypernetwork is used to generate parameters for all layers. Have you tried using a separate hypernetwork for each Transformer layer? Would this lead to a larger performance improvement?
Thanks!
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Valutazione
Questa issue non è ancora stata valutata.