Stored Object: paper:arxiv.2410.01131
- Dominant language
- JavaScript
- Stars
- 1
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Description
{
"primary_id": "arxiv.2410.01131",
"source": "arxiv",
"sourceId": "2410.01131",
"url": "https://arxiv.org/abs/2410.01131",
"title": "nGPT: Normalized Transformer with Representation Learning on the\n Hypersphere",
"authors": "Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, Boris Ginsburg",
"abstract": "We propose a novel neural network architecture, the normalized Transformer\n(nGPT) with representation learning on the hypersphere. In nGPT, all vectors\nforming the embeddings, MLP, attention matrices and hidden states are unit norm\nnormalized. The input stream of tokens travels on the surface of a hypersphere,\nwith each layer contributing a displacement towards the target output\npredictions. These displacements are defined by the MLP and attention blocks,\nwhose vector components also reside on the same hypersphere. Experiments show\nthat nGPT learns much faster, reducing the number of training steps required to\nachieve the same accuracy by a factor of 4 to 20, depending on the sequence\nlength.",
"timestamp": "2025-02-28T22:02:42.541Z",
"rating": "novote",
"arxivId": "2410.01131",
"arxiv_tags": [
"cs.LG",
"cs.AI"
],
"published_date": "2024-10-01T23:50:09Z"
}
---
*Copied from dmarx/papers-feed#1480*
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.