Asset Details

MbrlCatalogueTitleDetail

Do you wish to reserve the book?

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

by Xiang, Bingjie , Huang, Xingwang , Wu, Xinyi , Lu, Huaizheng , Huang, Weifang , Li, Chaopeng

in Comparative analysis / Computational linguistics / Deep learning / exploding gradients / Explosions / gradient normalization / Hypotheses / Language processing / Methods / Natural language interfaces / Neural networks / Neurons / probability distribution characteristics / recurrent neural networks / vanishing gradients

2024

Yes Please

Hey, we have placed the reservation for you!

By the way, why not check out events that you can attend while you pick your title.

Oops! Something went wrong.

Looks like we were not able to place the reservation. Kindly try again later.

Are you sure you want to remove the book from the shelf?

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

by Xiang, Bingjie , Huang, Xingwang , Wu, Xinyi , Lu, Huaizheng , Huang, Weifang , Li, Chaopeng

2024

Confirm

Do you wish to request the book?

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

by Xiang, Bingjie , Huang, Xingwang , Wu, Xinyi , Lu, Huaizheng , Huang, Weifang , Li, Chaopeng

2024

Please be aware that the book you have requested cannot be checked out. If you would like to checkout this book, you can reserve another copy

How would you like to get it?

Submit

We have requested the book for you!

Your request is successful and it will be processed during the Library working hours. Please check the status of your request in My Requests.

Oops! Something went wrong.

Looks like we were not able to place your request. Kindly try again later.

Journal Article

Optimizing Recurrent Neural Networks: A Study on Gradient Normalization of Weights for Enhanced Training Efficiency

Xiang, Bingjie,

Huang, Xingwang,

Wu, Xinyi,

Lu, Huaizheng,

Huang, Weifang,

Li, Chaopeng

2024

Overview

Recurrent Neural Networks (RNNs) are classical models for processing sequential data, demonstrating excellent performance in tasks such as natural language processing and time series prediction. However, during the training of RNNs, the issues of vanishing and exploding gradients often arise, significantly impacting the model’s performance and efficiency. In this paper, we investigate why RNNs are more prone to gradient problems compared to other common sequential networks. To address this issue and enhance network performance, we propose a method for gradient normalization of network weights. This method suppresses the occurrence of gradient problems by altering the statistical properties of RNN weights, thereby improving training effectiveness. Additionally, we analyze the impact of weight gradient normalization on the probability-distribution characteristics of model weights and validate the sensitivity of this method to hyperparameters such as learning rate. The experimental results demonstrate that gradient normalization enhances the stability of model training and reduces the frequency of gradient issues. On the Penn Treebank dataset, this method achieves a perplexity level of 110.89, representing an 11.48% improvement over conventional gradient descent methods. For prediction lengths of 24 and 96 on the ETTm1 dataset, Mean Absolute Error (MAE) values of 0.778 and 0.592 are attained, respectively, resulting in 3.00% and 6.77% improvement over conventional gradient descent methods. Moreover, selected subsets of the UCR dataset show an increase in accuracy ranging from 0.4% to 6.0%. The gradient normalization method enhances the ability of RNNs to learn from sequential and causal data, thereby holding significant implications for optimizing the training effectiveness of RNN-based models.

Share this book

Add to My Shelf

Publisher

MDPI AG

Subject

Comparative analysis

/ Computational linguistics

/ Deep learning

/ exploding gradients

/ Explosions

/ gradient normalization

/ Hypotheses