Study notes
From AdaGrad to RMSProp
Prof. Prabir Kumar Biswas, IIT Kharagpur · 2:35 · Watch the session
In short
- AdaGrad suffers from a vanishing learning rate due to the monotonic accumulation of squared gradients.
- RMSProp addresses this limitation by using an exponentially decaying average of past squared gradients.
- By avoiding extreme past history, RMSProp achieves faster convergence in convex error terrains.
Key concepts
Check yourself
1. What is the primary limitation of the AdaGrad algorithm mentioned in the session?
Answer: The learning rate vanishes over time due to monotonic accumulation
The session explicitly states that R t monotonically increases with time, which causes the learning rate to vanish in AdaGrad. 1:30
2. How does the RMSProp algorithm differ from AdaGrad regarding gradient calculation?
Answer: It uses an exponentially decaying average of squared gradients
RMSProp replaces the cumulative sum of squared gradients used in AdaGrad with an exponentially decaying average. 1:41
3. In the context of AdaGrad, what happens to the scaling factor R t as iteration t increases?
Answer: It monotonically increases
The professor notes that the accumulation process makes R t increase monotonically over time. 1:30
4. Why does RMSProp converge more rapidly than AdaGrad?
Answer: It discards extreme past gradient history
By using an exponentially decaying average, RMSProp stops considering extreme past history that hinders AdaGrad. 1:55
5. What is the result of using an exponentially decaying average in RMSProp?
Answer: Faster convergence in convex error surfaces
The session mentions that this approach allows the algorithm to converge rapidly once it reaches a locally convex error surface. 1:55
Made with Pravaha: ask your recordings, watch the answer.