Skip to content

← Concept map

AdaGrad Gradient Accumulation

3 explanations from your library.

  1. From AdaGrad to RMSProp, at 0:3010:30

    Prof. Prabir Kumar Biswas, IIT Kharagpur

    From AdaGrad to RMSProp

    “sum of the squares of the past gradients where this past gradient starts from time t equal to 0, ok. So, you are basically the operation that was done in Adagrad algorithm is r t”

    Closest moment in this session: it teaches the idea without listing it as a key concept.

    Open full session

    Moment · 0:30

    From AdaGrad to RMSProp

  2. Gradient Descent Variants and Momentum, at 1:3421:34

    IIT Madras (NPTEL)

    Gradient Descent Variants and Momentum

    “which will and then the Nesterov Accelerated Gradient, AdaGrad, AdaDelta, also RMSProp very similar initially, okay.”

    Closest moment in this session: it teaches the idea without listing it as a key concept.

    Open full session

    Moment · 1:34

    Gradient Descent Variants and Momentum

  3. Learning Rate Decay, at 0:0930:09

    IIT Madras (NPTEL)

    Learning Rate Decay

    “Epoch is when you have gone through your entire data set, epoch is when you have gone through your entire data set once, that is one epoch, right. Let us say you are using stochastic gradient descent or mini batch gradient descent, you have to run through your entire data set and that would be considered as one epoch.”

    Closest moment in this session: it teaches the idea without listing it as a key concept.

    Open full session

    Moment · 0:09

    Learning Rate Decay

Cloudinary under the hood2 Cloudinary URLs make this page. No render servers: each one is generated on request and cached.
  • Compare reel

    3 explanations of this concept, from different sessions, spliced into one video with fl_splice. Each clip is labelled with its speaker.

    Open ↗

    …/video/upload/so_29,eo_44,w_1280,h_720,c_fill/l_video:pravaha:ff2c7367-5718-4b86-9d17-04328cf12206,fl_spl…/so_92.5,eo_108.1,w_1280,h_720,c_fill/fl_layer_apply/fl_layer_apply,g_north_west,x_40,y_40,so_30.6,eo_48.5/f_auto:video,q_auto/pravaha/1ba747ae-bdb0-43b5-bd11-8a0eda32a02c.mp4

    so_29
    starts at 29 s · and 5 more like it
    eo_44
    ends at 44 s · and 5 more like it
    w_1280
    1280 px wide · and 2 more like it
    h_720
    720 px tall · and 2 more like it
    c_fill
    crops to fill the frame exactly · and 2 more like it
    l_video:pravaha:ff2c7367-57…
    another video (pravaha/ff2c7367-5718-4b86-9d17-04328cf12206) as a layer
    fl_splice
    joins the next clip onto the end of this one · and 1 more like it
    fl_layer_apply
    places the layer defined just before · and 4 more like it
    l_video:pravaha:2b17046e-fb…
    another video (pravaha/2b17046e-fb2d-4a2c-ad2c-e84a76f8c28a) as a layer
    l_text:arial_34_bold:1%20%C…
    text layer “1 · Prof. Prabir Kumar Biswas” · and 2 more like it
    co_white
    text colour white · and 2 more like it
    b_rgb:0f766ecc
    background #0f766ecc · and 2 more like it
    g_north_west
    anchored to the top-left corner · and 2 more like it
    x_40
    40 px from the side · and 2 more like it
    y_40
    40 px from the edge · and 2 more like it
    f_auto:video
    best format for each device (e.g. AV1, WebM, MP4, WebP)
    q_auto
    AI-chosen quality: smallest file that still looks right
  • Explanation thumbnail

    The frame at that moment, cropped around the speaker by AI.

    Open ↗

    …/video/upload/so_30.5,c_fill,ar_16:9,w_640,g_auto/f_auto,q_auto/pravaha/1ba747ae-bdb0-43b5-bd11-8a0eda32a02c.jpg

    so_30.5
    starts at 30.5 s
    c_fill
    crops to fill the frame exactly
    ar_16:9
    aspect ratio 16 : 9
    w_640
    640 px wide
    g_auto
    AI picks the focus (the speaker or slide), not a blind centre crop
    f_auto
    best format for each device (e.g. AV1, WebM, MP4, WebP)
    q_auto
    AI-chosen quality: smallest file that still looks right

Want a guided order, basics first? Turn AdaGrad Gradient Accumulation into a short course →