{
  "id": 583016,
  "title": "LR Decay For MAE",
  "url": "/competitions/waveform-inversion/discussion/583016",
  "author_name": "",
  "post_date": "2025-06-04T08:51:12.145764100Z",
  "votes": 39,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> shared it, and it was the first thing I modified in public notebooks as well: it is important to have a decreasing learning rate when optimizing MAE.</p>\n<p>Here is  a very (too?) simple example to show why.</p>\n<p>Let's say we optimize the simplest possible model: a constant prediction, using MAE.</p>\n<p>More formally, our model is a constant function <code>f(x) = a</code> for all input <code>x</code>. Let's assume the label is <code>b = 0.75</code>.</p>\n<p>Then we want to minimize <code>MAE(f(x), b).</code></p>\n<p>Let's assume we use gradient descent, with a learning rate of <code>1</code>.</p>\n<p>The gradient of <code>f</code>is:</p>\n<ul>\n<li><code>-1</code>if <code>a &lt;b</code></li>\n<li>undefined for <code>a = b</code></li>\n<li><code>1</code> if <code>a &gt; b</code></li>\n</ul>\n<p>Lets assume that <code>f</code>is  initialized with <code>a = 0</code>.</p>\n<p>The gradient of <code>f</code> at <code>0</code> is <code>-1</code>. We make a step of minus gradient times the learning rate, i.e. a step of <code>1</code>. This yields <code>a = 1</code>.</p>\n<p>The gradient of <code>f</code> at <code>1</code> is <code>1</code> hence we make a step of <code>-1.</code> This yields <code>a = 0</code>.</p>\n<p>The value of a will oscillate between <code>0</code> and <code>1</code> without ever converging to <code>b</code>. </p>\n<p>It is easy to check that there will be an oscillation with whatever constant learning rate is used, except if by chance <code>a</code> becomes equal to <code>b</code> at some point. And when oscillation starts it keeps on forever.</p>\n<p>My example is simplistic, but this oscillation happens with more complex models as well.</p>\n<p>Reducing the learning rate fixes that. Say learning rate is decreased by <code>0.1</code> at each iteration. Then we would get these steps:</p>\n<table>\n<thead>\n<tr>\n<th>a</th>\n<th>gradient</th>\n<th>lr</th>\n<th>step</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>-1</td>\n<td>1</td>\n<td>1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>0.9</td>\n<td>-0.9</td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>-1</td>\n<td>0.8</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>0.9</td>\n<td>1</td>\n<td>0.7</td>\n<td>-0.7</td>\n</tr>\n<tr>\n<td>0.2</td>\n<td>-1</td>\n<td>0.6</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>0.8</td>\n<td>1</td>\n<td>0.5</td>\n<td>-0.5</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>-1</td>\n<td>0.4</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>0.7</td>\n<td>-1</td>\n<td>0.3</td>\n<td>0.3</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>0.2</td>\n<td>-0.2</td>\n</tr>\n<tr>\n<td>0.8</td>\n<td>1</td>\n<td>0.1</td>\n<td>-0.1</td>\n</tr>\n<tr>\n<td>0.7</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>We end up much closer to the ground truth.</p>\n<p>Conclusion: use a learning rate scheduler to decrease your learning rate over time.</p>",
  "messages": [
    {
      "id": "3216895",
      "postDate": "06/04/2025 08:51:12",
      "content": "<p><a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> and <a href=\"https://www.kaggle.com/harshitsheoran\" target=\"_blank\">@harshitsheoran</a> shared it, and it was the first thing I modified in public notebooks as well: it is important to have a decreasing learning rate when optimizing MAE.</p>\n<p>Here is  a very (too?) simple example to show why.</p>\n<p>Let's say we optimize the simplest possible model: a constant prediction, using MAE.</p>\n<p>More formally, our model is a constant function <code>f(x) = a</code> for all input <code>x</code>. Let's assume the label is <code>b = 0.75</code>.</p>\n<p>Then we want to minimize <code>MAE(f(x), b).</code></p>\n<p>Let's assume we use gradient descent, with a learning rate of <code>1</code>.</p>\n<p>The gradient of <code>f</code>is:</p>\n<ul>\n<li><code>-1</code>if <code>a &lt;b</code></li>\n<li>undefined for <code>a = b</code></li>\n<li><code>1</code> if <code>a &gt; b</code></li>\n</ul>\n<p>Lets assume that <code>f</code>is  initialized with <code>a = 0</code>.</p>\n<p>The gradient of <code>f</code> at <code>0</code> is <code>-1</code>. We make a step of minus gradient times the learning rate, i.e. a step of <code>1</code>. This yields <code>a = 1</code>.</p>\n<p>The gradient of <code>f</code> at <code>1</code> is <code>1</code> hence we make a step of <code>-1.</code> This yields <code>a = 0</code>.</p>\n<p>The value of a will oscillate between <code>0</code> and <code>1</code> without ever converging to <code>b</code>. </p>\n<p>It is easy to check that there will be an oscillation with whatever constant learning rate is used, except if by chance <code>a</code> becomes equal to <code>b</code> at some point. And when oscillation starts it keeps on forever.</p>\n<p>My example is simplistic, but this oscillation happens with more complex models as well.</p>\n<p>Reducing the learning rate fixes that. Say learning rate is decreased by <code>0.1</code> at each iteration. Then we would get these steps:</p>\n<table>\n<thead>\n<tr>\n<th>a</th>\n<th>gradient</th>\n<th>lr</th>\n<th>step</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>-1</td>\n<td>1</td>\n<td>1</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>0.9</td>\n<td>-0.9</td>\n</tr>\n<tr>\n<td>0.1</td>\n<td>-1</td>\n<td>0.8</td>\n<td>0.8</td>\n</tr>\n<tr>\n<td>0.9</td>\n<td>1</td>\n<td>0.7</td>\n<td>-0.7</td>\n</tr>\n<tr>\n<td>0.2</td>\n<td>-1</td>\n<td>0.6</td>\n<td>0.6</td>\n</tr>\n<tr>\n<td>0.8</td>\n<td>1</td>\n<td>0.5</td>\n<td>-0.5</td>\n</tr>\n<tr>\n<td>0.3</td>\n<td>-1</td>\n<td>0.4</td>\n<td>0.4</td>\n</tr>\n<tr>\n<td>0.7</td>\n<td>-1</td>\n<td>0.3</td>\n<td>0.3</td>\n</tr>\n<tr>\n<td>1</td>\n<td>1</td>\n<td>0.2</td>\n<td>-0.2</td>\n</tr>\n<tr>\n<td>0.8</td>\n<td>1</td>\n<td>0.1</td>\n<td>-0.1</td>\n</tr>\n<tr>\n<td>0.7</td>\n<td>1</td>\n<td>0</td>\n<td>0</td>\n</tr>\n</tbody>\n</table>\n<p>We end up much closer to the ground truth.</p>\n<p>Conclusion: use a learning rate scheduler to decrease your learning rate over time.</p>",
      "rawMarkdown": "brendanartley and @harshitsheoran shared it, and it was the first thing I modified in public notebooks as well: it is important to have a decreasing learning rate when optimizing MAE.\n\nHere is  a very (too?) simple example to show why.\n\nLet's say we optimize the simplest possible model: a constant prediction, using MAE.\n\nMore formally, our model is a constant function `f(x) = a` for all input `x`. Let's assume the label is `b = 0.75`.\n\nThen we want to minimize `MAE(f(x), b).`\n\nLet's assume we use gradient descent, with a learning rate of `1`.\n\nThe gradient of `f `is:\n- `-1 `if `a <b`\n- undefined for `a = b`\n- `1` if `a > b`\n\nLets assume that `f `is  initialized with `a = 0`.\n\nThe gradient of `f` at `0` is `-1`. We make a step of minus gradient times the learning rate, i.e. a step of `1`. This yields `a = 1`.\n\nThe gradient of `f` at `1` is `1` hence we make a step of `-1.` This yields `a = 0`.\n\nThe value of a will oscillate between `0` and `1` without ever converging to `b`. \n\nIt is easy to check that there will be an oscillation with whatever constant learning rate is used, except if by chance `a` becomes equal to `b` at some point. And when oscillation starts it keeps on forever.\n\nMy example is simplistic, but this oscillation happens with more complex models as well.\n\nReducing the learning rate fixes that. Say learning rate is decreased by `0.1` at each iteration. Then we would get these steps:\n\n|  a |  gradient | lr | step |\n| --- | --- | -- | -- |\n| 0  |  -1  | 1 | 1 |\n| 1 | 1 | 0.9 | -0.9 |\n| 0.1 | -1 | 0.8 | 0.8 |\n| 0.9 | 1 | 0.7 | -0.7 |\n| 0.2 | -1 | 0.6 | 0.6 |\n| 0.8 | 1 | 0.5 | -0.5 |\n| 0.3 | -1 | 0.4 | 0.4 |\n| 0.7 | -1 | 0.3 | 0.3 |\n| 1 | 1 | 0.2 | -0.2 |\n| 0.8 | 1 | 0.1 | -0.1 |\n| 0.7 | 1 | 0 | 0|\n\nWe end up much closer to the ground truth.\n\nConclusion: use a learning rate scheduler to decrease your learning rate over time.",
      "votes": null
    },
    {
      "id": "3217031",
      "postDate": "06/04/2025 12:45:37",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  - for the explanation … I didnt know earlier why the decreasing learning rate is working in this problem . </p>\n<p>Also thank you for the nice ensemble implementation , it helped almost + 0.5 CV for me .. </p>",
      "rawMarkdown": "Thank you @cpmpml  - for the explanation ... I didnt know earlier why the decreasing learning rate is working in this problem . \n\nAlso thank you for the nice ensemble implementation , it helped almost + 0.5 CV for me ..",
      "votes": null
    },
    {
      "id": "3217041",
      "postDate": "06/04/2025 13:05:21",
      "content": "<p>Glad what I share helps! Thanks for the feedback.</p>",
      "rawMarkdown": "Glad what I share helps! Thanks for the feedback.",
      "votes": null
    },
    {
      "id": "3218954",
      "postDate": "06/07/2025 01:12:47",
      "content": "<p>Theoretically, warm-up can also be helpful, as it prevents erratic behavior when the optimizer first navigates sharp transitions or abrupt changes in the loss landscape, where gradients are non-smooth or even discontinuous. Additionally, gradient clipping can complement warm-up by helping to prevent instability during the early stages of training.</p>",
      "rawMarkdown": "Theoretically, warm-up can also be helpful, as it prevents erratic behavior when the optimizer first navigates sharp transitions or abrupt changes in the loss landscape, where gradients are non-smooth or even discontinuous. Additionally, gradient clipping can complement warm-up by helping to prevent instability during the early stages of training.",
      "votes": null
    },
    {
      "id": "3218968",
      "postDate": "06/07/2025 01:49:05",
      "content": "<p>Sure, but warming up helping is not specific to MAE.</p>",
      "rawMarkdown": "Sure, but warming up helping is not specific to MAE.",
      "votes": null
    },
    {
      "id": "3228317",
      "postDate": "06/20/2025 03:29:31",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> it is someway hard to implement in spark due to limited naive support. any notebook suggestion would be perfect</p>",
      "rawMarkdown": "thanks for sharing @cpmpml it is someway hard to implement in spark due to limited naive support. any notebook suggestion would be perfect",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217031,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "06/04/2025 12:45:37",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  - for the explanation … I didnt know earlier why the decreasing learning rate is working in this problem . </p>\n<p>Also thank you for the nice ensemble implementation , it helped almost + 0.5 CV for me .. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3217041,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/04/2025 13:05:21",
          "content": "<p>Glad what I share helps! Thanks for the feedback.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3218954,
      "author_name": "yekenot",
      "author_url": "",
      "post_date": "06/07/2025 01:12:47",
      "content": "<p>Theoretically, warm-up can also be helpful, as it prevents erratic behavior when the optimizer first navigates sharp transitions or abrupt changes in the loss landscape, where gradients are non-smooth or even discontinuous. Additionally, gradient clipping can complement warm-up by helping to prevent instability during the early stages of training.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3218968,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "06/07/2025 01:49:05",
          "content": "<p>Sure, but warming up helping is not specific to MAE.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3228317,
      "author_name": "ibozkurt79",
      "author_url": "",
      "post_date": "06/20/2025 03:29:31",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> it is someway hard to implement in spark due to limited naive support. any notebook suggestion would be perfect</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3216895": "brendanartley and @harshitsheoran shared it, and it was the first thing I modified in public notebooks as well: it is important to have a decreasing learning rate when optimizing MAE.\n\nHere is  a very (too?) simple example to show why.\n\nLet's say we optimize the simplest possible model: a constant prediction, using MAE.\n\nMore formally, our model is a constant function `f(x) = a` for all input `x`. Let's assume the label is `b = 0.75`.\n\nThen we want to minimize `MAE(f(x), b).`\n\nLet's assume we use gradient descent, with a learning rate of `1`.\n\nThe gradient of `f `is:\n- `-1 `if `a <b`\n- undefined for `a = b`\n- `1` if `a > b`\n\nLets assume that `f `is  initialized with `a = 0`.\n\nThe gradient of `f` at `0` is `-1`. We make a step of minus gradient times the learning rate, i.e. a step of `1`. This yields `a = 1`.\n\nThe gradient of `f` at `1` is `1` hence we make a step of `-1.` This yields `a = 0`.\n\nThe value of a will oscillate between `0` and `1` without ever converging to `b`. \n\nIt is easy to check that there will be an oscillation with whatever constant learning rate is used, except if by chance `a` becomes equal to `b` at some point. And when oscillation starts it keeps on forever.\n\nMy example is simplistic, but this oscillation happens with more complex models as well.\n\nReducing the learning rate fixes that. Say learning rate is decreased by `0.1` at each iteration. Then we would get these steps:\n\n|  a |  gradient | lr | step |\n| --- | --- | -- | -- |\n| 0  |  -1  | 1 | 1 |\n| 1 | 1 | 0.9 | -0.9 |\n| 0.1 | -1 | 0.8 | 0.8 |\n| 0.9 | 1 | 0.7 | -0.7 |\n| 0.2 | -1 | 0.6 | 0.6 |\n| 0.8 | 1 | 0.5 | -0.5 |\n| 0.3 | -1 | 0.4 | 0.4 |\n| 0.7 | -1 | 0.3 | 0.3 |\n| 1 | 1 | 0.2 | -0.2 |\n| 0.8 | 1 | 0.1 | -0.1 |\n| 0.7 | 1 | 0 | 0|\n\nWe end up much closer to the ground truth.\n\nConclusion: use a learning rate scheduler to decrease your learning rate over time.",
    "3217031": "Thank you @cpmpml  - for the explanation ... I didnt know earlier why the decreasing learning rate is working in this problem . \n\nAlso thank you for the nice ensemble implementation , it helped almost + 0.5 CV for me ..",
    "3217041": "Glad what I share helps! Thanks for the feedback.",
    "3218954": "Theoretically, warm-up can also be helpful, as it prevents erratic behavior when the optimizer first navigates sharp transitions or abrupt changes in the loss landscape, where gradients are non-smooth or even discontinuous. Additionally, gradient clipping can complement warm-up by helping to prevent instability during the early stages of training.",
    "3218968": "Sure, but warming up helping is not specific to MAE.",
    "3228317": "thanks for sharing @cpmpml it is someway hard to implement in spark due to limited naive support. any notebook suggestion would be perfect"
  },
  "source": "meta"
}