{
  "id": 488083,
  "title": "How To Tune Learning Schedule",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/488083",
  "author_name": "Chris Deotte",
  "post_date": "2024-04-01T02:59:06.038000",
  "votes": 181,
  "comment_count": 39,
  "views": 0,
  "content": "<p>Hi everyone. I was rereading my discussion posts and I saw many people asked about tuning learning schedules for NN. So, I thought I would provide some guidance. In this competition, everyone is exploring many models including 2D NN image models, 1D NN time series models, and ML models like GBDT.</p>\n<h1>Step 1 - Find Best Starting LR</h1>\n<p>The first thing I like to do is try constant learning rates for total of 5-10 epochs (or more if needed) to \"get a feel\" of learning. So pick LR = 1e-5, or 1e-4, or 1e-3, or 1e-2 (and maybe the midpoints too) and train with constant learning rate for 5-10 epochs (or more if validation loss never degrades).</p>\n<p>Observe what the best validation loss achieved is and observe at what epoch the validation loss stops improving and begins getting worse. Also observe which <code>LR = X</code> achieves the best validation loss (when using constant LR).</p>\n<h1>Step 2 - Find LR Steps</h1>\n<p>If the validation loss is something like <code>KL-Div = 0.40, 0.35, 0.30, 0.35, 0.37</code> (for epochs 1-5 respectively), then we observe that validation loss improved for epochs 1-3 and got worse in epochs 4 and 5+ (using <code>LR = X</code> for some <code>X</code>)</p>\n<p>So, next try a single step learning schedule. Use <code>LR = X</code> where <code>X</code> is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use <code>LR = X/10</code> for epochs 5-10. Then run this and watch.</p>\n<h1>Step 3 - Find LR Steps</h1>\n<p>We now pay attention to epochs <code>5-10</code> or whatever epochs correspond to your decreased <code>LR = X/10.</code>. For these epochs we see <code>KL-Div = 0.29, 0.28, 0.29, 0.30, 0.31, 0.32</code>. So we see that now epoch 5 and 6 improve compared to epochs 1-4 but epoch 6 gets worse. So we now set the <code>LR = X, X, X, X, X/10, X/10, X/10, X/100, X/100 X/100</code> (for epochs 1-10)</p>\n<p>In other words, we use <code>LR = X</code> for epochs 1-4 from our observation earlier. We use <code>LR = X/10</code> for epochs 5-7 (include 1 bad epoch), from our observations earlier. And we add <code>LR = X/100</code> for epochs 8-10. Then run this and watch.</p>\n<h1>Step 4 - Observe Results</h1>\n<p>By now, we see the pattern. We keep using constant learning rate with steps where needed and run. Then we see where the validation loss stops improving. Then we add another step using <code>LR_new = LR_old/10.</code> and we run again. In total, we will add 2 or 3 steps (or more until validation stops improving). So that our LR hits <code>X, X/10. X/100, and X/1000</code> or so. (Where <code>X</code> is our starting rate found in step 1 above).</p>\n<p>At this point, we have a well trained model. We can use this as our final model or we can explore cosine learning schedule</p>\n<h1>Step 5 - Convert to Cosine Schedule</h1>\n<p>If we wish to explore cosine learning schedule, then we can use Steps 1-4 above to see how many epochs we were able to keep improving the validation loss.</p>\n<p>Imagine that we we able to achieve validation loss = <code>KL-Div = 0.40, 0.35, 0.30, 0.35, 0.29, 0.28, 0.29, 0.27, 0.28, 0.28</code> when using 3 steps and 10 epochs with <code>LR = X, X, X,  X, X/10, X/10, X/10, X/100, X/100, X/100</code>. We observe that LR ranged from <code>X</code> to <code>X/100</code> from <code>epoch=1</code> to <code>epoch=8</code> (where validation loss was improving).</p>\n<p>Therefore we can now try cosine learning rate with <code>begin : LR = X</code> and <code>end : LR = X/100.</code> for 8 epochs. Often using cosine learning rate will perform better than using steps above but not always.</p>\n<p>After trying <code>start = X, end = X/100, epochs = 8</code>, we can also try <code>epochs = 7 and 9</code> because some times using cosine learning rate requires slightly different epochs than constant.</p>\n<h1>Conclusion</h1>\n<p>In conclusion we saw how to tune a step wise learning rate. In step 1 we find the optimal starting learning rate. Then in steps 2-4 we find where to add the steps.</p>\n<p>In step 5, we explore converting the step wise learning rate into a cosine learning rate for possible improvement.</p>",
  "messages": [
    {
      "id": 2726171,
      "postDate": "2024-04-01T02:59:06.037Z",
      "content": "<p>Hi everyone. I was rereading my discussion posts and I saw many people asked about tuning learning schedules for NN. So, I thought I would provide some guidance. In this competition, everyone is exploring many models including 2D NN image models, 1D NN time series models, and ML models like GBDT.</p>\n<h1>Step 1 - Find Best Starting LR</h1>\n<p>The first thing I like to do is try constant learning rates for total of 5-10 epochs (or more if needed) to \"get a feel\" of learning. So pick LR = 1e-5, or 1e-4, or 1e-3, or 1e-2 (and maybe the midpoints too) and train with constant learning rate for 5-10 epochs (or more if validation loss never degrades).</p>\n<p>Observe what the best validation loss achieved is and observe at what epoch the validation loss stops improving and begins getting worse. Also observe which <code>LR = X</code> achieves the best validation loss (when using constant LR).</p>\n<h1>Step 2 - Find LR Steps</h1>\n<p>If the validation loss is something like <code>KL-Div = 0.40, 0.35, 0.30, 0.35, 0.37</code> (for epochs 1-5 respectively), then we observe that validation loss improved for epochs 1-3 and got worse in epochs 4 and 5+ (using <code>LR = X</code> for some <code>X</code>)</p>\n<p>So, next try a single step learning schedule. Use <code>LR = X</code> where <code>X</code> is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use <code>LR = X/10</code> for epochs 5-10. Then run this and watch.</p>\n<h1>Step 3 - Find LR Steps</h1>\n<p>We now pay attention to epochs <code>5-10</code> or whatever epochs correspond to your decreased <code>LR = X/10.</code>. For these epochs we see <code>KL-Div = 0.29, 0.28, 0.29, 0.30, 0.31, 0.32</code>. So we see that now epoch 5 and 6 improve compared to epochs 1-4 but epoch 6 gets worse. So we now set the <code>LR = X, X, X, X, X/10, X/10, X/10, X/100, X/100 X/100</code> (for epochs 1-10)</p>\n<p>In other words, we use <code>LR = X</code> for epochs 1-4 from our observation earlier. We use <code>LR = X/10</code> for epochs 5-7 (include 1 bad epoch), from our observations earlier. And we add <code>LR = X/100</code> for epochs 8-10. Then run this and watch.</p>\n<h1>Step 4 - Observe Results</h1>\n<p>By now, we see the pattern. We keep using constant learning rate with steps where needed and run. Then we see where the validation loss stops improving. Then we add another step using <code>LR_new = LR_old/10.</code> and we run again. In total, we will add 2 or 3 steps (or more until validation stops improving). So that our LR hits <code>X, X/10. X/100, and X/1000</code> or so. (Where <code>X</code> is our starting rate found in step 1 above).</p>\n<p>At this point, we have a well trained model. We can use this as our final model or we can explore cosine learning schedule</p>\n<h1>Step 5 - Convert to Cosine Schedule</h1>\n<p>If we wish to explore cosine learning schedule, then we can use Steps 1-4 above to see how many epochs we were able to keep improving the validation loss.</p>\n<p>Imagine that we we able to achieve validation loss = <code>KL-Div = 0.40, 0.35, 0.30, 0.35, 0.29, 0.28, 0.29, 0.27, 0.28, 0.28</code> when using 3 steps and 10 epochs with <code>LR = X, X, X,  X, X/10, X/10, X/10, X/100, X/100, X/100</code>. We observe that LR ranged from <code>X</code> to <code>X/100</code> from <code>epoch=1</code> to <code>epoch=8</code> (where validation loss was improving).</p>\n<p>Therefore we can now try cosine learning rate with <code>begin : LR = X</code> and <code>end : LR = X/100.</code> for 8 epochs. Often using cosine learning rate will perform better than using steps above but not always.</p>\n<p>After trying <code>start = X, end = X/100, epochs = 8</code>, we can also try <code>epochs = 7 and 9</code> because some times using cosine learning rate requires slightly different epochs than constant.</p>\n<h1>Conclusion</h1>\n<p>In conclusion we saw how to tune a step wise learning rate. In step 1 we find the optimal starting learning rate. Then in steps 2-4 we find where to add the steps.</p>\n<p>In step 5, we explore converting the step wise learning rate into a cosine learning rate for possible improvement.</p>",
      "rawMarkdown": "Hi everyone. I was rereading my discussion posts and I saw many people asked about tuning learning schedules for NN. So, I thought I would provide some guidance. In this competition, everyone is exploring many models including 2D NN image models, 1D NN time series models, and ML models like GBDT.\n\n# Step 1 - Find Best Starting LR\nThe first thing I like to do is try constant learning rates for total of 5-10 epochs (or more if needed) to \"get a feel\" of learning. So pick LR = 1e-5, or 1e-4, or 1e-3, or 1e-2 (and maybe the midpoints too) and train with constant learning rate for 5-10 epochs (or more if validation loss never degrades).\n\nObserve what the best validation loss achieved is and observe at what epoch the validation loss stops improving and begins getting worse. Also observe which `LR = X` achieves the best validation loss (when using constant LR).\n\n# Step 2 - Find LR Steps\nIf the validation loss is something like `KL-Div = 0.40, 0.35, 0.30, 0.35, 0.37` (for epochs 1-5 respectively), then we observe that validation loss improved for epochs 1-3 and got worse in epochs 4 and 5+ (using `LR = X` for some `X`)\n\nSo, next try a single step learning schedule. Use `LR = X` where `X` is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use `LR = X/10` for epochs 5-10. Then run this and watch.\n\n# Step 3 - Find LR Steps\nWe now pay attention to epochs `5-10` or whatever epochs correspond to your decreased `LR = X/10.`. For these epochs we see `KL-Div = 0.29, 0.28, 0.29, 0.30, 0.31, 0.32`. So we see that now epoch 5 and 6 improve compared to epochs 1-4 but epoch 6 gets worse. So we now set the `LR = X, X, X, X, X/10, X/10, X/10, X/100, X/100 X/100` (for epochs 1-10)\n\nIn other words, we use `LR = X` for epochs 1-4 from our observation earlier. We use `LR = X/10` for epochs 5-7 (include 1 bad epoch), from our observations earlier. And we add `LR = X/100` for epochs 8-10. Then run this and watch.\n\n# Step 4 - Observe Results\nBy now, we see the pattern. We keep using constant learning rate with steps where needed and run. Then we see where the validation loss stops improving. Then we add another step using `LR_new = LR_old/10.` and we run again. In total, we will add 2 or 3 steps (or more until validation stops improving). So that our LR hits `X, X/10. X/100, and X/1000` or so. (Where `X` is our starting rate found in step 1 above).\n\nAt this point, we have a well trained model. We can use this as our final model or we can explore cosine learning schedule\n\n# Step 5 - Convert to Cosine Schedule\nIf we wish to explore cosine learning schedule, then we can use Steps 1-4 above to see how many epochs we were able to keep improving the validation loss.\n\nImagine that we we able to achieve validation loss = `KL-Div = 0.40, 0.35, 0.30, 0.35, 0.29, 0.28, 0.29, 0.27, 0.28, 0.28` when using 3 steps and 10 epochs with `LR = X, X, X,  X, X/10, X/10, X/10, X/100, X/100, X/100`. We observe that LR ranged from `X` to `X/100` from `epoch=1` to `epoch=8` (where validation loss was improving).\n\nTherefore we can now try cosine learning rate with `begin : LR = X` and `end : LR = X/100.` for 8 epochs. Often using cosine learning rate will perform better than using steps above but not always.\n\nAfter trying `start = X, end = X/100, epochs = 8`, we can also try `epochs = 7 and 9` because some times using cosine learning rate requires slightly different epochs than constant.\n\n# Conclusion\nIn conclusion we saw how to tune a step wise learning rate. In step 1 we find the optimal starting learning rate. Then in steps 2-4 we find where to add the steps.\n\nIn step 5, we explore converting the step wise learning rate into a cosine learning rate for possible improvement.",
      "votes": 181
    },
    {
      "id": 2726290,
      "postDate": "2024-04-01T05:10:19.950Z",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for this amazing sharing!</p>\n<blockquote>\n  <p>So, next try a single step learning schedule. Use LR = X where X is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use LR = X/10 for epochs 5-10. Then run this and watch.</p>\n</blockquote>\n<p>Can I have a curious question why do we need to include one bad epoch when decreasing the LR?</p>",
      "rawMarkdown": "Hello Chris,\n\nThanks for this amazing sharing!\n>So, next try a single step learning schedule. Use LR = X where X is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use LR = X/10 for epochs 5-10. Then run this and watch.\n\nCan I have a curious question why do we need to include one bad epoch when decreasing the LR?",
      "votes": 7,
      "replies": [
        {
          "id": 2726295,
          "postDate": "2024-04-01T05:13:42.440Z",
          "content": "<p>We can experiment with</p>\n<ul>\n<li>include 1 bad epoch</li>\n<li>dont include 1 bad epoch</li>\n</ul>\n<p>and see which is better. In general, i think \"including 1 bad epoch\" makes the validation loss better when the drop occurs (but i could be wrong).</p>",
          "rawMarkdown": "We can experiment with\n* include 1 bad epoch\n* dont include 1 bad epoch\n\nand see which is better. In general, i think \"including 1 bad epoch\" makes the validation loss better when the drop occurs (but i could be wrong).",
          "votes": 3,
          "replies": [
            {
              "id": 2726314,
              "postDate": "2024-04-01T05:39:31.777Z",
              "content": "<p>Thanks for the answer! I'll do an experiment on that.🫡</p>",
              "rawMarkdown": "Thanks for the answer! I'll do an experiment on that.🫡",
              "votes": 3
            }
          ]
        }
      ]
    },
    {
      "id": 2726197,
      "postDate": "2024-04-01T03:36:45.580Z",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for sharing this guidance. I'll try in future experiments!<br>\nI would like to ask a question about <code>batch_size</code>. For my local experiments, I observe <code>batch_size</code> matters a lot.<br>\nFor example, <code>batch_size=16</code> and <code>batch_size=32</code> can have CV score difference within the range of 0.02 ~ 0.05. Do you fix a commonly chosen <code>batch_size</code> during lr and epoch tuning, then switch to <code>batch_size</code> tuning after lr and epoch are well-tuned? Thanks a lot.</p>",
      "rawMarkdown": "Hello Chris,\n\nThanks for sharing this guidance. I'll try in future experiments!\nI would like to ask a question about `batch_size`. For my local experiments, I observe `batch_size` matters a lot.\nFor example, `batch_size=16` and `batch_size=32` can have CV score difference within the range of 0.02 ~ 0.05. Do you fix a commonly chosen `batch_size` during lr and epoch tuning, then switch to `batch_size` tuning after lr and epoch are well-tuned? Thanks a lot.",
      "votes": 7,
      "replies": [
        {
          "id": 2726201,
          "postDate": "2024-04-01T03:43:33.633Z",
          "content": "<p>Great question. I do explore batch size. In some competitions, batch size is the magic for large improvements.</p>\n<p>I will usually start with <code>batch_size = 32</code> and follow the steps in my discussion post. After tuning for <code>batch_size = 32</code>, i will use the same schedule and try <code>batch_size_new = 32 * R</code> and <code>LR_new = LR_Schedule_old * R</code>, to explore <code>batch_size = 8, 16, 64, 128</code>. (Note that some people suggest <code>B_new = 32*R</code> and <code>L_new = L*sqrt(R)</code>, i'm not sure what is best).</p>\n<p>Sometimes we are surprised to see a boost with batch size different than 32. Unfortunately in this competition I have not seen a significant difference with batch sizes other than 32.</p>",
          "rawMarkdown": "Great question. I do explore batch size. In some competitions, batch size is the magic for large improvements.\n\nI will usually start with `batch_size = 32` and follow the steps in my discussion post. After tuning for `batch_size = 32`, i will use the same schedule and try `batch_size_new = 32 * R` and `LR_new = LR_Schedule_old * R`, to explore `batch_size = 8, 16, 64, 128`. (Note that some people suggest `B_new = 32*R` and `L_new = L*sqrt(R)`, i'm not sure what is best).\n\nSometimes we are surprised to see a boost with batch size different than 32. Unfortunately in this competition I have not seen a significant difference with batch sizes other than 32.",
          "votes": 7,
          "replies": [
            {
              "id": 2726218,
              "postDate": "2024-04-01T03:52:55.453Z",
              "content": "<p>Hi Chris, </p>\n<p>Thanks for sharing your secret sauce as always. I'll try it out in my experiments. <br>\nHope you can stay on top in the private leaderboard, good luck!</p>",
              "rawMarkdown": "Hi Chris, \n\nThanks for sharing your secret sauce as always. I'll try it out in my experiments. \nHope you can stay on top in the private leaderboard, good luck!",
              "votes": 1
            },
            {
              "id": 2726239,
              "postDate": "2024-04-01T04:29:17.900Z",
              "content": "<p>I think the \"L_new = L * R\" part works well with SGD. <br>\nWhile \"L_new = L*sqrt(R)\" works well with ADAM.</p>\n<p>Reference:<br>\n<a href=\"https://www.cs.princeton.edu/~smalladi/blog/2024/01/22/SDEs-ScalingRules/\" target=\"_blank\">https://www.cs.princeton.edu/~smalladi/blog/2024/01/22/SDEs-ScalingRules/</a></p>",
              "rawMarkdown": "I think the \"L_new = L * R\" part works well with SGD. \nWhile \"L_new = L*sqrt(R)\" works well with ADAM.\n\nReference:\nhttps://www.cs.princeton.edu/~smalladi/blog/2024/01/22/SDEs-ScalingRules/",
              "votes": 9
            },
            {
              "id": 2733856,
              "postDate": "2024-04-03T20:32:08.190Z",
              "content": "<p>Does the <code>R</code> variable above stand for what ever scaling factor you're adjusting your current batch size by? As in if you want to change your batch size from 32 to 64, you're <code>R</code> factor is 2 and you should multiple you're old LR by 2?</p>",
              "rawMarkdown": "Does the `R` variable above stand for what ever scaling factor you're adjusting your current batch size by? As in if you want to change your batch size from 32 to 64, you're `R` factor is 2 and you should multiple you're old LR by 2?",
              "votes": 1
            },
            {
              "id": 2733861,
              "postDate": "2024-04-03T20:34:52.553Z",
              "content": "<p>So can we think of our parameter search for finding the ideal learning rate actually as finding the ideal ratio between batch size and learning rate? Once we've discovered that ratio, we can then search for the ideal batch size while maintaining the learning rate ratio?</p>",
              "rawMarkdown": "So can we think of our parameter search for finding the ideal learning rate actually as finding the ideal ratio between batch size and learning rate? Once we've discovered that ratio, we can then search for the ideal batch size while maintaining the learning rate ratio?"
            },
            {
              "id": 2733865,
              "postDate": "2024-04-03T20:37:03.067Z",
              "content": "<blockquote>\n  <p>Does the R variable above stand for what ever scaling factor you're adjusting your current batch size by? As in if you want to change your batch size from 32 to 64, you're R factor is 2 and you should multiple you're old LR by 2?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>So can we think of our parameter search for finding the ideal learning rate actually as finding the ideal ratio between batch size and learning rate? Once we've discovered that ratio, we can then search for the ideal batch size while maintaining the learning rate ratio?</p>\n</blockquote>\n<p>Yes and No. First we need to find the correct learning rate for batch size 32. Afterward we can apply ratio <code>R</code> to explore different batch sizes with different learning rates.</p>",
              "rawMarkdown": ">Does the R variable above stand for what ever scaling factor you're adjusting your current batch size by? As in if you want to change your batch size from 32 to 64, you're R factor is 2 and you should multiple you're old LR by 2?\n\nYes.\n\n>So can we think of our parameter search for finding the ideal learning rate actually as finding the ideal ratio between batch size and learning rate? Once we've discovered that ratio, we can then search for the ideal batch size while maintaining the learning rate ratio?\n\nYes and No. First we need to find the correct learning rate for batch size 32. Afterward we can apply ratio `R` to explore different batch sizes with different learning rates.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2728838,
      "postDate": "2024-04-02T12:32:17.740Z",
      "content": "<p>I usually use Adan optimizer with Onecycle lr, decreased LR = X/10 or increased LR=10*X.</p>",
      "rawMarkdown": "I usually use Adan optimizer with Onecycle lr, decreased LR = X/10 or increased LR=10*X.",
      "votes": 6,
      "replies": [
        {
          "id": 2730349,
          "postDate": "2024-04-02T16:58:09.157Z",
          "content": "<p>Good idea. When <code>number of epochs &gt; 20</code> or so then I use OneCycleLR (i.e. cosine) and search for best starting LR and best number of epochs (and ignore step seach). When <code>number of epochs &lt; 20</code> or so then I explore StepLR as described in my post above (and try converting to OneCycleLR afterward).</p>",
          "rawMarkdown": "Good idea. When `number of epochs > 20` or so then I use OneCycleLR (i.e. cosine) and search for best starting LR and best number of epochs (and ignore step seach). When `number of epochs < 20` or so then I explore StepLR as described in my post above (and try converting to OneCycleLR afterward).",
          "votes": 9
        }
      ]
    },
    {
      "id": 2733871,
      "postDate": "2024-04-03T20:39:09.473Z",
      "content": "<p>Would you conduct a new LR and batch size parameter search for each experiment you run to validate new data augmentation techniques? I was trying out different ways to augment my dataset but found that even simple tricks that seemed to work in the literature weren't improving my score. I was wondering if it was because I wasn't finding the ideal parameters for each data augmentation experiment.</p>\n<p>In general, I was trying to figure out some sort of pattern for when I should conduct a new LR/batch size search. How do you approach this?</p>",
      "rawMarkdown": "Would you conduct a new LR and batch size parameter search for each experiment you run to validate new data augmentation techniques? I was trying out different ways to augment my dataset but found that even simple tricks that seemed to work in the literature weren't improving my score. I was wondering if it was because I wasn't finding the ideal parameters for each data augmentation experiment.\n\nIn general, I was trying to figure out some sort of pattern for when I should conduct a new LR/batch size search. How do you approach this?",
      "votes": 3,
      "replies": [
        {
          "id": 2735601,
          "postDate": "2024-04-04T19:37:04.533Z",
          "content": "<p>In general, I do not tune the LR and batch size for each augmentation experiment.</p>\n<p>When an augmentation is beneficial, it will usually improve the baseline using the <strong>same</strong> LR and batch size as before. After finding successful augmentation, it is then sometimes beneficial to run more epochs with same starting LR because the augmentation allows us to train longer (without overfitting and achieve an even better CV score than experiment demonstrated).</p>",
          "rawMarkdown": "In general, I do not tune the LR and batch size for each augmentation experiment.\n\nWhen an augmentation is beneficial, it will usually improve the baseline using the **same** LR and batch size as before. After finding successful augmentation, it is then sometimes beneficial to run more epochs with same starting LR because the augmentation allows us to train longer (without overfitting and achieve an even better CV score than experiment demonstrated).",
          "votes": 3
        }
      ]
    },
    {
      "id": 2727164,
      "postDate": "2024-04-01T16:19:00.127Z",
      "content": "<p>Thank you for all your help.  What is wrong with Adam or AdamW?  None of the classes I took talked about LR schedules.</p>",
      "rawMarkdown": "Thank you for all your help.  What is wrong with Adam or AdamW?  None of the classes I took talked about LR schedules.",
      "votes": 3,
      "replies": [
        {
          "id": 2727173,
          "postDate": "2024-04-01T16:24:56.753Z",
          "content": "<p>Yes Adam and AdamW will adjust the learning rate internally to some extent for us (compared to other optimizers like SGD which don't). None-the-less, we still benefit from choosing an optimal starting learning rate. And we benefit by adding steps and/or decrease from cosine schedule (when validation loss stops improving).</p>",
          "rawMarkdown": "Yes Adam and AdamW will adjust the learning rate internally to some extent for us (compared to other optimizers like SGD which don't). None-the-less, we still benefit from choosing an optimal starting learning rate. And we benefit by adding steps and/or decrease from cosine schedule (when validation loss stops improving).",
          "votes": 5
        }
      ]
    },
    {
      "id": 2742517,
      "postDate": "2024-04-09T00:33:08.250Z",
      "content": "<p>i never tried this, but meta just release schedule free optimizer</p>\n<p><a href=\"https://github.com/facebookresearch/schedule_free\" target=\"_blank\">https://github.com/facebookresearch/schedule_free</a></p>",
      "rawMarkdown": "i never tried this, but meta just release schedule free optimizer\n\nhttps://github.com/facebookresearch/schedule_free",
      "votes": 4
    },
    {
      "id": 2727175,
      "postDate": "2024-04-01T16:26:16.950Z",
      "content": "<p>Hi Chris, <br>\nThanks for sharing great discussion.<br>\nI have a question about the balance of <code>batch_size</code>, <code>lr</code> and <code>number_of_epochs</code>.</p>\n<p>In what order should we tune<code>batch_size</code>, <code>lr</code> and <code>number_of_epochs</code> ? <br>\nThe score may depend on the balance of <code>batch_size</code>, <code>lr</code> and <code>number_of_epochs</code>.<br>\nI think the balance should be tuned to maximize the score and speed of improving cycle, but honestly I don' know about it. </p>",
      "rawMarkdown": "Hi Chris, \nThanks for sharing great discussion.\nI have a question about the balance of `batch_size`, `lr` and `number_of_epochs`.\n\nIn what order should we tune`batch_size`, `lr` and `number_of_epochs` ? \nThe score may depend on the balance of `batch_size`, `lr` and `number_of_epochs`.\nI think the balance should be tuned to maximize the score and speed of improving cycle, but honestly I don' know about it. ",
      "votes": 4,
      "replies": [
        {
          "id": 2727215,
          "postDate": "2024-04-01T16:45:10.160Z",
          "content": "<p><strong>Short Answer</strong>: I tune B, LR, E quickly and approximately at first. Then run lots of experiments exploring other model related stuff. Then finally I tune B, LR, E more carefully later with my best discovered model pipeline.</p>\n<p><strong>Long Answer</strong>: Personally, I begin with <code>batch_size = 32</code> and then tune <code>starting LR</code> first. Once I find a good starting LR, I then tune <code>number_of_epochs</code> second (with steps or cosine added). And finally I explore <code>batch size</code> third using the formula <code>new batch = old batch * scale</code> and <code>new LR = old LR * scale</code> OR <code>new LR = old LR * sqrt(scale)</code>. </p>\n<p>Some people prefer to first use <code>batch size = Largest Possible</code> where we increase batch size using mixed precision (i.e. AMP in torch) and multiple GPUs and utilize maximum GPU VRAM. By using largest batch size possible first, all experiments to identify LR and epochs are sped up. Afterward, we can explore smaller batch sizes (with scaling formulas above).</p>\n<p>Also note, during the experimentation of exploring  the preprocess, data augmentation, model architecture, etc etc. We often use less <code>number of epochs</code> and possibly <code>larger batch sizes</code> for increased speed of experiments. Then after we find optimal preprocess, data augmentation, model architecture, etc etc we then tune the <code>LR</code>, <code>number of epochs</code>, and <code>batch size</code> and train long and hard for maximum LB score and CV score!</p>",
          "rawMarkdown": "**Short Answer**: I tune B, LR, E quickly and approximately at first. Then run lots of experiments exploring other model related stuff. Then finally I tune B, LR, E more carefully later with my best discovered model pipeline.\n\n**Long Answer**: Personally, I begin with `batch_size = 32` and then tune `starting LR` first. Once I find a good starting LR, I then tune `number_of_epochs` second (with steps or cosine added). And finally I explore `batch size` third using the formula `new batch = old batch * scale` and `new LR = old LR * scale` OR `new LR = old LR * sqrt(scale)`. \n\nSome people prefer to first use `batch size = Largest Possible` where we increase batch size using mixed precision (i.e. AMP in torch) and multiple GPUs and utilize maximum GPU VRAM. By using largest batch size possible first, all experiments to identify LR and epochs are sped up. Afterward, we can explore smaller batch sizes (with scaling formulas above).\n\nAlso note, during the experimentation of exploring  the preprocess, data augmentation, model architecture, etc etc. We often use less `number of epochs` and possibly `larger batch sizes` for increased speed of experiments. Then after we find optimal preprocess, data augmentation, model architecture, etc etc we then tune the `LR`, `number of epochs`, and `batch size` and train long and hard for maximum LB score and CV score!",
          "votes": 17
        }
      ]
    },
    {
      "id": 2726431,
      "postDate": "2024-04-01T07:39:34.613Z",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for this amazing sharing!</p>\n<p>May I ask how to tune lr when using kfold? Because each fold may perform different with others.</p>",
      "rawMarkdown": "Hello Chris,\n\nThanks for this amazing sharing!\n\nMay I ask how to tune lr when using kfold? Because each fold may perform different with others.",
      "votes": 4,
      "replies": [
        {
          "id": 2726522,
          "postDate": "2024-04-01T08:46:53.333Z",
          "content": "<p>Hi. Great question. I look at all the folds and take the best approximation (which works for the most folds). (Or if I'm in a rush, I run the first N and look at the first N out of K folds where N&lt;K). Each experiment, I use the same step schedule (or cosine schedule) for every epoch. For later experiments, I will let all K folds finish and earlier experiments I may stop early with only N out of K folds if I see a pattern.</p>",
          "rawMarkdown": "Hi. Great question. I look at all the folds and take the best approximation (which works for the most folds). (Or if I'm in a rush, I run the first N and look at the first N out of K folds where N<K). Each experiment, I use the same step schedule (or cosine schedule) for every epoch. For later experiments, I will let all K folds finish and earlier experiments I may stop early with only N out of K folds if I see a pattern.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2742878,
      "postDate": "2024-04-09T06:01:00.150Z",
      "content": "<p>Thank you for sharing your ideas! I learned a lot from you through this competition!!</p>",
      "rawMarkdown": "Thank you for sharing your ideas! I learned a lot from you through this competition!!",
      "votes": 1
    },
    {
      "id": 2741155,
      "postDate": "2024-04-08T07:10:02.743Z",
      "content": "<p>Hi Chris, What about warmup steps, What's the best way to tune num_warmup_steps required??</p>",
      "rawMarkdown": "Hi Chris, What about warmup steps, What's the best way to tune num_warmup_steps required??",
      "votes": 1,
      "replies": [
        {
          "id": 2741285,
          "postDate": "2024-04-08T08:19:14.903Z",
          "content": "<p>I'm not sure. Most times, I do not use warm up steps. However sometimes I have added them and they helped.</p>",
          "rawMarkdown": "I'm not sure. Most times, I do not use warm up steps. However sometimes I have added them and they helped.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2738889,
      "postDate": "2024-04-06T17:24:26.810Z",
      "content": "<p>Thank you for sharing Chris!</p>\n<p>I was wondering if you guys use libraries like ray-tune to finetune your models? I'm not sure if it'd be more efficient to tun the parameters manually by intuition or use algorithms to do hyperparameter search. Do you have any advice regarding this? </p>\n<p>Thanks!</p>",
      "rawMarkdown": "Thank you for sharing Chris!\n\nI was wondering if you guys use libraries like ray-tune to finetune your models? I'm not sure if it'd be more efficient to tun the parameters manually by intuition or use algorithms to do hyperparameter search. Do you have any advice regarding this? \n\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2739102,
          "postDate": "2024-04-06T20:15:10.110Z",
          "content": "<p>I do everything manually. I think there are some helpful software for hyperparameter search but i don't have experience with it.</p>",
          "rawMarkdown": "I do everything manually. I think there are some helpful software for hyperparameter search but i don't have experience with it.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2735790,
      "postDate": "2024-04-04T22:40:18.287Z",
      "content": "<p>Thank you for sharing Chris.</p>\n<p>I wonder what is the explanation behind including 1 bad epoch! Maybe pushing the model to as much as possible before lowering the learning rate.</p>",
      "rawMarkdown": "Thank you for sharing Chris.\n\nI wonder what is the explanation behind including 1 bad epoch! Maybe pushing the model to as much as possible before lowering the learning rate.",
      "votes": 1
    },
    {
      "id": 2733284,
      "postDate": "2024-04-03T15:10:47.837Z",
      "content": "<p>Thank you for this brilliant share, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Bookmarking this for future use</p>",
      "rawMarkdown": "Thank you for this brilliant share, @cdeotte. Bookmarking this for future use",
      "votes": 1
    },
    {
      "id": 2732579,
      "postDate": "2024-04-03T08:00:06.850Z",
      "content": "<p>thanks for sharing, I tried a larger batch size on your open-source solution to speed up the training process, but the scores were lower. Even after adjusting the learning rate, I didn't get better results. I think I will try combining your shared insights in my experiments. Thanks again for sharing.🥳🥳🥳🥳🥳🥳🥳</p>",
      "rawMarkdown": "thanks for sharing, I tried a larger batch size on your open-source solution to speed up the training process, but the scores were lower. Even after adjusting the learning rate, I didn't get better results. I think I will try combining your shared insights in my experiments. Thanks again for sharing.🥳🥳🥳🥳🥳🥳🥳",
      "votes": 1
    },
    {
      "id": 2731461,
      "postDate": "2024-04-02T18:40:53.003Z",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for this amazing sharing!</p>",
      "rawMarkdown": "Hello Chris,\n\nThanks for this amazing sharing!",
      "votes": 1
    },
    {
      "id": 2728753,
      "postDate": "2024-04-02T11:38:21.937Z",
      "content": "<p>Great explanation! Interesting approach to Number of epochs! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Great explanation! Interesting approach to Number of epochs! @cdeotte ",
      "votes": 1
    },
    {
      "id": 2739240,
      "postDate": "2024-04-06T22:33:52.617Z",
      "content": "<p>This is really helpful Chris, cheers!</p>\n<p>This may be a silly question but if your model takes a long time to train even one epoch (and it does not appear to be overfitting), is it worth tuning your hyperparameters on a subset of your data first to speed up the iteration process?</p>",
      "rawMarkdown": "This is really helpful Chris, cheers!\n\nThis may be a silly question but if your model takes a long time to train even one epoch (and it does not appear to be overfitting), is it worth tuning your hyperparameters on a subset of your data first to speed up the iteration process?",
      "replies": [
        {
          "id": 2739267,
          "postDate": "2024-04-06T23:12:36.553Z",
          "content": "<p>I'm thinking it may not be useful in getting the right schedule but at least might be useful in setting the initial lr.</p>",
          "rawMarkdown": "I'm thinking it may not be useful in getting the right schedule but at least might be useful in setting the initial lr."
        }
      ]
    },
    {
      "id": 2742866,
      "postDate": "2024-04-09T05:51:44.527Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2726349,
      "postDate": "2024-04-01T06:23:00.863Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2742966,
      "postDate": "2024-04-09T06:50:10.493Z",
      "content": "<p>Thanks Chris. Very helpful!</p>",
      "rawMarkdown": "Thanks Chris. Very helpful!",
      "votes": 1
    },
    {
      "id": 2742705,
      "postDate": "2024-04-09T03:16:02.483Z",
      "content": "<p>thank you for sharing chris</p>",
      "rawMarkdown": "thank you for sharing chris",
      "votes": 1
    },
    {
      "id": 2735583,
      "postDate": "2024-04-04T19:25:53.507Z",
      "content": "<p>thanks, it`s very interesting</p>",
      "rawMarkdown": "thanks, it`s very interesting",
      "votes": 1
    },
    {
      "id": 2726351,
      "postDate": "2024-04-01T06:23:53.603Z",
      "content": "<p>You are my ML God and Goal, thanks!</p>",
      "rawMarkdown": "You are my ML God and Goal, thanks!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2726290,
      "author_name": "LHK",
      "author_url": "",
      "post_date": "2024-04-01T05:10:19.950000",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for this amazing sharing!</p>\n<blockquote>\n  <p>So, next try a single step learning schedule. Use LR = X where X is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use LR = X/10 for epochs 5-10. Then run this and watch.</p>\n</blockquote>\n<p>Can I have a curious question why do we need to include one bad epoch when decreasing the LR?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2726295,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-01T05:13:42.440000",
          "content": "<p>We can experiment with</p>\n<ul>\n<li>include 1 bad epoch</li>\n<li>dont include 1 bad epoch</li>\n</ul>\n<p>and see which is better. In general, i think \"including 1 bad epoch\" makes the validation loss better when the drop occurs (but i could be wrong).</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2726314,
              "author_name": "LHK",
              "author_url": "",
              "post_date": "2024-04-01T05:39:31.777000",
              "content": "<p>Thanks for the answer! I'll do an experiment on that.🫡</p>",
              "votes": 3,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2726197,
      "author_name": "AbaoJiang",
      "author_url": "",
      "post_date": "2024-04-01T03:36:45.580000",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for sharing this guidance. I'll try in future experiments!<br>\nI would like to ask a question about <code>batch_size</code>. For my local experiments, I observe <code>batch_size</code> matters a lot.<br>\nFor example, <code>batch_size=16</code> and <code>batch_size=32</code> can have CV score difference within the range of 0.02 ~ 0.05. Do you fix a commonly chosen <code>batch_size</code> during lr and epoch tuning, then switch to <code>batch_size</code> tuning after lr and epoch are well-tuned? Thanks a lot.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2726201,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-01T03:43:33.633000",
          "content": "<p>Great question. I do explore batch size. In some competitions, batch size is the magic for large improvements.</p>\n<p>I will usually start with <code>batch_size = 32</code> and follow the steps in my discussion post. After tuning for <code>batch_size = 32</code>, i will use the same schedule and try <code>batch_size_new = 32 * R</code> and <code>LR_new = LR_Schedule_old * R</code>, to explore <code>batch_size = 8, 16, 64, 128</code>. (Note that some people suggest <code>B_new = 32*R</code> and <code>L_new = L*sqrt(R)</code>, i'm not sure what is best).</p>\n<p>Sometimes we are surprised to see a boost with batch size different than 32. Unfortunately in this competition I have not seen a significant difference with batch sizes other than 32.</p>",
          "votes": 7,
          "replies": [
            {
              "id": 2726218,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2024-04-01T03:52:55.453000",
              "content": "<p>Hi Chris, </p>\n<p>Thanks for sharing your secret sauce as always. I'll try it out in my experiments. <br>\nHope you can stay on top in the private leaderboard, good luck!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2726239,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2024-04-01T04:29:17.900000",
              "content": "<p>I think the \"L_new = L * R\" part works well with SGD. <br>\nWhile \"L_new = L*sqrt(R)\" works well with ADAM.</p>\n<p>Reference:<br>\n<a href=\"https://www.cs.princeton.edu/~smalladi/blog/2024/01/22/SDEs-ScalingRules/\" target=\"_blank\">https://www.cs.princeton.edu/~smalladi/blog/2024/01/22/SDEs-ScalingRules/</a></p>",
              "votes": 9,
              "replies": []
            },
            {
              "id": 2733856,
              "author_name": "megalodon",
              "author_url": "",
              "post_date": "2024-04-03T20:32:08.190000",
              "content": "<p>Does the <code>R</code> variable above stand for what ever scaling factor you're adjusting your current batch size by? As in if you want to change your batch size from 32 to 64, you're <code>R</code> factor is 2 and you should multiple you're old LR by 2?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2733861,
              "author_name": "megalodon",
              "author_url": "",
              "post_date": "2024-04-03T20:34:52.553000",
              "content": "<p>So can we think of our parameter search for finding the ideal learning rate actually as finding the ideal ratio between batch size and learning rate? Once we've discovered that ratio, we can then search for the ideal batch size while maintaining the learning rate ratio?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2733865,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-04-03T20:37:03.067000",
              "content": "<blockquote>\n  <p>Does the R variable above stand for what ever scaling factor you're adjusting your current batch size by? As in if you want to change your batch size from 32 to 64, you're R factor is 2 and you should multiple you're old LR by 2?</p>\n</blockquote>\n<p>Yes.</p>\n<blockquote>\n  <p>So can we think of our parameter search for finding the ideal learning rate actually as finding the ideal ratio between batch size and learning rate? Once we've discovered that ratio, we can then search for the ideal batch size while maintaining the learning rate ratio?</p>\n</blockquote>\n<p>Yes and No. First we need to find the correct learning rate for batch size 32. Afterward we can apply ratio <code>R</code> to explore different batch sizes with different learning rates.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2728838,
      "author_name": "Quan Vu",
      "author_url": "",
      "post_date": "2024-04-02T12:32:17.740000",
      "content": "<p>I usually use Adan optimizer with Onecycle lr, decreased LR = X/10 or increased LR=10*X.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 2730349,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-02T16:58:09.157000",
          "content": "<p>Good idea. When <code>number of epochs &gt; 20</code> or so then I use OneCycleLR (i.e. cosine) and search for best starting LR and best number of epochs (and ignore step seach). When <code>number of epochs &lt; 20</code> or so then I explore StepLR as described in my post above (and try converting to OneCycleLR afterward).</p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 2733871,
      "author_name": "megalodon",
      "author_url": "",
      "post_date": "2024-04-03T20:39:09.473000",
      "content": "<p>Would you conduct a new LR and batch size parameter search for each experiment you run to validate new data augmentation techniques? I was trying out different ways to augment my dataset but found that even simple tricks that seemed to work in the literature weren't improving my score. I was wondering if it was because I wasn't finding the ideal parameters for each data augmentation experiment.</p>\n<p>In general, I was trying to figure out some sort of pattern for when I should conduct a new LR/batch size search. How do you approach this?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2735601,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-04T19:37:04.533000",
          "content": "<p>In general, I do not tune the LR and batch size for each augmentation experiment.</p>\n<p>When an augmentation is beneficial, it will usually improve the baseline using the <strong>same</strong> LR and batch size as before. After finding successful augmentation, it is then sometimes beneficial to run more epochs with same starting LR because the augmentation allows us to train longer (without overfitting and achieve an even better CV score than experiment demonstrated).</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2727164,
      "author_name": "Idith Haber",
      "author_url": "",
      "post_date": "2024-04-01T16:19:00.127000",
      "content": "<p>Thank you for all your help.  What is wrong with Adam or AdamW?  None of the classes I took talked about LR schedules.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2727173,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-01T16:24:56.753000",
          "content": "<p>Yes Adam and AdamW will adjust the learning rate internally to some extent for us (compared to other optimizers like SGD which don't). None-the-less, we still benefit from choosing an optimal starting learning rate. And we benefit by adding steps and/or decrease from cosine schedule (when validation loss stops improving).</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 2742517,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-09T00:33:08.250000",
      "content": "<p>i never tried this, but meta just release schedule free optimizer</p>\n<p><a href=\"https://github.com/facebookresearch/schedule_free\" target=\"_blank\">https://github.com/facebookresearch/schedule_free</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2727175,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-04-01T16:26:16.950000",
      "content": "<p>Hi Chris, <br>\nThanks for sharing great discussion.<br>\nI have a question about the balance of <code>batch_size</code>, <code>lr</code> and <code>number_of_epochs</code>.</p>\n<p>In what order should we tune<code>batch_size</code>, <code>lr</code> and <code>number_of_epochs</code> ? <br>\nThe score may depend on the balance of <code>batch_size</code>, <code>lr</code> and <code>number_of_epochs</code>.<br>\nI think the balance should be tuned to maximize the score and speed of improving cycle, but honestly I don' know about it. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2727215,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-01T16:45:10.160000",
          "content": "<p><strong>Short Answer</strong>: I tune B, LR, E quickly and approximately at first. Then run lots of experiments exploring other model related stuff. Then finally I tune B, LR, E more carefully later with my best discovered model pipeline.</p>\n<p><strong>Long Answer</strong>: Personally, I begin with <code>batch_size = 32</code> and then tune <code>starting LR</code> first. Once I find a good starting LR, I then tune <code>number_of_epochs</code> second (with steps or cosine added). And finally I explore <code>batch size</code> third using the formula <code>new batch = old batch * scale</code> and <code>new LR = old LR * scale</code> OR <code>new LR = old LR * sqrt(scale)</code>. </p>\n<p>Some people prefer to first use <code>batch size = Largest Possible</code> where we increase batch size using mixed precision (i.e. AMP in torch) and multiple GPUs and utilize maximum GPU VRAM. By using largest batch size possible first, all experiments to identify LR and epochs are sped up. Afterward, we can explore smaller batch sizes (with scaling formulas above).</p>\n<p>Also note, during the experimentation of exploring  the preprocess, data augmentation, model architecture, etc etc. We often use less <code>number of epochs</code> and possibly <code>larger batch sizes</code> for increased speed of experiments. Then after we find optimal preprocess, data augmentation, model architecture, etc etc we then tune the <code>LR</code>, <code>number of epochs</code>, and <code>batch size</code> and train long and hard for maximum LB score and CV score!</p>",
          "votes": 17,
          "replies": []
        }
      ]
    },
    {
      "id": 2726431,
      "author_name": "Wisp Vale",
      "author_url": "",
      "post_date": "2024-04-01T07:39:34.613000",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for this amazing sharing!</p>\n<p>May I ask how to tune lr when using kfold? Because each fold may perform different with others.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2726522,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-01T08:46:53.333000",
          "content": "<p>Hi. Great question. I look at all the folds and take the best approximation (which works for the most folds). (Or if I'm in a rush, I run the first N and look at the first N out of K folds where N&lt;K). Each experiment, I use the same step schedule (or cosine schedule) for every epoch. For later experiments, I will let all K folds finish and earlier experiments I may stop early with only N out of K folds if I see a pattern.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2742878,
      "author_name": "mono takumi sato",
      "author_url": "",
      "post_date": "2024-04-09T06:01:00.150000",
      "content": "<p>Thank you for sharing your ideas! I learned a lot from you through this competition!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2741155,
      "author_name": "Gowri Shankar Penugonda",
      "author_url": "",
      "post_date": "2024-04-08T07:10:02.743000",
      "content": "<p>Hi Chris, What about warmup steps, What's the best way to tune num_warmup_steps required??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2741285,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-08T08:19:14.903000",
          "content": "<p>I'm not sure. Most times, I do not use warm up steps. However sometimes I have added them and they helped.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2738889,
      "author_name": "Steven_Y",
      "author_url": "",
      "post_date": "2024-04-06T17:24:26.810000",
      "content": "<p>Thank you for sharing Chris!</p>\n<p>I was wondering if you guys use libraries like ray-tune to finetune your models? I'm not sure if it'd be more efficient to tun the parameters manually by intuition or use algorithms to do hyperparameter search. Do you have any advice regarding this? </p>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2739102,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-06T20:15:10.110000",
          "content": "<p>I do everything manually. I think there are some helpful software for hyperparameter search but i don't have experience with it.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2735790,
      "author_name": "Danial Zakaria",
      "author_url": "",
      "post_date": "2024-04-04T22:40:18.287000",
      "content": "<p>Thank you for sharing Chris.</p>\n<p>I wonder what is the explanation behind including 1 bad epoch! Maybe pushing the model to as much as possible before lowering the learning rate.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2733284,
      "author_name": "Sahir Maharaj",
      "author_url": "",
      "post_date": "2024-04-03T15:10:47.837000",
      "content": "<p>Thank you for this brilliant share, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. Bookmarking this for future use</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2732579,
      "author_name": "RogerOcean",
      "author_url": "",
      "post_date": "2024-04-03T08:00:06.850000",
      "content": "<p>thanks for sharing, I tried a larger batch size on your open-source solution to speed up the training process, but the scores were lower. Even after adjusting the learning rate, I didn't get better results. I think I will try combining your shared insights in my experiments. Thanks again for sharing.🥳🥳🥳🥳🥳🥳🥳</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2731461,
      "author_name": "KRISHNA CHAUHAN",
      "author_url": "",
      "post_date": "2024-04-02T18:40:53.003000",
      "content": "<p>Hello Chris,</p>\n<p>Thanks for this amazing sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2728753,
      "author_name": "Igor Volianiuk",
      "author_url": "",
      "post_date": "2024-04-02T11:38:21.937000",
      "content": "<p>Great explanation! Interesting approach to Number of epochs! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2739240,
      "author_name": "Cael Hasse",
      "author_url": "",
      "post_date": "2024-04-06T22:33:52.617000",
      "content": "<p>This is really helpful Chris, cheers!</p>\n<p>This may be a silly question but if your model takes a long time to train even one epoch (and it does not appear to be overfitting), is it worth tuning your hyperparameters on a subset of your data first to speed up the iteration process?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2739267,
          "author_name": "Cael Hasse",
          "author_url": "",
          "post_date": "2024-04-06T23:12:36.553000",
          "content": "<p>I'm thinking it may not be useful in getting the right schedule but at least might be useful in setting the initial lr.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2742866,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-09T05:51:44.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2726349,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-01T06:23:00.863000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2742966,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2024-04-09T06:50:10.493000",
      "content": "<p>Thanks Chris. Very helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2742705,
      "author_name": "Ganesh Talwar",
      "author_url": "",
      "post_date": "2024-04-09T03:16:02.483000",
      "content": "<p>thank you for sharing chris</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2735583,
      "author_name": "Maksim Aleksandrov",
      "author_url": "",
      "post_date": "2024-04-04T19:25:53.507000",
      "content": "<p>thanks, it`s very interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2726351,
      "author_name": "kerry sun",
      "author_url": "",
      "post_date": "2024-04-01T06:23:53.603000",
      "content": "<p>You are my ML God and Goal, thanks!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2726171": "Hi everyone. I was rereading my discussion posts and I saw many people asked about tuning learning schedules for NN. So, I thought I would provide some guidance. In this competition, everyone is exploring many models including 2D NN image models, 1D NN time series models, and ML models like GBDT.\n\n# Step 1 - Find Best Starting LR\nThe first thing I like to do is try constant learning rates for total of 5-10 epochs (or more if needed) to \"get a feel\" of learning. So pick LR = 1e-5, or 1e-4, or 1e-3, or 1e-2 (and maybe the midpoints too) and train with constant learning rate for 5-10 epochs (or more if validation loss never degrades).\n\nObserve what the best validation loss achieved is and observe at what epoch the validation loss stops improving and begins getting worse. Also observe which `LR = X` achieves the best validation loss (when using constant LR).\n\n# Step 2 - Find LR Steps\nIf the validation loss is something like `KL-Div = 0.40, 0.35, 0.30, 0.35, 0.37` (for epochs 1-5 respectively), then we observe that validation loss improved for epochs 1-3 and got worse in epochs 4 and 5+ (using `LR = X` for some `X`)\n\nSo, next try a single step learning schedule. Use `LR = X` where `X` is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use `LR = X/10` for epochs 5-10. Then run this and watch.\n\n# Step 3 - Find LR Steps\nWe now pay attention to epochs `5-10` or whatever epochs correspond to your decreased `LR = X/10.`. For these epochs we see `KL-Div = 0.29, 0.28, 0.29, 0.30, 0.31, 0.32`. So we see that now epoch 5 and 6 improve compared to epochs 1-4 but epoch 6 gets worse. So we now set the `LR = X, X, X, X, X/10, X/10, X/10, X/100, X/100 X/100` (for epochs 1-10)\n\nIn other words, we use `LR = X` for epochs 1-4 from our observation earlier. We use `LR = X/10` for epochs 5-7 (include 1 bad epoch), from our observations earlier. And we add `LR = X/100` for epochs 8-10. Then run this and watch.\n\n# Step 4 - Observe Results\nBy now, we see the pattern. We keep using constant learning rate with steps where needed and run. Then we see where the validation loss stops improving. Then we add another step using `LR_new = LR_old/10.` and we run again. In total, we will add 2 or 3 steps (or more until validation stops improving). So that our LR hits `X, X/10. X/100, and X/1000` or so. (Where `X` is our starting rate found in step 1 above).\n\nAt this point, we have a well trained model. We can use this as our final model or we can explore cosine learning schedule\n\n# Step 5 - Convert to Cosine Schedule\nIf we wish to explore cosine learning schedule, then we can use Steps 1-4 above to see how many epochs we were able to keep improving the validation loss.\n\nImagine that we we able to achieve validation loss = `KL-Div = 0.40, 0.35, 0.30, 0.35, 0.29, 0.28, 0.29, 0.27, 0.28, 0.28` when using 3 steps and 10 epochs with `LR = X, X, X,  X, X/10, X/10, X/10, X/100, X/100, X/100`. We observe that LR ranged from `X` to `X/100` from `epoch=1` to `epoch=8` (where validation loss was improving).\n\nTherefore we can now try cosine learning rate with `begin : LR = X` and `end : LR = X/100.` for 8 epochs. Often using cosine learning rate will perform better than using steps above but not always.\n\nAfter trying `start = X, end = X/100, epochs = 8`, we can also try `epochs = 7 and 9` because some times using cosine learning rate requires slightly different epochs than constant.\n\n# Conclusion\nIn conclusion we saw how to tune a step wise learning rate. In step 1 we find the optimal starting learning rate. Then in steps 2-4 we find where to add the steps.\n\nIn step 5, we explore converting the step wise learning rate into a cosine learning rate for possible improvement.",
    "2726290": "Hello Chris,\n\nThanks for this amazing sharing!\n>So, next try a single step learning schedule. Use LR = X where X is what we used previously for epochs 1-4 (yes, we include 1 bad epoch) and use LR = X/10 for epochs 5-10. Then run this and watch.\n\nCan I have a curious question why do we need to include one bad epoch when decreasing the LR?",
    "2726197": "Hello Chris,\n\nThanks for sharing this guidance. I'll try in future experiments!\nI would like to ask a question about `batch_size`. For my local experiments, I observe `batch_size` matters a lot.\nFor example, `batch_size=16` and `batch_size=32` can have CV score difference within the range of 0.02 ~ 0.05. Do you fix a commonly chosen `batch_size` during lr and epoch tuning, then switch to `batch_size` tuning after lr and epoch are well-tuned? Thanks a lot.",
    "2728838": "I usually use Adan optimizer with Onecycle lr, decreased LR = X/10 or increased LR=10*X.",
    "2733871": "Would you conduct a new LR and batch size parameter search for each experiment you run to validate new data augmentation techniques? I was trying out different ways to augment my dataset but found that even simple tricks that seemed to work in the literature weren't improving my score. I was wondering if it was because I wasn't finding the ideal parameters for each data augmentation experiment.\n\nIn general, I was trying to figure out some sort of pattern for when I should conduct a new LR/batch size search. How do you approach this?",
    "2727164": "Thank you for all your help.  What is wrong with Adam or AdamW?  None of the classes I took talked about LR schedules.",
    "2742517": "i never tried this, but meta just release schedule free optimizer\n\nhttps://github.com/facebookresearch/schedule_free",
    "2727175": "Hi Chris, \nThanks for sharing great discussion.\nI have a question about the balance of `batch_size`, `lr` and `number_of_epochs`.\n\nIn what order should we tune`batch_size`, `lr` and `number_of_epochs` ? \nThe score may depend on the balance of `batch_size`, `lr` and `number_of_epochs`.\nI think the balance should be tuned to maximize the score and speed of improving cycle, but honestly I don' know about it. ",
    "2726431": "Hello Chris,\n\nThanks for this amazing sharing!\n\nMay I ask how to tune lr when using kfold? Because each fold may perform different with others.",
    "2742878": "Thank you for sharing your ideas! I learned a lot from you through this competition!!",
    "2741155": "Hi Chris, What about warmup steps, What's the best way to tune num_warmup_steps required??",
    "2738889": "Thank you for sharing Chris!\n\nI was wondering if you guys use libraries like ray-tune to finetune your models? I'm not sure if it'd be more efficient to tun the parameters manually by intuition or use algorithms to do hyperparameter search. Do you have any advice regarding this? \n\nThanks!",
    "2735790": "Thank you for sharing Chris.\n\nI wonder what is the explanation behind including 1 bad epoch! Maybe pushing the model to as much as possible before lowering the learning rate.",
    "2733284": "Thank you for this brilliant share, @cdeotte. Bookmarking this for future use",
    "2732579": "thanks for sharing, I tried a larger batch size on your open-source solution to speed up the training process, but the scores were lower. Even after adjusting the learning rate, I didn't get better results. I think I will try combining your shared insights in my experiments. Thanks again for sharing.🥳🥳🥳🥳🥳🥳🥳",
    "2731461": "Hello Chris,\n\nThanks for this amazing sharing!",
    "2728753": "Great explanation! Interesting approach to Number of epochs! @cdeotte ",
    "2739240": "This is really helpful Chris, cheers!\n\nThis may be a silly question but if your model takes a long time to train even one epoch (and it does not appear to be overfitting), is it worth tuning your hyperparameters on a subset of your data first to speed up the iteration process?",
    "2742866": "",
    "2726349": "",
    "2742966": "Thanks Chris. Very helpful!",
    "2742705": "thank you for sharing chris",
    "2735583": "thanks, it`s very interesting",
    "2726351": "You are my ML God and Goal, thanks!"
  }
}