{
  "id": 195915,
  "title": "Is my model overfitting?",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/195915",
  "author_name": "",
  "post_date": "2020-11-08T07:53:02.250852100Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>I have trained my <a href=\"https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model\" target=\"_blank\">R34_baseline_model</a> for about 2M samples (~68k batches of 32 each) and this is the resulting training, validation curve I get.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa72ef5d5c1e9603d889f10385813d3e%2Fmodel_train_val_curves.jpg?generation=1604820520692494&amp;alt=media\" alt=\"Training_validation_curves\"></p>\n<p>Since I am using only kaggle for training, I average the training score (using previous iterations avg error). I read this <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">discussion</a> where it's mentioned I should have got a validation loss &lt; 30 if I had tuned the model correctly.</p>\n<p>But in my case, I'm tuning the learning rate just by my gut feel (0.4e-3 to 1e-3) at end of every ~7k batches manually and not using any learning rate schedulers. (Just read upon great posts - <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186770\" target=\"_blank\">lr_finder</a>, <a href=\"https://www.kaggle.com/isbhargav/guide-to-pytorch-learning-rate-scheduling\" target=\"_blank\">lr_schedulers</a> and <a href=\"https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee\" target=\"_blank\">Cyclical LR</a>)</p>\n<p>I haven't submitted this model to the leaderboard yet, still I feel I have overfitted my model. Just wanted to know how much improvement I could have made if I had made decent implementation of lr_schedulers?</p>",
  "messages": [
    {
      "id": "1072424",
      "postDate": "11/08/2020 07:53:02",
      "content": "<p>Hi,</p>\n<p>I have trained my <a href=\"https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model\" target=\"_blank\">R34_baseline_model</a> for about 2M samples (~68k batches of 32 each) and this is the resulting training, validation curve I get.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa72ef5d5c1e9603d889f10385813d3e%2Fmodel_train_val_curves.jpg?generation=1604820520692494&amp;alt=media\" alt=\"Training_validation_curves\"></p>\n<p>Since I am using only kaggle for training, I average the training score (using previous iterations avg error). I read this <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">discussion</a> where it's mentioned I should have got a validation loss &lt; 30 if I had tuned the model correctly.</p>\n<p>But in my case, I'm tuning the learning rate just by my gut feel (0.4e-3 to 1e-3) at end of every ~7k batches manually and not using any learning rate schedulers. (Just read upon great posts - <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186770\" target=\"_blank\">lr_finder</a>, <a href=\"https://www.kaggle.com/isbhargav/guide-to-pytorch-learning-rate-scheduling\" target=\"_blank\">lr_schedulers</a> and <a href=\"https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee\" target=\"_blank\">Cyclical LR</a>)</p>\n<p>I haven't submitted this model to the leaderboard yet, still I feel I have overfitted my model. Just wanted to know how much improvement I could have made if I had made decent implementation of lr_schedulers?</p>",
      "rawMarkdown": "Hi,\n\nI have trained my [R34_baseline_model](https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model) for about 2M samples (~68k batches of 32 each) and this is the resulting training, validation curve I get.\n\n![Training_validation_curves](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa72ef5d5c1e9603d889f10385813d3e%2Fmodel_train_val_curves.jpg?generation=1604820520692494&alt=media)\n\nSince I am using only kaggle for training, I average the training score (using previous iterations avg error). I read this [discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762) where it's mentioned I should have got a validation loss < 30 if I had tuned the model correctly.\n\nBut in my case, I'm tuning the learning rate just by my gut feel (0.4e-3 to 1e-3) at end of every ~7k batches manually and not using any learning rate schedulers. (Just read upon great posts - [lr_finder](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186770), [lr_schedulers](https://www.kaggle.com/isbhargav/guide-to-pytorch-learning-rate-scheduling) and [Cyclical LR](https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee))\n\nI haven't submitted this model to the leaderboard yet, still I feel I have overfitted my model. Just wanted to know how much improvement I could have made if I had made decent implementation of lr_schedulers?",
      "votes": null
    },
    {
      "id": "1072599",
      "postDate": "11/08/2020 13:49:10",
      "content": "<p>It does not look like it, but then again, you have too few iterations to tell for sure. </p>",
      "rawMarkdown": "It does not look like it, but then again, you have too few iterations to tell for sure.",
      "votes": null
    },
    {
      "id": "1072675",
      "postDate": "11/08/2020 15:23:45",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/moodoki\" target=\"_blank\">@moodoki</a>, thanks for your feedback. I thought that 2M samples should be enough for achieving at least ~25 on the LB score. It seems that I have to play with the learning rate and other hyper parameters to get that performance range. </p>",
      "rawMarkdown": "Hi @moodoki, thanks for your feedback. I thought that 2M samples should be enough for achieving at least ~25 on the LB score. It seems that I have to play with the learning rate and other hyper parameters to get that performance range.",
      "votes": null
    },
    {
      "id": "1072814",
      "postDate": "11/08/2020 18:00:22",
      "content": "<p><a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a>, why don't you submit the result to see the LB score as well? The CV/LB difference is an additional variable you can use to make that judgement.</p>",
      "rawMarkdown": "suryajrrafl, why don't you submit the result to see the LB score as well? The CV/LB difference is an additional variable you can use to make that judgement.",
      "votes": null
    },
    {
      "id": "1072848",
      "postDate": "11/08/2020 19:04:53",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a>, I did submit to the LB and found score to be similar to the validation (~65). I am trying to  check impact of learning schedulers (ReduceLRonPlateau and <a href=\"https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee\" target=\"_blank\">cyclicLR</a>.<br>\nFor that I have to fix my max LR. This is what I have done so far :</p>\n<p>`<br>\nlr_start = 1e-7<br>\nlr_end   = 0.1<br>\nnum_steps = 100</p>\n<p>test_optimizer = optim.Adam(model.parameters(), lr=lr_start)<br>\nlr_lambda = lambda x: math.exp(x * math.log(lr_end / lr_start) / num_steps)<br>\nscheduler = torch.optim.lr_scheduler.LambdaLR(test_optimizer, lr_lambda)</p>\n<p>lr_find_loss = []<br>\nlr_find_lr = []</p>\n<p>test_optimizer.zero_grad()<br>\nloss.backward()<br>\ntest_optimizer.step()</p>\n<h3>Update LR</h3>\n<p>scheduler.step()<br>\nlr_step = test_optimizer.state_dict()[\"param_groups\"][0][\"lr\"]<br>\nlr_find_lr.append(lr_step)</p>\n<p>def exp_avg(data, factor = 0.05):<br>\n    smooth_data = [data[0]]<br>\n    for i in range(1, len(data)):<br>\n        smooth_data.append(factor * data[i] + (1-factor) * smooth_data[-1])<br>\n    return smooth_data</p>\n<p>smooth_losses = exp_avg(lr_find_loss, 0.025)<br>\n`</p>\n<p>This is the result I got.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fcc5e2e8f6edd4e34ebfaa71c64651085%2Flr_vs_loss_curve.jpg?generation=1604862162801702&amp;alt=media\" alt=\"lr_curve\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa3bdd5ba5a3a82b243bc94edccb7a1f%2Flr_curve.jpg?generation=1604862016408187&amp;alt=media\" alt=\"lr_vs_loss\"></p>\n<p>From this result, can I conclude that my max learning rate to be ~exp(-5). Or Am I missing something here? How do I select my max learning rate? Or is there any inbuilt function available?</p>",
      "rawMarkdown": "Hi @sheriytm, I did submit to the LB and found score to be similar to the validation (~65). I am trying to  check impact of learning schedulers (ReduceLRonPlateau and [cyclicLR](https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee).\nFor that I have to fix my max LR. This is what I have done so far :\n\n`\nlr_start = 1e-7\nlr_end   = 0.1\nnum_steps = 100\n\ntest_optimizer = optim.Adam(model.parameters(), lr=lr_start)\nlr_lambda = lambda x: math.exp(x * math.log(lr_end / lr_start) / num_steps)\nscheduler = torch.optim.lr_scheduler.LambdaLR(test_optimizer, lr_lambda)\n\nlr_find_loss = []\nlr_find_lr = []\n\ntest_optimizer.zero_grad()\nloss.backward()\ntest_optimizer.step()\n\n### Update LR\nscheduler.step()\nlr_step = test_optimizer.state_dict()[\"param_groups\"][0][\"lr\"]\nlr_find_lr.append(lr_step)\n\ndef exp_avg(data, factor = 0.05):\n    smooth_data = [data[0]]\n    for i in range(1, len(data)):\n        smooth_data.append(factor * data[i] + (1-factor) * smooth_data[-1])\n    return smooth_data\n\nsmooth_losses = exp_avg(lr_find_loss, 0.025)\n`\n\nThis is the result I got.\n![lr_curve](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fcc5e2e8f6edd4e34ebfaa71c64651085%2Flr_vs_loss_curve.jpg?generation=1604862162801702&alt=media)\n\n![lr_vs_loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa3bdd5ba5a3a82b243bc94edccb7a1f%2Flr_curve.jpg?generation=1604862016408187&alt=media)\n\nFrom this result, can I conclude that my max learning rate to be ~exp(-5). Or Am I missing something here? How do I select my max learning rate? Or is there any inbuilt function available?",
      "votes": null
    },
    {
      "id": "1074123",
      "postDate": "11/10/2020 09:30:45",
      "content": "<p>I have tried to implement a simple lr_scheduler in this <a href=\"https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model\" target=\"_blank\">notebook</a>. If possible, can someone give feedback on the parameters I have chosen. My model runs for about 8000 iterations in one commit (Batch size 32). I find the initial learning rate manually as of now using the method described above. Considering this I have chosen</p>\n<ol>\n<li>reduction factor = 0.5</li>\n<li>patience  = 200 (batches)</li>\n<li>scheduler operates on smoothened loss  (exp weighted average of 0.04 ~ last 25 samples)</li>\n<li>threshold = 0.1 </li>\n</ol>\n<p>Ideally, how are these parameters estimated? </p>",
      "rawMarkdown": "I have tried to implement a simple lr_scheduler in this [notebook](https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model). If possible, can someone give feedback on the parameters I have chosen. My model runs for about 8000 iterations in one commit (Batch size 32). I find the initial learning rate manually as of now using the method described above. Considering this I have chosen\n\n1.  reduction factor = 0.5\n2. patience  = 200 (batches)\n3. scheduler operates on smoothened loss  (exp weighted average of 0.04 ~ last 25 samples)\n4. threshold = 0.1 \n\nIdeally, how are these parameters estimated?",
      "votes": null
    },
    {
      "id": "1079123",
      "postDate": "11/15/2020 16:35:48",
      "content": "<p>I have managed to implement a simple learning rate finder <a href=\"https://www.kaggle.com/suryajrrafl/lyft-resnet18-baseline\" target=\"_blank\">R18_with_lr_finder</a>. Suggestions are most welcome</p>",
      "rawMarkdown": "I have managed to implement a simple learning rate finder [R18_with_lr_finder](https://www.kaggle.com/suryajrrafl/lyft-resnet18-baseline). Suggestions are most welcome",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1072599,
      "author_name": "moodoki",
      "author_url": "",
      "post_date": "11/08/2020 13:49:10",
      "content": "<p>It does not look like it, but then again, you have too few iterations to tell for sure. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1072675,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "11/08/2020 15:23:45",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/moodoki\" target=\"_blank\">@moodoki</a>, thanks for your feedback. I thought that 2M samples should be enough for achieving at least ~25 on the LB score. It seems that I have to play with the learning rate and other hyper parameters to get that performance range. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1072814,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/08/2020 18:00:22",
      "content": "<p><a href=\"https://www.kaggle.com/suryajrrafl\" target=\"_blank\">@suryajrrafl</a>, why don't you submit the result to see the LB score as well? The CV/LB difference is an additional variable you can use to make that judgement.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1072848,
          "author_name": "suryajrrafl",
          "author_url": "",
          "post_date": "11/08/2020 19:04:53",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a>, I did submit to the LB and found score to be similar to the validation (~65). I am trying to  check impact of learning schedulers (ReduceLRonPlateau and <a href=\"https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee\" target=\"_blank\">cyclicLR</a>.<br>\nFor that I have to fix my max LR. This is what I have done so far :</p>\n<p>`<br>\nlr_start = 1e-7<br>\nlr_end   = 0.1<br>\nnum_steps = 100</p>\n<p>test_optimizer = optim.Adam(model.parameters(), lr=lr_start)<br>\nlr_lambda = lambda x: math.exp(x * math.log(lr_end / lr_start) / num_steps)<br>\nscheduler = torch.optim.lr_scheduler.LambdaLR(test_optimizer, lr_lambda)</p>\n<p>lr_find_loss = []<br>\nlr_find_lr = []</p>\n<p>test_optimizer.zero_grad()<br>\nloss.backward()<br>\ntest_optimizer.step()</p>\n<h3>Update LR</h3>\n<p>scheduler.step()<br>\nlr_step = test_optimizer.state_dict()[\"param_groups\"][0][\"lr\"]<br>\nlr_find_lr.append(lr_step)</p>\n<p>def exp_avg(data, factor = 0.05):<br>\n    smooth_data = [data[0]]<br>\n    for i in range(1, len(data)):<br>\n        smooth_data.append(factor * data[i] + (1-factor) * smooth_data[-1])<br>\n    return smooth_data</p>\n<p>smooth_losses = exp_avg(lr_find_loss, 0.025)<br>\n`</p>\n<p>This is the result I got.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fcc5e2e8f6edd4e34ebfaa71c64651085%2Flr_vs_loss_curve.jpg?generation=1604862162801702&amp;alt=media\" alt=\"lr_curve\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa3bdd5ba5a3a82b243bc94edccb7a1f%2Flr_curve.jpg?generation=1604862016408187&amp;alt=media\" alt=\"lr_vs_loss\"></p>\n<p>From this result, can I conclude that my max learning rate to be ~exp(-5). Or Am I missing something here? How do I select my max learning rate? Or is there any inbuilt function available?</p>",
          "votes": null,
          "replies": [
            {
              "id": 1074123,
              "author_name": "suryajrrafl",
              "author_url": "",
              "post_date": "11/10/2020 09:30:45",
              "content": "<p>I have tried to implement a simple lr_scheduler in this <a href=\"https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model\" target=\"_blank\">notebook</a>. If possible, can someone give feedback on the parameters I have chosen. My model runs for about 8000 iterations in one commit (Batch size 32). I find the initial learning rate manually as of now using the method described above. Considering this I have chosen</p>\n<ol>\n<li>reduction factor = 0.5</li>\n<li>patience  = 200 (batches)</li>\n<li>scheduler operates on smoothened loss  (exp weighted average of 0.04 ~ last 25 samples)</li>\n<li>threshold = 0.1 </li>\n</ol>\n<p>Ideally, how are these parameters estimated? </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 1079123,
      "author_name": "suryajrrafl",
      "author_url": "",
      "post_date": "11/15/2020 16:35:48",
      "content": "<p>I have managed to implement a simple learning rate finder <a href=\"https://www.kaggle.com/suryajrrafl/lyft-resnet18-baseline\" target=\"_blank\">R18_with_lr_finder</a>. Suggestions are most welcome</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1072424": "Hi,\n\nI have trained my [R34_baseline_model](https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model) for about 2M samples (~68k batches of 32 each) and this is the resulting training, validation curve I get.\n\n![Training_validation_curves](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa72ef5d5c1e9603d889f10385813d3e%2Fmodel_train_val_curves.jpg?generation=1604820520692494&alt=media)\n\nSince I am using only kaggle for training, I average the training score (using previous iterations avg error). I read this [discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762) where it's mentioned I should have got a validation loss < 30 if I had tuned the model correctly.\n\nBut in my case, I'm tuning the learning rate just by my gut feel (0.4e-3 to 1e-3) at end of every ~7k batches manually and not using any learning rate schedulers. (Just read upon great posts - [lr_finder](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/186770), [lr_schedulers](https://www.kaggle.com/isbhargav/guide-to-pytorch-learning-rate-scheduling) and [Cyclical LR](https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee))\n\nI haven't submitted this model to the leaderboard yet, still I feel I have overfitted my model. Just wanted to know how much improvement I could have made if I had made decent implementation of lr_schedulers?",
    "1072599": "It does not look like it, but then again, you have too few iterations to tell for sure.",
    "1072675": "Hi @moodoki, thanks for your feedback. I thought that 2M samples should be enough for achieving at least ~25 on the LB score. It seems that I have to play with the learning rate and other hyper parameters to get that performance range.",
    "1072814": "suryajrrafl, why don't you submit the result to see the LB score as well? The CV/LB difference is an additional variable you can use to make that judgement.",
    "1072848": "Hi @sheriytm, I did submit to the LB and found score to be similar to the validation (~65). I am trying to  check impact of learning schedulers (ReduceLRonPlateau and [cyclicLR](https://towardsdatascience.com/adaptive-and-cyclical-learning-rates-using-pytorch-2bf904d18dee).\nFor that I have to fix my max LR. This is what I have done so far :\n\n`\nlr_start = 1e-7\nlr_end   = 0.1\nnum_steps = 100\n\ntest_optimizer = optim.Adam(model.parameters(), lr=lr_start)\nlr_lambda = lambda x: math.exp(x * math.log(lr_end / lr_start) / num_steps)\nscheduler = torch.optim.lr_scheduler.LambdaLR(test_optimizer, lr_lambda)\n\nlr_find_loss = []\nlr_find_lr = []\n\ntest_optimizer.zero_grad()\nloss.backward()\ntest_optimizer.step()\n\n### Update LR\nscheduler.step()\nlr_step = test_optimizer.state_dict()[\"param_groups\"][0][\"lr\"]\nlr_find_lr.append(lr_step)\n\ndef exp_avg(data, factor = 0.05):\n    smooth_data = [data[0]]\n    for i in range(1, len(data)):\n        smooth_data.append(factor * data[i] + (1-factor) * smooth_data[-1])\n    return smooth_data\n\nsmooth_losses = exp_avg(lr_find_loss, 0.025)\n`\n\nThis is the result I got.\n![lr_curve](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Fcc5e2e8f6edd4e34ebfaa71c64651085%2Flr_vs_loss_curve.jpg?generation=1604862162801702&alt=media)\n\n![lr_vs_loss](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3481038%2Ffa3bdd5ba5a3a82b243bc94edccb7a1f%2Flr_curve.jpg?generation=1604862016408187&alt=media)\n\nFrom this result, can I conclude that my max learning rate to be ~exp(-5). Or Am I missing something here? How do I select my max learning rate? Or is there any inbuilt function available?",
    "1074123": "I have tried to implement a simple lr_scheduler in this [notebook](https://www.kaggle.com/suryajrrafl/lyft-resnet34-baseline-model). If possible, can someone give feedback on the parameters I have chosen. My model runs for about 8000 iterations in one commit (Batch size 32). I find the initial learning rate manually as of now using the method described above. Considering this I have chosen\n\n1.  reduction factor = 0.5\n2. patience  = 200 (batches)\n3. scheduler operates on smoothened loss  (exp weighted average of 0.04 ~ last 25 samples)\n4. threshold = 0.1 \n\nIdeally, how are these parameters estimated?",
    "1079123": "I have managed to implement a simple learning rate finder [R18_with_lr_finder](https://www.kaggle.com/suryajrrafl/lyft-resnet18-baseline). Suggestions are most welcome"
  },
  "source": "meta"
}