{
  "id": 212091,
  "title": "Things to consider this late (mostly for beginners like me)",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/212091",
  "author_name": "",
  "post_date": "2021-01-17T14:22:27.713989600Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey kagglers!</p>\n<p>As some of you, this is my first \"serious\" competition, where i'm focusing on learning and trying really hard to get at least a bronze medal or simply the best position i possibly can.<br>\nAs per the other discussions, i was a bit overwhelmed at first because implementing custom loss functions or trying out 100 different augmentations took a lot of my time.<br>\nBut despite this things helping out, what i mostly overlooked was a couple Hyperparameters which helped me go from about 88.8 LB (single model, single fold) to ~89.6 LB (single model, single fold) was the following two things:<br>\n1) Different schedulers have different update rules: you should really be careful, in my case for instance. I read up on CosineAnnealingWarmRestarts and it said that it should be updated per batch, but you should call scheduler.step(epoch + batch_no/total_batches) instead of simply scheduler.step() (which i was initially doing).<br>\nThis lead to my scheduler resetting every 10 batches and my LR was all over the place and not converging to an appropriate result.</p>\n<p>2) LR was the most volatile hyperparameter which i severely underlooked. The combination of Batch Size, Scheduler and Learning Rate is what breaks or makes a model for me at the moment. If you change your model or batch size you probably should adjust the learning rate. The learning rate improved my model far more than  tweaking T1 and T2 in the bitempered loss function.<br>\nJust to put things into perspective: an order of magnitude is the difference between the model converging in ~86-87 CV to 88~89 CV.</p>\n<p>Another tip is printing the LR with the train output so that you can check if it is being updated correctly.</p>\n<p>What i mostly will do for now (and please correct me if i'm wrong here) is settling into a good LR for a model, before i start tweaking again with other hyperparameters such as augmentations and different loss function options.</p>\n<p>Thanks for your attention!</p>",
  "messages": [
    {
      "id": "1156924",
      "postDate": "01/17/2021 14:22:27",
      "content": "<p>Hey kagglers!</p>\n<p>As some of you, this is my first \"serious\" competition, where i'm focusing on learning and trying really hard to get at least a bronze medal or simply the best position i possibly can.<br>\nAs per the other discussions, i was a bit overwhelmed at first because implementing custom loss functions or trying out 100 different augmentations took a lot of my time.<br>\nBut despite this things helping out, what i mostly overlooked was a couple Hyperparameters which helped me go from about 88.8 LB (single model, single fold) to ~89.6 LB (single model, single fold) was the following two things:<br>\n1) Different schedulers have different update rules: you should really be careful, in my case for instance. I read up on CosineAnnealingWarmRestarts and it said that it should be updated per batch, but you should call scheduler.step(epoch + batch_no/total_batches) instead of simply scheduler.step() (which i was initially doing).<br>\nThis lead to my scheduler resetting every 10 batches and my LR was all over the place and not converging to an appropriate result.</p>\n<p>2) LR was the most volatile hyperparameter which i severely underlooked. The combination of Batch Size, Scheduler and Learning Rate is what breaks or makes a model for me at the moment. If you change your model or batch size you probably should adjust the learning rate. The learning rate improved my model far more than  tweaking T1 and T2 in the bitempered loss function.<br>\nJust to put things into perspective: an order of magnitude is the difference between the model converging in ~86-87 CV to 88~89 CV.</p>\n<p>Another tip is printing the LR with the train output so that you can check if it is being updated correctly.</p>\n<p>What i mostly will do for now (and please correct me if i'm wrong here) is settling into a good LR for a model, before i start tweaking again with other hyperparameters such as augmentations and different loss function options.</p>\n<p>Thanks for your attention!</p>",
      "rawMarkdown": "Hey kagglers!\n\nAs some of you, this is my first \"serious\" competition, where i'm focusing on learning and trying really hard to get at least a bronze medal or simply the best position i possibly can.\nAs per the other discussions, i was a bit overwhelmed at first because implementing custom loss functions or trying out 100 different augmentations took a lot of my time.\nBut despite this things helping out, what i mostly overlooked was a couple Hyperparameters which helped me go from about 88.8 LB (single model, single fold) to ~89.6 LB (single model, single fold) was the following two things:\n1) Different schedulers have different update rules: you should really be careful, in my case for instance. I read up on CosineAnnealingWarmRestarts and it said that it should be updated per batch, but you should call scheduler.step(epoch + batch_no/total_batches) instead of simply scheduler.step() (which i was initially doing).\nThis lead to my scheduler resetting every 10 batches and my LR was all over the place and not converging to an appropriate result.\n\n2) LR was the most volatile hyperparameter which i severely underlooked. The combination of Batch Size, Scheduler and Learning Rate is what breaks or makes a model for me at the moment. If you change your model or batch size you probably should adjust the learning rate. The learning rate improved my model far more than  tweaking T1 and T2 in the bitempered loss function.\nJust to put things into perspective: an order of magnitude is the difference between the model converging in ~86-87 CV to 88~89 CV.\n\nAnother tip is printing the LR with the train output so that you can check if it is being updated correctly.\n\nWhat i mostly will do for now (and please correct me if i'm wrong here) is settling into a good LR for a model, before i start tweaking again with other hyperparameters such as augmentations and different loss function options.\n\nThanks for your attention!",
      "votes": null
    },
    {
      "id": "1205683",
      "postDate": "02/16/2021 21:58:55",
      "content": "<p>Great pointers. This script I created to monitor Train and Val Loss and LR should also help understand what is going on in the model. <a href=\"https://www.kaggle.com/trushk/model-validation-learning-rate-monitoring\" target=\"_blank\">https://www.kaggle.com/trushk/model-validation-learning-rate-monitoring</a></p>\n<p>Looking at the val loss vs LR might help picking good range of LR as well. </p>",
      "rawMarkdown": "Great pointers. This script I created to monitor Train and Val Loss and LR should also help understand what is going on in the model. https://www.kaggle.com/trushk/model-validation-learning-rate-monitoring\n\nLooking at the val loss vs LR might help picking good range of LR as well.",
      "votes": null
    },
    {
      "id": "1206352",
      "postDate": "02/17/2021 08:46:29",
      "content": "<p>Yes I also noticed the importance of the learning rate (unfortunately a bit late). A good learning rate scheduler with some warm up and and slow decaying towards the end really improved my results</p>",
      "rawMarkdown": "Yes I also noticed the importance of the learning rate (unfortunately a bit late). A good learning rate scheduler with some warm up and and slow decaying towards the end really improved my results",
      "votes": null
    },
    {
      "id": "1207025",
      "postDate": "02/17/2021 17:26:04",
      "content": "<p>Great tip. CosineAnnealingWarmRestarts helped me as well. </p>",
      "rawMarkdown": "Great tip. CosineAnnealingWarmRestarts helped me as well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1205683,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/16/2021 21:58:55",
      "content": "<p>Great pointers. This script I created to monitor Train and Val Loss and LR should also help understand what is going on in the model. <a href=\"https://www.kaggle.com/trushk/model-validation-learning-rate-monitoring\" target=\"_blank\">https://www.kaggle.com/trushk/model-validation-learning-rate-monitoring</a></p>\n<p>Looking at the val loss vs LR might help picking good range of LR as well. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1206352,
      "author_name": "alexanderriedel",
      "author_url": "",
      "post_date": "02/17/2021 08:46:29",
      "content": "<p>Yes I also noticed the importance of the learning rate (unfortunately a bit late). A good learning rate scheduler with some warm up and and slow decaying towards the end really improved my results</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207025,
      "author_name": "krisho007",
      "author_url": "",
      "post_date": "02/17/2021 17:26:04",
      "content": "<p>Great tip. CosineAnnealingWarmRestarts helped me as well. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1156924": "Hey kagglers!\n\nAs some of you, this is my first \"serious\" competition, where i'm focusing on learning and trying really hard to get at least a bronze medal or simply the best position i possibly can.\nAs per the other discussions, i was a bit overwhelmed at first because implementing custom loss functions or trying out 100 different augmentations took a lot of my time.\nBut despite this things helping out, what i mostly overlooked was a couple Hyperparameters which helped me go from about 88.8 LB (single model, single fold) to ~89.6 LB (single model, single fold) was the following two things:\n1) Different schedulers have different update rules: you should really be careful, in my case for instance. I read up on CosineAnnealingWarmRestarts and it said that it should be updated per batch, but you should call scheduler.step(epoch + batch_no/total_batches) instead of simply scheduler.step() (which i was initially doing).\nThis lead to my scheduler resetting every 10 batches and my LR was all over the place and not converging to an appropriate result.\n\n2) LR was the most volatile hyperparameter which i severely underlooked. The combination of Batch Size, Scheduler and Learning Rate is what breaks or makes a model for me at the moment. If you change your model or batch size you probably should adjust the learning rate. The learning rate improved my model far more than  tweaking T1 and T2 in the bitempered loss function.\nJust to put things into perspective: an order of magnitude is the difference between the model converging in ~86-87 CV to 88~89 CV.\n\nAnother tip is printing the LR with the train output so that you can check if it is being updated correctly.\n\nWhat i mostly will do for now (and please correct me if i'm wrong here) is settling into a good LR for a model, before i start tweaking again with other hyperparameters such as augmentations and different loss function options.\n\nThanks for your attention!",
    "1205683": "Great pointers. This script I created to monitor Train and Val Loss and LR should also help understand what is going on in the model. https://www.kaggle.com/trushk/model-validation-learning-rate-monitoring\n\nLooking at the val loss vs LR might help picking good range of LR as well.",
    "1206352": "Yes I also noticed the importance of the learning rate (unfortunately a bit late). A good learning rate scheduler with some warm up and and slow decaying towards the end really improved my results",
    "1207025": "Great tip. CosineAnnealingWarmRestarts helped me as well."
  },
  "source": "meta"
}