{
  "id": 175386,
  "title": "Iterating faster by careful sampling (157th private)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175386",
  "author_name": "",
  "post_date": "2020-08-18T04:19:05.750466200Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi everyone. I just wanted to share my experience entering the competition pretty late (last 20 days) and the things that helped me iterate faster (using a PyTorch setup).</p>\n<p>Initially, it was a bit overwhelming to see the high public LB scores and so many highly voted discussions + public kernels. This was compounded by the fact that I have a very strong preference for PyTorch and so, I was reluctant to move away from that. However, upon going through the most upvoted ones step-by-step (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172003\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168152\" target=\"_blank\">this</a> was very helpful), especially the ones by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, where he kept adding links to high scoring public kernels (like <a href=\"https://www.kaggle.com/ajaykumar7778/efficientnet-cv/notebook?scriptVersionId=40043341\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">this</a>), along with several relevant discussions (like the one on <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">Triple Stratified Splits</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">Data from past competitions</a>), provided me with a good starting point. I also tried sticking to an advice given by Jeremy Howard, that one can get a pretty good result using images of lower resolution.</p>\n<p>So, I got started with images of 384x384 and instead of upsampling the melanoma class, for each epoch, I used to downsample an equal number of samples from the majority class to match the count of the melanoma class. To avoid wasting data, I sample a fresh set of images from the majority class in each epoch. This ensured that my epochs were shorter - on a V100, I was able to complete an epoch under 2 mins. This was key for me to iterate faster. I am a big fan of finding the right LR scheduler that can provide a major lift quickly. <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR\" target=\"_blank\">OneCycle</a> has consistently provided faster performance for classification tasks in my previous experiments and using this, I was able to arrive at a schedule that provided me with a decent starting point to keep iterating. In the final week, I just re-ran my best results with 384x384 images with 512x512 and ensembled a few models to get my best score (public LB 95.7, private LB 94.03, CV: 0.93).</p>\n<p>I plan to make a separate post explaining what exactly worked for me finally and also, some things that didn't work along with releasing the entire code for the same (after some cleanup). I hope this helps!</p>",
  "messages": [
    {
      "id": "974834",
      "postDate": "08/18/2020 04:19:05",
      "content": "<p>Hi everyone. I just wanted to share my experience entering the competition pretty late (last 20 days) and the things that helped me iterate faster (using a PyTorch setup).</p>\n<p>Initially, it was a bit overwhelming to see the high public LB scores and so many highly voted discussions + public kernels. This was compounded by the fact that I have a very strong preference for PyTorch and so, I was reluctant to move away from that. However, upon going through the most upvoted ones step-by-step (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172003\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168152\" target=\"_blank\">this</a> was very helpful), especially the ones by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, where he kept adding links to high scoring public kernels (like <a href=\"https://www.kaggle.com/ajaykumar7778/efficientnet-cv/notebook?scriptVersionId=40043341\" target=\"_blank\">this</a> and <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">this</a>), along with several relevant discussions (like the one on <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526\" target=\"_blank\">Triple Stratified Splits</a> and <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910\" target=\"_blank\">Data from past competitions</a>), provided me with a good starting point. I also tried sticking to an advice given by Jeremy Howard, that one can get a pretty good result using images of lower resolution.</p>\n<p>So, I got started with images of 384x384 and instead of upsampling the melanoma class, for each epoch, I used to downsample an equal number of samples from the majority class to match the count of the melanoma class. To avoid wasting data, I sample a fresh set of images from the majority class in each epoch. This ensured that my epochs were shorter - on a V100, I was able to complete an epoch under 2 mins. This was key for me to iterate faster. I am a big fan of finding the right LR scheduler that can provide a major lift quickly. <a href=\"https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR\" target=\"_blank\">OneCycle</a> has consistently provided faster performance for classification tasks in my previous experiments and using this, I was able to arrive at a schedule that provided me with a decent starting point to keep iterating. In the final week, I just re-ran my best results with 384x384 images with 512x512 and ensembled a few models to get my best score (public LB 95.7, private LB 94.03, CV: 0.93).</p>\n<p>I plan to make a separate post explaining what exactly worked for me finally and also, some things that didn't work along with releasing the entire code for the same (after some cleanup). I hope this helps!</p>",
      "rawMarkdown": "Hi everyone. I just wanted to share my experience entering the competition pretty late (last 20 days) and the things that helped me iterate faster (using a PyTorch setup).\n\nInitially, it was a bit overwhelming to see the high public LB scores and so many highly voted discussions + public kernels. This was compounded by the fact that I have a very strong preference for PyTorch and so, I was reluctant to move away from that. However, upon going through the most upvoted ones step-by-step ([this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172003) and [this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168152) was very helpful), especially the ones by @cdeotte, where he kept adding links to high scoring public kernels (like [this](https://www.kaggle.com/ajaykumar7778/efficientnet-cv/notebook?scriptVersionId=40043341) and [this](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords)), along with several relevant discussions (like the one on [Triple Stratified Splits](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) and [Data from past competitions](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910)), provided me with a good starting point. I also tried sticking to an advice given by Jeremy Howard, that one can get a pretty good result using images of lower resolution.\n\nSo, I got started with images of 384x384 and instead of upsampling the melanoma class, for each epoch, I used to downsample an equal number of samples from the majority class to match the count of the melanoma class. To avoid wasting data, I sample a fresh set of images from the majority class in each epoch. This ensured that my epochs were shorter - on a V100, I was able to complete an epoch under 2 mins. This was key for me to iterate faster. I am a big fan of finding the right LR scheduler that can provide a major lift quickly. [OneCycle](https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR) has consistently provided faster performance for classification tasks in my previous experiments and using this, I was able to arrive at a schedule that provided me with a decent starting point to keep iterating. In the final week, I just re-ran my best results with 384x384 images with 512x512 and ensembled a few models to get my best score (public LB 95.7, private LB 94.03, CV: 0.93).\n\nI plan to make a separate post explaining what exactly worked for me finally and also, some things that didn't work along with releasing the entire code for the same (after some cleanup). I hope this helps!",
      "votes": null
    },
    {
      "id": "974839",
      "postDate": "08/18/2020 04:23:58",
      "content": "<p>Downsampling and using a different portion of the majority class each epoch is very smart. Great job. Congratulations on silver medal.</p>",
      "rawMarkdown": "Downsampling and using a different portion of the majority class each epoch is very smart. Great job. Congratulations on silver medal.",
      "votes": null
    },
    {
      "id": "974849",
      "postDate": "08/18/2020 04:27:35",
      "content": "<p>Thank you, Chris. Really grateful for everything you have done for the community. Congratulations to you as well! :)</p>",
      "rawMarkdown": "Thank you, Chris. Really grateful for everything you have done for the community. Congratulations to you as well! :)",
      "votes": null
    },
    {
      "id": "975928",
      "postDate": "08/18/2020 14:35:52",
      "content": "<p>Congratulations ! How did you fix the parameters LR schedule ? Did you try a number of discrete values and then seeing what makes the model converge fast (How to measure this ? Difference of loss between epochs ?) ?</p>",
      "rawMarkdown": "Congratulations ! How did you fix the parameters LR schedule ? Did you try a number of discrete values and then seeing what makes the model converge fast (How to measure this ? Difference of loss between epochs ?) ?",
      "votes": null
    },
    {
      "id": "976152",
      "postDate": "08/18/2020 17:08:16",
      "content": "<p>LR range test really helps here. Essentially, you keep a very small learning rate as min LR and a reasonably large learning rate as max LR and keep increasing the LR from min to max over 1-2 epochs. You track the learning rate after each batch and plot the LR vs loss plot. The LR till which the loss keeps decreasing clearly should be chosen as maxLR. More details can be found <a href=\"https://towardsdatascience.com/finding-good-learning-rate-and-the-one-cycle-policy-7159fe1db5d6\" target=\"_blank\">here</a>. </p>",
      "rawMarkdown": "LR range test really helps here. Essentially, you keep a very small learning rate as min LR and a reasonably large learning rate as max LR and keep increasing the LR from min to max over 1-2 epochs. You track the learning rate after each batch and plot the LR vs loss plot. The LR till which the loss keeps decreasing clearly should be chosen as maxLR. More details can be found [here](https://towardsdatascience.com/finding-good-learning-rate-and-the-one-cycle-policy-7159fe1db5d6).",
      "votes": null
    },
    {
      "id": "976716",
      "postDate": "08/19/2020 04:02:37",
      "content": "<p>Thanks you for clarifying !</p>",
      "rawMarkdown": "Thanks you for clarifying !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 974839,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/18/2020 04:23:58",
      "content": "<p>Downsampling and using a different portion of the majority class each epoch is very smart. Great job. Congratulations on silver medal.</p>",
      "votes": null,
      "replies": [
        {
          "id": 974849,
          "author_name": "themlenthusiast",
          "author_url": "",
          "post_date": "08/18/2020 04:27:35",
          "content": "<p>Thank you, Chris. Really grateful for everything you have done for the community. Congratulations to you as well! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 975928,
      "author_name": "realsid",
      "author_url": "",
      "post_date": "08/18/2020 14:35:52",
      "content": "<p>Congratulations ! How did you fix the parameters LR schedule ? Did you try a number of discrete values and then seeing what makes the model converge fast (How to measure this ? Difference of loss between epochs ?) ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 976152,
          "author_name": "themlenthusiast",
          "author_url": "",
          "post_date": "08/18/2020 17:08:16",
          "content": "<p>LR range test really helps here. Essentially, you keep a very small learning rate as min LR and a reasonably large learning rate as max LR and keep increasing the LR from min to max over 1-2 epochs. You track the learning rate after each batch and plot the LR vs loss plot. The LR till which the loss keeps decreasing clearly should be chosen as maxLR. More details can be found <a href=\"https://towardsdatascience.com/finding-good-learning-rate-and-the-one-cycle-policy-7159fe1db5d6\" target=\"_blank\">here</a>. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 976716,
          "author_name": "realsid",
          "author_url": "",
          "post_date": "08/19/2020 04:02:37",
          "content": "<p>Thanks you for clarifying !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "974834": "Hi everyone. I just wanted to share my experience entering the competition pretty late (last 20 days) and the things that helped me iterate faster (using a PyTorch setup).\n\nInitially, it was a bit overwhelming to see the high public LB scores and so many highly voted discussions + public kernels. This was compounded by the fact that I have a very strong preference for PyTorch and so, I was reluctant to move away from that. However, upon going through the most upvoted ones step-by-step ([this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172003) and [this](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/168152) was very helpful), especially the ones by @cdeotte, where he kept adding links to high scoring public kernels (like [this](https://www.kaggle.com/ajaykumar7778/efficientnet-cv/notebook?scriptVersionId=40043341) and [this](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords)), along with several relevant discussions (like the one on [Triple Stratified Splits](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/165526) and [Data from past competitions](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164910)), provided me with a good starting point. I also tried sticking to an advice given by Jeremy Howard, that one can get a pretty good result using images of lower resolution.\n\nSo, I got started with images of 384x384 and instead of upsampling the melanoma class, for each epoch, I used to downsample an equal number of samples from the majority class to match the count of the melanoma class. To avoid wasting data, I sample a fresh set of images from the majority class in each epoch. This ensured that my epochs were shorter - on a V100, I was able to complete an epoch under 2 mins. This was key for me to iterate faster. I am a big fan of finding the right LR scheduler that can provide a major lift quickly. [OneCycle](https://pytorch.org/docs/stable/optim.html#torch.optim.lr_scheduler.OneCycleLR) has consistently provided faster performance for classification tasks in my previous experiments and using this, I was able to arrive at a schedule that provided me with a decent starting point to keep iterating. In the final week, I just re-ran my best results with 384x384 images with 512x512 and ensembled a few models to get my best score (public LB 95.7, private LB 94.03, CV: 0.93).\n\nI plan to make a separate post explaining what exactly worked for me finally and also, some things that didn't work along with releasing the entire code for the same (after some cleanup). I hope this helps!",
    "974839": "Downsampling and using a different portion of the majority class each epoch is very smart. Great job. Congratulations on silver medal.",
    "974849": "Thank you, Chris. Really grateful for everything you have done for the community. Congratulations to you as well! :)",
    "975928": "Congratulations ! How did you fix the parameters LR schedule ? Did you try a number of discrete values and then seeing what makes the model converge fast (How to measure this ? Difference of loss between epochs ?) ?",
    "976152": "LR range test really helps here. Essentially, you keep a very small learning rate as min LR and a reasonably large learning rate as max LR and keep increasing the LR from min to max over 1-2 epochs. You track the learning rate after each batch and plot the LR vs loss plot. The LR till which the loss keeps decreasing clearly should be chosen as maxLR. More details can be found [here](https://towardsdatascience.com/finding-good-learning-rate-and-the-one-cycle-policy-7159fe1db5d6).",
    "976716": "Thanks you for clarifying !"
  },
  "source": "meta"
}