{
  "id": 174622,
  "title": "With more complex models (b4,b5) both CV and LB go down",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174622",
  "author_name": "",
  "post_date": "2020-08-14T09:47:53.488307200Z",
  "votes": 2,
  "comment_count": 18,
  "views": 0,
  "content": "<p>So it seemed that I build a polished and consistent basic Pytorch model. I use GPU, mixed precision, BCE, custom scheduler, TTA, a moderate set of standard augmentations (rotation, shift, scale, brightness etc.) plus coarse dropout (max_holes=8, p=0.5). </p>\n<p>Scheduler is</p>\n<pre><code>self.lr_start = 0.000005\nself.lr_min = 0.000001\nself.lr_ramp_ep = 5\nself.lr_sus_ep = 0\nself.lr_decay = 0.8 \nself.lr_max = 0.00000125 * 4 * batch_size\n</code></pre>\n<p>For (effnet b3, size 256, batch size=128) my model delivers the following: CV 0.87843,  CV TTA 0.88557, LB <strong>0.8942</strong>. But things are not getting better with more complex models and larger image sizes.</p>\n<p>(effnet b4, size 256, batch size=96): CV 0.8772,  CV TTA 0.88237, LB <strong>0.8903</strong><br>\n(effnet b5, size 256, batch size=64): CV 0.85882,  CV TTA 0.85967,  LB <strong>0.8761</strong><br>\n(effnet b4, size 384, batch size=40): CV TTA 0.8628, LB <strong>0.8821</strong></p>\n<p>I tried changing parameters of the scheduler (increase lr_min and lr_max, reduce lr_max by changing factor=4 to 2 or 1, etc.) - no big difference. Whatever I do, my model doesn't take advantage of larger effnet models and image sizes. What's more, it gets worse.</p>\n<p>Based on your experience, what could be wrong here? I understand that smaller batch size implies applying a lower learning rate, for example. But what else? E.g. if you switch from b3 to b4 or from 256 to 384, what would you adjust as a rule? </p>",
  "messages": [
    {
      "id": "970252",
      "postDate": "08/14/2020 09:47:53",
      "content": "<p>So it seemed that I build a polished and consistent basic Pytorch model. I use GPU, mixed precision, BCE, custom scheduler, TTA, a moderate set of standard augmentations (rotation, shift, scale, brightness etc.) plus coarse dropout (max_holes=8, p=0.5). </p>\n<p>Scheduler is</p>\n<pre><code>self.lr_start = 0.000005\nself.lr_min = 0.000001\nself.lr_ramp_ep = 5\nself.lr_sus_ep = 0\nself.lr_decay = 0.8 \nself.lr_max = 0.00000125 * 4 * batch_size\n</code></pre>\n<p>For (effnet b3, size 256, batch size=128) my model delivers the following: CV 0.87843,  CV TTA 0.88557, LB <strong>0.8942</strong>. But things are not getting better with more complex models and larger image sizes.</p>\n<p>(effnet b4, size 256, batch size=96): CV 0.8772,  CV TTA 0.88237, LB <strong>0.8903</strong><br>\n(effnet b5, size 256, batch size=64): CV 0.85882,  CV TTA 0.85967,  LB <strong>0.8761</strong><br>\n(effnet b4, size 384, batch size=40): CV TTA 0.8628, LB <strong>0.8821</strong></p>\n<p>I tried changing parameters of the scheduler (increase lr_min and lr_max, reduce lr_max by changing factor=4 to 2 or 1, etc.) - no big difference. Whatever I do, my model doesn't take advantage of larger effnet models and image sizes. What's more, it gets worse.</p>\n<p>Based on your experience, what could be wrong here? I understand that smaller batch size implies applying a lower learning rate, for example. But what else? E.g. if you switch from b3 to b4 or from 256 to 384, what would you adjust as a rule? </p>",
      "rawMarkdown": "So it seemed that I build a polished and consistent basic Pytorch model. I use GPU, mixed precision, BCE, custom scheduler, TTA, a moderate set of standard augmentations (rotation, shift, scale, brightness etc.) plus coarse dropout (max_holes=8, p=0.5). \n\nScheduler is\n```\nself.lr_start = 0.000005\nself.lr_min = 0.000001\nself.lr_ramp_ep = 5\nself.lr_sus_ep = 0\nself.lr_decay = 0.8 \nself.lr_max = 0.00000125 * 4 * batch_size\n```\n \nFor (effnet b3, size 256, batch size=128) my model delivers the following: CV 0.87843,  CV TTA 0.88557, LB **0.8942**. But things are not getting better with more complex models and larger image sizes.\n\n(effnet b4, size 256, batch size=96): CV 0.8772,  CV TTA 0.88237, LB **0.8903**\n(effnet b5, size 256, batch size=64): CV 0.85882,  CV TTA 0.85967,  LB **0.8761**\n(effnet b4, size 384, batch size=40): CV TTA 0.8628, LB **0.8821**\n\nI tried changing parameters of the scheduler (increase lr_min and lr_max, reduce lr_max by changing factor=4 to 2 or 1, etc.) - no big difference. Whatever I do, my model doesn't take advantage of larger effnet models and image sizes. What's more, it gets worse.\n\nBased on your experience, what could be wrong here? I understand that smaller batch size implies applying a lower learning rate, for example. But what else? E.g. if you switch from b3 to b4 or from 256 to 384, what would you adjust as a rule?",
      "votes": null
    },
    {
      "id": "970883",
      "postDate": "08/15/2020 00:10:51",
      "content": "<p>anything above b5 is probably overfitting</p>",
      "rawMarkdown": "anything above b5 is probably overfitting",
      "votes": null
    },
    {
      "id": "970931",
      "postDate": "08/15/2020 02:37:39",
      "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> but you are already at 379 with .9622 so why bother, is this just a side model for fun?  I personally have not found that scheduler to work well, and I have tried those exact parameters.</p>\n<p>I would adjust the learning rate.  Why not run a few epochs with RLRoP, try .001, .0008, .0005, .0001, etc.  maybe patience 2 and decay it .8, .5 or .2 …………just try somethings and get an idea of where it wants to be.  What data are you using? Any upsampling?</p>",
      "rawMarkdown": "dunklerwald but you are already at 379 with .9622 so why bother, is this just a side model for fun?  I personally have not found that scheduler to work well, and I have tried those exact parameters.\n\nI would adjust the learning rate.  Why not run a few epochs with RLRoP, try .001, .0008, .0005, .0001, etc.  maybe patience 2 and decay it .8, .5 or .2 ............just try somethings and get an idea of where it wants to be.  What data are you using? Any upsampling?",
      "votes": null
    },
    {
      "id": "970966",
      "postDate": "08/15/2020 03:48:36",
      "content": "<p>A larger model has more parameters and has a higher chance of overfitting. When increasing to larger models, you need to increase regularization to prevent overfitting.</p>",
      "rawMarkdown": "A larger model has more parameters and has a higher chance of overfitting. When increasing to larger models, you need to increase regularization to prevent overfitting.",
      "votes": null
    },
    {
      "id": "971122",
      "postDate": "08/15/2020 07:40:49",
      "content": "<p>I would assume 9622 is public kernel…</p>",
      "rawMarkdown": "I would assume 9622 is public kernel...",
      "votes": null
    },
    {
      "id": "971167",
      "postDate": "08/15/2020 08:29:54",
      "content": "<p>Your numbers are very low, try first to get better CV/LB with a simple model, say effnet b0.</p>",
      "rawMarkdown": "Your numbers are very low, try first to get better CV/LB with a simple model, say effnet b0.",
      "votes": null
    },
    {
      "id": "971198",
      "postDate": "08/15/2020 09:24:03",
      "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> hey, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  is right. I have an ensemble of my TF models (~0.950), based on the kernel of Chris. So yesterday I just mixed my ensemble with the infamous public ensemble (0.9619).</p>\n<p>I decided to use this competition to learn Pytorch. Looking at dramatic performance of my current Pytorch model, yes, it's more for fun. Actually I started with RLRoP and played out with different learning rates and patience numbers. Somehow a custom scheduler proved to be better in my case (learning process is more smooth and less overfitting).<br>\nI use external data but no upsampling. Minority upsampling didn't work for me also.</p>",
      "rawMarkdown": "brianfeeny hey, @philippsinger  is right. I have an ensemble of my TF models (~0.950), based on the kernel of Chris. So yesterday I just mixed my ensemble with the infamous public ensemble (0.9619).\n\nI decided to use this competition to learn Pytorch. Looking at dramatic performance of my current Pytorch model, yes, it's more for fun. Actually I started with RLRoP and played out with different learning rates and patience numbers. Somehow a custom scheduler proved to be better in my case (learning process is more smooth and less overfitting).\nI use external data but no upsampling. Minority upsampling didn't work for me also.",
      "votes": null
    },
    {
      "id": "971205",
      "postDate": "08/15/2020 09:28:43",
      "content": "<p>B1 showed slightly better performance. Around LB 0.902. So I'm sure I miss some fundamental thing(s) about scaling\\regularization and still need to learn a lot.</p>",
      "rawMarkdown": "B1 showed slightly better performance. Around LB 0.902. So I'm sure I miss some fundamental thing(s) about scaling\\regularization and still need to learn a lot.",
      "votes": null
    },
    {
      "id": "971219",
      "postDate": "08/15/2020 09:48:18",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Chris, thank you. I learned a lot from you during last year and also in the course of this competition.<br>\nSpeaking of regularization, it could be:</p>\n<ul>\n<li>adding dropout layers in the tail</li>\n<li>heavier and more diverse augmentations<br>\nCorrect? What else should I try?</li>\n</ul>\n<p>Is there any rule of thumb about learning rate when we switch to larger models? Or is it a process of trial and error?</p>",
      "rawMarkdown": "cdeotte Chris, thank you. I learned a lot from you during last year and also in the course of this competition.\nSpeaking of regularization, it could be:\n- adding dropout layers in the tail\n- heavier and more diverse augmentations\nCorrect? What else should I try?\n\nIs there any rule of thumb about learning rate when we switch to larger models? Or is it a process of trial and error?",
      "votes": null
    },
    {
      "id": "971225",
      "postDate": "08/15/2020 09:52:35",
      "content": "<p><a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a>, I just realized that my upvote of your comment vanished since someone downvoted your answer again:-)</p>",
      "rawMarkdown": "cpmp, I just realized that my upvote of your comment vanished since someone downvoted your answer again:-)",
      "votes": null
    },
    {
      "id": "971231",
      "postDate": "08/15/2020 09:58:53",
      "content": "<p>Don't bother.  every 6 months I get angry about it and comment it, but this is what it is.</p>",
      "rawMarkdown": "Don't bother.  every 6 months I get angry about it and comment it, but this is what it is.",
      "votes": null
    },
    {
      "id": "971234",
      "postDate": "08/15/2020 10:01:13",
      "content": "<p>what optimizer do you use?</p>\n<p>I'd start with Adam, and a simple scheduler among the predefined ones from pytorch.  I also find your lr to be very low.  I bet your models are underfitted.</p>",
      "rawMarkdown": "what optimizer do you use?\n\nI'd start with Adam, and a simple scheduler among the predefined ones from pytorch.  I also find your lr to be very low.  I bet your models are underfitted.",
      "votes": null
    },
    {
      "id": "971240",
      "postDate": "08/15/2020 10:10:16",
      "content": "<p>I use Adam. Yes, that's my feeling. That, being too precautions, I'd been <strong>over</strong>trying with <strong>over</strong>fitting all the time and as a result I <strong>under</strong>fitting. I tested this schema with no improvement:</p>\n<p>self.lr_start = 0.000005<br>\nself.lr_min = <strong>0.00005</strong><br>\nself.lr_ramp_ep = 5<br>\nself.lr_sus_ep = 0<br>\nself.lr_decay = 0.8 <br>\nself.lr_max = 0.00000125 * <strong>8</strong> * batch_size</p>\n<p>We are nearly at the edge of the time but I will try with larger rates. Thank you!</p>",
      "rawMarkdown": "I use Adam. Yes, that's my feeling. That, being too precautions, I'd been **over**trying with **over**fitting all the time and as a result I **under**fitting. I tested this schema with no improvement:\n\nself.lr_start = 0.000005\nself.lr_min = **0.00005**\nself.lr_ramp_ep = 5\nself.lr_sus_ep = 0\nself.lr_decay = 0.8 \nself.lr_max = 0.00000125 * **8** * batch_size\n\n We are nearly at the edge of the time but I will try with larger rates. Thank you!",
      "votes": null
    },
    {
      "id": "971537",
      "postDate": "08/15/2020 16:37:43",
      "content": "<p>Do you use early stopping, you would be able to detect when it starts overfitting and if its too early.</p>\n<p>For me lowering the learning rate was enough to achieve higher LB scores with large models, but I also have a very low batch-size of 5 (Due to using 8 Cores parallely)</p>",
      "rawMarkdown": "Do you use early stopping, you would be able to detect when it starts overfitting and if its too early.\n\nFor me lowering the learning rate was enough to achieve higher LB scores with large models, but I also have a very low batch-size of 5 (Due to using 8 Cores parallely)",
      "votes": null
    },
    {
      "id": "972753",
      "postDate": "08/16/2020 19:52:53",
      "content": "<p>Yes, I monitor the training process and use early stopping. I tried lower rates, higher rates, different batch sizes, more intense augmentations, added dropout in the head. All in vain so far. In average, every time the model learns slowly and then gets stuck in a local trap after 50% of epochs (12-14 out of 25). I think it's time to relax and switch to waiting mode. Looking forward to seeing pytorch solutions after the end of the competition. </p>",
      "rawMarkdown": "Yes, I monitor the training process and use early stopping. I tried lower rates, higher rates, different batch sizes, more intense augmentations, added dropout in the head. All in vain so far. In average, every time the model learns slowly and then gets stuck in a local trap after 50% of epochs (12-14 out of 25). I think it's time to relax and switch to waiting mode. Looking forward to seeing pytorch solutions after the end of the competition.",
      "votes": null
    },
    {
      "id": "972773",
      "postDate": "08/16/2020 20:12:59",
      "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> best of luck to you!</p>",
      "rawMarkdown": "dunklerwald best of luck to you!",
      "votes": null
    },
    {
      "id": "972811",
      "postDate": "08/16/2020 20:55:26",
      "content": "<p>Using on Plataeu LR Scheduler did not help either? That seems really weird. What is your initial LR?</p>",
      "rawMarkdown": "Using on Plataeu LR Scheduler did not help either? That seems really weird. What is your initial LR?",
      "votes": null
    },
    {
      "id": "974210",
      "postDate": "08/17/2020 19:45:42",
      "content": "<p><a href=\"https://www.kaggle.com/Signal\" target=\"_blank\">@Signal</a>, best of luck to you too!</p>",
      "rawMarkdown": "Signal, best of luck to you too!",
      "votes": null
    },
    {
      "id": "974214",
      "postDate": "08/17/2020 19:49:58",
      "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>, initial LR was 1e-04. I'm almost sure that's some stupid mistake I can't identify. </p>",
      "rawMarkdown": "aliabdin1, initial LR was 1e-04. I'm almost sure that's some stupid mistake I can't identify.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 970883,
      "author_name": "tkrsh09",
      "author_url": "",
      "post_date": "08/15/2020 00:10:51",
      "content": "<p>anything above b5 is probably overfitting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 970931,
      "author_name": "brianfeeny",
      "author_url": "",
      "post_date": "08/15/2020 02:37:39",
      "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> but you are already at 379 with .9622 so why bother, is this just a side model for fun?  I personally have not found that scheduler to work well, and I have tried those exact parameters.</p>\n<p>I would adjust the learning rate.  Why not run a few epochs with RLRoP, try .001, .0008, .0005, .0001, etc.  maybe patience 2 and decay it .8, .5 or .2 …………just try somethings and get an idea of where it wants to be.  What data are you using? Any upsampling?</p>",
      "votes": null,
      "replies": [
        {
          "id": 971122,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/15/2020 07:40:49",
          "content": "<p>I would assume 9622 is public kernel…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971198,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/15/2020 09:24:03",
          "content": "<p><a href=\"https://www.kaggle.com/brianfeeny\" target=\"_blank\">@brianfeeny</a> hey, <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a>  is right. I have an ensemble of my TF models (~0.950), based on the kernel of Chris. So yesterday I just mixed my ensemble with the infamous public ensemble (0.9619).</p>\n<p>I decided to use this competition to learn Pytorch. Looking at dramatic performance of my current Pytorch model, yes, it's more for fun. Actually I started with RLRoP and played out with different learning rates and patience numbers. Somehow a custom scheduler proved to be better in my case (learning process is more smooth and less overfitting).<br>\nI use external data but no upsampling. Minority upsampling didn't work for me also.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970966,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/15/2020 03:48:36",
      "content": "<p>A larger model has more parameters and has a higher chance of overfitting. When increasing to larger models, you need to increase regularization to prevent overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 971219,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/15/2020 09:48:18",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Chris, thank you. I learned a lot from you during last year and also in the course of this competition.<br>\nSpeaking of regularization, it could be:</p>\n<ul>\n<li>adding dropout layers in the tail</li>\n<li>heavier and more diverse augmentations<br>\nCorrect? What else should I try?</li>\n</ul>\n<p>Is there any rule of thumb about learning rate when we switch to larger models? Or is it a process of trial and error?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 971167,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/15/2020 08:29:54",
      "content": "<p>Your numbers are very low, try first to get better CV/LB with a simple model, say effnet b0.</p>",
      "votes": null,
      "replies": [
        {
          "id": 971205,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/15/2020 09:28:43",
          "content": "<p>B1 showed slightly better performance. Around LB 0.902. So I'm sure I miss some fundamental thing(s) about scaling\\regularization and still need to learn a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971225,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/15/2020 09:52:35",
          "content": "<p><a href=\"https://www.kaggle.com/cpmp\" target=\"_blank\">@cpmp</a>, I just realized that my upvote of your comment vanished since someone downvoted your answer again:-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971231,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2020 09:58:53",
          "content": "<p>Don't bother.  every 6 months I get angry about it and comment it, but this is what it is.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971234,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/15/2020 10:01:13",
          "content": "<p>what optimizer do you use?</p>\n<p>I'd start with Adam, and a simple scheduler among the predefined ones from pytorch.  I also find your lr to be very low.  I bet your models are underfitted.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971240,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/15/2020 10:10:16",
          "content": "<p>I use Adam. Yes, that's my feeling. That, being too precautions, I'd been <strong>over</strong>trying with <strong>over</strong>fitting all the time and as a result I <strong>under</strong>fitting. I tested this schema with no improvement:</p>\n<p>self.lr_start = 0.000005<br>\nself.lr_min = <strong>0.00005</strong><br>\nself.lr_ramp_ep = 5<br>\nself.lr_sus_ep = 0<br>\nself.lr_decay = 0.8 <br>\nself.lr_max = 0.00000125 * <strong>8</strong> * batch_size</p>\n<p>We are nearly at the edge of the time but I will try with larger rates. Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 971537,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "08/15/2020 16:37:43",
      "content": "<p>Do you use early stopping, you would be able to detect when it starts overfitting and if its too early.</p>\n<p>For me lowering the learning rate was enough to achieve higher LB scores with large models, but I also have a very low batch-size of 5 (Due to using 8 Cores parallely)</p>",
      "votes": null,
      "replies": [
        {
          "id": 972753,
          "author_name": "dunklerwald",
          "author_url": "",
          "post_date": "08/16/2020 19:52:53",
          "content": "<p>Yes, I monitor the training process and use early stopping. I tried lower rates, higher rates, different batch sizes, more intense augmentations, added dropout in the head. All in vain so far. In average, every time the model learns slowly and then gets stuck in a local trap after 50% of epochs (12-14 out of 25). I think it's time to relax and switch to waiting mode. Looking forward to seeing pytorch solutions after the end of the competition. </p>",
          "votes": null,
          "replies": [
            {
              "id": 972773,
              "author_name": "brianfeeny",
              "author_url": "",
              "post_date": "08/16/2020 20:12:59",
              "content": "<p><a href=\"https://www.kaggle.com/dunklerwald\" target=\"_blank\">@dunklerwald</a> best of luck to you!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 972811,
              "author_name": "aliabdin1",
              "author_url": "",
              "post_date": "08/16/2020 20:55:26",
              "content": "<p>Using on Plataeu LR Scheduler did not help either? That seems really weird. What is your initial LR?</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 974210,
              "author_name": "dunklerwald",
              "author_url": "",
              "post_date": "08/17/2020 19:45:42",
              "content": "<p><a href=\"https://www.kaggle.com/Signal\" target=\"_blank\">@Signal</a>, best of luck to you too!</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 974214,
              "author_name": "dunklerwald",
              "author_url": "",
              "post_date": "08/17/2020 19:49:58",
              "content": "<p><a href=\"https://www.kaggle.com/aliabdin1\" target=\"_blank\">@aliabdin1</a>, initial LR was 1e-04. I'm almost sure that's some stupid mistake I can't identify. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "970252": "So it seemed that I build a polished and consistent basic Pytorch model. I use GPU, mixed precision, BCE, custom scheduler, TTA, a moderate set of standard augmentations (rotation, shift, scale, brightness etc.) plus coarse dropout (max_holes=8, p=0.5). \n\nScheduler is\n```\nself.lr_start = 0.000005\nself.lr_min = 0.000001\nself.lr_ramp_ep = 5\nself.lr_sus_ep = 0\nself.lr_decay = 0.8 \nself.lr_max = 0.00000125 * 4 * batch_size\n```\n \nFor (effnet b3, size 256, batch size=128) my model delivers the following: CV 0.87843,  CV TTA 0.88557, LB **0.8942**. But things are not getting better with more complex models and larger image sizes.\n\n(effnet b4, size 256, batch size=96): CV 0.8772,  CV TTA 0.88237, LB **0.8903**\n(effnet b5, size 256, batch size=64): CV 0.85882,  CV TTA 0.85967,  LB **0.8761**\n(effnet b4, size 384, batch size=40): CV TTA 0.8628, LB **0.8821**\n\nI tried changing parameters of the scheduler (increase lr_min and lr_max, reduce lr_max by changing factor=4 to 2 or 1, etc.) - no big difference. Whatever I do, my model doesn't take advantage of larger effnet models and image sizes. What's more, it gets worse.\n\nBased on your experience, what could be wrong here? I understand that smaller batch size implies applying a lower learning rate, for example. But what else? E.g. if you switch from b3 to b4 or from 256 to 384, what would you adjust as a rule?",
    "970883": "anything above b5 is probably overfitting",
    "970931": "dunklerwald but you are already at 379 with .9622 so why bother, is this just a side model for fun?  I personally have not found that scheduler to work well, and I have tried those exact parameters.\n\nI would adjust the learning rate.  Why not run a few epochs with RLRoP, try .001, .0008, .0005, .0001, etc.  maybe patience 2 and decay it .8, .5 or .2 ............just try somethings and get an idea of where it wants to be.  What data are you using? Any upsampling?",
    "970966": "A larger model has more parameters and has a higher chance of overfitting. When increasing to larger models, you need to increase regularization to prevent overfitting.",
    "971122": "I would assume 9622 is public kernel...",
    "971167": "Your numbers are very low, try first to get better CV/LB with a simple model, say effnet b0.",
    "971198": "brianfeeny hey, @philippsinger  is right. I have an ensemble of my TF models (~0.950), based on the kernel of Chris. So yesterday I just mixed my ensemble with the infamous public ensemble (0.9619).\n\nI decided to use this competition to learn Pytorch. Looking at dramatic performance of my current Pytorch model, yes, it's more for fun. Actually I started with RLRoP and played out with different learning rates and patience numbers. Somehow a custom scheduler proved to be better in my case (learning process is more smooth and less overfitting).\nI use external data but no upsampling. Minority upsampling didn't work for me also.",
    "971205": "B1 showed slightly better performance. Around LB 0.902. So I'm sure I miss some fundamental thing(s) about scaling\\regularization and still need to learn a lot.",
    "971219": "cdeotte Chris, thank you. I learned a lot from you during last year and also in the course of this competition.\nSpeaking of regularization, it could be:\n- adding dropout layers in the tail\n- heavier and more diverse augmentations\nCorrect? What else should I try?\n\nIs there any rule of thumb about learning rate when we switch to larger models? Or is it a process of trial and error?",
    "971225": "cpmp, I just realized that my upvote of your comment vanished since someone downvoted your answer again:-)",
    "971231": "Don't bother.  every 6 months I get angry about it and comment it, but this is what it is.",
    "971234": "what optimizer do you use?\n\nI'd start with Adam, and a simple scheduler among the predefined ones from pytorch.  I also find your lr to be very low.  I bet your models are underfitted.",
    "971240": "I use Adam. Yes, that's my feeling. That, being too precautions, I'd been **over**trying with **over**fitting all the time and as a result I **under**fitting. I tested this schema with no improvement:\n\nself.lr_start = 0.000005\nself.lr_min = **0.00005**\nself.lr_ramp_ep = 5\nself.lr_sus_ep = 0\nself.lr_decay = 0.8 \nself.lr_max = 0.00000125 * **8** * batch_size\n\n We are nearly at the edge of the time but I will try with larger rates. Thank you!",
    "971537": "Do you use early stopping, you would be able to detect when it starts overfitting and if its too early.\n\nFor me lowering the learning rate was enough to achieve higher LB scores with large models, but I also have a very low batch-size of 5 (Due to using 8 Cores parallely)",
    "972753": "Yes, I monitor the training process and use early stopping. I tried lower rates, higher rates, different batch sizes, more intense augmentations, added dropout in the head. All in vain so far. In average, every time the model learns slowly and then gets stuck in a local trap after 50% of epochs (12-14 out of 25). I think it's time to relax and switch to waiting mode. Looking forward to seeing pytorch solutions after the end of the competition.",
    "972773": "dunklerwald best of luck to you!",
    "972811": "Using on Plataeu LR Scheduler did not help either? That seems really weird. What is your initial LR?",
    "974210": "Signal, best of luck to you too!",
    "974214": "aliabdin1, initial LR was 1e-04. I'm almost sure that's some stupid mistake I can't identify."
  },
  "source": "meta"
}