{
  "id": 217967,
  "title": "Larger models, lower score?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/217967",
  "author_name": "",
  "post_date": "2021-02-09T02:09:59.222617Z",
  "votes": 7,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Like many of you I have been trying different model architectures in order to get a better score. Also like most of you, my first port of call was efficientnet.</p>\n<p>I noticed that b0 scored a little worse than b3, but after that the scores either stayed the same or decreased. Again, many of you have made the same observations.</p>\n<p>This is a bit strange to me, since whenever I read a paper, say <a href=\"url\" target=\"_blank\">https://arxiv.org/abs/2010.01412</a>, I notice that they <em>always</em> get better scores with larger models. The best score is always achieved by Effnet-L2 or some other behemoth.</p>\n<p>Is there something we are doing wrong? Do you need more computing power or something in order to get the most out of the really large models?</p>\n<p>My suspicion is that large batch sizes become more important for larger models but it's just a hunch.</p>",
  "messages": [
    {
      "id": "1192145",
      "postDate": "02/09/2021 02:09:59",
      "content": "<p>Like many of you I have been trying different model architectures in order to get a better score. Also like most of you, my first port of call was efficientnet.</p>\n<p>I noticed that b0 scored a little worse than b3, but after that the scores either stayed the same or decreased. Again, many of you have made the same observations.</p>\n<p>This is a bit strange to me, since whenever I read a paper, say <a href=\"url\" target=\"_blank\">https://arxiv.org/abs/2010.01412</a>, I notice that they <em>always</em> get better scores with larger models. The best score is always achieved by Effnet-L2 or some other behemoth.</p>\n<p>Is there something we are doing wrong? Do you need more computing power or something in order to get the most out of the really large models?</p>\n<p>My suspicion is that large batch sizes become more important for larger models but it's just a hunch.</p>",
      "rawMarkdown": "Like many of you I have been trying different model architectures in order to get a better score. Also like most of you, my first port of call was efficientnet.\n\nI noticed that b0 scored a little worse than b3, but after that the scores either stayed the same or decreased. Again, many of you have made the same observations.\n\nThis is a bit strange to me, since whenever I read a paper, say [https://arxiv.org/abs/2010.01412](url), I notice that they *always* get better scores with larger models. The best score is always achieved by Effnet-L2 or some other behemoth.\n\nIs there something we are doing wrong? Do you need more computing power or something in order to get the most out of the really large models?\n\nMy suspicion is that large batch sizes become more important for larger models but it's just a hunch.",
      "votes": null
    },
    {
      "id": "1192270",
      "postDate": "02/09/2021 04:18:21",
      "content": "<p>Hello!</p>\n<p>Not an expert here, but I have a strong feeling that due to the label noise (which is, again, one of the main issues of this competition) larger models tend more to memorize the noise rather than generalize until you use really hard augmentations or use the larger model for distillation from smaller model(s). </p>",
      "rawMarkdown": "Hello!\n\nNot an expert here, but I have a strong feeling that due to the label noise (which is, again, one of the main issues of this competition) larger models tend more to memorize the noise rather than generalize until you use really hard augmentations or use the larger model for distillation from smaller model(s).",
      "votes": null
    },
    {
      "id": "1193331",
      "postDate": "02/09/2021 15:19:29",
      "content": "<p>For me large model may increase CV, but LB decreases</p>",
      "rawMarkdown": "For me large model may increase CV, but LB decreases",
      "votes": null
    },
    {
      "id": "1193746",
      "postDate": "02/09/2021 20:45:07",
      "content": "<p>Also larger models tend to overfit the data faster, make sure to detect it using e.g. early-stopping.<br>\nDont forget to use regularisation techniques like e.g. weight-decay for larger models.</p>",
      "rawMarkdown": "Also larger models tend to overfit the data faster, make sure to detect it using e.g. early-stopping.\nDont forget to use regularisation techniques like e.g. weight-decay for larger models.",
      "votes": null
    },
    {
      "id": "1194796",
      "postDate": "02/10/2021 11:26:36",
      "content": "<p>In this case, for me larger models work better. Looks like might be because of  augmentation techniques I have used in the training. </p>",
      "rawMarkdown": "In this case, for me larger models work better. Looks like might be because of  augmentation techniques I have used in the training.",
      "votes": null
    },
    {
      "id": "1194818",
      "postDate": "02/10/2021 11:40:15",
      "content": "<p>Their is no rule of thumbs </p>\n<p>I got better results with B7 than B5 and B6.  But training process is important.  You wouldn't  train B7 and B4 with the same settings. </p>",
      "rawMarkdown": "Their is no rule of thumbs \n\nI got better results with B7 than B5 and B6.  But training process is important.  You wouldn't  train B7 and B4 with the same settings.",
      "votes": null
    },
    {
      "id": "1209585",
      "postDate": "02/19/2021 00:43:05",
      "content": "<p>That's interesting. How do you select new hyperparameters for larger models though? Is it a matter of decreasing learning rate? Increasing batch size?</p>\n<p>Larger models have more parameters, so they are more prone to overfitting, so presumably a lower learning rate and larger batch size? More regularization also?</p>\n<p>Larger models are also more complex, so again, a lower learning rate is probably less likely to have parameters diverging to unpleasant locations. </p>\n<p>Is that about right?</p>",
      "rawMarkdown": "That's interesting. How do you select new hyperparameters for larger models though? Is it a matter of decreasing learning rate? Increasing batch size?\n\nLarger models have more parameters, so they are more prone to overfitting, so presumably a lower learning rate and larger batch size? More regularization also?\n\nLarger models are also more complex, so again, a lower learning rate is probably less likely to have parameters diverging to unpleasant locations. \n\nIs that about right?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1192270,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "02/09/2021 04:18:21",
      "content": "<p>Hello!</p>\n<p>Not an expert here, but I have a strong feeling that due to the label noise (which is, again, one of the main issues of this competition) larger models tend more to memorize the noise rather than generalize until you use really hard augmentations or use the larger model for distillation from smaller model(s). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1193746,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "02/09/2021 20:45:07",
          "content": "<p>Also larger models tend to overfit the data faster, make sure to detect it using e.g. early-stopping.<br>\nDont forget to use regularisation techniques like e.g. weight-decay for larger models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1193331,
      "author_name": "hanson0910",
      "author_url": "",
      "post_date": "02/09/2021 15:19:29",
      "content": "<p>For me large model may increase CV, but LB decreases</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1194796,
      "author_name": "vickygoyal",
      "author_url": "",
      "post_date": "02/10/2021 11:26:36",
      "content": "<p>In this case, for me larger models work better. Looks like might be because of  augmentation techniques I have used in the training. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1194818,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "02/10/2021 11:40:15",
      "content": "<p>Their is no rule of thumbs </p>\n<p>I got better results with B7 than B5 and B6.  But training process is important.  You wouldn't  train B7 and B4 with the same settings. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1209585,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "02/19/2021 00:43:05",
          "content": "<p>That's interesting. How do you select new hyperparameters for larger models though? Is it a matter of decreasing learning rate? Increasing batch size?</p>\n<p>Larger models have more parameters, so they are more prone to overfitting, so presumably a lower learning rate and larger batch size? More regularization also?</p>\n<p>Larger models are also more complex, so again, a lower learning rate is probably less likely to have parameters diverging to unpleasant locations. </p>\n<p>Is that about right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1192145": "Like many of you I have been trying different model architectures in order to get a better score. Also like most of you, my first port of call was efficientnet.\n\nI noticed that b0 scored a little worse than b3, but after that the scores either stayed the same or decreased. Again, many of you have made the same observations.\n\nThis is a bit strange to me, since whenever I read a paper, say [https://arxiv.org/abs/2010.01412](url), I notice that they *always* get better scores with larger models. The best score is always achieved by Effnet-L2 or some other behemoth.\n\nIs there something we are doing wrong? Do you need more computing power or something in order to get the most out of the really large models?\n\nMy suspicion is that large batch sizes become more important for larger models but it's just a hunch.",
    "1192270": "Hello!\n\nNot an expert here, but I have a strong feeling that due to the label noise (which is, again, one of the main issues of this competition) larger models tend more to memorize the noise rather than generalize until you use really hard augmentations or use the larger model for distillation from smaller model(s).",
    "1193331": "For me large model may increase CV, but LB decreases",
    "1193746": "Also larger models tend to overfit the data faster, make sure to detect it using e.g. early-stopping.\nDont forget to use regularisation techniques like e.g. weight-decay for larger models.",
    "1194796": "In this case, for me larger models work better. Looks like might be because of  augmentation techniques I have used in the training.",
    "1194818": "Their is no rule of thumbs \n\nI got better results with B7 than B5 and B6.  But training process is important.  You wouldn't  train B7 and B4 with the same settings.",
    "1209585": "That's interesting. How do you select new hyperparameters for larger models though? Is it a matter of decreasing learning rate? Increasing batch size?\n\nLarger models have more parameters, so they are more prone to overfitting, so presumably a lower learning rate and larger batch size? More regularization also?\n\nLarger models are also more complex, so again, a lower learning rate is probably less likely to have parameters diverging to unpleasant locations. \n\nIs that about right?"
  },
  "source": "meta"
}