{
  "id": 167206,
  "title": "Does Over fitting mean model has enough parameters?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/167206",
  "author_name": "",
  "post_date": "2020-07-15T16:01:05.514792Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey all,\nI am having a small confusion. Assuming training on Imagenet, let's say we trained EfnetB3 with 99.9% final train accuracy and 90% final validation accuracy showing clear signs of overfitting. Does this mean there is no point in increasing the size of the model to EfnetB7 or B6 as we have enough parameters in B3 itself to represent the whole train dataset? </p>",
  "messages": [
    {
      "id": "930638",
      "postDate": "07/15/2020 16:01:05",
      "content": "<p>Hey all,\nI am having a small confusion. Assuming training on Imagenet, let's say we trained EfnetB3 with 99.9% final train accuracy and 90% final validation accuracy showing clear signs of overfitting. Does this mean there is no point in increasing the size of the model to EfnetB7 or B6 as we have enough parameters in B3 itself to represent the whole train dataset? </p>",
      "rawMarkdown": "Hey all,\nI am having a small confusion. Assuming training on Imagenet, let's say we trained EfnetB3 with 99.9% final train accuracy and 90% final validation accuracy showing clear signs of overfitting. Does this mean there is no point in increasing the size of the model to EfnetB7 or B6 as we have enough parameters in B3 itself to represent the whole train dataset?",
      "votes": null
    },
    {
      "id": "930647",
      "postDate": "07/15/2020 16:10:28",
      "content": "<p>I wouldn't say so, as you can very well fit a B0 to 100% training accuracy in this dataset, specially if you remove all regularisation (like the implicit drop_connect in EfficientNet), anneal your learning rate, and train long enough.</p>\n\n<p>Even a relatively small deep learning model will have enough parameters to fully fit our small dataset :/</p>\n\n<p>btw, that's probably true even if you randomly swap all the training labels in the dataset ;)\nThere was a famous paper on that (hopefully someone here will know that paper)</p>",
      "rawMarkdown": "I wouldn't say so, as you can very well fit a B0 to 100% training accuracy in this dataset, specially if you remove all regularisation (like the implicit drop_connect in EfficientNet), anneal your learning rate, and train long enough.\n\nEven a relatively small deep learning model will have enough parameters to fully fit our small dataset :/\n\nbtw, that's probably true even if you randomly swap all the training labels in the dataset ;)\nThere was a famous paper on that (hopefully someone here will know that paper)",
      "votes": null
    },
    {
      "id": "930668",
      "postDate": "07/15/2020 16:23:57",
      "content": "<p>Short answer: Yes and No</p>\n\n<p>Long answer: Depends. If I was you, my first port of call would be to retain the B3 model and try to reduce overfitting using multiple techniques (increasing dropouts, doing multi-fold validation, performing augmentations, increasing data size set, training on multiple resolution of images, etc.). Migrating to B7 is a different question and the answer to that can be obtained by doing it - often times there are restrictions such as compute or memory availability which makes migrating to a larger model more difficult. However, if that's not the case, then migrating to a large model typically always helps. </p>",
      "rawMarkdown": "Short answer: Yes and No\n\nLong answer: Depends. If I was you, my first port of call would be to retain the B3 model and try to reduce overfitting using multiple techniques (increasing dropouts, doing multi-fold validation, performing augmentations, increasing data size set, training on multiple resolution of images, etc.). Migrating to B7 is a different question and the answer to that can be obtained by doing it - often times there are restrictions such as compute or memory availability which makes migrating to a larger model more difficult. However, if that's not the case, then migrating to a large model typically always helps.",
      "votes": null
    },
    {
      "id": "930688",
      "postDate": "07/15/2020 16:41:47",
      "content": "<p>I think <a href=\"/hmendonca\">@hmendonca</a> refers to this paper : <a href=\"https://arxiv.org/pdf/1611.03530.pdf\">https://arxiv.org/pdf/1611.03530.pdf</a></p>",
      "rawMarkdown": "I think @hmendonca refers to this paper : https://arxiv.org/pdf/1611.03530.pdf",
      "votes": null
    },
    {
      "id": "930861",
      "postDate": "07/15/2020 18:54:28",
      "content": "<p>Thanks…. :)</p>",
      "rawMarkdown": "Thanks.... :)",
      "votes": null
    },
    {
      "id": "930875",
      "postDate": "07/15/2020 19:13:06",
      "content": "<p>Thanks for the answer :). <br>\nI am still a bit confused actually. So, can B0 get 100% train on Imagenet? If yes doesn't it mean the model can find all the necessary features required for classifying the training data?  So generalizing these features itself should be helping for better datasets right? Also, I have seen cases of pruning where you initially train with a large model then prune &gt;50% of its parameters maintaining same val accuracy. <br>\nTo put it more clearly, I would initially train with B7 then prune to a model with the same parameters as B0 with the same accuracy. Why couldn't this model be learned from the B0 number of parameters itself? Are there any theoretical explanations?</p>",
      "rawMarkdown": "Thanks for the answer :). \nI am still a bit confused actually. So, can B0 get 100% train on Imagenet? If yes doesn't it mean the model can find all the necessary features required for classifying the training data?  So generalizing these features itself should be helping for better datasets right? Also, I have seen cases of pruning where you initially train with a large model then prune &gt;50% of its parameters maintaining same val accuracy. \nTo put it more clearly, I would initially train with B7 then prune to a model with the same parameters as B0 with the same accuracy. Why couldn't this model be learned from the B0 number of parameters itself? Are there any theoretical explanations?",
      "votes": null
    },
    {
      "id": "930925",
      "postDate": "07/15/2020 20:17:11",
      "content": "<p>Neural Networks are universal approximators, so there is a NN approaching any (reasonable) function (see here for a nice explanation <a href=\"http://neuralnetworksanddeeplearning.com/chap4.html\">http://neuralnetworksanddeeplearning.com/chap4.html</a> ).\nThe hard part is how to find the smallest approximating NN for a specific function or even harder for an unknown function  (on which you know only a few points).\nRecent convnets have enough capacity to approximate a lot of datasets (even with random labels). What is not fully understood is why do they generalize so well, i.e are able to predict correctly on unseen data and why don't they just learn to overfit. Even a simple one layer MLP can overfit CIFAR-10 as shown in the above paper.</p>\n\n<p>Training bigger networks usually give better results, this is an empirical fact. Most of modern architectures are over-parametrized (they have more parameters than necessary), after training they have redundancy and dead neurons but still performs better. </p>\n\n<p>Pruning is a way to get closer to the smallest NN that solves your problem, but as it's not well understood theoretically, I think there is no known answer to your questions yet. \nWe simply don't know any other way that CNN + SGD + over parametrization to get to a good solution for images (and all the training tricks we are all trying to use here to get a better generalization, like dropout, cutmix etc...).</p>\n\n<p>Would be happy to have other viewpoints on this!</p>",
      "rawMarkdown": "Neural Networks are universal approximators, so there is a NN approaching any (reasonable) function (see here for a nice explanation http://neuralnetworksanddeeplearning.com/chap4.html ).\nThe hard part is how to find the smallest approximating NN for a specific function or even harder for an unknown function  (on which you know only a few points).\nRecent convnets have enough capacity to approximate a lot of datasets (even with random labels). What is not fully understood is why do they generalize so well, i.e are able to predict correctly on unseen data and why don't they just learn to overfit. Even a simple one layer MLP can overfit CIFAR-10 as shown in the above paper.\n\nTraining bigger networks usually give better results, this is an empirical fact. Most of modern architectures are over-parametrized (they have more parameters than necessary), after training they have redundancy and dead neurons but still performs better. \n\nPruning is a way to get closer to the smallest NN that solves your problem, but as it's not well understood theoretically, I think there is no known answer to your questions yet. \nWe simply don't know any other way that CNN + SGD + over parametrization to get to a good solution for images (and all the training tricks we are all trying to use here to get a better generalization, like dropout, cutmix etc...).\n\nWould be happy to have other viewpoints on this!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 930647,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/15/2020 16:10:28",
      "content": "<p>I wouldn't say so, as you can very well fit a B0 to 100% training accuracy in this dataset, specially if you remove all regularisation (like the implicit drop_connect in EfficientNet), anneal your learning rate, and train long enough.</p>\n\n<p>Even a relatively small deep learning model will have enough parameters to fully fit our small dataset :/</p>\n\n<p>btw, that's probably true even if you randomly swap all the training labels in the dataset ;)\nThere was a famous paper on that (hopefully someone here will know that paper)</p>",
      "votes": null,
      "replies": [
        {
          "id": 930688,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/15/2020 16:41:47",
          "content": "<p>I think <a href=\"/hmendonca\">@hmendonca</a> refers to this paper : <a href=\"https://arxiv.org/pdf/1611.03530.pdf\">https://arxiv.org/pdf/1611.03530.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930875,
          "author_name": "josealways123",
          "author_url": "",
          "post_date": "07/15/2020 19:13:06",
          "content": "<p>Thanks for the answer :). <br>\nI am still a bit confused actually. So, can B0 get 100% train on Imagenet? If yes doesn't it mean the model can find all the necessary features required for classifying the training data?  So generalizing these features itself should be helping for better datasets right? Also, I have seen cases of pruning where you initially train with a large model then prune &gt;50% of its parameters maintaining same val accuracy. <br>\nTo put it more clearly, I would initially train with B7 then prune to a model with the same parameters as B0 with the same accuracy. Why couldn't this model be learned from the B0 number of parameters itself? Are there any theoretical explanations?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 930925,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/15/2020 20:17:11",
          "content": "<p>Neural Networks are universal approximators, so there is a NN approaching any (reasonable) function (see here for a nice explanation <a href=\"http://neuralnetworksanddeeplearning.com/chap4.html\">http://neuralnetworksanddeeplearning.com/chap4.html</a> ).\nThe hard part is how to find the smallest approximating NN for a specific function or even harder for an unknown function  (on which you know only a few points).\nRecent convnets have enough capacity to approximate a lot of datasets (even with random labels). What is not fully understood is why do they generalize so well, i.e are able to predict correctly on unseen data and why don't they just learn to overfit. Even a simple one layer MLP can overfit CIFAR-10 as shown in the above paper.</p>\n\n<p>Training bigger networks usually give better results, this is an empirical fact. Most of modern architectures are over-parametrized (they have more parameters than necessary), after training they have redundancy and dead neurons but still performs better. </p>\n\n<p>Pruning is a way to get closer to the smallest NN that solves your problem, but as it's not well understood theoretically, I think there is no known answer to your questions yet. \nWe simply don't know any other way that CNN + SGD + over parametrization to get to a good solution for images (and all the training tricks we are all trying to use here to get a better generalization, like dropout, cutmix etc...).</p>\n\n<p>Would be happy to have other viewpoints on this!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 930668,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "07/15/2020 16:23:57",
      "content": "<p>Short answer: Yes and No</p>\n\n<p>Long answer: Depends. If I was you, my first port of call would be to retain the B3 model and try to reduce overfitting using multiple techniques (increasing dropouts, doing multi-fold validation, performing augmentations, increasing data size set, training on multiple resolution of images, etc.). Migrating to B7 is a different question and the answer to that can be obtained by doing it - often times there are restrictions such as compute or memory availability which makes migrating to a larger model more difficult. However, if that's not the case, then migrating to a large model typically always helps. </p>",
      "votes": null,
      "replies": [
        {
          "id": 930861,
          "author_name": "josealways123",
          "author_url": "",
          "post_date": "07/15/2020 18:54:28",
          "content": "<p>Thanks…. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "930638": "Hey all,\nI am having a small confusion. Assuming training on Imagenet, let's say we trained EfnetB3 with 99.9% final train accuracy and 90% final validation accuracy showing clear signs of overfitting. Does this mean there is no point in increasing the size of the model to EfnetB7 or B6 as we have enough parameters in B3 itself to represent the whole train dataset?",
    "930647": "I wouldn't say so, as you can very well fit a B0 to 100% training accuracy in this dataset, specially if you remove all regularisation (like the implicit drop_connect in EfficientNet), anneal your learning rate, and train long enough.\n\nEven a relatively small deep learning model will have enough parameters to fully fit our small dataset :/\n\nbtw, that's probably true even if you randomly swap all the training labels in the dataset ;)\nThere was a famous paper on that (hopefully someone here will know that paper)",
    "930668": "Short answer: Yes and No\n\nLong answer: Depends. If I was you, my first port of call would be to retain the B3 model and try to reduce overfitting using multiple techniques (increasing dropouts, doing multi-fold validation, performing augmentations, increasing data size set, training on multiple resolution of images, etc.). Migrating to B7 is a different question and the answer to that can be obtained by doing it - often times there are restrictions such as compute or memory availability which makes migrating to a larger model more difficult. However, if that's not the case, then migrating to a large model typically always helps.",
    "930688": "I think @hmendonca refers to this paper : https://arxiv.org/pdf/1611.03530.pdf",
    "930861": "Thanks.... :)",
    "930875": "Thanks for the answer :). \nI am still a bit confused actually. So, can B0 get 100% train on Imagenet? If yes doesn't it mean the model can find all the necessary features required for classifying the training data?  So generalizing these features itself should be helping for better datasets right? Also, I have seen cases of pruning where you initially train with a large model then prune &gt;50% of its parameters maintaining same val accuracy. \nTo put it more clearly, I would initially train with B7 then prune to a model with the same parameters as B0 with the same accuracy. Why couldn't this model be learned from the B0 number of parameters itself? Are there any theoretical explanations?",
    "930925": "Neural Networks are universal approximators, so there is a NN approaching any (reasonable) function (see here for a nice explanation http://neuralnetworksanddeeplearning.com/chap4.html ).\nThe hard part is how to find the smallest approximating NN for a specific function or even harder for an unknown function  (on which you know only a few points).\nRecent convnets have enough capacity to approximate a lot of datasets (even with random labels). What is not fully understood is why do they generalize so well, i.e are able to predict correctly on unseen data and why don't they just learn to overfit. Even a simple one layer MLP can overfit CIFAR-10 as shown in the above paper.\n\nTraining bigger networks usually give better results, this is an empirical fact. Most of modern architectures are over-parametrized (they have more parameters than necessary), after training they have redundancy and dead neurons but still performs better. \n\nPruning is a way to get closer to the smallest NN that solves your problem, but as it's not well understood theoretically, I think there is no known answer to your questions yet. \nWe simply don't know any other way that CNN + SGD + over parametrization to get to a good solution for images (and all the training tricks we are all trying to use here to get a better generalization, like dropout, cutmix etc...).\n\nWould be happy to have other viewpoints on this!"
  },
  "source": "meta"
}