{
  "id": 135059,
  "title": "NASNetLarge fails to improve validation accuracy.",
  "url": "/competitions/flower-classification-with-tpus/discussion/135059",
  "author_name": "",
  "post_date": "2020-03-11T20:32:07.947210400Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi, after using a fairly successful DenseNet + Efficient Net ensemble, I have been experimenting with NASNetLarge, randomly initialized to use 512x512 images. (Had to reduce the batch size to 12 otherwise it gives me a resource error. )</p>\n\n<p>Using a learning rate scheduler to reduce learning rate after training for 5 epochs with a randomized learning rate.</p>\n\n<p>Training accuracy improves but validation accuracy remains very poor and does not improve after the first 5ish epochs. \nDid anyone have more success with it or knows why it's failing to generalize? </p>",
  "messages": [
    {
      "id": "769356",
      "postDate": "03/11/2020 20:32:07",
      "content": "<p>Hi, after using a fairly successful DenseNet + Efficient Net ensemble, I have been experimenting with NASNetLarge, randomly initialized to use 512x512 images. (Had to reduce the batch size to 12 otherwise it gives me a resource error. )</p>\n\n<p>Using a learning rate scheduler to reduce learning rate after training for 5 epochs with a randomized learning rate.</p>\n\n<p>Training accuracy improves but validation accuracy remains very poor and does not improve after the first 5ish epochs. \nDid anyone have more success with it or knows why it's failing to generalize? </p>",
      "rawMarkdown": "Hi, after using a fairly successful DenseNet + Efficient Net ensemble, I have been experimenting with NASNetLarge, randomly initialized to use 512x512 images. (Had to reduce the batch size to 12 otherwise it gives me a resource error. )\n\nUsing a learning rate scheduler to reduce learning rate after training for 5 epochs with a randomized learning rate.\n\nTraining accuracy improves but validation accuracy remains very poor and does not improve after the first 5ish epochs. \nDid anyone have more success with it or knows why it's failing to generalize?",
      "votes": null
    },
    {
      "id": "769465",
      "postDate": "03/11/2020 23:48:27",
      "content": "<p>Such large model, randomly initialized, trained on a small dataset is going to overfit.</p>",
      "rawMarkdown": "Such large model, randomly initialized, trained on a small dataset is going to overfit.",
      "votes": null
    },
    {
      "id": "771141",
      "postDate": "03/13/2020 19:44:27",
      "content": "<p>Very possible. I hadn't fully considered how large the model is. Unfortunately it appears to have issues loading the image-net weights at the moment. </p>",
      "rawMarkdown": "Very possible. I hadn't fully considered how large the model is. Unfortunately it appears to have issues loading the image-net weights at the moment.",
      "votes": null
    },
    {
      "id": "771240",
      "postDate": "03/13/2020 22:46:38",
      "content": "<p><a href=\"/atamazian\">@atamazian</a> published a <a href=\"https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge\">NASNETLarge notebook.</a>. There were a couple of workarounds needed to load it.</p>",
      "rawMarkdown": "atamazian published a [NASNETLarge notebook.](https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge). There were a couple of workarounds needed to load it.",
      "votes": null
    },
    {
      "id": "771459",
      "postDate": "03/14/2020 07:56:04",
      "content": "<p>Hello Martin, Your lecture \"Tensorflow and deep learning - without a PhD\" was amazing, big fan!!!.\nSince the validation accuracy is low, we can conclude the model is Over-fitting? </p>\n\n<p>I am getting low val accuracy with my notebook, tried regularization, and dropouts, after initializing VGG16 weights</p>\n\n<p><a href=\"https://www.kaggle.com/greatcodes/transfer-learning-using-vgg16\">https://www.kaggle.com/greatcodes/transfer-learning-using-vgg16</a> </p>\n\n<p>What do you suggest?</p>",
      "rawMarkdown": "Hello Martin, Your lecture \"Tensorflow and deep learning - without a PhD\" was amazing, big fan!!!.\nSince the validation accuracy is low, we can conclude the model is Over-fitting? \n\nI am getting low val accuracy with my notebook, tried regularization, and dropouts, after initializing VGG16 weights\n\nhttps://www.kaggle.com/greatcodes/transfer-learning-using-vgg16 \n\nWhat do you suggest?",
      "votes": null
    },
    {
      "id": "775672",
      "postDate": "03/16/2020 23:48:44",
      "content": "<p>Stop using VGG16 ! EfficientNet is the new state of the art and comes in various sizes. Other good alternatives: DenseNet, NasNet.</p>\n\n<p>Thank you for your kind comments on TF w/o a PhD 😊</p>",
      "rawMarkdown": "Stop using VGG16 ! EfficientNet is the new state of the art and comes in various sizes. Other good alternatives: DenseNet, NasNet.\n\nThank you for your kind comments on TF w/o a PhD 😊",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 769465,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "03/11/2020 23:48:27",
      "content": "<p>Such large model, randomly initialized, trained on a small dataset is going to overfit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 771141,
          "author_name": "sebastiankoenig",
          "author_url": "",
          "post_date": "03/13/2020 19:44:27",
          "content": "<p>Very possible. I hadn't fully considered how large the model is. Unfortunately it appears to have issues loading the image-net weights at the moment. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 771240,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/13/2020 22:46:38",
          "content": "<p><a href=\"/atamazian\">@atamazian</a> published a <a href=\"https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge\">NASNETLarge notebook.</a>. There were a couple of workarounds needed to load it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 771459,
          "author_name": "greatcodes",
          "author_url": "",
          "post_date": "03/14/2020 07:56:04",
          "content": "<p>Hello Martin, Your lecture \"Tensorflow and deep learning - without a PhD\" was amazing, big fan!!!.\nSince the validation accuracy is low, we can conclude the model is Over-fitting? </p>\n\n<p>I am getting low val accuracy with my notebook, tried regularization, and dropouts, after initializing VGG16 weights</p>\n\n<p><a href=\"https://www.kaggle.com/greatcodes/transfer-learning-using-vgg16\">https://www.kaggle.com/greatcodes/transfer-learning-using-vgg16</a> </p>\n\n<p>What do you suggest?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775672,
          "author_name": "mgorner",
          "author_url": "",
          "post_date": "03/16/2020 23:48:44",
          "content": "<p>Stop using VGG16 ! EfficientNet is the new state of the art and comes in various sizes. Other good alternatives: DenseNet, NasNet.</p>\n\n<p>Thank you for your kind comments on TF w/o a PhD 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "769356": "Hi, after using a fairly successful DenseNet + Efficient Net ensemble, I have been experimenting with NASNetLarge, randomly initialized to use 512x512 images. (Had to reduce the batch size to 12 otherwise it gives me a resource error. )\n\nUsing a learning rate scheduler to reduce learning rate after training for 5 epochs with a randomized learning rate.\n\nTraining accuracy improves but validation accuracy remains very poor and does not improve after the first 5ish epochs. \nDid anyone have more success with it or knows why it's failing to generalize?",
    "769465": "Such large model, randomly initialized, trained on a small dataset is going to overfit.",
    "771141": "Very possible. I hadn't fully considered how large the model is. Unfortunately it appears to have issues loading the image-net weights at the moment.",
    "771240": "atamazian published a [NASNETLarge notebook.](https://www.kaggle.com/atamazian/100-flowers-on-tpu-with-nasnetlarge). There were a couple of workarounds needed to load it.",
    "771459": "Hello Martin, Your lecture \"Tensorflow and deep learning - without a PhD\" was amazing, big fan!!!.\nSince the validation accuracy is low, we can conclude the model is Over-fitting? \n\nI am getting low val accuracy with my notebook, tried regularization, and dropouts, after initializing VGG16 weights\n\nhttps://www.kaggle.com/greatcodes/transfer-learning-using-vgg16 \n\nWhat do you suggest?",
    "775672": "Stop using VGG16 ! EfficientNet is the new state of the art and comes in various sizes. Other good alternatives: DenseNet, NasNet.\n\nThank you for your kind comments on TF w/o a PhD 😊"
  },
  "source": "meta"
}