{
  "id": 22598,
  "title": "Did anyone succeed in training resnet-50 from scratch?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22598",
  "author_name": "",
  "post_date": "2016-08-01T15:43:33.220Z",
  "votes": null,
  "comment_count": 6,
  "views": 1741,
  "content": "<p>Hi All,</p>\n\n<p>I have tried to train resnet-50 (the largest can be handled by my GPU :( ) from scratch but without success. I can always get training loss drop to very low like 0.001, but the validation loss stays very high on 4~3.  The training time is about 7 hours on single fold.</p>\n\n<p>I know it is overfitting, but I am surprised that it overfitted so much. Maybe someone can give me some hints about why this happen? I use rotation, zooming some classical data augmentation techniques. Do I have to use pretrained weights?</p>\n\n<p>Thanks</p>",
  "messages": [
    {
      "id": "129673",
      "postDate": "08/01/2016 15:43:33",
      "content": "<p>Hi All,</p>\n\n<p>I have tried to train resnet-50 (the largest can be handled by my GPU :( ) from scratch but without success. I can always get training loss drop to very low like 0.001, but the validation loss stays very high on 4~3.  The training time is about 7 hours on single fold.</p>\n\n<p>I know it is overfitting, but I am surprised that it overfitted so much. Maybe someone can give me some hints about why this happen? I use rotation, zooming some classical data augmentation techniques. Do I have to use pretrained weights?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hi All,\r\n\r\nI have tried to train resnet-50 (the largest can be handled by my GPU :( ) from scratch but without success. I can always get training loss drop to very low like 0.001, but the validation loss stays very high on 4~3.  The training time is about 7 hours on single fold.\r\n\r\nI know it is overfitting, but I am surprised that it overfitted so much. Maybe someone can give me some hints about why this happen? I use rotation, zooming some classical data augmentation techniques. Do I have to use pretrained weights?\r\n\r\nThanks",
      "votes": null
    },
    {
      "id": "129676",
      "postDate": "08/01/2016 16:21:28",
      "content": "<p>I think there is simply not enough training data to create a model from scratch - we ended up using Facebook's pretrained model with Resnet 200.</p>\n\n<p>It's also possible that you are using a very high learning rate, so this may be something that you have to change (the facebook one defaults to 0.1, whereas something like 0.001 is more reasonable)</p>",
      "rawMarkdown": "I think there is simply not enough training data to create a model from scratch - we ended up using Facebook's pretrained model with Resnet 200.\r\n\r\nIt's also possible that you are using a very high learning rate, so this may be something that you have to change (the facebook one defaults to 0.1, whereas something like 0.001 is more reasonable)",
      "votes": null
    },
    {
      "id": "129677",
      "postDate": "08/01/2016 16:28:29",
      "content": "<p>I trained a variant of resnet50 in lasagne from scratch, alternating between augmented (rotated, shifted, flipped , zoomed using Keras's Image Data Generator) &amp; normal data. Also, in each image I would randomly zero out a block of 32x32. Also subtracted the mean RGB values from each image. This gave a leaderboard score of around 0.64</p>",
      "rawMarkdown": "I trained a variant of resnet50 in lasagne from scratch, alternating between augmented (rotated, shifted, flipped , zoomed using Keras's Image Data Generator) & normal data. Also, in each image I would randomly zero out a block of 32x32. Also subtracted the mean RGB values from each image. This gave a leaderboard score of around 0.64",
      "votes": null
    },
    {
      "id": "129687",
      "postDate": "08/01/2016 19:04:30",
      "content": "<p>Thanks for  reply\nI use very small initial learning rate 1e-5 to 1e-6.  Amount of data seems crucial for generalization.</p>",
      "rawMarkdown": "Thanks for  reply\r\nI use very small initial learning rate 1e-5 to 1e-6.  Amount of data seems crucial for generalization.",
      "votes": null
    },
    {
      "id": "129714",
      "postDate": "08/01/2016 22:59:12",
      "content": "<p>With a pre-trained ResNet-152 and no data augmentation i got 0.57 on the LB. I think there isn't enough data to fine tune large networks like this one.</p>",
      "rawMarkdown": "With a pre-trained ResNet-152 and no data augmentation i got 0.57 on the LB. I think there isn't enough data to fine tune large networks like this one.",
      "votes": null
    },
    {
      "id": "130059",
      "postDate": "08/03/2016 14:35:54",
      "content": "<p>I got 0.41 with a ResNet trained from scratch. </p>\n\n<p>From my notes:</p>\n\n<pre><code>128x128 Image size\nInitial filter num - 16\nInitial filter size - 5x5, stride 1\nNo maxpool after first filter\nFullPreActivation\nBatch size - 64\nL2 regularization - 0.00001 (less than paper)\nDropout after GlobalPoolLayer p=0.25\nADAM for 100 epoch - lr_schedule = {0:0.001, 60:0.0001, 80:0.00001}\nHeavy Train Augmentations\nTrans TTA\nMean-pixel centering\nProjection option\nIndividual (old local cv) accuracy - 99.9%\nLocal CV loss - 0.003972 -&gt; 0.023\nSubmission score - 0.41797 -&gt; 0.49150\nTime per epoch - 174.7 seconds\n</code></pre>",
      "rawMarkdown": "I got 0.41 with a ResNet trained from scratch. \r\n\r\nFrom my notes:\r\n\r\n    128x128 Image size\r\n    Initial filter num - 16\r\n    Initial filter size - 5x5, stride 1\r\n    No maxpool after first filter\r\n    FullPreActivation\r\n    Batch size - 64\r\n    L2 regularization - 0.00001 (less than paper)\r\n    Dropout after GlobalPoolLayer p=0.25\r\n    ADAM for 100 epoch - lr_schedule = {0:0.001, 60:0.0001, 80:0.00001}\r\n    Heavy Train Augmentations\r\n    Trans TTA\r\n    Mean-pixel centering\r\n    Projection option\r\n    Individual (old local cv) accuracy - 99.9%\r\n    Local CV loss - 0.003972 -> 0.023\r\n    Submission score - 0.41797 -> 0.49150\r\n    Time per epoch - 174.7 seconds",
      "votes": null
    },
    {
      "id": "130066",
      "postDate": "08/03/2016 14:58:57",
      "content": "<p>I too trained a Resnet from scratch. When I say resnet I mean a Pre-activation resnet from the 2016 paper. For the 32x32 case the code is here:</p>\n\n<p><a href=\"https://bitly.com/cifar10-resnet\">https://bitly.com/cifar10-resnet</a></p>\n\n<p>I trained a network with a  <code>depth = 9*n+2</code> for <code>n=5</code> and <code>n=6</code> and got a LB of 0.42</p>\n\n<p>I wanted to also try a Wide Residual Networks but was too busy at work etc. to devote much time unfortunately. The code for the 32x32 case is here:</p>\n\n<p><a href=\"https://bit.ly/cifar10-wide-resnet\">https://bit.ly/cifar10-wide-resnet</a></p>\n\n<p>Hope this helps! </p>",
      "rawMarkdown": "I too trained a Resnet from scratch. When I say resnet I mean a Pre-activation resnet from the 2016 paper. For the 32x32 case the code is here:\r\n\r\nhttps://bitly.com/cifar10-resnet\r\n\r\nI trained a network with a  `depth = 9*n+2` for `n=5` and `n=6` and got a LB of 0.42\r\n\r\nI wanted to also try a Wide Residual Networks but was too busy at work etc. to devote much time unfortunately. The code for the 32x32 case is here:\r\n\r\nhttps://bit.ly/cifar10-wide-resnet\r\n\r\nHope this helps!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129676,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "08/01/2016 16:21:28",
      "content": "<p>I think there is simply not enough training data to create a model from scratch - we ended up using Facebook's pretrained model with Resnet 200.</p>\n\n<p>It's also possible that you are using a very high learning rate, so this may be something that you have to change (the facebook one defaults to 0.1, whereas something like 0.001 is more reasonable)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129677,
      "author_name": "nikhil16",
      "author_url": "",
      "post_date": "08/01/2016 16:28:29",
      "content": "<p>I trained a variant of resnet50 in lasagne from scratch, alternating between augmented (rotated, shifted, flipped , zoomed using Keras's Image Data Generator) &amp; normal data. Also, in each image I would randomly zero out a block of 32x32. Also subtracted the mean RGB values from each image. This gave a leaderboard score of around 0.64</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129687,
      "author_name": "yangyang",
      "author_url": "",
      "post_date": "08/01/2016 19:04:30",
      "content": "<p>Thanks for  reply\nI use very small initial learning rate 1e-5 to 1e-6.  Amount of data seems crucial for generalization.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129714,
      "author_name": "rajgot",
      "author_url": "",
      "post_date": "08/01/2016 22:59:12",
      "content": "<p>With a pre-trained ResNet-152 and no data augmentation i got 0.57 on the LB. I think there isn't enough data to fine tune large networks like this one.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130059,
      "author_name": "florianm",
      "author_url": "",
      "post_date": "08/03/2016 14:35:54",
      "content": "<p>I got 0.41 with a ResNet trained from scratch. </p>\n\n<p>From my notes:</p>\n\n<pre><code>128x128 Image size\nInitial filter num - 16\nInitial filter size - 5x5, stride 1\nNo maxpool after first filter\nFullPreActivation\nBatch size - 64\nL2 regularization - 0.00001 (less than paper)\nDropout after GlobalPoolLayer p=0.25\nADAM for 100 epoch - lr_schedule = {0:0.001, 60:0.0001, 80:0.00001}\nHeavy Train Augmentations\nTrans TTA\nMean-pixel centering\nProjection option\nIndividual (old local cv) accuracy - 99.9%\nLocal CV loss - 0.003972 -&gt; 0.023\nSubmission score - 0.41797 -&gt; 0.49150\nTime per epoch - 174.7 seconds\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130066,
      "author_name": "krasul",
      "author_url": "",
      "post_date": "08/03/2016 14:58:57",
      "content": "<p>I too trained a Resnet from scratch. When I say resnet I mean a Pre-activation resnet from the 2016 paper. For the 32x32 case the code is here:</p>\n\n<p><a href=\"https://bitly.com/cifar10-resnet\">https://bitly.com/cifar10-resnet</a></p>\n\n<p>I trained a network with a  <code>depth = 9*n+2</code> for <code>n=5</code> and <code>n=6</code> and got a LB of 0.42</p>\n\n<p>I wanted to also try a Wide Residual Networks but was too busy at work etc. to devote much time unfortunately. The code for the 32x32 case is here:</p>\n\n<p><a href=\"https://bit.ly/cifar10-wide-resnet\">https://bit.ly/cifar10-wide-resnet</a></p>\n\n<p>Hope this helps! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129673": "Hi All,\r\n\r\nI have tried to train resnet-50 (the largest can be handled by my GPU :( ) from scratch but without success. I can always get training loss drop to very low like 0.001, but the validation loss stays very high on 4~3.  The training time is about 7 hours on single fold.\r\n\r\nI know it is overfitting, but I am surprised that it overfitted so much. Maybe someone can give me some hints about why this happen? I use rotation, zooming some classical data augmentation techniques. Do I have to use pretrained weights?\r\n\r\nThanks",
    "129676": "I think there is simply not enough training data to create a model from scratch - we ended up using Facebook's pretrained model with Resnet 200.\r\n\r\nIt's also possible that you are using a very high learning rate, so this may be something that you have to change (the facebook one defaults to 0.1, whereas something like 0.001 is more reasonable)",
    "129677": "I trained a variant of resnet50 in lasagne from scratch, alternating between augmented (rotated, shifted, flipped , zoomed using Keras's Image Data Generator) & normal data. Also, in each image I would randomly zero out a block of 32x32. Also subtracted the mean RGB values from each image. This gave a leaderboard score of around 0.64",
    "129687": "Thanks for  reply\r\nI use very small initial learning rate 1e-5 to 1e-6.  Amount of data seems crucial for generalization.",
    "129714": "With a pre-trained ResNet-152 and no data augmentation i got 0.57 on the LB. I think there isn't enough data to fine tune large networks like this one.",
    "130059": "I got 0.41 with a ResNet trained from scratch. \r\n\r\nFrom my notes:\r\n\r\n    128x128 Image size\r\n    Initial filter num - 16\r\n    Initial filter size - 5x5, stride 1\r\n    No maxpool after first filter\r\n    FullPreActivation\r\n    Batch size - 64\r\n    L2 regularization - 0.00001 (less than paper)\r\n    Dropout after GlobalPoolLayer p=0.25\r\n    ADAM for 100 epoch - lr_schedule = {0:0.001, 60:0.0001, 80:0.00001}\r\n    Heavy Train Augmentations\r\n    Trans TTA\r\n    Mean-pixel centering\r\n    Projection option\r\n    Individual (old local cv) accuracy - 99.9%\r\n    Local CV loss - 0.003972 -> 0.023\r\n    Submission score - 0.41797 -> 0.49150\r\n    Time per epoch - 174.7 seconds",
    "130066": "I too trained a Resnet from scratch. When I say resnet I mean a Pre-activation resnet from the 2016 paper. For the 32x32 case the code is here:\r\n\r\nhttps://bitly.com/cifar10-resnet\r\n\r\nI trained a network with a  `depth = 9*n+2` for `n=5` and `n=6` and got a LB of 0.42\r\n\r\nI wanted to also try a Wide Residual Networks but was too busy at work etc. to devote much time unfortunately. The code for the 32x32 case is here:\r\n\r\nhttps://bit.ly/cifar10-wide-resnet\r\n\r\nHope this helps!"
  },
  "source": "meta"
}