{
  "id": 40664,
  "title": "Single model, single crop Top LB score?",
  "url": "/competitions/cdiscount-image-classification-challenge/discussion/40664",
  "author_name": "",
  "post_date": "2017-10-06T08:03:16.553106800Z",
  "votes": 5,
  "comment_count": 35,
  "views": 0,
  "content": "<p>As above.</p>\n\n<p>Mine: Resnet50 - single 160x160 crop  - 0.62316 LB</p>",
  "messages": [
    {
      "id": "228253",
      "postDate": "10/06/2017 08:03:16",
      "content": "<p>As above.</p>\n\n<p>Mine: Resnet50 - single 160x160 crop  - 0.62316 LB</p>",
      "rawMarkdown": "As above.\n\nMine: Resnet50 - single 160x160 crop  - 0.62316 LB",
      "votes": null
    },
    {
      "id": "228276",
      "postDate": "10/06/2017 09:30:00",
      "content": "<p>My SE-ResNet-50 model scored 0.64358, although I resized the test images from 180x180 to 160x160 so technically speaking I did not use a crop.</p>\n\n<p>With 10 crops on each test image, the LB score increased to 0.68159.</p>",
      "rawMarkdown": "My SE-ResNet-50 model scored 0.64358, although I resized the test images from 180x180 to 160x160 so technically speaking I did not use a crop.\n\nWith 10 crops on each test image, the LB score increased to 0.68159.",
      "votes": null
    },
    {
      "id": "228285",
      "postDate": "10/06/2017 10:02:06",
      "content": "<p>To clarify: the model I used is SE-ResNet-50 with a global average pooling layer at the end, followed by a fully-connected layer with 5270 neurons.</p>",
      "rawMarkdown": "To clarify: the model I used is SE-ResNet-50 with a global average pooling layer at the end, followed by a fully-connected layer with 5270 neurons.",
      "votes": null
    },
    {
      "id": "228292",
      "postDate": "10/06/2017 10:44:48",
      "content": "<p>Pretrained or from scratch? </p>",
      "rawMarkdown": "Pretrained or from scratch?",
      "votes": null
    },
    {
      "id": "228293",
      "postDate": "10/06/2017 10:46:13",
      "content": "<p>Pretrained, but with the last block fine-tuned. No data augmentation except for taking random 160x160 crops and randomly flipping them.</p>",
      "rawMarkdown": "Pretrained, but with the last block fine-tuned. No data augmentation except for taking random 160x160 crops and randomly flipping them.",
      "votes": null
    },
    {
      "id": "228294",
      "postDate": "10/06/2017 10:49:19",
      "content": "<p>So you did freeze the conv-layers?</p>",
      "rawMarkdown": "So you did freeze the conv-layers?",
      "votes": null
    },
    {
      "id": "228311",
      "postDate": "10/06/2017 11:59:51",
      "content": "<p>Yes, all conv layers were frozen except for the last block of layers.</p>",
      "rawMarkdown": "Yes, all conv layers were frozen except for the last block of layers.",
      "votes": null
    },
    {
      "id": "228312",
      "postDate": "10/06/2017 12:22:29",
      "content": "<p>Interesting. Looks like quite a good score for only training the last layer!</p>",
      "rawMarkdown": "Interesting. Looks like quite a good score for only training the last layer!",
      "votes": null
    },
    {
      "id": "228318",
      "postDate": "10/06/2017 12:39:59",
      "content": "<p>i can get lb 0.62978 for resnet50 using:</p>\n\n<ul>\n<li><p>train all layers (no freezing)</p></li>\n<li><p>0.01 for epoch 1,2;  0.001 for epoch 3;  0.0001 till 3.25 </p></li>\n</ul>\n\n<p>actually loss still decreasing for rate 0.0001, but i have to stop for some other experiments. I think resnet50 can get maybe around 0.631 ?</p>\n\n<p>for se-resnet, my experiments show that it is about 0.015 to 0.020 better then resnet. So I think if you continue to fintune all layers, you may get better results. I am using 160x160 crop from 180x180 images in my experiments</p>",
      "rawMarkdown": "i can get lb 0.62978 for resnet50 using:\n\n- train all layers (no freezing)\n\n- 0.01 for epoch 1,2;  0.001 for epoch 3;  0.0001 till 3.25 \n\nactually loss still decreasing for rate 0.0001, but i have to stop for some other experiments. I think resnet50 can get maybe around 0.631 ?\n\nfor se-resnet, my experiments show that it is about 0.015 to 0.020 better then resnet. So I think if you continue to fintune all layers, you may get better results. I am using 160x160 crop from 180x180 images in my experiments",
      "votes": null
    },
    {
      "id": "228321",
      "postDate": "10/06/2017 12:43:00",
      "content": "<p>I've done additional experiments to try to deal with the class imbalances (using <code>class_weight</code> and also using class-aware sampling) but these make it much harder for the network to learn. But I believe that if we deal with the class imbalances better, that SE-ResNet-50 should be able to get a higher accuracy.</p>",
      "rawMarkdown": "I've done additional experiments to try to deal with the class imbalances (using `class_weight` and also using class-aware sampling) but these make it much harder for the network to learn. But I believe that if we deal with the class imbalances better, that SE-ResNet-50 should be able to get a higher accuracy.",
      "votes": null
    },
    {
      "id": "228326",
      "postDate": "10/06/2017 12:46:11",
      "content": "<p>you may also want to refer to this paper:</p>\n\n<p><a href=\"https://www.cs.toronto.edu/~hinton/absps/distillation.pdf\">https://www.cs.toronto.edu/~hinton/absps/distillation.pdf</a>\n\"Distilling the Knowledge in a Neural Network\" - Geoffrey Hinton</p>\n\n<p>see:  5 Training ensembles of specialists on very big datasets\n\"Waiting for several years to train an ensemble of models was not an option, so we needed a much faster way to improve the baseline model.\"</p>\n\n<hr>\n\n<p>my idea of class balancing is to make resnet, se-resnet leans faster. The increase per iteration is very small, which i am trying to find out why.</p>",
      "rawMarkdown": "you may also want to refer to this paper:\n\nhttps://www.cs.toronto.edu/~hinton/absps/distillation.pdf\n\"Distilling the Knowledge in a Neural Network\" - Geoffrey Hinton\n\nsee:  5 Training ensembles of specialists on very big datasets\n\"Waiting for several years to train an ensemble of models was not an option, so we needed a much faster way to improve the baseline model.\"\n\n------\nmy idea of class balancing is to make resnet, se-resnet leans faster. The increase per iteration is very small, which i am trying to find out why.",
      "votes": null
    },
    {
      "id": "228454",
      "postDate": "10/06/2017 18:50:16",
      "content": "<p>How much time do you need to train the model ? On which hardware ?</p>",
      "rawMarkdown": "How much time do you need to train the model ? On which hardware ?",
      "votes": null
    },
    {
      "id": "228460",
      "postDate": "10/06/2017 19:00:43",
      "content": "<p>Hi, I'm trying to use class_weight/undersample/oversample strategies to deal with the imbalance problem. But the thing is, it did improve the precision on small classes, but it didn't improve the overall accuracy. What do you think of this?</p>",
      "rawMarkdown": "Hi, I'm trying to use class_weight/undersample/oversample strategies to deal with the imbalance problem. But the thing is, it did improve the precision on small classes, but it didn't improve the overall accuracy. What do you think of this?",
      "votes": null
    },
    {
      "id": "228533",
      "postDate": "10/07/2017 00:40:50",
      "content": "<p>cool. After the training stage, I also tried a second fine-tuning stage where each class is represented in equal proportion in the training sample, hoping to boost accuracy for some rare class, but this seems to decrease the accuracy overall. Though, I could see an increase in the number of class presented in the prediction (from 3k -&gt; 4k)</p>",
      "rawMarkdown": "cool. After the training stage, I also tried a second fine-tuning stage where each class is represented in equal proportion in the training sample, hoping to boost accuracy for some rare class, but this seems to decrease the accuracy overall. Though, I could see an increase in the number of class presented in the prediction (from 3k -&gt; 4k)",
      "votes": null
    },
    {
      "id": "228534",
      "postDate": "10/07/2017 00:42:39",
      "content": "<p>around 80k secs/epoch on Resnet50 (all layers)  + Keras + TF  on Titan X</p>",
      "rawMarkdown": "around 80k secs/epoch on Resnet50 (all layers)  + Keras + TF  on Titan X",
      "votes": null
    },
    {
      "id": "228615",
      "postDate": "10/07/2017 08:45:09",
      "content": "<p>Did anybody try with smaller models e.g. MobileNets or ResNet18, or do you think their representation power is not enough for this competition?</p>",
      "rawMarkdown": "Did anybody try with smaller models e.g. MobileNets or ResNet18, or do you think their representation power is not enough for this competition?",
      "votes": null
    },
    {
      "id": "228656",
      "postDate": "10/07/2017 11:50:27",
      "content": "<p>I tried MobileNet with the input resized to 128x128 but after 3 epochs of training the validation accuracy was only 0.35, and it did not seem to be improving. I did not spend a lot of time on this, as MobileNet is unable to break the top-3 leaderboard scores anyway.</p>",
      "rawMarkdown": "I tried MobileNet with the input resized to 128x128 but after 3 epochs of training the validation accuracy was only 0.35, and it did not seem to be improving. I did not spend a lot of time on this, as MobileNet is unable to break the top-3 leaderboard scores anyway.",
      "votes": null
    },
    {
      "id": "228772",
      "postDate": "10/07/2017 20:16:42",
      "content": "<p>Hi, Analog. Where do you find the pre-trained weights of SE-ResNet? I implemented it in Keras, but I have to train it from scratch, which is too slow for my GPU. Thx a lot!</p>",
      "rawMarkdown": "Hi, Analog. Where do you find the pre-trained weights of SE-ResNet? I implemented it in Keras, but I have to train it from scratch, which is too slow for my GPU. Thx a lot!",
      "votes": null
    },
    {
      "id": "228784",
      "postDate": "10/07/2017 20:44:00",
      "content": "<p>Hey, I used the weights from <a href=\"https://github.com/shicai/SENet-Caffe\">https://github.com/shicai/SENet-Caffe</a> which is a Caffe model. Here's the code I used to convert the model to Keras, if you're interested: <a href=\"https://gist.github.com/hollance/8d30bf5c1622036d16c4f27bd0ec88bf\">https://gist.github.com/hollance/8d30bf5c1622036d16c4f27bd0ec88bf</a></p>",
      "rawMarkdown": "Hey, I used the weights from https://github.com/shicai/SENet-Caffe which is a Caffe model. Here's the code I used to convert the model to Keras, if you're interested: https://gist.github.com/hollance/8d30bf5c1622036d16c4f27bd0ec88bf",
      "votes": null
    },
    {
      "id": "228789",
      "postDate": "10/07/2017 20:58:18",
      "content": "<p>Really appreciate! Why don't you use Caffe directly? I think it may have better speed.</p>",
      "rawMarkdown": "Really appreciate! Why don't you use Caffe directly? I think it may have better speed.",
      "votes": null
    },
    {
      "id": "228798",
      "postDate": "10/07/2017 21:58:35",
      "content": "<p>fyi, my results for se-resent-50. I train will all layers unfreeze. it is interetsing to note that your training is much faster and uses less memory becuase you freeze all layers except the last layer.</p>\n\n<p>LB 0.63358(160x160 single center crop) or 0.63839(180x180) or 0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.</p>",
      "rawMarkdown": "fyi, my results for se-resent-50. I train will all layers unfreeze. it is interetsing to note that your training is much faster and uses less memory becuase you freeze all layers except the last layer.\n\nLB 0.63358(160x160 single center crop) or 0.63839(180x180) or 0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.",
      "votes": null
    },
    {
      "id": "228918",
      "postDate": "10/08/2017 09:17:15",
      "content": "<p>@Heng CherKeng that's curious that you get better results with 180 resized to 160 than with just 180 (I didn't do a proper comparison myself yet). Is this with models trained until performance on validation plateaus, or stopped earlier?</p>",
      "rawMarkdown": "Heng CherKeng that's curious that you get better results with 180 resized to 160 than with just 180 (I didn't do a proper comparison myself yet). Is this with models trained until performance on validation plateaus, or stopped earlier?",
      "votes": null
    },
    {
      "id": "228920",
      "postDate": "10/08/2017 09:25:34",
      "content": "<p>\"Is this with models trained until performance on validation plateaus, or stopped earlier?\"</p>\n\n<p>the rate of improvement is very small from 3 epoch, lr=0.0001, onward. so i just stopped around 3.5 epoch. you may also want to read this: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780</a></p>",
      "rawMarkdown": "\"Is this with models trained until performance on validation plateaus, or stopped earlier?\"\n\nthe rate of improvement is very small from 3 epoch, lr=0.0001, onward. so i just stopped around 3.5 epoch. you may also want to read this: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780",
      "votes": null
    },
    {
      "id": "228922",
      "postDate": "10/08/2017 09:28:57",
      "content": "<p>Thanks for the answer! I'm also getting only marginal improvements after the third epoch.</p>",
      "rawMarkdown": "Thanks for the answer! I'm also getting only marginal improvements after the third epoch.",
      "votes": null
    },
    {
      "id": "230122",
      "postDate": "10/11/2017 10:35:06",
      "content": "<p>Try reduce momentum. It works magic for me. see :<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40934\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40934</a></p>",
      "rawMarkdown": "Try reduce momentum. It works magic for me. see :https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40934",
      "votes": null
    },
    {
      "id": "230124",
      "postDate": "10/11/2017 10:41:01",
      "content": "<p>I am using Keras. When I try and use an input size of 180x180 on the ResNet50 model, I get an error saying the minimum image size is 197x197. It seems this is due to the final layer of the model:</p>\n\n<pre><code>x = AveragePooling2D((7, 7), name='avg_pool')(x)\n</code></pre>\n\n<p>which needs to be changed to (6,6) if it is to work with 180x180 (or 5,5 for 160x160 images).\nAre people changing this last layer or re-scaling the images back up to 224x224, for example?\nInception and VGG16 seem quite happy with a 180x180 image size.</p>",
      "rawMarkdown": "I am using Keras. When I try and use an input size of 180x180 on the ResNet50 model, I get an error saying the minimum image size is 197x197. It seems this is due to the final layer of the model:\n\n\n    x = AveragePooling2D((7, 7), name='avg_pool')(x)\n\n\nwhich needs to be changed to (6,6) if it is to work with 180x180 (or 5,5 for 160x160 images).\nAre people changing this last layer or re-scaling the images back up to 224x224, for example?\nInception and VGG16 seem quite happy with a 180x180 image size.",
      "votes": null
    },
    {
      "id": "230132",
      "postDate": "10/11/2017 11:07:36",
      "content": "<p>I changed that layer to a <code>GlobalAveragePooling2D()</code> layer instead. You also need to change the parameter <code>min_size</code> in the call to <code>_obtain_input_shape()</code>.</p>",
      "rawMarkdown": "I changed that layer to a `GlobalAveragePooling2D()` layer instead. You also need to change the parameter `min_size` in the call to `_obtain_input_shape()`.",
      "votes": null
    },
    {
      "id": "230145",
      "postDate": "10/11/2017 11:36:06",
      "content": "<p>You retrained the 'last block' as in 'the last identiy block' (layer 162 onwards) or as in the last conv_block (layer 140 onwards?) </p>",
      "rawMarkdown": "You retrained the 'last block' as in 'the last identiy block' (layer 162 onwards) or as in the last conv_block (layer 140 onwards?)",
      "votes": null
    },
    {
      "id": "230147",
      "postDate": "10/11/2017 11:39:13",
      "content": "<p>There are 5 stages in SE-ResNet-50. I trained stage 5 plus the classifier. Stage 5 consists of a conv_block and two identity_blocks.</p>",
      "rawMarkdown": "There are 5 stages in SE-ResNet-50. I trained stage 5 plus the classifier. Stage 5 consists of a conv_block and two identity_blocks.",
      "votes": null
    },
    {
      "id": "230154",
      "postDate": "10/11/2017 11:52:30",
      "content": "<p>So, you changed the <strong>AveragePool</strong> to a <strong>GlobalPool</strong>. Seems very sensible.</p>\n\n<p>Looking at the resnet50.py code: if you set the 'pooling' parameter to 'avg', it adds a <strong>GlobalAveragePooling2D</strong> directly after the AveragePool.</p>\n\n<p>Does that even make any sense? Have the resnet50.py designers made a mistake? Did they mean that when the 'pooling' parameter was set to 'avg' that the <strong>Average</strong> should have been <strong>replaced</strong> by <strong>Global</strong>?</p>\n\n<p>Hang on - you are using the SE-ResNet-50 code, maybe that is different.</p>",
      "rawMarkdown": "So, you changed the **AveragePool** to a **GlobalPool**. Seems very sensible.\n\nLooking at the resnet50.py code: if you set the 'pooling' parameter to 'avg', it adds a **GlobalAveragePooling2D** directly after the AveragePool.\n\nDoes that even make any sense? Have the resnet50.py designers made a mistake? Did they mean that when the 'pooling' parameter was set to 'avg' that the **Average** should have been **replaced** by **Global**?\n\nHang on - you are using the SE-ResNet-50 code, maybe that is different.",
      "votes": null
    },
    {
      "id": "230176",
      "postDate": "10/11/2017 12:49:35",
      "content": "<p>Yeah, that seems a bit weird. I took that resnet50.py and added in the SENet stuff myself.</p>",
      "rawMarkdown": "Yeah, that seems a bit weird. I took that resnet50.py and added in the SENet stuff myself.",
      "votes": null
    },
    {
      "id": "230215",
      "postDate": "10/11/2017 14:16:11",
      "content": "<p>Can someone explain to me why cropping/flipping is an improvement when you are only training for 3 or 4 epochs. With 12M images per epoch, I would have thought that you didn't need to go to the trouble of generating more data. </p>\n\n<p>With 50 or 10 epochs, or with a more limited training set, I would understand the need for data augmentation.</p>",
      "rawMarkdown": "Can someone explain to me why cropping/flipping is an improvement when you are only training for 3 or 4 epochs. With 12M images per epoch, I would have thought that you didn't need to go to the trouble of generating more data. \n\nWith 50 or 10 epochs, or with a more limited training set, I would understand the need for data augmentation.",
      "votes": null
    },
    {
      "id": "230248",
      "postDate": "10/11/2017 15:56:42",
      "content": "<p>Refer to the paper <a href=\"https://arxiv.org/pdf/1409.1556.pdf\">https://arxiv.org/pdf/1409.1556.pdf</a></p>",
      "rawMarkdown": "Refer to the paper https://arxiv.org/pdf/1409.1556.pdf",
      "votes": null
    },
    {
      "id": "230330",
      "postDate": "10/11/2017 18:24:24",
      "content": "<p>Note that there are many classes with only a handful of images. For these classes data augmentation may be more important than for the classes with 80,000 images.</p>",
      "rawMarkdown": "Note that there are many classes with only a handful of images. For these classes data augmentation may be more important than for the classes with 80,000 images.",
      "votes": null
    },
    {
      "id": "230380",
      "postDate": "10/11/2017 20:52:40",
      "content": "<p>Hi Human Analog,\nWhen you said \"last block of layers\" mean you unfreeze the last conv block or just the softmax layer?</p>\n\n<blockquote>\n  <p><strong>Human Analog wrote</strong></p>\n  \n  <blockquote>\n    <p>Yes, all conv layers were frozen except for the last block of layers.</p>\n  </blockquote>\n</blockquote>",
      "rawMarkdown": "Hi Human Analog,\nWhen you said \"last block of layers\" mean you unfreeze the last conv block or just the softmax layer?\n\n\n&gt; **Human Analog wrote**\n&gt; \n&gt; &gt; Yes, all conv layers were frozen except for the last block of layers.",
      "votes": null
    },
    {
      "id": "230560",
      "postDate": "10/12/2017 08:01:51",
      "content": "<p>I meant the last residual stage, which includes one conv block and two identity blocks.</p>",
      "rawMarkdown": "I meant the last residual stage, which includes one conv block and two identity blocks.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 228276,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "10/06/2017 09:30:00",
      "content": "<p>My SE-ResNet-50 model scored 0.64358, although I resized the test images from 180x180 to 160x160 so technically speaking I did not use a crop.</p>\n\n<p>With 10 crops on each test image, the LB score increased to 0.68159.</p>",
      "votes": null,
      "replies": [
        {
          "id": 228285,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/06/2017 10:02:06",
          "content": "<p>To clarify: the model I used is SE-ResNet-50 with a global average pooling layer at the end, followed by a fully-connected layer with 5270 neurons.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228292,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/06/2017 10:44:48",
          "content": "<p>Pretrained or from scratch? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228293,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/06/2017 10:46:13",
          "content": "<p>Pretrained, but with the last block fine-tuned. No data augmentation except for taking random 160x160 crops and randomly flipping them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228294,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/06/2017 10:49:19",
          "content": "<p>So you did freeze the conv-layers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228311,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/06/2017 11:59:51",
          "content": "<p>Yes, all conv layers were frozen except for the last block of layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228312,
          "author_name": "timjoseph",
          "author_url": "",
          "post_date": "10/06/2017 12:22:29",
          "content": "<p>Interesting. Looks like quite a good score for only training the last layer!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228318,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/06/2017 12:39:59",
          "content": "<p>i can get lb 0.62978 for resnet50 using:</p>\n\n<ul>\n<li><p>train all layers (no freezing)</p></li>\n<li><p>0.01 for epoch 1,2;  0.001 for epoch 3;  0.0001 till 3.25 </p></li>\n</ul>\n\n<p>actually loss still decreasing for rate 0.0001, but i have to stop for some other experiments. I think resnet50 can get maybe around 0.631 ?</p>\n\n<p>for se-resnet, my experiments show that it is about 0.015 to 0.020 better then resnet. So I think if you continue to fintune all layers, you may get better results. I am using 160x160 crop from 180x180 images in my experiments</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228321,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/06/2017 12:43:00",
          "content": "<p>I've done additional experiments to try to deal with the class imbalances (using <code>class_weight</code> and also using class-aware sampling) but these make it much harder for the network to learn. But I believe that if we deal with the class imbalances better, that SE-ResNet-50 should be able to get a higher accuracy.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228326,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/06/2017 12:46:11",
          "content": "<p>you may also want to refer to this paper:</p>\n\n<p><a href=\"https://www.cs.toronto.edu/~hinton/absps/distillation.pdf\">https://www.cs.toronto.edu/~hinton/absps/distillation.pdf</a>\n\"Distilling the Knowledge in a Neural Network\" - Geoffrey Hinton</p>\n\n<p>see:  5 Training ensembles of specialists on very big datasets\n\"Waiting for several years to train an ensemble of models was not an option, so we needed a much faster way to improve the baseline model.\"</p>\n\n<hr>\n\n<p>my idea of class balancing is to make resnet, se-resnet leans faster. The increase per iteration is very small, which i am trying to find out why.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228460,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/06/2017 19:00:43",
          "content": "<p>Hi, I'm trying to use class_weight/undersample/oversample strategies to deal with the imbalance problem. But the thing is, it did improve the precision on small classes, but it didn't improve the overall accuracy. What do you think of this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228533,
          "author_name": "vinhnguyen",
          "author_url": "",
          "post_date": "10/07/2017 00:40:50",
          "content": "<p>cool. After the training stage, I also tried a second fine-tuning stage where each class is represented in equal proportion in the training sample, hoping to boost accuracy for some rare class, but this seems to decrease the accuracy overall. Though, I could see an increase in the number of class presented in the prediction (from 3k -&gt; 4k)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228772,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/07/2017 20:16:42",
          "content": "<p>Hi, Analog. Where do you find the pre-trained weights of SE-ResNet? I implemented it in Keras, but I have to train it from scratch, which is too slow for my GPU. Thx a lot!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228784,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/07/2017 20:44:00",
          "content": "<p>Hey, I used the weights from <a href=\"https://github.com/shicai/SENet-Caffe\">https://github.com/shicai/SENet-Caffe</a> which is a Caffe model. Here's the code I used to convert the model to Keras, if you're interested: <a href=\"https://gist.github.com/hollance/8d30bf5c1622036d16c4f27bd0ec88bf\">https://gist.github.com/hollance/8d30bf5c1622036d16c4f27bd0ec88bf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228789,
          "author_name": "brianlzm",
          "author_url": "",
          "post_date": "10/07/2017 20:58:18",
          "content": "<p>Really appreciate! Why don't you use Caffe directly? I think it may have better speed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228798,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/07/2017 21:58:35",
          "content": "<p>fyi, my results for se-resent-50. I train will all layers unfreeze. it is interetsing to note that your training is much faster and uses less memory becuase you freeze all layers except the last layer.</p>\n\n<p>LB 0.63358(160x160 single center crop) or 0.63839(180x180) or 0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228918,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "10/08/2017 09:17:15",
          "content": "<p>@Heng CherKeng that's curious that you get better results with 180 resized to 160 than with just 180 (I didn't do a proper comparison myself yet). Is this with models trained until performance on validation plateaus, or stopped earlier?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228920,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/08/2017 09:25:34",
          "content": "<p>\"Is this with models trained until performance on validation plateaus, or stopped earlier?\"</p>\n\n<p>the rate of improvement is very small from 3 epoch, lr=0.0001, onward. so i just stopped around 3.5 epoch. you may also want to read this: <a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 228922,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "10/08/2017 09:28:57",
          "content": "<p>Thanks for the answer! I'm also getting only marginal improvements after the third epoch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230122,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 10:35:06",
          "content": "<p>Try reduce momentum. It works magic for me. see :<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40934\">https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40934</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230145,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "10/11/2017 11:36:06",
          "content": "<p>You retrained the 'last block' as in 'the last identiy block' (layer 162 onwards) or as in the last conv_block (layer 140 onwards?) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230147,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/11/2017 11:39:13",
          "content": "<p>There are 5 stages in SE-ResNet-50. I trained stage 5 plus the classifier. Stage 5 consists of a conv_block and two identity_blocks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230215,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "10/11/2017 14:16:11",
          "content": "<p>Can someone explain to me why cropping/flipping is an improvement when you are only training for 3 or 4 epochs. With 12M images per epoch, I would have thought that you didn't need to go to the trouble of generating more data. </p>\n\n<p>With 50 or 10 epochs, or with a more limited training set, I would understand the need for data augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230248,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "10/11/2017 15:56:42",
          "content": "<p>Refer to the paper <a href=\"https://arxiv.org/pdf/1409.1556.pdf\">https://arxiv.org/pdf/1409.1556.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230330,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/11/2017 18:24:24",
          "content": "<p>Note that there are many classes with only a handful of images. For these classes data augmentation may be more important than for the classes with 80,000 images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230380,
          "author_name": "lamdang",
          "author_url": "",
          "post_date": "10/11/2017 20:52:40",
          "content": "<p>Hi Human Analog,\nWhen you said \"last block of layers\" mean you unfreeze the last conv block or just the softmax layer?</p>\n\n<blockquote>\n  <p><strong>Human Analog wrote</strong></p>\n  \n  <blockquote>\n    <p>Yes, all conv layers were frozen except for the last block of layers.</p>\n  </blockquote>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230560,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/12/2017 08:01:51",
          "content": "<p>I meant the last residual stage, which includes one conv block and two identity blocks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228454,
      "author_name": "thomasseleck",
      "author_url": "",
      "post_date": "10/06/2017 18:50:16",
      "content": "<p>How much time do you need to train the model ? On which hardware ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 228534,
          "author_name": "vinhnguyen",
          "author_url": "",
          "post_date": "10/07/2017 00:42:39",
          "content": "<p>around 80k secs/epoch on Resnet50 (all layers)  + Keras + TF  on Titan X</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 228615,
      "author_name": "vzsolt",
      "author_url": "",
      "post_date": "10/07/2017 08:45:09",
      "content": "<p>Did anybody try with smaller models e.g. MobileNets or ResNet18, or do you think their representation power is not enough for this competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 228656,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/07/2017 11:50:27",
          "content": "<p>I tried MobileNet with the input resized to 128x128 but after 3 epochs of training the validation accuracy was only 0.35, and it did not seem to be improving. I did not spend a lot of time on this, as MobileNet is unable to break the top-3 leaderboard scores anyway.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 230124,
      "author_name": "jinkos",
      "author_url": "",
      "post_date": "10/11/2017 10:41:01",
      "content": "<p>I am using Keras. When I try and use an input size of 180x180 on the ResNet50 model, I get an error saying the minimum image size is 197x197. It seems this is due to the final layer of the model:</p>\n\n<pre><code>x = AveragePooling2D((7, 7), name='avg_pool')(x)\n</code></pre>\n\n<p>which needs to be changed to (6,6) if it is to work with 180x180 (or 5,5 for 160x160 images).\nAre people changing this last layer or re-scaling the images back up to 224x224, for example?\nInception and VGG16 seem quite happy with a 180x180 image size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 230132,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/11/2017 11:07:36",
          "content": "<p>I changed that layer to a <code>GlobalAveragePooling2D()</code> layer instead. You also need to change the parameter <code>min_size</code> in the call to <code>_obtain_input_shape()</code>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230154,
          "author_name": "jinkos",
          "author_url": "",
          "post_date": "10/11/2017 11:52:30",
          "content": "<p>So, you changed the <strong>AveragePool</strong> to a <strong>GlobalPool</strong>. Seems very sensible.</p>\n\n<p>Looking at the resnet50.py code: if you set the 'pooling' parameter to 'avg', it adds a <strong>GlobalAveragePooling2D</strong> directly after the AveragePool.</p>\n\n<p>Does that even make any sense? Have the resnet50.py designers made a mistake? Did they mean that when the 'pooling' parameter was set to 'avg' that the <strong>Average</strong> should have been <strong>replaced</strong> by <strong>Global</strong>?</p>\n\n<p>Hang on - you are using the SE-ResNet-50 code, maybe that is different.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 230176,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "10/11/2017 12:49:35",
          "content": "<p>Yeah, that seems a bit weird. I took that resnet50.py and added in the SENet stuff myself.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "228253": "As above.\n\nMine: Resnet50 - single 160x160 crop  - 0.62316 LB",
    "228276": "My SE-ResNet-50 model scored 0.64358, although I resized the test images from 180x180 to 160x160 so technically speaking I did not use a crop.\n\nWith 10 crops on each test image, the LB score increased to 0.68159.",
    "228285": "To clarify: the model I used is SE-ResNet-50 with a global average pooling layer at the end, followed by a fully-connected layer with 5270 neurons.",
    "228292": "Pretrained or from scratch?",
    "228293": "Pretrained, but with the last block fine-tuned. No data augmentation except for taking random 160x160 crops and randomly flipping them.",
    "228294": "So you did freeze the conv-layers?",
    "228311": "Yes, all conv layers were frozen except for the last block of layers.",
    "228312": "Interesting. Looks like quite a good score for only training the last layer!",
    "228318": "i can get lb 0.62978 for resnet50 using:\n\n- train all layers (no freezing)\n\n- 0.01 for epoch 1,2;  0.001 for epoch 3;  0.0001 till 3.25 \n\nactually loss still decreasing for rate 0.0001, but i have to stop for some other experiments. I think resnet50 can get maybe around 0.631 ?\n\nfor se-resnet, my experiments show that it is about 0.015 to 0.020 better then resnet. So I think if you continue to fintune all layers, you may get better results. I am using 160x160 crop from 180x180 images in my experiments",
    "228321": "I've done additional experiments to try to deal with the class imbalances (using `class_weight` and also using class-aware sampling) but these make it much harder for the network to learn. But I believe that if we deal with the class imbalances better, that SE-ResNet-50 should be able to get a higher accuracy.",
    "228326": "you may also want to refer to this paper:\n\nhttps://www.cs.toronto.edu/~hinton/absps/distillation.pdf\n\"Distilling the Knowledge in a Neural Network\" - Geoffrey Hinton\n\nsee:  5 Training ensembles of specialists on very big datasets\n\"Waiting for several years to train an ensemble of models was not an option, so we needed a much faster way to improve the baseline model.\"\n\n------\nmy idea of class balancing is to make resnet, se-resnet leans faster. The increase per iteration is very small, which i am trying to find out why.",
    "228454": "How much time do you need to train the model ? On which hardware ?",
    "228460": "Hi, I'm trying to use class_weight/undersample/oversample strategies to deal with the imbalance problem. But the thing is, it did improve the precision on small classes, but it didn't improve the overall accuracy. What do you think of this?",
    "228533": "cool. After the training stage, I also tried a second fine-tuning stage where each class is represented in equal proportion in the training sample, hoping to boost accuracy for some rare class, but this seems to decrease the accuracy overall. Though, I could see an increase in the number of class presented in the prediction (from 3k -&gt; 4k)",
    "228534": "around 80k secs/epoch on Resnet50 (all layers)  + Keras + TF  on Titan X",
    "228615": "Did anybody try with smaller models e.g. MobileNets or ResNet18, or do you think their representation power is not enough for this competition?",
    "228656": "I tried MobileNet with the input resized to 128x128 but after 3 epochs of training the validation accuracy was only 0.35, and it did not seem to be improving. I did not spend a lot of time on this, as MobileNet is unable to break the top-3 leaderboard scores anyway.",
    "228772": "Hi, Analog. Where do you find the pre-trained weights of SE-ResNet? I implemented it in Keras, but I have to train it from scratch, which is too slow for my GPU. Thx a lot!",
    "228784": "Hey, I used the weights from https://github.com/shicai/SENet-Caffe which is a Caffe model. Here's the code I used to convert the model to Keras, if you're interested: https://gist.github.com/hollance/8d30bf5c1622036d16c4f27bd0ec88bf",
    "228789": "Really appreciate! Why don't you use Caffe directly? I think it may have better speed.",
    "228798": "fyi, my results for se-resent-50. I train will all layers unfreeze. it is interetsing to note that your training is much faster and uses less memory becuase you freeze all layers except the last layer.\n\nLB 0.63358(160x160 single center crop) or 0.63839(180x180) or 0.64356 (180x180 resize to 160x160) for se-resnet50 trained on 160x160 crops.",
    "228918": "Heng CherKeng that's curious that you get better results with 180 resized to 160 than with just 180 (I didn't do a proper comparison myself yet). Is this with models trained until performance on validation plateaus, or stopped earlier?",
    "228920": "\"Is this with models trained until performance on validation plateaus, or stopped earlier?\"\n\nthe rate of improvement is very small from 3 epoch, lr=0.0001, onward. so i just stopped around 3.5 epoch. you may also want to read this: https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40780",
    "228922": "Thanks for the answer! I'm also getting only marginal improvements after the third epoch.",
    "230122": "Try reduce momentum. It works magic for me. see :https://www.kaggle.com/c/cdiscount-image-classification-challenge/discussion/40934",
    "230124": "I am using Keras. When I try and use an input size of 180x180 on the ResNet50 model, I get an error saying the minimum image size is 197x197. It seems this is due to the final layer of the model:\n\n\n    x = AveragePooling2D((7, 7), name='avg_pool')(x)\n\n\nwhich needs to be changed to (6,6) if it is to work with 180x180 (or 5,5 for 160x160 images).\nAre people changing this last layer or re-scaling the images back up to 224x224, for example?\nInception and VGG16 seem quite happy with a 180x180 image size.",
    "230132": "I changed that layer to a `GlobalAveragePooling2D()` layer instead. You also need to change the parameter `min_size` in the call to `_obtain_input_shape()`.",
    "230145": "You retrained the 'last block' as in 'the last identiy block' (layer 162 onwards) or as in the last conv_block (layer 140 onwards?)",
    "230147": "There are 5 stages in SE-ResNet-50. I trained stage 5 plus the classifier. Stage 5 consists of a conv_block and two identity_blocks.",
    "230154": "So, you changed the **AveragePool** to a **GlobalPool**. Seems very sensible.\n\nLooking at the resnet50.py code: if you set the 'pooling' parameter to 'avg', it adds a **GlobalAveragePooling2D** directly after the AveragePool.\n\nDoes that even make any sense? Have the resnet50.py designers made a mistake? Did they mean that when the 'pooling' parameter was set to 'avg' that the **Average** should have been **replaced** by **Global**?\n\nHang on - you are using the SE-ResNet-50 code, maybe that is different.",
    "230176": "Yeah, that seems a bit weird. I took that resnet50.py and added in the SENet stuff myself.",
    "230215": "Can someone explain to me why cropping/flipping is an improvement when you are only training for 3 or 4 epochs. With 12M images per epoch, I would have thought that you didn't need to go to the trouble of generating more data. \n\nWith 50 or 10 epochs, or with a more limited training set, I would understand the need for data augmentation.",
    "230248": "Refer to the paper https://arxiv.org/pdf/1409.1556.pdf",
    "230330": "Note that there are many classes with only a handful of images. For these classes data augmentation may be more important than for the classes with 80,000 images.",
    "230380": "Hi Human Analog,\nWhen you said \"last block of layers\" mean you unfreeze the last conv block or just the softmax layer?\n\n\n&gt; **Human Analog wrote**\n&gt; \n&gt; &gt; Yes, all conv layers were frozen except for the last block of layers.",
    "230560": "I meant the last residual stage, which includes one conv block and two identity blocks."
  },
  "source": "meta"
}